A comparison of a novel optimized GSDMM Model with K-means clustering for topic modelling of free text

ORCID

M Wojtys: 0000-0002-6598-9572
C McNeile: 0000-0003-0305-2028

Abstract

Statistical topic modelling has become an important tool in the text processing field, because more applications are using it to handle the increasing amount of available text data, e.g. from social media platforms. The aim of topic modelling is to discover the main themes or topics from a collection of text documents. While several models have been developed, there is no consensus on evaluating the models, and how to determine the best hyper-parameters of the model. In this research, we develop a method for evaluating topic models for short text that employs word embedding and measuring within-topic variability and separation between topics. We focus on the Dirichlet Mixture Model and tuning its hyper-parameters. We also investigate using the K-means clustering algorithm. In empirical experiments, we present a novel case study on short text datasets related to the telecommunication industry. We find that the optimal values of hyper-parameters, obtained from our evaluation method, do not agree with the fixed values typically used in the literature and lead to different clustering of the text corpora. Moreover, we compare the discovered topics with those obtained from the K-means clustering.

DOI Link

10.11159/jmids.2023.007

Publication Date

2023-12-06

Publication Title

Journal of Machine Intelligence and Data Science

Acceptance Date

2023-12-02

Deposit Date

2023-04-12

Embargo Period

2024-01-10

Recommended Citation

Abdelmotaleb, H., Wojtys, M., & McNeile, C. (2023) 'A comparison of a novel optimized GSDMM Model with K-means clustering for topic modelling of free text', Journal of Machine Intelligence and Data Science, . Available at: 10.11159/jmids.2023.007

School of Engineering, Computing and Mathematics

A comparison of a novel optimized GSDMM Model with K-means clustering for topic modelling of free text

ORCID

Abstract

DOI Link

Publication Date

Publication Title

Acceptance Date

Deposit Date

Embargo Period

Recommended Citation

Search

Browse

About

Links

School of Engineering, Computing and Mathematics

A comparison of a novel optimized GSDMM Model with K-means clustering for topic modelling of free text

Authors

ORCID

Abstract

DOI Link

Publication Date

Publication Title

Acceptance Date

Deposit Date

Embargo Period

Recommended Citation

Share

Search

Browse

About

Links