Abstract
Social media has become one of the main places where people talk about brands every day. Customers use platforms such as Twitter (X) to praise, complain and share their opinions, and these public messages can quickly change how a brand is seen. Because of this, many companies now rely on sentiment analysis to follow what people are saying and to react to problems and opportunities in real time (Liu, 2012). At the same time, the informal, sarcastic and context-dependent language used on these platforms makes sentiment analysis difficult for standard text classification methods (Cambria, 2016). This study compares six sentiment classification approaches—Logistic Regression, Naive Bayes, Support Vector Machines, Random Forest, a Long Short-Term Memory (LSTM) network and a stacked ensemble—on the widely used Sentiment140 dataset of 1.6 million labelled tweets (Go, Bhayani & Huang, 2009), with the specific goal of informing model choice for brand reputation monitoring. Under matched data and preprocessing conditions, LSTM achieved the highest individual accuracy (79.20%), with the ensemble model close behind (78.69%), while traditional models remained competitive around 75–77% at much lower computational cost. Error analysis shows that all approaches still struggle with sarcasm, context-dependent phrasing and ambiguous language. For practitioners, the results offer a concrete trade-off: where speed and low resource usage matter most, traditional models and Random Forest are reasonable choices; where accuracy is critical and computing power is available, LSTM is preferable. The findings also point to sarcasm detection and richer contextual representations as key areas for future work in brand reputation monitoring.
Authors
Abhishen Raju
De Montfort University, United Kingdom
Keywords
Sentiment Analysis, Brand Reputation Management, Machine Learning, Deep Learning, LSTM, Ensemble Learning, Social Media Analytics