
From Roots to Tweets: Leveraging Linguistic Features for Arabic Social Media Content Classification
The proliferation of disinformation and hate speech across Arabic-speaking social media requires more effective content classification. This paper introduces a novel framework that constructs morpho-syntactic knowledge graphs Arabic Knowledge Graph (ARKG) specifically designed to exploit Arabic’s rich morphological structure for social media content classification. Our approach deconstructs tweets into their core linguistic components-words, lemmas, and roots-and uses Term Frequency-Inverse Document Frequency (TF-IDF) to weigh their relationships, creating rich, interpretable feature vectors that we evaluate using both traditional machine learning classifiers and graph neural network embeddings via metapath2vec. We evaluate the ARKG on three distinct tasks: fake news, hate speech, and sarcasm detection. Experimental results demonstrate that our linguistically-motivated features consistently outperform graph embeddings, achieving F1-scores of 95% for fake news, 84% for hate speech, and 82% for sarcasm detection-with our approach proving competitive with state-of-the-art transformer models while offering greater interpretability. Notably, our graph-based method outperforms a Masked Arabic Bidirectional Encoder Representations from Transformers (MARBERT) baseline by 5% in sarcasm detection while maintaining comparable performance on the other tasks. This work demonstrates that morphologically-aware graph representations can match transformer performance while providing interpretable insights into Arabic text classification.
[1] Abdul -Mageed M., Elmadany A., and Nagoudi E., “ARBERT and MARBERT: Deep Bidirectional Transformers for Arabic, ” in Proceedings of the 59 th Annual Meeting of the Association for Computational Linguistics and the 11 th International Joint Conference on Natural Language Processing , Online, pp. 7088 -7105, 2021. DOI:10.18653/v1/2021.acl -long.551
[2] Abu Farha I. and Magdy W., “From Arabic Sentiment An alysis to Sarcasm Detection: The Ar - Sarcasm Dataset, ” in Proceedings of the 4 th Workshop on Open -Source Arabic Corpora and Processing Tools, with a Shared Task on Offensive Language Detection , Marseille, pp. 32 -39, 2020. https://aclanthology.org/2020.osact -1.5/
[3] Al -Arfaj A. and Al -Salman A., “Arabic NLP Tools for Ontology Construction from Arabic Text: An Overview, ” in Proceedings of the International Conference on Electrical and Information Technologi es, Marrakech, pp. 246 -251, 2015. https://doi.org/10.1109/eitech.2015.7162980
[4] Al -Ayyoub M., Khamaiseh A., Jararweh Y., and Al -Kabi M., “A Comprehensive Survey of Arabic Sentiment Analysis, ” Information Processing and Management , vol. 56, no. 2, pp. 320 -342 , 2019. DOI:10.1016/j.ipm.2018.07.006
[5] Al -Ghadhban D., Alnkhilan E., Tatwany L., and Alrazgan M., “Arabic Sarcasm Detection in Twitter, ” in Proceedings of the International Conference on Engineering and MIS , Monastir, pp. 1 -7, 2017. https://doi.org/10.1109/icemis.2017.8272990
[6] Al -Hassan A. and Al -Dossari H., “Detection of Hate Speech in Social Networks: A Survey on Multilingual Corpus, ” In Proceedings of The 6 th International Conference on Computer Science and Information Technology , Sydney, pp. 1 -18, 2019. https://doi.org/10.5121/csit.2019.90208
[7] Al -Kabi M., Al -Radaideh Q., and Akkawi K., “Benchmarking and Assessing the Performance of Arabic Stemmers, ” Journal of Information Science , vol. 37, no. 2, pp. 111 -119, 2011. https://doi.org/1 0.1177/0165551510392305
[8] Al -Shawakfa E., Al -Badarneh A., Shatnawi S., Al - Rabab ’ah K., and Bani -Ismail B., “A Comparison Study of Some Arabic Root Finding Algorithms, ” Journal of the American society for information science and technology , vol. 61, no. 5, pp . 1015 - 1024, 2010. https://doi.org/10.1002/asi.21301
[9] Al -Yahya M., Al -Khalifa H., Al -Baity H., AlSaeed D., and Essam A., “Arabic Fake News Detection: Comparative Study of Neural Networks and Transformer -Based Approaches, ” Complexity , vol. 2021, pp. 1 -10, 20 21. https://doi.org/10.1155/2021/5516945
[10] Albadi N., Kurdi M., and Mishra S., “Are They Our Brothers? Analysis and Detection of Religious Hate Speech in The Arabic Twittersphere, ” in Proceedings of The IEEE/ACM International Conference on Advances in Social Networks Analysis and Mining , Barcelona, pp. 69 - 76, 2018. DOI:10.1109/asonam.2018.8508247
[11] Aljlayl M. and Frieder O., “On Arabic Search: Improving the Retrieval Effectiveness Via a Light Stemming Approach, ” In Proceedings of The 11 th International Conferen ce on Information and Knowledge Management , McLean, pp. 340 -347, 2002. https://doi.org/10.1145/584792.584848
[12] Alruily M., “Classification of Ara bic Tweets: A Review, ” Electronics , vol. 10, no. 10, pp. 1 -31, 2021. DOI:10.3390/electronics10101143
[13] Alzanin S., Azmi A., and Aboalsamh H., “Short Text Classification for Arabic Social Media Tweets, ” Journal of King Saud University - Computer and Information Sciences , vol. 34, no. 9, pp. 6595 -6604, 2022. https://doi.org/10.1016/j.jksuci.2022.03.020
[14] Anezi F., “Arabic Ha te Speech Detection Using Deep Recurrent Neural Networks, ” Applied Sciences , vol. 12, no. 12, pp. 1 -15, 2022. https://doi.org/10.3390/app12126010
[15] Antoun W., Baly F., and Hajj H., “AraELECTRA: Pre -Training Text Discriminators for Arabic Language Understanding, ” in Proceedings of the 6th Arabic Natural Language Processing Workshop , Kyiv, pp. 191 -195, 2020. https://aclanthology.org/2021.wanlp -1.20/
[16] Aoun Barakat K., Dabbous A., and Tarhini A., “An Empirical Approach to Understanding Users ’ Fake News Identification On Social Media, ” Online Information Review , vol. 45, no. 6, pp. 1080 -1096, 2021. https://doi.org/10.1108/oir -08 -2020 -0333
[17] Aref A., Al Mahm oud R., Taha K., Al -Sharif M., and et al, “Hate Speech Detection of Arabic Shorttext ”; in Proceedings of The 9th International Conference on Information Technology Convergence and Services , Copenhagen, pp. 81 - 94, 2020. https://doi.org/10.5121/csit.2020.100507
[18] Askool S., “The Use of Social Media in Arab Countries: A Case of Saudi Arabia, ” in Proceedings of the Web Information Systems and Technologies: 8 th International Conference , Porto, pp. 201 -219, 2012. 2013. https://d oi.org/10.1007/978 -3-642 -36608 -6_13
[19] Atwan J., Mohd M., Rashaideh H., and Kanaan G., “Semantically Enhanced Pseudo Relevance Feedback for Arabic Information Retrieval, ” Journal of Information Science , vol. 42, no. 2, pp. 246 -260, 2016. https://doi.org/10.11 77/0165551515594722
[20] Ben Guirat S., Bounhas I., Slimani Y., “Meta - Search Based Approach for Arabic Information Retrieval, ” Online Information Review , vol. 46, no. 7, pp. 1257 -1274, 2022. https://doi.org/10.1108/oir -11 -2020 -0515
[21] Boudchiche M., Mazroui A., Bebah M., Lakhouaja A., and Boudlal A., “AlKhalil Morpho Sys 2: A Robust Arabic Morpho -Syntactic Analyzer ,” Journal of King Saud University - Computer and Information Sciences , vol. 29, no. 2, pp. 141 -146, 2017. https://doi.org/10. 1016/j.jksuci.2016.05.002
[22] Bounhas I., Soudani N., and Slimani Y., “Building A Morpho -Semantic Knowledge Graph for Arabic Information Retrieval, ” Information Processing and Management , vol. 57, no. 6, pp. 102124, 2020. https://doi.org/10.1016/j.ipm.2019.102124
[23] Bounhas I., “On the Usage of a Classical Arabic Corpus as a Language Resource: Related Research and Key Challenges, ” ACM Transactions on Asian and Low -Resource Language Information Processing , vol. 18, no. 3, pp. 1 -45, 2019. https://doi.org/10.1145/3277591
[24] Castillo C., Mendoza M., and Poblete B., “Information Credibility on Twitter ,” in Proceedings of the 20 th International Conference on World Wide Web , Hyderabad, pp. 675 -684, 2011. https://doi.org/10.1145/1963405.1963500
[25] Chandra S., Mishra P., Yannakoudakis H., Nimishakavi M., and et al, “Graphbased Modeling of Online Communities for Fake News Detection, ” arXiv Pre -prints , vol. arXiv:2008.06274v4, pp. 1 -16, 2020. https://doi.org/10.48550/arXiv.2008.06274
[26] Devlin J., Chang M., Lee K., and Toutanova K., “BERT: Pre -Training of Deep Bidirectional Transformers for Language Understanding, ” in Proceedings of the Conference of the North American Chapter of the Association f or Computational Linguistics: Human Language Technologies , Minneapolis, pp. 4171 -4186, 2019. https://doi.org/10.18653/v1/n19 -1423
[27] Dong Y., Chawla N., and Swami A., “Metapath2vec: Scalable Representation Learning for Heterogeneous Networks, ” in Proceedings of the 23 rd ACM SIGKDD International Conference On Knowledge Discovery and Data Mining , Halifax, 2017, pp. 135 -144, 2017. https://doi.org/10.1145/3097983.3098036
[28] El Mahdaouy A., Gaussier E., and Ouatik El Alaoui S., “Should One Use Term Proximity or Multi -Word Terms for Arabic Information Retrieval?, ” Computer Speech and Language, vol. 58, pp. 76 -97, 2019. https://doi.org/10.1016/j.csl.2019.04.002
[29] El Mahdaouy A., El Alaoui S., and Gaussier E., “Word -Embedding -Based Pseudo -Relevance Feedback for Arabic Information Retrieval, ” Journal of information science , vol. 45, no. 4, pp. 429 -442, 2019. https://doi.org/10.1177/0165551518792210
[30] Faraj D. and Abdullah M., “SarcasmDet at Sarcasm Detection Task 2021 in Arabic Using AraBERT Pretrained Model, ” in Proceedi ngs of the 6 th Arabic Natural Language Processing Workshop , Kyiv, pp. 345 -350, 2021. https://aclanthology.org/2021.wanlp -1.44/
[31] Grover A. and Leskovec J., “Node2vec: Scalable Feature Learning for Networks, ” in Proceedings of The 22 nd ACM SIGKDD International Conference On Knowledge Discovery and Data Mining , San Francisco, pp. 855 -864, 2016. https://doi.org/10.1145/2939672.2939754
[32] Guirat S., Bounhas I., and Slimani Y., “Combining Indexing Units for Arabic Information Retrieval, ” International Journal of Software Innovation , vol. 4, no. 4, pp. 1 -14, 2016. https://doi.org/10.4018/ijsi.2016100101
[33] Guirat S., Bounhas I., and Slimani Y., “Enhancing Hybrid Indexing for Arabic Information Retrieval, ” in Proceedings of the 32 nd International Symposium , Poznan, pp. 247 -254, 2018. https://doi.org/10.1007/978 -3-030 -00840 - 6_27
[34] Guirat S., Bounhas I., and Slimani, Y., “Pre - Indexing Techniques in Arabic Information Retrieval, ” in Proceedings of the 11 th International Conference on Agents and Artificial Intelligence , Prague, pp. 237 -246, 2019. https://doi.org/10.5220/000739340237024 6
[35] Habash N., Introduction to Arabic Natural Language Processing , Springer Nature, 2022. https://doi.org/10.1007/978 -3-031 -02139 -8
[36] Hadni M., Lachkar A., and Ouatik S., “A New and Efficient Stemming Technique for Arabic Text Categorization, ” in Proceedings of the International Conference on Multimedia Computing and Systems , Tangier, pp. 791 -796, 2012. https://doi.org/10.1109/icmcs.2012.6320308
[37] Hamdi T., Slimi H., Bounhas I., and Slimani Y., “A Hybrid Approach for Fake News Detection in Twitter Based on User Features and Graph Embedding, ” in Proceedings of the 16 th International Conference on Distributed Computing and Internet Technology , Bhubane swar, pp. 266 -280, 2020. https://doi.org/10.1007/978 -3-030 -36987 -3_17
[38] Han Y., Silva A., Luo L., Karunasekera S., and From Roots to Tweets: Leveraging Linguistic Features for Arabic Social Media Content 891 Leckie C., “Knowledge Enhanced Multi -Modal Fake News Detection, ” arXiv Pre -prints , vol. arXiv:2108.04418v1, pp. 1 -10, 2021. https://doi.org /10.48550/arXiv.2108.04418
[39] Haouari F., Hasanain M., Suwaileh R., and Elsayed T., “ArCOV19 -Rumors: An Arabic COVID -19 Twitter Dataset for Misinformation Detection, ” in: Proceedings of the 6 th Arabic Natural Language Processing Workshop , Kyiv, pp. 72 -81, 2021. https://aclanthology.org/2021.wanlp -1.8/
[40] Harrag F. and Djahli M., “Arabic Fake News Detection: A Fact Checking Based Deep Learning Approach, ” Transactions on Asian and Low - Re source Language Information Processing , vol. 21, no. 4, pp. 1 -34, 2022. https://doi.org/10.1145/3501401
[41] Himdi H., Weir G., Assiri F., and Al -Barhamtoshy H., “Arabic Fake News Detection Based on Textual Analysis, ” Arabian Journal for Science and Engineering , vol. 47, no. 8, pp. 10453 -10469, 2022. https://doi.org/10.1007/s13369 -021 -06449 -y
[42] Himdi H., Alhayan F., and Shaalan K., “Neural Networks and Sentiment Features for Extremist Content Detection in Arabic Social Media, ” The International Arab Journal of Information Technology , vol. 22, no. 3, pp. 522 -534, 2025. https://doi.org/10.34028/iajit/22/3/8
[43] Jollife I. and Cadima J., “Principal Component Analysis: A Review and Recent Developments, ” Philosophical Transactions of the Royal Society A: Mathematical, Ph ysical and Engineering Sciences , vol. 374, pp. 1 -16, 2016. https://doi.org/10.1098/rsta.2015.0202
[44] Larkey L. and Connell M., “Arabic Information Retrieval at UMass in TREC -10,” in Proceedings of the Tenth Text REtrieval Conference , Gaithersburg, pp. 1 -9, 2001. https://doi.org/10.21236/ada456273
[45] Larkey L., Ballesteros L., and Connell M., Arabic Computational Morphology: Knowledge -Based and Empirical Methods , Springer, pp. 221 -243, 2007. https://doi.org/10.1007/978 -1-4020 -6046 - 5_12
[46] Liu Y., Ott M., Goyal N., Du J., and et al, “RoBERTa: A Robustly Optimized BERT Pretraining Approach ,” arXiv preprint , vol. arXiv:1907.11692, pp. 1 -13, 2019. https://doi.org/10.48550/arXiv.1907.11692
[47] Matsumoto H., Yoshida S., and Muneyasu M., “Propagation -Based Fake News Detection Using Graph Ne ural Networks with Transformer, ” in Proceedings of The IEEE 10 th Global Conference on Consumer Electronics , Kyoto, pp. 19 -20, 2021. https://doi.org/10.1109/gcce53005.2021.9621803
[48] Mohammed A. and Kora R., “An Effective Ensemble Deep Learning F ramework for Text Classification ”; Journal of King Saud University - Computer and Information Sciences , vol. 34, no. 10, pp. 8825 -8837, 2022. https://doi.org/10.1016/j.jksuci.2021.11.001
[49] Mulki H., Haddad H., Bechikh Ali C., and Alshabani H., “L-HSAB: A Levantine Twitter Dataset for Ha te Speech and Abusive Language, ” in Proceedings of the 3 rd Workshop on Abusive Language Online , Florence, pp. 111 -118, 2019. https://doi.org/10.18653/v1/w19 -3512
[50] Nabil M., Aly M., and Atiya A., “ASTD: Arabic Sentiment Tweets Dataset, ” in Proceedings of the Conference on Empirical Methods in Natural Language Processing , Lisbon, pp. 2515 -2519, 2015. https://doi.org/10.18653/v1/d15 -1299
[51] Obeid O., Zalmout N., Khalifa S., Taji D., and et al, “CAMeL Tools: An Open Source Python Too lkit for Arabic Natural Language Processing, ” in Proceedings of the 12 th Language Resources and Evaluation Conference , Marseille, pp. 7022 - 7032, 2020. https://aclanthology.org/2020.lrec - 1.868/
[52] Pasha A., Al -Badrashiny M., Diab M., El Kholy A., and et al, “Madamira: A Fast, Comprehensive Tool for Morphological Analysis and Disambiguation of Arabic, ” in Proceedings of the 9th International Conference on Language Resources and Evaluation , Reykjavik, pp. 1094 - 1101, 2014. http://www.lrec - conf.org/proceedings/lrec2014/pdf/593_Paper.pdf
[53] Ren Y. and Zhang J., “Fake News Detection on News -Oriented Heterogeneous Information Networks Through Hierarchical Graph Attention, ” in Proceedings of the International J oint Conference on Neural Networks , Shenzhen, pp. 1 - 8, 2021. DOI:10.1109/ijcnn52387.2021.9534362
[54] Rosenthal S., Farra, N., and Nakov P., “SemEval - 2017 Task 4: Sentiment Analysis in Twitter, ” in Proceedings of the 11 th International Workshop on Semantic Evaluation , Vancouver, pp. 502 -518, 2017. https://doi.org/10.18653/v1/S17 -2088
[55] Scarselli F., Gori M., Tsoi A., Hagenbuchner M., and Monfardini G., “The Graph Neural Network Model, ” IEEE Transactions On Neural Networks , vol. 20, no. 1, pp. 61 -80, 2008. https://doi.org/10.1109/tnn.2008.2005605
[56] Shoukry A. and Rafea A., “Sentence -Level Arabic Sentiment Analysis, ” in Proceedings of the International Conference on Collaboration Technologies and Systems , Denver, pp. 546 -550, 201 2. DOI:10.1109/cts.2012.6261103
[57] Slimi H., Bounhas I., and Slimani Y., “Twitter Users Crediblity Evaluation Based on Social Graph Impression, ” in Proceedings of the 16 th International Conference on Applied Computing , Cagliari, pp. 1 -8, 2019. https://doi.org/10.33965/ac2019_201912l002
[58] Slimi H., Bounhas I., Slimani Y., “URL -Based Tweet Credibility Evaluation, ” in Proceedings of the IEEE/ACS 16 th International Conference on Computer Systems and Applications , Abu Dhabi, pp. 1 -6, 2019. https://doi. org/10.1109/aiccsa47632.2019.9035255
[59] Slimi H., Bounhas I., and Slimani Y., “Adapting Pre -Trained Language Models to Rumor Detection on Twitter, ” Journal of Universal Computer Science , vol. 27, no. 10, pp. 1128 -1148, 2021. https://doi.org/10.3897/jucs.65918
[60] Yamaguchi S., “Why Are There So Many Extreme Opinions Online?: An Empirical, Comparative Analysis of Japan, Korea and the USA, ” Online information review , vol. 47, no. 1, pp. 1 -19, 2023. https://doi.org/10.1108/oir -07 -2020 -0310
[61] Yang Z., Dai Z., Yang Y., C arbonell J., and et al, “XLNet: Generalized Autoregressive Pretraining for Language Understanding, ” in Proceedings of the 33 rd International Conference on Neural Information Processing Systems , NewYork, pp. 1 - 18, 2019. https://doi.org/10.48550/arXiv.1906.08237