The International Arab Journal of Information Technology (IAJIT)

..............................
..............................
..............................


Assessing the Effectiveness of LLMs as Autonomous Solvers for University Course Timetabling

Basma Alharbi,

Large Language Models (LLMs) have shown promise in complex reasoning, yet their capability to solve Nondeterministic Polynomial -time Ha rd (NP -Ha rd) combinatorial optimization problems remains largely unverified. This paper evaluates the performance of high -reasoning LLMs on the Curriculum -Based Course Time Tabling (CB -CTT) problem. We conduct a systematic evaluation across four dimensions: baseline pe rformance, the impact of prompt engineering, the effectiveness of iterative refinement with feedback, and scalability with problem complexity. Our investigation revealed that a leading model can solve small -scale instances with an 80% feasibility rate. The results showed that: 1) standard prompt engineering techniques do not improve the performance of the models, 2) Iterative feedback can trigger powerful one -shot corrections, yet unstable and unreliable, 3) model ’s performance degrades as problem complexit y grows. The findings indicate that LLMs, in their current state, are not reliable autonomous solvers for complex, real -world scheduling problems due to their inability to consistently enforce hard constraints and maintain solution feasibility as problem s ize and complexity increase. However, their ability to support problem formulation, heuristic guidance, and partial solution generation makes them valuabl e components within hybrid scheduling frameworks. We conclude that future work should focus on hybrid approaches that integrate LLMs with established optimization techniques, leveraging LLMs for problem modeling, decomposition, and heuristic guidance while relying on classical solvers to ensure feasibility and scalability.

 

[1] Abgaryan H., Harutyunyan A., and Cazenave T., “LLMs Can Schedule, ” arXiv preprint , pp. 1 -12, 2024. https://doi.org/10.48550/arXiv.2408.06993

[2] Akkan C., Gülcü A., and Kus Z., “Bi -Criteria Simulated Annealing for The Curriculum -Based Course Timetabling Problem with Robustness Approximation, ” Journal of Scheduling , vol. 25, no. 4, pp. 477 -501, 2022. DOI:10.1007/s10951 - 022 -00722 -0

[3] Alonge O., Sakpere W., and Adediran E., “An Automated Timetable Scheduler Using Nsga II for Optimized Scheduling in Educational Institutions, ” Advance Journal of Science, Engineering and Technology , vol. 10, no. 4, pp. 79 -97, 202 5. https://aspjournals.net/ajset/index.php/ajset/articl e/view/130

[4] Atanda O., Adebiyi M., Lawrence M., and Adeniyi E., “Enhancing Timetable Scheduling Efficiency with a Firefly Algorithm -Based Automated System: A Case Study of Landmark University, ” in Proceedi ngs of the 3 rd International Conference on Advancement in Computation and Computer Technologies , Punjab, pp. 78 -82, 2025. DOI:10.1109/InCACCT65424.2025.11011428

[5] Awad F., Al -Kubaisi A., and Mahmood M., “Large -Scale Timetabling Problems with Adaptive Tabu Se arch, ” Journal of Intelligent Systems , vol. 31, no. 1, pp. 168 -176, 2022. DOI: 10.1515/jisys -2022 -0003

[6] Bashab A., Ibrahim A ., Tarigo Hashem I., and Aggarwal K., “Optimization Techniques in University Timetabling Problem: Constraints, Methodologies, Benchmarks, And Open Issues, ” Computers, Materials and Continua , vol. 74, no. 3, pp. 6461 -6484, 2023. DOI:10.32604/cmc.2023.034051

[7] Bettinelli A., Cacchiani V., Roberti R., and Toth P., “An Overview of Curriculum -based Course Timetabling, ” TOP , vol. 23, no. 2, pp. 313 -349, 2015. DOI:10.1007/s11750 -015 -0366 -z

[8] Burke E., Mareˇcek J., Parkes A., and Rudová H., “A Branch -And -Cut Procedure for the Udine Course Timetabling Problem, ” Annals of Operations Research , vol. 194, no. 1, pp. 71 -87, 2012. DOI:10.1007/s10479 -010 -0828 -5

[9] Ceschia S., Di Gaspero L., and Schaerf A., “Design, Engineering, and Experimental Analysis of a Simulated Annealing Ap proach to The Post - Enrolment Course Timetabling Problem, ” Computers and Operations Research , vol. 39, no. 7, pp. 1615 -1624, 2012. htt ps://doi.org/10.1016/j.cor.2011.09.014

[10] Chen M., Sze S., Goh S., Sabar N., and Kendall G., “A Survey of University Course Timetabling Problem: Perspectives, Trends and Opportunities, ” IEEE Access , vol. 9, pp. 106515 - 106529, 2021. DOI:10.1109/ACCESS.2021.310 0613

[11] Di Gaspero L., McCollum B., and Schaerf A., “The Second International Timetabling Competition (itc -2007): Curriculum -Based Course Timetablin g (track 3),” in P roceedings of the 1 st International Workshop on Scheduling a Scheduling Competition , Rhode Island, pp. 1 -4, 2007. https://api.semanticscholar.org/CorpusID:5872498

[12] Fan L., Hua W., Li L., Ling H., and Zhang Y., “Nphardeval: Dynamic Benchmark on Reasoning Ability of Large Language Models Via Complexity Classes, ” arXiv preprint , vol. arXiv:2 312.14890, pp. 1 -23, 2023. https://doi.org/10.48550/arXiv.2312.14890

[13] Fonseca G., Santos H., Carrano E., and Stidsen T., “Integer Programming Techniques for Educational Timetabling, ” European Journal of Operational Research , vol. 262, no. 1, pp. 28 -39, 2017 . DOI:10.1016/j.ejor.2017.03.020

[14] Giannoulis P., Yorgos P., and Tzamos C., “Teaching Transformers to Solve Combinatorial Problems through Efficient Trial and Error. ” arXiv Preprint , vol. arXiv:2509.22023 , pp. 1 -33, 2025. https://doi.org/10.48550/arXiv.2509. 22023

[15] Gu X., Krish M., Sohail S., Thakur S., and et al, “From Integer Programming to Machine Learning: A Technical Review on Solving University Timetabling Problems, ” Computation , vol. 13, no. 1, pp. 1 -27, 2025. https://doi.org/10.3390/computation13010010

[16] Hafsa M., Wattebled P., Jacques J., and Jourdan L., “Solving a Multi -Objective Professional Timetabling Problem Using Evolutionary Algorithms at Mandarine Academy, ” International Transactions in Operational Research , vol. 32, no. 1, pp. 244 -269, 2025. https://doi.org/10.1111/itor.13276

[17] Li T., Chiang W. , Frick E., Dunlap L., and et al, “From Crowdsourced Data to High -Quality Benchmarks: Arena -Hard and Benchbuilder Pipeline, ” in Proceedings of the 42 nd International Conference on Machine Learning , pp. 34209 - 34231, 2024.

[18] Lin B., Le Bras R., Richardson K., Sabharwal A., and et al, “Zebralogic: On The Scaling Limits of LLMs for Logical Reasoning, ” in P roceedings of the 42 nd International Conference on Machine Learning , pp. 1 -17, Vancouver, 2025. https://hf.co/spaces/WildEval/ZebraLogic

[19] Liu B., Bubeck S., Eldan R., and Kulkarni J., “TinyGSM: Achieving> 80% on Gsm8k with Small Language Models, ” arXiv prep rint , vol. arXiv:2312.09241, pp. 1 -15, 2023. https://doi.org/10.48550/arXiv.2312.09241

[20] Lu Z. and Hao J ., “ Adaptive Tabu Search for Course Timetabling, ” European Journal of Operational Research , vol. 200, no. 1, pp. 235 - 244, 2010. DOI:10.1016/j.ejor.2008.12 .007

[21] McCollum B. and Ireland N., “University Timetabling: Bridging The Gap Between Research and Practice, ” in Proceedings of the Practice and Theory of Automated Timetabling Conference , Brno, pp. 15 -35, 2006. https://patatconference.org/patat2006/proceedings/ 1_2.pdf

[22] Phan L., Gatti A., Han Z., Li N., and et al, “Humanity ’s Last Exam, ” arXiv preprint , vol. arXiv:2501.14249, pp. 1 -29, 2025. https://doi.org/10.48550/arXiv.2501.14249

[23] Premananda I., Tjahyanto A., and Muklason A., “Timetabling Problems and the Effort Towards Generic Algorithms: A Comprehensive Survey, ” IEEE Access , vol. 12, pp. 143854 -143868, 2024. DOI:10.1109/ACCESS.2024.3463721

[24] Qintong Li., Cui L., Zhao X., and Kong L., “GSM - Plus: A Comprehensive Benchmark for Evaluating the Robustness of LLMs as Mathematical Problem Solvers, ” in Proceedings of the 62 nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Pape rs). Association for Computational Linguistics , Bangkok, pp. 2961 -2984, 2024. DOI:10.18653/v1/2024.acl - long.163

[25] Rappos E., Thiémard E., Robert S., and Hêche J - F., “A Mixed -Integer Programming Approach for Solving University Course Timetabling Problems, ” Jo urnal of Scheduling , vol. 25, no. 4, pp. 391 -404, 2022. DOI:10.1007/s10951 -021 - 00715 -5

[26] Respati R. and Ramadhani D., “Optimizing Lecture Scheduling Using Genetic Algorithm: A Case Study at Universitas Riau, ” International Journal of Electrical, Energy and P ower System Engineering , vol. 8, no. 2, pp. 220 -232, 2025. DOI: 10.31258/ijeepse.8.2.220 -232

[27] Selvaraj A., R. Prabakaran, and Kanimozhi R., “A Novel Resource Scheduler for Resource Allo cation and Scheduling in Big Data Using Hybrid Optimization Algorithm at Cloud Environment, ” The International Arab Journal of Information Technology , vol. 20, no. 6, pp. 863 - 873, 2023. https://doi.org/10.34028/iajit/20/6/3

[28] Wang Y., Ma X., Zhang G., Ni Y., and et al, “MMLU -Pro: A More Robust and Challenging Multi -Task Language Understanding Benchmark, ” in Proceedings of the 38 th International Conference on Neural Information Processing Systems , Vancouver, pp. 95266 -95290, 2024. DOI: 10.52202/079017 -3018

[29] Xiao Y., Li X., Jiang L., and Wang P., “Towards Dynamic University Course Timetabling Problem: An Automated Approach Augmented Via Reinforcement Learning, ” in Proceedings o f The IEEE International Conference on Data Mining , Abu Dhabi, pp. 520 -529, 2024. doi:10.1109/ICDM59182.2024.00059

[30] Yuksekgonul M., Chandrasekaran V., Jones E., Gunasekar S., and et al, “Attention Satisfies: A Constraint -Satisfaction Lens on Factual Errors of Language Models, ” in Proceedings of the 12 th International Conference on Learning Representations , Vienna, pp. 1 -25, 2024. https://openreview.net/pdf?id=gfFVATffPd

[31] Zhang H., Da J., Lee D., Robinson V., and et al, “A Careful Examination of Large Language Model Performance on Grade School Arithmetic, ” in Proceedings of the 38 th International Conference on Neural Information Processing System , New York, pp. 46819 -46836, 2024. https://dl.acm.org/doi/10.5555/3737916.3739401

[32] Zhou J., Lu T., Mi shra S., and Brahma S., “Instruction -Following Evaluation for Large Language Models, ” arXiv preprint , vol. arXiv:2311.07911v1 , pp. 1 -43, 2023. https://doi.org/10.48550/arXiv.2311.07911

[33] Zhou Y., Ye J., Ling Z., Han Y., and et al , “Dissecting Logical Reasoni ng in LLMs: A Fine - Grained Evaluation and Supervision Study, ” arXiv preprint , vol. arXiv:2506.04810, pp. 1 -24, 2025. DOI:10.48550/arXiv.2506.04810