The International Arab Journal of Information Technology (IAJIT)

..............................
..............................
..............................


Towards Resilient Distributed Computing: A Self-Healing Fault Tolerance Framework Using Hybrid AI Models

In today’s distributed computing paradigm, achieving resilience against unpredictable faults is still a challenging task. Several state-of-the-art fault tolerance algorithms are static and rule-based, lacking the ability to foresee future events. This work investigates a novel self-healing framework that employs Long Short-Term Memory (LSTM) networks for fault prediction and a Q-learning-based Reinforcement Learning (RL) methodology for adaptive fault recovery. In this paper, we model the recovery decision process as a Markov Decision Process (MDP) whereby the framework learns the best actions to take autonomously from the system state transitions, as well as any feedback received. A high-fidelity simulation environment was utilized to inject realistic errors so that we could rigorously evaluate the model’s performance. The comparative experimental results demonstrate significant improvements in accuracy (96.3%) and Mean Time to Recovery (MTTR) (34 sec) as well as improved availability (97.2%) compared to baseline and state-of-the-art models. This work highlights the opportunity for hybrid Artificial Intelligence (AI) models to promote resilience in cloud-edge distributed systems. The modular design of this framework also presents learning opportunities for deploying as a lesson plan in courses that focus on both distributed systems and intelligent automation in computer science curricula.


[1] Abdel Aziz N., Fouad G., Al-Saeed S., and Fawzy A., “Deep Q-Network (DQN) Model for Disease Prediction Using Electronic Health Records (EHRs),” Sci, vol. 7, no. 1, pp. 1-15, 2025. https://doi.org/10.3390/sci7010014

[2] Ahmad F., Haroon M., and Siddiqui Z., “Evaluating Fault Tolerance in Distributed Systems Using Predictive Analytics with Gated Recurrent Unit and Long Short-Term Memory Models,” Journal of Information Systems Engineering and Management, vol. 10, no. 27s, pp. 378-399, 2025. https://doi.org/10.52783/jisem.v10i27s.4421

[3] Ahmad F., Haroon M., and Siddiqui Z., “Predictive Analytics for Fault Tolerance in Distributed Systems Using Logistic Regression and Support Vector Machine,” International Journal of Environmental Sciences, vol. 11, no. 22s, pp. 5275-5289, 2025. https://doi.org/10.64252/yed3d546

[4] Ahmed W. and Wu Y., “A Survey on Reliability in Distributed Systems,” Journal of Computer and System Sciences, vol. 79, no. 8, pp. 1243-1255, 2013. https://doi.org/10.1016/j.jcss.2013.02.006

[5] Amin Z., Sethi N., and Singh H., “Review on Fault Tolerance Techniques in Cloud Computing,” International Journal of Computer Applications, vol. 116, no. 18, pp. 11-17, 2015. https://research.ijcaonline.org/volume116/number 18/pxc3902768.pdf

[6] Bampoula X., Nikolakis N., and Alexopoulos K., “Condition Monitoring and Predictive Maintenance of Assets in Manufacturing Using LSTM-Autoencoders and Transformer Encoders,” Sensors, vol. 24, no. 10, 12-5, 2024. https://doi.org/10.3390/s24103215

[7] Bank D., Koenigstein N., and Giryes R., Data Mining and Knowledge Discovery Handbook, Springer International Publishing, 2023. https://doi.org/10.1007/978-3-031-24628-9_16

[8] Bappy F., Islam T., Zaman T., Hasan R., and Caicedo C., “A Deep Dive into the Google Cluster Workload Traces: Analyzing the Application Failure Characteristics and User Behaviors,” in Proceedings of the 10th IEEE International Conference on Future Internet of Things and Cloud, Marrakesh, pp. 103-108, 2023. DOI: 10.1109/FiCloud58648.2023.00023

[9] Berman R., Azar O., and Lichtman Y., “Simulated Fault Injection Environments for Teaching System Dependability,” Computer Applications in Engineering Education, vol. 29, no. 6, pp. 1532- 1545, 2021. DOI: 10.1002/cae.22411

[10] Carmona R., Lauriere M., and Tan Z., “Model- Free Mean-Field Reinforcement Learning: Mean- Field MDP and Mean-Field Q-Learning,” The Annals of Applied Probability, vol. 33, no. 6B, 5334-5381, 2023. DOI: 10.1214/23-AAP1949

[11] Cheng Y., Elsayed E., and Huang Z., “Systems Resilience Assessments: A Review, Framework and Metrics,” International Journal of Production Research, vol. 60, no. 2, pp. 595-622, 2022. https://doi.org/10.1080/00207543.2021.1971789

[12] Choe C., Baek S., Woon B., and Kong S., “Deep Q learning with LSTM for Traffic Light Control,” in Proceedings of the IEEE 24th Asia-Pacific Conference on Communications, Ningbo, pp. 331- 336. 2018. DOI: 10.1109/APCC.2018.8633520

[13] Clifton J. and Laber E., “Q-Learning: Theory and Applications,” Annual Review of Statistics and its Application, vol. 7, no. 1, pp. 279-301, 2020. https://doi.org/10.1146/annurev-statistics- 031219-041220

[14] Coulouris G., Dollimore J., Kindberg T., and Gordon B., Distributed Systems: Concepts and Design, Pearson Education, 2011. https://books.google.co.in/books?id=3ZouAAAA QBAJ

[15] Der Kiureghian A., Ditlevsen O., and Song J., “Availability, Reliability and Downtime of Systems with Repairable Components,” Reliability Engineering and System Safety, vol. 92, no. 2, pp. 231-242, 2007. https://doi.org/10.1016/j.ress.2005.12.003

[16] Fan J., Wang Z., Xie Y., and Yang Z., “A Theoretical Analysis of Deep Q-Learning,” in Proceedings of the 2nd Conference on Learning for Dynamics and Control PMLR, Berkeley, pp. 486- 489, 2020. https://proceedings.mlr.press/v120/yang20a.html

[17] Gogineni A., “Artificial Intelligence-Driven Fault Tolerance Mechanisms for Distributed Systems Using Deep Learning Model,” Journal of Artificial Intelligence, Machine Learning and Data Science, vol. 1, no 4, pp. 1-6, 2023. doi.org/10.51219/JAIMLD/anila-gogineni/519

[18] Goyal A. and Lavenberg S., “Modeling and Analysis of Computer System Availability,” IBM Journal of Research and Development, vol. 31, no. 6, pp. 651-664, 1987. DOI: 10.1147/rd.316.0651

[19] Heda L. and Sahare P., “QLGWYB: Design of an Efficient Model for Analyzing Crowd Behavior Through Quad LSTM and Quad GRU Fusion Enhanced by Q-Learning and YOLO,” Iran Journal of Computer Science, vol. 8, pp. 1463- 1483, 2025. https://doi.org/10.1007/s42044-025- 00272-6

[20] Kambala G., “Intelligent Fault Detection and Self- Healing Architectures in Distributed Software Systems for Mission-Critical Applications,” International Journal of Scientific Research and Management, vol. 12, no. 10, 1647-1657, 2024. DOI: 10.18535/ijsrm/v12i10.ec11

[21] Karamzadeh A. and Shameli-Sendi A., “Reducing Cold Start Delay in Serverless Computing Using Lightweight Virtual Machines,” Journal of Network and Computer Applications, vol. 232, no. c, pp 104030, 2024. https://doi.org/10.1016/j.jnca.2024.104030

[22] Khowaja S. and Khuwaja P., “Q-Learning and LSTM Based Deep Active Learning Strategy for Malware Defense in Industrial IOT Applications,” Multimedia Tools and Applications, vol. 80, pp. 14637-14663, 2021. https://doi.org/10.1007/s11042-020-10371-0

[23] Liang Y., Ruan N., Yi L., and Su X., “An Approach to Workload Generation for Modern Data Centers: A View from Alibaba Trace,” BenchCouncil Transactions on Benchmarks, Standards and Evaluations, vol. 4, no. 1, pp. 100-164, 2024. https://doi.org/10.1016/j.tbench.2024.1001

[24] Liu Y., Young R., and Jafarpour B., “Long-Short- Term Memory Encoder-Decoder with Regularized Hidden Dynamics for Fault Detection in Industrial Processes,” Journal of Process Control, vol. 124, pp. 166-178, 2023. https://doi.org/10.1016/j.jprocont.2023.01.015

[25] Pasham S., “Fault-Tolerant Distributed Computing for Real-Time Applications in Critical Systems,” The Computertech, vol. 6, pp. 1-29, 2020. https://www.yuktabpublisher.com/index.php/TCT/a rticle/view/142/127

[26] Punia S., Nikolopoulos K., Singh S., Madaan J., and Litsiou K., “Deep Learning with Long Short- Term Memory Networks and Random Forests for Demand Forecasting in Multi-Channel Retail,” International Journal of Production Research, vol. 58, no. 16, pp. 4964-4979, 2020. https://doi.org/10.1080/00207543.2020.1735666

[27] Puterman M., Handbooks in Operations Research and Management Science, Elsevier, 1990. https://doi.org/10.1016/S0927-0507(05)80172-0

[28] Reiss C., Wilkes J., and Hellerstein J., Google Cluster-Usage Traces: Format+Schema, White Paper, Google Inc, 2011. https://www.researchgate.net/profile/Auday-Al- Dulaimy/post/Are-there-any-datasets-for- cloudSim/attachment/59d61de379197b807797be3e/A S%3A273823268573184%401442295962733/downlo ad/Google+cluster+usage+traces.pdf

[29] Salman H., Kalakech A., and Steiti A., “Random Forest Algorithm Overview,” Babylonian Journal of Machine Learning, vol. 2024. pp, 69-79, 2024. https://doi.org/10.58496/BJML/2024/007

[30] Samir A., Dagenborg H., and Johansen D., “QMConn: A Self-Healing Controller for Microservices Using Q-Learning and Markov Decision Processes,” in Proceedings of the IEEE/ACM 17th International Conference on Utility and Cloud Computing, Sharjah, pp. 389- 398, 2024. DOI: 10.1109/UCC63386.2024.00060

[31] Sekar J. and Aquilanz L., “Autonomous Cloud Management Using AI: Techniques for Self- Healing and Self-Optimization,” Journal of Emerging Technologies and Innovative Research, vol. 10, no. 5, pp. 571-580, 2023. http://www.jetir.org/papers/JETIR2305G78.pdf

[32] Sen P., Hajra M., and Ghosh M., Emerging Technology in Modelling and Graphics, Springer, 2020. https://doi.org/10.1007/978-981-13-7403- 6_11

[33] Shah H. and Patel J., “Self-Healing AI: Leveraging Cloud Computing for Autonomous Software Recovery,” International Journal of Intelligent Systems and Applications in Engineering, vol. 10, no. 3s, pp. 341-351, 2022. https://ijisae.org/index.php/IJISAE/article/view/7502/ 6515

[34] Shahid M., Islam N., Alam M., Mazliham M., and Musa S., “Towards Resilient Method: An Exhaustive Survey of Fault Tolerance Methods in the Cloud Computing Environment,” Computer Science Review, vol. 40, no. 100398, pp. 1-38, 2021. DOI: 10.1016/j.cosrev.2021.100398

[35] Shukla A., Kumar S., and Singh H, “Fault Tolerance Based Load Balancing Approach for Web Resources in Cloud Environment,” The International Arab Journal of Information Technology, vol. 17, no. 2, pp. 225-232, 2020. https://www.iajit.org/portal/PDF/Vol%2017,%20No .%202/17514.pdf

[36] Sun X., Cui B., and Cai Z, “Deep Q-Learning Based Circuit Breaking Method for Micro-Services in Cloud Native Systems,” in Proceedings of the CCF Conference on Computer Supported Cooperative Work and Social Computing, Singapore, pp. 348-362, 2024. https://doi.org/10.1007/978-981-99-9637-7_26

[37] Syed A. and Anazagasty E., “AI-Driven Infrastructure Automation: Leveraging AI and ML for Self-Healing and Auto-Scaling Cloud Environments,” International Journal of Artificial Intelligence Data Science, and Machine Learning, vol. 5, no. 1, pp. 32-43, 2024. https://doi.org/10.63282/3050-9262.IJAIDSML- V5I1P104

[38] Torabi H., Mirtaheri S., and Greco S. “Practical Autoencoder Based Anomaly Detection by Using Vector Reconstruction Error,” Cybersecurity, vol. 6, no. 1, pp. 1-13, 2023. https://doi.org/10.1186/s42400-022-00134-9

[39] Vadisetty R., Polamarasetti A., Butani J., Prajapati S., and et al., “AI-Powered Self-Healing and Fault-Tolerant Cloud Infrastructures for Improved Resilience and Reliability,” SSRN Electronic Journal, vol. 9, no. 1, pp. 1-10, 2024. http://dx.doi.org/10.2139/ssrn.5286332

[40] Varma S., “Artificial Intelligence in Cloud Computing: Building Intelligent, Distributed, and Fault-Tolerant Systems,” International Journal of AI, BigData, Computational and Management Studies, vol. 3, no. 1, pp. 37-45, 2022. DOI: 10.63282/3050-9416.IJAIBDCMS-V3I1P105

[41] Wiering M. and Otterlo M., Reinforcement Learning: State-of-the-Art, Adaptation, Learning, and Optimization, Springer, 2012. https://doi.org/10.1007/978-3-642-27645-3

[42] Xiao W., Fang X., Liu B., Wang J., and Zhu X., “UNION: Fault-Tolerant Cooperative Computing in Opportunistic Mobile Edge Cloud,” ACM Transactions on Internet Technology, vol. 23, no. 4, pp. 1-27, 2023. https://doi.org/10.1145/3617994

[43] Yang Y., Gao Y., Ding Z., Wu J., and et al., “Advancements in Q‐learning Meta‐Heuristic Optimization Algorithms: A Survey,” Wiley Interdisciplinary Reviews: Data Mining and Knowledge Discovery, vol. 14, no. 6, pp. 1-37, 2024. https://doi.org/10.1002/widm.1548

[44] Yang Z., Jin Y., Liu J., and Xu X., “An Intelligent Fault Self-Healing Mechanism for Cloud AI Systems via Integration of Large Language Models and Deep Reinforcement Learning,” arXiv Preprint, vol. arXiv:2506.07411v1, pp. 1-6, 2025. https://doi.org/10.48550/arXiv.2506.07411

[45] Zadakbar O., Imtiaz S., and Khan F., “Dynamic Risk Assessment and Fault Detection Using a Multivariate Technique,” Process Safety Progress, vol. 32, no. 4, pp. 365-375, 2013. https://doi.org/10.1002/prs.11609

[46] Zhang L., Jia T., Jia M., Wu Y., and et al., “A Survey of AIOps for Failure Management in the Era of Large Language Models,” arXiv Preprint, vol. arXiv:2406.11213v4, 2024. https://doi.org/10.48550/arXiv.2406.11213

[47] Zhang S., Huang Z., and Lang Y., “Application of Video Game Algorithm Based on Deep Q- Network Learning in Music Rhythm Teaching,” The International Arab Journal of Information Technology, vol. 22, no. 1, pp. 124-138, 2025. https://doi.org/10.34028/iajit/22/1/10

[48] Zhang S., Wang Y., Liu M., and Bao Z., “Data- Based Line Trip Fault Prediction in Power Systems Using LSTM Networks and SVM,” IEEE Access, vol. 6, pp. 7675-7686, 2017. DOI: 10.1109/ACCESS.2017.2785763

[49] Zhao Z., Chen W., Wu X., Chen P., and Liu J., “LSTM Network: A Deep Learning Approach for Short‐Term Traffic Forecast,” IET Intelligent Transport Systems, vol. 11, pp. 68-75, 2017. https://doi.org/10.1049/iet-its.2016.0208