1 Seattle, USA.
2 Texas, USA.
Received on 15 May 2026; revised on 25 June 2026; accepted on 27 June 2026
Cloud computing systems help to maintain massive distributed applications as well as data-intensive services in various fields. Nonetheless, the distributed and complex nature of the cloud infrastructures render them susceptible to a number of system failures, such as hardware failures, software failures, network failures, virtualization failures, and ineffective resource allocation. Such failures may have a great impact on the availability of services, the reliability of the system, and the general performance. This paper aims to provide a systematic review on the topic of preventing failures and recovery methods in cloud systems in a systematic way to offer an organized view of the existing methods of promoting network reliability. Based on the systematic search strategy in the ScienceDirect database, the sources related to interventions that were published in the last ten years (2015 to 2025) were searched and filtered according to the pre-established inclusion and exclusion criteria, which have led to the identification of a final dataset of 68 studies. The review evaluates proactive failure prevention methods including predictive monitoring, machine learning-based anomaly detection, redundancy plan, load balancing plan and intelligent resource scheduling. Moreover, recovery strategies such as checkpointing, replication, failover systems, container migration and disaster recovery systems are discussed. The authors find that proactive monitoring strategies in combination with effective recovery strategies are highly effective in enhancing the resilience and operational continuity of distributed cloud infrastructures. The paper also finds significant research issues associated with scalability, security risk, real-time monitoring, and multi-cloud reliability. Such results indicate that there is a necessity of coherent resilience frameworks which integrate intelligent monitoring, intelligent fault-tolerance schemes, and automatic recovery systems. The paper offers an organized insight into the approaches of preventing failures and recovery mechanisms of cloud computing infrastructure and the fact that smart and adaptable resilience models are necessary.
Cloud computing; Failure prevention; Fault tolerance; Failure recovery; Predictive monitoring
Preview Article PDF
Oluwafemi Oluwagboyega Fabiyi, Mary Magdalene Yeboah. Failure Prevention and Recovery Techniques in Cloud-Based Systems: A Systematic Review. Magna Scientia Advanced Research and Reviews, 2026, 17(01), 394-405. Article DOI: https://doi.org/10.30574/msarr.2026.17.1.0114