Deep Learning for Failure Prediction and Self-Healing in Large- Scale Distributed Computing Environments

Authors

  • Manoj Parsuram Cisco Products & Solutions, USA Author

DOI:

https://doi.org/10.21590/

Keywords:

deep learning, failure prediction, self-healing systems, large-scale distributed computing, anomaly detection, long short-term memory, Transformer, graph neural networks, log analysis, multivariate time series, explainable AI, event streaming, zero-downtime deployment, GPU clusters

Abstract

Large-scale distributed computing environments generate telemetry at a volume and
complexity that simple thresholds and hand-built features cannot fully exploit: thousands of
correlated metrics per node, millions of log lines per hour, and dense dependency graphs
between services. Deep learning can learn directly from this raw, high-dimensional data.
Recurrent and stochastic recurrent networks model multivariate time series, attention-based
models read log sequences, and graph neural networks follow faults along service
dependencies. Yet deep models also bring difficulties in production: logs change with every
software release, labeled failures are scarce, inference must keep pace with streaming data,
predictions are opaque, and models go stale. This article proposes a deep learning framework
for failure prediction and self-healing that addresses these difficulties. Telemetry flows
through a high-volume event stream into an ensemble of deep models for metrics, logs, and
dependency graphs. Their outputs are fused, localized to a probable root component, and
explained. A recovery orchestrator acts on them through a risk-governed workflow, and
models are retrained and released without downtime. Using a design-oriented approach
grounded in eleven studies published between 1997 and 2023, the article maps the evidence
base, compares deep architectures, sets out production challenges and responses, and maps
deep model outputs to recovery. It illustrates the design with a multi-tenant GPU cluster used
for training machine learning models. It argues that deep learning's value for self-healing lies
in localizing failures across complex systems, provided its outputs are robust, explained, and
governed.

Downloads

Published

2023-12-14

How to Cite

Parsuram, M. (2023). Deep Learning for Failure Prediction and Self-Healing in Large- Scale Distributed Computing Environments. International Journal of Technology, Management and Humanities, 9(04), 388-397. https://doi.org/10.21590/

Similar Articles

31-40 of 298

You may also start an advanced similarity search for this article.