Reinforcement Learning for Auto-Scaling and Cost Optimization in Multi-Cloud
Abstract
Auto-scaling decides how much capacity to provision for a workload and — in a multi-cloud setting — where to place it across providers, regions, and pricing models. Traditional threshold and forecast-based autoscalers react with fixed rules and tend to over-provision, wasting money, or under-provision, breaching service-level agreements (SLAs). This article presents reinforcement learning (RL) as a principled alternative: an agent that learns a cost-aware, SLA-constrained scaling and placement policy directly from interaction with the environment. We formalize multi-cloud auto-scaling as a Markov decision process, design a composite reward that trades resource cost against SLA penalties, utilization, and scaling churn, and extend the action space to multi-cloud placement — choosing provider, region, and on-demand versus spot capacity to exploit cross-cloud price differences and spot discounts while respecting interruption risk and data-residency constraints. We review the dominant algorithms (DQN, PPO, DDPG), discuss safe exploration via bootstrapping from a conventional autoscaler, and give illustrative results consistent with the published literature, where RL autoscalers report on the order of 20% cost savings and improved tail latency relative to threshold-based baselines. Public technical sources are cited throughout.
References
RESOURCE, A. O. M. C. (2024). ALLOCATION FOR COST-EFFICIENT COMPUTING. International Journal of Information Technology (IJIT), 5(2).
Nerella, V. M. L. G., Mahavratayajula, S., & Janardhanan, H. (2023). Machine Learning-Driven Finops Strategies: Adaptive Scaling Models For Balancing Reliability And Cost In Multi-Cloud Data Platforms. Machine Learning, 6(4).
Sharma, R., Singh, C., & Kaur, P. (2026, February). AI-Driven Cloud Resource Auto-Scaling in Multi-Cloud Environments Using Reinforcement Learning. In 2026 9th International Conference on Electronics, Materials Engineering & Nano-Technology (IEMENTech) (Vol. 9, pp. 1-6). IEEE.
Aashish, K. C., Sultana, N., Ahmed, K. R., Raihan, K. A., Ali, A. H. M. M., & Salam, T. (2026, April). CARL: A Cost-Aware Reinforcement Learning Framework for Efficient and SLA-Aware Multi-Cloud Auto-Scaling. In 2026 IEEE 2nd International Conference on Quantum Photonics, Artificial Intelligence & Networking (QPAIN) (pp. 1-6). IEEE.
Aashish, K. C., Sultana, N., Ahmed, K. R., Raihan, K. A., Ali, A. H. M. M., & Salam, T. (2026, April). CARL: A Cost-Aware Reinforcement Learning Framework for Efficient and SLA-Aware Multi-Cloud Auto-Scaling. In 2026 IEEE 2nd International Conference on Quantum Photonics, Artificial Intelligence & Networking (QPAIN) (pp. 1-6). IEEE.
Alharthi, S., Alshamsi, A., Alseiari, A., & Alwarafy, A. (2024). Auto-scaling techniques in cloud computing: Issues and research directions. Sensors, 24(17), 5551.
Refbacks
- There are currently no refbacks.