Алгоритм динамического планирования на основе функции Ляпунова для гетерогенных вычислительных кластеров

Моделирование вычислительных, телекоммуникационных и управляющих систем
Авторы:
Аннотация:

В статье рассматривается алгоритм динамического планирования на основе функции Ляпунова для гетерогенных вычислительных кластеров, ориентированный на точное управление ресурсами при импульсных и чувствительных к задержкам рабочих нагрузках. Путем построения квадратичной функции Ляпунова и применения подхода drift-plus-penalty задача планирования формулируется как задача двухкритериальной оптимизации для стабильности очереди и задержки планирования. Вводится параметр динамического управления V для количественного регулирования компромисса между стабильностью очереди и минимизацией задержки. Анализ чувствительности демонстрирует компромисс между O (1/V) очереди и O (V) задержки. Эксперименты, проведенные на наборе данных трассировки кластера GPU Alibaba, показывают, что при импульсных рабочих нагрузках предложенный метод снижает среднюю задержку планирования до 0,2663 сек, при этом достигая коэффициента использования ресурсов 0,5459 и индекса справедливости 0,6489. Данный метод особенно хорошо подходит для чувствительных к задержкам и динамически изменяющихся рабочих окружений.

  • Список литературы

    1. Ismail A.A., Khalifa N.E., El-Khoribi R.A. A survey on resource scheduling approaches in multi-access edge computing environment: a deep reinforcement learning study. Cluster Computing, 2025, Vol. 28, Art. no. 184. DOI: 10.1007/s10586-024-04893-7

    2. Polo J., Castillo C., Carrera D., Becerra Y., Whalley I., Steinder M., Torres J., Ayguadé E. Resource-aware adaptive scheduling for MapReduce clusters. In: Middleware 2011: Lecture Notes in Computer Science (eds. F. Kon, A.M. Kermarrec), 2011, Vol. 7049, Pp. 187–207. DOI: 10.1007/978-3-642-25821-3_10

    3. Chen Y., Griffith R., Liu J., Katz R.H., Joseph A.D. Understanding TCP incast throughput collapse in datacenter networks. Proceedings of the 1st ACM Workshop on Research on Enterprise Networking, 2009, Pp. 73–82. DOI: 10.1145/1592681.1592693

    4. Hindman B., Konwinski A., Zaharia M., Ghodsi A., Joseph A.D., Katz R., Shenker S., Stoica I. Mesos: A platform for fine-grained resource sharing in the data center. Proceedings of the 8th USENIX Symposium on Networked Systems Design and Implementation, 2011, Pp. 295–308.

    5. Neely M.J. Stochastic Network Optimization with Application to Communication and Queueing Systems. Cham: Springer, 2010. DOI: 10.1007/978-3-031-79995-2

    6. Shi Y., Yang K., Jiang T., Zhang J., Letaief K.B. Communication-efficient edge AI: Algorithms and systems. arXiv:2002.09668, 2020. DOI: 10.48550/arXiv.2002.09668

    7. Shahrad M., Fonseca R., Goiri Í., Chaudhry G., Batum P., Cooke J., Laureano E., Tresness C., Russinovich M., Bianchini R. Serverless in the wild: characterizing and optimizing the serverless workload at a large cloud provider. Proceedings of the 2020 USENIX Conference on Usenix Annual Technical Conference, 2020, Pp. 205–218.

    8. Zhang J., Zhai Y., Liu Z., Wang Y. A Lyapunov-based resource allocation method for edge-assisted industrial internet of things. IEEE Internet of Things Journal, 2024, Vol. 11, No. 24, Pp. 39464–39472. DOI: 10.1109/JIOT.2024.3446722

    9. Gao Y., Liu L., Zheng X., Zhang C., Ma H. Federated sensing: Edge-cloud elastic collaborative learning for intelligent sensing. IEEE Internet of Things Journal, 2021, Vol. 8, No. 14, Pp. 11100–11111. DOI: 10.1109/JIOT.2021.3053055

    10. Tang S., He B.-S., Zhang S., Niu Z. Elastic multi-resource fairness: balancing fairness and efficiency in coupled CPU-GPU architectures. Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis, 2016, Pp. 875–886. DOI: 10.1109/SC.2016.74

    11. Verma A., Pedrosa L., Korupolu M., Oppenheimer D., Tune E., Wilkes J. Large-scale cluster management at Google with Borg. Proceedings of the 10th European Conference on Computer Systems, 2015, Art. no. 18. DOI: 10.1145/2741948.2741964

    12. Reiss C., Wilkes J. Google cluster-usage traces: format + schema. Google Inc. Technical Report, 2011.

    13. Burns B., Grant B., Oppenheimer D., Brewer E., Wilkes J. Borg, Omega, and Kubernetes. Communications of the ACM, 2016, Vol. 59, No. 5, Pp. 50–57. DOI: 10.1145/2890784

    14. Ghodsi A., Zaharia M., Hindman B., Konwinski A., Shenker S., Stoica I. Dominant resource fairness: fair allocation of multiple resource types. Proceedings of the 8th USENIX Conference on Networked Systems Design and Implementation, 2011, Pp. 323–336.

    15. Xiao W., Bhardwaj R., Ramjee R. et al. Gandiva: introspective cluster scheduling for deep learning workloads. Proceedings of the 13th USENIX Conference on Operating Systems Design and Implementation, 2018, Pp. 595–610.

    16. Zhao X., Yao J., Gao P., Guan H. Efficient sharing and fine-grained scheduling of virtualized GPU resources. 2018 IEEE 38th International Conference on Distributed Computing Systems (ICDCS), 2018, Pp. 742–752. DOI: 10.1109/ICDCS.2018.00077

    17. Sukhoroslov O. Building web-based services for practical exercises in parallel and distributed computing. Journal of Parallel and Distributed Computing, 2018, Vol. 118 (1), Pp. 177–188. DOI: 10.1016/j.jpdc.2018.02.024

    18. Mao Y., You C., Zhang J., Huang K., Letaief K.B. A survey on mobile edge computing: The communication perspective. IEEE Communications Surveys & Tutorials, 2017, Vol. 19, No. 4, Pp. 2322–2358. DOI: 10.1109/COMST.2017.2745201

    19. Beloglazov A., Buyya R. Optimal online deterministic algorithms and adaptive heuristics for energy and performance efficient dynamic consolidation of virtual machines in cloud data centers. Concurrency and Computation: Practice and Experience, 2012, Vol. 24, No. 13, Pp. 1397–1420. DOI: 10.1002/cpe.1867

    20. Smorodnikov G., Zolotarev R., Rykova A., Sabutkevich A., Samochadin A. Elastic cloud resource allocation using short-term long short-term memory-based workload prediction. Proceedings of the 4th International Conference on Optics, Computer Applications, and Materials Science (CMSD-IV 2024), 2025, Vol. 13651, Art. no. 136510J. DOI: 10.1117/12.3060861

    21. Sukhoroslov O., Nazarenko A., Aleksandrov R. An experimental study of scheduling algorithms for many-task applications. The Journal of Supercomputing, 2019, Vol. 75, Pp. 7857–7871. DOI: 10.1007/s11227-018-2553-9

    22. Sukhoroslov O. Supporting efficient execution of workflows on Everest platform. Supercomputing (RuSCDays), 2019, Pp. 713–724. DOI: 10.1007/978-3-030-36592-9_58

    23. Peng Y., Bao Y., Chen Y., Wu C., Guo C. Optimus: an efficient dynamic resource scheduler for deep learning clusters. Proceedings of the 13th EuroSys Conference, 2018, Art. no. 3. DOI: 10.1145/3190508.3190517

This work is licensed under a Creative Commons Attribution-NonCommercial 4.0 International License
Предыдущая статьяСледующая статья