Publications Memory & Shared-Resource Management
Profile-driven memory bandwidth management for accelerators and CPUs in QoS-enabled platforms
Real-Time Systems, December 2022
Abstract
The proliferation of multi-core, accelerator-enabled embedded systems has introduced new opportunities to consolidate real-time systems of increasing complexity. But the road to build confidence on the temporal behavior of co-running applications has presented formidable challenges. Most prominently, the main memory subsystem represents a performance bottleneck for both CPUs and accelerators. And industry-viable frameworks for full-system main memory management and performance analysis are past due. In this paper, we propose our Envelope-aWare Predictive model , or E-WarP for short. E-WarP is a methodology and technological framework to: (1) analyze the memory demand of applications following a profile-driven approach; (2) make realistic predictions on the temporal behavior of workload deployed on CPUs and accelerators; and (3) perform saturation-aware system consolidation. This work aims at providing the technological foundations as well as the theoretical grassroots for truly workload-aware analysis of real-time systems. This work combines traditional CPU-centric bandwidth regulation techniques with state-of-the-art hardware support for memory traffic shaping via the ARM QoS extensions. We make three key observations. First, our profile-driven methodology achieves, on average, 6% over-prediction on the runtime of bandwidth-regulated applications. Second, we experimentally validate that the calculated bounds hold system-wide if the main memory subsystem operates below saturation. Third, we show that the E-WarP methodology is practical even when applications exhibit input-dependent memory access patterns. We provide a full implementation of our techniques on a commercial platform (NXP S32V234).
Code and hardware artifacts
Cite this paper
@article{profile_QoS_RTSJ22,
title = {{Profile-driven memory bandwidth management for accelerators and CPUs in QoS-enabled platforms}},
author = {Sohal, Parul and Tabish, Rohan and Drepper, Ulrich and Mancuso, Renato},
journal = {Real-Time Systems},
year = 2022,
month = dec,
publisher = {Springer},
doi = {10.1007/s11241-022-09382-x},
url = {https://doi.org/10.1007/s11241-022-09382-x}
}
TY - JOUR AU - Sohal, Parul AU - Tabish, Rohan AU - Drepper, Ulrich AU - Mancuso, Renato TI - Profile-driven memory bandwidth management for accelerators and CPUs in QoS-enabled platforms JO - Real-Time Systems PY - 2022 DA - 2022/12// PB - Springer DO - 10.1007/s11241-022-09382-x UR - https://doi.org/10.1007/s11241-022-09382-x AB - The proliferation of multi-core, accelerator-enabled embedded systems has introduced new opportunities to consolidate real-time systems of increasing complexity. But the road to build confidence on the temporal behavior of co-running applications has presented formidable challenges. Most prominently, the main memory subsystem represents a performance bottleneck for both CPUs and accelerators. And industry-viable frameworks for full-system main memory management and performance analysis are past due. In this paper, we propose our Envelope-aWare Predictive model , or E-WarP for short. E-WarP is a methodology and technological framework to: (1) analyze the memory demand of applications following a profile-driven approach; (2) make realistic predictions on the temporal behavior of workload deployed on CPUs and accelerators; and (3) perform saturation-aware system consolidation. This work aims at providing the technological foundations as well as the theoretical grassroots for truly workload-aware analysis of real-time systems. This work combines traditional CPU-centric bandwidth regulation techniques with state-of-the-art hardware support for memory traffic shaping via the ARM QoS extensions. We make three key observations. First, our profile-driven methodology achieves, on average, 6% over-prediction on the runtime of bandwidth-regulated applications. Second, we experimentally validate that the calculated bounds hold system-wide if the main memory subsystem operates below saturation. Third, we show that the E-WarP methodology is practical even when applications exhibit input-dependent memory access patterns. We provide a full implementation of our techniques on a commercial platform (NXP S32V234). ER -
%0 Journal Article %A Sohal, Parul %A Tabish, Rohan %A Drepper, Ulrich %A Mancuso, Renato %T Profile-driven memory bandwidth management for accelerators and CPUs in QoS-enabled platforms %J Real-Time Systems %D 2022 %I Springer %R 10.1007/s11241-022-09382-x %U https://doi.org/10.1007/s11241-022-09382-x %X The proliferation of multi-core, accelerator-enabled embedded systems has introduced new opportunities to consolidate real-time systems of increasing complexity. But the road to build confidence on the temporal behavior of co-running applications has presented formidable challenges. Most prominently, the main memory subsystem represents a performance bottleneck for both CPUs and accelerators. And industry-viable frameworks for full-system main memory management and performance analysis are past due. In this paper, we propose our Envelope-aWare Predictive model , or E-WarP for short. E-WarP is a methodology and technological framework to: (1) analyze the memory demand of applications following a profile-driven approach; (2) make realistic predictions on the temporal behavior of workload deployed on CPUs and accelerators; and (3) perform saturation-aware system consolidation. This work aims at providing the technological foundations as well as the theoretical grassroots for truly workload-aware analysis of real-time systems. This work combines traditional CPU-centric bandwidth regulation techniques with state-of-the-art hardware support for memory traffic shaping via the ARM QoS extensions. We make three key observations. First, our profile-driven methodology achieves, on average, 6% over-prediction on the runtime of bandwidth-regulated applications. Second, we experimentally validate that the calculated bounds hold system-wide if the main memory subsystem operates below saturation. Third, we show that the E-WarP methodology is practical even when applications exhibit input-dependent memory access patterns. We provide a full implementation of our techniques on a commercial platform (NXP S32V234).
[
{
"id": "profile_QoS_RTSJ22",
"type": "article-journal",
"title": "Profile-driven memory bandwidth management for accelerators and CPUs in QoS-enabled platforms",
"author": [
{
"family": "Sohal",
"given": "Parul"
},
{
"family": "Tabish",
"given": "Rohan"
},
{
"family": "Drepper",
"given": "Ulrich"
},
{
"family": "Mancuso",
"given": "Renato"
}
],
"container-title": "Real-Time Systems",
"issued": {
"date-parts": [
[
2022,
12
]
]
},
"publisher": "Springer",
"DOI": "10.1007/s11241-022-09382-x",
"abstract": "The proliferation of multi-core, accelerator-enabled embedded systems has introduced new opportunities to consolidate real-time systems of increasing complexity. But the road to build confidence on the temporal behavior of co-running applications has presented formidable challenges. Most prominently, the main memory subsystem represents a performance bottleneck for both CPUs and accelerators. And industry-viable frameworks for full-system main memory management and performance analysis are past due. In this paper, we propose our Envelope-aWare Predictive model , or E-WarP for short. E-WarP is a methodology and technological framework to: (1) analyze the memory demand of applications following a profile-driven approach; (2) make realistic predictions on the temporal behavior of workload deployed on CPUs and accelerators; and (3) perform saturation-aware system consolidation. This work aims at providing the technological foundations as well as the theoretical grassroots for truly workload-aware analysis of real-time systems. This work combines traditional CPU-centric bandwidth regulation techniques with state-of-the-art hardware support for memory traffic shaping via the ARM QoS extensions. We make three key observations. First, our profile-driven methodology achieves, on average, 6% over-prediction on the runtime of bandwidth-regulated applications. Second, we experimentally validate that the calculated bounds hold system-wide if the main memory subsystem operates below saturation. Third, we show that the E-WarP methodology is practical even when applications exhibit input-dependent memory access patterns. We provide a full implementation of our techniques on a commercial platform (NXP S32V234)."
}
]
P. Sohal, R. Tabish, U. Drepper, and R. Mancuso, “Profile-driven memory bandwidth management for accelerators and CPUs in QoS-enabled platforms,” Real-Time Systems, Dec. 2022, doi: 10.1007/s11241-022-09382-x.
Sohal, P., Tabish, R., Drepper, U., & Mancuso, R. (2022). Profile-driven memory bandwidth management for accelerators and CPUs in QoS-enabled platforms. Real-Time Systems. Springer. https://doi.org/10.1007/s11241-022-09382-x