Publications Partially Reconfigurable Platforms
CAESAR: Coherence-Aided Elective and Seamless Alternative Routing via on-chip FPGA
43rd IEEE Real-Time Systems Symposium (RTSS 2022), 2022, Houston, TX, USA
Abstract
Prompted by the ever-growing demand for high-performance System-on-Chip (SoC) and the plateauing of CPU frequencies, the SoC design landscape is shifting. In a quest to offer programmable specialization, the adoption of tightlycoupled FPGAs co-located with traditional compute clusters has been embraced by major vendors. This CPU+FPGA architectural paradigm opens the door to novel hardware/software co-design opportunities. The key principle is that CPU-originated memory traffic can be re-routed through the FPGA for analysis and management purposes. Albeit promising, the side-effect of this approach is that time-critical operations—such as cache-line refills—are fulfilled by moving data over slower interconnects meant for I/O traffic. In this article, we introduce a novel principle named Cache Coherence Backstabbing to precisely tackle these shortcomings. The technique leverages the ability to include the FGPA in the same coherence domain as the core processing elements. Importantly, this enables Coherence-Aided Elective and Seamless Alternative Routing (CAESAR), i.e., seamless inspection and routing of memory transactions, especially cache-line refills, through the FPGA. CAESAR allows the definition of new memory programming paradigms. We discuss the intrinsic potentials of the approach and evaluate it with a full-stack prototype implementation on a commercial platform. Our experiments show an improvement of up to 29% in read bandwidth, 23% in latency, and 13% in pragmatic workloads over the state of the art. Furthermore, we showcase the first in-coherence-domain run-time profiler design as a use-case of the CAESAR approach.
Code and hardware artifacts
Cite this paper
@inproceedings{CAESAR_RTSS22,
title = {{CAESAR: Coherence-Aided Elective and Seamless Alternative Routing via on-chip FPGA}},
author = {Roozkhosh, Shahin and Hoornaert, Denis and Mancuso, Renato},
booktitle = {43rd IEEE Real-Time Systems Symposium (RTSS 2022)},
year = 2022,
pages = {356--369},
address = {Houston, TX, USA},
doi = {10.1109/RTSS55097.2022.00038},
url = {https://doi.org/10.1109/RTSS55097.2022.00038}
}
TY - CONF AU - Roozkhosh, Shahin AU - Hoornaert, Denis AU - Mancuso, Renato TI - CAESAR: Coherence-Aided Elective and Seamless Alternative Routing via on-chip FPGA T2 - 43rd IEEE Real-Time Systems Symposium (RTSS 2022) PY - 2022 CY - Houston, TX, USA SP - 356 EP - 369 DO - 10.1109/RTSS55097.2022.00038 UR - https://doi.org/10.1109/RTSS55097.2022.00038 AB - Prompted by the ever-growing demand for high-performance System-on-Chip (SoC) and the plateauing of CPU frequencies, the SoC design landscape is shifting. In a quest to offer programmable specialization, the adoption of tightlycoupled FPGAs co-located with traditional compute clusters has been embraced by major vendors. This CPU+FPGA architectural paradigm opens the door to novel hardware/software co-design opportunities. The key principle is that CPU-originated memory traffic can be re-routed through the FPGA for analysis and management purposes. Albeit promising, the side-effect of this approach is that time-critical operations—such as cache-line refills—are fulfilled by moving data over slower interconnects meant for I/O traffic. In this article, we introduce a novel principle named Cache Coherence Backstabbing to precisely tackle these shortcomings. The technique leverages the ability to include the FGPA in the same coherence domain as the core processing elements. Importantly, this enables Coherence-Aided Elective and Seamless Alternative Routing (CAESAR), i.e., seamless inspection and routing of memory transactions, especially cache-line refills, through the FPGA. CAESAR allows the definition of new memory programming paradigms. We discuss the intrinsic potentials of the approach and evaluate it with a full-stack prototype implementation on a commercial platform. Our experiments show an improvement of up to 29% in read bandwidth, 23% in latency, and 13% in pragmatic workloads over the state of the art. Furthermore, we showcase the first in-coherence-domain run-time profiler design as a use-case of the CAESAR approach. ER -
%0 Conference Paper %A Roozkhosh, Shahin %A Hoornaert, Denis %A Mancuso, Renato %T CAESAR: Coherence-Aided Elective and Seamless Alternative Routing via on-chip FPGA %B 43rd IEEE Real-Time Systems Symposium (RTSS 2022) %D 2022 %P 356-369 %C Houston, TX, USA %R 10.1109/RTSS55097.2022.00038 %U https://doi.org/10.1109/RTSS55097.2022.00038 %X Prompted by the ever-growing demand for high-performance System-on-Chip (SoC) and the plateauing of CPU frequencies, the SoC design landscape is shifting. In a quest to offer programmable specialization, the adoption of tightlycoupled FPGAs co-located with traditional compute clusters has been embraced by major vendors. This CPU+FPGA architectural paradigm opens the door to novel hardware/software co-design opportunities. The key principle is that CPU-originated memory traffic can be re-routed through the FPGA for analysis and management purposes. Albeit promising, the side-effect of this approach is that time-critical operations—such as cache-line refills—are fulfilled by moving data over slower interconnects meant for I/O traffic. In this article, we introduce a novel principle named Cache Coherence Backstabbing to precisely tackle these shortcomings. The technique leverages the ability to include the FGPA in the same coherence domain as the core processing elements. Importantly, this enables Coherence-Aided Elective and Seamless Alternative Routing (CAESAR), i.e., seamless inspection and routing of memory transactions, especially cache-line refills, through the FPGA. CAESAR allows the definition of new memory programming paradigms. We discuss the intrinsic potentials of the approach and evaluate it with a full-stack prototype implementation on a commercial platform. Our experiments show an improvement of up to 29% in read bandwidth, 23% in latency, and 13% in pragmatic workloads over the state of the art. Furthermore, we showcase the first in-coherence-domain run-time profiler design as a use-case of the CAESAR approach.
[
{
"id": "CAESAR_RTSS22",
"type": "paper-conference",
"title": "CAESAR: Coherence-Aided Elective and Seamless Alternative Routing via on-chip FPGA",
"author": [
{
"family": "Roozkhosh",
"given": "Shahin"
},
{
"family": "Hoornaert",
"given": "Denis"
},
{
"family": "Mancuso",
"given": "Renato"
}
],
"container-title": "43rd IEEE Real-Time Systems Symposium (RTSS 2022)",
"issued": {
"date-parts": [
[
2022
]
]
},
"page": "356-369",
"publisher-place": "Houston, TX, USA",
"DOI": "10.1109/RTSS55097.2022.00038",
"abstract": "Prompted by the ever-growing demand for high-performance System-on-Chip (SoC) and the plateauing of CPU frequencies, the SoC design landscape is shifting. In a quest to offer programmable specialization, the adoption of tightlycoupled FPGAs co-located with traditional compute clusters has been embraced by major vendors. This CPU+FPGA architectural paradigm opens the door to novel hardware/software co-design opportunities. The key principle is that CPU-originated memory traffic can be re-routed through the FPGA for analysis and management purposes. Albeit promising, the side-effect of this approach is that time-critical operations—such as cache-line refills—are fulfilled by moving data over slower interconnects meant for I/O traffic. In this article, we introduce a novel principle named Cache Coherence Backstabbing to precisely tackle these shortcomings. The technique leverages the ability to include the FGPA in the same coherence domain as the core processing elements. Importantly, this enables Coherence-Aided Elective and Seamless Alternative Routing (CAESAR), i.e., seamless inspection and routing of memory transactions, especially cache-line refills, through the FPGA. CAESAR allows the definition of new memory programming paradigms. We discuss the intrinsic potentials of the approach and evaluate it with a full-stack prototype implementation on a commercial platform. Our experiments show an improvement of up to 29% in read bandwidth, 23% in latency, and 13% in pragmatic workloads over the state of the art. Furthermore, we showcase the first in-coherence-domain run-time profiler design as a use-case of the CAESAR approach."
}
]
S. Roozkhosh, D. Hoornaert, and R. Mancuso, “CAESAR: Coherence-Aided Elective and Seamless Alternative Routing via on-chip FPGA,” in 43rd IEEE Real-Time Systems Symposium (RTSS 2022), pp. 356–369, 2022, doi: 10.1109/RTSS55097.2022.00038.
Roozkhosh, S., Hoornaert, D., & Mancuso, R. (2022). CAESAR: Coherence-Aided Elective and Seamless Alternative Routing via on-chip FPGA. In 43rd IEEE Real-Time Systems Symposium (RTSS 2022) (pp. 356–369). https://doi.org/10.1109/RTSS55097.2022.00038