Publications Partially Reconfigurable Platforms
Relational Fabric: Transparent Data Transformation
2023 IEEE International Conference on Data Engineering (ICDE), April 2023, Anaheim, California, USA
Abstract
A key design decision for data systems is whether they follow the row-store or the column-store paradigm. The former supports transactional workloads, while the latter is better for analytical queries. This decision has a profound impact on the entire data system architecture. The multiple-decadelong journey of these two designs has led to a new family of hybrid transactional/analytical processing (HTAP) architectures. Several efforts have been proposed to reap the benefits of both worlds by proposing systems that maintain multiple copies of data (in different physical layouts) and convert them into the desired layout as required. Due to data duplication, the additional necessary bookkeeping, and the cost of converting data between different layouts, these systems compromise between efficient analytics and data freshness. We depart from existing designs by proposing a radically new approach. We ask the question: “What if we could access any layout and ship only the relevant data through the memory hierarchy by transparently converting rows to (arbitrary groups of) columns?” To achieve this functionality, we capitalize on the reinvigorated trend of hardware specialization (that has been accelerated due to the tapering of Moore’s law) to propose Relational Fabric, a near-data vertical partitioner that allows memory or storage component to perform on-the-fly transparent data transformation. By exposing an intuitive API, Relational Fabric pushes vertical partitioning to the hardware, which has a profound impact on the process of designing and building data systems. (A) There is no need for data duplication and layout conversion, making HTAP systems viable using a single layout. (B) It simplifies the memory and storage manager that needs to maintain and update a single data layout. (C) It reduces unnecessary data movement through the memory hierarchy allowing for better hardware utilization, and ultimately better performance. In this paper, we present Relational Fabric for both memory and storage. We present our initial results on Relational Fabric for in-memory systems and discuss the challenges of building this hardware, as well as the opportunities it brings for simplicity and innovation in the data system software stack, including physical design, query optimization, query evaluation, and concurrency control.
Cite this paper
@inproceedings{RelFab_ICDE23,
title = {{Relational Fabric: Transparent Data Transformation}},
author = {Papon, Tarikul Islam and Mun, Ju-Hyoung and Roozkhosh, Shahin and Hoornaert, Denis and Sanaullah, Ahmed and Drepper, Ulrich and Mancuso, Renato and Athanassoulis, Manos},
booktitle = {2023 IEEE International Conference on Data Engineering (ICDE)},
year = 2023,
month = apr,
address = {Anaheim, California, USA},
url = {https://cs-people.bu.edu/rmancuso/files/papers/RelFab_ICDE23.pdf}
}
TY - CONF AU - Papon, Tarikul Islam AU - Mun, Ju-Hyoung AU - Roozkhosh, Shahin AU - Hoornaert, Denis AU - Sanaullah, Ahmed AU - Drepper, Ulrich AU - Mancuso, Renato AU - Athanassoulis, Manos TI - Relational Fabric: Transparent Data Transformation T2 - 2023 IEEE International Conference on Data Engineering (ICDE) PY - 2023 DA - 2023/04// CY - Anaheim, California, USA AB - A key design decision for data systems is whether they follow the row-store or the column-store paradigm. The former supports transactional workloads, while the latter is better for analytical queries. This decision has a profound impact on the entire data system architecture. The multiple-decadelong journey of these two designs has led to a new family of hybrid transactional/analytical processing (HTAP) architectures. Several efforts have been proposed to reap the benefits of both worlds by proposing systems that maintain multiple copies of data (in different physical layouts) and convert them into the desired layout as required. Due to data duplication, the additional necessary bookkeeping, and the cost of converting data between different layouts, these systems compromise between efficient analytics and data freshness. We depart from existing designs by proposing a radically new approach. We ask the question: “What if we could access any layout and ship only the relevant data through the memory hierarchy by transparently converting rows to (arbitrary groups of) columns?” To achieve this functionality, we capitalize on the reinvigorated trend of hardware specialization (that has been accelerated due to the tapering of Moore’s law) to propose Relational Fabric, a near-data vertical partitioner that allows memory or storage component to perform on-the-fly transparent data transformation. By exposing an intuitive API, Relational Fabric pushes vertical partitioning to the hardware, which has a profound impact on the process of designing and building data systems. (A) There is no need for data duplication and layout conversion, making HTAP systems viable using a single layout. (B) It simplifies the memory and storage manager that needs to maintain and update a single data layout. (C) It reduces unnecessary data movement through the memory hierarchy allowing for better hardware utilization, and ultimately better performance. In this paper, we present Relational Fabric for both memory and storage. We present our initial results on Relational Fabric for in-memory systems and discuss the challenges of building this hardware, as well as the opportunities it brings for simplicity and innovation in the data system software stack, including physical design, query optimization, query evaluation, and concurrency control. ER -
%0 Conference Paper %A Papon, Tarikul Islam %A Mun, Ju-Hyoung %A Roozkhosh, Shahin %A Hoornaert, Denis %A Sanaullah, Ahmed %A Drepper, Ulrich %A Mancuso, Renato %A Athanassoulis, Manos %T Relational Fabric: Transparent Data Transformation %B 2023 IEEE International Conference on Data Engineering (ICDE) %D 2023 %C Anaheim, California, USA %X A key design decision for data systems is whether they follow the row-store or the column-store paradigm. The former supports transactional workloads, while the latter is better for analytical queries. This decision has a profound impact on the entire data system architecture. The multiple-decadelong journey of these two designs has led to a new family of hybrid transactional/analytical processing (HTAP) architectures. Several efforts have been proposed to reap the benefits of both worlds by proposing systems that maintain multiple copies of data (in different physical layouts) and convert them into the desired layout as required. Due to data duplication, the additional necessary bookkeeping, and the cost of converting data between different layouts, these systems compromise between efficient analytics and data freshness. We depart from existing designs by proposing a radically new approach. We ask the question: “What if we could access any layout and ship only the relevant data through the memory hierarchy by transparently converting rows to (arbitrary groups of) columns?” To achieve this functionality, we capitalize on the reinvigorated trend of hardware specialization (that has been accelerated due to the tapering of Moore’s law) to propose Relational Fabric, a near-data vertical partitioner that allows memory or storage component to perform on-the-fly transparent data transformation. By exposing an intuitive API, Relational Fabric pushes vertical partitioning to the hardware, which has a profound impact on the process of designing and building data systems. (A) There is no need for data duplication and layout conversion, making HTAP systems viable using a single layout. (B) It simplifies the memory and storage manager that needs to maintain and update a single data layout. (C) It reduces unnecessary data movement through the memory hierarchy allowing for better hardware utilization, and ultimately better performance. In this paper, we present Relational Fabric for both memory and storage. We present our initial results on Relational Fabric for in-memory systems and discuss the challenges of building this hardware, as well as the opportunities it brings for simplicity and innovation in the data system software stack, including physical design, query optimization, query evaluation, and concurrency control.
[
{
"id": "RelFab_ICDE23",
"type": "paper-conference",
"title": "Relational Fabric: Transparent Data Transformation",
"author": [
{
"family": "Papon",
"given": "Tarikul Islam"
},
{
"family": "Mun",
"given": "Ju-Hyoung"
},
{
"family": "Roozkhosh",
"given": "Shahin"
},
{
"family": "Hoornaert",
"given": "Denis"
},
{
"family": "Sanaullah",
"given": "Ahmed"
},
{
"family": "Drepper",
"given": "Ulrich"
},
{
"family": "Mancuso",
"given": "Renato"
},
{
"family": "Athanassoulis",
"given": "Manos"
}
],
"container-title": "2023 IEEE International Conference on Data Engineering (ICDE)",
"issued": {
"date-parts": [
[
2023,
4
]
]
},
"publisher-place": "Anaheim, California, USA",
"abstract": "A key design decision for data systems is whether they follow the row-store or the column-store paradigm. The former supports transactional workloads, while the latter is better for analytical queries. This decision has a profound impact on the entire data system architecture. The multiple-decadelong journey of these two designs has led to a new family of hybrid transactional/analytical processing (HTAP) architectures. Several efforts have been proposed to reap the benefits of both worlds by proposing systems that maintain multiple copies of data (in different physical layouts) and convert them into the desired layout as required. Due to data duplication, the additional necessary bookkeeping, and the cost of converting data between different layouts, these systems compromise between efficient analytics and data freshness. We depart from existing designs by proposing a radically new approach. We ask the question: “What if we could access any layout and ship only the relevant data through the memory hierarchy by transparently converting rows to (arbitrary groups of) columns?” To achieve this functionality, we capitalize on the reinvigorated trend of hardware specialization (that has been accelerated due to the tapering of Moore’s law) to propose Relational Fabric, a near-data vertical partitioner that allows memory or storage component to perform on-the-fly transparent data transformation. By exposing an intuitive API, Relational Fabric pushes vertical partitioning to the hardware, which has a profound impact on the process of designing and building data systems. (A) There is no need for data duplication and layout conversion, making HTAP systems viable using a single layout. (B) It simplifies the memory and storage manager that needs to maintain and update a single data layout. (C) It reduces unnecessary data movement through the memory hierarchy allowing for better hardware utilization, and ultimately better performance. In this paper, we present Relational Fabric for both memory and storage. We present our initial results on Relational Fabric for in-memory systems and discuss the challenges of building this hardware, as well as the opportunities it brings for simplicity and innovation in the data system software stack, including physical design, query optimization, query evaluation, and concurrency control."
}
]
T. I. Papon, J. H. Mun, S. Roozkhosh, D. Hoornaert, A. Sanaullah, U. Drepper, R. Mancuso, and M. Athanassoulis, “Relational Fabric: Transparent Data Transformation,” in 2023 IEEE International Conference on Data Engineering (ICDE), Apr. 2023.
Papon, T. I., Mun, J. H., Roozkhosh, S., Hoornaert, D., Sanaullah, A., Drepper, U., Mancuso, R., & Athanassoulis, M. (2023). Relational Fabric: Transparent Data Transformation. In 2023 IEEE International Conference on Data Engineering (ICDE).