Open Access Open Access  Restricted Access Subscription Access

Design and Evaluation of a Lakehouse Architecture Integrating Data Lakes and Data Warehouses

Mission Franklin

Abstract


The increasing volume and complexity of organizational data have exposed the limitations of traditional data management platforms in supporting modern analytical workloads. While data warehouses provide high-performance querying, strong governance, and structured data management, they are less suitable for handling diverse and rapidly growing datasets. Conversely, data lakes offer scalable and cost-effective storage for structured, semi-structured, and unstructured data but often face challenges related to data quality, governance, and query performance. Lakehouse architecture has emerged as a unified approach that integrates the strengths of data lakes and data warehouses to support both business intelligence and advanced analytics on a single platform. This study aims to design and evaluate a Lakehouse architecture integrating data lakes and data warehouses. The proposed architecture will be implemented using cloud-native technologies and evaluated through standardized analytical workloads. Performance will be assessed using metrics including query execution time, throughput, scalability, storage efficiency, resource utilization, and cost efficiency. Comparative analysis will be conducted against traditional data lake and data warehouse architectures to determine the effectiveness of the proposed solution. The study is expected to demonstrate that the Lakehouse architecture provides a balanced platform that combines the scalability and flexibility of data lakes with the reliability, governance, and analytical performance of data warehouses. The findings will contribute to the advancement of cloud-native data management and provide practical guidance for organizations seeking efficient and scalable analytical platforms.


Full Text:

PDF

References


Chaudhuri, S., & Dayal, U. (1997). An overview of data warehousing and OLAP technology. ACM SIGMOD Record, 26(1), 65–74. https://doi.org/10.1145/248603.248616

Hai, R., Koutras, C., Quix, C., & Jarke, M. (2023). Data lakes: A survey of functions and systems. IEEE Transactions on Knowledge and Data Engineering, 35(12), 12571–12590. https://doi.org/10.1109/TKDE.2023.3270101

Harby, A. A., & Zulkernine, F. H. (2025). Data lakehouse: A survey and experimental study. Information Systems, 127, 102460. https://doi.org/10.1016/j.is.2024.102460

Kimball, R., & Ross, M. (2013). The data warehouse toolkit: The definitive guide to dimensional modeling (3rd ed.). Wiley.

Oliveira e Sá, J., Gonçalves, R., & Kaldeich, C. (2024). Benchmark of market cloud data warehouse technologies. Procedia Computer Science, 246, 4077–4086. https://doi.org/10.1016/j.procs.2024.10.350

Ouda, A. (2024). Data lakes: A survey of concepts and architectures. Computers, 13(7), 183. https://doi.org/10.3390/computers13070183

Sawadogo, P. N., & Darmont, J. (2023). DLBench+: A benchmark for quantitative and qualitative data lake assessment. Data & Knowledge Engineering, 144, 102154. https://doi.org/10.1016/j.datak.2023.102154

Schneider, J., Gröger, C., Lutsch, A., Schwarz, H., & Mitschang, B. (2024). The lakehouse: State of the art on concepts and technologies. SN Computer Science, 5, 449. https://doi.org/10.1007/s42979-024-02737-0

Szőke, M.-A. (2026). Modern alternatives to Hive: A systematic review and single-node benchmark of SQL-on-Hadoop and lakehouse engines. Information Systems, 136, 102635. https://doi.org/10.1016/j.is.2025.102635

Zaharia, M., Xin, R. S., Wendell, P., Das, T., Armbrust, M., Dave, A., Meng, X., Rosen, J., Venkataraman, S., Franklin, M. J., Ghodsi, A., Gonzalez, J. E., Shenker, S., & Stoica, I. (2016). Apache Spark: A unified engine for big data processing. Communications of the ACM, 59(11), 56–65. https://doi.org/10.1145/2934664


Refbacks

  • There are currently no refbacks.