Design and Evaluation of a Lakehouse Architecture Integrating Data Lakes and Data Warehouses
Abstract
The increasing volume and complexity of organizational data have exposed the limitations of traditional data management platforms in supporting modern analytical workloads. While data warehouses provide high-performance querying, strong governance, and structured data management, they are less suitable for handling diverse and rapidly growing datasets. Conversely, data lakes offer scalable and cost-effective storage for structured, semi-structured, and unstructured data but often face challenges related to data quality, governance, and query performance. Lakehouse architecture has emerged as a unified approach that integrates the strengths of data lakes and data warehouses to support both business intelligence and advanced analytics on a single platform. This study aims to design and evaluate a Lakehouse architecture integrating data lakes and data warehouses. The proposed architecture will be implemented using cloud-native technologies and evaluated through standardized analytical workloads. Performance will be assessed using metrics including query execution time, throughput, scalability, storage efficiency, resource utilization, and cost efficiency. Comparative analysis will be conducted against traditional data lake and data warehouse architectures to determine the effectiveness of the proposed solution. The study is expected to demonstrate that the Lakehouse architecture provides a balanced platform that combines the scalability and flexibility of data lakes with the reliability, governance, and analytical performance of data warehouses. The findings will contribute to the advancement of cloud-native data management and provide practical guidance for organizations seeking efficient and scalable analytical platforms.
References
Chaudhuri, S., & Dayal, U. (1997). An overview of data warehousing and OLAP technology. ACM SIGMOD Record, 26(1), 65–74. https://doi.org/10.1145/248603.248616
Hai, R., Koutras, C., Quix, C., & Jarke, M. (2023). Data lakes: A survey of functions and systems. IEEE Transactions on Knowledge and Data Engineering, 35(12), 12571–12590. https://doi.org/10.1109/TKDE.2023.3270101
Harby, A. A., & Zulkernine, F. H. (2025). Data lakehouse: A survey and experimental study. Information Systems, 127, 102460. https://doi.org/10.1016/j.is.2024.102460
Kimball, R., & Ross, M. (2013). The data warehouse toolkit: The definitive guide to dimensional modeling (3rd ed.). Wiley.
Oliveira e Sá, J., Gonçalves, R., & Kaldeich, C. (2024). Benchmark of market cloud data warehouse technologies. Procedia Computer Science, 246, 4077–4086. https://doi.org/10.1016/j.procs.2024.10.350
Ouda, A. (2024). Data lakes: A survey of concepts and architectures. Computers, 13(7), 183. https://doi.org/10.3390/computers13070183
Sawadogo, P. N., & Darmont, J. (2023). DLBench+: A benchmark for quantitative and qualitative data lake assessment. Data & Knowledge Engineering, 144, 102154. https://doi.org/10.1016/j.datak.2023.102154
Schneider, J., Gröger, C., Lutsch, A., Schwarz, H., & Mitschang, B. (2024). The lakehouse: State of the art on concepts and technologies. SN Computer Science, 5, 449. https://doi.org/10.1007/s42979-024-02737-0
Szőke, M.-A. (2026). Modern alternatives to Hive: A systematic review and single-node benchmark of SQL-on-Hadoop and lakehouse engines. Information Systems, 136, 102635. https://doi.org/10.1016/j.is.2025.102635
Zaharia, M., Xin, R. S., Wendell, P., Das, T., Armbrust, M., Dave, A., Meng, X., Rosen, J., Venkataraman, S., Franklin, M. J., Ghodsi, A., Gonzalez, J. E., Shenker, S., & Stoica, I. (2016). Apache Spark: A unified engine for big data processing. Communications of the ACM, 59(11), 56–65. https://doi.org/10.1145/2934664
Refbacks
- There are currently no refbacks.