
Closed
Posted
Paid on delivery
Saya sedang mencari freelancer yang memiliki pengalaman kuat di bidang Big Data, Data Engineering, dan Performance Optimization untuk pendampingan sebuah proyek penelitian dan implementasi sistem analitik data skala besar. Kebutuhan utama meliputi: Analisis dan optimasi penyimpanan data skala besar. Desain arsitektur data berbasis lakehouse. Optimasi format penyimpanan kolumnar (Parquet, ORC, atau sejenisnya). Compression, partitioning, clustering, dan data layout optimization. Query performance tuning. Benchmarking dan performance testing. Analisis metrik performa seperti latency, throughput, storage footprint, bytes scanned/read, CPU dan memory utilization. Penyusunan dokumentasi teknis dan analisis hasil pengujian. Keahlian yang diharapkan: Apache Spark Delta Lake / Apache Iceberg Trino / Presto / Spark SQL Data Lakehouse Architecture Big Data Storage Optimization Data Pipeline dan Data Engineering Benchmarking dan Performance Analysis Nilai tambah apabila: Pernah mengerjakan proyek riset atau publikasi ilmiah. Berpengalaman melakukan eksperimen dan evaluasi performa sistem data skala besar. Memahami metodologi penelitian kuantitatif dan eksperimental. Silakan kirim: Profil dan pengalaman yang relevan. Portofolio proyek yang pernah dikerjakan. Estimasi biaya dan skema kerja yang ditawarkan. Saya mencari kerja sama jangka menengah dengan komunikasi yang baik dan kemampuan teknis yang kuat pada bidang Data Engineering dan Big Data.
Project ID: 40525324
14 proposals
Remote project
Active 4 days ago
Set your budget and timeframe
Get paid for your work
Outline your proposal
It's free to sign up and bid on jobs
14 freelancers are bidding on average $915 USD for this job

This looks like a great fit, Saya akan membantu desain arsitektur lakehouse, optimasi format kolumnar (Parquet/ORC), serta benchmarking performa — mencakup partitioning, clustering, compression, dan query performance tuning menggunakan Spark dan Trino. Satu hal yang akan saya prioritaskan: membandingkan Delta Lake vs Apache Iceberg secara langsung pada dataset Anda — mengukur latency, throughput, bytes scanned, dan storage footprint di setiap format. Hasil ini akan memberikan dasar eksperimental yang kuat untuk dokumentasi riset Anda, bukan sekadar rekomendasi teoretis. Questions: 1) Apakah data sudah tersedia di object storage (S3/GCS), atau perlu pipeline ingestion terlebih dahulu? 2) Apakah target output berupa laporan teknis internal atau publikasi ilmiah dengan metodologi formal? Ready to start whenever you are. Kamran
$859 USD in 13 days
7.8
7.8

Saya memiliki pengalaman dalam Data Engineering, Big Data Processing, dan optimasi performa menggunakan Apache Spark, Delta Lake, Iceberg, serta engine query seperti Trino dan Spark SQL. Saya dapat membantu mulai dari desain arsitektur lakehouse, optimasi storage dan query, benchmarking, analisis metrik performa, hingga dokumentasi teknis yang terstruktur untuk kebutuhan implementasi maupun penelitian. Regards, Muhammad Jibran Ahmed
$1,200 USD in 15 days
5.9
5.9

Hi there, Kami akan membangun arsitektur lakehouse, mengoptimasi format kolumnar (Parquet/ORC), dan menjalankan benchmarking performa secara menyeluruh untuk proyek penelitian Anda. Pendekatan kami: partitioning dan clustering dirancang berdasarkan pola query utama, bukan sekadar volume data. Ini memangkas bytes scanned secara signifikan dan mempercepat latency. Kami juga akan menyusun dokumentasi teknis lengkap beserta analisis hasil pengujian. A couple of quick things to confirm: 1) Apakah infrastruktur sudah tersedia (cloud atau on-premise), atau perlu kami rekomendasikan? 2) Apakah output akhir ditujukan untuk publikasi ilmiah, sehingga perlu mengikuti format metodologi tertentu? The number quoted here is a starting estimate. The exact cost and timeline will be confirmed after we go through the full scope together. Looking forward to talking through the details. Faizan
$844 USD in 13 days
4.3
4.3

Your lakehouse architecture will fail at scale if you don't implement proper Z-ordering and file compaction strategies. I've seen Delta Lake deployments where query times degraded from 2 seconds to 45 seconds after 6 months because of small file proliferation. Before designing the optimization strategy, I need clarity on two things: What's your current data volume and daily ingestion rate (TB/day)? Are you running on Databricks, EMR, or self-managed Spark clusters? Here's the architectural approach: - APACHE SPARK + DELTA LAKE: Implement adaptive query execution with dynamic partition pruning and optimize shuffle operations to reduce stage execution time by 60-70%. - PARQUET OPTIMIZATION: Design column pruning strategies and implement predicate pushdown with bloom filters to minimize bytes scanned from 500GB to under 50GB per query. - LAKEHOUSE ARCHITECTURE: Build a medallion architecture (bronze/silver/gold layers) with automated compaction jobs and vacuum operations to maintain optimal file sizes between 128MB-1GB. - PERFORMANCE BENCHMARKING: Set up TPC-DS or custom benchmark suites measuring query latency, CPU utilization, and I/O throughput across different compression codecs (Snappy vs ZSTD). - TRINO/PRESTO TUNING: Configure memory allocation, join distribution strategies, and cost-based optimizer settings to handle complex analytical queries on petabyte-scale datasets. I've architected similar lakehouse systems for 4 research projects that processed 50TB+ daily with sub-3-second query response times. I'll also deliver comprehensive technical documentation in both English and Bahasa Indonesia covering methodology, experimental results, and optimization recommendations suitable for academic publication. Let's schedule a technical discussion to review your current data pipeline bottlenecks before finalizing the implementation roadmap.
$1,020 USD in 30 days
4.7
4.7

Hi, I read through your project for Bantuan Proyek Big Data dan Optimalisasi Performa. You need help with bantuan proyek big data dan optimalisasi performa. I can build and deliver this efficiently based on the job requirements. This fits well with my experience in Excel. I can take what you outlined and turn it into a clean final delivery, with attention to Excel and the details in your brief. I will keep the implementation clean, practical, and aligned with your requested scope. Send me the details and I can start immediately.
$750 USD in 1 day
3.1
3.1

Halo, saya bisa membantu proyek Big Data dan optimalisasi performa Anda, khususnya pada desain lakehouse, optimasi storage, query tuning, benchmarking, dan dokumentasi teknis untuk kebutuhan riset maupun implementasi sistem analitik skala besar. Solusi terbaik adalah memulai dari review arsitektur dan workload saat ini, lalu menyusun eksperimen terukur untuk membandingkan format penyimpanan seperti Parquet/ORC, strategi compression, partitioning, clustering, data layout, serta performa query menggunakan Spark SQL, Trino/Presto, Delta Lake atau Apache Iceberg. Saya terbiasa dengan Apache Spark, data engineering, data lakehouse architecture, pipeline data, performance tuning, benchmarking, analisis latency, throughput, storage footprint, bytes scanned, CPU/memory usage, serta penyusunan laporan teknis yang rapi dan mudah dipahami. Deliverables dapat mencakup: * Review dan rekomendasi arsitektur data * Optimasi format dan layout penyimpanan * Query performance tuning * Benchmarking dan hasil eksperimen * Analisis metrik performa * Dokumentasi teknis dan ringkasan temuan * Rekomendasi untuk implementasi lanjutan Saya dapat bekerja secara terstruktur, komunikatif, dan mendukung kebutuhan riset/eksperimen dengan pendekatan yang reproducible. Best regards Ankit
$750 USD in 10 days
2.8
2.8

Hi, I read your project description regarding the research and implementation of a large-scale data analytics system. With 8+ years of experience in Data Engineering and Big Data optimization, I’m confident I can support both the research and practical implementation aspects of your project. For the past ~8 years I have been working with one of the largest telecommunications companies in the Netherlands, where I designed and managed large-scale data platforms. In one of the key projects, I implemented and optimized a Hadoop-based cluster with around 140 physical nodes, managing and processing approximately 3 TB of data with highly optimized storage and query performance. My experience closely matches the areas you mentioned: • Data Lakehouse architecture using technologies such as Apache Spark, Delta Lake, and Apache Iceberg • Optimization of columnar storage formats like Parquet and ORC (compression strategies, partitioning, clustering, and data layout optimization) • Query performance tuning using Spark SQL and Trino/Presto • Large-scale data pipeline design and performance optimization • Benchmarking and performance analysis using metrics such as latency, throughput, bytes scanned, CPU utilization, and memory usage I also have strong experience in analyzing system behavior through benchmarking experiments and translating those results into clear technical documentation and recommendations. For collaboration, I usually structure work in phases: Architecture review and system design Benchmarking experiments (storage formats, partitioning strategies, compression) Performance analysis and optimization Documentation of results and recommendations I am available for medium‑term collaboration and would be happy to discuss your research objectives and the scale of the data platform you are planning. Best regards Robert B
$1,250 USD in 6 days
2.6
2.6

The complexity of optimizing large-scale data storage and analytics demands a robust architecture and advanced methodologies. Implementing a data lakehouse model can drastically improve both storage efficiency and query performance. Leveraging tools like Apache Spark and Delta Lake will facilitate effective data manipulation and architecture design. With extensive experience in performance benchmarking, I can deliver thorough metrics analyses covering latency, throughput, and resource utilization. The expected timeframe for the initial deliverable is 30 days. Do you have a specific timeline for the first milestone?
$1,025 USD in 21 days
0.0
0.0

We've just completed a similar project with outstanding results in Big Data Performance Optimization. We can seamlessly assist in achieving your goal through a professional process, leveraging our expertise in Apache Spark, Delta Lake, Trino, and more. I am currently offering discounted pricing on Freelancer.com to build my profile. Remember, if the final work doesn't meet the agreed scope, payment is not required. I'd love to chat about your project! The worst that can happen is you walk away with a free consultation. Regards, Jabu.
$750 USD in 7 days
0.0
0.0

Hello, I came across your project and it immediately caught my attention. I’d be delighted to help you with Bantuan Proyek Big Data dan Optimalisasi Performa. With over 8 years of experience, I specialize in delivering high-quality, professional solutions tailored to client goals rather than generic templates. I handle data, automation, and backend tasks efficiently with a focus on accuracy, speed, and scalability. Please come over chat and discuss your requirement in a detailed way. Best regards, Khadija Amin freelancer.com/u/khadijaamin9
$750 USD in 2 days
2.6
2.6

Indonesia
Member since May 30, 2018
$30-250 USD
$30-250 USD
₹100-400 INR / hour
£10-20 GBP
£10-20 GBP
$250-750 USD
$250-750 USD
₹600-1500 INR
₹100-400 INR / hour
₹600-1500 INR
£10-20 GBP
₹100-400 INR / hour
$30-250 USD
$250-750 USD
₹600-1500 INR