
Closed
Posted
Paid on delivery
I have roughly 500 Arabic-language PDFs (books and administrative letters) stored in a single folder on my Windows machine. Some are digitally generated while others are scanned images, so reliable, high-accuracy OCR is essential; ordinary solutions have failed because the text quality is only medium and many pages need careful pre-processing before recognition. What I need the script to do: • Read every PDF in the folder in one pass. • For each file pull out four data points—رقم الكتاب (book number), تاريخ الكتاب (in the existing DD-MM-YYYY format), موضوع الكتاب (subject) and اسم الدائرة أو الجهة المرسلة (issuing department). • Append these fields to an existing Excel workbook, filling only the blank rows so nothing already entered is duplicated. Each new row must also contain a working hyperlink that opens the corresponding PDF. • Run fully offline on Windows; no cloud calls or external APIs. • Handle the mixed batch gracefully: where the PDF already contains selectable text, extract it directly; where it is an image, trigger high-quality OCR after image enhancement (noise removal, skew correction, contrast boost, etc.). • Deliver the full, well-documented Python source code, ready to execute, along with a concise “how to run” guide. For testing, I will supply the Excel file, a subset of real PDFs and sample expected output. I will consider the job complete when the script processes the test set end-to-end without manual intervention and shows near-perfect field accuracy.
Project ID: 40548068
29 proposals
Remote project
Active 2 days ago
Set your budget and timeframe
Get paid for your work
Outline your proposal
It's free to sign up and bid on jobs
29 freelancers are bidding on average $24 USD for this job

AVAILABLE TO START IMMEDIATELY..,,. I will deliver a robust Python script using advanced OCR and image processing to accurately extract specified Arabic data from your mixed PDFs and append it to Excel with hyperlinks, running fully offline on Windows. 10+ years Advanced Excel experience, Certified VBA Programmer, MBA.
$19 USD in 1 day
6.6
6.6

Hello I can develop a fully offline Python script for Windows that accurately processes both searchable and scanned Arabic PDFs using advanced OCR and image enhancement, extracts the required fields, updates only blank rows in your Excel file with PDF hyperlinks, and provides clean, well-documented code with an easy setup guide. Regards Muhammad
$30 USD in 1 day
6.1
6.1

Hi, You’re not just looking for OCR, you need a dependable offline pipeline that can separate clean PDFs from scanned ones and still extract the right fields with minimal manual cleanup. I’ve built Python workflows for document processing, PDF parsing, OCR preprocessing, and Excel automation, and I can structure this so each file is checked intelligently, enhanced when needed, then mapped into your workbook with duplicate-safe row filling and working PDF links. I’ve shared an initial estimate based on your description, and once we go over a few technical or functional details, I’ll confirm the exact cost and delivery schedule. What are the exact extraction rules for each field when the PDF layout varies between scanned and digitally generated pages? I’ll also provide clean, documented source code and a short run guide so you can test it end to end on Windows with your sample files. Looking forward to your reply so we can finalize the exact plan. Best regards, Asad
$30 USD in 3 days
5.8
5.8

I have experience building offline Python OCR and document-processing solutions using OpenCV, Tesseract, and PDF parsing to accurately extract structured data from mixed Arabic PDFs, enhance scanned documents, and automatically populate Excel with hyperlinks while preventing duplicate entries.
$30 USD in 1 day
5.5
5.5

Hi, I can definitely help with this project. I have experience working with PDF processing, data extraction, Excel . I can build a solution that processes your PDFs, extracts the required information, updates your existing Excel file without duplicates, and provides clean, well-documented code with an easy setup guide. Question: Could you please share a few sample PDFs and the Excel file so I can review the document format before getting started? Best Regards, Farah.
$30 USD in 1 day
5.0
5.0

Hello, I've built a very similar pipeline before a clinic records system extracting structured data from mixed digital/scanned documents into a linked database. Same problem, Arabic at scale. My approach: • Digital PDFs → direct text extraction via PyMuPDF • Scanned PDFs → OpenCV pre-processing (deskew, denoise, threshold) → Tesseract OCR with Arabic pack This pre-processing step is exactly what separates this from the solutions that have already failed on your batch. The script will extract all four fields, append to your existing Excel without duplicating, and add a working PDF hyperlink per row. Fully offline on Windows. Send me your test subset. I'll process it, share the output Excel, and you verify field accuracy before we call the job done. I won't consider it complete until your test set passes cleanly. What's the rough split between digital and scanned files? Best regards, Jawad Aziz
$30 USD in 1 day
4.6
4.6

Hello, I can develop a fully offline Python solution for Windows that processes your mixed Arabic PDF collection with high-accuracy extraction. The workflow will automatically detect whether a PDF contains selectable text or requires OCR, apply image preprocessing (deskewing, denoising, contrast enhancement, etc.), extract the required fields, and append only new records to your existing Excel workbook with working PDF hyperlinks. To maximize accuracy, could you please confirm whether the four fields (رقم الكتاب، تاريخ الكتاب، موضوع الكتاب، واسم الدائرة أو الجهة المرسلة) are typically identified by consistent Arabic keywords or appear in predictable locations within the documents? I will provide well-documented Python source code, a clear execution guide, and a fully automated workflow that runs entirely offline without any cloud services or external APIs. Best regards
$15 USD in 2 days
4.8
4.8

As an automation specialist fluent in Python, I am confident that my skillset and approach align perfectly with your project requirements. At Solves Inn, my team and I have gained extensive experience in leveraging OCR technology to extract structured data from complex documents. Particularly, we've developed robust solutions that handle a range of PDF qualities accurately and efficiently, while providing integrated hyperlink features similar to what you're seeking. Your instruction on graceful handling of mixed texts resonates well with me. Over the years, we've designed numerous preprocessing pipelines for difficult-to-read documents, which include enhancing image quality, removal of noise factors, correcting skewness and contrast among others. By combining these techniques with state-of-the-art OCR approaches, we have consistently achieved excellent extraction accuracies even from low-quality Arabic texts. Moreover, backing up our technical expertise is our commitment to delivering sustainable solutions that work offline. We can assure you a well-documented, fully self-sufficient Python script that can be effortlessly run on your Windows machine without any external cloud calls or dependencies.
$20 USD in 1 day
4.3
4.3

Hi there, I am A.R.M. MASUD, with a strong Data Science background. As a Python developer, I have extensive experience building robust, scalable, and efficient solutions that address various business needs. I understand the importance of delivering high-quality, well-architected code, and I am committed to working closely with you to ensure the success of this project. I implement core functionality using Python, utilizing relevant libraries and frameworks such as Pandas, NumPy, GUI, SciPy, Matplotlib, Seaborn, Plotly, Scikit-learn, TensorFlow, Keras, PyTorch, spaCy, Flask, Django, FastAPI, OpenCV, and Jupyter. I am a professional responsible for extracting actionable insights and knowledge from large volumes of data through Machine Learning models, including CNNs, RNNs, LSTMs, GANs, Transformers, FNNs, ANNs, and DNNs. I conduct comprehensive unit, integration, and performance testing to ensure the solution is error-free and optimized. https://www.freelancer.com/u/MZITSERVICES I appreciate the opportunity to submit this proposal and am excited about the possibility of working with you to bring your project to life. Thanks A.R.M MASUD
$20 USD in 7 days
4.5
4.5

Hi there! I understand you have a large batch of Arabic PDFs and current OCR tools are failing due to mixed scanned and digital content with low-quality text. This is causing inaccurate extraction and manual workload. I have experience building Python automation scripts for PDF data extraction, OCR pipelines, and structured Excel reporting. I have worked with mixed-format documents where both selectable text and scanned images need separate handling for accurate results. My approach will process all PDFs in a single run and intelligently detect text-based vs scanned pages. For scanned pages, I will apply image enhancement techniques like noise removal, skew correction, and contrast improvement before OCR. Then I will extract the required fields (رقم الكتاب, تاريخ الكتاب, موضوع الكتاب, اسم الدائرة) using structured parsing rules and append them into Excel without duplicating existing rows. Each entry will include a clickable PDF hyperlink. The script will run fully offline on Windows using optimized local OCR processing. check our work [https://www.freelancer.com/u/ayesha86664](https://www.freelancer.com/u/ayesha86664) Do all your PDFs follow a consistent layout, or do formats vary between departments? Let me know if you’re interested & we can discuss it. Best Regards Ayesha
$20 USD in 3 days
4.0
4.0

مرحباً، تحتاج نظاماً محلياً بالكامل على ويندوز قادر على استخراج بيانات دقيقة من ملفات PDF عربية مختلطة بين نصوص رقمية وصور ممسوحة دون الاعتماد على أي خدمات سحابية. في أعمال سابقة مرتبطة بالأتمتة وإدارة البيانات، مثل تحسين تدفقات البيانات في أنظمة تجارة إلكترونية مع الحفاظ على سلامة المعلومات أثناء النقل والمعالجة، تم التعامل مع تحديات مشابهة في الدقة والتكرار. سأبني سكربت بايثون يعالج كل ملف عبر استخراج النص مباشرة عند توفره، ثم تطبيق OCR عالي الدقة مع تحسين الصور (تنقية الضوضاء، تصحيح الميل، وتحسين التباين) للملفات الممسوحة، ثم استخراج الحقول المطلوبة وإدراجها في Excel مع منع التكرار وإضافة روابط تفتح كل ملف PDF. سيكون الحل مستقراً ويعمل دفعة واحدة على جميع الملفات مع ملف تشغيل واضح يتيح لك إعادة الاستخدام بسهولة دون أي تدخل يدوي. هل جميع النماذج والكتب لديك تتبع نفس تنسيق الحقول أم أن هناك اختلافات كبيرة بين الجهات المصدرة؟ Best regards, Fizza Nadeem K
$15 USD in 4 days
3.9
3.9

Hi, As an Automation Engineer experienced in data pipelines and offline text processing, I can deliver a high-accuracy Python script to automate your 500 Arabic PDFs directly on your Windows machine. My Execution Plan: ✅ Hybrid Extraction & Image Boosting: For digital PDFs, text is pulled directly. For scans, I will implement OpenCV pre-processing (skew correction, contrast boost, noise removal) followed by an offline Arabic OCR engine. ✅ Smart Arabic Parsing & Excel Link: Custom regular expressions will isolate the 4 required fields (رقم الكتاب، تاريخ الكتاب، موضوع الكتاب، الجهة المرسلة). The script will find the first blank row in your Excel file, append the data without overwriting, and add a working local hyperlink to the PDF. ✅ 100% Offline & Secure: The solution runs entirely local on your PC with zero external API calls, keeping your administrative data fully private. I will deliver clean, well-documented source code and a simple running guide. Ready to review your sample PDFs and Excel sheet to get started! Best regards, Zakaria L.
$100 USD in 7 days
3.8
3.8

With over a decade of full stack web and mobile application development experience, I offer highly reliable and scalable solutions, which are crucial for tackling your unique Arabic-language PDF extraction project. My strong suit, namely Python, would be fundamental for this job due to its capabilities in data manipulation and automation. I specialize in a variety of areas including image processing and OCR, ensuring a smooth handling of your mixed batch of PDFs. I have taken note of your specifications for the script: extracting four data points efficiently from each PDF, offline performance without cloud calls or external APIs, as well as properly integrating those fields into the existing Excel workbook. My wide-range expertise in database management and optimization would allow me to not only achieve this but also provide you with well-structured and easily readable output. Additionally, I can further improve the script by incorporating pre-processing functionalities like noise removal, contrast boost and skew correction to give you an even higher accuracy rate.
$15 USD in 7 days
3.1
3.1

As a Certified Data Analyst and Python Automation Specialist, I possess the precise skills you're seeking for your important project. With expertise in Python, I assure you that I can create a comprehensive script that will not only excel at efficiently extracting data from Arabic-language PDFs but also handle diverse files with ease. Having worked extensively on data cleaning and formatting, I can effectively resolve the issue of mixed-quality texts by implementing advanced techniques like noise removal, contrast boost, and skew correction. Moreover, with my training in Web Scraping using Selenium and Playwright, I am well-equipped to build a completely offline Python solution for your needs. My focus on reliable data extraction and management aligns perfectly with your requirement to append data to an existing Excel workbook seamlessly. Additionally, as an SQL Database Management expert, generating working hyperlinks that open corresponding PDFs won't be a problem. Over the course of my career, I have spearheaded numerous projects requiring high-accuracy OCR such as yours with great success. Having developed thick-skinned familiarity with different PDF formats and a discerning eyes for scanning errors is crucial to delivering near-perfect field accuracy. Thus, testing the subset of real PDFs and expected output will be done meticulously to ensure flawless end-to-end functioning before marking the task complete.
$26 USD in 7 days
2.0
2.0

Hi, Drop me a message — I'll share a quick prototype based on what I understood. If it matches your expectations, we can move forward. Thanks!
$20 USD in 7 days
1.8
1.8

Hi there, I hope you’re doing well. Building a reliable offline OCR automation for Arabic PDFs requires accuracy and smart preprocessing, and I’d be happy to develop a Python solution that extracts the required fields, updates your Excel file without duplicates, and handles both searchable and scanned PDFs seamlessly. Could you please share a few sample PDFs and the Excel template so I can validate the extraction logic before development? Let's have a quick chat to discuss more in detail. I am looking forward to hearing from you. Best, Sajid.
$20 USD in 7 days
0.8
0.8

Hi, I hope you're doing well. I have carefully reviewed your project, تطوير برنامج Python لاستخراج بيانات من ملفات PDF العربية باستخدام OCR إلى Excel, and I'm confident I can deliver a high quality solution tailored to your requirements. I'm a Full Stack Developer with 5+ years of experience building websites, SaaS platforms, AI powered applications, automation tools, web scrapers, lead generation systems, and custom software. I focus on delivering reliable, high quality solutions that meet business objectives while maintaining accuracy, performance, and scalability. I'd be happy to discuss your project in more detail and recommend the best approach before we get started. I look forward to working with you. Best regards, Adnan Hussain Full Stack Developer | Technical Fixes | AI Automation | Lead Generation & Extraction Expert | Websites Dev
$10 USD in 1 day
0.0
0.0

I'll use Tesseract OCR with Python's Pytesseract library, ensuring high accuracy even for scanned images. The script will read all PDFs in one pass, extract the required data points, and output them to Excel. Delivery includes the script, a README with setup instructions, and a brief guide on pre-processing for optimal OCR results. Done in 2 days. How do you want the extracted data formatted in Excel? Note: bidding below market rate — building my Freelancer portfolio with first quality deliveries. You get full work at a discount.
$50 USD in 7 days
0.0
0.0

We recently helped a diverse range of clients achieve streamlined data extraction and automation solutions. Our expertise lies in developing customized software to enhance operational efficiency and accuracy. Noticing your emphasis on high-quality OCR and seamless integration, we are well-equipped to assist in efficiently extracting key data points from your Arabic-language PDFs into Excel. Our services encompass advanced OCR capabilities, ensuring professional and accurate results. I understand the importance of reliable data processing, as mentioned in your detailed project description. With 75+ 5-star reviews on similar projects and a top 1% ranking among 75 million users, we guarantee exceptional service quality and client satisfaction. Feel free to reach out to discuss your project further. Looking forward to the opportunity to collaborate. Regards, Hamza
$15 USD in 7 days
0.0
0.0

✅It sounds like you've put a lot of thought into this project ✅and want to get it done right. I've successfully developed similar Python programs for PDF data extraction with OCR, ensuring high accuracy even with challenging text quality. Understanding the importance of reliable OCR for Arabic PDFs, I aim to deliver a script that precisely extracts the required data points. One specific detail I'll focus on is implementing robust pre-processing techniques for scanned images to enhance OCR accuracy. With my expertise in Python, I will prioritize performance, security, and user experience for seamless offline operation on Windows. Would it be a ridiculous idea for you to send me a message and see whether we're the right fit for this project? Kind regards, Curtley
$18 USD in 7 days
0.0
0.0

Baghdad, Iraq
Member since Jun 29, 2026
₹400-750 INR / hour
₹12500-37500 INR
₹12500-37500 INR
₹12500-37500 INR
$2-8 USD / hour
₹1500-12500 INR
₹1250-2500 INR / hour
₹750-1250 INR / hour
$2-8 USD / hour
$15-25 USD / hour
₹750-1250 INR / hour
$78 USD / hour
$250-750 USD
$10-30 USD
$30-250 USD
$30-250 USD
$250-750 USD
₹12500-37500 INR
₹750-1250 INR / hour
$250-750 USD