
Closed
Posted
Paid on delivery
I want to build a dataset for training a medical-assistant bot, so I need Image data that appears inside publicly available hospital reports across the web. The task is to locate those reports, extract every embedded image, and hand the files over to me in readable format I’m open to any solid, well-documented approach—focused web crawling, API use, or site-specific scraping—as long as it scales, respects each site’s terms of service, Please tell me which tools you prefer Deliverables 1. Working scraper or repeatable method with a brief setup guide. 2. First sample batch of at least 1,000 JPEG images for validation. 3. Short report describing sources, filtering logic Acceptance criteria • Only Image data is collected; no text extraction needed. • No protected or copyrighted material is included. • Method can be re-run by me without modification. If anything is unclear, let me know so we can sort it out quickly and move forward.
Project ID: 40530614
24 proposals
Remote project
Active 4 days ago
Set your budget and timeframe
Get paid for your work
Outline your proposal
It's free to sign up and bid on jobs
24 freelancers are bidding on average ₹1,104 INR for this job

Have over 18 years of experience in data mining/ Web scrapping/ Scraping Bots/ Chrome/Opera Extensions I have done it all. Tell us your source and we will put it in excel for you, Or we can even give you filtered results as per your requirement, In the format you want. You can also ask for data into a particular format - Excel, Json, Mysql, Databases, XMLs, you name them. Further Can help you with integrating it with ur databases, Can create json outputs. We are not only good with scraping but also with the tools that u may need after that. We can help you build you softwares round the data we have 99% Data Accuracy. We have Duplicate finder. etc., We can help with Statistics on the data We can help with creating Api's front the data We can create Softwares to manage that data We can build Sites round the data
₹1,050 INR in 7 days
6.9
6.9

Hi, I can build a repeatable Python and php scraper to collect publicly available medical-report images, deliver 1,000+ sample images, provide source documentation, filtering logic, and a complete setup guide for re-running the process. Ready to start immediately.
₹1,200 INR in 1 day
5.0
5.0

Hi, I can help build a scalable, repeatable solution for collecting images from publicly available hospital reports while respecting website terms and licensing requirements. My preferred stack is **Python** using **Scrapy/BeautifulSoup**, **Playwright** (where JavaScript rendering is needed), and **Pillow/OpenCV** for image validation and filtering. I'll provide a reusable scraper, a setup guide, a sample batch of 1,000+ images, and a report documenting data sources and the collection process. Before we begin, I'd like to clarify one point: should the dataset include **only images that are explicitly licensed for reuse (e.g., Creative Commons or public domain)**, or do you have a predefined list of approved sources? I can start immediately and ensure the solution is well-documented and easy to run.
₹600 INR in 7 days
3.3
3.3

With my expertise in AI, cloud data engineering, and real-time analytics, I'm ready to provide a seamless solution for your medical-assistant bot training dataset. My skillsets align perfectly with the requirements of this project, as I have an extensive background in building scalable systems that can collect massive datasets like you need here. Plus, my experience working in healthcare and enterprise environments means I understand the sensitive nature of medical data and always prioritize compliance and data security. I understand you’re open to any efficient approach, and I propose using Python, combined with cloud-based technologies like AWS Lambda and Azure Data Bricks. These tools allow for large-scale web crawling, adhering strictly to site-specific terms of service while ensuring high performance. My algorithms would filter out copyrighted or protected materials while collecting only image data to meet your acceptance criteria. Lastly, as a business-oriented professional, I am focused not just on implementations but on meaningful outcomes for my clients. By working with me, you can expect not just a scraper tool but a complete solution that comprises a step-by-step guide for setup, a considerable sample batch (minimum 1k JPEG images) for validation along with a comprehensive report that documents the sources and filtering logic used. Let's transform your vision into a reality by leveraging intelligent systems and gaining invaluable insights from sizable image data!
₹2,000 INR in 1 day
2.7
2.7

Hello, I understand you need a scalable and repeatable system to collect publicly available medical report images from hospital-related web sources, build a dataset of at least 1,000 valid JPEG images, and provide a documented scraping pipeline for reuse. The focus is strictly on image extraction with proper filtering and compliance with website policies. Here’s what I can provide: Re-runnable dataset generation script with logging, source tracking, and structured output in clean JPEG format. I bring over 4+ years of experience in Python, web scraping, data mining, and image processing pipelines. I have built scalable crawlers and dataset generation tools with a strong focus on clean data extraction, automation, and reproducibility. Just to clarify a few things: • Do you have a preferred list of hospital/report sources, or should the scraper discover them automatically using search-based crawling? • Should the dataset include any labeling/metadata (e.g., report type/source), or only raw filtered images? Please come to the chat box to discuss more about your project. Best regards Indresh Kushwaha
₹1,050 INR in 7 days
1.7
1.7

I noticed the key requirement isn’t just collecting images—it’s building a repeatable, compliant pipeline that can be re-run without modification while avoiding protected or copyrighted content. I’ve handled large-scale web data collection projects and would approach this by identifying publicly accessible hospital report repositories, applying source-level filtering, extracting only embedded image assets, validating file integrity, and exporting everything in a structured JPEG dataset. The workflow will include source attribution, duplicate detection, metadata logging, and compliance checks to ensure every image can be traced back to its origin and reviewed. For delivery, you’ll receive the scraper, setup guide, first 1,000-image validation batch, and a concise methodology report covering sources, filtering rules, and verification steps. My preferred stack is Python (Scrapy/Requests/BeautifulSoup), with site-specific handling where needed for reliability and scalability. Do you already have preferred source domains in mind, or should I begin by building a vetted source list as part of the process?
₹1,500 INR in 1 day
1.4
1.4

Hi, I can collect and organize the required medical report image data with accuracy and proper file management. I’ll ensure the information is well-structured, clearly labeled, and delivered according to your requirements with attention to detail. Ready to start immediately.
₹1,000 INR in 1 day
0.7
0.7

I went through your project details for scraping medical report images and am ready to start immediately. I will develop a scalable web crawler using Python (BeautifulSoup and Scrapy) to systematically locate public hospital reports across the web and extract all embedded image data without capturing text. I will provide a working, repeatable script with a clear setup guide, a validation sample of 1,000 clean JPEG images, and a brief report outlining the filtering logic to ensure no protected or copyrighted material is included
₹1,250 INR in 2 days
0.0
0.0

Thank you for the detailed description. I can build a repeatable Python-based pipeline to locate publicly available medical reports, extract embedded images, convert them to JPEG format, and generate a source report documenting where each image originated. Before proceeding, I would like to clarify one point regarding licensing. The requirement states that no protected or copyrighted material should be included. Many hospital reports are publicly accessible but still copyrighted. Would you like the dataset to be limited to reports and publications that are explicitly released under open licenses (for example CC-BY, CC0, or other reuse-permitted licenses), or is publicly accessible content acceptable? My preferred approach would be to use open-access medical repositories and publications with clearly documented licenses, then automate image extraction and filtering using Python tools such as Requests, BeautifulSoup, PyMuPDF, and Pillow. This would provide a repeatable workflow and reduce legal risk. Please let me know your licensing requirements and preferred image types, and I can propose the exact collection strategy. Zhai Kun.
₹1,200 INR in 2 days
0.0
0.0

I am create the your all work please send me your work details I am complete your work and send you when you see my work yo very happy
₹1,050 INR in 7 days
0.0
0.0

ฉันมีความรู้ความเข้าใจในด้านเทคโนโลยีอย่างยิ่งและฉันเป็นคนที่ให้คำแนะนำได้ดีนอกจากฉันเรียนในสาขาจิตวิทยาการศึกษาและการแนะแนวมีความรู้ความเข้าใจในด้านภาษาอังกฤษสื่อสารเก่ง
₹1,050 INR in 7 days
0.0
0.0

I have experience in Python web scraping and data collection. I can build a robust script to systematically extract medical report images from hospital portals. Delivered in clean, organized format with proper documentation.
₹1,000 INR in 7 days
0.0
0.0

Hi, I can complete this within 24 hours. I have experience building Python automation and web-scraping pipelines, and my approach here would be to create a repeatable workflow rather than a one-time extraction script. Before starting, I'd like to clarify one point: should the dataset include only images from reports that explicitly allow reuse (open-access/public-domain sources), or any publicly accessible hospital reports? My plan is to: - Discover and download relevant public hospital reports/PDFs. - Extract all embedded images directly from the reports. - Convert outputs to JPEG where needed. - Remove duplicates and filter out logos, icons, and other non-relevant graphics. - Organize everything in a clean, structured format. Tools: Python, Requests, BeautifulSoup, PyMuPDF, and Pillow. Deliverables: - A working scraper/repeatable extraction method with setup instructions. - An initial batch of 1,000+ JPEG images for validation. - A short report covering sources used, filtering logic, and how to rerun the process. The workflow will be documented so you can run it yourself later without modification. Let me know your preference regarding source licensing, and I can get started right away.
₹1,000 INR in 1 day
0.0
0.0

Hello, I can help you build a reliable Python-based solution to collect publicly available hospital report images and deliver them in a structured format. My approach would include: ✅ Python web scraping using BeautifulSoup, Requests, and Selenium (if required) ✅ API-based collection where available ✅ Image extraction and filtering logic ✅ Organized storage of JPEG images with metadata ✅ Re-runnable and well-documented code ✅ Setup guide and usage instructions I will ensure the scraper is modular, easy to maintain, and follows the target websites' terms of service and technical limitations. Deliverables: Working Python scraper First batch of images for validation Documentation and setup instructions Source list and filtering methodology I can start immediately and provide regular progress updates throughout the project. Best Regards, Manushri Raval
₹1,050 INR in 7 days
0.0
0.0

Hello, I am a Full Stack Developer and AI Engineer with experience in web scraping, data extraction, OCR systems, and large-scale dataset generation. I can develop a repeatable and well-documented pipeline to identify publicly available hospital reports, extract embedded images, filter the results, and organize them into a structured dataset suitable for validation and downstream AI training workflows. My preferred stack includes Python, Scrapy, Requests, BeautifulSoup, Selenium (when required), PDF processing libraries such as PyMuPDF/pdfplumber, and cloud storage solutions for scalable data collection. The pipeline will include automated source discovery, image extraction, duplicate detection, metadata tracking, and export to standard formats (JPEG/PNG) with a clear folder structure. For the initial milestone, I can deliver a working scraper, setup documentation, a sample batch of 1,000+ extracted images, and a concise report detailing source locations, extraction methodology, filtering criteria, and dataset statistics. The solution will be designed for easy re-execution without code changes and can be extended to support additional sources as requirements evolve. Before starting, I would like to confirm the target jurisdictions and approved source types to ensure compliance with licensing requirements and usage restrictions. I look forward to discussing the project further. Best regards, Atharva Naik
₹600 INR in 7 days
0.0
0.0

Hello I have 3 years experience in web scraping I can do image and data collection and I can download the images also
₹1,050 INR in 7 days
0.0
0.0

Hi, I can handle the crawling and image extraction workflow using Python-based scraping and PDF/image extraction tools. Before proceeding, I would like to clarify the requirement regarding "No protected or copyrighted material is included." Could you please confirm which source types are acceptable for image collection? Public domain sources Government hospital reports Open-access medical repositories Creative Commons licensed content Also, do you want me to verify and document the license/source information for each collected image batch? This will help ensure the dataset fully meets your acceptance criteria.
₹1,050 INR in 7 days
0.0
0.0

Delhi, India
Member since Jun 15, 2026
£250-750 GBP
$30-250 USD
₹12500-37500 INR
$1500-3000 USD
₹12500-37500 INR
$20-36 USD
$1500-3000 USD
$15-25 USD / hour
$30-250 USD
$2-8 AUD / hour
$10-30 USD
$750-1500 USD
₹600-1500 INR
₹1500-12500 INR
$30-250 CAD
$10-30 USD
$250-750 USD
$5000-10000 USD
₹600-1500 INR
$30-250 USD