PDF Parsing Pipeline
Built custom Python extraction scripts to clean, normalize, and validate unstructured PDF tables into a relational database schema.
Let's talk ↗
Structuring 77,000+ polytechnic admission cutoff records into a fast, mobile-first search engine for aspirants across Maharashtra.
Every admission season, over 100,000 diploma aspirants in Maharashtra face high-stakes choices. The State Common Entrance Test Cell released cutoff data scattered across massive 400+ page PDF booklets with dense tables.
Students and parents spent days scrolling through complex PDF tables on low-end smartphones trying to compare branch cutoffs, college codes, and caste category reservations. This friction resulted in missed preference deadlines and avoidable admission errors.
Built custom Python extraction scripts to clean, normalize, and validate unstructured PDF tables into a relational database schema.
Designed a 3-step filter hierarchy: Select Region → Select Branch (CS, IT, Mech, Civil) → Select Category (OPEN, OBC, SC, ST, EWS).
Engineered a lightweight frontend architecture optimized for low-bandwidth 3G networks and mid-range mobile devices.
Used plain Marathi and English labeling, clear college codes, and direct links to official DTE portal forms.
Instant search results dynamically filter by percentile cutoffs, autonomous vs non-autonomous status, and shift options without reloading the page.
Students can bookmark shortlist colleges and export preference lists to share directly with parents or cyber cafe counselors.
MahaPoly Search proves that when engineering and design work from the same brief, a complex public data archive can become a seamless, empowering tool in a matter of weeks.
— Bhavesh Patil · Founder, DigitAlly