Role
Extraction, data modeling, document processing, and search delivery
Context
Paid internal-tool development at OEG
Stack
PowerShell 5.1, JavaScript, Edge, Excel, PDF parsing
Status
Validated pipeline and accepted desktop tester workbook; tester feedback pending

A construction submittal can have several revisions, workflow responses, and attachments. Finding a phrase inside an attachment requires a different tool from finding its parent record. I built a pipeline that preserves those relationships and makes extracted document text searchable in Excel.

The environment shaped the implementation

The corporate Windows machine did not permit a new Python environment, a development SDK, or administrator-level installation. An Edge extension gathers the selected records through an authenticated Procore session. PowerShell 5.1 performs the local processing, using a pinned PDF parser and available Windows libraries for other supported formats.

Preserve first, derive second

The extractor freezes its filtered work queue before collection. Original source records and attachment files remain preserved. Normalized tables retain revisions, responses, and attachment relationships instead of flattening everything into an ambiguous document list.

File hashes establish physical identity. Extraction sidecars record each file's processing result, and relationship checks validate the derived tables. The completion manifest is published last so an interrupted run cannot appear complete.

A search surface people can use

The tester workbook uses ordinary Excel text filters and links back to Procore. It needs no macros or formulas. Extracted text is divided into bounded, overlapping chunks so long documents fit Excel's cell limit and short literal searches can cross a chunk boundary.

Explicit text typing prevents import from converting identifiers or document content into numbers and dates. The workbook is reconciled against the generated search data before acceptance.

Make incomplete extraction visible

Some PDF pages fail native extraction; others contain scanned content that needs OCR. Those gaps remain recorded. A search returning no match is not proof that the source document lacks the phrase. OCR is a separate future step.

The pipeline has passed its acceptance checks, and the desktop workbook has been accepted for tester handoff. Distribution and user feedback remain pending; browser-based Excel compatibility has not been established.

What this demonstrates

  • Practical engineering within a restricted corporate environment
  • Document processing with explicit extraction coverage
  • Relational modeling and traceability across source records and files
  • A familiar user interface backed by validated, rebuildable data