Context
Problem & Responsibility
Operational Problem
Manual customer-file processing was repetitive, time intensive and difficult to scale consistently.
My Responsibility
Designed and implemented automated intake, validation, transformation, spatial assignment and quality-control workflows.
System Design
Architecture & Workflow
- 01
Email-based file intake and Cloud Storage
- 02
Pub/Sub and Cloud Run processing
- 03
PostgreSQL/PostGIS and BigQuery workflows
- 04
Automated spatial assignment and quality control
Deep Dive
Inside the Build
Approach
Designed and implemented automated intake, validation, transformation, spatial assignment and quality-control workflows: email/attachment ingestion, customer/sample identification, cloud movement, spatial-layer sample association, received-versus-missing validation, spatial join/transformation and loading of clean, spatially usable soil results into BigQuery.
- 01
Designed sender/lab classification so files route through the correct cleaning logic.
- 02
Automated attachment capture and cloud storage instead of manual save-and-organize steps.
- 03
Linked laboratory values to the correct field/sample context using geospatial relationships and existing sample identifiers.
- 04
Standardized incoming chemistry data for reliable historical analytics and decision-support use.
- 05
Delivered analysis-ready data to BigQuery and downstream analytical frontends.
Operational Architecture
From Email Intake to Analytics-Ready Soil Data
The production workflow separates event capture, secure staging, validation and delivery so every step can be monitored, recovered and scaled independently.
- 01
Source
Email Intake
Soil-sample CSV files and their manifests arrive through a monitored inbox.
- 02
Ingestion
Event & Fetch
Inbox events travel through Pub/Sub and trigger a Cloud Run service to retrieve new files.
- 03
Buffer
Secure Staging
Files, manifests and processing logs are staged in Cloud Storage before transformation.
- 04
Processing
Validate & Assign
A second service validates schemas, deduplicates rows, assigns samples to management zones and transforms each record.
- 05
Delivery
Publish & Notify
Analytics-ready rows are loaded into BigQuery and processing status is sent to the operations team.
Shared Controls
Cloud Scheduler
Keeps inbox event registration active.
Secret Manager
Provides service credentials securely at runtime.
Observability & Recovery
Preserves processing logs and supports traceable reruns.
Interactive Diagrams
Explore the Full Architecture
Three views of the same pipeline: the end-to-end system map, the processing flow with its decision points and failure handling, and the QA/QC validation gates. Drag to pan, use Ctrl or ⌘ with scroll to zoom, or expand the viewer to full window.
Architecture adapted from the operational pipeline. Environment-specific resource names and credentials are intentionally omitted.
Evidence
Scale & Measurable Impact
- Reduced processing from approximately eight hours per customer to approximately 3–8 minutes
- Processed approximately 80,000 soil samples per season
- Covered approximately 300,000 acres
