Resume
Kris Kokomoor
kris.kokomoor@gmail.com Β· pysynapse.com Β· New London, CT
Principal Data, Platform & Systems Engineer
Cloud Architecture • Clinical Data & Imaging Platforms • Data Quality Engineering • Linux Systems • AI-Augmented Engineering • Regulated Environments
Engagement Models
Permanent leadership (Principal/Director), fractional head of data, and strategic consulting for regulated industries.
Typical Problems Solved
High-stakes reporting trust, clinical data & imaging platform reliability, audit-ready cloud architecture, and agentic pipeline management.
Common Stacks
Python, C++, SQL, Airflow, dbt, BigQuery, Kafka, DICOM/DCMTK, Terraform, Pulumi, AWS, GCP, Docker, Prometheus.
Professional Summary
Engineering leader and hands-on builder specializing in cloud-native data platforms, large-scale clinical and imaging ingestion, and quality-driven system design. Fourteen years managing regulated clinical data and imaging at pharmaceutical scale (1.5 PB image archive, global trial platforms), combined with modern cloud and data tooling (GCP, BigQuery, Airflow, dbt) and deep systems-level expertise in Linux, infrastructure, and performance engineering.
Track record of collapsing operational cycle times through automation: onboarding of new data sources reduced from one week to under 15 minutes, and clinical study image QC from 30 days to 5.
Currently developing formal methods for diagnostic evidence capture in autonomous data operations, and high-performance DICOM metadata tooling for large clinical image sets.
Core Skills
- Data Platforms: BigQuery, PostgreSQL, dbt, Airflow (2.x/3.x), Kafka, Snowflake, PySpark, PACS, pub/sub systems
- Cloud & Infrastructure: GCP, AWS, Cloud Functions (Gen2), Cloud Run, Cloud Storage, S3, Glacier, Terraform, Pulumi, Docker, Docker Compose
- Clinical & Imaging Standards: DICOM, DCMTK, HL7, FHIR, CDISC-adjacent workflows, clinical data anonymization, CTMS/CDMS integration
- Systems Engineering: Linux/UNIX, Solaris, networking, process & memory diagnostics, performance tuning, distributed systems, data center design
- AI & Agentic Systems: LLM integration, prompt engineering, API-driven workflows, local agent frameworks (OpenClaw), structured output pipelines, LLM-assisted diagnostics
- Quality & Reliability: Validation frameworks, schema drift detection, observability, alerting, reconciliation, validated system lifecycle (GxP)
- Languages & Tools: Python, SQL, C++, Bash, Git, REST APIs, FastAPI, SQLAlchemy, Pydantic, JSON workflows, Prometheus, Grafana
Experience
Principal Consultant & Independent Researcher β pySynapse / SURL (2025 β present)
Independent consulting and applied research at the intersection of data platform reliability, regulated data management, and AI-assisted engineering.
- Researching Evidence Packets, a source-independent framework for the diagnostic context required to identify and resolve data pipeline incidents; closed Phase I of the governed experimental program (representation utility and evidence durability), with 2 research papers published to date of a planned 5-part series and a public reference implementation (MVP) on GitHub.
- Formalizing the boundary between deterministic and probabilistic methods in incident identification, targeting controlled LLM use in regulated data operations.
- Building fastDICOMattrs β a C++/DCMTK library with Python bindings delivering native-speed root tag reads from DICOM files, addressing metadata query performance across large clinical image sets (public repo).
- Building fastDICOMarchive β a reference application using fastDICOMattrs to filter and persist images received via file, stream, and other transport modes (public repo).
Podimetrics β Principal Quality Engineer, Data & Platform Engineering (2024 β 2025)
Designed and implemented a quality-first data platform supporting patient-facing and internal analytics products across GCP.
- Reduced time-to-ingest for new data sources from one week of manual coding to under 15 minutes (polling interval) by combining generalized Airflow orchestration with a reusable Gen2 Cloud Function ingest layer β eliminating bespoke development as a precondition for onboarding a feed.
- Architected and deployed a production-grade Airflow 3.x environment on GCP (VM-based), provisioned end-to-end via Pulumi with a PostgreSQL metadata backend β full server, application, and supporting environment build.
- Authored and operated 40+ DAGs orchestrating batch ingest from 8 external medical and pharmaceutical claims providers on daily and weekly SFTP cadences, plus manufacturing data ingest.
- Designed, built, and maintained 200+ dbt models in a hub-and-spoke architecture serving BI consumers, tightly integrated with Airflow orchestration for governed, repeatable analytics pipelines.
- Designed, authored, and supported a suite of 12 Gen2 Cloud Functions with an associated CI/CD environment, managing ingest from Google Cloud Storage, AWS S3, and Google Drive.
- Defined a data validation taxonomy spanning schema integrity, correctness, freshness, reconciliation, and referential consistency.
- Developed an early-warning reliability framework combining Airflow, dbt tests, BigQuery metrics, and Slack alerting; led schema drift detection and impact analysis for high-volume external feeds.
- Explored LLM-assisted validation and anomaly detection to accelerate triage and improve signal detection in pipeline monitoring.
Hands-on throughout: system design, implementation, debugging, and operational support.
Founder / Principal Engineer β SURL (Secure URL) (2023 β 2024) | AWS
Designed and built a cloud-native platform for secure, auditable transfer of regulated datasets.
- Architected an event-driven system using Kafka for request handling, authorization, auditing, and lifecycle management.
- Designed a PostgreSQL-centered control plane governing metadata, access rules, retention policies, and state transitions.
- Built containerized Python services (FastAPI, SQLAlchemy, Pydantic) supporting full lifecycle APIs.
- Implemented quarantine β authorized object store workflows enforcing compliance and governance constraints.
- Deployed on AWS (EC2, S3) with Terraform-managed infrastructure; instrumented with Prometheus and Grafana across services and brokers.
- Designed explicit tracking of cloud ingress/egress costs for operational transparency.
- Built a BrightScript / Roku client application as a cross-fabric proof of concept.
- Provided consulting services to a medical imaging startup on secure regulated data transfer.
Owned full lifecycle: architecture, development, deployment, monitoring, and debugging.
Pfizer β Associate Director, Clinical Image Management (2015 β 2023) | Product Development
Led architecture and development of large-scale clinical data and imaging platforms supporting global trials.
- Cut clinical study image QC from 30 days to 5 β an 83% reduction in study QC cycle time β by designing and implementing automated image quality control across the archive.
- Managed a clinical image archive exceeding 1.5 petabytes, sustaining an average annual ingest of 150 TB of clinical imaging data.
- Led migration of on-premise image metadata management to AWS, augmented with demographic and diagnostic data.
- Built search and indexing systems for DICOM-based image archives, enabling clinician access to imaging and trial data.
- Established enterprise-scale strategies for image metadata processing, tiered archival, and governed access (Python, AWS, Glacier).
- Developed the Unified Clinical Data Hub (Snowflake), integrating imaging metadata, CTMS milestones, and CDMS extracts β enabling cross-domain analytics linking clinical outcomes, imaging biomarkers, and trial timelines.
- Owned first enterprise purchase and implementation of d-Wise Blur (2015) for clinical data anonymization; authored the governing procedure for release of clinical images with associated clinical and demographic data.
- Led clinical specimen management systems and supported regulatory reporting and patient safety analytics with near real-time data availability.
- Designed scalable PySpark/Spark pipelines for transformation and enrichment of clinical and imaging datasets.
Pfizer β Senior Engineering & Data Leadership Roles (2009 β 2015)
- Translated Python-based analytical workflows into PySpark pipelines optimized for scale and performance.
- Led modernization of global clinical and analytical platforms supporting regulated research environments.
- Directed implementation of data anonymization platforms, balancing regulatory requirements against technical scalability.
- Designed and supported large-scale data pipelines across heterogeneous systems and formats.
- Supported validation of analytical systems, including Monte Carloβbased verification approaches.
- Team leader for Linux/UNIX support of clinical development, with 125+ Sun Solaris servers under management.
- Partnered with quality and compliance stakeholders on validated system lifecycle management.
Earlier β Systems, Linux & Infrastructure Engineering
Built deep expertise in Linux/UNIX systems, infrastructure, and performance engineering across healthcare, federal, and enterprise environments.
- Medaphis: designed and built six corporate raised-floor data centers from bootstrap, consolidating East Coast medical billing operations.
- Designed and operated large-scale Linux/UNIX environments supporting mission-critical workloads.
- Managed bare-metal infrastructure β compute, storage, and networking β in high-availability systems.
- Developed strong Linux operational expertise: process management, memory and I/O diagnostics, system tuning, and failure analysis.
- Led deployment of distributed processing systems prior to modern cloud paradigms.
- Established enduring expertise in systems-level debugging and performance optimization, now applied to cloud-native platforms and data pipelines.
Education
- B.S. Mathematics, University of Florida
- M.S. Electrical Engineering, University of Florida
- M.B.A., University of Rhode Island