Principal Data Operations Engineer Austin, TX

Mayank Sethi.

Seventeen years building the systems that let regulated institutions prove where their data is — automatically, and at a scale no team could check by hand.

Mayank Sethi
PRTR_001 / 2026 ● REC
Currently Group1001 — Zionsville, IN
Specialization Enterprise data & database engineering
Available for Talks, advisory, papers
Local time -- CST
01 / About Engineer's note

Making the safe path the automatic one.

Located Austin, Texas Discipline Enterprise data governance · Cloud database platforms · Regulated financial services

Banks and insurers hold millions of pieces of personal information — social security numbers, policy details, account balances. Regulators require them to know exactly where all of it sits and who can reach it. Most institutions answer that question by hand, which makes the answer slow, expensive, and out of date the moment it is produced.

Mayank Sethi builds the systems that answer it automatically. At PIMCO, one of the world's largest fixed-income managers, and now at Group1001, an insurance holding company with $86.1 billion under management, he has turned manual compliance work into governed platforms that produce their own evidence.

The discipline is the same across everything he has built: use the cheapest method that can decide correctly, escalate only the genuinely hard cases to expensive technology, keep a human able to override without creating a second source of truth, and design so the thing survives being moved to another company under a different regulator. Three of his four frameworks were built at one institution and rebuilt at a second.

He is an IEEE Senior Member, a grade held by roughly ten percent of the institute's half-million members. He has published three peer-reviewed papers and serves the field as a reviewer, program committee member and awards judge.

He writes about what actually breaks at scale — and the patterns engineers use to keep it from breaking again.

0+
Years in production data systems
Fixed-income, insurance, asset management
0+
Columns under AI classification
PII / SPI governance across 62 Snowflake databases
0
Frameworks originated
Three rebuilt at a second institution under different regulation
02 / Experience Seventeen years

Inside the engine rooms of finance.

Feb 2025 → present
Group1001
$86.1B insurance holding company — Zionsville, IN

Principal Data Operations Engineer

Snowflake platform · Credential lifecycle automation · AI governance
  • Own the Snowflake platform underpinning an $86.1B insurance group — took it from manual click-ops to a governed production service, live June 2025.
  • Built the Terraform governance framework managing every Snowflake object as code — ~15,000 resources across 15 business domains, approved as organizational standard by the Architecture Review Board.
  • Designed an AI-driven PII/SPI classification system governing 1.5 million+ columns across 62 Snowflake databases.
  • Improved CIS benchmark compliance from 48% → 92% across the platform.
  • Recovered 450,000+ business attachments from Oracle Fusion backups after decommissioning left them unpreserved — ~1TB of encrypted exports, extracted via an 8-way parallel pipeline.
2014 → Feb 2025
PIMCO
One of the world's largest fixed-income investment managers — direct from Oct 2019, prior engagement via Gemini Solutions

Senior Database Engineer · from Oct 2019

Snowflake SME · Tech Lead · 3 PB of historical data · 170+ apps
  • Became the firm's designated Snowflake authority during its move off Oracle — the escalation point for a platform holding 3 PB and serving 170+ application teams.
  • Led Snowflake cost optimization across 21 app teams — $300K annually from warehouse right-sizing.
  • Engineered HVR environment refresh using zero-copy cloning, cutting refresh time from 36 hours → 45 minutes$120K/year in compute savings.
  • Built ServiceNow automation that completed 4,500+ requests — roughly 80% of all cloud database user requests — with no engineer involvement.
  • Led the firm's response to the 2024 Snowflake security incident — originated the SCALER framework and rotated credentials across 1,000+ service accounts spanning 150+ application teams, with no production outage.
  • Led Exadata consolidation, decommissioning 200+ physical machines and migrating workloads onto a consolidated platform — approximately $2M in infrastructure savings.
  • Reclaimed 190TB+ of Oracle storage in a single year through compression and reorganisation — including the firm's historical store, taken from 208TB to 96TB.
03 / Frameworks Built, deployed, published

Systems that outlived their first employer.

Tiered Classification
& Masking.

1.5 million+ columns across 62 Snowflake databases. Built solo, in nine days.

The industry's reflex answer to finding sensitive data is to point a large language model at every column. At this volume that is the expensive way to get a slow answer, which is why most organisations classify once and let the result go stale.

This engine resolves each column with the cheapest method capable of deciding correctly. Pattern matching handles the bulk, native classification catches what patterns miss, and generative AI is reserved for the genuinely ambiguous remainder. The hard problem was never detection. It was deciding what deserves the expensive tier.

Classification and enforcement are a single step: results write to tags that masking policies are already bound to, so a column is protected the moment it is identified. Data owners correct misclassifications through a self-service app, and those corrections travel the same orchestrated path as the automated run, so human and machine answers cannot diverge.

Featured — Dagster engineering blog · Published — IEEE ICOSAAS 2026
Pattern matching Native classify Cortex AI Override Tag Mask TWO INPUTS · ONE PATH

The ADCP
Control Plane.

Autonomous Database Change Platform, built on the Request–Policy–Execution–Audit model.

In a regulated firm, a request as small as read access to a table waits two to three business days. The delay is not difficulty. It is that nobody can prove the action was authorized unless a person signs off on it.

ADCP separates every request into four independent concerns, so authorization is evaluated mechanically and evidence produced automatically. The data owner still approves. The database engineer is no longer the bottleneck, and the automated path becomes more auditable than the manual one.

Built at PIMCO, then independently rebuilt at Group1001 under insurance regulation — only the policy and audit layers required reconfiguration. Over twelve months: 80% of ~6,200 requests resolved autonomously, latency cut from days to under five minutes, and zero incorrect authorizations across ~4,960 automated actions.

Published — American Journal of Technology, 2026 · Sole author
Request Policy Execution Audit ONLY THESE TWO CHANGED

The SCALER
Framework.

Scale · Classify · Assign · Lifecycle · Enforce · Rotate.

Most enterprises manage database credentials reactively — rotated after incidents, tracked in spreadsheets, governed by manual checklists.

SCALER treats credential lifecycle as a first-class engineering problem, and the ordering is the contribution. Scale → Classify → Assign → Lifecycle → Enforce → Rotate. Ownership is established before anything is touched, which is what separates remediation from an outage.

Applied across 1,000+ Snowflake service accounts spanning 150+ application teams, with automation handling 1,300+ related requests and no production outage. Derived entirely from production experience rather than vendor guidance.

Presenting — CornCon 12 Cybersecurity Conference, 2 October 2026
06 STAGES Scale Classify Assign Lifecycle Enforce Rotate
04 / Featured work Shipped to production

Frameworks and systems shipped to production.

/ 01 Production

Terraform Infrastructure Governance for Snowflake

Terraform pipelines governing every Snowflake object type as version-controlled code — databases, schemas, warehouses, roles, grants, network policies, storage integrations, masking policies, tags and tasks. Extending infrastructure-as-code past infrastructure and into the governance and data-protection layer is what makes the whole compliance posture reviewable as a diff. Built and adopted as organizational standard at PIMCO, then independently rebuilt at Group1001 and approved as standard by its Architecture Review Board.

Scale~15,000 resources · 15 domains
AdoptedOrg standard at two firms
ComplianceCIS 48% → 92%
/ 02 Migration

Oracle Estate Consolidation & Storage Reclamation

Reclaimed storage across PIMCO's Oracle estate through compression, reorganisation and decommissioning — including the firm's historical data store, taken from 208TB to 96TB. Contributed to the enterprise Data Containment programme and built the disaster recovery environment supporting PIMCO's participation in the Federal Reserve's TALF lending facility, delivered ahead of schedule.

StackOracle Exadata · ASM · ZFS
Reclaimed190TB+ in one year
Largest single DB208TB → 96TB
/ 03 Migration

Oracle Fusion Document Recovery

When Oracle Fusion was decommissioned, its attachments were not preserved — leaving 450,000+ business documents recoverable only from backups. Invoices, purchase orders, expense receipts, journal vouchers and supplier contracts, stored as BLOBs inside the WebCenter Content layer. Restored roughly a terabyte of encrypted Data Pump exports to AWS RDS, built an eight-way parallel extraction pipeline running at about five times the throughput of a single-cursor extraction, and landed everything into categorized S3 with a Snowflake metadata layer for search.

Recovered450,000+ documents
Source~1TB encrypted backups
Throughput~5x vs single-cursor
05 / Writing Selected publications

Research, frameworks, and field notes.

06 / Judging Peer review & program committees

Evaluating other people's work.

/01
Claro Awards for Tech Excellence — Judge, AI Category
Claro Awards · 2026
/02
2nd International Conference on AI and Robotics (AIR 2026) — Program Committee Member & Reviewer
Nazarbayev University · IOES · 2026
/03
IEEE International Conference on Contemporary Computing and Communications (InC4 2026) — Reviewer
IEEE Computer Society Bangalore Chapter · CHRIST University · 2026
/04
International Conference on High Performance Computing and Artificial Intelligence (HPCAI 2026) — Session Chair
Eastern Michigan Joint College of Engineering · September 2026
07 / Speaking Keynotes & talks

Sharing what actually works.

Press Coverage naming Mayank
08 / Connect Selectively open

Let's
connect.

Selectively open to conversations about enterprise data architecture, speaking invitations, and advisory discussions in financial data infrastructure.

If you're building something at scale — and hitting hard problems — let's talk.