Infrastructure Architecture & Systems Design
Cloud, hybrid, bare metal, or self-hosted — systems designed around your actual workflows, with capacity, cost, and failure modes decided on paper before they cost you in production.
Salt Lake City, remote across North America. Evenings, Mountain Time. That’s still afternoon on the West Coast.
Every engagement is scoped around outcomes — faster delivery, fewer incidents, lower spend, and systems your team can actually reason about.
Cloud, hybrid, bare metal, or self-hosted — systems designed around your actual workflows, with capacity, cost, and failure modes decided on paper before they cost you in production.
Cluster design, provisioning, upgrades, and day-2 operations for container platforms — including Nomad cluster hosting — run calmly at any scale.
Bitbucket, GitHub Actions runners, N8N, Mage AI, Proxmox, and the rest of your internal platform — deployed, secured, kept current, and off your team's plate.
Pipelines and internal platforms that collapse the distance between commit and production — safely, repeatably, and without tribal knowledge.
Troubleshooting, production support, and post-incident hardening. Incident response is a Watch Window, not included in every engagement. Recovery is the floor — the goal is making the same failure impossible twice.
Rigorous usage analysis, rightsizing, and capacity strategy. Most platforms carry meaningful waste — we find it without sacrificing performance.
Accurate maps of how your systems actually connect — dependencies, bottlenecks, and single points of failure made visible, then actionable.
AI applied where it earns its keep — faster troubleshooting, automated runbooks, and internal tooling that reduces operational toil. A tool in the kit, not the product.
Source, delivery, orchestration, and state — one observable path from commit to production. Internal platforms run beside your applications, not in their way.
One system, four movements — designed, connected, tested by failure, and strengthened by it. This is the operating loop behind every engagement.
Architecture first. Every system begins as a precise drawing — boundaries, dependencies, and failure domains decided before a single server exists.
Review the current systems, workflows, constraints, and goals — as they actually are, not as the wiki says.
Diagram dependencies, surface bottlenecks, and agree on a target architecture everyone can point at.
Implement infrastructure, automation, and platform improvements in small, reversible steps.
Harden reliability, observability, cost efficiency, and documentation for the long run.
Real architecture and data work — anonymized, with client names and identifying details removed. Each diagram is the kind of system we design, migrate, and run.
A full DR architecture for a financial institution: every workload re-platformed to AWS with Application Migration Service, firewall-inspected transit networking, replication staging, and phased cutover across two regions — with a component-level cost model to back the plan.
A person-centered core banking platform wired to an event bus — decoupled applications consuming domain events instead of point-to-point calls, with a data lake feeding analytics and GenAI while backup and business continuity stay built in.
Spark's execution model mapped end-to-end — resilient datasets, DAG scheduling, cluster managers — and how EMR topologies connect to S3, Redshift, and DynamoDB for batch and streaming workloads that stay cheap to run.
An end-to-end AI pipeline — raw sources collected and cleaned into golden data, LLM insight generation, automated proposal agents, and a model-evaluation loop with experiment tracking — from unstructured inputs to a deployed demo.
Client names and identifying details removed — the work, not the logo.
Published floors, scoped after Discovery — every engagement is set together.
Scoped triage · one-time
16 hours · set evening window
Fixed scope · Saturday delivery
On-call · defined window
standalone $1,800 / month
Start with Discovery. From there: a Build Project, a Platform Window, or both. Watch Window stacks on a Platform Window — or stands alone after Discovery or a Build Project.
Set windows in US Mountain Time. Async is the default; live work is the exception.
Mon–Thu 16:30–19:30 MT · Fri 16:30–18:00 MT
Live pairing and same-window work Monday–Thursday. Friday 16:30–18:00 MT is for the weekly status artifact only. 16:30 MT = 15:30 PT.
Mon–Thu 19:30–22:00 MT · Sat 09:00–13:00 MT
Defined-window on-call — never 24/7, no uptime SLA. In-window P1 ack ≤15 min. Nights after 22:00 MT and weekday daytime are out of window.
Written asks are the default. Live work is the exception, booked inside a published window.
Weekly Friday Loom or written equivalent, plus the hour ledger and next week’s queue.
Inquiry replies within 24 hours. In-window Platform Window ack ≤30 min.
Work waits for the next window, unless a Watch emergency at $225/hr with a 2-hour minimum (founder discretion).
Salt Lake City Saturdays by arrangement when the work needs a room — not a 15-minute on-site response.
Straight answers on scope, pricing, and how the work runs — the things teams usually ask before a first call.
Engagements start at $500 for Discovery (scoped triage), $3,500 per month for a 16-hour Platform Window, and $8,500 for a fixed-scope Build Project. Watch Window is $1,200 per month as an add-on to a Platform Window, or $1,800 per month standalone. Those are published floors — final scope is set together after Discovery. Platform Window is the ongoing operations engagement (formerly listed as Retainer).
Discovery is a scoped triage: 3–4 hours one-time, a 45-minute kickoff, and written findings within 7 days. It includes an infrastructure audit, an architecture map with gap analysis, a cost and usage review, and a written next-step — project, Platform Window, or not a fit.
BP Data is based in Salt Lake City, Utah, and works remote-first with teams across North America. On-site sessions in Salt Lake City are available Saturdays by arrangement when the work genuinely needs a room. This is not a 15-minute on-site response.
Set windows are in US Mountain Time. Platform Window: Monday–Thursday 16:30–19:30 MT, plus Friday 16:30–18:00 MT for the weekly status artifact. Watch Window: Monday–Thursday 19:30–22:00 MT and Saturday 09:00–13:00 MT. Async is the default; live work is the exception. The Friday artifact is a Loom or written equivalent, plus the hour ledger and next week's queue.
No. Watch Window is the incident-response product. A Platform Window is 16 hours of scheduled ops in a set evening window — it does not include on-call. Watch is never 24/7 and carries no uptime SLA.
Kubernetes fits broad ecosystem and hiring gravity; Nomad often fits mixed workloads and simpler day-2 operations. BP Data designs and operates either — and will say when a managed service is enough.
Cloud, hybrid, and bare metal — with deep focus on Kubernetes and Nomad, CI/CD and developer platforms, and self-hosted services such as Bitbucket, GitHub Actions runners, N8N, Mage AI, and Proxmox.
When workloads are steady, data control matters, or SaaS and egress costs outpace operations, self-hosting can be the calmer path. When elasticity dominates, cloud stays the right default — the tradeoff gets modeled on paper before anything moves.
Teams running self-hosted or hybrid platforms that have outgrown managed services but do not need — or cannot yet staff — a full platform organization. Typical fits include growing product companies and internal platform owners who need senior ops on Bitbucket, runners, N8N, Mage AI, Proxmox, Nomad, or Kubernetes.
Replies typically land within 24 hours, by email or Slack once a direct channel is open. No commitment is required to start a conversation.
Infrastructure needs, platform improvements, cost review, migrations, reliability — bring the problem. Start with Discovery if you are unsure which engagement fits. No commitment required, and replies typically land within 24 hours.
BP Data is the independent practice of Brandon Phillips — infrastructure & platform engineering, Salt Lake City.
Typically within 24 hours. In-window Platform Window ack ≤30 min.
Salt Lake City, Utah · Remote — North America · On-site Saturdays by arrangement