Infrastructure & Platform Engineering

I do self-hosted
and hybrid platform work.

Salt Lake City, remote across North America. Evenings, Mountain Time. That’s still afternoon on the West Coast.

Bitbucket GitHub Actions runners N8N Mage AI Proxmox Nomad Kubernetes
Scroll

What We Design, Build & Run

Every engagement is scoped around outcomes — faster delivery, fewer incidents, lower spend, and systems your team can actually reason about.

Architecture map on every engagement Runbooks & documentation as standard Direct channel — Slack or email Replies within 24 hours
01

Infrastructure Architecture & Systems Design

Cloud, hybrid, bare metal, or self-hosted — systems designed around your actual workflows, with capacity, cost, and failure modes decided on paper before they cost you in production.

02

Kubernetes & Nomad Operations

Cluster design, provisioning, upgrades, and day-2 operations for container platforms — including Nomad cluster hosting — run calmly at any scale.

03

Self-Hosted Platform Services

Bitbucket, GitHub Actions runners, N8N, Mage AI, Proxmox, and the rest of your internal platform — deployed, secured, kept current, and off your team's plate.

04

CI/CD & Developer Platforms

Pipelines and internal platforms that collapse the distance between commit and production — safely, repeatably, and without tribal knowledge.

05

Reliability & Incident Response

Troubleshooting, production support, and post-incident hardening. Incident response is a Watch Window, not included in every engagement. Recovery is the floor — the goal is making the same failure impossible twice.

06

Cloud Cost & Usage Optimization

Rigorous usage analysis, rightsizing, and capacity strategy. Most platforms carry meaningful waste — we find it without sacrificing performance.

07

Architecture Diagramming & System Mapping

Accurate maps of how your systems actually connect — dependencies, bottlenecks, and single points of failure made visible, then actionable.

08

AI-Assisted Operations

AI applied where it earns its keep — faster troubleshooting, automated runbooks, and internal tooling that reduces operational toil. A tool in the kit, not the product.

One Connected System

← scroll to explore →
GIT BITBUCKET · SELF-HOSTED CI / CD ACTIONS RUNNERS REGISTRY ARTIFACTS ORCHESTRATION — K8S / NOMAD APP SERVICES INTERNAL PLATFORMS N8N MAGE AI STATE DB MQ KV INGRESS PROD OBSERVABILITY · COST · RELIABILITY — ACROSS THE ENTIRE PATH

Source, delivery, orchestration, and state — one observable path from commit to production. Internal platforms run beside your applications, not in their way.

How We Operate

One system, four movements — designed, connected, tested by failure, and strengthened by it. This is the operating loop behind every engagement.

← scroll to explore →
Animated schematic: a delivery platform is drawn, code flows from source through CI into an orchestrated cluster and on to production, a workload fails and traffic shifts to healthy capacity, then the workload is restored and standby capacity is added. GIT SELF-HOSTED CI / CD RUNNERS INGRESS PROD ORCHESTRATION — K8S / NOMAD APP SERVICES + INTERNAL N8N MAGE STATE DB CACHE QUEUE

Architecture first. Every system begins as a precise drawing — boundaries, dependencies, and failure domains decided before a single server exists.

How Engagements Run

01

Understand

Review the current systems, workflows, constraints, and goals — as they actually are, not as the wiki says.

02

Map

Diagram dependencies, surface bottlenecks, and agree on a target architecture everyone can point at.

03

Deliver

Implement infrastructure, automation, and platform improvements in small, reversible steps.

04

Strengthen

Harden reliability, observability, cost efficiency, and documentation for the long run.

Systems We've Actually Designed

Real architecture and data work — anonymized, with client names and identifying details removed. Each diagram is the kind of system we design, migrate, and run.

Custom hand-drawn disaster recovery architecture: on-premise datacenters connect to two AWS regions via VPN and Direct Connect, with transit gateway routing, firewall inspection, MGN replication, staging and cutover, plus sizing and TCO panel
Infrastructure · Disaster Recovery

Disaster Recovery Architecture

A full DR architecture for a financial institution: every workload re-platformed to AWS with Application Migration Service, firewall-inspected transit networking, replication staging, and phased cutover across two regions — with a component-level cost model to back the plan.

Custom hand-drawn event-driven core platform: core platform emits domain events to an event bus consumed by decoupled apps, with a data lake feeding analytics and GenAI plus backup and business continuity
Banking · Event-Driven Architecture

Event-Driven Core Platform

A person-centered core banking platform wired to an event bus — decoupled applications consuming domain events instead of point-to-point calls, with a data lake feeding analytics and GenAI while backup and business continuity stay built in.

Custom hand-drawn distributed data processing diagram: Spark driver builds a plan through RDDs and DAG scheduling, cluster manager provisions workers, plus EMR master-core-task topology and storage connectors
Data · Distributed Compute

Distributed Data Processing

Spark's execution model mapped end-to-end — resilient datasets, DAG scheduling, cluster managers — and how EMR topologies connect to S3, Redshift, and DynamoDB for batch and streaming workloads that stay cheap to run.

Custom hand-drawn AI pipeline: raw data scraped and cleaned into golden data, LLM insight generation, proposal agents, a working demo, and an evaluation loop with model experiments
AI · LLM Systems

AI Service Pipeline

An end-to-end AI pipeline — raw sources collected and cleaned into golden data, LLM insight generation, automated proposal agents, and a model-evaluation loop with experiment tracking — from unstructured inputs to a deployed demo.

Client names and identifying details removed — the work, not the logo.

Engagement Models

Published floors, scoped after Discovery — every engagement is set together.

Discovery

Scoped triage · one-time

from $500
  • 3–4 hours, one-time
  • 45-min kickoff
  • Written findings within 7 days
  • Audit, architecture map, cost review, written next-step
Start with Discovery

Build Project

Fixed scope · Saturday delivery

from $8,500
  • Fixed-scope builds & migrations
  • Saturday delivery day
  • Documentation & runbooks
  • 30-day post-delivery support
  • One project at a time
Scope a Build Project

Watch Window

On-call · defined window

add-on $1,200 / month

standalone $1,800 / month

  • Mon–Thu 19:30–22:00 MT + Sat 09:00–13:00 MT
  • P1 ack ≤15 min in-window
  • 2 incidents / month included
  • Unused hours convert to scheduled prevention
  • Out-of-window emergency $225/hr (2-hr minimum) or next window
  • Never 24/7 · no uptime SLA
  • Standalone after Discovery or a Build Project so the stack is known
Discuss a Watch Window

Start with Discovery. From there: a Build Project, a Platform Window, or both. Watch Window stacks on a Platform Window — or stands alone after Discovery or a Build Project.

When We Are Present

Set windows in US Mountain Time. Async is the default; live work is the exception.

Platform Window

Mon–Thu 16:30–19:30 MT · Fri 16:30–18:00 MT

Live pairing and same-window work Monday–Thursday. Friday 16:30–18:00 MT is for the weekly status artifact only. 16:30 MT = 15:30 PT.

Watch Window

Mon–Thu 19:30–22:00 MT · Sat 09:00–13:00 MT

Defined-window on-call — never 24/7, no uptime SLA. In-window P1 ack ≤15 min. Nights after 22:00 MT and weekday daytime are out of window.

Async-first

Written asks are the default. Live work is the exception, booked inside a published window.

Friday artifact

Weekly Friday Loom or written equivalent, plus the hour ledger and next week’s queue.

Replies

Inquiry replies within 24 hours. In-window Platform Window ack ≤30 min.

Out of window

Work waits for the next window, unless a Watch emergency at $225/hr with a 2-hour minimum (founder discretion).

On-site

Salt Lake City Saturdays by arrangement when the work needs a room — not a 15-minute on-site response.

Direct Answers

Straight answers on scope, pricing, and how the work runs — the things teams usually ask before a first call.

How much does BP Data charge for infrastructure consulting?

Engagements start at $500 for Discovery (scoped triage), $3,500 per month for a 16-hour Platform Window, and $8,500 for a fixed-scope Build Project. Watch Window is $1,200 per month as an add-on to a Platform Window, or $1,800 per month standalone. Those are published floors — final scope is set together after Discovery. Platform Window is the ongoing operations engagement (formerly listed as Retainer).

What is included in the $500 Discovery assessment?

Discovery is a scoped triage: 3–4 hours one-time, a 45-minute kickoff, and written findings within 7 days. It includes an infrastructure audit, an architecture map with gap analysis, a cost and usage review, and a written next-step — project, Platform Window, or not a fit.

Does BP Data work remotely or on-site in Salt Lake City?

BP Data is based in Salt Lake City, Utah, and works remote-first with teams across North America. On-site sessions in Salt Lake City are available Saturdays by arrangement when the work genuinely needs a room. This is not a 15-minute on-site response.

When is the coverage window?

Set windows are in US Mountain Time. Platform Window: Monday–Thursday 16:30–19:30 MT, plus Friday 16:30–18:00 MT for the weekly status artifact. Watch Window: Monday–Thursday 19:30–22:00 MT and Saturday 09:00–13:00 MT. Async is the default; live work is the exception. The Friday artifact is a Loom or written equivalent, plus the hour ledger and next week's queue.

Does a Platform Window include on-call?

No. Watch Window is the incident-response product. A Platform Window is 16 hours of scheduled ops in a set evening window — it does not include on-call. Watch is never 24/7 and carries no uptime SLA.

Kubernetes or Nomad — which one fits?

Kubernetes fits broad ecosystem and hiring gravity; Nomad often fits mixed workloads and simpler day-2 operations. BP Data designs and operates either — and will say when a managed service is enough.

What infrastructure platforms does BP Data operate?

Cloud, hybrid, and bare metal — with deep focus on Kubernetes and Nomad, CI/CD and developer platforms, and self-hosted services such as Bitbucket, GitHub Actions runners, N8N, Mage AI, and Proxmox.

Is self-hosting worth it?

When workloads are steady, data control matters, or SaaS and egress costs outpace operations, self-hosting can be the calmer path. When elasticity dominates, cloud stays the right default — the tradeoff gets modeled on paper before anything moves.

Who is BP Data best suited for?

Teams running self-hosted or hybrid platforms that have outgrown managed services but do not need — or cannot yet staff — a full platform organization. Typical fits include growing product companies and internal platform owners who need senior ops on Bitbucket, runners, N8N, Mage AI, Proxmox, Nomad, or Kubernetes.

How quickly does BP Data respond to new inquiries?

Replies typically land within 24 hours, by email or Slack once a direct channel is open. No commitment is required to start a conversation.

Start a Conversation

Infrastructure needs, platform improvements, cost review, migrations, reliability — bring the problem. Start with Discovery if you are unsure which engagement fits. No commitment required, and replies typically land within 24 hours.

Practice

BP Data is the independent practice of Brandon Phillips — infrastructure & platform engineering, Salt Lake City.

Response

Typically within 24 hours. In-window Platform Window ack ≤30 min.

Location

Salt Lake City, Utah · Remote — North America · On-site Saturdays by arrangement