Senior Platform Engineer - Cloud Architecture, Fleet Delivery, and Engineering Infrastructure

Summary
This role owns the engineering spine underneath that: how the platform is designed, how software reaches machines already installed in plants, what we know about the state of every one of them, and how a problem found in the field becomes a tracked defect and a delivered fix. The center of gravity is the cloud and central services side. The real-time inspection software on the machines is built by product teams; this role owns the mechanisms that deliver it, record it, and watch over it, and the interfaces the two sides meet across.
On-site - Westborough, MA. Full-time, permanent, senior/staff-level position, with occasional travel to customer sites required.
Duties and Responsibilities
- Set and document the architecture of the central platform and every contract the products meet it across: service boundaries, data schemas, versioning and compatibility rules, edge-to-cloud protocols, and tenant isolation. Architecture here is written down, reviewed, and revisited - decision records, interface specifications, and diagrams that other teams build against. Holds review authority on changes that cross those boundaries.
- Lead the contract between the machines and the central platform, and the tooling that lets every product line meet it the same way, including on the constrained, real-time side. The goal is one supported path to the cloud rather than one integration per product.
- Take platform and infrastructure workstreams from written design to delivered, operated software. Decompose the work, sequence it across contributors and sites, run the review cadence, surface risk early, and be accountable for the date.
- Define and build the mechanism by which software reaches systems already running in production plants: release channels, compatibility checks, staged rollout to named sites, resumable transfer over constrained and intermittent links, tested rollback, and confirmation that the machine ended up running what it was told to run.
- Maintain a durable, queryable record of every deployed system - hardware, software versions, optics, and configuration, and the history of every change made to it. Own its design and the discipline that keeps drift visible rather than silent.
- Design and deliver the path by which installed systems report what is true of them - health, uptime, alarms, running versions, live configuration - back to a central view, over intermittent, locked-down, customer-controlled networks.
- Direct how evidence gets off a machine in the field: structured logs, diagnostic bundles, reproducible capture of the conditions around a fault, and secure retrieval that a support engineer can run without a developer present.
- Run one path from field observation to tracked defect to delivered fix: internal defect tracking, triage that assigns ownership, and a hotfix route that reaches a single site fast without leaving the record behind it inaccurately.
- Be accountable for the tooling engineers depend on: CI/CD, build agents, artifact and package hosting, code signing, test and simulation harnesses, and reproducible development environments, plus release engineering for the central services, including AI-agent tooling and harnesses, treated as production infrastructure.
- Govern the ground everything runs on: infrastructure as code for computing, managed databases, storage, networking, identity, and secrets; containerized service deployment; per-customer isolation enforced in infrastructure and proven by automated tests; metrics, logs, and alerting wired per service and per tenant.
Qualifications Required
- 7+ years of professional software engineering, several of them owning systems that ran in production for real users, not only contributing to them.
- Architecture ownership across more than one team or product, with a written record: designs, decision records, or specifications authored that others built against.
- Experience building or operating software delivery to installed/remote machines you cannot reach release channels, compatibility gating, staged rollout, rollback, and verification of end state, over intermittent links and customer networks you do not control.
- Experience building or operating a source of truth for deployed state, versions, configuration, change history, kept reconciled against reality.
- Infrastructure as code delivered to production (Terraform, Bicep, CloudFormation, Pulumi, or equivalent) that provisioned environments people depended on, with a defensible state, drift, and rollback story.
- Containers and CI/CD as daily working tools, containerized services operated under an orchestrator or managed container service, and pipelines built rather than inherited.
- Hands-on production use of at least one major public cloud (Azure, AWS, or GCP): networking, managed databases, storage, identity and access control, TLS, secrets handling, and least privilege.
- Strong programming in at least two languages used for real services or automation.
- Systems-level C or C++ experience, including shipping a library other teams embedded into products you did not own, with the versioning, memory discipline, and support that followed.
- Linux and Git as daily working tools.
- Observability practice, instrumented systems you operated and used what you built to find real faults.
- AI-assisted development ("vibe coding") as a working skill: directing an AI coding agent through real, multi-step work, and critically reviewing, testing, and correcting what it produces.
- Independent delivery and technical leadership, able to take an ambiguous problem, produce a written design, decompose and sequence work across a distributed team, and deliver on a cadence.
- Clear written communication and professionalism to have work reviewed critically and to review others' the same way.
- Eligible to work on-site in Westborough, MA, on a permanent basis.
Language Skills
Clear written and verbal English communication required. Must be able to write designs, decision records, and specifications that other teams build against, and communicate clearly across engineering, product, and field service teams spanning multiple sites and time zones.
Education
No specific degree required.
Traits/Skills Preferred: (the more of these, the stronger the application)
- Multi-tenant SaaS isolation done in earnest (per-tenant databases, credentials, network policy)
- Kubernetes in production
- A relational database operated at scale (provisioning, grants, backups, migrations)
- PKI, code signing, and mTLS beyond the basics (running a CA, certificate lifecycle)
- Open Telemetry or an equivalent observability stack
- Telemetry and OTA update systems for connected devices
- Configuration management or desired-state tooling operated at fleet scale
- Libraries or SDKs designed for other teams to embed, including the support burden that came with them
- Industrial, IoT, machine vision, or instrumentation systems experience
- Systems-level depth on the real-time side - memory management, scheduling, and performance tuning of deterministic systems
- Experience hosting other teams' services on a platform you ran
- Experience in regulated or security-reviewed environments
Work Environment
- Works from written designs and recorded decisions, expected to read them before changing things, and to write the same way.
- Everything moves through pull requests and review, infrastructure, documentation, and runbooks included, written as you go. No direct-to-production changes.
- Collaborates across sites and time zones with engineering, product, and field service teams who depend on what is built.
- Weekly delivery checkpoints against an agreed plan; risks and slips are raised when they appear, not at the deadline.
Location
On-site - Westborough, MA. Full-time, permanent, senior/staff-level position, with occasional travel to customer sites required.