Brad Clark
I lead teams that focus on maximizing the time engineers and stakeholders spend on solving business problems, and minimize the time they need to spend on technical details around implementation and operation of these solutions. I accomplish this through careful balance: when to accrue technical debt and when to pay it back, when to build or buy, when to question and to commit, and when to put business needs over technical correctness. Through a nuanced and thoughtful approach, I empower teams to focus on core problem-solving, enhancing overall productivity and achieving organizational goals.
Summary
- 15+ years combined experience in leadership and platform engineering; built engineering teams from scratch up to 6 engineers
- Maintained 99.98% service uptime by implementing best practices, observability tooling, and CI/CD improvements
- Detail-oriented compliance leader, securing year-over-year compliance certifications across multiple programs
- Able to quickly switch from a high-level business-centric view of a situation to a technical deep dive
- Open-source maintainer and contributor
Skills
- Management
-
Agile methodologies (Scrum, Kanban), Interviewing and hiring individual contributors and peer managers, Coaching and Mentoring Engineers, Vendor management, Building cross-team relationships with technical teams and business stakeholders
- Tools
-
Kubernetes, Terraform, GCP, AWS, Helm, Docker, Ansible, Git, PostgreSQL, MySQL, MongoDB, NGINX, Apache HTTP, Linux, BigQuery, dbt, Argo Workflows, Skyvia, Fivetran, Elasticsearch, Logstash, Datadog, Cloudflare, Doppler, Vault, Github
- Languages
-
Bash, Python, Ruby, Go, YAML, HCL, regex, node.js, JavaScript, Markdown
- Compliance
-
HIPAA, HITRUST, SOC 2, PCI-DSS
- Other
-
Jira, CI/CD, Capacity Planning, System Design
Experience
Author Health | Remote
Director of Platform Engineering | 3/2024 – Present
- Defined and led the infrastructure platform and InfoSec strategy, aligning cloud architecture and engineering execution with company goals
- Deployed Infrastructure as Code using Terraform in CI/CD to ensure consistent, auditable, and repeatable environments
- Led engineering efforts to improve documentation, security posture, and compliance adherence across the organization
- Built and operated secure, scalable, highly available cloud infrastructure with emphasis on reliability, observability, automation, and cost management
- Acted as subject matter expert on infrastructure, security, and compliance, providing technical direction to engineering teams
- As Security Officer, created and implemented security policies, processes, and guidelines to ensure compliance with applicable laws/regulations (including HIPAA and partner agreements)
- Achieved HITRUST i1 certification
- Managed cross-functional evidence collection, ensuring on-time certification renewals
LightTouch Technology
Owner | 2/2024 – Present
- Provided infrastructure consulting and engineering to early- and mid-stage SaaS companies, primarily in the healthcare space, spanning both advisory and hands-on implementation
- Built secure cloud foundations (landing zones) from scratch for core platforms on GCP using Terraform, adapting patterns from Google’s terraform-example-foundation, with ancillary services hosted on AWS
- Supported and integrated legacy and existing client applications across multiple cloud providers
- Helped clients pass HIPAA assessments and achieve HITRUST e1 and i1 certification on infrastructure built as part of engagements
- Led migration of multiple databases and applications from AWS and Aptible to GCP, consolidating data into a secure, BigQuery-based data warehouse
Eleanor Health | Remote
Head of Platform and Data Engineering | 8/2023 – 2/2024
- Built the data engineering organization from scratch, through hiring and internal development of existing engineers
- Managed members of four engineering teams across multiple disciplines: Software Development, Platform Engineering, and Data Engineering. Provided mentorship and coaching while balancing business needs and priorities
- Worked closely with business stakeholders, such as the Revenue Cycle, Clinical Outreach, and Partner Relations teams, to identify key projects that have the largest business impact
- Acting Product Manager for Data Engineering and Platform Engineering, balancing feature work with tech debt
- Reduced data processing times for tasks from 24-72 hours to approximately 15 minutes by improving pipelines, allowing for improved outcomes by shortening the time between events and member outreach
- Reduced manual task ingestion work by 98% through automation
Head of Infrastructure | 4/2021 – 8/2023
- Hired and built the Infrastructure Engineering team from scratch
- Co-authored, designed, and developed our engineering hiring guidelines and process, helping to provide a consistent candidate experience and shorten our hiring cycle, while maintaining quality
- Created career ladders and competency matrices to align salary and performance expectations, and provide clear growth paths for team members
- Created onboarding and 30/60/90 day goals for engineers, creating a clearly defined path through onboarding to improve the new hire experience
- Achieved a near-zero rate of bad production deployments (99.98% app uptime) via improved observability. This was accomplished through strategic mentorship of an engineer, guiding them to spearhead this project, which led to both increased service reliability and professional growth in key areas, such as cross-team collaboration and communication
- Authored and implemented a production-readiness process, providing clear guidelines for engineers to launch stable, observable, and maintainable services
- Ensured products and systems met HIPAA compliance by leading architecture and technical discussions and implementations. Helped achieve SOC2 Type 1 compliance by partnering with the CISO and Head of Product Operations to complete the work across various teams, delegating work as required to complete our audit
- Built highly-available (over 99.99%) infrastructure to host patient management and member apps. Defined SLOs (Service Level Objectives) and SLI (Service Level Indicators) to ensure we meet our SLA (Service Level Agreements)
- Designed systems to meet disaster recovery goals and compliance (HIPAA, SOC2) requirements
- Decreased page load times by over 50% by identifying areas for improvement and facilitating their implementation
Sun Life (Maxwell Health) | Remote
Associate Director, Senior Infrastructure Engineer | 1/2019 – 4/2021
- Rebuilt Infrastructure Engineering team from near-zero to 5 engineers
- Advised engineering teams on platform and architectural decisions
- Provided technical and professional leadership and mentorship to the Infrastructure Engineering team
- Wrote production readiness guidelines, establishing a base set of recommendations for engineers when creating products, and advised development teams on best practices for monitoring, operating, and maintaining stable, reliable services
- Maintained a standardized deployment pipeline built with Github Actions, Docker, Helm, and multiple Kubernetes environments
- Redesigned MongoDB deployment to allow for the automation of most of the manual work required to replace lost nodes or perform maintenance, reducing risk and greatly decreasing recovery time
Maxwell Health | Hybrid, Boston, MA
Infrastructure Engineer | 8/2016 – 1/2019
- Led project to manage all AWS resources as Infrastructure as Code, using Terraform, reducing differences between multiple environments, which ultimately cut down development time and reduced deployment risk
- Helped maintain compliance (HIPAA, NYDFS), improved security posture, and decreased effort and time required for disaster recovery with auditable infrastructure changes, RBAC (Role Based Access Control) managed through Terraform
- Designed event message store backed by Elasticsearch, enabling efficient and reliable communication between dozens of microservices, reducing technical debt while providing durability and the ability to replay messages as needed
- Reduced deployment time ~90% by migrating microservices from Elastic Beanstalk to Kubernetes
- Enhanced developer productivity by building templated Helm charts for Kubernetes
- Improved engineer efficiency, cross-team collaboration, and software security by standardizing our CI/CD build pipeline
Acquia | Hybrid, Boston, MA
Site Reliability Engineer | 8/2014 – 5/2016
- Ensured site availability during some of the world’s largest events: The Olympics, The Super Bowl, and more
- Worked in a high-touch relationship with large enterprise customer, including regular calls and on-site visits
- Reduced downtime to near-zero from multiple customers, by developing tools for auditing and automation to allow for capacity planning as needs changed
- Improved platform monitoring and created auditing tools to improve stability
- Built tooling to use metrics to provide proactive notifications of pending issues before they occur
Senior Cloud Systems Engineer | 7/2013 – 8/2014
- Managed and maintained 13,000+ Linux Systems on Amazon EC2
- Redesigned the monitoring system to accommodate 10,000+ systems and 300,000+ services
- Lowered MTTR (mean time to resolve) and improved engineer productivity by developing standardized tooling and development environments
PlumChoice | Lowell, MA
Linux Systems Administrator | 2/2011 – 7/2013
- Centralized management configuration and security updates for Red Hat, CentOS, and Ubuntu Servers
- Architected system and network monitoring across multiple sites worldwide, using Nagios, Icinga, and Splunk
- Acted as company Security Administrator for PCI-DSS certification. Assisted in writing policies, procedures, and reviewing all security incidents
- Created PC cloning and imaging solution using Clonezilla and DRBL
- Designed and implemented escalation procedures for Operations to Development
Education
- 2001-2002
-
Computer Science; Worcester Polytechnic Institute (Worcester, MA)
- 2002-2003
-
Management Information Systems; UMass Lowell (Lowell, MA)
- 2003-2006
-
Management Information Systems; Middlesex Community College (Lowell, MA)
[email protected] • +1-978-846-1533