Job Summary
We are looking for a DevOps Engineer to build, automate, and operate the cloud infrastructure and software delivery platforms that enable engineering teams to develop, deploy, and run secure, reliable, and scalable applications.
The ideal candidate will have strong experience with AWS, Infrastructure as Code, CI/CD, automation, containers, cloud operations, and DevSecOps. You will play a key role in improving deployment efficiency, system reliability, observability, security, and operational excellence across our technology environment.
You will also have the opportunity to support AI-enabled applications and platforms, including AI governance, model monitoring, and agent/LLM observability.
Key Responsibilities
- Design, build, and maintain AWS cloud infrastructure, CI/CD pipelines, and platform automation.
- Automate application deployments, environment provisioning, infrastructure management, monitoring, and operational processes.
- Develop and maintain Infrastructure as Code (IaC) using tools such as Terraform or similar technologies.
- Build reliable and scalable deployment pipelines that support modern software development practices.
- Support production environments, incident management, troubleshooting, and root cause analysis.
- Implement reliability, availability, resiliency, and disaster recovery best practices.
- Implement DevSecOps practices, security controls, compliance automation, and infrastructure security.
- Work closely with software engineers and architects to ensure applications are production-ready.
- Establish monitoring, logging, alerting, and observability capabilities across applications and infrastructure.
- Develop reusable DevOps and platform capabilities that improve developer productivity and operational consistency.
- Automate repetitive operational tasks and continuously improve engineering workflows.
- Participate in production readiness reviews and help establish operational standards and best practices.
- Support governance, auditability, and compliance requirements through automation and policy-based controls.
- Help support infrastructure and operational requirements for AI-enabled applications, agents, and LLM-based solutions.
Required Qualifications
- 10+ years of experience in DevOps, Cloud Engineering, Platform Engineering, Site Reliability Engineering (SRE), or a related field.
- Strong hands-on experience with AWS cloud services and infrastructure.
- Experience with Infrastructure as Code (IaC), preferably Terraform.
- Strong understanding of CI/CD pipelines, DevOps practices, automation, and cloud infrastructure.
- Experience with containers and container orchestration technologies such as Docker and Kubernetes.
- Experience working with APIs, distributed systems, and modern software delivery practices.
- Strong scripting and automation skills using Python, Bash, PowerShell, or similar languages.
- Understanding of cloud security, networking, IAM, monitoring, and production operations.
- Strong troubleshooting, problem-solving, and incident management skills.
- Ability to collaborate effectively with software engineers, architects, security teams, and other technical stakeholders.
Preferred Qualifications
- Experience with observability and monitoring platforms such as:
- AWS CloudWatch
- OpenTelemetry
- Datadog
- Grafana
- Prometheus
- Splunk
- Experience with governance automation, policy-as-code, auditability, and compliance controls.
- Experience implementing DevSecOps and security telemetry.
- Experience working in regulated or compliance-focused environments.
- Experience with security automation and cloud security best practices.
- Familiarity with AI governance, model monitoring, AI infrastructure, or agent/LLM observability.
- Experience supporting production AI/ML or AI-enabled applications.
Key Skills
AWS | DevOps | Cloud Engineering | Platform Engineering | CI/CD | Terraform | Infrastructure as Code | Docker | Kubernetes | Python | Bash | PowerShell | DevSecOps | Cloud Security | Observability | Monitoring | OpenTelemetry | CloudWatch | Grafana | Prometheus | Splunk | SRE | Automation | Infrastructure Automation | AI/LLM Observability
Pay: $70.00-$100.00 per hour
Application question(s):
- How many years of experience do you have in DevOps, Cloud Engineering, Platform Engineering, or SRE?
A. 8+
B. 10+
- How many years of hands-on AWS experience do you have?
A. 5+
B. 7+
C. 10+
- Which scripting languages do you use for automation?
A. Python
B. Bash,
C. PowerShell
D. Other
- Do you have experience implementing cloud security or DevSecOps practices?
A. Yes
B. No
- Have you worked with monitoring/observability tools such as CloudWatch, OpenTelemetry, Datadog, Grafana, Prometheus, or Splunk?
A. Yes
B. No
Work Location: Hybrid remote in Markham, ON L3R 9Z7