Questrade Financial Group (QFG), through its companies - Questrade, Questbank, Questrade Wealth Management, Community Trust Company, Zolo, and Flexiti, provides securities and foreign currency investment, professionally managed investment portfolios, mortgages, real estate services, financial services and more. We use cutting-edge technology to help Canadians become much more financially successful and secure.
At QFG, we combine human-centric collaboration with AI-driven innovation to redefine financial services. The ideal candidate will be a catalyst for change, using AI to transform and deliver unparalleled customer experiences and shaping a future where AI empowers our teams to do their best work.
Join our diverse, inclusive, and hybrid workplace to unleash your creativity and nurture your curiosity without limits. If you share this sense of infinite possibility, come shape your future at QFG.
What’s in it for you as an employee of QFG?
- Health & wellbeing resources and programs
- Paid vacation, personal, and sick days for work-life balance
- Competitive compensation and benefits packages
- Work-life balance in a hybrid environment with at least 3 days in office
- Career growth and development opportunities
- Opportunities to contribute to community causes
- Work with diverse team members in an inclusive and collaborative environment
This job posting is for an existing vacancy
We’re looking for our next Senior Principal Site Reliability Engineer. Could It Be You?
The Senior Principal Site Reliability Engineer is directly responsible for the stability, resiliency, and scalability of business-critical brokerage back-end applications running across a hybrid on-premises and cloud architecture.
This individual drives reliability engineering practices — SLOs/SLIs, observability, incident response, and capacity planning — while also making hands-on, code contributions directly into multiple applications across different technology stacks using an inner-source model. The role blends deep technical execution with cross-team influence: this person is expected to identify systemic reliability risks, fix them where they appear, and raise the operational bar for every team they touch.
This position is a strong fit for a hands-on senior principal engineer who is energized by fixing production reliability at the source, is comfortable navigating multiple codebases and cloud/on-prem environments, and wants to have an outsized impact on the resiliency of a regulated, high-availability brokerage platform.
Need more details? Keep reading…
Application Stability, Reliability & Growth
- Own the end-to-end reliability posture of critical brokerage back-end applications, driving measurable improvements in availability, latency, and error budgets across on-premises and cloud environments.
- Define and track Service Level Objectives (SLOs), Service Level Indicators (SLIs), and error budgets in partnership with application teams; use them to prioritize reliability work over feature work when warranted.
- Lead root cause analysis and blameless post-incident reviews for high-severity production incidents; drive remediation items to closure and identify systemic patterns across applications.
- Establish and mature observability practices (metrics, logging, tracing, alerting) so that failures are detected proactively and diagnosed quickly across a heterogeneous, multi-stack estate.
- Build capacity planning, load testing, and chaos/failure-injection practices to validate resilience before incidents occur.
- Champion a culture of operational excellence, toil reduction, and AI and automation-first thinking across engineering teams.
Cross-Stack Engineering & Inner Sourcing
- Make hands-on, code contributions directly into multiple applications spanning different languages, frameworks, and stacks, using an inner-source model to fix reliability defects, add instrumentation, and improve resiliency patterns.
- Partner with individual application teams to raise pull requests, follow their contribution standards, and pair with owning engineers so fixes land safely and are properly reviewed and owned long-term.
- Identify recurring reliability anti-patterns across codebases (e.g., missing timeouts/retries, unbounded queues, improper connection pooling) and drive standardized, reusable fixes or shared libraries.
- Contribute to and help govern internal reliability tooling, shared SDKs, and common patterns (circuit breakers, backoff/retry, health checks) that can be inner-sourced across teams.
Cloud Scalability & Hybrid Architecture
- Design and advise on cloud scalability strategies (auto-scaling, load balancing, multi-region/multi-AZ failover, caching, queuing) for workloads that span on-premises data centers and public cloud.
- Guide capacity and cost-aware scaling decisions, balancing performance, resiliency, and cloud spend across hybrid deployments.
- Evaluate and recommend cloud-native and hybrid resiliency patterns (e.g., disaster recovery, active-active/active-passive architectures, data replication strategies) appropriate for regulated brokerage workloads.
Organizational Awareness & Risk
- Bring strong organizational awareness of the operational, financial, regulatory, and reputational risk that production incidents pose to a brokerage business, and factor that into prioritization.
- Participate in risk assessments related to system reliability, availability, and disaster recovery, partnering with Risk, Compliance, and Information Security as needed.
- Contribute to change management and release governance practices that reduce the likelihood and blast radius of production incidents.
- Promptly identify, escalate, and help remediate reliability or security-related incidents in accordance with company policy.
Leadership & Influence
- Act as a technical reference and mentor for reliability engineering practices, coaching application teams on operational excellence without formal direct reports.
- Influence architecture and design decisions across multiple teams by bringing a reliability and scalability lens to reviews and planning.
- Document and evangelize reliability standards, runbooks, and best practices; lead or contribute to internal tech talks and communities of practice.
- Partner with engineering leadership to define the reliability roadmap and report on progress against stability goals
So are YOU our next Senior Principal Site Reliability Engineer.? You are if you…
- Bachelor's or Master's degree in Computer Science, Information Systems, Engineering, or a related field, or equivalent combination of education and experience.
- 8+ years of software engineering and/or site reliability engineering experience, including production ownership of business-critical applications; financial services or brokerage experience strongly preferred.
- Demonstrated ability to read, debug, and make minor-to-moderate code changes across multiple languages/stacks (e.g., Java, .NET, Node.js/TypeScript, Python) in an inner-source or cross-team contribution model.
- Deep experience with cloud scalability strategies on one or more major providers (AWS, Azure, GCP), including auto-scaling, load balancing, multi-region resiliency, and cost-aware capacity planning.
- Experience operating and supporting hybrid architectures spanning on-premises data centers and cloud environments.
- Strong background in observability tooling (e.g., Prometheus/Grafana, Datadog, Splunk, ELK, AppDynamics, Dynatrace) and building actionable alerting and dashboards.
- Practical experience defining and operating against SLOs/SLIs/error budgets and running blameless post-incident reviews.
- Experience with CI/CD pipelines and infrastructure-as-code (e.g., Terraform, Ansible, CloudFormation) in support of reliable, repeatable deployments.
- Solid understanding of microservices architecture, distributed systems failure modes, and resiliency patterns (circuit breakers, retries/backoff, bulkheads, timeouts).
- Familiarity with relational and NoSQL data stores and their operational/scaling characteristics.
- Experience with incident management and on-call practices (e.g., PagerDuty, Opsgenie) including leading major incident response
- Knowledge of security, audit, and regulatory considerations relevant to brokerage / financial services production systems.
- Excellent communication skills, with the ability to influence engineers and stakeholders across many teams without direct authority.
- Strong documentation, analytical, and problem-solving skills
Compensation Information:
- Base salary range: $150,000 - $190,000
- The final compensation package will be commensurate with the successful candidate's experience, skills, and geographic location (Canada). It includes a comprehensive benefits plan and a competitive incentive (bonus) program for Full-Time Permanent roles.
Sounds like you? Click below to apply!
#LI-Hybrid
At Questrade Financial Group of Companies, with multiple office locations around the world, we are committed to fostering a diverse, inclusive and accessible work environment. This is an environment where individuals are treated with dignity and respect. Here, the unique skills and experience you bring will be valued. You will be supported and motivated, so that you can harness your unlimited potential. Our team reflects the diversity of the communities we serve and operate in. Having a collaborative and diverse team helps us push boundaries to bring the future of fintech into existence—not only for the benefit of our customers, but for those who build their career with us.
Questrade Financial Group of companies Applicant Tracking System utilizes artificial intelligence (AI) for application screening. The AI system operates on predetermined criteria, with final decisions subject to human review.
Candidates selected for an interview will be contacted directly. If you require accommodation during the recruitment/selection process, please let us know and we will work with you to meet your needs.