Bluecopa
Back to job details
Bluecopa Full-time Hyderabad

Team Lead/EM – SRE

Team Lead/EM – SRE

Location: Hyderabad (Work from Office) | Type: Full-time | Experience: 7–10 years | Working Days: 5 Days a Week (Mon-Fri)


About Bluecopa

Bluecopa is building the next-generation finance operations platform for high-growth companies. Our mission is to help finance teams automate, analyse, and act faster on their data — enabling smarter decisions and accelerating business outcomes.

We're a fast-paced, product-driven team solving real-world data problems at scale. If you love running reliable production systems, leading strong engineering teams, and being the calm hand during an incident, we'd love to have you on board.


The Role

We're looking for a Team Lead – SRE to lead our Site Reliability and DevOps team. You'll own the reliability, availability, and operability of the Bluecopa platform across both enterprise (customer-hosted) and Bluecopa-managed deployments.

This role sits at the intersection of engineering and customers. You'll drive customer onboarding end-to-end — requirements, coordination, approvals, and provisioning — whether the platform runs in the customer's environment or ours. You'll also own monitoring, incident management, and platform support for our enterprise customers.

This is a hands-on leadership role. You'll manage a team of SRE and DevOps engineers, set the standards for how we operate production, and stay close enough to the systems to debug, unblock, and lead from the front when it matters.


What You'll Do

Team Leadership

  • Lead and manage a team of SRE and DevOps engineers — planning, prioritization, mentoring, and performance.
  • Define and enforce operational standards, runbooks, and on-call practices across the team.
  • Balance project work (onboarding, automation) with operational work (support, incidents) across the team.

Enterprise Customer Onboarding (Customer-Hosted)

  • Gather and validate infrastructure, security, and deployment requirements with enterprise customers.
  • Coordinate with customer IT, security, and infrastructure teams to plan and execute deployments.
  • Drive approvals — security reviews, access requests, compliance sign-offs — to keep onboarding on schedule.
  • Own provisioning and deployment of the Bluecopa platform in the customer's environment, end-to-end.

Bluecopa-Managed Customer Onboarding

  • Own requirements gathering and environment planning for customers hosted on Bluecopa-managed infrastructure.
  • Coordinate internally with Product, Platform, and Delivery teams to sequence onboarding activities.
  • Manage approvals and access provisioning across internal and customer stakeholders.
  • Provision, configure, and hand over customer environments on Bluecopa-managed cloud infrastructure.

Enterprise App / Product / Platform Support

  • Own app, product, and platform support for enterprise customers — triage, resolution, and escalation.
  • Work with Engineering and Product teams to resolve issues and feed recurring problems back into the roadmap.
  • Track and meet support SLAs; communicate clearly with customers during issues.

Monitoring Setup

  • Design and set up monitoring, alerting, and observability for all customer environments — enterprise and Bluecopa-managed.
  • Build dashboards and alerts that catch issues before customers do — infrastructure, application, and data-pipeline health.
  • Continuously tune alerting to reduce noise and improve signal.

Monitoring & Incident Management

  • Own the incident management process — detection, response, communication, and resolution.
  • Run structured RCAs (Root Cause Analysis) and post-incident reviews; drive corrective actions to closure.
  • Establish and manage on-call rotations, escalation paths, and incident severity frameworks.
  • Report on reliability metrics — uptime, MTTR, incident trends — and drive them in the right direction.

What We're Looking For

Core (All Mandatory)

  • 7–10 years of experience in SRE, DevOps, or infrastructure engineering roles, with at least 2 years leading a team.
  • Strong hands-on experience with Kubernetes and containerized production environments.
  • Hands-on with at least one major cloud platform: GCP, AWS, or Azure.
  • Proven experience with monitoring and observability stacks — Prometheus, Grafana, ELK, Datadog, or similar.
  • Strong incident management experience — running incidents, writing RCAs, and driving post-incident improvements.
  • Experience deploying and supporting software in customer-controlled (enterprise) environments.

Infrastructure & Automation

  • Experience with Infrastructure as Code — Terraform, Helm, or similar.
  • Solid grasp of CI/CD pipelines and release management.
  • Scripting proficiency in Python or Bash for automation and tooling.
  • Working knowledge of networking, security, and access management fundamentals — VPNs, firewalls, IAM, certificates.

Ways of Working

  • Comfortable working directly with enterprise customers — technical discussions, approvals, and escalations.
  • Able to coordinate across customer teams and internal teams to keep multi-stakeholder projects moving.
  • Clear, proactive communicator — especially under pressure during incidents.

Good to Have

  • Experience onboarding customers in regulated or security-sensitive environments (security reviews, VAPT, compliance approvals).
  • Exposure to SaaS platform operations and multi-tenant environments.
  • Familiarity with ITIL or similar service management frameworks.
  • Prior experience in a high-growth startup or product company.

30–60–90 Day Goals

  • First 30 Days — Learn & Stabilize: Understand the Bluecopa platform architecture, deployment models (enterprise and Bluecopa-managed), and current customer environments. Meet the SRE/DevOps team; understand each member's strengths, workload, and current commitments. Review existing monitoring, alerting, on-call, and incident processes; identify the biggest gaps. Shadow at least one customer onboarding end-to-end and take stock of all in-flight onboardings and open support issues.
  • By 60 Days — Own & Improve: Take full ownership of customer onboarding — requirements, coordination, approvals, and provisioning — for both deployment models. Standardize onboarding into a repeatable checklist/runbook with clear owners and timelines. Close the top monitoring gaps — dashboards and alerts in place for all active customer environments. Establish the incident management framework — severity levels, escalation paths, on-call rotation, and RCA template.
  • By 90 Days — Scale & Lead: Run onboarding and support predictably — customers onboarded on schedule, support SLAs consistently met. Every major incident has an RCA with corrective actions tracked to closure. Publish a monthly reliability report — uptime, MTTR, incident trends — with a clear improvement plan. Have a clear view of team structure, skill gaps, and automation priorities for the next two quarters.

Mindset

High ownership. Calm under pressure. You treat reliability as a product, not an afterthought — and you hold the bar for your team, your customers, and yourself.

You don't wait to be told what to do. You find the problem, scope it, and fix it — and you build a team that does the same.

Your application

Fields marked with * are required.

PDF, DOC, or DOCX — max 10 MB

Additional questions

Your information is kept confidential.