About Responsive

Responsive (formerly RFPIO) is the global leader in strategic response management software, transforming how organizations share and exchange critical information. The AI-powered Responsive Platform is purpose-built to manage responses at scale, empowering companies across the world to accelerate growth, mitigate risk and improve employee experiences. Nearly 2,000 customers have standardized on Responsive to respond to RFPs, RFIs, DDQs, ESGs, security questionnaires, ad hoc information requests and more. Learn more at responsive.io.

 

About the Role

We are seeking a highly skilled Senior DevOps Engineer to design, implement, and operate scalable, secure, and resilient cloud infrastructure that supports modern cloud-native applications. The ideal candidate will have deep expertise in Kubernetes, Infrastructure as Code (IaC), CI/CD automation, cloud platforms, observability, and emerging AI-powered DevOps practices.

This role will collaborate closely with Engineering, Security, SRE, and Product teams to build reliable deployment platforms, improve operational efficiency, enhance system reliability, and drive automation across the software delivery lifecycle.

What you will be doing

Essential Responsibilities

  • Design, implement, and manage secure, scalable, and highly available cloud infrastructure across AWS, Azure, or GCP, ensuring optimal performance, disaster recovery, business continuity, and cost efficiency.

  • Deploy, manage, and optimize production Kubernetes platforms, including workloads, networking, storage, Helm, autoscaling, and deployment standards while troubleshooting infrastructure and application issues.

  • Develop and maintain Infrastructure as Code using Terraform and automate infrastructure provisioning and operational workflows using Python or Bash across multiple environments.

  • Design, build, and optimize CI/CD pipelines using Jenkins and GitHub Actions, implementing automated testing, security scanning, and modern deployment strategies such as rolling, blue-green, and canary releases.

  • Implement end-to-end observability and reliability using Prometheus, Grafana, OpenTelemetry, ELK/OpenSearch, and Loki, while defining SLIs, SLOs, error budgets, and capacity planning.

  • Drive DevSecOps by implementing IAM, RBAC, vulnerability management, container security, secrets management, SSL/TLS, and infrastructure compliance best practices.

  • Leverage AI and LLM technologies to automate infrastructure management, generate deployment artifacts, enhance CI/CD workflows, and improve developer productivity.
  • Develop AI-powered operational capabilities for incident response, root cause analysis, anomaly detection, predictive scaling, and cloud cost optimization while evaluating emerging AI technologies.

What we are looking for

Education

Bachelor's degree in Computer Science, Information Technology, or related field.

Experience

  • 7+ years of experience in DevOps, Site Reliability Engineering (SRE), or Platform Engineering.
  • Strong hands-on experience managing production Kubernetes environments at scale.
  • Experience with GitOps tools such as Argo CD.
  • Experience with multi-cloud architectures and cloud governance.
  • Hands-on experience implementing observability solutions and troubleshooting distributed systems.
  • Exposure to AI/ML frameworks or cloud AI services is an added advantage
  • Strong experience with cloud platforms including AWS, Azure, or Google Cloud Platform (GCP).
  • Hands-on expertise in Docker, Kubernetes, and Helm for containerization and orchestration.
  • Experience building and maintaining CI/CD pipelines using Jenkins and GitHub Actions.
  • Experience with Git and GitHub for version control and collaboration.

Knowledge & Skills

  • Proficiency in Infrastructure as Code (IaC) using Terraform.
  • Strong knowledge of monitoring and observability tools including Prometheus, Grafana, OpenTelemetry, ELK/OpenSearch, and Loki.
  • Proficiency in Python and Bash scripting for automation.
  • Good understanding of DevSecOps practices, including Trivy, OSV, SonarQube, container security, vulnerability management, and security best practices.
  • Exposure to AI-powered DevOps technologies, including Prompt Engineering, LLM APIs, AI-assisted development tools, Retrieval-Augmented Generation (RAG), Vector Databases & Embeddings, AI-driven infrastructure automation, and AI-based incident analysis.
  • Strong analytical, troubleshooting, communication, and leadership skills.

Why Join Us?

  • Impact-Driven Work: Build innovative solutions that redefine strategic response management.
  • Collaborative Environment: Work with a passionate team of technologists, designers, and product leaders.
  • Career Growth: Be part of a company that values learning and professional development.
  • Competitive Benefits: We offer comprehensive compensation and benefits to support our employees.
  • Trusted by Industry Leaders: Be part of a product that is trusted by world-leading organizations.
  • Cutting-Edge Technology: Work on AI-driven solutions, cloud-native architectures, and large-scale data processing.
  • Diverse and Inclusive Workplace: Collaborate with a global team that values different perspectives and ideas.