Site Reliability Engineer III (SRE) - Guidewire Cloud Platform (Application)
ProNavigator
Software Engineering
Dublin, Ireland
Job Description
What You'll Do
Work with development teams to troubleshoot and resolve issues, minimizing customer impact.
Develop and maintain automated runbooks to manage issues proactively.
Apply engineering principles and automation to enhance our operating environments.
Monitor and improve the reliability and performance of applications on the Guidewire Cloud Platform.
Use your software engineering expertise to optimize systems and reduce manual toil.
Document incidents and develop processes to prevent future occurrences.
Stay current with industry trends, tools, and best practices in site reliability engineering.
Foster a culture of innovation, learning, and continuous improvement.
Participate in on-call rotations to ensure the availability and reliability of our services.
What You'll Bring
Experience as an SRE or similar role, with a focus on improving system reliability.
Strong problem-solving skills and the ability to analyze complex systems and devise effective solutions.
Effective collaboration and communication skills to work cross-functionally and document processes clearly.
Experience with automation, monitoring, and performance optimization tools and techniques.
Commitment to maximizing uptime, scalability, and delivering an exceptional end-user experience.
Passion for technology and a desire to continuously learn and grow your skills.
Alignment with Guidewire's mission to leverage technology to help protect and support others.
Required Skills:
Software engineering background with experience in Python, Go, or Java, following best practices (SOLID, DRY, KISS) and writing clean, testable code
Experience with designing and implementing SLI's, SLO's, and Error Budgets
Familiarity with application performance monitoring (APM) and telemetry tools to maintain expected service levels for applications
Experience troubleshooting and debugging distributed systems on cloud infrastructure
Experience with CICD pipelines within K8S and legacy ecosystems
Experience creating monitors, dashboards, and synthetic transactions in monitoring tools like Datadog
Experience deploying and managing scalable infrastructure within AWS and Kubernetes ecosystems using Terraform and other cloud-native approaches
Experience with infrastructure configuration management using tools such as GitOps, Puppet, or Ansible
Good understanding of cloud networking, security, and vulnerability management, with the ability to programmatically remediate infrastructure issues
Preferred Skills:
SRE Certification in one or more categories
AWS Certification in one or more categories
Experience with SQL, database administration, data pipelines, performance tuning, and schema design
Familiarity with pipelining tools such as Team City, Bitbucket Pipelines, Jenkins, or GitHub Actions
Exposure to open-source distributed data processing frameworks such as Hadoop, Apache Spark, AWS RedShift, etc.
Experience with distributed systems, including microservices and event-driven architectures
Why Guidewire
This is an opportunity to join a mission-driven company and make a real impact in the lives of people facing challenges. You'll work with cutting-edge technology, collaborate with talented peers, and grow your skills in a culture that values innovation, teamwork, and work-life balance. We offer competitive compensation, comprehensive benefits, and opportunities for career development.
If you're an SRE who combines deep technical expertise with a passion for problem-solving and a commitment to reliability, we'd love to hear from you. Join us in building the software that helps insurers care for their customers when they need it most.
This position requires participation in mandatory on-call rotations to ensure the availability and reliability of our services. This includes responding to incidents and alerts outside of regular business hours, on weekends, and during holidays, as per the established on-call schedule. Candidates must be willing and able to fulfill this critical responsibility.
#LI-AS3