職種概要
Key Responsibilities
Core DevOps Operations
– Manage and upgrade EKS clusters — version upgrades, node group rotations, add-on compatibility, and zero-downtime rollouts
– Troubleshoot and resolve Kubernetes issues: pod failures, OOMKills, scheduling problems, networking, ingress/ALB misconfigurations
– Handle Jenkins administration — pipeline debugging, plugin upgrades, controller/agent connectivity, build failure triage
– Manage AWS infrastructure: EC2 instance rehydration, ASG self-healing, ALB/ELBv2 target health, Route53 DNS, CloudTrail/CloudWatch investigation
– Perform instance rehydration and capacity management across auto-scaling groups
– Respond to and investigate infrastructure incidents; perform root cause analysis and post-incident documentation
– Maintain and improve Helm chart deployments; manage helm upgrade releases across environments
– Manage IAM roles, cross-account access (STS AssumeRole), and security group configurations
– Maintain SSL/TLS certificate lifecycle across internal and external endpoints
Platform & Tooling Development
– Write and maintain Python tooling that integrates with internal systems (Kubernetes API, AWS, Jira, Confluence, Jenkins, LDAP/AD)
– Contribute to an internal AI-powered operations platform — extending tool integrations, debugging issues, and adding capabilities as needed
– Maintain REST APIs built on FastAPI; write clean, typed Python (3.10+) following team standards
– Write and maintain tests using pytest; follow code quality standards (black, isort, mypy)
Collaboration & Process
– Document runbooks, operational procedures, and architecture decisions
– Review infrastructure and code changes; enforce operational best practices
—
Required Skills
DevOps & Infrastructure (Mandatory)
– 5+ years of hands-on Kubernetes experience — EKS preferred; cluster upgrades, node management, networking, RBAC
– Strong AWS skills — EC2, ASG, ALB/ELBv2, Route53, IAM, CloudWatch, CloudTrail, STS
– Helm — chart authoring, release management, values file strategy
– Jenkins — administration, pipeline authoring (Declarative/Scripted), build failure investigation
– Docker — image builds, multi-platform (linux/amd64), container registry (ECR)
– Infrastructure-as-code mindset; comfortable reading and modifying YAML-heavy configs
Python (Mandatory)
– Strong Python 3.10+ proficiency — type hints, pydantic v2, async patterns, httpx
– Experience building or maintaining REST APIs (FastAPI or similar)
– Comfortable writing scripts and tools that integrate with AWS (boto3), Kubernetes (kubernetes client), and internal REST APIs
– Able to read, extend, and debug an existing Python codebase without full prior context
Systems Integration
– REST API integration with enterprise tooling (Jira, Confluence, Jenkins, or similar)
– Authentication patterns: Bearer tokens, Basic Auth, AWS IRSA, cross-account IAM
General
– Strong incident investigation and debugging skills across distributed systems
– Able to work independently, manage priorities, and communicate blockers early
– PostgreSQL — basic query and operational familiarity
—
Nice to Have
– LangChain / LangGraph — multi-agent graphs, ReAct pattern, stateful graph checkpointing
– Any LLM API experience (Anthropic, OpenAI, or similar) — tool use and prompt construction
– Experience contributing to AI-assisted operations or internal developer tooling
—
What You Don’t Need
– ML or data science background
– Prior experience with LangGraph or AI agent frameworks — willingness to learn from existing code is sufficient
– Full-stack frontend experience
必要条件
About the Role
We are looking for a Senior DevOps Engineer to join our operations team. You will own day-to-day operations of our cloud infrastructure — Kubernetes cluster
management, CI/CD pipelines, AWS infrastructure, and incident response — while also contributing to an internal AI-powered operations platform built in Python. The
AI platform work is additive; your primary responsibility is keeping production infrastructure healthy and reliable
職務内容
Key Responsibilities
Core DevOps Operations
– Manage and upgrade EKS clusters — version upgrades, node group rotations, add-on compatibility, and zero-downtime rollouts
– Troubleshoot and resolve Kubernetes issues: pod failures, OOMKills, scheduling problems, networking, ingress/ALB misconfigurations
– Handle Jenkins administration — pipeline debugging, plugin upgrades, controller/agent connectivity, build failure triage
– Manage AWS infrastructure: EC2 instance rehydration, ASG self-healing, ALB/ELBv2 target health, Route53 DNS, CloudTrail/CloudWatch investigation
– Perform instance rehydration and capacity management across auto-scaling groups
– Respond to and investigate infrastructure incidents; perform root cause analysis and post-incident documentation
– Maintain and improve Helm chart deployments; manage helm upgrade releases across environments
– Manage IAM roles, cross-account access (STS AssumeRole), and security group configurations
– Maintain SSL/TLS certificate lifecycle across internal and external endpoints
Platform & Tooling Development
– Write and maintain Python tooling that integrates with internal systems (Kubernetes API, AWS, Jira, Confluence, Jenkins, LDAP/AD)
– Contribute to an internal AI-powered operations platform — extending tool integrations, debugging issues, and adding capabilities as needed
– Maintain REST APIs built on FastAPI; write clean, typed Python (3.10+) following team standards
– Write and maintain tests using pytest; follow code quality standards (black, isort, mypy)
Collaboration & Process
– Document runbooks, operational procedures, and architecture decisions
– Review infrastructure and code changes; enforce operational best practices
—
Required Skills
DevOps & Infrastructure (Mandatory)
– 5+ years of hands-on Kubernetes experience — EKS preferred; cluster upgrades, node management, networking, RBAC
– Strong AWS skills — EC2, ASG, ALB/ELBv2, Route53, IAM, CloudWatch, CloudTrail, STS
– Helm — chart authoring, release management, values file strategy
– Jenkins — administration, pipeline authoring (Declarative/Scripted), build failure investigation
– Docker — image builds, multi-platform (linux/amd64), container registry (ECR)
– Infrastructure-as-code mindset; comfortable reading and modifying YAML-heavy configs
Python (Mandatory)
– Strong Python 3.10+ proficiency — type hints, pydantic v2, async patterns, httpx
– Experience building or maintaining REST APIs (FastAPI or similar)
– Comfortable writing scripts and tools that integrate with AWS (boto3), Kubernetes (kubernetes client), and internal REST APIs
– Able to read, extend, and debug an existing Python codebase without full prior context
Systems Integration
– REST API integration with enterprise tooling (Jira, Confluence, Jenkins, or similar)
– Authentication patterns: Bearer tokens, Basic Auth, AWS IRSA, cross-account IAM
General
– Strong incident investigation and debugging skills across distributed systems
– Able to work independently, manage priorities, and communicate blockers early
– PostgreSQL — basic query and operational familiarity
—
Nice to Have
– LangChain / LangGraph — multi-agent graphs, ReAct pattern, stateful graph checkpointing
– Any LLM API experience (Anthropic, OpenAI, or similar) — tool use and prompt construction
– Experience contributing to AI-assisted operations or internal developer tooling
—
What You Don’t Need
– ML or data science background
– Prior experience with LangGraph or AI agent frameworks — willingness to learn from existing code is sufficient
– Full-stack frontend experience
私たちが提供するもの
Exciting Projects: We focus on industries like High-Tech, communication, media, healthcare, retail and telecom. Our customer list is full of fantastic global brands and leaders who love what we build for them.
Collaborative Environment: You Can expand your skills by collaborating with a diverse team of highly talented people in an open, laidback environment — or even abroad in one of our global centers or client facilities!
Work-Life Balance: GlobalLogic prioritizes work-life balance, which is why we offer flexible work schedules, opportunities to work from home, and paid time off and holidays.
Professional Development: Our dedicated Learning & Development team regularly organizes Communication skills training(GL Vantage, Toast Master),Stress Management program, professional certifications, and technical and soft skill trainings.
Excellent Benefits: We provide our employees with competitive salaries, family medical insurance, Group Term Life Insurance, Group Personal Accident Insurance , NPS(National Pension Scheme ), Periodic health awareness program, extended maternity leave, annual performance bonuses, and referral bonuses.
Fun Perks: We want you to love where you work, which is why we host sports events, cultural activities, offer food on subsidies rates, Corporate parties. Our vibrant offices also include dedicated GL Zones, rooftop decks and GL Club where you can drink coffee or tea with your colleagues over a game of table and offer discounts for popular stores and restaurants!
GlobalLogicについて
GlobalLogic, a Hitachi Group Company, is a trusted digital engineering partner to the world’s largest and most forward-thinking companies. Since 2000, we’ve been at the forefront of the digital revolution – helping create some of the most innovative and widely used digital products and experiences. Today we continue to collaborate with clients in transforming businesses and redefining industries through intelligent products, platforms, and services.


