Back to all roles

03 / CLOUD PLATFORMS AND RELIABILITY

Cloud Engineer

Build the conditions that let great software keep working.

INTERNSHIP → FULL-TIME OPPORTUNITYGRADUATE ENTRYFIRST-PRINCIPLES THINKING

STEP INSIDE THE WORK

AN ILLUSTRATIVE PROJECT

The traffic spike is coming. Make the system ready.

An AI application slows down when a campaign launches. What does the system need you to discover?

YOUR MOVE / 01

Understand the bottleneck.

Reproduce the load in a test environment. Follow logs, metrics and traces to see where requests spend time or fail. Understand the workload before adding capacity.

WHAT YOU MAKE VISIBLE

1

A reproducible load pattern

2

Latency and error evidence

3

The failing dependency or resource

Think from first principles.
Make your reasoning visible.

01 / YOUR MISSION

Make the system work when the real world shows up.

The application works locally. Then traffic rises, a dependency fails, or a configuration changes. Cloud engineering starts with understanding what happens next. At devx, you will help build the environments that let software and AI services run reliably. If you enjoy Linux, networks, automation and the satisfaction of finding a root cause, you can turn that curiosity into systems other people depend on.

YOUR FOCUS

The cloud environment, deployment automation and operational behaviour around the workload. You collaborate with application engineers on failures that cross infrastructure and service boundaries.

WHY DEVX

We reimagine customer interactions, business operations and enterprise architecture around AI. You bring fresh questions. Experienced practitioners bring context and judgment. The work connects both to a real customer outcome.

02 / WHAT YOU WILL WORK ON

Make your curiosity
useful.

01

Build reproducible environments

Deploy scoped services in development and test environments on the project's cloud platform. Learn how compute, storage, networking and managed services work together.

02

Automate the repeatable work

Write scripts and contribute to reviewed infrastructure-as-code and deployment pipelines. Make configuration changes traceable, repeatable and easier for another engineer to understand.

03

Make system behaviour visible

Use logs, metrics and traces to investigate latency, errors and resource use. Help create actionable alerts tied to user experience. For AI services, work with application engineers to understand external model limits, timeouts and capacity constraints.

04

Make changes safely

Apply least-privilege access and approved secret storage. Contribute reviewed network and deployment changes, and test backup restoration and rollback procedures in appropriate environments. Record recovery steps so another engineer can use them.

05

Improve reliability and cost together

Investigate incidents with experienced engineers, document causes and contribute fixes. Review idle resources and capacity choices while preserving the workload's performance and availability needs.

What good work looks like

  • An environment that can be recreated from reviewed configuration.
  • A reliability improvement supported by logs, load evidence and a cost comparison.
  • A deployment or recovery procedure that a teammate can follow and verify.
AI IN YOUR OWN WORK

Use AI to help explain unfamiliar logs, draft scripts and compare configuration options. Check suggestions against platform documentation, inspect permissions and changes, and test in a controlled environment before review.

03 / YOUR STARTING POINT

Bring a foundation.
Build the range.

Coursework, personal projects, research and student initiatives all count. Previous full-time experience is not required.

  • Linux and operating-system basics, including processes, files, permissions and the command line.
  • Networking fundamentals: IP, DNS, HTTP, ports and the path a request takes from a user to a service.
  • Basic Python or shell scripting, Git, and a systematic approach to debugging from observed evidence.
  • A lab, coursework or personal project where you deployed, configured or investigated a running service.

Useful exposure

One of AWS, Azure or Google Cloud; Docker, CI/CD, Terraform or monitoring tools. Certifications, multi-cloud expertise and production Kubernetes experience are optional, not entry requirements.

04 / HOW YOUR OWNERSHIP GROWS

Learn in the work.
Grow through the feedback.

Start with reproducible deployments and diagnostic tasks in reviewed environments. Progress towards owning a bounded automation or reliability improvement, including its documentation. You will build foundations for cloud engineering, platform engineering and site reliability; production changes and incident decisions involve experienced reviewers.

01

Explore

Understand the problem and ask useful questions.

02

Build

Make a defined contribution and learn through review.

03

Own

Take on broader responsibility as your readiness grows.

05 / THE DEVX CULTURE

The principles
show up in the work.

Understand what the customer is trying to change. Ask the extra question, surface the real obstacle and connect your work to their goals.

A feature shipped, but the workflow is still slow. Stay with the problem and find out why.

06 / START A CONVERSATION

Show us the way
you think.

Share a résumé and a lab or project note. Include the setup, a failure you investigated, the evidence you followed, and how someone else could reproduce your result.

Join through an internship with the opportunity to convert to a full-time role. The hiring team will share internship duration, work arrangements, campus eligibility and conversion criteria during recruitment.

KEEP EXPLORING

Data Engineer

Make the information behind the decision worth trusting.