03 / CLOUD PLATFORMS AND RELIABILITY
Cloud Engineer
Build the conditions that let great software keep working.
STEP INSIDE THE WORK
AN ILLUSTRATIVE PROJECTThe traffic spike is coming. Make the system ready.
An AI application slows down when a campaign launches. What does the system need you to discover?
YOUR MOVE / 01
Understand the bottleneck.
Reproduce the load in a test environment. Follow logs, metrics and traces to see where requests spend time or fail. Understand the workload before adding capacity.
WHAT YOU MAKE VISIBLE
A reproducible load pattern
Latency and error evidence
The failing dependency or resource
Think from first principles.
Make your reasoning visible.
01 / YOUR MISSION
Make the system work when the real world shows up.
The application works locally. Then traffic rises, a dependency fails, or a configuration changes. Cloud engineering starts with understanding what happens next. At devx, you will help build the environments that let software and AI services run reliably. If you enjoy Linux, networks, automation and the satisfaction of finding a root cause, you can turn that curiosity into systems other people depend on.
The cloud environment, deployment automation and operational behaviour around the workload. You collaborate with application engineers on failures that cross infrastructure and service boundaries.
We reimagine customer interactions, business operations and enterprise architecture around AI. You bring fresh questions. Experienced practitioners bring context and judgment. The work connects both to a real customer outcome.
02 / WHAT YOU WILL WORK ON
Make your curiosity
useful.
Build reproducible environments
Deploy scoped services in development and test environments on the project's cloud platform. Learn how compute, storage, networking and managed services work together.
Automate the repeatable work
Write scripts and contribute to reviewed infrastructure-as-code and deployment pipelines. Make configuration changes traceable, repeatable and easier for another engineer to understand.
Make system behaviour visible
Use logs, metrics and traces to investigate latency, errors and resource use. Help create actionable alerts tied to user experience. For AI services, work with application engineers to understand external model limits, timeouts and capacity constraints.
Make changes safely
Apply least-privilege access and approved secret storage. Contribute reviewed network and deployment changes, and test backup restoration and rollback procedures in appropriate environments. Record recovery steps so another engineer can use them.
Improve reliability and cost together
Investigate incidents with experienced engineers, document causes and contribute fixes. Review idle resources and capacity choices while preserving the workload's performance and availability needs.
What good work looks like
- An environment that can be recreated from reviewed configuration.
- A reliability improvement supported by logs, load evidence and a cost comparison.
- A deployment or recovery procedure that a teammate can follow and verify.
Use AI to help explain unfamiliar logs, draft scripts and compare configuration options. Check suggestions against platform documentation, inspect permissions and changes, and test in a controlled environment before review.
03 / YOUR STARTING POINT
Bring a foundation.
Build the range.
Coursework, personal projects, research and student initiatives all count. Previous full-time experience is not required.
- Linux and operating-system basics, including processes, files, permissions and the command line.
- Networking fundamentals: IP, DNS, HTTP, ports and the path a request takes from a user to a service.
- Basic Python or shell scripting, Git, and a systematic approach to debugging from observed evidence.
- A lab, coursework or personal project where you deployed, configured or investigated a running service.
Useful exposure
One of AWS, Azure or Google Cloud; Docker, CI/CD, Terraform or monitoring tools. Certifications, multi-cloud expertise and production Kubernetes experience are optional, not entry requirements.
04 / HOW YOUR OWNERSHIP GROWS
Learn in the work.
Grow through the feedback.
Start with reproducible deployments and diagnostic tasks in reviewed environments. Progress towards owning a bounded automation or reliability improvement, including its documentation. You will build foundations for cloud engineering, platform engineering and site reliability; production changes and incident decisions involve experienced reviewers.
Explore
Understand the problem and ask useful questions.
Build
Make a defined contribution and learn through review.
Own
Take on broader responsibility as your readiness grows.
05 / THE DEVX CULTURE
The principles
show up in the work.
Understand what the customer is trying to change. Ask the extra question, surface the real obstacle and connect your work to their goals.
A feature shipped, but the workflow is still slow. Stay with the problem and find out why.
Client obsession. Understand what the customer is trying to change. Ask the extra question, surface the real obstacle and connect your work to their goals.
AI-native everything. Use AI to improve how you research, build, analyse and document. Understand its limits, verify the output and handle customer information responsibly.
Hire and train exceptional talent. We bring young talent and experienced practitioners together. Bring ambition, seek direct feedback and put new understanding into practice.
Document to scale. Write down the decision, its context and what you learned. Make your work understandable enough for someone else to continue, question or improve it.
Only ever be honest. Flag risk while there is time to act. Separate what you know from what you assume. Give clear feedback and make uncertainty visible.
06 / START A CONVERSATION
Show us the way
you think.
Share a résumé and a lab or project note. Include the setup, a failure you investigated, the evidence you followed, and how someone else could reproduce your result.
Join through an internship with the opportunity to convert to a full-time role. The hiring team will share internship duration, work arrangements, campus eligibility and conversion criteria during recruitment.
An illustrative project
An AI-enabled application slows down during a campaign. In a test environment, you could reproduce the load, trace the bottleneck and compare a configuration or capacity change. Document response time, error rate and estimated cost before and after. Include a rollback plan so the team can review the improvement with confidence.