Back to all roles

04 / DATA FOUNDATIONS FOR AI AND BUSINESS DECISIONS

Data Engineer

Make the information behind the decision worth trusting.

INTERNSHIP → FULL-TIME OPPORTUNITYGRADUATE ENTRYFIRST-PRINCIPLES THINKING

STEP INSIDE THE WORK

AN ILLUSTRATIVE PROJECT

Two dashboards. Two answers. One truth to find.

Orders, cancellations and refunds arrive at different times. How do you build a daily view people can trust?

YOUR MOVE / 01

Define what the data means.

Trace the sources and define what one row represents. Agree keys, time windows and how cancellations or late refunds should be treated. A pipeline begins with business meaning.

WHAT YOU MAKE VISIBLE

1

Clear grain and keys

2

Source-to-target mapping

3

Rules for late events

Think from first principles.
Make your reasoning visible.

01 / YOUR MISSION

Give every ambitious AI idea a dependable starting point.

One dashboard says sales rose. Another says they fell. The model sees duplicated orders. Somebody needs to follow the data all the way back. At devx, you will help turn fragmented information into data that people and AI systems can use with confidence. This role rewards curiosity about how businesses work and precision about what a number, record or event actually means.

YOUR FOCUS

Dependable ingestion, transformation, data models and quality checks. You make the data usable and explainable so the AI, application and customer teams can rely on it.

WHY DEVX

We reimagine customer interactions, business operations and enterprise architecture around AI. You bring fresh questions. Experienced practitioners bring context and judgment. The work connects both to a real customer outcome.

02 / WHAT YOU WILL WORK ON

Make your curiosity
useful.

01

Bring source data into usable form

Build scoped ingestion and transformation tasks using SQL and Python. Work with files, APIs or databases, understand source behaviour, and document the intended destination and refresh pattern.

02

Model the business clearly

Define the meaning of a row, keys, relationships and important measures with the team. Agree time zones and treatment of late or corrected records. Create understandable datasets for operational workflows, analysis and AI use cases.

03

Make data quality visible

Check missing values, duplicates, invalid records and unexpected volume changes. Reconcile important totals to source data and investigate differences before a downstream team relies on them.

04

Make pipelines recoverable

Contribute scheduling, incremental loads and backfills within a reviewed design. Make reruns idempotent: processing the same input again should leave the intended result unchanged. Handle partial failures and source corrections without silently duplicating or losing records.

05

Build trust across the handoff

Document lineage, transformations and access expectations. Work with AI, application and consulting colleagues so they understand the dataset's meaning, freshness and known limitations.

What good work looks like

  • A dataset with clear grain, keys, lineage and freshness expectations.
  • Quality checks that catch important errors and reconcile results to the source.
  • A pipeline that can recover from a failed run without corrupting the result.
AI IN YOUR OWN WORK

Use AI to explore unfamiliar schemas, draft transformations and suggest edge cases. Inspect every generated query, test joins and totals against known examples, and keep sensitive data within approved tools and access boundaries.

03 / YOUR STARTING POINT

Bring a foundation.
Build the range.

Coursework, personal projects, research and student initiatives all count. Previous full-time experience is not required.

  • SQL fundamentals, including joins, grouping, filtering and reasoning about how a query changes the result.
  • Python for data manipulation, basic programming and debugging; comfort reading common file formats.
  • Relational database concepts, keys and data types, with attention to duplicate, missing and inconsistent data.
  • A coursework or personal project in which you collected, cleaned, transformed or modelled data and checked its correctness.

Useful exposure

Window functions, pandas, a cloud warehouse, dbt, orchestration or Spark. These can be learned progressively; a clear SQL project is more useful than an unexplained list of platforms.

04 / HOW YOUR OWNERSHIP GROWS

Learn in the work.
Grow through the feedback.

Start with a well-defined dataset or pipeline component and reviewed changes. Grow towards owning its quality checks, refresh behaviour and downstream contract. Develop data engineering judgment by connecting business meaning to storage, modelling and processing choices, and learning when distributed tools are actually needed.

01

Explore

Understand the problem and ask useful questions.

02

Build

Make a defined contribution and learn through review.

03

Own

Take on broader responsibility as your readiness grows.

05 / THE DEVX CULTURE

The principles
show up in the work.

Understand what the customer is trying to change. Ask the extra question, surface the real obstacle and connect your work to their goals.

A feature shipped, but the workflow is still slow. Stay with the problem and find out why.

06 / START A CONVERSATION

Show us the way
you think.

Share a résumé and one data project. Explain the source, what one row means, a quality issue you found, your checks, and how you would recover from a failed run.

Join through an internship with the opportunity to convert to a full-time role. The hiring team will share internship duration, work arrangements, campus eligibility and conversion criteria during recruitment.

KEEP EXPLORING

Outcome Manager

Keep the team connected to the result the customer needs.