AI & Data Engineer (Intermediate)
NextGenMaint
Ibadan, Nigeria · Full-time
₦750,000
We are hiring an intermediate AI & Data Engineer to build the data and machine learning layer behind our product. Most of the data we receive is messy. It arrives as spreadsheets, CSV exports and system dumps produced by teams who never expected anyone else to read them: inconsistent columns, duplicate records that are not quite duplicates, missing values, and the same thing named five different ways. Your job is to turn that into something clean and queryable, and then to build the matching and enrichment that makes it genuinely useful. What you will do Build ingestion and cleaning pipelines. Design workflows that take raw files and produce validated, standardised data — schema checks, sensible handling of missing values, and detection of the outliers that signal a broken export rather than an unusual record. Build matching and deduplication. Work out when two records describe the same real thing, using a mix of exact rules, fuzzy text matching and vector similarity. This is the core of the product and the part that is genuinely difficult. Write resilient parsers. Handle awkward real-world formats: nested structures, inconsistent delimiters, broken encodings, and files that change shape without warning. Apply language models where they earn their place. Use LLM APIs for the tasks that rules handle badly — standardising descriptions, filling gaps, classifying free text — with structured outputs and sensible checks. We are interested in engineers who can tell the difference between a problem that needs a model and one that needs a lookup table. Measure what you build. Define what good looks like for accuracy and speed, track it honestly, and iterate on the basis of real results rather than a hunch. Where humans review output, close that loop and learn from it. Build reusable components. Package your work as services with clear interfaces, versioned schemas and useful logging, so the next problem starts from something rather than nothing. Why join You will own the ful
Requirements
2-4 years in data engineering, machine learning engineering, or a similar hands-on role.,Strong Python, including confident use of a dataframe library (Pandas or Polars) on data that does not fit neatly.,Strong SQL — not just SELECT, but the joins, window functions and query tuning that large transformations need.,Practical experience with language model APIs in something that shipped, including an honest sense of cost, latency and where models are the wrong tool.,Experience with embeddings and similarity search for matching or classification, and comfort with the trade-offs between approaches.,The ability to turn exploratory notebook code into tested, deployable services under version control.,Comfort with containers and reproducible environments.,A bias toward shipping, balanced by enough architectural judgement that what you ship survives contact with more data.,Nice to have,Experience with vector databases and indexing strategies. Familiarity with data from operations, maintenance, inventory or supply chain systems, or with parsing exports from large enterprise systems. Experience tracking experiments and monitoring models in production, including detecting when the data has drifted underneath a model that used to work.