Data Scientist
We are seeking an experienced Data Scientist to join Attercop's Applied Science and Engineering team to design, build, and evaluate AI and machine learning solutions across the full project lifecycle. The majority of our client problems involve unstructured text (extracting structure from documents, retrieving the right information from large and messy corpora, and applying generative models reliably enough to put in front of a client's users), so we are looking for genuine depth in natural language processing and machine learning. You will own model and system design, experimentation, and evaluation, while productionisation and infrastructure sit with our AI Engineers, with whom you'll work closely.
Core Responsibilities
1. Applied Machine Learning
- •Problem Framing: Translate ambiguous client requirements into well-posed machine learning problems with defined success criteria and realistic scope.
- •Modelling: Select, implement, and tune models appropriate to the problem and the data available, resisting unnecessary complexity where a simpler approach performs.
- •Statistical Rigour: Apply sound experimental design, validation strategy, and error analysis, including investigation of bias, data leakage, and distribution shift.
- •Data Exploration and Feature Engineering: Conduct thorough exploration of structured and unstructured data, and perform the cleaning, transformation, and feature construction needed to make it modellable.
- •2. NLP and Language Technologies
- •Model Development: Design, fine-tune, and evaluate NLP models for tasks including classification, information extraction, entity recognition, summarisation, and semantic similarity.
- •Information Retrieval: Build and evaluate retrieval systems spanning sparse, dense, and hybrid approaches, including embedding model selection, chunking strategy, reranking, and query understanding.
- •Generative AI: Apply large language models to client problems through prompting, structured output, RAG architectures, and fine-tuning where warranted, with a clear understanding of each failure mode.
- •Evaluation Design: Define task-appropriate metrics and build evaluation harnesses for systems where ground truth is expensive, ambiguous, or subjective.
- •3. Knowledge Representation and Graph-Based Approaches (desirable)
- •Ontology and Schema Design: Contribute to the taxonomies, ontologies, and graph schemas used to structure client domain knowledge.
- •Graph Construction from Text: Apply entity recognition, entity linking, and relation extraction to build knowledge graphs from unstructured sources.
- •Graph-Augmented Retrieval and Reasoning: Explore graph-based approaches to retrieval and reasoning where relational structure adds value beyond vector search alone.
- •4. Delivery and Engineering Quality
- •Production-Ready Code: Write clean, well-tested, maintainable Python in line with Attercop's coding standards, contributing to shared codebases and reusable components.
- •Hand-Off and Integration: Work closely with AI Engineers to move models and pipelines into client systems, providing the documentation, tests, and interfaces required for a clean handover.
- •Tooling: Assist in evaluating and adopting new tools, frameworks, and techniques that enhance the team's capabilities and the quality of client deliverables.
- •Ownership: Take responsibility for the quality of your own output, seek peer review proactively, and escalate technical risks or blockers promptly and constructively.
- •5. Research and Innovation
- •Horizon Scanning: Stay current with developments in NLP, information retrieval, and generative AI, bringing relevant ideas and techniques to the attention of the wider team.
- •Prototyping: Contribute to internal R&D efforts, prototyping new approaches and helping to assess their viability for client applications.
- •Documentation: Record research findings, methodologies, and implementation details clearly, sharing knowledge with colleagues and, where appropriate, external audiences.
Candidate Requirements
Professional Experience
- •We'd expect either:
- •A minimum of 3 years of professional experience as a Data Scientist or Research Scientist with a substantial proportion of that time spent on NLP or language-centric problems in a commercial or applied research setting, alongside a degree (or equivalent demonstrable experience) in a quantitative discipline such as computer science, mathematics, statistics, physics, or engineering; or
- •A PhD in one of those disciplines with a research focus on NLP or machine learning.
Technical Proficiencies
- •Machine Learning (essential):
- •Solid grounding in machine learning fundamentals, model evaluation, and experimental design.
- •Demonstrable experience building and validating models on real, imperfect data
- •Working knowledge of scikit-learn and PyTorch, and comfort reading and adapting research code.
- •Natural Language Processing (essential):
- •Strong, demonstrable background in NLP, including hands-on work with transformer architectures, embeddings, and modern pre-trained models.
- •Proficiency with the current NLP toolchain, such as Hugging Face Transformers, spaCy, and sentence-transformers.
- •Practical experience with information retrieval and/or RAG systems, including how to measure and improve retrieval quality.
- •Demonstrated experience applying LLMs to real problems, with a clear-eyed view of where they work, where they fail, and how to tell the difference.
- •Python and Software Engineering (essential):
- •Strong proficiency in Python, including the core data stack (pandas, NumPy) and an understanding of how to structure code beyond a notebook.
- •Good software engineering practices, including version control (Git), automated testing, and code review.
- •Experience working with structured and unstructured data at scale, including cleaning, transformation, and pipeline construction.
- •Knowledge Graphs and Semantic Technologies (desirable):
- •Experience with knowledge graph construction or processing, ideally derived from text.
- •Familiarity with graph databases and query languages (e.g. Neo4j / Cypher, RDF / SPARQL) and with ontology or taxonomy design.
- •Awareness of graph-augmented retrieval approaches such as GraphRAG.
- •Cloud, Data, and MLOps (desirable):
- •Familiarity with cloud platforms (Azure preferred, or AWS/GCP) and with MLOps tooling for model deployment, versioning, and monitoring.
- •Experience with vector databases and semantic search infrastructure.
- •Experience with data engineering concepts, including pipelines, orchestration, and warehousing.
- •Wider Experience (desirable):
- •Experience working in a consultancy or client-facing environment.
- •Contributions to open-source projects, published research, or public writing on NLP or data science topics.
- •Familiarity with agentic frameworks such as LangChain and LangGraph.
- •Strategic and Collaborative Competencies
- •Cross-Functional Collaboration: Work effectively within cross-functional project teams, supporting the Lead Data Scientist in planning, scoping, and delivering technical workstreams alongside AI Engineers, consultants, and client stakeholders.
- •Technical Communication: Articulate complex technical concepts, model limitations, and performance metrics clearly to both technical peers and non-technical leadership, through presentations, written reports, and informal discussion.
- •Client Engagement: Represent Attercop professionally in client meetings and workshops, contributing technical input and building positive working relationships. Support the pre-sales process where required, providing technical input to proposals and helping to scope prospective engagements.
- •Knowledge Sharing: Participate actively in code reviews, knowledge-sharing sessions, and team retrospectives, contributing to a culture of continuous learning and improvement.
- •Documentation: Commitment to maintaining clear documentation for methodologies, experiments, evaluation results, and handover interfaces.
Important Information
English Language Requirement
All roles require excellent English. We work entirely in English for meetings, client calls, and business communications. This is non-negotiable.
No Recruitment Agencies
We do not work with recruitment agencies. Please do not contact us if you're representing candidates. We hire directly.
Ready to Apply?
Send us your CV and a brief note about what you do and why you're interested in joining Attercop.