Early Career Health Data Scientist , BHF Data Science Centre
Health Data Research UK
- Closing: 5:00pm, 23rd Oct 2026 BST
Perks and benefits
Flexible working hours
Work from home option
Wellness programs
Employee Assistance Programme
Enhanced maternity and paternity leave
Extra holiday
Professional development
Mentoring/coaching
Team social events
Candidate happiness
8.42 (9174)
8.42 (9174)
Job Description
Purpose of the post
The Early Career Health Data Scientist will join our Health Data Science team and will contribute to the development of scalable, reusable, and well-documented resources that help researchers efficiently curate high-quality, analysis-ready data for their research projects. The post-holder will work with senior colleagues and researchers to identify common data curation challenges, develop solutions that can be reused across projects, and share good practice through code, documentation and guidance. These resources may include:
Data dictionaries, dataset summaries and shared exploratory analyses and insights that help researchers understand the characteristics, quality and limitations of datasets and how they can appropriately be used on research projects.
Coding tutorials, guidance notes, and worked examples to help researchers develop the technical skills needed to curate data within Secure Data Environments.
Reusable code, functions, and data curation pipelines that researchers can adapt for their own projects, reducing duplication and accelerating the data curation phase of their project
Curated data methods that produce cleaned and enhanced views of datasets and can be integrated into data curation pipelines, reducing the need to repeatedly implement equivalent data processing and validation logic across projects.
The post-holder will provide direct, hands-on support to researchers, either by providing guidance and sign-posting to existing data curation resources relevant to their project, or by providing targeted, bespoke solutions where appropriate. This may include helping researchers identify suitable datasets and existing curation resources, adapting reusable code and pipelines, and developing project-specific data curation approaches to produce robust, analysis-ready data.
The post-holder will also undertake data quality and exploratory analyses to help assess the quality, completeness and utility of datasets, identify potential issues and limitations, and understand how data can be appropriately used to address specific research questions. Findings from this work will contribute to shared knowledge and resources that can benefit researchers across programmes.
This post is an attractive career development opportunity, which would suit a health data scientist, data analyst, data engineer who can demonstrate previous experience of data wrangling and curation of health data for research projects, and who wishes to expand and deepen their expertise in large-scale, linked health data, and collaborative research environments. Due to the high number of applications in previous recruitment rounds, we would encourage applications only from candidates with relevant experience of working with linked administrative healthcare data in a professional or research setting who are looking to build on this experience.
Main responsibilities
Providing data engineering and data curation support in secure data environments (SDEs) and trusted research environments (TREs) to produce robust, analysis-ready datasets.
Contributing to the development, testing, and maintenance of reproducible and well-documented data curation pipelines and shared resources under the supervision of senior colleagues.
Developing and applying expertise in the assessment of data quality, completeness, and data utility of the various routinely collected health datasets across the four devolved nations, including contributing to early feasibility and exploratory assessments to inform study design. Summarising and disseminating findings and lessons from data quality and data utility assessments to inform research design and appropriate use of routinely collected data.
Under the supervision of senior colleagues, writing, organising and maintaining support documentation for linked data resources (e.g. data dictionaries, variable mapping tables, data access process documentation, and Git repositories).
Carrying out technical validation checks on linked data sources (e.g. duplicates, linkage errors, temporal inconsistencies) and developing reusable functions to check these data rigorously for errors and inconsistencies.
Working with relevant researchers to identify and apply appropriate existing and novel phenotype definitions and algorithms from linked national health data.
Preparing clear numerical summaries and visualisations to communicate findings (e.g. data characteristics, quality, and implications for research and decision making) to researchers when curating data.
Preparing and presenting results in oral and written reports, technical notes, and contributing to academic publications.
Actively participating and attending regular Centre and project meetings, reporting on progress and presenting analytical results.
Demonstrating a strong commitment to open source, transparent, and reproducible research, as the post will involve releasing tools, code, documentation under an open-source licence.
Contributing to the sharing of knowledge and good practice within the Health Data Science Team and wider research community, including through presentations, demonstrations, guidance and training materials.
Developing technical and methodological expertise through training, mentoring and hands-on experience, and taking increasing responsibility for defined areas of work as experience develops.
Planning and organizing
The post-holder will join a small team that is embedded within HDR UK and will be responsible for planning and day-to-day management of their own workload and prioritising tasks across multiple projects. At the same time, the post-holder will require a flexible approach to work to changing demands, particularly external changes.
Problem solving
The post-holder will bring existing technical expertise and domain knowledge, alongside the expectation to develop a deep understanding of the data quality and utility of national routinely collected coded health data and its suitability for different types of research.
The post-holder will apply engineering and analytic approaches to complex health-related data requiring prior technical data science knowledge. The post-holder will also need to be able to resolve non-trivial data curation and analysis challenges, seeking input as required from other members of the team and external colleagues. The post-holder will make an effective judgement on when to escalate issues to senior colleagues’ attention and with what urgency.
Decision making
In collaboration with senior colleagues, the post-holder will make decisions about the most appropriate tools, technologies, and approaches for querying, processing, analysing, and documenting complex health-related data.
Key contacts/relationships
The post holder will work in close conjunction with the core BHF Data Science Centre team but primarily within the BHF Data Science Centre’s Health Data Science team. They will also work with external researchers, analysts, data managers, data custodians, clinicians, health data scientists, and epidemiologists across a number of programmes and projects supported by the Centre.
They will build and maintain effective working relationships across multiple HDR UK teams, partners within the British Heart Foundation, the wider cardiovascular and health data science communities, and other key stakeholders. The role may also contribute, where appropriate, to activities that support patient and public involvement and transparency in data-enabled research.
Removing bias from the hiring process
Removing bias from the hiring process
- Your application will be anonymously reviewed by our hiring team to ensure fairness
- You’ll need a CV/résumé, but it’ll only be considered if you score well on the anonymous review
Share
Share
Email
Share