Senior Data Engineer
Pie Insurance · United States
Johns Hopkins University · Baltimore, Maryland, United States
What they do, how they make money, why this role exists.
Reading up on Johns Hopkins University…
One proof-of-work idea, an hour to build, that shows this employer you get their business.
Sketching a proof-of-work idea…
Who likely hires for this role, a LinkedIn search pre-filled to find them, and a short intro written for this exact job — you send it yourself.
The Library Systems department of Sheridan Libraries, Archives, and Museums is seeking a Data Engineer II. This position has primary responsibility for data extract, transform, and load (ETL) processes across all Library operations except cataloging and collection management. The role is responsible for ensuring data capture workflows, translation processes, deposit methods, and quality assurance protocols meet the requirements of the Libraries.
Processed data and data systems support data access and reuse, analytics and assessment, strategic planning, marketing and communications, operational efficiency improvements, and metadata enrichment and crosswalks. The data engineer will follow an iterative design and development approach with continual review cycle conducted in collaboration with stakeholders and Library leadership.
The Data Engineer II will also design, produce, and maintain software infrastructure to automatically extract, transform, and load data from diverse sources. Additionally, this role will input/output data from databases, realize complex queries, and be responsible for the creation of scripts for data manipulation, cleaning, filtering, and preparation of data outputs to be used for data visualization or web display.
This work will also include creating systems to monitor and guarantee data quality, perform daily data quality assurance tasks, maintain, and troubleshoot the infrastructure, and audit data sources. Lastly, the data engineer will also contribute to exploration and analysis of data to answer policy-related questions as needed.
Collaborate with Business Analysts to manage roadmap for ETL and reporting projects.
Serve as team lead for ETL and data exchange software development life cycle.
Lead the development of tools and process to capture, normalize, and make available research and library collections data for study and research through the central data store and AI-ready data platform being developed by Library IT.
Support library catalog and institutional repository systems through work with bibliographic metadata, including quality assurance, large-scale enrichment and corrections, and creating crosswalks to exchange or migrate metadata between systems that use different metadata schemas.
Documentation of data pipelines for both developers and non-technical stakeholders following best practices for technical writing and systems diagrams.
Contribute to the design, production, and maintenance of data pipelines for data acquisition, management, transformation, and back-end code development to power data web applications and convert raw data into usable information.
Write and maintain ETL/ELTs that operate on a variety of structured and unstructured sources.
Develop and maintain web data scraping systems for automatic data acquisition.
Help design data architecture and provide ongoing support.
Input/output data from databases and perform queries.
Create scripts to clean, transform, and analyze data.
Put into production data pipelines using data warehousing systems.
Create and implement production software to monitor data quality and detect data anomalies.
Perform daily manual data quality assurance tasks.
Support, maintain, and troubleshoot the software infrastructure.
Source data, conduct analyses, visualize data, and generate insights to support ongoing research projects and other requests across the organization.
Collaborate with developers, analysts, data scientists, researchers, policy experts, and other partners.
Communicate with division leadership, and others on the team.
Collaborate with external partners, contractors, and vendors.
Other duties as assigned.
In addition to duties & responsibilities above Collaborate with Business Analysts to manage roadmap for ETL and reporting projects.
Serve as team lead for ETL and data exchange software development life cycle.
Lead the development of tools and process to capture, normalize, and make available research and library collections data for study and research through the central data store and AI-ready data platform being developed by Library IT.
Support library catalog and institutional repository systems through work with bibliographic metadata, including quality assurance, large-scale enrichment and corrections, and creating crosswalks to exchange or migrate metadata between systems that use different metadata schemas.
Documentation of data pipelines for both developers and non-technical stakeholders following best practices for technical writing and systems diagrams.
Minimum Qualifications
Bachelor’s Degree.
Five years of related work experience focused within database management and design and business requirement gathering.
Additional education may substitute for required experience, and additional related experience may substitute for required education beyond a high school diploma/graduation equivalent, to the extent permitted by the JHU equivalency formula.
Proficiency in working with bibliographic metadata standards including MARC and Dublin Core.
Familiarity with Business Intelligence systems and tools such as Power BI and Tableau.
Experience with AWS and Microsoft Azure data utilities Experience designing and implementing ETL processes using stored procedures in an enterprise SQL Server environment. Experience performing data normalization following traditional and dimensional data modelling.
Experience developing software and scripts (using SQL, Python, or other relevant languages) to automate ETL and other data analysis and manipulation tasks.
Demonstrated ability and willingness to learn, adopt, and apply emerging technologies, including AI-enabled tools, in support of professional responsibilities.
Technical Skills & Expected Level of Proficiency
Data Management and Analysis - Intermediate
Data Pipeline Architecture & ETL/ELT Development - Intermediate
Database Querying - Intermediate
Data Validation and Quality Assurance - Intermediate
Data Warehousing & Architecture - Intermediate Oral and written communications: Intermediate
Programming Languages - Intermediate
Web Scraping and Data Acquisition - Intermediate
The core technical skills listed are most essential; additional technical skills may be required based on specific division or department needs.
Classified Title: Data Engineer II
Role/Level/Range: ATP/04/PG
Starting Salary Range: $102,295 - $140,835 - $179,375 Annually (Commensurate w/exp.)
Employee group: Full Time
Schedule: Mon-Fri; 8:30am-5pm
FLSA Status: Exempt
Location: Remote
Department name: Library Systems
Personnel area: Libraries
Straight from the source — this role comes from Johns Hopkins University's own hiring system, not a scraped repost. iRocket links you directly; we don't host the posting or the application.
Pie Insurance · United States
Claritas Rx · United States
Livefront · United States