Managing Consultant Data Scientist
Excella · Arlington, Virginia, United States
Spokeo · Pasadena, California, United States
What they do, how they make money, why this role exists.
Reading up on Spokeo…
One proof-of-work idea, an hour to build, that shows this employer you get their business.
Sketching a proof-of-work idea…
Who likely hires for this role, a LinkedIn search pre-filled to find them, and a short intro written for this exact job — you send it yourself.
Join our mission to make the world more transparent with data.
Spokeo is a people intelligence platform that helps over 18 million monthly visitors reconnect with friends, reunite with families, and build trust in new relationships. Thousands of companies also trust Spokeo’s 60 billion public records to improve customer research, help verify information, and prevent fraud.
Founded in 2006, Spokeo has built a dedicated, remote-first team with an average tenure of 6.9 years. It has earned recognition from Comparably as a Best Company for Compensation, Employee Happiness, Perks and Benefits, Support for Women, Work-Life Balance, and CEO Leadership.
As a Senior Data Engineer at Spokeo, you will develop, optimize, and improve our data systems, including ETL pipelines, storage, and entity resolution. This involves working with infrastructure built in AWS, including Airflow, PySpark, EMR, S3, DynamoDB, and more. This role will help build and improve data products, automation platform features, analytical software packages, and data pipeline orchestration tools.
Build infrastructure and data automation pipelines to ingest, process, and load data from various sources. Automate and integrate new components into the data pipeline.
Collaborate with stakeholders and data science teams to develop data products, including entity resolution and best selection, to efficiently execute product vision and strategy in alignment with organizational goals and priorities.
Create unit and stress-test components to monitor technical performance and ensure that identified issues are resolved.
Develop data analysis tools to provide data insights and capture key metrics.
Research solutions and maintain technical documentation.
Follow best practices for data governance, quality, cleansing, and other ETL-related activities.
Required: 7+ years of development experience in data engineering within a production environment (internships and academic settings excluded).
Proven experience working with large datasets exceeding 100M+ records or multiple terabytes.
Required: 5+ years of development experience in highly scalable, distributed systems and cluster architectures using AWS and utilizing EMR.
Required: 5+ years of hands-on programming experience with Python.
Required: 5+ years of professional experience working in big data ecosystems; Spark is required; PySpark is preferable.
Required: 5+ years of experience with SQL, schema design, and dimensional data modeling.
Required: 5+ years of professional experience working with dataflow orchestration tools, such as Airflow.
Required: 2+ years of experience with non-relational databases (e.g., DynamoDB, Elasticsearch, etc.).
Required: A bachelor’s degree in Computer Science, Information Systems, Mathematics, or a related field is required.
Spokeo offers a bonus program, equity plans, and a 401 (k). Once a year, we do a discretionary, merit-based salary increase. Additional benefits include 100% medical/dental/vision coverage and unlimited employee PTO.
We extend written offers to candidates who successfully complete their selection process. Offers will depend on several factors, including, but not limited to, marketplace competition, job leveling, experience, and skills.
Straight from the source — this role comes from Spokeo's own hiring system, not a scraped repost. iRocket links you directly; we don't host the posting or the application.
Excella · Arlington, Virginia, United States
VivoSense, Inc. · United States
Airbnb · United States