Course Outline
Introduction
- Understanding the importance of data preparation in analytics and machine learning
- Data preparation pipeline and its role in the data lifecycle
- Exploring common challenges in raw data and the impact on analysis
Data Collection and Acquisition
- Sources of data: databases, APIs, spreadsheets, text files, and more
- Techniques for collecting data and ensuring data quality during collection
- Collecting data from various sources
Data Cleaning Techniques
- Identifying and handling missing values, outliers, and inconsistencies
- Dealing with duplicates and errors in the dataset
- Cleaning real-world datasets
Data Transformation and Standardization
- Data normalization and standardization techniques
- Categorical data handling: encoding, binning, and feature engineering
- Transforming raw data into usable formats
Data Integration and Aggregation
- Merging and combining datasets from different sources
- Resolving data conflicts and aligning data types
- Techniques for data aggregation and consolidation
Data Quality Assurance
- Methods for ensuring data quality and integrity throughout the process
- Implementing quality checks and validation procedures
- Case studies and practical applications of data quality assurance
Dimensionality Reduction and Feature Selection
- Understanding the need for dimensionality reduction
- Techniques like PCA, feature selection, and reduction strategies
- Implementing dimensionality reduction techniques
Summary and Next Steps
Requirements
- Basic understanding of data concepts
Audience
- Data analysts
- Database administrators
- IT professionals
Delivery Options
Private Group Training
Our identity is rooted in delivering exactly what our clients need.
- Pre-course call with your trainer
- Customisation of the learning experience to achieve your goals -
- Bespoke outlines
- Practical hands-on exercises containing data / scenarios recognisable to the learners
- Training scheduled on a date of your choice
- Delivered online, onsite/classroom or hybrid by experts sharing real world experience
Private Group Prices RRP from £3800 online delivery, based on a group of 2 delegates, £1200 per additional delegate (excludes any certification / exam costs). We recommend a maximum group size of 12 for most learning events.
Contact us for an exact quote and to hear our latest promotions
Public Training
Please see our public courses
Testimonials (3)
The ability to Engauge on a 1:1 basis and ensure I had clarity and understanding on the concepts discussed.
Dave - Sea
Course - Data Architecture Fundamentals
It's a hands-on session.
Vorraluck Sarechuer - Total Access Communication Public Company Limited (dtac)
Course - Talend Open Studio for ESB
I generally enjoyed the knowledge of the trainer.