Resources · LessonBy Neelima, SailPoint Architect · Published · IdentityIQ 8.4 · all levels
Datasets
Sample identity and account datasets for practice.
Quick answer
Synthetic datasets of identities, accounts and entitlements you can load to practice aggregation, correlation and certifications.
Key takeaways
- Synthetic identities and accounts
- Safe, non-real data
- Practice aggregation/correlation
- Test roles and certifications
These synthetic datasets of identities, accounts and entitlements let you populate a sandbox with realistic-looking but entirely fake data, so you can practise the full governance loop, aggregation, correlation, roles, certifications, without touching real systems or real personal data.
What you get
- Synthetic identity records
- Sample accounts and entitlements
- Data safe for training and testing
- Enough volume to exercise real workflows
How to use them
Load a dataset as a source, aggregate it, correlate the accounts, then build roles and run a certification against it to experience the whole flow safely. Ideal for training and for testing customisations.
Common pitfalls
- Mixing synthetic and real data in the same environment.
- Assuming synthetic data matches your real data quality.
- Using tiny datasets that hide performance behaviour.
Practice challenge
+0 XPStreak ×0
Question 1 of 3
What are the datasets?
Frequently asked questions
What does the term Datasets refer to in SailPoint?
These synthetic datasets of identities, accounts and entitlements let you populate a sandbox with realistic-looking but entirely fake data, so you can practise the full governance loop, aggregation, correlation, roles, certifications, without touching real systems or real personal data.
What is the role of aggregation in Datasets?
Load a dataset as a source, aggregate it, correlate the accounts, then build roles and run a certification against it to experience the whole flow safely.
What tends to go wrong with Datasets?
Mixing synthetic and real data in the same environment. Assuming synthetic data matches your real data quality. Using tiny datasets that hide performance behaviour.
Want this with a live instructor and a lab tenant?