
screenshot taken from: https://datascienceretreat.com/about
When
Tuesday, 21st July 2026, 5:30pm to 9:30pm
Where
SAP SE (BER01), George-Stephenson-Straße 7, Berlin, Germany
Hosting Organization
Data Science Retreat
Participation Fee
Free Entrance
Agenda
Host Intro 1, Host Intro 2, Talk 1, Talk 2, Talk 3, Talk 4, Short Break, Talk 5, Talk 6, Talk 7, Socializing
Topics Covered
Welcoming Words & Data Science at SAP LeanIX (Host Intro 1), Origin Story of Data Science Retreat (Host Intro 2), A Self-Service Web Application for Property Lifecycle Simulation for German Rental Investments (Talk 1), AI Detection Of Mispronunciation On Phoneme-Level For Language Learners (Talk 2), An ML-Powered Communication Intelligence Infrastructure for Navigating Regulatory Compliance Needs (Talk 3), An Intelligent Indoor Air Quality Management System Supporting Smarter And More Sustainable Building Management (Talk 4), Automated Seamless Pattern Generation For Textiles, Design, And 3D Modeling (Talk 5), An AI-Powered Mobile Application For Rapid Identification Of Valuable Stamps From Large Collections (Talk 6), Fine-Tuning An Open-Source Music Generation Model On A Curated Dataset Of Traditional Iranian Music (Talk 7)
I’ve learned something today
- SAP LeanIX is an enterprise architecture platform. It centralizes IT landscape data, automates reporting, and supports transformation, risk, and lifecycle planning. For Enterprise Architects, it helps govern applications, technologies, and roadmaps. For Data Architects, it helps map systems, dependencies, ownership, and key architectural decisions. In the past, and often still today, much of this work has been managed manually in Excel spreadsheets and PowerPoint presentations.
- Metadata scarcity in SAP LeanIX means that Fact Sheet data is often quite sparse. Key attributes, relationships, and lifecycle information are frequently missing because ownership is unclear, automated discovery is limited, and complex meta-model requirements discourage casual users from keeping records up to date.
- Data Science Retreat is a Berlin-based organization and one of Europe’s longest-running data science bootcamps. Its vision is to help close Europe’s AI and data science gap with the United States & China by developing talent and supporting real-world AI adoption.
- WavLM (see huggingface.co/docs/transformers/model_doc/wavlm) is a pretrained speech model that already provides strong audio representations for tasks such as speech recognition, speaker identification, and pronunciation analysis. It can be fine-tuned for a specific use case, and it expects 16 kHz audio so that recordings are processed in the same format used during training.
- DALTON (see ubinet-iitkgp.github.io/ubinet/pages/DALTON) is a large-scale indoor air quality dataset that combines 89.1 million pollution measurements from 30 sites in India with real-time annotations of daily human activities. It enables research on pollution sources, pollutant spread, healthy building design, and smarter control of ventilation, air purifiers, and other indoor systems.
- L2-ARCTIC (see psi.engr.tamu.edu/l2-arctic-corpus/) is a non-native English speech dataset created by researchers at Texas A&M University and Iowa State University. It contains recordings from 24 speakers across six native-language backgrounds, with word-level and phoneme-level transcriptions plus detailed annotations of pronunciation errors.
- In terms of flags SAP’s office in Berlin is well positioned:

picture taken at venue
Published:
Modified:
