MA · Research Assistant · Jožef Stefan Institute

Ivan Porupski

Speech Scientist & Corpus Phonetician

Acoustic phonetics · Paralinguistic feature extraction · ASR & speech transformer fine-tuning · Corpus construction · Machine learning for speech and language

About

I'm a speech scientist and corpus/computational phonetician working at the Jožef Stefan Institute in Ljubljana, where I research how people speak — prosody, disfluency, sentiment — and build the tools and corpora needed to study it at scale.

My work sits at the intersection of acoustic phonetics, machine learning, and large-scale corpus engineering. In practice this means: designing spoken corpora from parliamentary recordings across four Slavic languages, fine-tuning and deploying speech transformer models for paralinguistic tasks (primary stress, filled pauses, acoustic sentiment), running statistical analyses on the results, and building open-source hardware and software when the right tool doesn't exist yet.

I care about transparency: open data, open code, reproducible pipelines. When I build something, I try to make sure anyone else can pick it up and run with it.

Jožef Stefan Institute — E8 Knowledge Technologies MPŠ PhD Fellow CLASSLA / CLARIN Slovenia Member, HFD (Croatian Philological Society)

Projects

A mix of research tooling and personal hardware/software builds. Research Personal

Research

Nosey MEMS Mk2

Open-source single-channel MEMS mic preamp. Pair two boards for acoustic nasalance, or run one as a general-purpose balanced mic.

KiCad MEMS Phonetics hardware
Personal

Kompic̄ Mk1

Open-source, offline-first, ML-capable wrist-mounted sensor platform. ESP32-S3, dense sensor suite, on-device neural nets.

ESP32-S3 On-device ML Wearable
Research

Slavic Speech Pipeline Beta

Tutorial-first toolkit for training and fine-tuning speech transformer models on Slavic languages. Notebooks for learning, scripts for production.

Python HuggingFace Slavic NLP
Personal

Šuška Mk1

Project in progress — stay tuned!

ESP32-S3 Wireless audio

Research

ParlaSpeech

Spoken parliamentary corpora — Croatian, Serbian, Czech, Polish · 6,000+ hours

ParlaSpeech is the central thread of my research. It's a multilingual collection of spoken parliamentary corpora built by aligning ParlaMint transcripts to session recordings, now covering four Slavic languages and six thousand hours of speech. In its 3.0 release the corpora are enriched with five automatic annotation layers: linguistic annotation, sentiment, filled-pause detection, word-level alignment, and primary stress — all of which are products of work I've been directly involved in.

parlaspeech → clarinsi.github.io

Acoustic Phonetics

Prosody, primary stress identification, acoustic feature extraction from parliamentary and naturalistic speech.

Paralinguistic NLP

Filled pause detection, sentiment analysis and acoustic correlates of sentiment across Slavic languages.

Speech Transformer Models

Fine-tuning and deploying wav2vec2-family models for frame-level classification tasks; cross-lingual transfer.

Corpus Construction

Spoken corpus bootstrapping from parliamentary recordings; alignment pipelines; annotation layer design.

Data Science & Statistics

Cross-lingual model evaluation, SVM baselines, statistical comparison of model and human performance.

Open Hardware

Custom MEMS microphone preamps and wearable sensor platforms designed for phonetic and health research.

Research Projects

2026 – Active
EPIC-SI

Early parent–infant communication corpus research (J6-70222 / ARIS).

2025 – Active
Gravitacija / LLM4DH

Large language models for digital humanities research.

2025 Active
CLARIN

Infrastructure program, CLARIN.SI / CLARIN ERIC (I0-E004).

– 2025 Ended
MEZZANINE

Multimodal and multilingual parliamentary speech analysis.

Publications

Authors listed as published. My name in bold.

Accepted / In Press

InterSpeech 2026 2026 · accepted ✦ first author

Umm… With Transformers? Insights from Filled Pause Use across Four Slavic Parliaments

Ivan Porupski, Branimir Dropuljić, Nikola Ljubešić

CLARIN 2026 2026 · accepted ✦ first author

Past the Rapids: Downstream Research on ParlaSpeech

Ivan Porupski, Nikola Ljubešić

JTDH 2026 2026 · accepted ✦ first author

Opinionated, Hesitant and Stressed: Three Studies of How Politicians Speak in Four Slavic Parliaments

Ivan Porupski, Nikola Ljubešić

Published

LREC-COLING 2026 2026

ParlaSpeech 3.0: Richly Annotated Spoken Parliamentary Corpora of Croatian, Czech, Polish, and Serbian

Nikola Ljubešić, Peter Rupnik, Ivan Porupski, Taja Kuzman Pungeršek

arXiv
LLMs4SSH @ LREC 2026 2026

State of the Art in Text Classification for South Slavic Languages: Fine-Tuning or Prompting?

Taja Kuzman Pungeršek, Peter Rupnik, Ivan Porupski, Vuk Dinić, Nikola Ljubešić

arXiv
InterSpeech 2025 2025

Identifying Primary Stress Across Related Languages and Dialects with Transformer-based Speech Encoder Models

Nikola Ljubešić, Ivan Porupski, Peter Rupnik

Proc. Interspeech 2025, pp. 5768–5772

SlavNLP @ ACL 2025 2025

Identifying Filled Pauses in Speech Across South and West Slavic Languages

Nikola Ljubešić, Ivan Porupski, Peter Rupnik

Proc. SlavNLP 2025 (ACL), pp. 1–8, Vienna

ACL Anthology
CLARIN Annual Conference 2025 2025

The ParlaSpeech v3 Collection of Spoken Parliamentary Corpora from the Croatian, Czech, Polish and Serbian Parliament Enriched with Linguistic and Paralinguistic Annotation Layers

Nikola Ljubešić, Peter Rupnik, Ivan Porupski, Taja Kuzman Pungeršek

Dataset · CLARIN.SI 2025

Dataset for Primary Stress Identification in Croatian and Related Languages and Dialects

Nikola Ljubešić, Peter Rupnik, Ivan Porupski, Nejc Robida, Mirna Potočnjak

CLARIN.SI

Conferences & Workshops

Notable Conferences

2026 InterSpeech 2026 paper accepted
2026 LREC-COLING 2026 Palma, Spain
2025 InterSpeech 2025 Rotterdam, Netherlands
2025 ACL 2025 Vienna, Austria

Regional & Local

2026 JTDH 2026 Ljubljana, Slovenia · paper accepted
2026 HDPL 2026 Pula, Croatia · workshop instructor
2025 HDPL 2025 Zagreb, Croatia
2025 CLARIN Annual Conference 2025 paper presented
2024 JTDH 2024 Ljubljana, Slovenia · attendance

Workshops (Co-author & Instructor)

I'm a co-author of the CLASSLA-Express 3.0 workshop cycle — Speech and Web Corpora in the Study of Language Variation — where I designed the spoken-corpus (ParlaSpeech) component and its original teaching materials and hands-on exercises, and I deliver those workshops. The series provides hands-on training in NLP tools for South Slavic languages, run under the CLARIN Knowledge Centre umbrella.

  • Jun 2026 HDPL 2026 — Pula, Croatia
  • Sep 2026 JTDH 2026 — Ljubljana, Slovenia
  • TBD Standalone installment — Zagreb, Croatia

Skills

Speech Science

  • Acoustic analysis & feature extraction
  • Prosody & primary stress
  • Disfluency (filled pauses)
  • Paralinguistic feature modeling
  • Praat, Audacity

Machine Learning & Data Science

  • Fine-tuning & deploying speech transformers
  • SVM & classical ML baselines
  • Cross-lingual evaluation
  • Statistical analysis
  • HuggingFace ecosystem

Corpus Engineering

  • Spoken corpus bootstrapping from recordings
  • Forced alignment pipelines
  • Annotation layer design
  • RTTM, TRS formats

Programming & Tools

  • Python
  • Praat
  • VS Code
  • LLMs
  • Libre/MS
  • KiCad
  • ESP-IDF

Languages

  • Croatian — native
  • English — C1 / IELTS 8.0 (May 2016)

Education & Employment

Employment

Jan 2025 –
Research Assistant Jožef Stefan Institute, Dept. of Knowledge Technologies (E8), Ljubljana
Feb – Jul 2024
Student Teaching Assistant (Demonstrator) Faculty of Humanities and Social Sciences (FFZG), University of Zagreb

Education

2022 – 2024
Master's Degree — Phonetics and Linguistics Faculty of Humanities and Social Sciences (FFZG), University of Zagreb, Croatia
2019 – 2022
Bachelor's Degree — Phonetics and Comparative Literature Faculty of Humanities and Social Sciences (FFZG), University of Zagreb, Croatia
2016 – 2017
Propedeuse — Applied Chemistry (Chemie B) HZ University of Applied Sciences, Vlissingen, Netherlands
2010 – 2014
High School Diploma — Chemical Technology Prirodoslovna škola Vladimira Preloga (PŠVP), Zagreb, Croatia

Full CV

PDF version available on request, or see the full academic record on SICRIS.