Logo
User Name

Emina Alickovic

Associate Professor, Linköping University

Društvene mreže:

Institucija

Linköping University
Associate Professor
David Rannaleet, Victor Gunnarsson, Bo Bernhardsson, Martin A. Skoglund, E. Alickovic

Limited training data constrains deep learning models for Auditory Attention Decoding (AAD) in hearing aids (HAs). AAD uses electroencephalogram (EEG) data to decode listener's attention, enabling real-time tracking of specific sound sources. However, achieving high AAD performance with short time windows typical in HAs (<=1s) is challenging due to the scarcity of real-world speech-evoked EEG data. To address this issue, we investigate diffusion probabilistic models (DPMs) for generating synthetic speech-evoked EEG data. DPMs learn the underlying complex data structure through a denoising process and can generate realistic samples suitable for data augmentation. We evaluate the use of synthetic EEG data for augmenting datasets in locus-of-attention (LoA) classification tasks. Our experiments demonstrate that DPMs can generate realistic EEG signals and that incorporating synthetic data significantly improves AAD performance compared to models trained solely on measured EEG data (p<0.05). These results highlight the potential of diffusion-based data augmentation to mitigate training data limitations and improve the robustness of short-window AAD models in HA applications.

Sara Carta, E. Alickovic, Johannes Zaar, Alejandro López Valdés, Giovanni M. Di Liberto

Successful speech communication in multi-talker scenarios requires a skillful combination of sustained attention and rapid attention switching. While the neurophysiology literature offers detailed insights into the neural underpinnings of sustained attention, there remains considerable uncertainty on how attention switching takes place. In this study, using EEG recordings from normal-hearing adults in an immersive multi-talker environment, we measured the neural encoding of two competing speech streams amid background babble. Participants were cued to switch attention between streams every 15–30 s. Neural tracking was assessed via Temporal Response Functions (TRF), confirming reliable decoding of attentional focus. Our results indicate asymmetric disengagement and engagement processes during attention switches, where the neural tracking of the new target stream emerges before disengaging from the previous target, revealing a transient simultaneous encoding of two speech streams. That transition was closely mirrored by a reduction in EEG alpha power, informing on the cognitive effort during different phases of the attention switch. We then isolated cortical activity reflecting lexical prediction mechanisms to determine how lexical context is updated after an attention switch, comparing four context-accumulation strategies that were constructed using Large Language Models. Our findings elucidate both the temporal and contextual mechanisms underlying auditory attention shifts, pointing to the possibility that listeners carry out a reset in lexical context after switching attention. By focusing on dynamic attentional reallocation, this study offers insights into the brain’s capacity for flexible speech processing in complex listening environments.

Johanna Wilroth, Oskar Keding, Martin A. Skoglund, Maria Sandsten, Martin Enqvist, E. Alickovic

Everyday communication is dynamic and multisensory, often involving shifting attention, overlapping speech, and visual cues. Yet, most neural attention tracking studies are still limited to highly controlled lab settings, using clean, often audio‐only stimuli and requiring sustained attention to a single talker. This work addresses that gap by introducing a novel dataset from 24 normal‐hearing participants. We used a wearable electroencephalography (EEG) system (44 scalp electrodes and 20 cEEGrid electrodes) in an audiovisual (AV) paradigm with three conditions: sustained attention to a single talker in a two‐talker environment, attention switching between two talkers, and unscripted two‐talker conversations with a competing single talker. Analysis included temporal response functions (TRFs) modeling, optimal lag analysis, selective attention classification with decision windows ranging from 1.1 to 35 s, and comparisons of TRFs for attention to AV conversations versus side audio‐only talkers. Key findings show significant differences in the attention‐related P2 peak between attended and ignored speech across conditions for scalp EEG. Interestingly, our results revealed strong cross‐condition generalization, with models trained in one condition maintaining good performance when evaluated on the other two. No significant change in performance between switching and sustained attention suggests robustness for attention switches. Optimal lag analysis revealed a narrower peak for conversation compared to single‐talker AV stimuli, reflecting the additional complexity of multi‐talker processing. Classification of selective attention was consistently above chance (55%–70% accuracy) for scalp EEG, whereas cEEGrid data yielded lower correlations, highlighting the need for further methodological improvements. These results demonstrate that wearable EEG can reliably track selective attention in dynamic, multisensory listening scenarios and provide guidance for designing future AV paradigms and real‐world attention tracking applications.

Payam Shahsavari Baboukani, Rodrigo Ordoñez, Carina Gravesen, Jan Østergaard, M. Rank, E. Alickovic, A. F. Cabrera

Payam Shahsavari Baboukani, E. Alickovic, Jan Østergaard, Kasper Eskelund

This study examines how the signal‐to‐noise‐interference ratio (SNIR) influences auditory performance and neural responses associated with listening effort (LE). A new dataset was collected from individuals with moderate hearing loss, all fitted with hearing aids (HAs). Participants listened to two competing audiobooks presented via front‐facing loudspeakers, while 16‐talker babble noise was delivered from background speakers. Six SNIR levels (5.47, −$$ - $$ 3.55, −$$ - $$ 2.13, −$$ - $$ 1.19, −$$ - $$ 0.64, and −$$ - $$ 0.27 dB) were tested. Participants were instructed to attend to one audiobook while ignoring the competing speech and background noise and were subsequently assessed on content of the attended speech and perceived LE. The performance results revealed a significant linear effect of SNIR on subjective ratings of LE and a primarily quadratic effect on comprehension questionnaire accuracy, suggesting that perceived effort decreases steadily with improving SNIR, while comprehension questionnaire performance exhibits a plateau at higher SNIR levels. The EEG analyses demonstrated a significant relationship between SNIR and local connectivity, specifically in the parietal electrodes and in the alpha frequency band. Further analysis confirmed that parietal local connectivity correlates linearly with subjective LE ratings. Moreover, spectral power analysis showed that parietal alpha power is not significantly related to SNIR, indicating that local connectivity may serve as a more sensitive neural marker. While local connectivity and alpha power may share some neural underpinnings, they offer complementary, yet non‐identical insights. These findings highlight the potential of local EEG connectivity as a reliable estimate of LE in acoustically challenging environments.

Johanna Wilroth, Oskar Keding, Martin A. Skoglund, Maria Sandsten, Martin Enqvist, E. Alickovic

Everyday communication is dynamic and multisensory, often involving shifting attention, overlapping speech and visual cues. Yet, most neural attention tracking studies are still limited to highly controlled lab settings, using clean, often audio-only stimuli and requiring sustained attention to a single talker. This work addresses that gap by introducing a novel dataset from 24 normal-hearing participants. We used a mobile electroencephalography (EEG) system (44 scalp electrodes and 20 cEEGrid electrodes) in an audiovisual (AV) paradigm with three conditions: sustained attention to a single talker in a two-talker environment, attention switching between two talkers, and unscripted two-talker conversations with a competing single talker. Analysis included temporal response functions (TRFs) modeling, optimal lag analysis, selective attention classification with decision windows ranging from 1.1s to 35s, and comparisons of TRFs for attention to AV conversations versus side audio-only talkers. Key findings show significant differences in the attention-related P2-peak between attended and ignored speech across conditions for scalp EEG. No significant change in performance between switching and sustained attention suggests robustness for attention switches. Optimal lag analysis revealed narrower peak for conversation compared to single-talker AV stimuli, reflecting the additional complexity of multi-talker processing. Classification of selective attention was consistently above chance (55-70% accuracy) for scalp EEG, while cEEGrid data yielded lower correlations, highlighting the need for further methodological improvements. These results demonstrate that mobile EEG can reliably track selective attention in dynamic, multisensory listening scenarios and provide guidance for designing future AV paradigms and real-world attention tracking applications.

Heidi B Borges, Johannes Zaar, E. Alickovic, C. B. Christensen, P. Kidmose

OBJECTIVE Previous studies have demonstrated that the speech reception threshold (SRT) can be estimated using scalp electroencephalography (EEG), referred to as SRTneuro. The present study assesses the feasibility of using ear-EEG, which allows for discreet measurement of neural activity from in and around the ear, to estimate the SRTneuro. Approach: Twenty young normal-hearing participants listened to audiobook excerpts at varying signal-to-noise ratios (SNRs) whilst wearing a 66-channel EEG cap and 12 ear-EEG electrodes. A linear decoder was trained on different electrode configurations to estimate the envelope of the audio excerpts from the EEG recordings. The reconstruction accuracy was determined by calculating the Pearson's correlation between the actual and the estimated envelope. A sigmoid function was then fitted to the reconstruction-accuracy-vs-SNR data points, with the midpoint of the sigmoid serving as the SRTneuro estimate for each participant. Main results: Using only in-ear electrodes , the estimated SRTneuro was within 3 dB of the behaviorally measured SRT (SRTbeh) for 6 out of 20 participants (30%). With electrodes placed both in and around the ear, the SRTneuro was within 3 dB of the SRTbeh for 19 out of 20 participants (95%) and thus on par with the reference estimate obtained from full-scalp EEG. Using only electrodes in and around the ear from the right side of the head, the SRTneuro remained within 3 dB of the SRTbeh for 19 out of 20 participants. .

Heidi B Borges, E. Alickovic, C. B. Christensen, Preben Kidmose, Johannes Zaar

Previous studies have demonstrated the feasibility of estimating the speech reception threshold (SRT) based on electroencephalography (EEG), termed SRTneuro, in younger normal-hearing (YNH) participants. This method may support speech perception in hearing-aid users through continuous adaptation of noise-reduction algorithms. The prevalence of hearing impairment and thereby hearing-aid use increases with age. The SRTneuro estimation is based on envelope reconstruction accuracy, which has also been shown to increase with age, possibly due to excitatory/inhibitory imbalance or recruitment of additional cortical regions. This could affect the estimated SRTneuro. This study investigated the age-related changes in the temporal response function (TRF) and the feasibility of SRTneuro estimation across age. Twenty YNH and 22 older normal-hearing (ONH) participants listened to audiobook excerpts at various signal-to-noise ratios (SNRs) while EEG was recorded using 66 scalp electrodes and 12 in-ear-EEG electrodes. A linear decoder reconstructed the speech envelope, and the Pearson's correlation was calculated between the reconstructed and speech-stimulus envelopes. A sigmoid function was fitted to the reconstruction-accuracy-versus-SNR data points, and the midpoint was used as the estimated SRTneuro. The results show that the SRTneuro can be estimated with similar precision in both age groups, whether using all scalp electrodes or only those in and around the ear. This consistency across age groups was observed despite physiological differences, with the ONH participants showing higher reconstruction accuracies and greater TRF amplitudes. Overall, these findings demonstrate the robustness of the SRTneuro method in older individuals and highlight its potential for applications in age-related hearing loss and hearing-aid technology.

Payam Shahsavari Baboukani, E. Alickovic, Jan Østergaard

Hearing aid (HA) users often experience increased listening effort, particularly in noisy environments. While noise reduction (NR) algorithms aim to alleviate this, traditional electroencephalography (EEG) methods based on power analysis have limited success in assessing the listening effort in this population. This study proposes a novel method using a whole-head synchronization map analysis that uses local connectivity, a measure of statistical dependencies within localized brain regions. We use EEG electrodes to define a region based on the surrounding electrodes in the first-order neighborhood. This approach was tested using EEG data from 22 HA users with active or inactive NR engaged in a continuous speech-in-noise (SiN) task at low (3dB) and high (8dB) signal-to-noise ratio (SNR) levels. Whole-head synchronization was quantified using circular omega complexity (COC), a multivariate phase synchrony measure. Results showed increased local connectivity in the alpha band (8–12 Hz) within frontal and occipital regions during SiN condition compared to the background noise-only (NO) condition. Furthermore, NR activation impacted the synchronization map differently at the two SNRs of the experiment, with greater effect observed at low SNR, primarily in the left parietal region and alpha band. This behavior is in line with that of existing measures for listening effort, and therefore suggests that EEG local connectivity analysis holds promise as a tool for objectively assessing listening effort in HA users, especially in challenging listening environments.

...
...
...

Pretplatite se na novosti o BH Akademskom Imeniku

Ova stranica koristi kolačiće da bi vam pružila najbolje iskustvo

Saznaj više