Publications
2024

InteRead: An Eye Tracking Dataset of Interrupted Reading
Francesca Zermiani, Prajit Dhar, Ekta Sood, Fabian Kögel, Andreas Bulling, Maria Wirzberger
Proc. 31st Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING), pp. 9154--9169, 2024.
AbstractLinksBibTeXProject
Eye movements during reading offer a window into cognitive processes and language comprehension, but the scarcity of reading data with interruptions – which learners frequently encounter in their everyday learning environments – hampers advances in the development of intelligent learning technologies. We introduce InteRead – a novel 50-participant dataset of gaze data recorded during self-paced reading of real-world text. InteRead further offers fine-grained annotations of interruptions interspersed throughout the text as well as resumption lags incurred by these interruptions. Interruptions were triggered automatically once readers reached predefined target words. We validate our dataset by reporting interdisciplinary analyses on different measures of gaze behavior. In line with prior research, our analyses show that the interruptions as well as word length and word frequency effects significantly impact eye movements during reading. We also explore individual differences within our dataset, shedding light on the potential for tailored educational solutions. InteRead is accessible from our datasets web-page: https://www.ife.uni-stuttgart.de/en/llis/research/datasets/.
@inproceedings{zermiani24_coling,
title = {InteRead: An Eye Tracking Dataset of Interrupted Reading},
author = {Francesca Zermiani and Prajit Dhar and Ekta Sood and Fabian Kögel and Andreas Bulling and Maria Wirzberger},
year = {2024},
booktitle = {Proc. 31st Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING)},
pages = {9154--9169},
url = {https://aclanthology.org/2024.lrec-main.802/},
}
2023

Impact of Privacy Protection Methods of Lifelogs on Remembered Memories
Passant Elagroudy, Mohamed Khamis, Florian Mathis, Diana Irmscher, Ekta Sood, Andreas Bulling, Albrecht Schmidt
Proc. ACM SIGCHI Conference on Human Factors in Computing Systems (CHI), pp. 1--10, 2023.
AbstractLinksBibTeXProject
Lifelogging is traditionally used for memory augmentation. However, recent research shows that users’ trust in the completeness and accuracy of lifelogs might skew their memories. Privacy-protection alterations such as body blurring and content deletion are commonly applied to photos to circumvent capturing sensitive information. However, their impact on how users remember memories remain unclear. To this end, we conduct a white-hat memory attack and report on an iterative experiment (N=21) to compare the impact of viewing 1) unaltered lifelogs, 2) blurred lifelogs, and 3) a subset of the lifelogs after deleting private ones, on confidently remembering memories. Findings indicate that all the privacy methods impact memories’ quality similarly and that users tend to change their answers in recognition more than recall scenarios. Results also show that users have high confidence in their remembered content across all privacy methods. Our work raises awareness about the mindful designing of technological interventions.
@inproceedings{elagroudy23_chi,
title = {Impact of Privacy Protection Methods of Lifelogs on Remembered Memories},
author = {Passant Elagroudy and Mohamed Khamis and Florian Mathis and Diana Irmscher and Ekta Sood and Andreas Bulling and Albrecht Schmidt},
year = {2023},
booktitle = {Proc. ACM SIGCHI Conference on Human Factors in Computing Systems (CHI)},
pages = {1--10},
doi = {10.1145/3544548.3581565},
}

Multimodal Integration of Human-Like Attention in Visual Question Answering
Ekta Sood, Fabian Kögel, Philipp Müller, Dominike Thomas, Mihai Bâce, Andreas Bulling
Proc. Workshop on Gaze Estimation and Prediction in the Wild (GAZE), CVPRW, pp. 2647--2657, 2023.
AbstractLinksBibTeXProject Tobii Sponsor Award, Oral Presentation
Human-like attention as a supervisory signal to guide neural attention has shown significant promise but is currently limited to uni-modal integration – even for inherently multi-modal tasks such as visual question answering (VQA). We present the Multimodal Human-like Attention Network (MULAN) – the first method for multimodal integration of human-like attention on image and text during training of VQA models. MULAN integrates attention predictions from two state-of-the-art text and image saliency models into neural self-attention layers of a recent transformer-based VQA model. Through evaluations on the challenging VQAv2 dataset, we show that MULAN achieves a new state-of-the-art performance of 73.98% accuracy on test-std and 73.72% on test-dev and, at the same time, has approximately 80% fewer trainable parameters than prior work. Overall, our work underlines the potential of integrating multimodal human-like and neural attention for VQA.
@inproceedings{sood23_gaze,
title = {Multimodal Integration of Human-Like Attention in Visual Question Answering},
author = {Sood, Ekta and Fabian Kögel and Philipp Müller and Dominike Thomas and Mihai Bâce and Andreas Bulling},
year = {2023},
booktitle = {Proc. Workshop on Gaze Estimation and Prediction in the Wild (GAZE), CVPRW},
pages = {2647--2657},
url = {https://openaccess.thecvf.com/content/CVPR2023W/GAZE/papers/Sood_Multimodal_Integration_of_Human-Like_Attention_in_Visual_Question_Answering_CVPRW_2023_paper.pdf},
}

Improving Neural Saliency Prediction with a Cognitive Model of Human Visual Attention
Ekta Sood, Lei Shi, Matteo Bortoletto, Yao Wang, Philipp Müller, Andreas Bulling
Proc. Annual Meeting of the Cognitive Science Society (CogSci), pp. 3639--3646, 2023.
AbstractLinksBibTeXProject
We present a novel method for saliency prediction that leverages a cognitive model of visual attention as an inductive bias. This approach is in stark contrast to recent purely data-driven saliency models that achieve performance improvements mainly by increased capacity, resulting in high computational costs and the need for large-scale training datasets. We demonstrate that by using a cognitive model, our method achieves competitive performance to the state of the art across several natural image datasets while only requiring a fraction of the parameters. Furthermore, we set the new state of the art for saliency prediction on information visualizations, demonstrating the effectiveness of our approach for cross-domain generalization. We further provide augmented versions of the full MSCOCO dataset with synthetic gaze data using the cognitive model, which we used to pre-train our method. Our results are highly promising and underline the significant potential of bridging between cognitive and data-driven models, potentially also beyond attention.
@inproceedings{sood23_cogsci,
title = {Improving Neural Saliency Prediction with a Cognitive Model of Human Visual Attention},
author = {Ekta Sood and Lei Shi and Matteo Bortoletto and Yao Wang and Philipp Müller and Andreas Bulling},
year = {2023},
booktitle = {Proc. Annual Meeting of the Cognitive Science Society (CogSci)},
pages = {3639--3646},
}

Facial Composite Generation with Iterative Human Feedback
Florian Strohm, Ekta Sood, Dominike Thomas, Mihai Bâce, Andreas Bulling
Proc. The 1st Gaze Meets ML workshop, PMLR, pp. 165--183, 2023.
AbstractLinksBibTeXProject
We propose the first method in which human and AI collaborate to iteratively reconstruct the human’s mental image of another person’s face only from their eye gaze. Current tools for generating digital human faces involve a tedious and time-consuming manual design process. While gaze-based mental image reconstruction represents a promising alternative, previous methods still assumed prior knowledge about the target face, thereby severely limiting their practical usefulness. The key novelty of our method is a collaborative, it- erative query engine: Based on the user’s gaze behaviour in each iteration, our method predicts which images to show to the user in the next iteration. Results from two human studies (N=12 and N=22) show that our method can visually reconstruct digital faces that are more similar to the mental image, and is more usable compared to other methods. As such, our findings point at the significant potential of human-AI collaboration for recon- structing mental images, potentially also beyond faces, and of human gaze as a rich source of information and a powerful mediator in said collaboration.
@inproceedings{strohm23_gmml,
title = {Facial Composite Generation with Iterative Human Feedback},
author = {Strohm, Florian and Sood, Ekta and Thomas, Dominike and B{\^a}ce, Mihai and Bulling, Andreas},
editor = {Lourentzou, Ismini and Wu, Joy and Kashyap, Satyananda and Karargyris, Alexandros and Celi, Leo Anthony and Kawas, Ban and Talathi, Sachin},
year = {2023},
booktitle = {Proc. The 1st Gaze Meets ML workshop, PMLR},
volume = {210},
pages = {165--183},
url = {https://proceedings.mlr.press/v210/strohm23a.html},
publisher = {PMLR},
series = {Proceedings of Machine Learning Research},
pdf = {https://proceedings.mlr.press/v210/strohm23a/strohm23a.pdf},
}
2022

Video Language Co-Attention with Multimodal Fast-Learning Feature Fusion for VideoQA
Adnen Abdessaied, Ekta Sood, Andreas Bulling
Proc. of the 7th Workshop on Representation Learning for NLP (Repl4NLP), pp. 1--12, 2022.
AbstractLinksBibTeXProject
[Equal contribution by the first two authors.] We propose the Video Language Co-Attention Network (VLCN) – a novel memory-enhanced model for Video Question Answering (VideoQA). Our model combines two original contributions: A multimodal fast-learning feature fusion (FLF) block and a mechanism that uses self-attended language features to separately guide neural attention on both static and dynamic visual features extracted from individual video frames and short video clips. When trained from scratch, VLCN achieves competitive results with the state of the art on both MSVD-QA and MSRVTT-QA with 38.06% and 36.01% test accuracies, respectively. Through an ablation study, we further show that FLF improves generalization across different VideoQA datasets and performance for question types that are notoriously challenging in current datasets, such as long questions that require deeper reasoning as well as questions with rare answers.
@inproceedings{abdessaied22_repl4NLP,
title = {Video Language Co-Attention with Multimodal Fast-Learning Feature Fusion for VideoQA},
author = {Adnen Abdessaied and Ekta Sood and Andreas Bulling},
year = {2022},
booktitle = {Proc. of the 7th Workshop on Representation Learning for NLP (Repl4NLP)},
pages = {1--12},
}

Gaze-enhanced Crossmodal Embeddings for Emotion Recognition
Ahmed Abdou, Ekta Sood, Philipp Müller, Andreas Bulling
Proc. International Symposium on Eye Tracking Research and Applications (ETRA), pp. 1--18, 2022.
AbstractLinksBibTeXProject
Emotional expressions are inherently multimodal – integrating facial behavior, speech, and gaze – but their automatic recognition is often limited to a single modality, e.g. speech during a phone call. While previous work proposed crossmodal emotion embeddings to improve monomodal recognition performance, despite its importance, a representation of gaze was not included. We propose a new approach to emotion recognition that incorporates an explicit representation of gaze in a crossmodal emotion embedding framework. We show that our method outperforms the previous state of the art for both audio-only and video-only emotion classification on the popular One-Minute Gradual Emotion Recognition dataset. Furthermore, we report extensive ablation experiments and provide insights into the performance of different state-of-the-art gaze representations and integration strategies. Our results not only underline the importance of gaze for emotion recognition but also demonstrate a practical and highly effective approach to leveraging gaze information for this task.
@inproceedings{abdou22_etra,
title = {Gaze-enhanced Crossmodal Embeddings for Emotion Recognition},
author = {Ahmed Abdou and Ekta Sood and Philipp Müller and Andreas Bulling},
year = {2022},
booktitle = {Proc. International Symposium on Eye Tracking Research and Applications (ETRA)},
volume = {6},
pages = {1--18},
doi = {10.1145/3530879},
}
2021

Neural Photofit: Gaze-based Mental Image Reconstruction
Florian Strohm, Ekta Sood, Sven Mayer, Philipp Müller, Mihai Bâce, Andreas Bulling
Proc. IEEE International Conference on Computer Vision (ICCV), pp. 245-254, 2021.
AbstractLinksBibTeXProject
We propose a novel method that leverages human fixations to visually decode the image a person has in mind into a photofit (facial composite). Our method combines three neural networks: An encoder, a scoring network, and a decoder. The encoder extracts image features and predicts a neural activation map for each face looked at by a human observer. A neural scoring network compares the human and neural attention and predicts a relevance score for each extracted image feature. Finally, image features are aggregated into a single feature vector as a linear combination of all features weighted by relevance which a decoder decodes into the final photofit. We train the neural scoring network on a novel dataset containing gaze data of 19 participants looking at collages of synthetic faces. We show that our method significantly outperforms a mean baseline predictor and report on a human study that shows that we can decode photofits that are visually plausible and close to the observer's mental image. Code and dataset available upon request.
Code: Available upon request.
Dataset: Available upon request.
@inproceedings{strohm21_iccv,
title = {Neural Photofit: Gaze-based Mental Image Reconstruction},
author = {Florian Strohm and Ekta Sood and Sven Mayer and Philipp Müller and Mihai Bâce and Andreas Bulling},
year = {2021},
booktitle = {Proc. IEEE International Conference on Computer Vision (ICCV)},
pages = {245-254},
doi = {10.1109/ICCV48922.2021.00031},
}

VQA-MHUG: A gaze dataset to study multimodal neural attention in VQA
Ekta Sood, Fabian Kögel, Florian Strohm, Prajit Dhar, Andreas Bulling
Proc. ACL SIGNLL Conference on Computational Natural Language Learning (CoNLL), pp. 27--43, 2021.
AbstractLinksBibTeXProject Oral Presentation
We present VQA-MHUG - a novel 49-participant dataset of multimodal human gaze on both images and questions during visual question answering (VQA) collected using a high-speed eye tracker. We use our dataset to analyze the similarity between human and neural attentive strategies learned by five state-of-the-art VQA models: Modulated Co-Attention Network (MCAN) with either grid or region features, Pythia, Bilinear Attention Network (BAN), and the Multimodal Factorized Bilinear Pooling Network (MFB). While prior work has focused on studying the image modality, our analyses show - for the first time - that for all models, higher correlation with human attention on text is a significant predictor of VQA performance. This finding points at a potential for improving VQA performance and, at the same time, calls for further research on neural text attention mechanisms and their integration into architectures for vision and language tasks, including but potentially also beyond VQA.
@inproceedings{sood21_conll,
title = {VQA-MHUG: A gaze dataset to study multimodal neural attention in VQA},
author = {Sood, Ekta and Kögel, Fabian and Strohm, Florian and Dhar, Prajit and Bulling, Andreas},
year = {2021},
booktitle = {Proc. ACL SIGNLL Conference on Computational Natural Language Learning (CoNLL)},
pages = {27--43},
doi = {10.18653/v1/2021.conll-1.3},
publisher = {Association for Computational Linguistics},
}
2020

Anticipating Averted Gaze in Dyadic Interactions
Philipp Müller, Ekta Sood, Andreas Bulling
Proc. ACM International Symposium on Eye Tracking Research and Applications (ETRA), pp. 1-10, 2020.
AbstractLinksBibTeXProject
We present the first method to anticipate averted gaze in natural dyadic interactions. The task of anticipating averted gaze, i.e. that a person will not make eye contact in the near future, remains unsolved despite its importance for human social encounters as well as a number of applications, including human-robot interaction or conversational agents. Our multimodal method is based on a long short-term memory (LSTM) network that analyses non-verbal facial cues and speaking behaviour. We empirically evaluate our method for different future time horizons on a novel dataset of 121 YouTube videos of dyadic video conferences (74 hours in total). We investigate person-specific and person-independent performance and demonstrate that our method clearly outperforms baselines in both settings. As such, our work sheds light on the tight interplay between eye contact and other non-verbal signals and underlines the potential of computational modelling and anticipation of averted gaze for interactive applications.
@inproceedings{mueller20_etra,
title = {Anticipating Averted Gaze in Dyadic Interactions},
author = {Philipp Müller and Ekta Sood and Andreas Bulling},
year = {2020},
booktitle = {Proc. ACM International Symposium on Eye Tracking Research and Applications (ETRA)},
pages = {1-10},
doi = {10.1145/3379155.3391332},
}

Interpreting Attention Models with Human Visual Attention in Machine Reading Comprehension
Ekta Sood, Simon Tannert, Diego Frassinelli, Andreas Bulling, Ngoc Thang Vu
Proc. ACL SIGNLL Conference on Computational Natural Language Learning (CoNLL), pp. 12-25, 2020.
AbstractLinksBibTeXProject
While neural networks with attention mechanisms have achieved superior performance on many natural language processing tasks, it remains unclear to which extent learned attention resembles human visual attention. In this paper, we propose a new method that leverages eye-tracking data to investigate the relationship between human visual attention and neural attention in machine reading comprehension. To this end, we introduce a novel 23 participant eye tracking dataset - MQA-RC, in which participants read movie plots and answered pre-defined questions. We compare state of the art networks based on long short-term memory (LSTM), convolutional neural models (CNN) and XLNet Transformer architectures. We find that higher similarity to human attention and performance significantly correlates to the LSTM and CNN models. However, we show this relationship does not hold true for the XLNet models – despite the fact that the XLNet performs best on this challenging task. Our results suggest that different architectures seem to learn rather different neural attention strategies and similarity of neural to human attention does not guarantee best performance.
@inproceedings{sood20_conll,
title = {Interpreting Attention Models with Human Visual Attention in Machine Reading Comprehension},
author = {Sood, Ekta and Tannert, Simon and Frassinelli, Diego and Bulling, Andreas and Vu, Ngoc Thang},
year = {2020},
booktitle = {Proc. ACL SIGNLL Conference on Computational Natural Language Learning (CoNLL)},
pages = {12-25},
doi = {10.18653/v1/P17},
publisher = {Association for Computational Linguistics},
}

Improving Natural Language Processing Tasks with Human Gaze-Guided Neural Attention
Ekta Sood, Simon Tannert, Philipp Müller, Andreas Bulling
Advances in Neural Information Processing Systems (NeurIPS), pp. 1--15, 2020.
AbstractLinksBibTeXProject
A lack of corpora has so far limited advances in integrating human gaze data as a supervisory signal in neural attention mechanisms for natural language processing (NLP). We propose a novel hybrid text saliency model (TSM) that, for the first time, combines a cognitive model of reading with explicit human gaze supervision in a single machine learning framework. We show on four different corpora that our hybrid TSM duration predictions are highly correlated with human gaze ground truth. We further propose a novel joint modelling approach to integrate the predictions of the TSM into the attention layer of a network designed for a specific upstream task without the need for task-specific human gaze data. We demonstrate that our joint model outperforms the state of the art in paraphrase generation on the Quora Question Pairs corpus by more than 10% in BLEU-4 and achieves state-of-the-art performance for sentence compression on the challenging Google Sentence Compression corpus. As such, our work introduces a practical approach for bridging between data-driven and cognitive models and demonstrates a new way to integrate human gaze-guided neural attention into NLP tasks.
@inproceedings{sood20_neurips,
title = {Improving Natural Language Processing Tasks with Human Gaze-Guided Neural Attention},
author = {Sood, Ekta and Tannert, Simon and Müller, Philipp and Bulling, Andreas},
year = {2020},
booktitle = {Advances in Neural Information Processing Systems (NeurIPS)},
pages = {1--15},
url = {https://proceedings.neurips.cc/paper/2020/hash/460191c72f67e90150a093b4585e7eb4-Abstract.html},
}
2019

Predicting Gaze Patterns: Text Saliency for Integration into Machine Learning Tasks
Ekta Sood
Proc. International Workshop on Computational Cognition (ComCo), pp. 1--2, 2019.
LinksBibTeXProject Best Poster Award
@inproceedings{sood19_comco,
title = {Predicting Gaze Patterns: Text Saliency for Integration into Machine Learning Tasks},
author = {Sood, Ekta},
year = {2019},
booktitle = {Proc. International Workshop on Computational Cognition (ComCo)},
pages = {1--2},
url = {https://collaborative-ai.org/publications/sood19_comco_poster.pdf},
}
2018

Comparing Attention-Based Convolutional and Recurrent Neural Networks: Success and Limitations in Machine Reading Comprehension
Matthias and
Jagfeld, Glorianna and
Sood, Ekta and
Yu, Xiang and
Vu, Ngoc Thang Blohm
Proc. ACL SIGNLL Conference on Computational Natural Language Learning (CoNLL), pp. 108--118, 2018.
AbstractLinksBibTeXProject
We propose a machine reading comprehension model based on the compare-aggregate framework with two-staged attention that achieves state-of-the-art results on the MovieQA question answering dataset. To investigate the limitations of our model as well as the behavioral difference between convolutional and recurrent neural networks, we generate adversarial examples to confuse the model and compare to human performance. Furthermore, we assess the generalizability of our model by analyzing its differences to human inference, drawing upon insights from cognitive science.
@inproceedings{blohm18_conll,
title = {Comparing Attention-Based Convolutional and Recurrent Neural Networks: Success and Limitations in Machine Reading Comprehension},
author = {Blohm, Matthias and
Jagfeld, Glorianna and
Sood, Ekta and
Yu, Xiang and
Vu, Ngoc Thang},
year = {2018},
booktitle = {Proc. ACL SIGNLL Conference on Computational Natural Language Learning (CoNLL)},
pages = {108--118},
doi = {10.18653/v1/K18-1011},
url = {https://www.aclweb.org/anthology/K18-1011},
publisher = {Association for Computational Linguistics},
}