This year, the 35th International Joint Conference on Artificial Intelligence (IJCAI-ECAI 2026) took place in Bremen from 15–21 August, and the IML department was represented with ten contributions. Across the main conference and associated tracks, including the Demonstrations Track, 858 of 6,342 full submissions were accepted, corresponding to an overall acceptance rate of 13.53%.

At the demonstrations track, Siting Liang presented a unified generative sequence-to-sequence system for joint event extraction that supports both pipeline and end-to-end configurations, titled “A Scalable Cross-Domain Event Extraction System via a Unified Generative Training Framework”.

The author Siting Liang demonstrates how trained across diverse event datasets, a single model captures domain-specific knowledge while generalizing to large and evolving event ontologies. Live Video Demo

Rida Saghir presented “Visualizing and Interacting with Model Representation Space for Human-Centric Active Learning” as a demo. The work introduces an interactive tool that lets users engage directly with a model’s evolving internal representation space to guide which data gets labeled next, rather than leaving that choice entirely to the algorithm. In a pilot study on audio classification, human-guided sample selection matched the performance of standard automated strategies while giving annotators clearer insight into how the model was learning.

Novruz Mammadli presented the demo “Making Weak Supervision Interactive: Exploring Transfer from Sound Libraries to Passive Acoustic Monitoring Data”. The work addresses a key challenge in wildlife acoustic monitoring: while large sound archives contain valuable recordings and species-level annotations, they usually do not indicate exactly when a target vocalisation occurs. Our system uses Multiple Instance Learning (MIL) to learn approximate temporal locations from these weak labels and transfer the resulting acoustic detectors to Passive Acoustic Monitoring (PAM) data.

The demo brings this process into an interactive workflow in which users can inspect detected sound segments and refine the model through active learning. Using recordings from a real museum sound collection and the AnuraSet PAM dataset, our preliminary evaluation shows that weakly annotated archival recordings can provide a useful training signal for downstream species detection, while also highlighting the impact of model choice and domain differences.

Novruz Mammadli shows IML’s demo at the IJCAI conference

IML researchers also presented seven contributions at workshops held in conjunction with IJCAI-ECAI 2026.

At the Workshop on Explainable Artificial Intelligence (XAI), Siting Liang presented the study of “Evaluating Explanation-Driven Vision–Language Reasoning via Generation Order Interventions”. In this work, we systematically compare answer-first and rationale-first generation under a controlled single-step setting across knowledge-intensive QA, visual entailment, and compositional grounding tasks, eliminating confounding intermediate reasoning processes. Our results show that larger models are essential for effective rationale-first reasoning, while answer-first generation is more robust to formatting errors, with reasoning faithfulness and task performance jointly shaped by explanation order, model scale, pretraining knowledge, fine-tuning, and task structure.

Furthermore, “Beyond Heatmaps: Unsupervised Concept-Graph Reasoning for Interpretable Visual Explanation” was presented at the same workshop by Md Abdul Kadir. Concept Bottleneck Models (CBMs) provide an intrinsically interpretable alternative to post-hoc explanations. However, existing CBMs often rely on predefined concept vocabularies or supervised annotations, lack explicit concept grounding, and summarize each concept with a single image-level score. This work proposes a Graph-based Concept Bottleneck Model (G-CBM), an intrinsically interpretable framework that performs unsupervised concept discovery via Non-negative Matrix Factorization (NMF) and represents the discovered concepts as nodes in a per-image concept-graph representation.

At the EXPLIMED workshop on Explainable Artificial Intelligence for the Medical Domain, the paper TRACE: A Concept Bottleneck Model for Longitudinal 3D Glioblastoma Response Assessment was presented by Hasan Md Tusfiqur Alam. The work introduces TRACE, an interpretable AI framework for assessing treatment response in glioblastoma from longitudinal MRI. Instead of directly predicting a clinical response from medical images as a black box, TRACE follows the clinically established RANO 2.0 assessment process and represents intermediate information, such as tumor measurements and changes between baseline and follow-up scans, as human-interpretable clinical concepts.

A key aspect of TRACE is its focus on actionable interpretability. The intermediate concepts can not only be inspected but also corrected when necessary, with these corrections propagating through the structured reasoning process to the final treatment-response prediction. TRACE combines a 3D vision model with a structured concept bottleneck that captures clinically meaningful dependencies between measurements, derived concepts, and the final decision. Evaluated on the LUMIERE longitudinal glioblastoma dataset, the results demonstrate the potential of concept-based AI for building more transparent, clinically verifiable, and interactive decision-support systems for longitudinal medical imaging.

At the 1st Workshop on AI-Based Humanoid Robot Design and Control Through the Lens of HRI, Evolution, and Biomechanics, Tuan Tran presented an agentic approach for deploying humanoid robots in indoor tasks via language instructions, without expensive data collection or fine-tuning, titled “Training-Free Language-Guided Robot Scene Interaction: An Agentic Approach”.

Tuan Tran from IML presents at the IJCAI workshop

He also presented our report on the next generation of LLMs for software engineering, bridging SE4AI and AI4SE through dynamic visual grounding, at the Generative Code Intelligence (GeCoin) Workshop (“Beyond Text-Only Code Generation: Dynamic Visual Understanding in Software Engineering”).

Two IML contributions were accepted for the GenAIK-NORA Workshop: In the paper “PhotoGraph: Claim-Centric Knowledge Graphs for Personal Photo Management” by Omair Shahzad Bhatti et al. we present PhotoGraph, a claim-centric, spatially and temporally aware knowledge graph framework for personal photo understanding and photobook co-creation. Instead of treating model outputs as fixed facts, PhotoGraph stores every prediction as an evidence-grounded claim with provenance, confidence, and an explicit lifecycle state, so that users can accept, reject, or correct it. This makes contextual retrieval traceable; every result comes with the claims that justify it and supports event-based photobook creation with fact-grounded storylines.

In the second paper “Structuring Annotation Label Spaces by Natural Language Concept Elicitation and Ontology Grounding” by Pratik Sitapara et al., we present a demonstration system that implements Grounded Label Space Engineering (GLSE) for knowledge-centric annotation workflows, enabling domain experts to construct label spaces as evolving, ontology-linked semantic objects through natural language interaction, ground concepts in authoritative knowledge resources where possible, and represent genuinely novel concepts as modular ontology representations with provenance. The workflow is demonstrated in a bioacoustics annotation scenario, and a video of the demo is available.

Pratik Sitapara explains his demo and poster, Thiago Gouvêa on the right

Both of these contributions received best paper awards, see this news here, congratulations to the authors!


Categories: