2026
Authors
Campos, R; Jatowt, A; Lan, Y; Aliannejadi, M; Bauer, C; MacAvaney, S; Anand, A; Ren, Z; Verberne, S; Bai, N; Mansoury, M;
Publication
ECIR (4)
Abstract
2026
Authors
Reis, J; Areias, M; Barbosa, JG;
Publication
PROGRESS IN ARTIFICIAL INTELLIGENCE, EPIA 2025, PT I
Abstract
Log analysis is fundamental to modern software observability systems, playing a key role in improving system reliability. Recently, there has been a growing adoption of Large Language Models (LLMs) for log anomaly detection, due to their ability to learn complex patterns. In this work, we propose a model-agnostic framework that allows seamless plug-and-play integration of different LLMs, making it easy to experiment with and select the model that fits specific needs. These models are first fine-tuned on normal log data, learning their patterns. During inference, the model predicts the most probable next tokens based on the preceding context in each sequence. Anomaly detection is performed using Top-K predictions, where sequences are flagged as anomalous if the actual log entry does not appear among the K most probable next tokens, with K determined using the validation dataset. The proposed framework is evaluated on three widely-used benchmark datasets-HDFS, BGL, and Thunderbird-where it consistently achieves competitive results, outperforming state-of-the-art methods in multiple scenarios. These results highlight the effectiveness of LLM-based log analysis and the importance of flexibility when selecting models for specific operational contexts.
2026
Authors
Campos, R; Sequeira, R; Nerea, S; Cantante, I; Folques, D; Cunha, LF; Canavilhas, J; Branco, A; Jorge, A; Nunes, S; Guimaraes, N; Silvano, P;
Publication
ADVANCES IN INFORMATION RETRIEVAL, ECIR 2026, PT IV
Abstract
Fact-checking remains a demanding and time-consuming task, still largely dependent on manual verification and unable to match the rapid spread of misinformation online. This is particularly important because debunking false information typically takes longer to reach consumers than the misinformation itself; accelerating corrections through automation can therefore help counter it more effectively. Although many organizations perform manual fact-checking, this approach is difficult to scale given the growing volume of digital content. These limitations have motivated interest in automating fact-checking, where identifying claims is a crucial first step. However, progress has been uneven across languages, with English dominating due to abundant annotated data. Portuguese, like other languages, still lacks accessible, licensed datasets, limiting research, Natural Language Processing (NLP) developments, and applications. In this paper, we introduce ClaimPT, a dataset of European Portuguese news articles annotated for factual claims, comprising 1,308 articles and 6,875 individual annotations. Unlike most existing resources based on social media or parliamentary transcripts, ClaimPT focuses on journalistic content, collected through a partnership with LUSA, the Portuguese News Agency. To ensure annotation quality, two trained annotators labeled each article, with a curator validating all annotations according to a newly proposed scheme. We also provide baseline models for claim detection, establishing initial benchmarks and enabling future NLP and Information Retrieval (IR) applications. By releasing ClaimPT, we aim to advance research on low-resource fact-checking and enhance understanding of misinformation in news media.
2026
Authors
Gomes, G; Ribeiro, E; Pilarski, L; Pinto, T; Reis, A; Barroso, J;
Publication
DISTRIBUTED COMPUTING AND ARTIFICIAL INTELLIGENCE, SPECIAL SESSIONS II, 22ND INTERNATIONAL CONFERENCE
Abstract
The design and development of consumer products require an interdisciplinary approach, often constrained by time-consuming prototyping and manual decision-making processes. As product complexity increases and market demands evolve, the need for automation and intelligent collaboration becomes evident. This paper presents a case study on the design and virtual validation of a premium pen using a multi-agent system, leveraging the integration of large language models (LLMs) and software agents. This combination enables a rational representation of human expertise and interactions, streamlining the design process while enhancing adaptability. Using CrewAI, agents were configured with specialized tasks, collaborating to optimize design, select sustainable materials, and establish quality standards. The agents generated a markdown report and a 3D simulation using Blender and Python, ensuring efficient coordination for an ergonomic, sustainable, high-quality pen. By modeling the rational behavior of human experts, the system demonstrated how LLMs and multi-agent coordination can reduce decision overhead and improve collaboration. The results show that multi-agent systems streamline product development by reducing decision overhead, improving task delegation, and enhancing collaboration. The final design met strict virtual quality standards and aligned with market preferences. This study demonstrates the role of multi-agent systems and LLM integration in Industry 4.0, supporting digital prototyping and virtual simulations to replace traditional physical prototyping.
2026
Authors
Campos, R; Jatowt, A; Lan, Y; Aliannejadi, M; Bauer, C; MacAvaney, S; Anand, A; Ren, Z; Verberne, S; Bai, N; Mansoury, M;
Publication
ECIR (3)
Abstract
2026
Authors
Silva, R; Evans, J; Isidro, J; Marques, M; Fonseca, A; Morais, R; Canavilhas, J; Pasquali, A; Silvano, P; Jorge, A; Guimaraes, N; Nunes, S; Campos, R;
Publication
ADVANCES IN INFORMATION RETRIEVAL, ECIR 2026, PT IV
Abstract
City council minutes are typically lengthy and formal documents with a bureaucratic writing style. Although publicly available, their structure often makes it difficult for citizens or journalists to efficiently find information. In this demo, we present CitiLink, a platform designed to transform unstructured municipal meeting minutes into structured and searchable data, demonstrating how NLP and IR can enhance the accessibility and transparency of local government. The system employs LLMs to extract metadata, discussed subjects, and voting outcomes, which are then indexed in a database to support full-text search with BM25 ranking and faceted filtering through a user-friendly interface. The developed system was built over a collection of 120 min made available by six Portuguese municipalities. To assess its usability, CitiLink was tested through guided sessions with municipal personnel, providing insights into how real users interact with the system. In addition, we evaluated Geminis performance in extracting relevant information from the minutes, highlighting its performance in data extraction.
The access to the final selection minute is only available to applicants.
Please check the confirmation e-mail of your application to obtain the access code.