Show Me Examples: Inferring Visual Concepts from Image Sets
Authors Nick Strackeâ , Kolja Bauerâ , Josh Susskind, Miguel Angel Bautista Martin, Björn Ommerâ
Vision-language models (VLMs) can follow complex textual instructions, yet they struggle to reason from purely visual context. In particular, current models fail to infer shared concepts from sets of example images and apply them to new inputs. We introduce Visual Concept Inference from Sets (VICIS), a task that evaluates this capability. Given a small context set of images sharing a concept and a query image, the model must generate new images that preserve the context-defined concept while remaining consistent with the query. We show that state-of-the-art VLMs perform poorly on this task, often ignoring the visual context or defaulting to biased generations. To address this gap, we propose a training framework and architecture that learn to infer visual concepts from image sets and extract concept-specific embeddings from queries. Experiments on synthetic data and large-scale ImageNet/WordNet data show that our model generates more accurate and diverse outputs and generalizes to unseen concepts and modalities such as sketches.
October 12, 2021 research area Methods and Algorithms , research area Speech and Natural Language Processing conference BayLearn
In this work we study the presence of expert units in pre-trained Transformer Models (TM), and how they impact a modelâs performance. We define expert units to be neurons that are able to classify a concept with a given average precision, where a concept is represented by a binary set of sentences containing the concept (or not). Leveraging the OneSec dataset (Scarlini et al., 2019), we compile a dataset of 1641 concepts that allows diverseâ¦
Improving the Realism of Synthetic Images
Improving the Realism of Synthetic Images
July 7, 2017 research area Computer Vision