Junior
This master's thesis will explore Multimodal Feedback-Guided Demonstration Selection. The goal is to move beyond simple similarity-based retrieval and investigate whether demonstrations can be selected based on their actual usefulness to the model. The project will consider the complete demonstration - image, question/instruction, and answer; and will investigate how multiple complementary demonstrations can be selected efficiently.
The topic sits at the intersection of multimodal AI, retrieval, LLMs, in-context learning, and representation learning, with opportunities for both practical system development and research-oriented experimentation.
You will design, implement, and evaluate a demonstration-selection framework for Large Multimodal Models.
Your work will include:
Reviewing research on multimodal in-context learning, demonstration retrieval, vision-language models, and feedback-guided prompting
Implementing baseline retrieval methods such as random, visual similarity, textual, and multimodal retrieval
Building a feedback-guided retriever that learns which demonstrations are useful based on changes in LMM performance
Representing demonstrations using their image, question/instruction, and answer, rather than visual information alone
Investigating a selection of multiple demonstrations while considering usefulness, diversity, and redundancy
Evaluating the impact of demonstration ordering and context size
Conducting experiments on established multimodal benchmarks using open-source LMMs
Performing ablation studies to understand which components contribute most to performanc
Comparing the proposed approach with existing methods, including GRIP
Analyzing results and documenting findings in a master's thesis and, where appropriate, a research publication
Background in Computer Science, Artificial Intelligence, Data Science, Electrical Engineering, or a related field
Good programming skills in Python
Knowledge of machine learning and deep learning
Familiarity with PyTorch or similar ML frameworks
Basic understanding of Transformers and neural representation learning
Interest in Large Language Models, Computer Vision, NLP, or Multimodal AI
Experience with Hugging Face, vision-language models, GPU computing, or contrastive learning is an advantage
An analytical and research-oriented mindset, with the ability to design experiments and interpret results independently
At Ericsson, you´ll have an outstanding opportunity. The chance to use your skills and imagination to push the boundaries of what´s possible. To build solutions never seen before to some of the world’s toughest problems. You´ll be challenged, but you won’t be alone. You´ll be joining a team of diverse innovators, all driven to go beyond the status quo to craft what comes next.
What happens once you apply?
Click Here to find all you need to know about what our typical hiring process looks like.Encouraging a diverse and inclusive organization is core to our values at Ericsson, that's why we champion it in everything we do. We truly believe that by collaborating with people with different experiences we drive innovation, which is essential for our future growth. We encourage people from all backgrounds to apply and realize their full potential as part of our Ericsson team. Ericsson is proud to be an Equal Opportunity Employer. learn more.
Primary country and city: Sweden (SE) || Stockholm
Req ID: 790806
Sign up to apply and find out right away if you're a fit.
Your agent will tell you — in seconds.
Sign up and I'll tell you right away how well Ericsson matches you — what you already have, and what's missing. Then I stay on it: I search for you and only write when I find something worth your time.