Agrera Robotics

Computer vision for robots that work in unstructured places.

Agrera Robotics is a research and development company building perception systems for robots that operate outside the laboratory — in homes, in the field, and in facilities where conditions change from one day to the next.

Our work centers on learning from very few labeled examples. Real deployments rarely come with a large annotated dataset, and often the thing a robot most needs to recognize is the thing it has seen least. We combine foundation models with self-supervised and low-shot training so that a system can be brought to a new environment, or a new object of interest, without a labeling campaign first.

Applications

Domestic

A service robot that moves through the same space repeatedly can notice what changed since its last run, and use those changes to decide what needs doing. Our published work covers the full pipeline: aligning runs to a common frame, segmenting objects of interest, and comparing them across visits. The harder problem is knowing which changes matter, and we use vision-language models to set aside the ones that do not.

Industrial

The same comparison across repeated visits applies at larger scales. In a tunnel, a warehouse, or a plant, the changes worth flagging are objects left behind, doors that should not be open, and disturbances to a space that is expected to stay fixed. This is a direction the underlying technique extends to naturally, and one we are actively developing.

Agricultural

Crop disease is a low-shot problem by nature: symptoms are sparse in the field, expensive to annotate, and easily confused with ordinary seasonal change. We are building multi-spectral detectors that combine color and thermal imagery to find disease in vineyard canopy, trained on largely unlabeled field collections.

From the work

A framed painting on a wall, detected with a bounding box across two separate runs, with the model output reading object: painting, is_pickup: false.
Deciding which changes matter. A vision-language model classifies the painting as a fixture, so the shift between runs is filtered out rather than reported.
A three-dimensional reconstruction of a living room, with the recovered camera positions from the robot's path shown as wireframe frustums.
Aligning one run to the next. Each pass is reconstructed in three dimensions and registered against the last, recovering the robot's camera path.
A grid comparing SAM v3, CLIP-Seg and DINO on small-object segmentation and on the resulting delta change masks for the same living room scene.
Comparing segmentation backbones. Different foundation models recover different subsets of a cluttered scene, which sets a ceiling on what change can be detected at all.
A field of sunflowers in bloom under a partly cloudy Michigan sky.
Field conditions. Uncontrolled light, dense occlusion, and a narrow window each season in which to collect anything at all.

Select any figure to view it at full size.

Selected publications

2026
E. Martinson, I. Fishta. Change Detection Filtering with Visual Language Models. Ground Vehicle Systems Engineering and Technology Symposium (GVSETS), Novi, MI.
2026
E. Martinson, I. Fishta, D. Butani. Integrating foundation models with change detection to identify tasking for service robots. Frontiers in Robotics and AI, 13:1772005.
2025
E. Martinson, H. Indurthi. Augmenting Open Vocabulary Object Detection with Large Language Models for Home Service Robots. IEEE 21st International Conference on Automation Science and Engineering (CASE), Los Angeles, CA.
2025
Z. Galymzhankyzy, E. Martinson. Lightweight Multispectral Crop-Weed Segmentation for Precision Agriculture. ICRA Workshop on Agricultural Robotics and Automation, Atlanta, GA.
2024
E. Martinson, P. Lauren. Meaningful Change Detection in Indoor Environments Using CLIP Models and NeRF-based Image Synthesis. International Conference on Ubiquitous Robots, New York, NY.
2024
E. Martinson, F. Alladkani. Interactive, Privacy-Aware Semantic Mapping for Homes. International Conference on Ubiquitous Robots, New York, NY.
2021
E. Martinson, B. Furlong, A. Gillies. Training Rare Object Detection in Satellite Imagery with Synthetic GAN Images. CVPR Workshop on Learning from Limited Data.
2018
E. Martinson. Interactive Training of Object Detection Without ImageNet. IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS).
2016
E. Martinson, G. Yalla. Augmenting Deep Convolutional Neural Networks with Depth-Based Layered Detection for Human Detection. IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS).

A complete list is available at ORCID 0000-0001-8194-7770.