Vision-Based Weed Detection in Row Crops

Computer vision and lasers replace broad herbicide spraying with precise weed targeting.

Columnist · · 11 min read
Cover illustration for “Vision-Based Weed Detection in Row Crops”
Agricultural Robots · October 1, 2026 · 11 min read · 2,394 words

Weed pressure competes with row crops for light, nutrients, water, and space, and the cost of fighting that competition with broadcast herbicide has become difficult to sustain. That spending buys less than it used to. Few new herbicide modes of action have reached the market in recent decades, none appears to be in active development, and herbicide resistance among weed populations keeps expanding, so farmers are paying rising prices for chemistry that works less reliably every season.

The logic that follows from those two facts is straightforward. A system that sprays only where a weed actually grows, instead of blanketing an entire field regardless of weed presence, cuts input cost, slows the rate at which resistance develops, and reduces the herbicide load reaching soil and water. Targeted intervention does not require inventing new chemistry to solve a chemistry problem. It requires knowing where the weeds are, and knowing it precisely enough, and fast enough, to act on that knowledge in the same pass a sprayer already makes across the field.

The vision pipeline: from field image to spray decision

The first is image capture. Cameras mount on ground robots, spray booms, or UAVs, and the altitude, resolution, and speed of motion at the moment of capture set a hard ceiling on what any downstream model can resolve. A camera moving too fast at too great a height will blur or miss the small, low-contrast weed seedlings that most need identifying, no matter how capable the neural network behind it is.

The second layer is inference: the point where a neural network looks at that image and decides, pixel by pixel or box by box, what is crop, what is weed, and what is bare soil. Single-stage object detection methods built on convolutional neural networks, with the YOLO family the most widely used example, dominate this layer because they balance detection accuracy against inference speed in a way that meets the real-time demands of a moving field machine. A YOLO-family model has to classify a frame while the equipment keeps moving, and that inference has to finish before the sprayer boom or laser head passes the point in question. Segmentation approaches, which produce pixel-level masks rather than bounding boxes, offer finer spatial detail, but that precision costs more compute, and compute is scarce on a piece of equipment running off a field-mounted processor rather than a data center.

Crop row geometry gives the inference layer a shortcut. Because row crops are planted in predictable, repeating lines, a detection system can treat the row itself as a spatial prior: anything found outside the expected row band becomes a candidate weed before the network does any fine-grained classification. That shortcut only works if the row-finding step is itself reliable. It must handle weed occlusion, changing light, and rows that are discontinuous or damaged.

The third layer is actuation, the point where a detection decision becomes a physical action, whether that is a herbicide nozzle firing, a laser pulsing, or a mechanical tool moving. The precision of that final step determines how much of the promised reduction in chemical use is actually delivered in the field, rather than simply promised on a spec sheet. A detection model that correctly identifies a weed accomplishes nothing if the nozzle downstream sprays a swath wide enough to cover the crop next to it.

How detection accuracy holds up in practice

Benchmark results for YOLO-family detectors look strong on paper, but those numbers describe performance on the specific dataset a model was trained and tested against, not performance in a field the model has never seen. In one benchmark, a dataset built around corn and four weed species produced a YOLOv7 model with strong mean average precision, improving further once data augmentation was applied. In another, a model called CSCW-YOLOv7, purpose-built for weed detection in wheat, returned precision, recall, and mean average precision scores among the highest reported for any task-specific detector.

Those figures describe a ceiling. An autonomous weeding robot tested under real-field conditions, rather than against a curated benchmark dataset, achieved roughly 85 percent classification accuracy, a meaningful step down from the numbers reported in controlled studies. That gap between benchmark and field is the single fact that everything else in this article has to explain. The reason the gap exists is that performance drops when models face variation in growth stage, lighting, sensor characteristics, or regional field conditions, all of which are the normal texture of a working farm rather than an edge case.

A separate limitation compounds the accuracy gap. Most of these detection models function as black boxes, producing a classification without offering any account of how that classification was reached. That opacity matters practically, not just philosophically: a farmer or agronomist evaluating whether to trust a system's output has no way to interrogate a wrong answer, and an engineer trying to fix a model that fails in a new field has no diagnostic trail to follow.

Diagram: Benchmark vs. Field: The Accuracy Gap That Matters. Visualizes: Show a stark magnitude contrast between two figures: the benchmark accuracy reported for task-specific detectors (the CSCW-YOLOv7 model returning precision, recall, and mean…

How deployed commercial systems handle the pipeline's real constraints

Commercial deployment forces engineering trade-offs that a benchmark table never has to confront, including throughput at working speed, compatibility across crop types, the limits of edge-computing hardware, and the choice of actuation method.

John Deere's See & Spray system mounts cameras on the spray boom itself, scanning the field at working speed and triggering individual ExactApply nozzles in real time, so the actuation remains a conventional herbicide spray applied only where a weed is detected. The model year 2027 version upgrades the camera hardware to operate at higher travel speeds while distinguishing crop from weed with four times the accuracy of the system's earlier versions. The system covered millions of acres in 2025 and now supports corn, soybeans, cotton, wheat, barley, canola, peanuts, sugar beets, and milo, nearly doubling the number of crops it previously handled. A 2025 software update added above-canopy spray support, letting the system target weed escapes and volunteer corn that grow above the crop canopy later in the season, extending its usefulness past the point where most targeted systems stop being effective. See & Spray can reduce non-residual herbicide use by more than two-thirds by target-spraying weeds.

Carbon Robotics takes a different approach to actuation entirely. Its LaserWeeder uses computer vision paired with a Large Plant Model, trained on 150 million labeled plant images gathered by LaserWeeder units operating in the field, and destroys weeds with high-powered lasers rather than herbicide, at high throughput. The G2 generation, announced in early 2025, adds optics capable of sub-millimeter weed detection and real-time image processing across more than 100 crop models. A newer 40-foot version of the LaserWeeder extends this laser-based approach into organic corn and soybeans, a far larger segment of commodity row-crop acreage than the specialty crops the technology first targeted, and runs day and night. Carbon Robotics has a substantial number of units in commercial service and surpassed $100 million in annual revenue for its fiscal year ending January 31, 2026, making it the first field-robotics agtech company built around eliminating herbicide use to reach that revenue level. NVentures, NVIDIA's venture capital arm, has backed the company, and the underlying model's ability to recognize a new weed species from a single image means each additional unit deployed adds to a compounding data advantage rather than simply doing its own job in isolation.

Dimensions Agri Technologies takes a third path with its DAT Ecopatch system, an edge-based vision platform that finds weeds using deep learning and camera hardware and targets them with herbicide, built to integrate with existing sprayers through standard ISOBUS infrastructure rather than requiring farmers to buy proprietary equipment. The company, founded in 1999 in Ski, Norway, began precision spraying work in 2015 using hard-coded algorithms before shifting to AI-based detection, and brought the DAT Ecopatch to commercial launch in 2022. That interoperability choice represents a genuinely different bet than the vertically integrated hardware of See & Spray or LaserWeeder: it lowers the cost for a farmer to adopt the system, at the price of requiring the detection pipeline to perform reliably across whatever sprayer configuration it happens to be bolted onto.

The annotation bottleneck is the pipeline's most stubborn constraint

Compute and camera hardware are not the deepest constraint in this pipeline. Labeled data is. Building a model that reliably tells a specific weed apart from a specific crop at a specific growth stage requires a large, carefully annotated dataset, and that annotation work does not carry over from one field, crop, or region to the next.

Models built for one crop-weed combination generally have to be rebuilt, or at minimum retrained, whenever the crop, the region, the growth stage, or the weed species changes, and this is why the autonomous weeding robot's real-field accuracy falls short of the numbers reported in controlled benchmark studies. The black-box nature of most detection models makes this worse rather than better. When a model fails in a field it was not trained on, the lack of interpretability leaves no way to tell whether the failure traces back to a gap in the labeled training data, an unusual lighting condition, or a genuine failure of the model to generalize, so the default response tends to be gathering more annotated data rather than making a targeted fix.

A further trade-off between accuracy and deployability produces both problems. Models accurate enough to be useful tend to carry large numbers of parameters, and those parameters make the models expensive, or in some cases simply infeasible, to run on the edge hardware that a field machine actually carries. Real-time actuation demands edge deployment, so a model too large to run fast enough on that hardware is not a viable field product regardless of its benchmark accuracy. Researchers are pursuing knowledge distillation and attention mechanisms as ways to shrink model size without losing too much accuracy, and both remain active areas of study without a settled standard.

Every piece of this problem, the annotation cost, the poor generalization, the black-box opacity, and the edge-compute ceiling, comes down to the same root condition: these models know only what they have already been shown, in the form they were shown it.

What vision-language models reveal about the annotation bottleneck, and its limits

A study published in Frontiers in Plant Science tests a direct challenge to that condition: whether vision-language models can detect weeds in soybean UAV imagery without any labeled agricultural training data at all. The study evaluates six commercial VLMs, ChatGPT-4.1, ChatGPT-4o, Gemini Flash 2.5, Gemini Flash Lite 2.5, LLaMA-4 Scout, and LLaMA-4 Maverick, under a single unified prompt designed to elicit weed presence, spatial localization, reasoning, crop growth stage, and crop type.

The results split the six models along different strengths rather than producing one clear winner. Gemini Flash 2.5 delivers the most consistent zero-shot performance along with the highest interpretability of the group. ChatGPT-4o lands between the extremes, offering a balanced profile rather than excelling or lagging at any one dimension. Gemini Flash Lite 2.5 runs efficiently but fails under the study's Error-Probing Prompting stress test, a finding that exposes brittle reasoning once the model is pushed with a counterfactual challenge. That stress test, introduced by the study itself, is a counterfactual follow-up prompt that forces the model to re-analyze an image under the assumption that weeds are present, and the study measures how well a model corrects or defends its original answer using expert-rated scores for Grounding, Specificity, Plausibility, Non-Hallucination, and Actionability.

The methodological finding that matters most is that interpretability tracks spatial correctness rather than simply tracking model size. A larger model that cannot explain where it believes a weed is located offers less field reliability than a smaller model that can point to its reasoning and be checked. That single result reframes what "better" means in this context: scale alone does not buy the trust a farmer or agronomist needs before acting on a model's output.

The practical implication cuts in both directions. If a foundation model can reason about a weed species it has never explicitly been trained to recognize, without a field-specific annotation campaign behind it, the assumption that bespoke crop-weed datasets are the only viable path forward becomes harder to defend, and VLMs earn a real place as promising tools for low-annotation precision weed management. None of that means the annotation bottleneck has been resolved. Zero-shot localization accuracy from these models still lags task-specific YOLO benchmarks, the inference cost of running them on edge hardware has not been demonstrated at the throughput commercial field equipment requires, and brittle reasoning under stress, as Gemini Flash Lite 2.5 showed, would produce costly misidentification errors if deployed at commercial scale.

Continuous in-field learning is changing the data problem at commercial scale

A third path exists between static, hand-annotated datasets and zero-shot vision-language reasoning: a deployed fleet of machines that feeds new field imagery back into its own model continuously. That architecture builds a compounding data advantage that neither the static-dataset paradigm nor current zero-shot VLMs can match, and Carbon Robotics' Large Plant Model is the clearest commercial example of it in operation today.

The LPM trains on what the company describes as the world's largest and fastest-growing agricultural dataset, built from imagery collected continuously by LaserWeeder units operating in fields around the world, and its capacity to recognize a new weed species from a single image means the model's generalization gap narrows as the fleet expands, rather than requiring a fresh annotation project every time a new weed or region enters the picture. NVIDIA Ventures' investment in Carbon Robotics signals that outside capital is treating this flywheel structure as infrastructure in its own right, not merely a feature bolted onto a laser weeding product.

The structural consequence follows directly from how the flywheel compounds. A company that reaches commercial fleet scale first accumulates a training advantage that keeps growing every day its machines run in new fields, under new lighting, against new weed populations. A competitor entering later, whether a startup building a static dataset from scratch or a research lab refining a zero-shot foundation model, has to close a gap that widens the longer the incumbent's fleet keeps operating. The annotation bottleneck does not disappear under this model. It shifts from a cost every new entrant has to pay upfront to an advantage that accrues, acre by acre, to whichever system got its cameras into the most fields first.

Sources

  1. Frontiers | Vision-language models for zero-shot weed detection and visual reasoning in UAV-based precision agriculture
  2. A Vision-Guided Autonomous Weeding Robot Using Image-Based Weed Classification for Sustainable Agriculture | Springer Nature Link
  3. Frontiers | Enhancing weed detection through knowledge distillation and attention mechanism
  4. Frontiers | Deep learning–based approaches for weed detection in crops

More in Agricultural Robots