Foundry Unveils SmartRoto: An AI-Powered Paradigm Shift for Visual Effects Rotoscoping
LONDON — Rotoscoping—often regarded as the ultimate endurance test in visualeffects—may finally be shedding its reputation as the industry’s most tedious grind. Foundry, the creative software powerhouse behind the industry-standard compositing application Nuke, has officially announced the release of SmartRoto. Designed as an AI-driven plugin integrated natively across the Nuke ecosystem (including Nuke, NukeX, Nuke Studio, and Nuke Indie), SmartRoto aims to fundamentally transform how artists approach frame-by-frame shape creation and animation.
Priced at an annual subscription of $499 with both node-locked and floating license options available, SmartRoto arrives after years of intensive R&D. Ahead of the launch, industry publications sat down with Adam Cherbetji, Foundry’s Director of Product for AI Research, to dissect the technology, its underlying philosophy, and what it means for the future of digital compositing.
Main Facts: What SmartRoto Is—and What It Isn’t
To understand SmartRoto, industry professionals must first unlearn expectations set by preceding automated matting tools. It is neither a blanket "magic bullet" nor a standard segmentation-based matting tool that wraps an automated spline around a coarse AI mask.
Instead of replacing the artist’s cognitive workload, SmartRoto is engineered to accelerate it.
The Core Mechanism: SmartRoto operates directly on image features rather than relying on optical flow or coarse pixel segmentation. It tracks and propagates subpixel-accurate Bézier splines across an image sequence.
The Two-Tier Keyframe System: The workflow revolves around a strict hierarchy of keyframes. Smart keys are predicted automatically by the AI model. User keys (originally dubbed "gold keys" during development) are explicitly placed or adjusted by the artist and are treated by the system as an absolute source of truth.
Non-Destructive Guidance: User keys are never overwritten by the model. Corrections made by an artist feed dynamically back into the system, updating the local surrounding frame range without requiring a user to restart their work from scratch.
Hardware Requirements: Built to be lightweight compared to massive cloud-dependent foundation models, SmartRoto runs entirely locally and requires approximately 8 gigabytes of VRAM, depending on resolution and active shape counts.
Commercial Safety: Trained exclusively on fully licensed data—including custom footage shot by Foundry’s director of research and annotated both internally and externally—the model is completely secure for commercial studio pipelines. No telemetry or image data is transmitted to the cloud.
Chronology: Four Years in the Making
The journey to SmartRoto spans nearly a decade of technical iteration, evolving from academic explorations into a focused commercial product.
6–7 Years Ago: Foundry first explored automated rotoscoping concepts in partnership with academic institutions, notably collaborating with DNEG and the University of Bath. While these early experiments laid intellectual groundwork, the technology was not yet mature enough for production demands.
4 Years Ago: Realizing the unique constraints of production pipelines, Foundry reset the project internally. A dedicated research team built an entirely novel machine-learning model and custom training datasets from the ground up.
Beta Testing Phase: Foundry embedded professional roto artists into its offices for intensive two-week testing cycles. Artists were subjected to blind tests comparing plain vanilla Nuke workflows against SmartRoto across diverse shot types.
Present Day: Foundry officially commercializes SmartRoto, rolling it out as an annual subscription plugin for the global Nuke ecosystem.
Supporting Data and Production Metrics
The true metric of any visual effects tool is its impact on turnaround times in a live-fire pipeline. During Foundry’s controlled beta testing, professional roto artists were clocked working up to four times faster on benchmarked shots by the end of the second week.
While Adam Cherbetji cautioned against using the fourfold speedup as a universal production estimate—noting that it represented the higher ceiling of efficiency gains—he confirmed that artists consistently operated twice as fast, frequently seeing even higher multipliers depending on shot complexity.
Perhaps more telling than the raw quantitative data was the qualitative feedback. Beta participants reported that the tool fundamentally altered their workflow comfort, with many stating they "never wanted to roto without it again."
Benjamin Bratt, a compositor and author of Rotoscoping: Techniques and Tools for the Aspiring Artist, participated in the beta testing group and shared high praise:
"I’m shocked at how good SmartRoto is. It uses principles of rotoscoping and incorporates an intuitive, easily adaptable key frame approach that meshes with manual roto practices. It’s easy to fix auto-generated keys without needing to start from scratch."
Official Responses: The Philosophy of Splines Over Pixels
The primary engineering hurdle in automated rotoscoping has always been the edge: maintaining silhouette precision, temporal coherence, and an editable vector output. Speaking on why traditional approaches fail, Cherbetji dismissed optical flow outright:
"Not optical flow, that’s been tried. It was the first thing we attempted. Optical flow breaks down at the edge. That’s exactly where you need the precision."
Cherbetji was equally critical of popular market solutions that rely purely on segmentation masks wrapped in automated splines:
"A lot of the solutions we’ve seen in the market so far are using a segmentation model and drawing a spline around that. We don’t think that’s fit for purpose either."
By forcing the machine learning model to latch onto image features and propagate splines directly, SmartRoto preserves the core currency of the rotoscoping artist: subpixel-accurate vectors. Furthermore, the model uses spatial relationships between multiple shapes as mutual constraints. This means drawing a comprehensive breakdown of smaller, articulated shapes actively informs the AI where elements belong, preventing errors that occur when an artist relies on a single, lazy outline.
Regarding edge cases, SmartRoto handles motion blur remarkably well, typically placing edges directly at the midpoint of the blur while allowing artists to render motion blur natively into final mattes. While partial occlusions are smoothly accounted for, full occlusions require standard lifetime management. In scenarios featuring visually identical subjects—such as rows of Stormtroopers—distinct localized image features maintain tracking integrity, though edge cases can occasionally require manual intervention.
Implications for the VFX Industry and Competitive Landscape
SmartRoto steps into an industry-wide race to solve one of post-production’s most stubborn bottlenecks. On paper, rotoscoping appears to be a prime candidate for artificial intelligence: it is repetitive, highly defined, and backed by decades of ground-truth data locked inside studio archives. Yet, delivering an output that satisfies precision, temporal stability, and artist editability has proven notoriously elusive.
Foundry is not alone in tackling this frontier. Industry developers like Sam Hodge at Kognat have spent years cultivating AI-assisted roto workflows with impressive early results. Meanwhile, the academic research community continues to push parallel boundaries. At SIGGRAPH, papers like RotoShop—authored by Sirak Ghebremusse at OTOY—tackle the raster-to-vector pipeline by fitting pose-rigged Bézier splines onto segmentation masks derived from models like SAMv2.
The contrast in technical philosophy highlights a growing ideological split in the pipeline space:
The Segmentation-First Route (e.g., RotoShop): Works backward from raster masks, fitting vector data to pixel footprints. Its primary use case leans toward dataset generation and extreme compression.
The Spline-First Route (e.g., SmartRoto): Bypasses the intermediate matte entirely, propagating human-drawn vector shapes directly using image-feature tracking.
For compositors on the ground, the distinction is vital. As industry consensus dictates, any AI rotoscoping solution that concludes with uneditable pixel mattes rather than high-fidelity, subpixel-accurate vectors fundamentally misunderstands the demands of professional compositing.
By anchoring its AI model in the traditional methodology of breaking down shapes, establishing extremes, and locking user truths, Foundry’s SmartRoto signals a mature bridge between machine learning and human artistry—one where the computer takes over the exhausting mechanical labor, leaving the creative judgment firmly in the hands of the artist.