Video-Based Robotic Assembly Instruction Generation Without Object Profiles

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Programming robotic machines for assembly tasks is time-consuming and resource-intensive, requiring multiple iterations and consuming significant power and processing resources.

Innovation Solution

A method that uses a video encoding frames of an assembly process to determine spatio-temporal features, identify actions, and generate an assembly plan, which includes combining output from a point cloud model and a color embedding model to calculate coordinates and perform object segmentation to estimate grip points and widths.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If traditional programming methods are used to program robotic machines for assembly tasks, then the robotic machines can perform assembly operations, but the process is time-consuming and consumes significant power and processing resources

Engineering Contradiction:
Improveassembly programming efficiencyVSAvoidprogramming time
Core Design Contradiction:
ProductivityVSLoss of time

Solution Approach 1:

The system captures video footage of human assembly operations and uses computer vision to automatically generate robotic programming instructions by copying the demonstrated actions. This eliminates manual programming time while preserving the assembly expertise embedded in human operators' movements and techniques.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The robotic system performs self-programming by automatically analyzing video data of assembly operations and generating its own control instructions without requiring external manual programming. The system extracts motion trajectories, tool paths, and assembly sequences autonomously from the video content.

Inventive Principle:
Principle #25Self-service

2Extent of automation

If complex hardware like AR markers and motion sensors are used to program robotic machines, then assembly operations can be captured, but power and processing resources are consumed

Engineering Contradiction:
Improveassembly capture capabilityVSAvoidpower consumption
Core Design Contradiction:
Extent of automationVSUse of energy by moving object

Solution Approach 1:

The system extracts assembly information directly from standard video footage without requiring specialized tracking hardware like AR markers or motion sensors. By removing these additional components, the system significantly reduces power consumption and processing requirements while maintaining the ability to capture assembly operations.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The system uses ordinary video recording technology instead of expensive, power-intensive specialized hardware. Standard cameras and video processing algorithms replace costly motion capture systems, making the solution more energy-efficient and accessible.

Inventive Principle:
Principle #27Cheap short-living objects (Disposable)

3Manufacturing precision

If pre-existing profiles for sub-objects are required to generate robotic instructions, then accurate manipulation can be achieved, but memory space and processing resources are needed

Engineering Contradiction:
Improvemanipulation accuracyVSAvoidmemory space
Core Design Contradiction:
Manufacturing precisionVSQuantity of substance

Solution Approach 1:

The system performs preliminary analysis of sub-objects directly from video frames by detecting edges, contours, and geometric features. This preliminary extraction of object properties from visual data eliminates the need to store pre-existing profiles in memory, as all necessary information is derived on-demand from the video content.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system uses video frames as an intermediary to convey sub-object information. Instead of storing detailed profiles in memory, the system processes visual information from video frames to extract necessary geometric and physical properties, reducing memory requirements while maintaining manipulation accuracy.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS12343884B2Robotic assembly instruction generation from a video
Publication Date: 2025.07.01 ACCENTURE GLOBAL SOLUTIONS LTD
  • US12343884B2 patent drawing
  • US12343884B2 patent drawing
  • US12343884B2 patent drawing

AI summary

In some implementations, a robot host may receive a video associated with assembly using a plurality of sub-objects. The robot host may determine spatio-temporal features based on the video and may identify a plurality of actions represented in the video based on the spatio-temporal features. The robot host may map the plurality of actions to the plurality of sub-objects to generate an assembly plan and may combine output from a point cloud model and output from a color embedding model to generate a plurality of sets of coordinates corresponding to the plurality of sub-objects. The robot host may perform object segmentation to estimate a plurality of grip points and a plurality of widths corresponding to the plurality of sub-objects. Accordingly, the robot host may generate instructions, for robotic machines, based on the assembly plan, the plurality of sets of coordinates, the plurality of grip points, and the plurality of widths.