Flagman traffic gesture recognition

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Autonomous vehicles face challenges in interpreting real-time gestures from flagmen for safe navigation, as existing methods rely on pre-programmed data or maps, which are inadequate for immediate and variable environmental conditions.

Innovation Solution

A system utilizing a camera and neural networks to generate encoded hand vectors, skeletons, and representation vectors from images of flagmen, combined with recurrent neural networks to predict gestures and operate the vehicle accordingly, allowing for immediate and accurate interpretation of flagmen commands.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If pre-programmed methods or map databases are used to determine traffic control commands, then the system has simple processing logic, but it cannot respond to variable environmental conditions and flagmen gestures

Engineering Contradiction:
Improveability to respond to flagmen gesturesVSAvoidprocessing system complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent replaces traditional mechanical rule-based traffic control systems with a neural network-based vision system. The neural network processes images of flagmen and gestures to determine traffic control commands, enabling the autonomous vehicle to adapt to variable environmental conditions and understand human gestures without complex pre-programming of every possible scenario

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system uses image copying and processing to capture and analyze visual information about flagmen and their gestures. By converting real-world visual scenes into digital images and processing them through neural networks, the system can recognize gestures and determine appropriate responses without physical interaction or complex mechanical sensors

Inventive Principle:
Principle #26Copying

2Measurement precision

If complex neural network processing is used to recognize gestures accurately, then measurement precision of gestures improves, but processing time and latency increase

Engineering Contradiction:
Improvegesture recognition accuracyVSAvoidprocessing latency
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The neural network processing is segmented into multiple specialized networks: a first neural network extracts gesture features from images, a second neural network determines gesture meaning from those features, and a third neural network translates gestures into traffic control commands. This segmentation allows each network to be optimized for its specific task, improving overall accuracy while reducing total processing time compared to a single monolithic network

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system performs preliminary action by continuously capturing and pre-processing images of the environment, even before a specific gesture recognition task is initiated. The neural networks are pre-trained with large datasets of gestures and scenarios, so when a gesture needs to be recognized, the system can quickly process the input without extensive computation, reducing latency while maintaining high accuracy

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS11893801B2Flagman traffic gesture recognition
Publication Date: 2024.02.06 GM GLOBAL TECHNOLOGY OPERATIONS LLC
  • US11893801B2 patent drawing
  • US11893801B2 patent drawing
  • US11893801B2 patent drawing

AI summary

A vehicle and a system and method of operating the vehicle based on a gesture made by a traffic director. The system includes a camera and at least one neural network. The camera obtains an image of a flag operator. The at least one neural network is to generates an encoded hand vector based on a configuration of a hand of the traffic director from the image, combines a skeleton of the traffic director generated from the image and the encoded hand vector to generate a representation vector, and predicts a gesture of the traffic director from the representation vector.