Early Fusion Neural Ray Graph Networks for 3D Object Detection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing vehicle assistance systems struggle with providing enhanced situational awareness to drivers, particularly in situations where distractions or rapid environmental changes occur, leading to potential collisions.

Innovation Solution

The implementation of a multi-camera setup that uses early fusion of images represented as neural rays, organized into an ordered set and processed through a graph network, to create a unified feature set for improved 3D object detection and vehicle assistance.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If traditional image processing methods are used, then the system complexity is lower, but the 3D object detection accuracy and situational awareness are insufficient

Engineering Contradiction:
Improve3D object detection accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent segments the image processing task into distinct stages: neural ray generation from image frames, graph network construction for spatial relationship modeling, and feature set determination for 3D object detection. This segmentation allows each component to specialize in specific computations, improving overall accuracy while managing complexity through modular architecture.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent transitions from 2D image frames to 3D neural rays representing spatial positions of pixels. By projecting pixels onto 3D space and organizing them into rays extending from camera centers, the system captures depth information and spatial relationships, enabling accurate 3D object detection while maintaining computational efficiency through structured representation.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Productivity

If early fusion of images is implemented, then the 3D object detection efficiency is improved, but the processing complexity increases

Engineering Contradiction:
Improve3D object detection efficiencyVSAvoidprocessing complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent performs preliminary actions by generating neural rays and constructing graph networks before final feature extraction. Neural rays are pre-computed from image frames, establishing 3D spatial relationships in advance. Graph networks are built to encode spatial proximity between points, enabling efficient subsequent processing and 3D object detection without re-computing spatial relationships.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent introduces neural rays and graph networks as intermediary structures between raw image frames and final 3D object detection outputs. Neural rays serve as mediators that transform 2D pixel data into 3D spatial representations. Graph networks act as intermediaries that model spatial relationships, enabling efficient feature extraction and object detection through structured data representation.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Reliability

If multi-camera setup with neural rays is used, then the situational awareness is enhanced, but the computational requirements increase

Engineering Contradiction:
Improvesituational awarenessVSAvoidcomputational requirements
Core Design Contradiction:
ReliabilityVSUse of energy by moving object

Solution Approach 1:

The patent merges information from multiple cameras by projecting pixels from different camera views onto a unified 3D space. Neural rays from multiple cameras are combined and organized into a coherent graph network structure, integrating spatial information from all cameras into a single representation that enhances situational awareness while sharing computational workload.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent changes the representation parameters from raw pixel data to neural rays defined by 3D positions and spatial relationships. By transforming data into this structured format with explicit spatial encoding, the system reduces redundant computations and optimizes energy usage while maintaining enhanced situational awareness through accurate 3D object detection.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS20250157178A1Early fusion of neural ray graph networks for multi-view camera setups
Publication Date: 2025.05.15 QUALCOMM INC
  • US20250157178A1 patent drawing
  • US20250157178A1 patent drawing
  • US20250157178A1 patent drawing

AI summary

This disclosure provides systems, methods, and devices for vehicle driving assistance systems that support image processing. In a first aspect, an image processing method includes receiving image frames; determining an ordered set of neural rays based on the image frames; determining a graph network that represents each neural ray of the ordered set of neural rays as a sequence of points; and determining a feature set based on the graph network. Each neural ray of the ordered set of neural rays represents three-dimensional positions of pixels of an image frame. Each point on the graph network is associated with a node of a plurality of nodes of the graph network. The feature set includes features of each of the image frames. Other aspects and features are also claimed and described.