Workspace Image Representation With Geometric Reference for Neural Networks
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing neural networks for image processing in automated vehicle or robot environments require extensive manual labeling for supervised training and lack clear geometric information, making them inefficient and error-prone.
Innovation Solution
Represent input images as a superposition of location-dependent functions, enabling unsupervised training and providing clearer semantic meaning for task networks, allowing for better preparation and plausibility checks.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If neural networks use convolutional layers to transform input images into feature maps with lower dimensionality, then the network can process images efficiently, but the representations lose clear geometric reference to reality
Solution Approach 1:
The patent introduces an intermediary representation layer between the convolutional feature maps and the task network. This intermediary uses location-dependent basis functions (such as Gaussian functions centered at image locations) to transform the abstract feature maps into a representation that maintains geometric reference to the original image. The basis functions act as a mediator that preserves spatial information while enabling efficient processing.
Solution Approach 2:
The patent transforms the problem from the standard image coordinate system into a new dimensional space defined by location-dependent basis functions. Instead of working directly with pixel coordinates or feature map indices, the representation is expressed in terms of spatially-varying basis functions with parameters describing their location, scale, and orientation. This dimensional transformation preserves geometric information in a more interpretable form.
2Measurement precision
If neural networks are trained in a supervised manner with manually labeled training examples, then the network can learn to perform tasks accurately, but the training process becomes expensive and time-consuming
Solution Approach 1:
The patent enables the neural network to perform self-supervised learning by leveraging the geometric structure of the image data itself. The location-dependent basis function representation provides an inherent supervisory signal through its spatial coherence properties, allowing the network to learn meaningful representations without requiring external manual labels. The system serves itself by using the image's own geometric structure as the training objective.
Solution Approach 2:
The patent changes the parameter space of the representation from abstract feature activations to physically interpretable parameters of basis functions (location, scale, orientation). This parameter transformation enables unsupervised learning objectives based on spatial consistency and geometric priors, eliminating the need for expensive manual labeling while maintaining learning effectiveness.
3Adaptability or versatility
If the representation in the workspace has abstract feature maps, then the network can generalize to unseen situations, but the representations lack clear semantic meaning for plausibility checks
Solution Approach 1:
The patent segments the image representation into location-specific basis function components, each with clear semantic meaning related to its spatial position and properties. Rather than having monolithic feature maps, the representation is divided into contributions from individual basis functions centered at different locations, making it possible to interpret and verify the plausibility of each component separately while maintaining overall generalization capability.
Data Source
AI summary
A method for processing an input image using a task network trained to produce output with regard to a specified task from a representation of the input image in a workspace. The method includes: representing the input image as a superposition of functions that provide location-dependent contributions to the input image; generating a representation of the input image in a workspace from parameters that characterize this superposition; feeding this representation to the task network so that the task network ascertains the output with regard to the specified task. A method for transforming an input image into a superposition of functions, and a method for training a decomposition network for use in the method, are also described.


