Multimodal Asset Recognition Using Heterogeneous Transformer Models
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing asset recognition technologies face challenges such as labor-intensive manual labeling, high costs, and limitations in accuracy due to variations in lighting and viewing angles, as well as reliance on high-quality image systems.
Innovation Solution
A multimodal data heterogeneous Transformer-based asset recognition method that combines ALBERT, ViT, and CLIP models to extract features from text and image information, and uses discriminative loss for class discrimination learning to improve recognition accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Extent of automation
If identifier-based asset recognition technology is used, then asset identification can be achieved, but manual labeling is labor-intensive and identifiers can become defunct or damaged
Solution Approach 1:
The patent replaces mechanical/physical identifier systems (barcodes, RFID tags requiring manual labeling and scanning devices) with an image recognition system using cameras and computer vision algorithms. The system captures images of assets and automatically identifies them through feature extraction and pattern recognition, eliminating the need for manual labeling while avoiding the limitations of physical identifiers.
2Extent of automation
If image recognition-based asset recognition technology is used, then automatic identification without identifiers is achieved, but recognition accuracy fluctuates under variations in lighting conditions and viewing angles
Solution Approach 1:
The patent implements a feedback mechanism where the system captures multiple images of the same asset from different angles and under varying lighting conditions. The image recognition algorithm processes these images, and the system uses feedback from comparing multiple views to refine and verify identification results, thereby maintaining high accuracy despite variations in lighting and viewing angles.
Solution Approach 2:
The patent transitions from two-dimensional image analysis to three-dimensional spatial analysis by capturing images from multiple angles and positions. This multi-view geometry approach allows the system to reconstruct the asset's appearance in 3D space, providing more comprehensive features for accurate recognition regardless of the specific viewing angle or lighting condition in any single image.
3Productivity
If image recognition technology is used, then asset classification can be performed, but high computational power and huge storage resources are required
Solution Approach 1:
The patent performs preliminary actions by extracting and storing only the most salient features from images during the capture phase. Instead of storing and processing entire high-resolution images, the system pre-processes images to identify and store key visual features (such as shape descriptors, color histograms, or extracted patterns), significantly reducing the storage requirements and computational load for subsequent classification tasks.
Data Source
AI summary
This invention discloses a multimodal data heterogeneous Transformer-based asset recognition method, system, and device, the method including: collecting various-modal information of an asset, including text information and image information; building an ALBERT model, a ViT model, and a CLIP model; by the ALBERT model, extracting a text information feature; by the ViT model, extracting an image information feature; by the CLIP model, extracting image-text matching information feature; by different channels, applying asset type recognition to information in different modalities; outputting classification information from the different channels; by the CLIP model, generating asset void information; and discriminatively fusing the classification information from the different channels with the matching degree between the image information and the text information obtained by the CLIP model, and outputting final asset class information. This invention realizes comprehensive discrimination by drawing from multiple modalities to improve the accuracy of asset recognition.


