Multimodal Asset Recognition Using Heterogeneous Transformer Models

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing asset recognition technologies face challenges such as labor-intensive manual labeling, high costs, and limitations in accuracy due to variations in lighting and viewing angles, as well as reliance on high-quality image systems.

Innovation Solution

A multimodal data heterogeneous Transformer-based asset recognition method that combines ALBERT, ViT, and CLIP models to extract features from text and image information, and uses discriminative loss for class discrimination learning to improve recognition accuracy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Extent of automation

If identifier-based asset recognition technology is used, then asset identification can be achieved, but manual labeling is labor-intensive and identifiers can become defunct or damaged

Engineering Contradiction:
Improveasset identificationVSAvoidmanual labeling time
Core Design Contradiction:
Extent of automationVSLoss of time

Solution Approach 1:

The patent replaces mechanical/physical identifier systems (barcodes, RFID tags requiring manual labeling and scanning devices) with an image recognition system using cameras and computer vision algorithms. The system captures images of assets and automatically identifies them through feature extraction and pattern recognition, eliminating the need for manual labeling while avoiding the limitations of physical identifiers.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Extent of automation

If image recognition-based asset recognition technology is used, then automatic identification without identifiers is achieved, but recognition accuracy fluctuates under variations in lighting conditions and viewing angles

Engineering Contradiction:
Improveautomatic identificationVSAvoidrecognition accuracy
Core Design Contradiction:
Extent of automationVSMeasurement precision

Solution Approach 1:

The patent implements a feedback mechanism where the system captures multiple images of the same asset from different angles and under varying lighting conditions. The image recognition algorithm processes these images, and the system uses feedback from comparing multiple views to refine and verify identification results, thereby maintaining high accuracy despite variations in lighting and viewing angles.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The patent transitions from two-dimensional image analysis to three-dimensional spatial analysis by capturing images from multiple angles and positions. This multi-view geometry approach allows the system to reconstruct the asset's appearance in 3D space, providing more comprehensive features for accurate recognition regardless of the specific viewing angle or lighting condition in any single image.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

3Productivity

If image recognition technology is used, then asset classification can be performed, but high computational power and huge storage resources are required

Engineering Contradiction:
Improveasset classification capabilityVSAvoidcomputational power consumption
Core Design Contradiction:
ProductivityVSUse of energy by moving object

Solution Approach 1:

The patent performs preliminary actions by extracting and storing only the most salient features from images during the capture phase. Instead of storing and processing entire high-resolution images, the system pre-processes images to identify and store key visual features (such as shape descriptors, color histograms, or extracted patterns), significantly reducing the storage requirements and computational load for subsequent classification tasks.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS12236699B1Multimodal data heterogeneous transformer-based asset recognition method, system, and device
Publication Date: 2025.02.25 JINAN UNIVERSITY
  • US12236699B1 patent drawing
  • US12236699B1 patent drawing
  • US12236699B1 patent drawing

AI summary

This invention discloses a multimodal data heterogeneous Transformer-based asset recognition method, system, and device, the method including: collecting various-modal information of an asset, including text information and image information; building an ALBERT model, a ViT model, and a CLIP model; by the ALBERT model, extracting a text information feature; by the ViT model, extracting an image information feature; by the CLIP model, extracting image-text matching information feature; by different channels, applying asset type recognition to information in different modalities; outputting classification information from the different channels; by the CLIP model, generating asset void information; and discriminatively fusing the classification information from the different channels with the matching degree between the image information and the text information obtained by the CLIP model, and outputting final asset class information. This invention realizes comprehensive discrimination by drawing from multiple modalities to improve the accuracy of asset recognition.