Cross-Modal Token Recognition via Associated Tokens

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing token recognition methods lack generality and accuracy, particularly when dealing with different modal data types, and fail to effectively recognize cross-modal tokens.

Innovation Solution

A method for recognizing tokens involves obtaining and processing first and second modal data, determining associated tokens between them, and recognizing a target shared token through fine-grained fusion and alignment using pre-trained neural networks and transformers, enhancing the generality and accuracy of token recognition across various data modalities.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If existing token recognition methods are used, then the process is simple, but the accuracy and generality of token recognition deteriorate when dealing with different modal data types

Engineering Contradiction:
Improvetoken recognition accuracyVSAvoidmethod complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent segments the token recognition process into distinct stages: obtaining multi-modal data, determining tokens for each modality separately, determining associated tokens between modalities, and recognizing target shared tokens. This segmentation allows each stage to be optimized independently, improving overall accuracy while managing complexity through structured processing.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces an intermediary mechanism by determining 'associated tokens' that bridge different modalities. These associated tokens serve as intermediaries that connect first modal data tokens and second modal data tokens, enabling accurate cross-modal token recognition while maintaining a systematic approach to handling complexity.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Adaptability or versatility

If existing token recognition methods are used, then the implementation is straightforward, but the ability to recognize cross-modal tokens deteriorates

Engineering Contradiction:
Improvecross-modal recognition capabilityVSAvoidimplementation ease
Core Design Contradiction:
Adaptability or versatilityVSEase of operation

Solution Approach 1:

The patent creates a universal token recognition framework that handles multiple data modalities (text, image, audio, etc.) through a unified process. The method determines tokens for different modalities and identifies shared tokens across modalities, enabling the system to adapt to various cross-modal scenarios while maintaining a consistent implementation approach.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The 'associated token' determination step serves as an intermediary mechanism that enables cross-modal recognition. By establishing associations between tokens of different modalities, the system achieves versatile cross-modal recognition capability while the structured four-step process maintains implementation ease through clear procedural guidance.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Measurement precision

If fine-grained fusion and alignment is applied, then the generality and accuracy of token recognition improve, but the computational complexity increases

Engineering Contradiction:
Improvetoken recognition accuracyVSAvoidcomputational resource consumption
Core Design Contradiction:
Measurement precisionVSUse of energy by moving object

Solution Approach 1:

The patent applies segmentation by dividing the computational process into four distinct steps: obtaining data, determining modality-specific tokens, determining associated tokens, and recognizing shared tokens. This segmentation allows computational resources to be allocated efficiently to each stage, achieving fine-grained fusion and alignment while managing computational complexity through structured processing.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent performs preliminary actions by first determining tokens for each modality separately before performing the more computationally intensive fusion and alignment operations. This preliminary token determination for each modality independently reduces the complexity of subsequent cross-modal matching, achieving accurate fine-grained recognition while optimizing computational resource usage.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS20230114673A1Method for recognizing token, electronic device and storage medium
Publication Date: 2023.04.13 BEIJING BAIDU NETCOM SCI & TECH CO LTD
  • US20230114673A1 patent drawing
  • US20230114673A1 patent drawing
  • US20230114673A1 patent drawing

AI summary

A method for recognizing a token is performed by an electronic device. The method includes: obtaining first modal data and second modal data; determining a first token of the first modal data and a second token of the second modal data; determining an associated token between the first token and the second token; and recognizing a target shared token between the first modal data and the second modal data based on the first token, the second token and the associated token.