Cross-Modal Token Recognition via Associated Tokens
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing token recognition methods lack generality and accuracy, particularly when dealing with different modal data types, and fail to effectively recognize cross-modal tokens.
Innovation Solution
A method for recognizing tokens involves obtaining and processing first and second modal data, determining associated tokens between them, and recognizing a target shared token through fine-grained fusion and alignment using pre-trained neural networks and transformers, enhancing the generality and accuracy of token recognition across various data modalities.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If existing token recognition methods are used, then the process is simple, but the accuracy and generality of token recognition deteriorate when dealing with different modal data types
Solution Approach 1:
The patent segments the token recognition process into distinct stages: obtaining multi-modal data, determining tokens for each modality separately, determining associated tokens between modalities, and recognizing target shared tokens. This segmentation allows each stage to be optimized independently, improving overall accuracy while managing complexity through structured processing.
Solution Approach 2:
The patent introduces an intermediary mechanism by determining 'associated tokens' that bridge different modalities. These associated tokens serve as intermediaries that connect first modal data tokens and second modal data tokens, enabling accurate cross-modal token recognition while maintaining a systematic approach to handling complexity.
2Adaptability or versatility
If existing token recognition methods are used, then the implementation is straightforward, but the ability to recognize cross-modal tokens deteriorates
Solution Approach 1:
The patent creates a universal token recognition framework that handles multiple data modalities (text, image, audio, etc.) through a unified process. The method determines tokens for different modalities and identifies shared tokens across modalities, enabling the system to adapt to various cross-modal scenarios while maintaining a consistent implementation approach.
Solution Approach 2:
The 'associated token' determination step serves as an intermediary mechanism that enables cross-modal recognition. By establishing associations between tokens of different modalities, the system achieves versatile cross-modal recognition capability while the structured four-step process maintains implementation ease through clear procedural guidance.
3Measurement precision
If fine-grained fusion and alignment is applied, then the generality and accuracy of token recognition improve, but the computational complexity increases
Solution Approach 1:
The patent applies segmentation by dividing the computational process into four distinct steps: obtaining data, determining modality-specific tokens, determining associated tokens, and recognizing shared tokens. This segmentation allows computational resources to be allocated efficiently to each stage, achieving fine-grained fusion and alignment while managing computational complexity through structured processing.
Solution Approach 2:
The patent performs preliminary actions by first determining tokens for each modality separately before performing the more computationally intensive fusion and alignment operations. This preliminary token determination for each modality independently reduces the complexity of subsequent cross-modal matching, achieving accurate fine-grained recognition while optimizing computational resource usage.
Data Source
AI summary
A method for recognizing a token is performed by an electronic device. The method includes: obtaining first modal data and second modal data; determining a first token of the first modal data and a second token of the second modal data; determining an associated token between the first token and the second token; and recognizing a target shared token between the first modal data and the second modal data based on the first token, the second token and the associated token.


