Cross-modal data retrieval method based on voice intention
By improving the Paraformer algorithm and optimizing speech recognition technology with a domain-specific lexicon, and combining multimodal feature extraction with the Elasticsearch database, the problems of speech recognition adaptability and retrieval efficiency in cross-modal data retrieval are solved, achieving accurate and efficient multimodal data retrieval.
Patent Information
- Application Number
- CN202510774994.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-11
- Publication Date
- 2025-11-25
AI Technical Summary
Existing technologies suffer from poor adaptability to speech recognition, insufficient support for multimodal retrieval, and low retrieval efficiency in cross-modal data retrieval, especially in terms of low accuracy and weak generalization ability when processing speech queries.
An improved Paraformer speech recognition algorithm is adopted and combined with a domain lexicon to optimize speech recognition. Structured intent representation is generated through intent parsing, and multimodal feature vectors are extracted by combining Sentence-Transformer and EfficientNet models. Cross-modal feature association and matching are performed using the Elasticsearch database, and the model is optimized based on user feedback.
It achieves accurate semantic understanding and efficient multimodal retrieval, supports multimodal retrieval of images, text, videos and structured libraries, adapts to the needs of specific domains, provides personalized retrieval services, and has high retrieval timeliness.
Smart Images

Figure CN121009215A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of speech recognition, natural language processing, computer vision and multimedia retrieval, and in particular to a cross-modal data retrieval method based on speech intent. BACKGROUND
[0002] With the rapid growth of multi-modal data such as images, videos, texts and structured data, traditional single-modal retrieval methods have been unable to meet the needs of users for accurate retrieval. The current mainstream cross-modal retrieval methods are mostly based on deep learning, using double-flow neural networks or shared embedding spaces to project data of different modalities into a unified semantic space for similarity calculation. However, existing methods mainly rely on static modalities such as text or images as query inputs, and there is relatively little research on speech as a query entry, and there are bottlenecks such as low accuracy and weak generalization ability in processing the complex semantic intent implied in speech:
[0003] 1) Poor adaptability in the field of speech recognition: general speech recognition models perform poorly in specific fields and cannot accurately recognize field-related vocabulary.
[0004] 2) Insufficient support for multi-modal retrieval: existing systems usually only support single-modal (such as image or text) retrieval, lacking unified retrieval capabilities for multi-modal data.
[0005] 3) Low retrieval efficiency: the retrieval efficiency of modal data is low, making it difficult to meet the timeliness requirements of heterogeneous data retrieval. SUMMARY
[0006] The technical problem to be solved by the present application is to provide a cross-modal data retrieval method based on speech intent to address the shortcomings of the prior art.
[0007] To solve the above technical problems, the present application discloses a cross-modal data retrieval method based on speech intent, which specifically comprises:
[0008] Step 1: input the speech corresponding to the user query request, convert the speech into text by using an improved Paraformer speech recognition algorithm, and optimize the recognition effect by combining a domain vocabulary library;
[0009] Step 2: perform intent analysis on the text generated by speech recognition to output a structured intent representation, perform word segmentation on the keywords of the speech recognition result, and extract domain labels by combining a domain vocabulary library;
[0010] Step 3: Extract image features and video temporal features from local image and video data to generate image feature vectors and visual feature vectors; semantically encode local text data using the Sentence-Transformer model to generate text feature vectors; semantically encode local structured data using the Embedding model to generate database feature vectors, i.e., structured features; establish a cross-modal feature association space to map visual feature vectors, text feature vectors, and structured feature vectors to the same semantic space; store the feature vectors, text feature vectors, and structured feature vectors of image and video keyframes, along with their domain labels, in the Elasticsearch ES vector database, with each vector associated with a corresponding domain label.
[0011] Step 4: Perform two-stage matching based on the query information entered by the user, and conduct multimodal retrieval to obtain candidate results;
[0012] Step 5: Filter the candidate results based on the intent parsing results (i.e. intent conditions). Using the filtering conditions ("HD" (video resolution ≥1080p)), filter out videos with the tags "tracked engineering vehicle" and "desert terrain" to obtain the final query results (including images, text, structured database tables, or audio and video).
[0013] Step 6: Optimize the model based on user feedback and maintain and update the vector library in real time through the online learning module.
[0014] Step 2 is as follows:
[0015] Step 2-1: Use the Fine-tuned BERT algorithm to perform intent parsing on the text generated by speech recognition, extract the semantic information of the query, extract the user's core needs (such as "desert" and "construction vehicle"), identify the user's implicit needs (such as "high-definition" and "latest"), and output a structured intent representation; for example:
[0016] {
[0017] "intent":"search_video",
[0018] "keywords":["desert","engineering vehicle"],
[0019] "filters":{"quality":"High Definition","date":"Latest"}
[0020] }
[0021] The symbols {} and [] serve as identifiers and have no special meaning.
[0022] Step 2-2, use the HanLP tokenizer to tokenize the keywords in the intent analysis result, and extract domain labels from the structured intent representation combined with the domain lexicon. For example, extract the labels from "find the video of the engineering truck in the desert": "desert", "engineering truck", "video".
[0023] The image feature extraction in step 3 is specifically:
[0024] Use the EfficientNet model to extract image feature vectors, optimize the depth, width and resolution of the model through compound scaling, which can extract high-quality features at a lower computational cost, and output 2048-dimensional image feature vectors.
[0025] The video feature extraction in step 3 is specifically:
[0026] Step 3-1, detect the shot boundary in the video by frame difference method;
[0027] Step 3-2, in each shot, use K-Means clustering algorithm to select the most representative key frame;
[0028] Step 3-3, for the extracted key frame, use EfficientNet network to extract feature vector.
[0029] Step 4 is specifically:
[0030] Step 4-1, use Sentence-Transformer to convert the structured representation of the user query into a query vector;
[0031] Step 4-2, use cosine similarity for similarity search in the ES vector database, find the most similar feature vector to the query vector, and return the result through approximate nearest neighbor search (ANN);
[0032] Step 4-3, extract the domain label of the candidate result, match with the domain label of the user query, and filter out the most relevant result;
[0033] Filtering and sorting: filter the candidate results by combining vector similarity and label matching degree, keep the data above the threshold, comprehensive score = a · vector similarity + (1-a) · label matching degree comprehensive score; Where, a is the weight parameter, usually set to 0.6.
[0034] Step 5 is specifically:
[0035] Step 5-1: Extract implicit requirements (such as "high definition" and "latest") from the intent analysis results as filtering conditions for the search. Convert the filtering conditions into search parameters (such as "high definition" → video resolution ≥ 1080p; "latest" → release date within the past 30 days).
[0036] Step 5-2: Calculate the matching degree of the filtering conditions and candidate results using Jaccard similarity. Return the refined search results that meet the filtering conditions.
[0037] The improved speech recognition Paraformer algorithm in step 1 is specifically:
[0038] Freeze the last 2 layers of the Paraformer encoder, add a nonlinear activation function layer, and improve the model's expression ability.
[0039] Step 1-1-1: Pre-emphasize and frame the input 16kHz sampling rate audio signal, extract 80-dimensional Fbank acoustic features, and perform energy normalization.
[0040] Step 1-1-2: Map the acoustic features to a 512-dimensional vector through a linear projection layer and add position encoding. Input them into a 12-layer Transformer encoder for context modeling. The first 10 layers are kept frozen, and only the last 2 layers are fine-tuned. Insert a Swish nonlinear activation function in the feedforward network to enhance the model's expression ability.
[0041] Step 1-1-3: Use the encoder's output as the input of the LSTM predictor. Fuse the encoder's output with the historical information of the 256-dimensional unidirectional LSTM predictor through a joint network. Train end-to-end using the CTC loss function and the forward-backward algorithm.
[0042] Step 1-1-4: Select the maximum probability token frame by frame through the greedy decoding strategy, remove repeated symbols and white spaces, and re-score (weight α = 0.3) combined with the 4-gram language model to output the final text recognition result.
[0043] The intent understanding in step 2-1 is specifically:
[0044] Step 2-1-1: Text preprocessing and segmentation. Standardize the input user text (remove special symbols, unify case), use the WordPiece segmentation algorithm to divide the text into sub-word units, add special markers [CLS] and [SEP] to form the model input sequence, and through padding pooling and truncation, unify the sequence length to 512 tokens, while generating an attention mask to identify valid positions.
[0045] Step 2-1-2: BERT encoder feature extraction, the tokenized sequence is mapped to token_ids by the vocabulary, and the position encoding and paragraph encoding are added to form the input embedding vector, which is input into the 12-layer Transformer encoder for bidirectional context modeling, each layer containing 768-dimensional hidden states and 12 multi-head attention mechanisms Attention (Q, K, V):
[0046]
[0047] Q (Query) is obtained by projecting the input embedding through the matrix W Q , which represents the current word's query of other words; K (Key): the input embedding is projected through the matrix W K , which represents the representation of each word itself; V (Value): the input embedding is projected through the matrix W V , which represents the information passed to other words by each word; d k : is the dimension of the Key vector, used to scale the dot product to avoid softmax gradient vanishing (usually d k = 64); W: represents the weight matrix in the attention module, such as W Q , W K , W V , whose dimensions are (BERT-base).
[0048] On the basis of the attention weight matrix W, a low-rank matrix ΔW = BA T is injected, where the small matrix represents the mapping from the original space (768 dimensions) to a low-dimensional space (r = 8 dimensions), and the small matrix maps back from the low-dimensional space (r = 8 dimensions) to the original space (768 dimensions).
[0049] During training, only the low-rank matrix parameters are updated, and the original BERT parameters are frozen, and the domain-specific data is used for comparative learning fine-tuning.
[0050] Step 2-1-3: intent classification head calculation, extract the 768-dimensional context representation vector of the [CLS] position as the global semantic feature, pass it through a linear classification layer for intent category mapping, and apply the softmax activation function to calculate the probability distribution P (intent | text) = softmax (logits) of each intent category, select the class with the highest probability as the predicted intent:
[0051]
[0052] Wherein, text represents the input text, intent represents the intent category (such as "play music", "query weather"), Logits represents the hidden state vector of the [CLS] position in the BERT output, and represents the semantic representation of the whole sentence; represents the weight matrix of the intent classification head, which maps the [CLS] vector to C intent categories;
[0053] represents the bias term of the classification head; represents the unnormalized score corresponding to each intent category; C represents the total number of intent categories; softmax represents the activation function for normalizing the logits vector into a probability distribution;
[0054] Step 2-1-4: Slot annotation and entity extraction, BIO annotation is performed on the hidden state of each token position in the sequence (as shown in Table 1), and the transition probability between labels is modeled through the CRF (Conditional Random Field) layer, and the loss function is L_slot = -log P(tag_sequence|token_sequence), wherein tag_sequence is the true label sequence (such as [B-date, I-date, O, B-city]), represents the slot label of each token, and token_sequence represents the input token sequence (such as ["tomorrow", "go", "Beijing"]), the key entity information (such as time, place, name, etc.) in the user input is extracted, and the semantic alignment verification is performed with the intent category;
[0055] Table 1. Meaning of BIO annotation
[0056]
[0057] Step 2-1-5: Structured intent representation generation, the intent classification result, slot annotation result and confidence score are integrated into a JSON format structured representation: {"intent":"predicted intent", "confidence":0.95, "slots":[{"entity":"entity value", "type":"entity type", "start":start position, "end":end position}]}, the rule engine is used for post-processing and consistency checking, and finally the complete user intent understanding result is output for downstream tasks.
[0058] The combination of domain keyword label extraction in step 2-2 includes a three-level filtering mechanism:
[0059] Step 2-2-1: Field word preliminary screening based on TF-IDF (Term Frequency-Inverse Document Frequency), retaining keywords with weight values >0.7;
[0060] Step 2-2-2: Calculate semantic similarity through pre-trained word vectors, and select candidate words with cosine similarity >0.85 with the field word library;
[0061] Step 2-2-3: Key importance evaluation based on attention weight, retaining words with semantic head attention weight >0.3.
[0062] The two-stage matching in Step 4 includes:
[0063] Coarse-grained matching, calculate the cosine similarity of query vector to feature library, select TOP-K candidate set, where K can be set, usually 50;
[0064] Fine-grained matching, construct a screening matrix M ∈ R {N×T} where N is the number of candidates, T is the label dimension; through M × Q T weighted ranking, Q is the query label vector.
[0065] The online learning module in Step 6 adopts a double-buffer incremental update mechanism, including:
[0066] Step 6-1: Real-time maintenance and update, periodically (e.g. every 24 hours) collect new data to build vector library and update field word library and label weight;
[0067] Step 6-2: Batch check update: periodically (e.g. every 7 days) use sliding window sampling mechanism, use KL divergence (Kullback-Leibler divergence) to detect data distribution shift, when KL divergence meets threshold >1.5, trigger model retraining;
[0068] Step 6-3: Version verification rollback: when the quality of random validation set decreases, you can choose to roll back to the previous stable version.
[0069] Beneficial effects:
[0070] 1. Accurate semantic understanding: through Paraformer algorithm and field word fine-tuning, the system can improve the domain adaptability of speech recognition and accurately understand closed professional semantic information.
[0071] 2. Efficient retrieval performance: based on ES vector database and field label filtering, significantly improve retrieval efficiency and accuracy.
[0072] 3. Multi-modal search support: Supports multi-modal search of images, text, video and structured libraries, meeting diverse user needs.
[0073] 4. Domain adaptability: Through closed-source professional domain vocabulary, the system can adapt to the needs of specialized fields and provide personalized search services.
[0074] 5. Search timeliness: Based on the ES vector library retrieval mechanism, vector matching is extremely fast, supporting fast retrieval and timely response of large-scale data. BRIEF DESCRIPTION OF DRAWINGS
[0075] Figure 1 The system architecture of the present application.
[0076] Figure 2 The workflow diagram of the present application. DETAILED DESCRIPTION
[0077] The present application proposes a cross-modal data retrieval method based on speech intent combining multi-modal vector extraction, domain word fine-tuning and vector matching retrieval.
[0078] The multi-modal data includes text data, image data, video data and structured data;
[0079] Embodiment:
[0080] This embodiment takes the scenario of searching for video by voice as an example, the query request is "help me find some videos of tracked engineering vehicles in the desert, preferably in high definition", the local data includes images, text, videos, etc. Local data is data on the server locally, which can be changed arbitrarily, increased or decreased, and is not controlled by the algorithm, which belongs to user behavior guidance. The purpose of this patent is to find the data the user wants through voice search for the user's messy data.
[0081] Step 1, input the user input voice corresponding to the user query request "help me find some videos of tracked engineering vehicles in the desert, preferably in high definition", convert the voice into text by using the improved Paraformer speech recognition algorithm, and optimize the recognition effect by combining the domain vocabulary;
[0082] Step 2, perform intent analysis on the text generated by voice recognition to output structured intent representation, perform word segmentation on the keywords of the voice recognition result, and extract domain labels by combining the domain vocabulary;
[0083] Step 3: Extract image features and video temporal features from local image and video data to generate image feature vectors and visual feature vectors; Semantically encode local text data using the Sentence-Transformer model to generate text feature vectors; Semantically encode local structured data using the Embedding model to generate database feature vectors, i.e., structured features; Establish a cross-modal feature association space to map visual feature vectors, text feature vectors, and structured feature vectors to the same semantic space.
[0084] Local text data is converted into high-dimensional features, semantic information is extracted to generate strong representation vectors, structured data (such as equipment databases) is feature-encoded, and embedding technology is used to convert it into vector representation;
[0085] The feature vectors, text feature vectors, and structured feature vectors of keyframes in images and videos, along with their domain labels, are stored in the Elasticsearch ES vector database, with each vector associated with a corresponding domain label.
[0086] Step 4: Perform two-stage matching based on the query information entered by the user, conduct multimodal retrieval, and obtain candidate results;
[0087] Step 5: Filter the candidate results based on the intent parsing results (i.e. intent conditions). Using the filtering conditions ("HD" (video resolution ≥ 1080p)), select videos with the tags "tracked engineering vehicle" and "desert terrain" to obtain the final query results (including images, text, structured database tables, or audio and video), and display videos that highly match the query semantics and visual features.
[0088] Step 6: Optimize the model based on user feedback. Maintain and update the vector library in real time through the online learning module to achieve version iteration or rollback of the search algorithm.
[0089] Step 2 is as follows:
[0090] Step 2-1: Use the Fine-tuned BERT algorithm to perform intent parsing on the text generated by speech recognition, extract the semantic information of the query, extract the user's core needs (such as "desert" and "construction vehicle"), identify the user's implicit needs (such as "high-definition" and "latest"), and output a structured intent representation.
[0091] {
[0092] "intent":"search_video",
[0093] "keywords":["desert","engineering vehicle"],
[0094] "filters":{"quality":"High Definition","date":"Latest"}
[0095] }
[0096] Step 2-2: Use the HanLP word segmenter to segment the keywords in the intent parsing results, and extract domain tags from the structured intent representation using a domain thesaurus. Extract the tags "desert", "engineering vehicle", and "video" from "find videos of construction vehicles in the desert".
[0097] The image feature extraction described in step 3 specifically involves:
[0098] Using the EfficientNet model to extract image feature vectors, and optimizing the model's depth, width, and resolution through compound scaling, high-quality features can be extracted with low computational cost, outputting 2048-dimensional image feature vectors.
[0099] The video feature extraction described in step 3 specifically involves:
[0100] Step 3-1: Detect shot boundaries in the video using the frame difference method;
[0101] Step 3-2: In each shot, use the K-Means clustering algorithm to select the most representative keyframes;
[0102] Step 3-3: Use the EfficientNet network to extract feature vectors from the extracted keyframes.
[0103] Step 4 specifically involves:
[0104] Step 4-1: Use Sentence-Transformer to convert the structured representation of the user query into a query vector;
[0105] Step 4-2: Use cosine similarity to perform a similarity search in the ES vector database to find the feature vector most similar to the query vector, and return the result through Approximate Nearest Neighbor (ANN) search.
[0106] Step 4-3: Extract the domain labels of candidate results and match them with the domain labels queried by the user to filter out the most relevant results. Filtering and sorting: Combine vector similarity and label matching to filter candidate results, retaining data above the threshold. Overall score = α·vector similarity + (1-α)·label similarity overall score; where α is a weight parameter, usually set to 0.6.
[0107] Step 5 specifically involves:
[0108] Step 5-1: Extract implicit requirements (such as "high definition" and "latest") from the intent analysis results as filtering conditions for the search. Convert the filtering conditions into search parameters (such as "high definition" → video resolution ≥ 1080p; "latest" → release date within the past 30 days).
[0109] Step 5-2: Calculate the matching degree of the filtering conditions and candidate results using Jaccard similarity. Return the refined search results that meet the filtering conditions.
[0110] The improved speech recognition Paraformer algorithm in step 1 is specifically:
[0111] Freeze the last 2 layers of the Paraformer encoder, add a nonlinear activation function layer, and improve the model's expression ability.
[0112] Step 1-1-1: Pre-emphasize and frame the input 16kHz sampling rate audio signal, extract 80-dimensional Fbank acoustic features, and perform energy normalization.
[0113] Step 1-1-2: Map the acoustic features to a 512-dimensional vector through a linear projection layer and add position encoding. Input them into a 12-layer Transformer encoder for context modeling. The first 10 layers are kept frozen, and only the last 2 layers are fine-tuned. Insert a Swish nonlinear activation function in the feedforward network to enhance the model's expression ability.
[0114] Step 1-1-3: Use the encoder's output as the input of the LSTM predictor. Feature fusion is performed through a joint network using the encoder's output and the historical information of the 256-dimensional unidirectional LSTM predictor. End-to-end training is performed using the CTC loss function and the forward-backward algorithm.
[0115] Step 1-1-4: Select the maximum probability token frame by frame through the greedy decoding strategy, remove repeated symbols and white spaces, and re-score (weight α = 0.3) combined with the 4-gram language model to output the final text recognition result.
[0116] The intent understanding in step 2-1 is specifically:
[0117] Step 2-1-1: Text preprocessing and segmentation. Standardize the input user text (remove special symbols, unify case), use the WordPiece segmentation algorithm to divide the text into subword units, add special markers [CLS] and [SEP] to form the model input sequence, and through padding pooling and truncation, unify the sequence length to 512 tokens, while generating attention mask to identify valid positions.
[0118] Step 2-1-2: BERT encoder feature extraction, the tokenized sequence is mapped to token_ids by the vocabulary, and the position encoding and paragraph encoding are added to form the input embedding vector, which is input into the 12-layer Transformer encoder for bidirectional context modeling, each layer containing 768-dimensional hidden states and 12 multi-head attention mechanisms Attention (Q, K, V):
[0119]
[0120] Q (Query) is obtained by projecting the input embedding through the matrix W Q , which represents the current word's query of other words; K (Key): obtained by projecting the input embedding through the matrix W K , which represents the representation of each word itself; V (Value): obtained by projecting the input embedding through the matrix W V , which represents the information passed by each word to other words; d k : is the dimension of the Key vector, used to scale the dot product to avoid softmax gradient disappearance (usually d k = 64); W: represents the weight matrix in the attention module, such as W Q , W K , W V , which have dimensions (BERT-base).
[0121] Inject a low-rank matrix rank r = 8) into the attention weight matrix W, freeze the original BERT parameters during training and only update the low-rank matrix parameters, and fine-tune through comparative learning of domain-specific data;
[0122] Step 2-1-3: text represents the input text, intent represents the intent category (such as "play music", "query weather"), and the intent classification head is calculated to extract the 768-dimensional context representation vector at the [CLS] position as the global semantic feature, which is mapped to the intent category through a linear classification layer, and the probability distribution P (intent | text) = softmax (logits) is calculated by applying the softmax activation function, and the class with the highest probability is selected as the predicted intent:
[0123]
[0124] where Logits represents the hidden state vector at the [CLS] position in the BERT output, which represents the semantic representation of the entire sentence; represents the weight matrix of the intent classification head, which maps the [CLS] vector to C intent categories; represents the bias term of the classification head; represents the unnormalized score corresponding to each intent category; C represents the total number of intent categories; softmax represents an activation function for normalizing logits vector into probability distribution;
[0125] Step 2-1-4: Slot labeling and entity extraction, BIO labeling is performed on the hidden state of each token position in the sequence (as shown in Table 1), and the transition probability between labels is modeled through a CRF (Conditional Random Field) layer, and the loss function is L_slot = -log P(tag_sequence|token_sequence), wherein tag_sequence is the real label sequence (such as [B-date, I-date, O, B-city]), representing the slot label of each token, and token_sequence represents the input token sequence (such as ["tomorrow", "go", "Beijing"]), and the key entity information (such as time, place, and name) in the user input is extracted and verified for semantic alignment with the intent category;
[0126] Step 2-1-5: Structured intent representation generation, the intent classification result, slot labeling result, and confidence score are integrated into a JSON format structured representation: {"intent":"predicted intent", "confidence":0.95, "slots":[{"entity":"entity value", "type":"entity type", "start":start position, "end":end position}]}, and post-processing and consistency verification are performed through a rule engine, and finally the complete user intent understanding result is output for use by downstream tasks.
[0127] The domain keyword label extraction described in step 2-2 includes a three-level filtering mechanism:
[0128] Step 2-2-1: Initial screening of domain words based on TF-IDF (Term Frequency-Inverse Document Frequency), retaining keywords with a weight value >0.7;
[0129] Step 2-2-2: Calculate semantic similarity through pre-trained word vectors, and select candidate words with a cosine similarity to the domain word library >0.85;
[0130] Step 2-2-3: Key importance evaluation based on attention weight, retaining words with a semantic head attention weight >0.3.
[0131] The two-stage matching described in step 4 includes:
[0132] Coarse-grained matching, calculate the cosine similarity of the query vector to the feature library, and select the TOP-K candidate set, where K can be set, usually 50;
[0133] Fine-grained matching, constructing a screening matrix M ∈ R {N×T} where N is the number of candidates, T is the label dimension; through M × Q T Weighted sorting, Q is the query label vector.
[0134] The online learning module in step 6 adopts a double-buffer incremental update mechanism, including:
[0135] Step 6-1: Real-time maintenance and update, periodically (such as every 24 hours) collect new data to construct the vector library and update the domain word library language label weight;
[0136] Step 6-2: Batch check update: periodically (such as every 7 days) through a sliding window sampling mechanism, use KL divergence to detect data distribution offset, and trigger model retraining when the KL divergence meets the threshold > 1.5;
[0137] Step 6-3: Version verification rollback: when the quality of the random verification set decreases, the previous stable version can be selected to roll back.
[0138] The result output in step 6 is specifically:
[0139] According to the comprehensive score, the search results are sorted, and the sorted data is displayed to the user, supporting mixed display of multi-modal content.
[0140] The present application includes feedback and optimization:
[0141] User feedback data collection: collect user behavior data such as clicking, collecting, and scoring on search results;
[0142] Model optimization: according to user feedback data, use online learning (Online Learning) method to continuously optimize Paraformer model, intent recognition model and feature extraction model;
[0143] Vector library update: regularly update the feature vectors and domain labels in the ES vector database to ensure the timeliness of the system.
[0144] The present application provides a cross-modal data retrieval method based on voice intent, and there are many methods and ways to realize this technical solution. The above description is only the preferred embodiment of the present application, and it should be pointed out that for ordinary technical personnel in this technical field, without departing from the principle of the present application, some improvements and refinements can be made, which should be regarded as the protection scope of the present application. The components not explicitly described in the embodiment can be realized by existing technology.
Claims
1. A cross-modal data retrieval method based on voice intent, characterized in that, Specifically, it includes: Step 1: Input the voice corresponding to the user's query request, convert the voice into text through a speech recognition algorithm, and optimize the recognition effect by combining it with a domain-specific vocabulary. Step 2: Perform intent parsing on the text generated by speech recognition to output a structured intent representation, segment the keywords in the speech recognition results, and extract domain tags by combining with the domain lexicon; Step 3: Extract image features and video temporal features from local image and video data to generate image feature vectors and visual feature vectors; perform semantic encoding on local text data to generate text feature vectors; perform semantic encoding on local structured data to generate database feature vectors, i.e., structured features; establish a cross-modal feature association space to map visual feature vectors, text feature vectors, and structured feature vectors to the same semantic space; store the feature vectors, text feature vectors, and structured feature vectors of image and video keyframes, along with their domain labels, in the ES vector database, with each vector associated with a corresponding domain label; Step 4: Perform two-stage matching based on the query information entered by the user, and conduct multimodal retrieval to obtain candidate results; Step 5: Filter the candidate results based on the intent parsing results, and use the filtering conditions to obtain the final query results; Step 6: Optimize the model based on user feedback and maintain and update the vector library in real time through the online learning module.
2. The method for cross-modal data retrieval based on voice intent according to claim 1, characterized in that, Step 2 is as follows: Step 2-1: Perform intent parsing on the text generated by speech recognition, extract the semantic information of the query, extract the user's core needs, identify the user's implicit needs, and output a structured intent representation; Step 2-2: Segment the keywords in the intent parsing results and extract domain tags from the structured intent representation by combining the domain thesaurus.
3. The method for cross-modal data retrieval based on voice intent according to claim 1, characterized in that, In step 3, the image feature extraction specifically involves: Extract image feature vectors, optimize model depth, width and resolution through compound scaling, and output image feature vectors; The video feature extraction specifically involves: Step 3-1: Detect shot boundaries in the video using the inter-frame difference method; Step 3-2: In each shot, use a clustering algorithm to select the most representative keyframes; Step 3-3: Extract feature vectors from the extracted keyframes.
4. The method for cross-modal data retrieval based on voice intent according to claim 1, characterized in that, Step 4 is as follows: Step 4-1: Convert the structured representation of the user query into a query vector; Step 4-2: Use cosine similarity to perform a similarity search in the ES vector database to find the feature vector most similar to the query vector, and return the result through an approximate nearest neighbor search. Step 4-3: Extract the domain tags of the candidate results, match them with the domain tags queried by the user, and filter out the most relevant results.
5. The method for cross-modal data retrieval based on voice intent according to claim 1, characterized in that, Step 5 specifically involves: Step 5-1: Extract implicit requirements from the intent parsing results and use them as search filtering conditions; Step 5-2: Calculate the matching degree between the filtering conditions and the candidate results, and return the refined search results that meet the filtering conditions.
6. The method for cross-modal data retrieval based on voice intent according to claim 1, characterized in that, The improved speech recognition algorithm mentioned in step 1 is specifically as follows: Step 1-1-1: The input audio signal is processed by pre-emphasis filtering and frame segmentation to extract acoustic features and perform energy normalization. Step 1-1-2: The acoustic features are mapped to vectors through a linear projection layer and positional encoding is added. The vectors are then fed into the encoder for context modeling. The parameters of the earlier layers are kept frozen, and only the last few layers are fine-tuned and a nonlinear activation function is inserted into the feedforward network. Step 1-1-3: Use the encoder output as the input of the LSTM predictor, and fuse the encoder output with the context information of the unidirectional LSTM predictor through a joint network. Use the CTC loss function and the forward-backward algorithm for end-to-end training. Step 1-1-4: Select the token with the highest probability frame by frame using a greedy decoding strategy and remove duplicate symbols and whitespace characters. Combine the language model to re-score and output the final text recognition result.
7. The method for cross-modal data retrieval based on voice intent according to claim 1, characterized in that, The intent understanding described in step 2-1 specifically refers to: Step 2-1-1: Text preprocessing and word segmentation. The input user text is standardized and segmented into sub-word units. Special markers [CLS] and [SEP] are added to form the model input sequence. Padding pooling and truncation are used to unify the sequence length. Attention mask labels are generated to mark the effective positions. Step 2-1-2: BERT encoder feature extraction. The segmented sequence is mapped to token_ids through the vocabulary, and added to the positional encoding and paragraph encoding to form the input embedding vector, which is fed into the encoder for bidirectional context modeling. Each layer includes hidden states and a multi-head attention mechanism Attention(Q, K, V): Where Q is obtained by projecting the input embedding through the projection matrix WQ, representing the query of the current word to other words; K is obtained by projecting the input embedding through the projection matrix WK, representing the representation of each word itself; V is obtained by projecting the input embedding through the projection matrix WV, representing the information passed from each word to other words; dk is the dimension of the Key vector; and W represents the weight matrices in the attention module, such as WQ, WK, and WV, which have dimensions of... A low-rank matrix is injected into the attention weight matrix W. During training, the original parameters of the BERT encoder are frozen and only the parameters of the low-rank matrix are updated. Comparative learning and fine-tuning are performed using domain-specific data. Step 2-1-3: Intent classification head calculation. Extract the context representation vector of the [CLS] position as the global semantic feature. Map the intent category through a linear classification layer. Apply the softmax activation function to calculate the probability distribution P(intent|text) = softmax(logits) for each intent category. Select the category with the highest probability as the predicted intent. Where text represents the input text, intent represents the intent category, logits is the raw score of the model output, represents the hidden state vector at the [CLS] position in the BERT encoder output, and represents the semantic representation of the entire sentence; The weight matrix represents the intent classification head, mapping the [CLS] vector to C intent categories; Indicates the bias term of the classification header; represents the unnormalized score corresponding to each intent category; C represents the total number of intent categories; softmax represents the activation function that normalizes the logits vector to a probability distribution; Step 2-1-4: Slot labeling and entity extraction. The hidden state of each token position in the sequence is labeled using BIO. The transition probability between tags is modeled through a CRF layer, with the loss function being L_slot = -log P(tag_sequence|token_sequence). Key entity information in the user input is extracted and semantically aligned with the intent category for verification. Here, tag_sequence is the real tag sequence, representing the slot label of each token, and token_sequence represents the input token sequence. Step 2-1-5: Structured intent representation generation. The intent classification results, slot labeling results, and confidence scores are integrated into a structured representation in JSON format. The representation is then post-processed and validated using a rule engine to finally output a complete user intent understanding result.
8. The method for cross-modal data retrieval based on voice intent according to claim 1, characterized in that, Step 2-2, which involves extracting domain keyword tags, includes a three-level filtering mechanism: Step 2-2-1: Initial screening of domain terms based on TF-IDF; Step 2-2-2: Calculate semantic similarity using pre-trained word vectors and filter candidate words whose cosine similarity with the domain vocabulary exceeds a threshold; Step 2-2-3: Based on the critical evaluation of attention weights, retain words whose semantic head attention exceeds the threshold.
9. A cross-modal data retrieval method based on voice intent according to claim 1, characterized in that, The two-stage matching described in step 4 includes: Coarse-grained matching: Calculate the cosine similarity between the query vector and the feature library, and select the TOP-K candidate set; Fine-grained matching, constructing a filtering matrix M∈R based on domain labels {N×T} Where N is the number of candidates and T is the label dimension; through M×Q T Perform a weighted sort, where Q is the query tag vector.
10. A cross-modal data retrieval method based on speech intent according to claim 1, characterized in that, The online learning module described in step 6 employs a double-buffered incremental update mechanism, including: Step 6-1: Real-time maintenance and updates, regularly collecting new data to build a vector library, and updating the domain thesaurus and tag weights; Step 6-2: Batch check and update: Periodically use the sliding window sampling mechanism to detect data distribution shift using KL divergence. When the KL divergence meets the threshold, trigger model retraining. Step 6-3: Version Verification Rollback: When the quality of the random validation set degrades, roll back to the previous stable version.
Citation Information
Cited By
Cross-modal conference information association retrieval method and system and medium
CN121808072A