Multi-mode-based data retrieval enhancement method
Through multi-step data preprocessing, feature extraction and multi-modal data processing methods, problems such as cross-modal semantic alignment in multi-modal applications are solved, the quality and accuracy of the retrieval system are improved, and more efficient data retrieval and processing are achieved.
Patent Information
- Application Number
- CN202411857390.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-17
- Publication Date
- 2025-05-09
AI Technical Summary
In multimodal applications, problems such as cross-modal semantic alignment, retrieval system quality, external data chunking, vector search matching accuracy, and large-model neglect of intermediate information have affected the output quality of the RAG model and the overall performance of the retrieval system.
Multi-step methods are used to pre-process data, feature extraction, build multi-modal feature space, define similarity metrics, build index structures, introduce dynamic search layers, realize cross-modal semantic alignment, real-time data source integration, and use complex mathematical formulas to optimize search algorithms to improve the quality and efficiency of the search system.
Through these measures, semantic alignment between data of different modality is achieved, the quality and accuracy of the retrieval system are improved, the external data blocking strategy is optimized, the accuracy of vector retrieval matching is improved, and the problem of large models neglecting intermediate information is solved.
Smart Images

Figure CN119961461A_ABST
Abstract
Description
Technical Field
[0001] The present invention specifically relates to a multimodal-based data retrieval enhancement method. Background Art
[0002] In the field of artificial intelligence (AI), the term "modality" usually refers to different types or formats of data, such as text, images, audio, and video. Each modality represents a unique form or channel of information. Multimodality refers to the use or combination of two or more modalities at the same time. In AI systems, multimodality usually means that the model is able to process and integrate information from different sensory channels. For example, a multimodal system can analyze images and text at the same time to better understand the content of news reports. With the rapid development of AI technology, multimodal retrieval enhancement generation technology has emerged. This technology retrieves information related to the input from a large knowledge base and inputs this information as context and questions to the large language model (LLM), allowing the LLM to generate answers based on this information. This approach enables the LLM to connect with the latest external data or knowledge, and then answer questions based on the latest knowledge and data.
[0003] The prior art in this field currently has the following problems:
[0004] Cross-modal semantic alignment problem: In multimodal applications, it is crucial to ensure consistency between different forms of information (such as text, images, and videos). However, since data from different modalities have different representations and characteristics, it is difficult to map them onto a unified basis, which affects the output quality and coherence of the RAG model.
[0005] Retrieval system quality: The quality of the retrieval system directly affects the final result generation. If the retrieved content is irrelevant or not accurate enough, the output of the large model will also be affected.
[0006] External data segmentation: External data usually needs to be segmented before vectorization. However, how to segment data properly to ensure that the segmented data contains rich logical information but is not too lengthy is a complex problem. In addition, the choice of segmentation strategy will also affect the effect of vectorization and the quality of subsequent retrieval.
[0007] Vector retrieval matching accuracy: Even if external knowledge and user input are vectorized and stored in a vector database, vector matching is still a challenge because in high-dimensional space, the distances between all points become very close, making it very difficult to distinguish between "near" and "far" points.
[0008] Big models ignore intermediate information: Studies have found that retrieval-enhanced generation usually uses retrieval results sorted by relevance and sent to big models to generate results. However, many of the current mainstream big models pay more attention to the beginning and end of the content, and may ignore the middle part even if it is valuable information.
[0009] In summary, the present application proposes a multimodal data retrieval enhancement method to solve the above problems. Summary of the invention
[0010] The purpose of the present invention is to provide a multi-modal data retrieval enhancement method to address the deficiencies of the prior art. The multi-modal data retrieval enhancement method can well solve the above-mentioned problems.
[0011] In order to achieve the above requirements, the technical solution adopted by the present invention is: to provide a multi-modal data retrieval enhancement method, the multi-modal data retrieval enhancement method comprising the following steps:
[0012] S1: Data collection and preprocessing steps; collect data from different modalities, such as text, images, videos, and audio. The collected data often contains noise and redundant information, so preprocessing is required. Preprocessing includes data cleaning, deleting duplicate data, filling missing data, data conversion, i.e. converting raw data into a unified format, and data standardization, i.e. scaling the data to the same range, to improve the efficiency and accuracy of subsequent algorithms;
[0013] S2: Steps for feature extraction; for text data, natural language processing technology is used to extract keywords, phrases, and sentence embedding features; for image and video data, computer vision technology is used to extract color, texture, shape, and motion features; for audio data, features such as spectrum, pitch, and rhythm are extracted. The goal of feature extraction is to retain important information in the data while reducing the dimension and complexity of the data;
[0014] S3: Perform the steps of constructing a multimodal feature space; this is achieved through joint matrix decomposition or deep learning technology. Joint matrix decomposition decomposes a high-dimensional matrix into the product of several matrices with lower dimensions, thereby reducing the complexity of calculation; deep learning technology learns the potential relationship between different modalities by training a neural network and generates a unified feature representation;
[0015] S4: Perform the steps of defining similarity metrics; in multimodal data retrieval, it is necessary to define a metric method that can calculate similarity across modalities. Similarity metrics include cosine similarity, Euclidean distance, and Manhattan distance, or use deep learning-based methods such as twin networks or matching networks to learn more complex similarity metrics;
[0016] S5: Perform the step of building an index structure; in order to improve retrieval efficiency, an index structure is built to organize and manage data. In multimodal data retrieval, a hierarchical index or a hybrid index structure is used to support cross-modal retrieval. Hierarchical indexes store data of different modalities in separate indexes and perform queries through a unified retrieval interface. Hybrid indexes merge data of different modalities into a unified index to achieve more efficient cross-modal retrieval.
[0017] S6: Perform the steps of introducing a dynamic retrieval layer; establish a multi-level index and use a dynamic retrieval algorithm to quickly find the most relevant data in the massive data. The dynamic retrieval layer dynamically adjusts the retrieval strategy according to the user's query and context information to optimize the retrieval results. The adaptive retrieval model based on reinforcement learning is used to adjust the retrieval strategy according to the user's preference and data background.
[0018] S7: Steps for cross-modal semantic alignment; this is achieved through the shared latent space method, that is, mapping data from different modalities onto a unified basis, using deep learning techniques to learn the potential relationships between different modalities and generate shared semantic representations, using pre-trained language models and visual models to generate cross-modal semantic representations, and improving the accuracy and consistency of the representations through fine-tuning;
[0019] S8: The step of integrating real-time data sources. In a production environment, static knowledge bases are prone to becoming outdated. In order to maintain the accuracy and relevance of retrieval results, it is necessary to integrate real-time data sources to dynamically update the knowledge base. This can be achieved through crawler technology, API interfaces, or data streams. Real-time data sources provide the latest information, thereby enhancing the timeliness and accuracy of the retrieval system.
[0020] S9: Perform the steps of using complex mathematical formulas to optimize the retrieval algorithm; in order to improve the efficiency of the retrieval algorithm, mathematical formulas are used for optimization. In the sparse-dense hybrid retrieval mechanism, sparse retrieval and dense retrieval methods are combined to balance precision and recall. Sparse retrieval quickly processes keywords, and dense retrieval improves the relevance of data through semantic understanding. The following formula is used to represent the hybrid retrieval mechanism:
[0021] Score(q, d) = α·SparseScore(q, d) + (1-α)·DenseScore(q, d); where q represents the query, d represents the document, α is the weight coefficient, SparseScore and DenseScore represent the scores of sparse retrieval and dense retrieval respectively;
[0022] By adjusting the value of α, the precision and recall rate are balanced to optimize the retrieval results;
[0023] S10: Perform post-processing and result optimization steps; remove duplicate results, adjust ranking, relevance scoring and summary generation steps, automatically optimize user queries through intelligent query reconstruction technology, ensure that the results returned by the retrieval process are more relevant and accurate, use feedback-based retrieval optimization methods, continuously improve retrieval strategies through user feedback, and improve system personalization and effectiveness.
[0024] The advantages of this multimodal data retrieval enhancement method are as follows:
[0025] (1) Cross-modal semantic alignment: By adopting the shared latent space method or other advanced techniques, this method may achieve semantic alignment between different modal data, thereby improving the output quality and coherence of the RAG model.
[0026] (2) Improving the quality of the retrieval system: This method may improve the quality and accuracy of the retrieval system by optimizing the retrieval algorithm, introducing dynamic retrieval layers and other technologies, thereby ensuring that the retrieved content is highly relevant to the user input.
[0027] (3) Optimizing external data partitioning strategy: This method may adopt a more reasonable external data partitioning strategy to ensure that the partitioned data contains rich logical information but is not too lengthy. At the same time, it may also improve the efficiency of subsequent processing or reference by storing metadata related to each block.
[0028] (4) Improving the accuracy of vector retrieval and matching: This method may improve the accuracy of vector matching by introducing more advanced vector retrieval algorithms or technologies. For example, a more accurate similarity measurement method or deep learning technology can be used to optimize the vector matching process.
[0029] (5) Solve the problem of large models ignoring intermediate information: This method may ensure that the large model can make full use of the intermediate information in the retrieval results by adjusting the input order of the large model or introducing other mechanisms. For example, important information can be placed at the beginning or end of the input, or an attention mechanism can be introduced to make the large model pay more attention to the information in the middle. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The same reference numerals are used in these drawings to represent the same or similar parts. The exemplary embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:
[0031] Figure 1 The figure schematically shows a multi-modal data retrieval enhancement method according to an embodiment of the present application. DETAILED DESCRIPTION
[0032] In order to make the objectives, technical solutions and advantages of the present application more clear, the present application is further described in detail below in conjunction with the accompanying drawings and specific embodiments.
[0033] In the following description, references to "one embodiment", "an embodiment", "an example", "an example", etc. indicate that the embodiment or example described in this way may include specific features, structures, characteristics, properties, elements or limitations, but not every embodiment or example necessarily includes the specific features, structures, characteristics, properties, elements or limitations. In addition, repeated use of the phrase "according to one embodiment of the present application" may refer to the same embodiment, but does not necessarily refer to the same embodiment.
[0034] For the sake of simplicity, certain technical features well known to those skilled in the art are omitted in the following description.
[0035] According to one embodiment of the present application, a multi-modal data retrieval enhancement method is provided, such as Figure 1 As shown, the following steps are included:
[0036] S1: Data collection and preprocessing steps;
[0037] Data collection is the first step in any data retrieval task. In this step, data from different modalities, such as text, images, videos, and audio, needs to be collected. The collected data often contains noise and redundant information, so it needs to be preprocessed. Preprocessing includes data cleaning (such as removing duplicate data and filling missing data), data conversion (such as converting raw data into a unified format), and data standardization (such as scaling data to the same range) to improve the efficiency and accuracy of subsequent algorithms.
[0038] S2: step of feature extraction;
[0039] Feature extraction is the process of converting raw data into a form that can be used by algorithms. For text data, natural language processing techniques can be used to extract features such as keywords, phrases, and sentence embeddings; for image and video data, computer vision techniques can be used to extract color, texture, shape, and motion features; for audio data, features such as spectrum, pitch, and rhythm can be extracted. The goal of feature extraction is to retain the important information of the data while reducing the dimensionality and complexity of the data.
[0040] S3: Perform the step of constructing a multimodal feature space;
[0041] When constructing a multimodal feature space, it is necessary to map the features of different modalities into the same feature space for cross-modal retrieval. This is usually achieved through joint matrix decomposition or deep learning technology. Joint matrix decomposition can decompose a high-dimensional matrix into the product of several matrices with lower dimensions, thereby reducing the complexity of the calculation. For example, techniques such as singular value decomposition (SVD) or non-negative matrix factorization (NMF) can be used. Deep learning technology can learn the potential relationship between different modalities by training neural networks and generate a unified feature representation.
[0042] S4: performing a step of defining a similarity measure;
[0043] Similarity measurement is a method to evaluate the similarity between two data points. In multimodal data retrieval, it is necessary to define a measurement method that can calculate similarity across modalities. Commonly used similarity measures include cosine similarity, Euclidean distance, and Manhattan distance. In addition, deep learning-based methods such as Siamese Network or Matching Network can be used to learn more complex similarity measures.
[0044] S5: Perform the step of building an index structure;
[0045] In order to improve retrieval efficiency, it is necessary to build an index structure to organize and manage data. Common index structures include hash tables, B-trees, inverted indexes, and KD trees. In multimodal data retrieval, hierarchical indexes or hybrid index structures can be used to support cross-modal retrieval. Hierarchical indexes store data of different modalities in separate indexes and query them through a unified retrieval interface. Hybrid indexes merge data of different modalities into a unified index to achieve more efficient cross-modal retrieval.
[0046] S6: performing a step of introducing a dynamic retrieval layer;
[0047] The dynamic retrieval layer is the key to improving retrieval speed and accuracy. By building multi-level indexes and utilizing dynamic retrieval algorithms, the most relevant data can be quickly found in massive data. The dynamic retrieval layer can dynamically adjust the retrieval strategy based on the user's query and context information to optimize the retrieval results. For example, an adaptive retrieval model based on reinforcement learning can be used to adjust the retrieval strategy based on the user's preferences and data background.
[0048] S7: Steps for cross-modal semantic alignment;
[0049] Cross-modal semantic alignment is a key step to ensure that data from different modalities are semantically consistent. This is usually achieved through a shared latent space approach, which maps data from different modalities onto a unified basis. Deep learning techniques can be used to learn the latent relationships between different modalities and generate shared semantic representations. For example, pre-trained language models (PLMs) and visual models can be used to generate cross-modal semantic representations, and fine-tuned to improve the accuracy and consistency of the representations.
[0050] S8: Steps for real-time data source integration;
[0051] In a production environment, static knowledge bases are prone to becoming outdated. In order to maintain the accuracy and relevance of search results, it is necessary to integrate real-time data sources to dynamically update the knowledge base. This can be achieved through crawler technology, API interfaces, or data streams. Real-time data sources can provide the latest information, such as news, social media updates, and user feedback, thereby enhancing the timeliness and accuracy of the retrieval system.
[0052] S9: performing a step of optimizing the search algorithm using a complex mathematical formula;
[0053] In order to improve the efficiency of the retrieval algorithm, complex mathematical formulas can be used for optimization. For example, in a sparse-dense hybrid retrieval mechanism, sparse retrieval and dense retrieval methods can be combined to balance precision and recall. Sparse retrieval can quickly process keywords, while dense retrieval can improve the relevance of data through semantic understanding. The following formula can be used to represent this hybrid retrieval mechanism:
[0054] Score(q, d) = α·SparseScore(q, d) + (1-α)·DenseScore(q, d); where q represents the query, d represents the document, α is the weight coefficient, SparseScore and DenseScore represent the scores of sparse retrieval and dense retrieval respectively. By adjusting the value of α, you can balance precision and recall, thereby optimizing the retrieval results.
[0055] S10: post-processing and result optimization steps;
[0056] Post-processing is the process of further optimizing the search results. This includes steps such as removing duplicate results, ranking adjustment, relevance scoring, and summary generation. Through intelligent query reconstruction technology, the user's query can be automatically optimized to ensure that the results returned by the search process are more relevant and accurate. In addition, feedback-based search optimization methods can be used to continuously improve search strategies through user feedback to improve the personalization and effectiveness of the system.
[0057] According to an embodiment of the present application, the multimodal data retrieval enhancement method specifically includes the following steps:
[0058] 1. Data collection and organization
[0059] Data collection: Obtain image, text, audio, and video data from public datasets, social media platforms, online news sites, academic databases, etc. Use crawler technology or API interfaces to automatically collect data to ensure the timeliness and diversity of the data.
[0060] Data collation: Perform preliminary screening of the collected data to remove duplicate, invalid or poor quality data. Categorize and store the data according to different modalities for subsequent processing. Check the integrity and consistency of the data to ensure that each data item contains the necessary metadata (such as timestamp, source, etc.).
[0061] 2. Data Preprocessing
[0062] Text preprocessing: Use natural language processing tools (such as NLTK, SpaCy, etc.) to perform word segmentation, part-of-speech tagging, and named entity recognition. Remove stop words, punctuation marks, and special characters from the text to reduce data sparsity. Perform stem extraction or lemmatization on the text to reduce the impact of vocabulary diversity on retrieval results.
[0063] Image preprocessing: Use image processing libraries (such as OpenCV, PIL, etc.) to crop, scale, and denoise the images. Convert the images to grayscale or perform color space conversion to reduce computational complexity. Normalize the images to ensure that different images have the same scale when extracting features.
[0064] Audio / video preprocessing: Use audio processing libraries (such as Librosa, pydub, etc.) to sample, denoise, and segment audio. Perform frame extraction, key frame detection, and image processing on the video so that it can be converted into an image modality for processing.
[0065] 3. Feature Extraction
[0066] Text feature extraction: Use statistical methods such as Bag of Words and TF-IDF to extract text features. Use word embedding technology (such as Word2Vec, GloVe, etc.) to convert text into dense vector representation. Use deep learning models (such as BERT, GPT, etc.) to extract semantic features of text.
[0067] Image feature extraction: Use convolutional neural networks (CNN) to extract local and global features of images. Use pre-trained CNN models (such as VGG, ResNet, etc.) for feature extraction to improve efficiency.
[0068] Perform feature encoding on the image and convert high-dimensional features into low-dimensional representations for easy subsequent processing.
[0069] Audio / video feature extraction: Use audio feature extraction methods such as Mel-frequency cepstral coefficients (MFCC) to extract the spectral features of the audio. Use video processing libraries (such as FFmpeg) to extract information such as key frames and audio tracks of the video. Synchronize the audio and video data to ensure consistency during feature extraction.
[0070] 4. Modal Fusion
[0071] Feature-level fusion: concatenate or weighted sum feature vectors of different modalities to form a unified feature representation. Use feature transformation techniques (such as PCA, LDA, etc.) to reduce the dimensionality of the fused features.
[0072] Decision-level fusion: Sort or score the retrieval results of different modalities, and then fuse them using strategies such as weighted averaging and voting. Use ensemble learning methods (such as random forests, gradient boosting trees, etc.) to further optimize the fused results.
[0073] 5. Build an index
[0074] Select index structure: Select a suitable index structure according to data characteristics and retrieval requirements, such as inverted index, hash index, etc. Consider factors such as index storage efficiency, retrieval speed, and update cost.
[0075] Build an index: Use the extracted feature vector or fused feature representation as the index key, and the original data or data ID as the index value. Use index building tools (such as Lucene, Elasticsearch, etc.) to automatically build the index.
[0076] 6. Query Processing
[0077] Query parsing: Parse the query entered by the user and extract key information (such as keywords, images, audio, etc.). Select the corresponding processing flow according to the modality type of the query.
[0078] Query transformation: Convert the query into a form that matches the index structure, such as converting a text query into a feature vector. Preprocess and extract features for image, audio, and other queries to match them with the data in the index.
[0079] 7. Similarity Calculation
[0080] Select similarity measurement method: Select appropriate similarity measurement method according to data characteristics and retrieval requirements, such as cosine similarity, Euclidean distance, etc. Consider factors such as the efficiency and accuracy of similarity calculation.
[0081] Calculate similarity: Use similarity metrics to calculate the similarity between the query and the data in the index. Sort or score the calculated similarities so that the most relevant data is returned as the search result.
[0082] 8. Result Fusion and Sorting
[0083] Result fusion: If the query involves multiple modalities, the retrieval results of each modality need to be fused. Use strategies such as weighted average and voting to fuse the results to ensure that the fused results can reflect the information of different modalities.
[0084] Result sorting: Sort the fused results according to similarity or scoring results. You can use sorting learning algorithms (such as LambdaRank, ListNet, etc.) to further optimize the sorting results.
[0085] 9. Optimization and Adjustment
[0086] Performance evaluation: Use indicators such as accuracy, recall, and F1 score to evaluate retrieval performance. Analyze the accuracy and efficiency of retrieval results to identify performance bottlenecks and improvement directions.
[0087] Parameter tuning: Tune the parameters in steps such as feature extraction, modality fusion, and index construction. Use strategies such as grid search and random search to find the optimal parameter combination.
[0088] Algorithm improvement: Try to use new algorithms or models to replace existing algorithms or models. Introduce advanced technologies such as deep learning to improve the accuracy of feature extraction and modality fusion.
[0089] 10. System deployment and maintenance
[0090] System deployment: Deploy the multimodal data retrieval system on the server. Configure the corresponding network environment and security measures to ensure the stability and security of the system.
[0091] System maintenance: Regularly maintain and update the system, including fixing vulnerabilities, updating algorithms, optimizing performance, etc. Monitor the system's operating status and performance indicators to identify and resolve problems in a timely manner.
[0092] User support: Provide user manuals and operation guides to help users quickly get started with the system. Collect user feedback and opinions to continuously optimize system functions and user experience.
[0093] The above-mentioned embodiments only represent several implementation methods of the present invention, and the description thereof is relatively specific and detailed, but it cannot be understood as limiting the scope of the present invention. It should be pointed out that, for a person skilled in the art, several modifications and improvements can be made without departing from the concept of the present invention, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the claims.
Claims
1. A multimodal data retrieval enhancement method, characterized in that: The steps include: S1: Data collection and preprocessing steps; Collect data from different modalities, such as text, images, videos, and audio. The collected data often contains noise and redundant information, so preprocessing is required. Preprocessing includes data cleaning, deleting duplicate data, filling missing data, data conversion (converting raw data into a unified format), and data standardization (scaling data to the same range) to improve the efficiency and accuracy of subsequent algorithms. S2: step of feature extraction; For text data, natural language processing technology is used to extract keywords, phrases, and sentence embedding features; for image and video data, computer vision technology is used to extract color, texture, shape, and motion features; for audio data, features such as spectrum, pitch, and rhythm are extracted. The goal of feature extraction is to retain important information in the data while reducing the dimension and complexity of the data; S3: Perform the step of constructing a multimodal feature space; This is achieved through joint matrix decomposition or deep learning technology. Joint matrix decomposition decomposes a high-dimensional matrix into the product of several matrices with lower dimensions, thereby reducing the complexity of calculation; deep learning technology learns the potential relationship between different modalities by training neural networks and generates a unified feature representation; S4: performing a step of defining a similarity measure; In multimodal data retrieval, it is necessary to define a metric method that can calculate similarity across modalities. Similarity metrics include cosine similarity, Euclidean distance, and Manhattan distance, or use deep learning-based methods such as twin networks or matching networks to learn more complex similarity metrics; S5: Perform the step of building an index structure; In order to improve retrieval efficiency, an index structure is constructed to organize and manage data. In multimodal data retrieval, a hierarchical index or a hybrid index structure is used to support cross-modal retrieval. Hierarchical indexes store data of different modalities in separate indexes and query them through a unified retrieval interface. Hybrid indexes merge data of different modalities into a unified index to achieve more efficient cross-modal retrieval. S6: performing a step of introducing a dynamic retrieval layer; Establish multi-level indexes and use dynamic retrieval algorithms to quickly find the most relevant data in massive data. The dynamic retrieval layer dynamically adjusts the retrieval strategy based on the user's query and context information to optimize the retrieval results. Use an adaptive retrieval model based on reinforcement learning to adjust the retrieval strategy based on user preferences and data background; S7: Steps for cross-modal semantic alignment; This is achieved through the shared latent space method, that is, mapping data from different modalities onto a unified basis, using deep learning techniques to learn the potential relationships between different modalities and generate shared semantic representations, using pre-trained language models and visual models to generate cross-modal semantic representations, and improving the accuracy and consistency of the representations through fine-tuning; S8: Steps for real-time data source integration; In a production environment, static knowledge bases are prone to becoming outdated. In order to maintain the accuracy and relevance of retrieval results, it is necessary to integrate real-time data sources to dynamically update the knowledge base. This can be achieved through crawler technology, API interfaces, or data streams. Real-time data sources provide the latest information, thereby enhancing the timeliness and accuracy of the retrieval system. S9: performing a step of optimizing the search algorithm using a complex mathematical formula; In order to improve the efficiency of the retrieval algorithm, mathematical formulas are used for optimization. In the sparse-dense hybrid retrieval mechanism, sparse retrieval and dense retrieval methods are combined to balance precision and recall. Sparse retrieval quickly processes keywords, while dense retrieval improves the relevance of data through semantic understanding. The following formula is used to represent the hybrid retrieval mechanism: Score(q,d)=α·SparseScore(q,d)+(1-α)·DenseScore(q,d); Among them, q represents the query, d represents the document, α is the weight coefficient, SparseScore and DenseScore Represent the scores of sparse retrieval and dense retrieval respectively; By adjusting the value of α, the precision and recall rate are balanced to optimize the retrieval results; S10: post-processing and result optimization steps; Remove duplicate results, adjust rankings, relevance scoring and summary generation steps, automatically optimize user queries through intelligent query reconstruction technology, ensure that the results returned by the retrieval process are more relevant and accurate, use feedback-based retrieval optimization methods, continuously improve retrieval strategies through user feedback, and improve system personalization and effectiveness.
2. The multimodal data retrieval enhancement method according to claim 1, characterized in that: Step S1 specifically includes: Data collection: Obtain image, text, audio, and video data from public datasets, social media platforms, online news sites, and academic databases. Use APIs to automatically collect data to ensure the timeliness and diversity of the data. Data collation: Preliminary screening of collected data, removal of duplicate, invalid or poor quality data, classification and storage of data according to different modalities for subsequent processing, checking the integrity and consistency of data, and ensuring that each data item contains necessary metadata, such as timestamps and sources; Step S2 specifically includes: Text preprocessing: Use NLTK natural language processing tools to perform word segmentation, part-of-speech tagging, and named entity recognition, remove stop words, punctuation marks, and special characters in the text, reduce data sparsity, and perform stem extraction or word form restoration on the text to reduce the impact of vocabulary diversity on retrieval results; Image preprocessing: Use the OpenCV image processing library to crop, scale, and denoise the image, convert the image to grayscale or perform color space conversion to reduce computational complexity; normalize the image to ensure that different images have the same scale when extracting features; Audio and video preprocessing: Use the Librosa audio processing library to sample, denoise, and segment audio, and perform frame extraction, key frame detection, and image processing on video so as to convert it into image mode for processing; Step S3 specifically includes: Text feature extraction: Use TF-IDF statistical method to extract text features, use Word2Vec word embedding technology to convert text into dense vector representation, and use BERT deep learning model to extract semantic features of text; Image feature extraction: Use convolutional neural networks to extract local and global features of images, and use pre-trained CNN models for feature extraction to improve efficiency; Perform feature encoding on the image and convert high-dimensional features into low-dimensional representations for easy subsequent processing. Audio and video feature extraction: Use the Mel-frequency cepstral coefficient audio feature extraction method to extract the spectral features of the audio, use the FFmpeg video processing library to extract the key frames and audio track information of the video, and synchronize the audio and video data to ensure consistency during feature extraction.
3. The multimodal data retrieval enhancement method according to claim 1, characterized in that: Step S4 specifically includes: Feature-level fusion: concatenate feature vectors of different modalities to form a unified feature representation, and use PCA feature transformation technology to reduce the dimensionality of the fused features; Decision-level fusion: sort or score the retrieval results of different modalities, then fuse them using weighted average and voting strategies, and further optimize the fused results using random forest ensemble learning methods; Step S5 specifically includes: Select index structure: select appropriate index structure according to data characteristics and retrieval requirements; Build an index: Use the extracted feature vector or fused feature representation as the index key, the original data or data ID as the index value, and use the Lucene index building tool to automatically build the index.
4. The multimodal data retrieval enhancement method according to claim 1, characterized in that: Step S6 specifically includes: Query parsing: Parse the query entered by the user, extract key information including keywords, images, and audio, and select the corresponding processing flow according to the query modality type; Query conversion: convert the query into a form that matches the index structure, preprocess and extract features for image, audio and other queries so that they can be matched with the data in the index; Step S7 specifically includes: Select similarity measurement method: Select the appropriate similarity measurement method based on data characteristics and retrieval requirements; Calculate similarity: Use similarity measurement methods to calculate the similarity between the query and the data in the index, and sort or score the calculated similarities to return the most relevant data as the search results; Step S8 specifically includes: Result fusion: If the query involves multiple modalities, the retrieval results of each modality need to be fused, and the results are fused using weighted average and voting strategies to ensure that the fused results can reflect the information of different modalities; Result sorting: Sort the fused results according to similarity or scoring results. The LambdaRank sorting learning algorithm can be used to further optimize the sorting results.
5. The multimodal data retrieval enhancement method according to claim 1, characterized in that: Step S9 specifically includes: Performance evaluation: Use accuracy, recall, and F1 score indicators to evaluate retrieval performance, analyze the accuracy and efficiency of retrieval results, and identify performance bottlenecks and improvement directions; Parameter tuning: Tune the parameters in the steps of feature extraction, modality fusion, index construction, etc., and use grid search and random search strategies to find the optimal parameter combination; Algorithm improvement: Try to use new algorithms or models to replace existing algorithms or models. Introduce advanced technologies such as deep learning to improve the accuracy of feature extraction and modality fusion; Step S10 specifically includes: System deployment: deploy the multimodal data retrieval system on the server, configure the corresponding network environment and security measures to ensure the stability and security of the system; System maintenance: Regularly maintain and update the system, including fixing vulnerabilities, updating algorithms, optimizing performance, etc., monitor the system's operating status and performance indicators, and promptly identify and resolve problems; User support: Provide user manuals and operation guides to help users quickly get started with the system, collect user feedback and opinions, and continuously optimize system functions and user experience.
Citation Information
Cited By
Video image data cleaning method and system for urban rail transit engineering
CN120182636A
RAG knowledge retrieval method based on B + tree structure
CN120371839A
Rag knowledge retrieval method based on b+ tree structure
CN120371839B
Server retrieval result generation method and device, computer equipment and medium
CN120541277A
Method, device, computer equipment and medium for generating search results of server
CN120541277B