A Multimodal Information Intelligent Retrieval Method and System
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-24
- Publication Date
- 2026-08-14
AI Technical Summary
[0005]有鉴于此,本申请实施例提供了一种多模态信息智能检索方法及系统,旨在解决现有技术中存在的检索场景适配精准度不足、检索精度偏低、检索效率受限的问题
[0015]本申请实施例的第四方面提供了一种计算机可读存储介质,包括:存储有计算机程序,所述计算机程序被处理器执行时实现如上述第一方面中所述多模态信息智能检索方法的步骤。
Smart Images

Figure CN122570702A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of data processing technology, and in particular relates to a method and system for intelligent retrieval of multimodal information. Background Technology
[0002] With the rapid iteration of artificial intelligence and multimedia data technologies, heterogeneous multimodal data such as text, images, audio, and video are experiencing explosive growth. Intelligent retrieval of multimodal information has become a core technology direction in fields such as intelligent search, knowledge services, and smart media.
[0003] Existing technologies mainly employ a single-level cross-modal feature fusion architecture, which encodes modal data such as text and images through a unified vector space and relies on a basic semantic mapping library to match retrieval intent with multimodal data.
[0004] However, existing technologies cannot accurately match user search modality preferences with search scenarios, and the granularity of cross-modal semantic matching is coarse, resulting in low search accuracy and difficulty in improvement. Summary of the Invention
[0005] In view of this, embodiments of this application provide a multimodal information intelligent retrieval method and system, aiming to solve the problems of insufficient accuracy in adapting retrieval scenarios, low retrieval precision, and limited retrieval efficiency in the prior art.
[0006] The first aspect of this application provides a multimodal information intelligent retrieval method, including:
[0007] Acquire multiple user search modality preference information, multiple current search modality information, multiple current search content information, multiple historical search modality information, multiple historical search content information, multiple historical search content information, and a search content generation model;
[0008] Based on a preset search scenario matching model, multiple current search scenario matching information is generated according to the multiple user search modality preference information, multiple historical search modality information, multiple current search modality information, and multiple current search content information.
[0009] Based on the multiple current search scenario matching information, multiple current search modality information, multiple current search content information, multiple historical search modality information, multiple historical search content information, multiple historical search content information, search content generation model, and multiple preset modality search content semantic mapping information databases, multiple current search content information are generated.
[0010] A second aspect of this application provides a multimodal information intelligent retrieval system, comprising:
[0011] The information acquisition module is used to acquire multiple user search modality preference information, multiple current search modality information, multiple current search content information, multiple historical search modality information, multiple historical search content information, multiple historical search content information, and a search content generation model;
[0012] The current search scenario matching information generation module is used to generate multiple current search scenario matching information based on a preset search scenario matching model, according to multiple user search modality preference information, multiple historical search modality information, multiple current search modality information, and multiple current search content information.
[0013] The current search content information generation module is used to generate multiple current search content information based on the multiple current search scenario matching information, multiple current search modality information, multiple current search content information, multiple historical search modality information, multiple historical search content information, multiple historical search content information, search content generation model, and multiple preset modality search content semantic mapping information databases.
[0014] A third aspect of this application provides a terminal device, the terminal device including a memory and a processor, the memory storing a computer program executable on the processor, and the processor executing the computer program to implement the steps of the multimodal information intelligent retrieval method described in the first aspect above.
[0015] A fourth aspect of this application provides a computer-readable storage medium, comprising: storing a computer program, wherein when executed by a processor, the computer program implements the steps of the multimodal information intelligent retrieval method described in the first aspect above.
[0016] The beneficial effects of this application embodiment compared with the prior art are as follows: This application adopts a multimodal data fusion retrieval mechanism, which fully combines user retrieval preferences, historical retrieval data and real-time retrieval scenarios to complete retrieval matching, achieves accurate cross-modal alignment, effectively solves the problems of single retrieval mode and poor scenario adaptability in traditional retrieval, and significantly improves the matching accuracy and retrieval response efficiency of multimodal information retrieval. Attached Figure Description
[0017] To more clearly illustrate the technical solutions in the embodiments of this application, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0018] Figure 1 This is a schematic diagram illustrating the implementation process of the multimodal information intelligent retrieval method provided in Embodiment 1 of this application;
[0019] Figure 2 This is a schematic diagram illustrating the implementation process of the multimodal information intelligent retrieval method provided in Embodiment 2 of this application;
[0020] Figure 3 This is a schematic diagram illustrating the implementation process of the multimodal information intelligent retrieval method provided in Embodiment 3 of this application;
[0021] Figure 4 This is a schematic diagram illustrating the implementation process of the multimodal information intelligent retrieval method provided in Embodiment 4 of this application;
[0022] Figure 5 This is a schematic diagram illustrating the implementation process of the multimodal information intelligent retrieval method provided in Embodiment 5 of this application;
[0023] Figure 6 This is a schematic diagram illustrating the implementation process of the multimodal information intelligent retrieval method provided in Embodiment Six of this application;
[0024] Figure 7 This is a schematic diagram illustrating the implementation process of the multimodal information intelligent retrieval method provided in Embodiment 7 of this application;
[0025] Figure 8 This is a schematic diagram of the structure of the multimodal information intelligent retrieval system provided in the embodiments of this application;
[0026] Figure 9 This is a schematic diagram of the terminal device provided in the embodiments of this application. Detailed Implementation
[0027] In the following description, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of this application. However, those skilled in the art will understand that this application may also be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, apparatuses, circuits, and methods have been omitted so as not to obscure the description of this application with unnecessary detail.
[0028] To illustrate the technical solution described in this application, specific embodiments are provided below.
[0029] Figure 1 The flowchart illustrating the implementation of the multimodal information intelligent retrieval method provided in Embodiment 1 of this application is shown below in detail:
[0030] Step S101: Obtain multiple user search modality preference information, multiple current search modality information, multiple current search content information, multiple historical search modality information, multiple historical search content information, multiple historical search content information, and a search content generation model.
[0031] In this embodiment, multiple user search modality preference information can be user-selected text search priority mode, image search priority mode, audio / video search priority mode, or multimodal balanced search mode. This information can be used to match user-specific multimodal search adaptation rules and can be manually selected by the user through the personalized settings interface of the user's search terminal or adaptively recorded by the system. Multiple current search modality information can refer to heterogeneous modal data information such as text, images, audio, and video corresponding to the current search task. This information is used to define the modal input type and data dimension of the current search and can be acquired in real time through multimodal data acquisition devices. Multiple current search content information can refer to search requirement data such as search text keywords, image features to be matched, audio semantics to be searched, and video content to be searched, which represent the user's real-time search intent and can be captured in real time through the search front-end interaction module. Multiple historical search modality information can be data on modality selection types and modality input combinations corresponding to all past search tasks of the system. This information can be used to analyze the user's normal search modality usage habits and can be continuously recorded and periodically updated through the search system's backend database. Multiple historical search content information can include various historical search requests, search keywords, and modal material data recorded by the system in the past. This information can be used to optimize the search intent recognition logic and can be generated through daily cumulative iteration updates in the search system backend. Multiple historical search content information can include the final search results, search matching materials, and search ranking content data output by the system after previous searches. This information can be used to construct multimodal search matching prior templates and can be continuously archived and generated through storage. The search content generation model can be a multimodal search semantic matching generation model built on the Transformer architecture. This model can be used to dynamically output accurately matched multimodal search content based on multi-dimensional search scenario data. It can be obtained by pre-training a dataset mapping multimodal search scenario features to search results.
[0032] Step S102: Based on the preset search scenario matching model, generate multiple current search scenario matching information according to the multiple user search modality preference information, multiple historical search modality information, multiple current search modality information, and multiple current search content information.
[0033] In this embodiment, the preset retrieval scenario matching model can be manually preset and can adopt a hybrid deep learning model that combines CNN and Transformer. The preset scenario matching confidence threshold can be set to 0.82. This allows for the integration of user-personalized retrieval preferences, historical retrieval behavior data, and current real-time retrieval modality and retrieval content data. This enables the fusion verification and matching analysis of multi-dimensional retrieval scenario features, and then the selection of scenario adaptation features that meet the preset confidence threshold. This generates multiple current retrieval scenario matching information that can accurately represent the current retrieval environment and user needs.
[0034] Step S103: Generate multiple current search content information based on the multiple current search scenario matching information, multiple current search modality information, multiple current search content information, multiple historical search modality information, multiple historical search content information, multiple historical search content information, search content generation model, and multiple preset modality search content semantic mapping information databases.
[0035] In this embodiment, multiple preset modal retrieval content semantic mapping information databases can be manually preset and may include text-image semantic mapping databases, text-audio semantic mapping databases, text-video semantic mapping databases, and cross-modal general semantic mapping databases. The preset semantic matching accuracy threshold can be 0.9, and the preset retrieval result ranking weight coefficient can be 0.75. Multiple current retrieval scenario matching information, multiple current retrieval modality information, multiple current retrieval content information, multiple historical retrieval modality information, multiple historical retrieval content information, and multiple historical retrieval content information are then uniformly feature-encoded. The encoded multi-dimensional retrieval features are then input into the retrieval content generation model. Combined with multiple preset modal retrieval content semantic mapping information databases, cross-modal semantic matching and content retrieval are completed. Then, according to the preset retrieval result ranking weight coefficient, retrieval content filtering and priority calibration are completed, thereby generating multiple current retrieval content information containing multimodal matching materials, accurate retrieval results, and ordered retrieval ranking. Finally, multiple current retrieval content information can be used to complete the intelligent retrieval output and display processing of multimodal information.
[0036] The multimodal information intelligent retrieval method provided in this application adopts a multimodal data fusion retrieval mechanism, which fully combines user retrieval preferences, historical retrieval data and real-time retrieval scenarios to complete retrieval matching, achieves accurate cross-modal alignment, effectively solves the problems of single retrieval mode and poor scenario adaptability in traditional retrieval, and significantly improves the matching accuracy and retrieval response efficiency of multimodal information retrieval.
[0037] Figure 2 The flowchart illustrating the implementation of the multimodal information intelligent retrieval method provided in Embodiment 2 of this application is shown. Its difference from Embodiment 1 described above lies in:
[0038] The preset retrieval scenario matching model includes a preset retrieval modality feature generation sub-model, a preset user retrieval preference generation sub-model, and a preset retrieval intent keyword recognition sub-model;
[0039] The current search scenario matching information includes current search modality feature information, current user search preference information, and current search intent keyword identification information;
[0040] Step S102 specifically includes:
[0041] Step S201: Generate a sub-model based on preset retrieval modality features, and generate multiple current retrieval modality feature information according to the multiple current retrieval modality information.
[0042] In this embodiment, the preset retrieval modality feature generation sub-model can be manually preset and can adopt a lightweight CNN deep learning model. The preset modality feature extraction dimension threshold can be 128 dimensions, and the preset modality feature effective filtering threshold can be 0.78. Then, full-dimensional feature extraction and dimension normalization processing can be performed on multiple current retrieval modality information, thereby eliminating invalid and redundant modality interference feature information, and retaining core feature data that is higher than the preset modality feature effective filtering threshold, thereby generating multiple current retrieval modality feature information with unified dimensions and effective features.
[0043] Step S202: Generate a sub-model based on preset user search preferences, and generate multiple current user search preference information according to the multiple user search modality preference information, multiple historical search modality information and multiple current search modality information.
[0044] In this embodiment, the preset user search preference generation sub-model can be manually preset and can adopt a Bi-LSTM deep learning model. The preset preference matching fit threshold can be set to 0.8, and the preset historical preference weight ratio parameter can be set to 0.6. Then, time-series feature fusion analysis can be performed on multiple user search modality preference information and multiple historical search modality information. Then, real-time preference adaptation verification is completed by combining multiple current search modality information. Finally, preference feature weighted fusion is completed based on the preset historical preference weight ratio parameter, and adaptation results higher than the preset preference matching fit threshold are retained, thereby generating multiple current user search preference information that fit user habits and are adapted to the current search scenario.
[0045] Step S203: Based on the preset search intent keyword recognition sub-model, generate multiple current search intent keyword recognition information according to the multiple current search modality feature information, multiple current user search preference information, and multiple current content to be searched information.
[0046] In this embodiment, the preset search intent keyword recognition sub-model can be manually preset, and can adopt a BERT pre-trained deep learning model. The preset keyword semantic matching threshold can be 0.88, and the preset effective number threshold of intent keywords can be 3. Then, multiple current search modality feature information and multiple current user search preference information can be combined to adapt to the corresponding scenario keyword thesaurus. Then, semantic parsing and keyword extraction are performed on multiple current search content information. Then, effective search keywords that are higher than the preset keyword semantic matching threshold and meet the quantity requirements are selected, thereby generating multiple current search intent keyword recognition information that can accurately represent the user's core search needs.
[0047] The multimodal information intelligent retrieval method provided in this application embodiment achieves refined and hierarchical accurate matching of retrieval scenarios, effectively reduces the probability of invalid and mismatched retrievals, effectively improves the accuracy of intent recognition and scenario adaptation of multimodal retrieval, thereby improving the accuracy and adaptability of multimodal information intelligent retrieval and ensuring the stability of retrieval services in various complex multimodal retrieval scenarios.
[0048] Figure 3 The flowchart illustrating the implementation of the multimodal information intelligent retrieval method provided in Embodiment 3 of this application is shown. The difference between this method and Embodiment 1 is that step S103 specifically includes:
[0049] Step S301: Based on the multiple historical retrieval modal information, multiple historical content to be retrieved information, multiple historical retrieval content information, and multiple preset modal retrieval content semantic mapping information databases, the retrieval content generation model is trained to generate a trained retrieval content generation model.
[0050] In this embodiment, multiple preset modal retrieval content semantic mapping information databases can be pre-set manually and may include text-image semantic mapping databases, text-audio semantic mapping databases, text-video semantic mapping databases, and cross-modal general semantic mapping databases. The preset cross-modal semantic matching similarity threshold can be 0.85, and the preset model iteration training number can be 1000 times. Then, multiple historical retrieval modal information, multiple historical retrieval content information, and multiple historical retrieval content information can be globally vectorized and encoded. Subsequently, the encoded multi-dimensional historical retrieval features can be bound and mapped with the multiple preset modal retrieval content semantic mapping information databases. Then, the mapped standardized dataset can be input into the retrieval content generation model for iterative training. The internal weight parameters of the model are continuously updated through the backpropagation mechanism until the model loss function tends to stabilize and converge, thereby generating a trained retrieval content generation model that adapts to user retrieval habits and multi-modal semantic matching rules.
[0051] Step S302: Generate multiple current search content information based on the multiple current search scenario matching information, multiple current search modality information, multiple current search content information, trained search content generation model, and multiple preset modality search content semantic mapping information databases.
[0052] In this embodiment, the preset search content output priority rule can be manually preset, with the priority from high to low being: precise semantic matching search content, modality-adapted search content, and similar association search content. The preset upper limit for the number of search results output can be 20. This allows multiple current search scenario matching information, multiple current search modality information, and multiple current search content information to be uniformly organized into a search input feature sequence that the model can recognize. Then, the feature sequence can be input into the search content generation model after training. Combined with multiple preset modality search content semantic mapping information databases, cross-modal semantic reasoning and search content matching are completed. Then, the search results can be filtered, sorted, and deduplicated according to the manually preset search content output priority rule, thereby generating multiple current search content information containing precise matching materials, associated search content, and an ordered sorted list. Finally, the results of multimodal information intelligent retrieval are output and displayed through multiple current search content information.
[0053] The multimodal information intelligent retrieval method provided in this application embodiment enables the post-trained retrieval content generation model to fully learn the semantic association rules of multimodal retrieval and user retrieval habits, effectively improving the adaptability of the post-trained retrieval content generation model to complex retrieval scenarios, making the final generated retrieval content more in line with the user's real-time retrieval needs, greatly improving the matching accuracy and content adaptability of multimodal retrieval, and at the same time improving the generalization ability of the retrieval model and the stability of the retrieval service.
[0054] Figure 4 The flowchart illustrating the implementation of the multimodal information intelligent retrieval method provided in Embodiment 4 of this application is shown. The difference between this method and Embodiment 3 above is that step S301 specifically includes:
[0055] Step S401: Perform matching processing based on the multiple historical retrieval modal information, multiple historical content to be retrieved information, and multiple preset modal retrieval content semantic mapping information databases to generate multiple historical modal content to be retrieved semantic mapping information.
[0056] In this embodiment, the preset historical retrieval semantic matching model can be manually preset, and a multilayer perceptron deep learning model can be used. The preset retrieval content semantic matching threshold can be 0.82. Then, modal features of multiple historical retrieval modal information and semantic features of multiple historical retrieval content information can be extracted. Then, the modal features of multiple historical retrieval modal information and the semantic features of multiple historical retrieval content information are globally matched and compared with multiple preset modal retrieval content semantic mapping information databases. Then, effective semantic association data higher than the preset retrieval content semantic matching threshold are selected, thereby generating multiple historical modal retrieval content semantic mapping information with modal association characteristics.
[0057] Step S402: Perform matching processing based on the multiple historical retrieval modality information, multiple historical retrieval content information, and multiple preset modality retrieval content semantic mapping information databases to generate multiple historical modality retrieval content semantic mapping information.
[0058] In this embodiment, the preset historical search content semantic matching model can be manually preset or can adopt a deep learning model of deep semantic matching network. The preset search content semantic matching threshold can be 0.86. Then, feature fusion encoding can be performed on multiple historical search modal information and multiple historical search content information. The fused features are then cross-modal semantic matching verification with multiple preset modal search content semantic mapping information databases. Valid mapping data higher than the preset search content semantic matching threshold is retained, thereby generating multiple historical modal search content semantic mapping information with historical search association characteristics.
[0059] Step S403: Perform splicing processing on the multiple historical retrieval modal information, multiple historical content to be retrieved information, and multiple historical modal content to be retrieved semantic mapping information to generate multiple historical modal content to be retrieved semantic splicing information.
[0060] In this embodiment, the preset feature dimension alignment parameters can be preset manually, and the unified feature dimension can be 256 dimensions. Then, multiple historical retrieval modal information, multiple historical content information to be retrieved, and multiple historical modal content semantic mapping information to be retrieved can be normalized in dimensions. Then, they are sequentially spliced according to a fixed order of modal features, content features, and semantic mapping features. Finally, feature fusion calibration processing is completed to generate multiple historical modal content semantic splicing information with unified dimensions and complete features.
[0061] Step S404: Perform splicing processing on the multiple historical retrieval modal information, multiple historical retrieval content information and multiple historical modal retrieval content semantic mapping information to generate multiple historical modal retrieval content semantic splicing information.
[0062] In this embodiment, the preset feature fusion calibration coefficient can be manually preset and can be set to 0.92. This allows for feature dimension alignment processing of multiple historical retrieval modal information, multiple historical retrieval content information, and multiple historical modal retrieval content semantic mapping information. Then, feature weight adaptation is completed according to the preset feature fusion calibration coefficient, and the overall splicing and fusion is completed according to a fixed feature order, thereby generating semantic splicing information of multiple historical modal retrieval content with complete semantic association and balanced feature distribution.
[0063] Step S405: Based on the semantic splicing information of the multiple historical modal content to be retrieved and the semantic splicing information of the multiple historical modal content to be retrieved, the retrieval content generation model is trained to generate a trained retrieval content generation model.
[0064] In this embodiment, the preset model training learning rate can be manually preset and can be set to 0.001. The preset model convergence loss threshold can be set to 0.0001. Then, the semantic concatenation information of multiple historical modalities to be retrieved and the semantic concatenation information of multiple historical modalities to be retrieved can be input into the retrieval content generation model in batches. Then, the retrieval semantic matching loss value is calculated through forward inference of the model. Then, the model parameters are updated iteratively by backpropagation based on the preset model training learning rate until the model loss value is lower than the preset model convergence loss threshold, thereby generating a trained retrieval content generation model with stronger feature extraction ability and higher semantic matching accuracy.
[0065] The multimodal information intelligent retrieval method provided in this application fully explores the deep semantic association features of multimodal retrieval data, effectively enhances the learning ability of the retrieval content generation model for modal association and semantic matching, thereby improving the semantic matching accuracy of multimodal retrieval and reducing the semantic bias problem of cross-modal retrieval.
[0066] Figure 5 The flowchart illustrating the implementation of the multimodal information intelligent retrieval method provided in Embodiment 5 of this application is shown. Its difference from Embodiment 3 described above lies in:
[0067] The preset modal retrieval content semantic mapping information includes multiple preset modal information of retrieval content to be mapped and multiple preset modal retrieval content semantic mapping relationships;
[0068] Step S301 specifically includes:
[0069] Step S501: Perform time-series feature conversion and splicing processing on the multiple historical content information to be retrieved and the multiple historical content information to generate multiple historical content time-series feature information.
[0070] In this embodiment, the preset temporal feature transformation model can be manually preset, and a temporal convolutional network deep learning model can be adopted. The preset temporal feature sampling interval can be 1 second, and the preset temporal feature fusion weight can be 0.85. Then, continuous feature sampling and temporal encoding can be performed on multiple historical content information to be retrieved and multiple historical content information according to the time series. Then, the dimensional alignment and weighted fusion of the two types of temporal features are completed. Then, redundant temporal features and invalid interference features are removed, thereby generating multiple historical content temporal feature information with complete temporal correlation and effective features.
[0071] Step S502: Calculate similarity based on the multiple historical retrieval modal information and multiple preset retrieval content modal information to be mapped, and generate multiple retrieval modal similarity information.
[0072] In this embodiment, the preset modal similarity calculation model can be preset by humans, and a cosine similarity calculation model can be used. The preset modal feature similarity calculation precision can be set to four decimal places. Then, the core modal features of multiple historical retrieval modal information can be extracted, and then a global similarity comparison calculation can be performed with multiple preset retrieval content modal information to be mapped. Then, the matching similarity between each modality is quantified and output, thereby generating accurate quantified multiple retrieval modal similarity information.
[0073] Step S503: Based on the multiple retrieval modality similarity information and the preset retrieval modality similarity threshold, the multiple preset modality retrieval content semantic mapping relationships are filtered to obtain multiple filtered modality retrieval content semantic mapping relationships and multiple unfiltered modality retrieval content semantic mapping relationships.
[0074] In this embodiment, the preset search modality similarity threshold can be manually preset and can be set to 0.75. Then, multiple search modality similarity information can be compared and verified one by one with the preset search modality similarity threshold. Then, the semantic mapping relationship of the modality search content with similarity value higher than the threshold is filtered out, generating multiple filtered modality search content semantic mapping relationships. Then, the semantic mapping relationship of the modality search content with similarity value lower than the threshold is filtered and retained, generating multiple unfiltered modality search content semantic mapping relationships.
[0075] Step S504: Based on the temporal feature information of the multiple historical search contents, the semantic mapping relationship of the multiple filtered modal search contents, and the semantic mapping relationship of the multiple unfiltered modal search contents, generate temporal mapping feature information of multiple filtered search contents and temporal mapping feature information of multiple unfiltered search contents.
[0076] In this embodiment, the preset temporal mapping feature transformation model can be preset manually, and can adopt a fully connected neural network deep learning model. The preset feature mapping dimension can be 512 dimensions. Then, the temporal feature information of multiple historical search content can be associated and mapped with the semantic mapping relationship of multiple selected modal search content and the semantic mapping relationship of multiple unselected modal search content, thereby completing the unified transformation and fusion adaptation of feature dimensions, and then generating corresponding matching feature sequences, thereby generating temporal mapping feature information of multiple selected search content and temporal mapping feature information of multiple unselected search content.
[0077] Step S505: Based on the temporal mapping feature information of the multiple filtered search contents and the temporal mapping feature information of the multiple unfiltered search contents, the search content generation model is trained to generate a trained search content generation model.
[0078] In this embodiment, the preset model hierarchical training strategy can be preset manually. The preset positive training weight can be 0.7, and the preset negative training weight can be 0.3. Then, multiple time-series mapping feature information of the selected search content can be used as positive training samples, and multiple time-series mapping feature information of the unselected search content can be used as negative training samples. Then, the search content generation model is iteratively trained by mixing the inputs according to the preset hierarchical training weight ratio. Then, the feature discrimination ability and semantic matching ability of the model are optimized by alternating training with positive and negative samples, and the model parameters are iteratively converged and the weights are solidified, thereby generating a post-trained search content generation model that is adapted to all modal scenarios and has stronger anti-interference ability.
[0079] The multimodal information intelligent retrieval method provided in this application enhances the accuracy of the retrieval content generation model in highly correlated cross-modal retrieval scenarios and improves the discrimination ability of the retrieval content generation model in low-correlation modal scenarios. It effectively solves the problems of insufficient temporal feature mining and low modality discrimination accuracy of traditional multimodal retrieval models, thereby improving the accuracy, adaptability and robustness of multimodal information intelligent retrieval in complex scenarios.
[0080] Figure 6 The flowchart illustrating the implementation of the multimodal information intelligent retrieval method provided in Embodiment Six of this application is shown. The difference between this method and Embodiment Five is that step S504 specifically includes:
[0081] Step S601: Perform temporal feature transformation processing based on the semantic mapping relationships of the multiple filtered modal search contents and the multiple unfiltered modal search contents to generate semantic mapping feature information of the multiple filtered modal search contents and the multiple unfiltered modal search contents.
[0082] In this embodiment, the preset temporal feature transformation processing model can be manually preset, and can adopt a temporal residual network deep learning model. The preset temporal feature uniformity dimension can be 512 dimensions, and the preset temporal feature smoothing coefficient can be 0.8. Then, the semantic features of multiple selected modal retrieval content semantic mapping relationships can be extracted in the whole domain and the temporal dimension normalization transformation can be performed, thereby completing feature denoising and smoothing processing, eliminating invalid semantic interference features, and generating multiple selected modal retrieval content semantic mapping feature information with uniform dimension and stable features. Then, according to the same model rules and preset parameters, the temporal feature transformation and calibration processing of multiple unselected modal retrieval content semantic mapping relationships can be performed, thereby generating multiple unselected modal retrieval content semantic mapping feature information.
[0083] Step S602: The temporal feature information of the multiple historical search contents and the semantic mapping feature information of the multiple selected modal search contents are concatenated to generate multiple temporal mapping feature information of the selected search contents.
[0084] In this embodiment, the preset feature splicing alignment rule can be preset manually, and the preset feature splicing overlap rate can be set to 0.1. This allows for precise dimensional alignment of the temporal feature information of multiple historical search content with the semantic mapping feature information of multiple selected modal search content. Then, the features are spliced in an ordered manner according to a fixed order of temporal features first and semantic mapping features second. Finally, feature fusion calibration is performed according to the preset feature splicing overlap rate to eliminate the feature discontinuity problem at the splicing boundary, thereby generating multiple temporal mapping feature information of selected search content with complete feature fusion and high temporal semantic matching degree.
[0085] Step S603: The temporal feature information of the multiple historical search contents and the semantic mapping feature information of the multiple unfiltered modal search contents are concatenated to generate multiple temporal mapping feature information of the unfiltered search contents.
[0086] In this embodiment, the preset negative feature splicing calibration parameter can be preset manually and can be 0.86. This allows for global dimensional adaptation processing of the temporal feature information of multiple historical search content and the semantic mapping feature information of multiple unfiltered modal search content, thereby completing the ordered splicing and weight calibration of the two types of features. Then, the complete temporal features and non-matching semantic mapping features are retained, thereby generating multiple unfiltered search content temporal mapping feature information that can represent low-association modal search features.
[0087] The multimodal information intelligent retrieval method provided in this application refines the feature differentiation expression of positive and negative samples, strengthens the ability of the retrieval content generation model to distinguish between effective and invalid retrieval semantics, effectively improves the semantic discrimination accuracy of the retrieval content generation model in complex cross-modal retrieval scenarios, and thus optimizes the matching accuracy and anti-interference ability of multimodal information intelligent retrieval.
[0088] Figure 7 The flowchart illustrating the implementation of the multimodal information intelligent retrieval method provided in Embodiment Seven of this application is shown. The difference between this method and Embodiment Five is that step S504 specifically includes:
[0089] Step S701: Based on the semantic mapping relationship of the multiple filtered modal search contents and the semantic mapping relationship of the multiple unfiltered modal search contents, perform semantic feature extraction processing to generate semantic feature information of the semantic mapping of the multiple filtered modal search contents and semantic feature information of the semantic mapping of the multiple unfiltered modal search contents.
[0090] In this embodiment, the preset semantic feature extraction model can be manually preset, and an improved BERT deep learning model can be used. The preset semantic feature extraction confidence threshold can be 0.87, and the preset core semantic feature retention ratio can be 0.9. Then, deep semantic feature analysis and core feature extraction can be performed on the semantic mapping relationship of multiple selected modal retrieval contents. Then, core semantic features higher than the preset semantic feature extraction confidence threshold are selected. Feature selection is completed according to the preset core semantic feature retention ratio, thereby generating high-precision semantic feature information of multiple selected modal retrieval contents semantic mapping. Then, the same model and parameters are used to perform semantic extraction and selection processing on the semantic mapping relationship of multiple unselected modal retrieval contents, thereby generating semantic feature information of multiple unselected modal retrieval contents semantic mapping.
[0091] Step S702: Cross-concatenate the temporal feature information of the multiple historical search contents and the semantic feature information of the semantic mapping of the multiple filtered modal search contents to generate multiple temporal mapping feature information of the filtered search contents.
[0092] In this embodiment, the preset cross-splitting step size parameter can be preset manually, and can be set to 16. The preset temporal semantic fusion weight can be set to 0.9. Then, the temporal feature information of multiple historical search content and the semantic mapping feature information of multiple selected modal search content can be cross-splitting in a temporal semantic alternation manner. Then, the preset temporal semantic fusion weight is used to complete the bidirectional feature weighted fusion. Finally, the dimensional calibration and redundancy removal are performed on the spliced feature sequence to generate multiple temporal mapping feature information of selected search content with strong temporal correlation and high semantic matching degree.
[0093] Step S703: Generate multiple unfiltered search content temporal mapping feature information based on the semantic feature information of the multiple unfiltered modal search content semantic mapping.
[0094] In this embodiment, the preset non-matching feature temporal completion model can be manually preset and can adopt a lightweight temporal filling network deep learning model. The preset temporal feature completion accuracy threshold can be set to 0.82. Then, temporal dimension completion and feature adaptation processing can be performed on the semantic feature information of semantic mapping of multiple unfiltered modal retrieval content. Then, the temporal feature distribution rules of historical retrieval content are matched to complete the feature standardization transformation. Then, abnormal interference features are removed, thereby generating temporal mapping feature information of multiple unfiltered retrieval content with unified temporal dimension and standardized features.
[0095] The multimodal information intelligent retrieval method provided in this application enhances the temporal semantic fusion effect of highly matched retrieval samples, while performing standardized temporal completion processing on non-matching samples to widen the feature differences between positive and negative training samples. This enables the retrieval content generation model to learn the semantic matching rules and non-matching discrimination logic of cross-modal retrieval more accurately, significantly improving the model's deep semantic understanding ability and cross-modal retrieval robustness, and effectively adapting to various complex and highly interfering multimodal intelligent retrieval scenarios.
[0096] Corresponding to the method in the above embodiments, Figure 8 The diagram shows a structural block diagram of the multimodal information intelligent retrieval system provided in the embodiments of this application. For ease of explanation, only the parts related to the embodiments of this application are shown. Figure 8 The example multimodal information intelligent retrieval system can be the execution subject of the multimodal information intelligent retrieval method provided in the aforementioned embodiment one.
[0097] Reference Figure 8 The multimodal information intelligent retrieval system includes:
[0098] The information acquisition module 810 is used to acquire multiple user search modality preference information, multiple current search modality information, multiple current search content information, multiple historical search modality information, multiple historical search content information, multiple historical search content information, and a search content generation model;
[0099] The current search scenario matching information generation module 820 is used to generate multiple current search scenario matching information based on a preset search scenario matching model, according to multiple user search modality preference information, multiple historical search modality information, multiple current search modality information, and multiple current search content information.
[0100] The current search content information generation module 830 is used to generate multiple current search content information based on the multiple current search scenario matching information, multiple current search modality information, multiple current search content information, multiple historical search modality information, multiple historical search content information, multiple historical search content information, search content generation model, and multiple preset modality search content semantic mapping information databases.
[0101] The process by which each module in the multimodal information intelligent retrieval system provided in this application implements its respective function can be found in the foregoing. Figure 1 The description of Embodiment 1 shown will not be repeated here.
[0102] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.
[0103] It should be understood that, when used in this application specification and the appended claims, the term "comprising" indicates the presence of the described features, integrals, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or a collection thereof.
[0104] It should also be understood that the term “and / or” as used in this application specification and the appended claims means any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.
[0105] As used in this application specification and the appended claims, the term "if" may be interpreted, depending on the context, as "when," "once," "in response to determination," or "in response to detection." Similarly, the phrase "if determined" or "if detected [the described condition or event]" may be interpreted, depending on the context, as meaning "once determined," "in response to determination," "once detected [the described condition or event]," or "in response to detection [the described condition or event]."
[0106] Furthermore, in the description of this application and the appended claims, the terms "first," "second," "third," etc., are used only for distinguishing descriptions and should not be construed as indicating or implying relative importance. It should also be understood that although the terms "first," "second," etc., are used in the text to describe various elements in some embodiments of this application, these elements should not be limited by these terms. These terms are merely used to distinguish one element from another. For example, a first table may be named a second table, and similarly, a second table may be named a first table, without departing from the scope of the various described embodiments. Both the first table and the second table are tables, but they are not the same table.
[0107] References to "one embodiment" or "some embodiments" as described in this specification mean that one or more embodiments of this application include a specific feature, structure, or characteristic described in connection with that embodiment. Therefore, the phrases "in one embodiment," "in some embodiments," "in other embodiments," "in still other embodiments," etc., appearing in different parts of this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized. The terms "comprising," "including," "having," and variations thereof mean "including but not limited to," unless otherwise specifically emphasized.
[0108] The multimodal information intelligent retrieval method provided in this application embodiment can be applied to terminal devices such as mobile phones, tablets, wearable devices, in-vehicle devices, augmented reality / virtual reality devices, laptops, super mobile personal computers, netbooks, and personal digital assistants. This application embodiment does not impose any restrictions on the specific type of terminal device.
[0109] For example, the terminal device may be a station in a WLAN, a cellular phone, a cordless phone, a session initiation protocol phone, a wireless local loop station, a personal digital processing device, a handheld device with wireless communication capabilities, a computing device or other processing device connected to a wireless modem, an in-vehicle device, a vehicle networking terminal, a computer, a laptop computer, a handheld communication device, a handheld computing device, a satellite wireless device, a wireless modem card, a set-top box, a user premises equipment, and / or other devices for communication over a wireless system, as well as next-generation communication systems, such as mobile terminals in 5G networks or mobile terminals in future evolved public terrestrial mobile networks, etc.
[0110] As an example and not a limitation, when the terminal device is a wearable device, the term "wearable device" can also refer to any device that utilizes wearable technology to intelligently design and develop everyday wearables, such as glasses, gloves, watches, clothing, and shoes. Wearable devices are portable devices worn directly on the body or integrated into a user's clothing or accessories. Wearable devices are not merely hardware devices; they achieve powerful functions through software support, data interaction, and cloud interaction. Broadly defined, wearable smart devices include those with comprehensive functions, large sizes, and the ability to perform complete or partial functions without relying on a smartphone, such as smartwatches or smart glasses, as well as those focused on a specific application function that require interaction with other devices such as smartphones, such as various smart bracelets and smart jewelry for vital sign monitoring.
[0111] Figure 9 This is a schematic diagram of the structure of a terminal device provided in an embodiment of this application. For example... Figure 9 As shown, the terminal device 9 of this embodiment includes: at least one processor 90 ( Figure 9 (Only one is shown in the image) A memory 91 stores a computer program 92 that can run on the processor 90. When the processor 90 executes the computer program 92, it implements the steps in the various embodiments of the multimodal information intelligent retrieval method described above, for example... Figure 1 Steps S101 to S103 are shown. Alternatively, when the processor 90 executes the computer program 92, it implements the functions of each module / unit in the above system embodiments, for example... Figure 8 The functions of modules 810 to 830 are shown.
[0112] The terminal device 9 can be a desktop computer, laptop, handheld computer, or cloud server, etc. The terminal device may include, but is not limited to, a processor 90 and a memory 91. Those skilled in the art will understand that... Figure 9 This is merely an example of terminal device 9 and does not constitute a limitation on terminal device 9. It may include more or fewer components than shown, or combine certain components, or different components. For example, the terminal device may also include input transmission devices, network access devices, buses, etc.
[0113] The processor 90 may be a central processing unit, or it may be other general-purpose processors, digital signal processors, application-specific integrated circuits, off-the-shelf programmable gate arrays or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor, etc.
[0114] In some embodiments, the memory 91 may be an internal storage unit of the terminal device 9, such as a hard disk or memory of the terminal device 9. The memory 91 may also be an external storage device of the terminal device 9, such as a plug-in hard disk, smart memory card, secure digital card, flash memory card, etc., equipped on the terminal device 9. Furthermore, the memory 91 may include both internal and external storage units of the terminal device 9. The memory 91 is used to store operating systems, applications, bootloaders, data, and other programs, such as the program code of computer programs. The memory 91 can also be used to temporarily store data that has been sent or will be sent.
[0115] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0116] This application also provides a terminal device, which includes at least one memory, at least one processor, and a computer program stored in the at least one memory and executable on the at least one processor. When the processor executes the computer program, it causes the terminal device to implement the steps in any of the above method embodiments.
[0117] This application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps described in the various method embodiments above.
[0118] This application provides a computer program product that, when run on a terminal device, enables the terminal device to implement the steps described in the various method embodiments above.
[0119] If the integrated module / unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the methods of the above embodiments can also be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium can include: any entity or device capable of carrying the computer program code, a recording medium, a USB flash drive, a portable hard drive, a magnetic disk, an optical disk, a computer memory, a read-only memory, a random access memory, an electrical carrier signal, a telecommunication signal, and a software distribution medium, etc.
[0120] In the above embodiments, the descriptions of each embodiment have different focuses. For parts that are not described in detail or recorded in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0121] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0122] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0123] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included within the protection scope of this application.
Claims
1. A multimodal information intelligent retrieval method, characterized in that, include: Acquire multiple user search modality preference information, multiple current search modality information, multiple current search content information, multiple historical search modality information, multiple historical search content information, multiple historical search content information, and a search content generation model; Based on a preset search scenario matching model, multiple current search scenario matching information is generated according to the multiple user search modality preference information, multiple historical search modality information, multiple current search modality information, and multiple current search content information. Based on the multiple current search scenario matching information, multiple current search modality information, multiple current search content information, multiple historical search modality information, multiple historical search content information, multiple historical search content information, search content generation model, and multiple preset modality search content semantic mapping information databases, multiple current search content information are generated.
2. The multimodal information intelligent retrieval method as described in claim 1, characterized in that, The preset retrieval scenario matching model includes a preset retrieval modality feature generation sub-model, a preset user retrieval preference generation sub-model, and a preset retrieval intent keyword recognition sub-model; The current search scenario matching information includes current search modality feature information, current user search preference information, and current search intent keyword identification information; The step of generating multiple current search scenario matching information based on the preset search scenario matching model, according to the multiple user search modality preference information, multiple historical search modality information, multiple current search modality information, and multiple current search content information, specifically includes: Based on the preset retrieval modality feature generation sub-model, multiple current retrieval modality feature information are generated according to the multiple current retrieval modality information; Based on a preset user search preference generation sub-model, multiple current user search preference information is generated according to the multiple user search modality preference information, multiple historical search modality information, and multiple current search modality information. Based on the preset search intent keyword recognition sub-model, multiple current search intent keyword recognition information are generated according to the multiple current search modality feature information, multiple current user search preference information, and multiple current content to be searched information.
3. The multimodal information intelligent retrieval method as described in claim 1, characterized in that, The step of generating multiple current search content information based on the multiple current search scenario matching information, multiple current search modality information, multiple current search content information, multiple historical search modality information, multiple historical search content information, multiple historical search content information, search content generation model, and multiple preset modality search content semantic mapping information databases specifically includes: Based on the multiple historical retrieval modal information, multiple historical content to be retrieved information, multiple historical retrieval content information, and multiple preset modal retrieval content semantic mapping information databases, the retrieval content generation model is trained to generate a trained retrieval content generation model. Based on the multiple current retrieval scenario matching information, multiple current retrieval modality information, multiple current retrieval content information, the trained retrieval content generation model, and multiple preset modality retrieval content semantic mapping information databases, multiple current retrieval content information are generated.
4. The multimodal information intelligent retrieval method as described in claim 3, characterized in that, The step of training the retrieval content generation model based on the multiple historical retrieval modality information, multiple historical retrieval content information, multiple historical retrieval content information, and multiple preset modality retrieval content semantic mapping information databases to generate a trained retrieval content generation model specifically includes: The matching process is performed based on the multiple historical retrieval modal information, multiple historical content to be retrieved information, and multiple preset modal retrieval content semantic mapping information databases to generate multiple historical modal content to be retrieved semantic mapping information. The matching process is performed based on the multiple historical retrieval modal information, multiple historical retrieval content information, and multiple preset modal retrieval content semantic mapping information databases to generate multiple historical modal retrieval content semantic mapping information. The information is concatenated based on the multiple historical retrieval modal information, multiple historical content to be retrieved information, and multiple semantic mapping information of the historical content to be retrieved to generate semantic concatenation information of the historical content to be retrieved. The information is concatenated based on the multiple historical retrieval modal information, multiple historical retrieval content information, and multiple historical modal retrieval content semantic mapping information to generate multiple historical modal retrieval content semantic concatenation information. Based on the semantic splicing information of the multiple historical modalities of the content to be retrieved and the semantic splicing information of the multiple historical modalities of the retrieved content, the retrieval content generation model is trained to generate a trained retrieval content generation model.
5. The multimodal information intelligent retrieval method as described in claim 3, characterized in that, The preset modal retrieval content semantic mapping information includes multiple preset modal information of retrieval content to be mapped and multiple preset modal retrieval content semantic mapping relationships; The step of training the retrieval content generation model based on the multiple historical retrieval modality information, multiple historical retrieval content information, multiple historical retrieval content information, and multiple preset modality retrieval content semantic mapping information databases to generate a trained retrieval content generation model specifically includes: Based on the multiple historical content information to be retrieved and the multiple historical content information to be retrieved, time-series feature transformation and splicing are performed to generate multiple historical content time-series feature information. Similarity is calculated based on the multiple historical retrieval modal information and multiple preset retrieval content modal information to be mapped, and multiple retrieval modal similarity information is generated. Based on the multiple retrieval modality similarity information and the preset retrieval modality similarity threshold, the multiple preset modality retrieval content semantic mapping relationships are filtered to obtain multiple filtered modality retrieval content semantic mapping relationships and multiple unfiltered modality retrieval content semantic mapping relationships. Based on the temporal feature information of the multiple historical search contents, the semantic mapping relationship of the multiple filtered modal search contents, and the semantic mapping relationship of the multiple unfiltered modal search contents, multiple temporal mapping feature information of the filtered search contents and multiple temporal mapping feature information of the unfiltered search contents are generated. Based on the temporal mapping feature information of the multiple filtered search contents and the temporal mapping feature information of the multiple unfiltered search contents, the search content generation model is trained to generate a trained search content generation model.
6. The multimodal information intelligent retrieval method as described in claim 5, characterized in that, The step of generating multiple times-series mapping feature information of filtered search content and multiple times-series mapping feature information of unfiltered search content based on the times-series feature information of multiple historical search content, the semantic mapping relationship of multiple filtered modal search content, and the semantic mapping relationship of multiple unfiltered modal search content specifically includes: Based on the semantic mapping relationships of the multiple selected modal search contents and the multiple unselected modal search contents, a temporal feature transformation process is performed to generate semantic mapping feature information of multiple selected modal search contents and semantic mapping feature information of multiple unselected modal search contents. The temporal feature information of the multiple historical search contents and the semantic mapping feature information of the multiple selected modal search contents are concatenated to generate multiple temporal mapping feature information of the selected search contents. The temporal feature information of multiple historical search contents and the semantic mapping feature information of multiple unfiltered modal search contents are concatenated to generate multiple temporal mapping feature information of unfiltered search contents.
7. The multimodal information intelligent retrieval method as described in claim 5, characterized in that, The step of generating multiple times-series mapping feature information of filtered search content and multiple times-series mapping feature information of unfiltered search content based on the times-series feature information of multiple historical search content, the semantic mapping relationship of multiple filtered modal search content, and the semantic mapping relationship of multiple unfiltered modal search content specifically includes: Semantic feature extraction is performed based on the semantic mapping relationships of the multiple selected modal search contents and the multiple unselected modal search contents to generate semantic feature information of the semantic mapping of the multiple selected modal search contents and the semantic feature information of the semantic mapping of the multiple unselected modal search contents. The time-series feature information of the multiple historical search contents and the semantic feature information of the semantic mapping of the multiple selected modal search contents are cross-concatenated to generate multiple time-series mapping feature information of the selected search contents. Based on the semantic feature information of the semantic mapping of the multiple unfiltered modal search content, multiple temporal mapping feature information of the unfiltered search content is generated.
8. A multimodal intelligent information retrieval system, characterized in that, include: The information acquisition module is used to acquire multiple user search modality preference information, multiple current search modality information, multiple current search content information, multiple historical search modality information, multiple historical search content information, multiple historical search content information, and a search content generation model; The current search scenario matching information generation module is used to generate multiple current search scenario matching information based on a preset search scenario matching model, according to multiple user search modality preference information, multiple historical search modality information, multiple current search modality information, and multiple current search content information. The current search content information generation module is used to generate multiple current search content information based on the multiple current search scenario matching information, multiple current search modality information, multiple current search content information, multiple historical search modality information, multiple historical search content information, multiple historical search content information, search content generation model, and multiple preset modality search content semantic mapping information databases.
9. A terminal device, characterized in that, The terminal device includes a memory and a processor. The memory stores a computer program that can run on the processor. When the processor executes the computer program, it implements the steps of the method as described in any one of claims 1 to 7.
10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method as described in any one of claims 1 to 7.