Multi-modal data search method and device, storage medium and computer equipment

By constructing a multimodal knowledge base and combining user behavior information, the shortcomings of traditional search methods in multimodal data processing are solved, and more accurate and personalized information retrieval is achieved.

CN120561349APending Publication Date: 2025-08-29PING AN INT FINANCIAL LEASING CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510686926.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-27
Publication Date
2025-08-29

AI Technical Summary

Technical Problem

Traditional search methods are difficult to effectively integrate and analyze multimodal data, resulting in reduced accuracy and correlation of search results in the financial and medical fields.

Method used

Through the multimodal data fusion search mechanism, a multimodal knowledge base is built, and the convolutional neural network and natural language processing methods are used to uniformly encode different modal data, and the search results are sorted based on user historical behavior information.

Benefits of technology

It improves the accuracy and relevance of information retrieval, provides a personalized search experience, and optimizes search efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120561349A_ABST
    Figure CN120561349A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of data processing, can be applied to the fields of financial information search, multi-dimensional information retrieval in intelligent investment advisers, medical information search and the like, and provides a multi-modal data search method and device, a storage medium and computer equipment. Performing multi-modal unified coding on different modal data in the knowledge base to obtain a multi-modal knowledge base; acquiring input information of a user; retrieving multi-modal information in a multi-modal knowledge base, wherein the similarity between the multi-modal information and the input information meets the requirement; the multi-modal information comprises feature vectors corresponding to the information of different modals, wherein the similarity of the information and the input information meets the requirement; and sorting the feature vectors corresponding to the information of different modes based on the historical behavior information of the user, and determining target search information corresponding to the input information based on a sorting result. According to the embodiment of the invention, a fusion retrieval mechanism of multi-modal data is introduced, so that the precision and correlation of information retrieval are effectively improved, and the user experience is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of data processing technology, and in particular to a multimodal data search method, apparatus, storage medium, and computer equipment. Background Art

[0002] With the rapid development of information technology and the continuous innovation of artificial intelligence and big data technologies, web search, as a core means of obtaining information, has become widely integrated into people's daily lives. However, with the diversification of data types and information carriers, traditional search methods face many challenges when processing complex multimodal data.

[0003] Traditional search methods, primarily based on keyword matching and shallow semantic analysis, are effective for processing single-modal data, but face significant limitations when processing multimodal data. Multimodal data encompasses a variety of formats, including images, audio, and video, and contains rich and complex information. Traditional search methods struggle to deeply analyze and effectively integrate this information, resulting in reduced accuracy and relevance in search results.

[0004] This problem is particularly acute in the financial sector, where financial data contains a vast amount of multimodal information, including charts, reports, audio and video analysis, and more. Investors often need to integrate this multimodal data when analyzing market trends, but traditional search engines cannot provide comprehensive, accurate, and personalized search results, potentially hindering the efficiency and accuracy of investment decisions. Similarly, fields like medical information search face similar challenges. Summary of the Invention

[0005] The embodiments of the present disclosure at least provide a multimodal data search method, apparatus, storage medium, and computer device, which effectively improve the accuracy and relevance of information retrieval by introducing a fusion retrieval mechanism for multimodal data, thereby enhancing the user experience.

[0006] The present disclosure provides a multimodal data search method, including:

[0007] Obtaining a knowledge base, and performing multimodal unified encoding on different modal data in the knowledge base to obtain a multimodal knowledge base; wherein the multimodal knowledge base includes feature vectors corresponding to different modal data in a unified semantic space;

[0008] Obtaining user input information; wherein the user input information includes at least one of an image, text, audio, and video;

[0009] Retrieving multimodal information from the multimodal knowledge base, the multimodal information having a similarity to the input information meeting the requirements; wherein the multimodal information includes feature vectors corresponding to different modal data having a similarity to the input information meeting the requirements;

[0010] The feature vectors corresponding to the different modal data are sorted based on the user's historical behavior information, and the target search information corresponding to the input information is determined based on the sorting result.

[0011] The present disclosure provides a multimodal data search device, comprising:

[0012] A knowledge base construction module is used to obtain a knowledge base and perform multimodal unified encoding on different modal data in the knowledge base to obtain a multimodal knowledge base; wherein the multimodal knowledge base includes feature vectors corresponding to different modal data in a unified semantic space;

[0013] An information acquisition module, configured to acquire user input information; wherein the user input information includes at least one of images, text, audio, and video;

[0014] an information retrieval module, configured to retrieve multimodal information from the multimodal knowledge base that meets the requirements for similarity with the input information; wherein the multimodal information includes feature vectors corresponding to different modal data that meet the requirements for similarity with the input information;

[0015] An information determination module is used to sort the feature vectors corresponding to the different modal data based on the user's historical behavior information, and determine the target search information corresponding to the input information based on the sorting result.

[0016] An embodiment of the present disclosure provides a computer device, comprising: a processor, a memory, and a bus, wherein the memory stores machine-readable instructions executable by the processor. When the computer device is running, the processor and the memory communicate via the bus, and when the machine-readable instructions are executed by the processor, a multimodal data search method as described in any possible embodiment described above is performed.

[0017] An embodiment of the present disclosure provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the multimodal data search method as described in any of the possible implementations described above is implemented.

[0018] The multimodal data search method, device, storage medium, and computer equipment provided in the embodiments of the present disclosure perform multimodal unified encoding of different modal data in the knowledge base, construct a multimodal knowledge base containing a unified semantic space feature vector, and sort the search results in combination with the user's historical behavior information during the retrieval process. This can achieve efficient retrieval of cross-modal data, accurately match the multimodal information input by the user with the knowledge base content, and optimize the order of search result presentation based on the user's personalized preferences. In this way, by breaking down the semantic barriers between different modal data and making full use of user behavior data to mine personalized needs, the relevance and accuracy of search results can be greatly improved, providing a more personalized search experience, and effectively optimizing search efficiency, thereby improving user experience.

[0019] In order to make the above-mentioned objectives, features and advantages of the present disclosure more obvious and easy to understand, preferred embodiments are given below and described in detail with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] In order to more clearly illustrate the technical solutions of the embodiments of the present disclosure, the following briefly introduces the drawings that need to be cited in the embodiments. The drawings herein are incorporated into and constitute a part of the specification. These drawings illustrate embodiments consistent with the present disclosure and, together with the specification, are used to illustrate the technical solutions of the present disclosure. It should be understood that the following drawings only illustrate certain embodiments of the present disclosure and should not be regarded as limiting the scope. For those of ordinary skill in the art, other relevant drawings can be obtained based on these drawings without inventive effort.

[0021] Figure 1 A schematic diagram of an application environment of a multimodal data search method provided by an embodiment of the present disclosure is shown;

[0022] Figure 2 A flowchart of a multimodal data search method provided by an embodiment of the present disclosure is shown;

[0023] Figure 3 A flowchart of a method for uniformly encoding multimodal data provided by an embodiment of the present disclosure is shown;

[0024] Figure 4 A flowchart of a multimodal information retrieval method provided by an embodiment of the present disclosure is shown;

[0025] Figure 5 A flow chart of a target search information optimization method provided by an embodiment of the present disclosure is shown;

[0026] Figure 6 A schematic structural diagram of a multimodal data search device provided by an embodiment of the present disclosure is shown;

[0027] Figure 7 A schematic structural diagram of a computer device provided by an embodiment of the present disclosure is shown. DETAILED DESCRIPTION

[0028] In order to make the purpose, technical solutions and advantages of the embodiments of the present disclosure clearer, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below in conjunction with the drawings in the embodiments of the present disclosure. Obviously, the described embodiments are only part of the embodiments of the present disclosure, not all of the embodiments. The components of the embodiments of the present disclosure generally described and shown in the drawings herein can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present disclosure provided in the drawings is not intended to limit the scope of the disclosure for which protection is sought, but merely represents selected embodiments of the present disclosure. Based on the embodiments of the present disclosure, all other embodiments obtained by those skilled in the art without making creative work are within the scope of protection of the present disclosure.

[0029] It should be noted that similar reference numerals and letters denote similar items in the following drawings, and therefore, once an item is defined in one drawing, it does not need to be further defined or explained in subsequent drawings.

[0030] The term "and / or" herein simply describes an association relationship, indicating that three relationships can exist. For example, A and / or B can represent the existence of A alone, the simultaneous existence of A and B, and the existence of B alone. In addition, the term "at least one" herein refers to any combination of at least two of any one or more of a plurality of items. For example, "at least one of A, B, and C" can represent any one or more elements selected from the set consisting of A, B, and C.

[0031] To facilitate understanding of this embodiment, the execution subject of the multimodal data search method provided by the embodiment of the present disclosure is first introduced in detail. The multimodal data search method provided by the embodiment of the present invention can be applied to Figure 1 In an application environment, a client communicates with a server via a network. The client can be a mobile device, user terminal, terminal, handheld device, computing device, etc. The server can be a standalone physical server, a server cluster or distributed system consisting of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud storage, big data, and artificial intelligence platforms.

[0032] The following describes in detail the multimodal data search method provided by the embodiment of the present application in conjunction with the accompanying drawings. Figure 2FIG. 2 is a flow chart of a multimodal data search method provided by an embodiment of the present disclosure, the method comprising the following steps S201 to S204:

[0033] S201: Acquire a knowledge base, and perform multimodal unified encoding on different modal data in the knowledge base to obtain a multimodal knowledge base.

[0034] Here, the knowledge base can be obtained from multiple different sources, such as academic databases, encyclopedias, or open source resources on the Internet. By integrating information from these different channels, a knowledge base with rich content and wide coverage can be obtained.

[0035] It is understandable that since the data in the knowledge base often exists in different modalities, the data of different modalities have different ways and characteristics in expressing information. In order to achieve knowledge integration and deep mining, the data of different modalities in the knowledge base can be multimodally unified encoded. Here, modality refers to the form of expression of data. Common different modal data include text, images, audio, video, etc. Multimodal unified encoding is to convert data of different modalities into an encoding form in a unified semantic space. After the different modal data in the knowledge base are processed by multimodal unified encoding, a multimodal knowledge base can be obtained. Among them, the multimodal knowledge base is a database containing multiple types of data (images, text, audio, video, etc.). It encodes data of different modalities into a unified semantic space and integrates the feature vectors corresponding to different modal data for use in searching.

[0036] For example, referring to Figure 3 As shown, when uniformly encoding the modal data in the knowledge base, the following steps S301 to S304 may be included:

[0037] S301: For image data, use a convolutional neural network to extract visual features of the image data to obtain image visual features of the image data, and map the image visual features to a unified semantic space to obtain an image feature vector.

[0038] It's understandable that convolutional neural networks can be used to extract visual features from image data. Convolutional neural networks automatically identify different levels of image features, such as edges, shapes, and textures. Through multiple layers of convolutional processing, they can capture complex features ranging from local details to global structures. Ultimately, these visual features are mapped into a unified semantic space, transforming them into image feature vectors corresponding to the image data.

[0039] S302: For the text data, perform semantic analysis on the text data using a natural language processing method, and generate a text feature vector in a unified semantic space based on the semantic analysis result.

[0040] Here, for text data, natural language processing methods can be used to perform semantic analysis on the text data, and based on the semantic analysis results, text feature vectors in a unified semantic space can be generated. Natural language processing methods can include steps such as lexical analysis, syntactic analysis, and semantic analysis. Through natural language processing methods, the semantic content of the text can be understood, and then the semantic information is encoded into a vector form and mapped to a unified semantic space. For example, word embedding technology (such as Word2Vec, BERT, etc.) is used to convert words into vectors, and then the entire text is converted into a vector representation through sentence-level encoding methods.

[0041] S303 : For the audio data, extract the spectrum features of the audio data based on the acoustic model, identify and extract the semantic labels of the audio data, and encode the spectrum features and the semantic labels into an audio feature vector in a unified semantic space.

[0042] Specifically, processing audio data primarily involves extracting acoustic features and semantic labels. This process involves: first, performing spectral analysis on the audio data using an acoustic model to extract frequency and time domain features, which describe the fundamental properties of the audio signal. Next, speech recognition technology can be used to identify the speech content in the audio and extract the corresponding semantic labels. Finally, the spectral features and semantic labels are fused and mapped into a unified semantic space to generate a feature vector for the audio data.

[0043] S304: For video data, visual features of video frames in the video data are extracted based on a convolutional neural network, and features of the audio part of the video data are extracted in combination with audio processing technology, and the visual features and audio features of the video data are fused and mapped to a unified semantic space to obtain a video feature vector.

[0044] It is understandable that video data includes two types of data: images (i.e., video frames) and audio. Therefore, when processing video data, the image and audio data types can be processed separately and then fused. Specifically, the visual features of the video frame can be extracted through a convolutional neural network, similar to the image data processing process in step S301 above. Then, combined with audio processing technology, the acoustic features and semantic labels of the audio part of the video can be extracted. By fusing these visual features with audio features, the video data is mapped to a unified semantic space to generate a video feature vector. This feature vector integrates the visual information and audio information of the video.

[0045] In this way, through the above-mentioned unified encoding method of multimodal data, effective fusion and unified encoding of multimodal data can be achieved, thereby providing richer and more accurate data representation for tasks such as cross-modal analysis and information retrieval.

[0046] Here, the present disclosure integrates the various modal data processing processes involved in the above steps S301 to S304 based on a deep learning model, and the deep learning model is capable of end-to-end processing of data of different modalities. It first receives the original data of different modalities as input, and then extracts features through the corresponding sub-modules (such as convolutional neural network module for image and video frame processing, natural language processing module for text processing, acoustic model and speech recognition module for audio processing), then fuses the extracted different modal features, and finally maps the fused features into a unified semantic space. For example, in an intelligent customer service system in the financial field, when a user uploads a consultation content containing voice and video (such as a user showing relevant files while speaking), the integrated model can quickly and accurately convert different modal data into feature vectors of a unified semantic space, thereby more accurately understanding the user's needs and providing more effective replies and solutions. At the same time, the uniformly encoded feature vectors are easy to store and manage, providing a solid foundation for subsequent cross-modal search, knowledge reasoning and other tasks.

[0047] In some possible embodiments, a knowledge base may include a static knowledge database and a dynamic knowledge database, and a multimodal knowledge base includes a static knowledge base and a dynamic knowledge base. The static knowledge database is used to construct a static knowledge base, and the dynamic knowledge database is used to construct a dynamic knowledge base. Specifically, a static knowledge base is primarily used to store relatively stable, infrequently changing data, such as relatively constant factual information, academic resources, technical documents, historical data, etc. The static knowledge base can serve as a basic knowledge base, providing long-term and effective information support for information search. The dynamic knowledge base focuses on the collection and processing of real-time data, and can promptly reflect the latest information changes, so that when searching for information, updated dynamic information, real-time data, or highly timely content, such as news, social media updates, market trends, etc., can be synchronously searched. In this way, by combining static and dynamic knowledge bases, multi-dimensional information support can be achieved, the accuracy and flexibility of the search engine can be improved, and users can be provided with more comprehensive, timely, and efficient information query services, thereby improving the efficiency of decision support and knowledge acquisition.

[0048] Specifically, the static knowledge base can be constructed through the following steps (1) to (2):

[0049] (1) Acquiring a static knowledge database; wherein the static knowledge database includes image data, text data, audio data, and video data;

[0050] (2) performing multimodal unified encoding on the image data, the text data, the audio data, and the video data, respectively, and performing feature fusion processing and semantic alignment processing on the encoding results of each modal data to obtain the static knowledge base; wherein the static knowledge base includes feature vectors corresponding to each modal data in the unified semantic space.

[0051] It is understood that static knowledge databases can be collected from a variety of reliable data sources, such as company websites, public information platforms, and open source databases. Their data types can include multimodal data such as images, text, audio, and video. The application scope of static knowledge bases includes, but is not limited to, academic research, enterprise knowledge management, and educational resource storage.

[0052] Furthermore, since the collected data of different modalities (i.e., static knowledge databases) have different formats and features, they need to be uniformly encoded (see steps S301 to S304 above for details) so that they can be processed in a unified semantic space. Moreover, after completing the multimodal unified encoding, the encoding results of each modal data can be subjected to feature fusion processing to combine the feature information of different modalities to form a more comprehensive and richer data representation. For example, in a financial application scenario, text data may include a customer's credit report and bank communication records, and audio data may include the content of a customer's telephone conversation. By fusing these features, it can help build a more accurate financial knowledge base and provide multi-dimensional auxiliary decision support. Similarly, in the medical field, text modal data can include text information such as patient symptom descriptions, diagnosis results, treatment plans, etc. recorded in electronic medical records; image modal data can include X-rays, CT scan images, MRI images, etc.; audio modal data may be recordings of doctors' consultations, recordings of patients' descriptions of their conditions, etc.

[0053] Since data from different modalities may differ in their semantic expression, after data fusion, semantic alignment can be performed on the fused information to ensure semantic consistency across modalities. For example, the semantic meaning of "stock price rise" may appear as an upward curve in a stock chart in an image, while the text may simply describe it as "stock price rise." Semantic alignment allows these two modal expressions to be aligned in a unified semantic space. After completing these steps, a static knowledge base containing the corresponding feature vectors for each modality's data in the unified semantic space is obtained.

[0054] In some possible implementations, in order to more accurately process and associate data from different modalities through a multimodal knowledge base and improve the accuracy of cross-modal search and understanding, a multi-tuple dataset can be constructed and a contrastive loss function (such as InfoNCE) can be used to optimize the vector distance of data from different modalities, thereby further improving the effect of semantic alignment. For example, when processing the text description of "cheerful music," it should be kept close to high-frequency sound wave features and bright-toned images in the semantic space to ensure that the model can correctly understand its semantic relevance when processing cross-modal data.

[0055] Specifically, the dynamic knowledge base can be constructed through the following steps (a) to (c):

[0056] (a) establishing a real-time data acquisition channel; the real-time data acquisition channel supports access to multi-source and multi-modal real-time data;

[0057] (b) acquiring the dynamic knowledge database based on the real-time data acquisition channel; wherein the dynamic knowledge database includes real-time image data, real-time text data, real-time audio data, and real-time video data;

[0058] (c) performing multimodal unified encoding on the real-time image data, real-time text data, real-time audio data, and real-time video data, respectively, and performing feature fusion processing and semantic alignment processing on the encoding results of each modality of real-time data to obtain the dynamic knowledge base; wherein the dynamic knowledge base includes feature vectors corresponding to each modality of real-time data in a unified semantic space.

[0059] It is understandable that in order to obtain multi-source and multi-modal real-time data in a timely manner, a real-time data acquisition channel can be established to access real-time data in different modalities from multiple data sources. For example, real-time trading data (price, trading volume, price fluctuation, etc.) of financial products such as stocks and futures can be obtained from the real-time trading system of the financial market; the latest financial news reports (in text, image, video format) can be obtained from the API interface of news websites; and user discussions and comments on financial topics (in text, audio, video, etc.) can be obtained from the open interface of social media platforms.

[0060] Furthermore, after obtaining the dynamic knowledge database (i.e., real-time image, text, audio, and video data) through the real-time data acquisition channel, similar to the construction process of the static knowledge base, it is subjected to multimodal unified encoding (please refer to steps S301 to S304 for details), feature fusion processing, and semantic alignment processing to obtain a dynamic knowledge base containing the feature vectors corresponding to the real-time data of each modality in the unified semantic space.

[0061] In some possible embodiments, in order to achieve efficient data processing and integration, a real-time data channel can be established through the Kafka message queue, which supports the access of multi-source data, including patent data change logs, user behavior streams, and data pushed by external APIs. Through Kafka's distributed messaging mechanism, high throughput and low-latency data transmission can be ensured, so that changes in various data sources can be obtained in real time. After the data is accessed, the Schema Registry can be used to verify the data structure. The role of the Schema Registry is to perform standardized verification on the format of the input data to ensure that all data meets the predetermined structure and format requirements, avoiding errors caused by inconsistent formats when multiple data sources interact. After the data passes the verification, it will enter the Flink SQL stream processing platform, where real-time cleaning operations are performed on the dynamic data stream. Flink SQL can perform data cleaning operations such as deduplication and null value filling during the streaming data transmission process, ensuring that subsequent data processing and analysis can be carried out on the basis of accurate and clean data, thereby improving data quality and processing efficiency.

[0062] S202: Obtain user input information.

[0063] It is understood that user input information is the sentence or information provided by the user for search, and can include data in at least one modality: image, text, audio, and video. Image input can include photos taken by the user, scanned document images, and the like. For example, a user might upload an image containing a company logo and hope to search for information related to that company. Text input is textual content entered by the user, such as "2023 financial analysis of a leading company in a certain industry" to obtain relevant information. Audio input can include a recorded voice message, which may contain information about a financial product or condition. Video input can include video clips of financial lectures shot by the user.

[0064] In some possible embodiments, to ensure the security and privacy of user input information and prevent the leakage of sensitive information, the following measures may be taken: first, the user's initial input information is obtained. After obtaining the initial input information, a target encryption algorithm is determined based on the data security level of the initial input information. The data security level is divided according to the sensitivity and importance of the input information, or according to factors such as the user's identity and level of the input information. For example, if the input information contains highly sensitive information such as the user's bank account number, password, and ID number, then its data security level is high; while if it is just some ordinary financial product consultation information, the data security level is relatively low. Based on the different data security levels, an appropriate target encryption algorithm is selected to encrypt the initial input information. Through encryption, the input information can be effectively prevented from being stolen or tampered with during transmission and storage. Here, the encryption algorithm can use symmetric encryption algorithms (such as the AES algorithm) and asymmetric encryption algorithms (such as the RSA algorithm), etc., without specific limitations.

[0065] At the same time, sensitive information can be identified from the initial input. Sensitive information identification technology can identify sensitive content in the input information through methods such as keyword matching, pattern recognition, and machine learning. For example, in the financial sector, sensitive information can include customers' personal identification information, transaction records, account passwords, etc. If sensitive information is present in the initial input, it can be anonymized through methods such as data desensitization and data masking. Taking a user's ID number as an example, partial masking can be used, retaining only some of the digits and replacing the rest with asterisks. This protects the user's privacy while not affecting normal information retrieval and analysis. The final input information is then determined based on the anonymized sensitive information and the encrypted initial input information. If no sensitive information is present in the initial input information, the input information is determined directly based on the encrypted initial input information.

[0066] In this way, by encrypting input information and processing sensitive information, the user's information security is guaranteed, which can enhance the user's trust. At the same time, it also reduces the security risks faced by the system, which helps to improve the compliance and reliability of the system.

[0067] S203: Retrieve multimodal information whose similarity with the input information meets the requirements from the multimodal knowledge base.

[0068] It can be understood that the multimodal knowledge base stores the feature vectors corresponding to each modal data after unified encoding. These feature vectors are located in a unified semantic space and represent the semantic information of the different modal data. Here, multimodal information refers to the feature vectors corresponding to the different modal data retrieved from the multimodal knowledge base and meeting the required similarity with the input information.

[0069] Specifically, when retrieving multimodal information in a multimodal knowledge base whose similarity with the input information meets the requirements, the input information can be first converted into a vector form, and the similarity between the feature vector of the user input information and the feature vector of each data in the multimodal knowledge base is calculated, and then the multimodal information (that is, the feature vector corresponding to the different modal data whose similarity with the input information meets the requirements) is determined. Here, the method of similarity calculation can use cosine similarity, Euclidean distance, etc., which are not specifically limited here. For example: in the financial field, suppose the user enters a text query "stock market performance in the fourth quarter of 2023", the query information can be calculated by the similarity between the feature vectors in the multimodal knowledge base to find the feature vectors corresponding to the multimodal data such as financial news, stock trend charts, analysis reports, etc. that meet the similarity requirements in the knowledge base, so as to help users find relevant financial data and charts. For example, in the medical field, when a user enters a text query for "typical symptoms and CT imaging manifestations of acute appendicitis", the query information can be retrieved through similarity calculation between the feature vectors in the multimodal knowledge base, and the feature vectors corresponding to the multimodal data in the multimodal knowledge base that meet the similarity requirements, such as medical textbook text descriptions of acute appendicitis, audio explanations of doctors' clinical diagnoses, medical imaging data containing typical CT images of acute appendicitis, and related surgical operation videos, can be retrieved. This can further provide doctors or patients with relevant medical information, assist doctors in making diagnostic decisions, or help patients better understand their own condition.

[0070] For example, referring to Figure 4 As shown, in order to more efficiently retrieve multimodal information related to the input information in the multimodal knowledge base, when retrieving multimodal information whose similarity with the input information meets the requirements, the following steps S401 to S403 may be included:

[0071] S401 : Encode the input information into a query feature vector in a unified semantic space; and determine a candidate data set matching the query feature vector based on the static knowledge base and the dynamic knowledge base.

[0072] Here, in order to quickly and accurately locate data related to the input information, the input information can be first encoded into a query feature vector in a unified semantic space, and then candidate data sets matching the query feature vector are determined in the static knowledge base and the dynamic knowledge base.

[0073] Exemplarily, when determining a candidate data set based on input information, a static knowledge base, and a dynamic knowledge base, the following steps (I) to (III) may be included:

[0074] (I) performing semantic parsing on the input information, and decomposing the input information into a plurality of query subtasks based on the semantic parsing results; wherein each query subtask corresponds to one or more data shards in the static knowledge base and / or the dynamic knowledge base;

[0075] (II) allocating the plurality of query subtasks to a plurality of computing nodes for parallel processing, including:

[0076] For each query subtask, encoding the input information fragment corresponding to the query subtask into a query feature subvector in a unified semantic space;

[0077] Using a vector similarity calculation method, searching, in the data shard corresponding to the query subtask, feature vectors corresponding to different modal data whose similarity to the query feature subvector meets a preset initial similarity threshold, to obtain an initial candidate data set corresponding to the computing node;

[0078] (III) Summarizing the initial candidate data sets corresponding to the various computing nodes to obtain the candidate data sets.

[0079] Specifically, the goal of semantic parsing is to understand the semantics of the input information and decompose the input text into components with clear semantic meaning using natural language processing techniques. For example, the input information "Query the recent performance and industry prospects of a certain technology company's stock" can be decomposed into two query subtasks after semantic parsing: "Query the recent performance of a certain technology company's stock" and "Query the prospects of the industry in which the certain technology company is located." Each query subtask corresponds to one or more data shards in the static knowledge base and / or the dynamic knowledge base. Data shards are subsets of the knowledge base data divided according to certain rules, such as by financial product type, time range, etc. For example, in a financial knowledge base, data can be divided according to financial product type (such as stocks, funds, bonds, etc.), time range (such as daily, weekly, monthly, etc.), and region (such as domestic market, international market, etc.). For the subtask "Query the recent performance of a certain technology company's stock," it may correspond to a data shard in the static knowledge base that stores the technology company's historical stock price data and financial statement data, and it may also correspond to a data shard in the dynamic knowledge base that updates the latest trading data and market commentary of the stock in real time. Through this correspondence, the subset that may contain relevant information can be quickly located, thereby improving retrieval efficiency.

[0080] Furthermore, in order to make full use of computing resources and improve the retrieval speed, the present disclosure proposes to assign multiple query subtasks to multiple computing nodes for parallel processing. For each query subtask, it is first necessary to encode the corresponding input information fragment into a query feature subvector in a unified semantic space. Through a specific encoding algorithm, the input information fragment can be converted into a vector representation with a fixed dimension, and these vectors contain the semantic features of the input information. Then, using the vector similarity calculation method, the feature vectors corresponding to different modal data whose similarity with the query feature subvector meets the preset initial similarity threshold (such as 0.5~1 or 0.75~1) are retrieved in the data shard corresponding to the query subtask. During the retrieval process, all data feature vectors in the data shard corresponding to the query subtask are traversed, their similarity with the query feature subvector is calculated, and the data feature vectors whose similarity meets the preset initial similarity threshold are screened out to form an initial candidate data set corresponding to the computing node. For example, in the stock data shard, feature vectors corresponding to different modal data such as stock price trend charts and financial indicator data related to the recent performance of the technology company's stock may be screened out.

[0081] Finally, as query subtasks are distributed across multiple compute nodes for parallel processing, each compute node generates an initial candidate dataset. Aggregating these initial candidate datasets yields a more comprehensive candidate dataset, encompassing feature vectors corresponding to various modalities of data that may be relevant to the input information.

[0082] S402: Calculate the similarity between the query feature vector and the feature vectors corresponding to each modality data in the candidate data set based on a vector similarity calculation method.

[0083] Here, after obtaining the candidate data set, the similarity between the query feature vector and the feature vector corresponding to each modal data in the candidate data set can be calculated to evaluate the relevance of the candidate data to the input information.

[0084] S403 , determining feature vectors corresponding to different modal data whose similarity to the input information meets the requirements based on a preset similarity threshold and the similarity between the query feature vector and the feature vectors corresponding to each modal data in the candidate data set.

[0085] Specifically, if the similarity of the candidate data exceeds a preset similarity threshold (which differs from the initial similarity threshold, which is used to normalize the similarity between the candidate data and the query feature vector, while the similarity value here is used to normalize the similarity between the candidate data and the query feature vector), the candidate data is considered to have a high correlation with the input information and meets the retrieval requirements. By screening out feature vectors that meet the similarity requirements, we can ultimately determine different modal data that meet the requirements for similarity with the input information. This data can include text reports, image materials, audio explanations, video presentations, and other forms.

[0086] S203: sorting the feature vectors corresponding to the different modal data based on the user's historical behavior information, and determining the target search information corresponding to the input information based on the sorting result.

[0087] It's understandable that a user's historical behavior information records their preferences and needs, and can include information such as click history, dwell time, and search history. Based on this historical user behavior information, the feature vectors corresponding to retrieved data of different modalities can be ranked. For example, if a user has frequently searched for and clicked on news and reports related to technology stocks, then the feature vectors of images, text, audio, and video data related to technology stocks will receive a higher ranking weight in this search result.

[0088] For example, for text data, we can count the number of clicks on text related to different financial topics in the user's historical behavior and assign a weight to each topic. When sorting, text feature vectors related to the user's high-weighted topics will be ranked first. For image data, we can assign corresponding weights to image feature vectors based on the user's past viewing frequency of different types of financial images (such as stock charts, company financial statements, etc.). Audio and video data can also be weighted based on the user's historical listening and viewing records.

[0089] It is understood that after sorting the feature vectors of the different modal data, the target search information corresponding to the input information can be determined based on the sorting results. Here, the multimodal data corresponding to the top-ranked feature vectors can be selected as the target search information. This target search information can better meet the user's search needs because it not only has a high degree of semantic similarity with the input information, but also takes into account the user's historical preferences and behavioral habits.

[0090] For example, in the search scenario of "future stock trends of a certain technology company," after sorting, the target search information that may be determined includes: the stock price chart (image) of the technology company over the past period, the latest rating report (text) of the technology company's stock by a professional institution, audio analysis of the future stock trends of the technology company by industry experts, and promotional videos of the company's latest product launch. This multimodal target search information can provide users with more comprehensive and personalized financial information services, helping them make more accurate investment decisions.

[0091] For example, in a medical scenario, a user's historical behavior information records the user's preferences and demand clues for medical content, including click history (such as multiple clicks to view popular science articles related to a certain disease), length of stay (staying on a page explaining a treatment plan for a certain disease for a long time), search history (frequent searches for specific rare disease information), and other information. Based on this historical behavior information, the feature vectors corresponding to the retrieved data of different modalities can be sorted. For example, if a user has frequently searched and clicked to view health information and research reports related to cardiovascular disease prevention in the past, then in the search results for "How to Reduce the Risk of Cardiovascular Disease", the feature vectors of images (such as healthy diet maps, exercise demonstration maps), text (professional prevention guidelines, expert interpretation articles), audio (health lecture recordings), and video data (prevention exercise demonstration videos) related to cardiovascular disease prevention will obtain higher ranking weights, and then determine the target search information that best matches the user's input information based on the ranking results, providing users with medical content that better suits their needs and preferences.

[0092] In some possible embodiments, in order to improve the fit between search results and user needs, the final results can be sorted and optimized based on the similarity relationship between user interest characteristics and multimodal information. Figure 5 As shown, the following steps S501 to S504 may be included:

[0093] S501: Determine a user interest feature vector based on user historical behavior information.

[0094] Specifically, user historical behavior information reflects user interests and preferences, including but not limited to the webpage content and browsing time, searched keywords, clicked product links, watched video types, and saved article topics. By analyzing historical behavior information, we can create a user interest profile and obtain a corresponding user interest feature vector.

[0095] Here, historical user behavior data can be cleaned and preprocessed to remove noise and irrelevant information to improve data quality. Machine learning algorithms, such as clustering and association rule mining, are then applied to the processed data for feature extraction and pattern recognition. For example, clustering algorithms can categorize web pages viewed by topic, thereby revealing user interests in different areas. Association rule mining algorithms can identify combinations of items that users frequently search or click on simultaneously, revealing potential connections between user interests. Finally, these mined interest features are converted into user interest feature vectors in a unified semantic space. This vector represents the user's level of interest in a particular topic or feature, with the numerical value reflecting the strength of interest. For example, in a financial investment system, a user interest feature vector might include interest weights for different financial products, such as stocks, funds, and bonds, as well as the degree of attention paid to different information dimensions, such as macroeconomic conditions and industry dynamics.

[0096] S502: Calculate the similarity between the user interest feature vector and the feature vector corresponding to the different modal data.

[0097] It is understandable that after determining the user interest feature vector, the similarity between it and the feature vector corresponding to the different modal data obtained in the above steps can be calculated. During the calculation process, the feature vectors corresponding to all different modal data can be traversed and the similarity between them and the user interest feature vector can be calculated using a similarity calculation method (such as the above-mentioned cosine similarity, Euclidean distance, Pearson correlation coefficient, etc.). For example, the text feature vector of a financial news report, the image feature vector showing the stock market trend, and the audio feature vector of an expert interpretation will all be calculated for similarity with the user interest feature vector. In this way, the degree of match between different modal data and user interests can be quantified, providing a basis for subsequent data sorting.

[0098] S503 , based on the similarity calculation result, sorting the feature vectors corresponding to the different modal data to obtain a sorting result.

[0099] It can be understood that according to the similarity between the user interest feature vector calculated in the above steps and the feature vectors corresponding to different modal data, these feature vectors can be sorted, with the aim of putting the data most relevant to the user's interests in front so that the user can find the content of their interest more quickly.

[0100] Here, feature vectors can be sorted in descending order of similarity. The data corresponding to feature vectors with higher similarity will be ranked higher in the sorting results. For example, if the feature vector of a financial analysis report has the highest similarity to the feature vector of a user's interest, then this report will be ranked first in the sorting results. Data with lower similarity, such as irrelevant advertising information or outdated news reports, will be ranked lower.

[0101] In this way, through this sorting method, users can directly see the data that best suits their interests without having to filter through a large number of related search results one by one, thereby improving the efficiency of information acquisition.

[0102] S504: Determine target search information corresponding to the input information according to the ranking result.

[0103] Here, after obtaining the ranking results, the target search information corresponding to the input information can be determined based on the results. Since the ranking results have been arranged according to the similarity with the user's interests, the top-ranked data can be directly selected as the target search information.

[0104] For example, when determining the target search information, the amount of target search information to be selected can be determined based on specific needs and scenarios. For example, in scenarios where speed of information acquisition is required, only the first few pieces of data in the sorted results can be selected as the target search information; whereas in scenarios where a more comprehensive understanding of the information is required, a larger amount of data can be selected for display.

[0105] In some other embodiments, in order to further improve the quality and accuracy of the target search information, other factors may be combined for comprehensive judgment. For example, considering the timeliness of the data, the most recently released data may be given priority; considering the authority of the data, data from reliable sources may be selected.

[0106] The multimodal data search method, device, storage medium, and computer equipment provided in the embodiments of the present disclosure perform multimodal unified encoding of different modal data in the knowledge base, construct a multimodal knowledge base containing a unified semantic space feature vector, and sort the search results in combination with the user's historical behavior information during the retrieval process. This can achieve efficient retrieval of cross-modal data, accurately match the multimodal information input by the user with the knowledge base content, and optimize the order of search result presentation based on the user's personalized preferences. In this way, by breaking down the semantic barriers between different modal data and making full use of user behavior data to mine personalized needs, the relevance and accuracy of search results can be greatly improved, providing a more personalized search experience, and effectively optimizing search efficiency, thereby improving user experience.

[0107] Those skilled in the art will understand that in the above-mentioned method of the specific implementation method, the writing order of each step does not mean a strict execution order and does not constitute any limitation on the implementation process. The specific execution order of each step should be determined by its function and possible internal logic.

[0108] Based on the same inventive concept, the embodiments of the present disclosure also provide a multimodal data search device corresponding to the multimodal data search method. Since the principle of solving the problem by the device in the embodiments of the present disclosure is similar to the above-mentioned multimodal data search method in the embodiments of the present disclosure, the implementation of the device can refer to the implementation of the method, and the repeated parts will not be repeated.

[0109] Reference Figure 6 FIG. 1 is a schematic diagram of a multimodal data search device 600 provided in an embodiment of the present disclosure, wherein the device includes:

[0110] The knowledge base construction module 601 is used to obtain a knowledge base and perform multimodal unified encoding on different modal data in the knowledge base to obtain a multimodal knowledge base; wherein the multimodal knowledge base includes feature vectors corresponding to different modal data in a unified semantic space;

[0111] The information acquisition module 602 is used to acquire user input information; wherein the user input information includes at least one of images, text, audio and video;

[0112] An information retrieval module 603 is configured to retrieve multimodal information from the multimodal knowledge base that meets the requirements for similarity with the input information; wherein the multimodal information includes feature vectors corresponding to different modal data that meet the requirements for similarity with the input information;

[0113] The information determination module 604 is used to sort the feature vectors corresponding to the different modal data based on the user's historical behavior information, and determine the target search information corresponding to the input information based on the sorting result.

[0114] In some possible embodiments, the knowledge base includes a static knowledge database and a dynamic knowledge database, and the multimodal knowledge base includes a static knowledge base and a dynamic knowledge base; the knowledge base construction module 601 is specifically used to:

[0115] Acquire the static knowledge database; wherein the static knowledge database includes image data, text data, audio data and video data;

[0116] Performing multimodal unified encoding on the image data, the text data, the audio data, and the video data, respectively, and performing feature fusion processing and semantic alignment processing on the encoding results of each modal data to obtain the static knowledge base; wherein the static knowledge base includes feature vectors corresponding to each modal data in a unified semantic space;

[0117] Establishing a real-time data acquisition channel; the real-time data acquisition channel supports access to multi-source and multi-modal real-time data;

[0118] Acquiring the dynamic knowledge database based on the real-time data acquisition channel; wherein the dynamic knowledge database includes real-time image data, real-time text data, real-time audio data and real-time video data;

[0119] The real-time image data, real-time text data, real-time audio data and real-time video data are respectively subjected to multimodal unified encoding, and the encoding results of each modal real-time data are subjected to feature fusion processing and semantic alignment processing to obtain the dynamic knowledge base; wherein, the dynamic knowledge base includes feature vectors corresponding to each modal real-time data in a unified semantic space.

[0120] In some possible embodiments, the knowledge base construction module 601 is specifically used to:

[0121] For image data, a convolutional neural network is used to extract visual features of the image data to obtain image visual features of the image data, and the image visual features are mapped to a unified semantic space to obtain an image feature vector;

[0122] For text data, using a natural language processing method to perform semantic analysis on the text data, and generating a text feature vector in a unified semantic space based on the semantic analysis result;

[0123] For audio data, extracting spectral features of the audio data based on an acoustic model, identifying and extracting semantic labels of the audio data, and encoding the spectral features and the semantic labels into audio feature vectors in a unified semantic space;

[0124] For video data, visual features of the video frames in the video data are extracted based on a convolutional neural network, and features of the audio part of the video data are extracted in combination with audio processing technology. In addition, the visual features and audio features of the video data are fused and mapped to a unified semantic space to obtain a video feature vector.

[0125] In some possible embodiments, the information retrieval module 603 is specifically configured to:

[0126] Encoding the input information into a query feature vector in a unified semantic space; and determining a candidate data set matching the query feature vector based on the static knowledge base and the dynamic knowledge base;

[0127] Calculating the similarity between the query feature vector and the feature vectors corresponding to each modality data in the candidate data set based on a vector similarity calculation method;

[0128] According to a preset similarity threshold and the similarity between the query feature vector and the feature vectors corresponding to each modal data in the candidate data set, the feature vectors corresponding to the different modal data whose similarity with the input information meets the requirements are determined.

[0129] In some possible embodiments, the information retrieval module 603 is specifically configured to:

[0130] Performing semantic parsing on the input information, and decomposing the input information into a plurality of query subtasks based on the semantic parsing results; wherein each query subtask corresponds to one or more data shards in the static knowledge base and / or the dynamic knowledge base;

[0131] Allocating the multiple query subtasks to multiple computing nodes for parallel processing includes:

[0132] For each query subtask, encoding the input information fragment corresponding to the query subtask into a query feature subvector in a unified semantic space;

[0133] Using a vector similarity calculation method, searching, in the data shard corresponding to the query subtask, feature vectors corresponding to different modal data whose similarity to the query feature subvector meets a preset initial similarity threshold, to obtain an initial candidate data set corresponding to the computing node;

[0134] The initial candidate data sets corresponding to the various computing nodes are aggregated to obtain the candidate data sets.

[0135] In some possible embodiments, the information determination module 604 is specifically configured to:

[0136] Determine a user interest feature vector based on user historical behavior information; wherein the user historical behavior information includes click history, stay time, and search history;

[0137] Calculating the similarity between the user interest feature vector and the feature vectors corresponding to the different modal data;

[0138] Based on the similarity calculation result, the feature vectors corresponding to the different modal data are sorted to obtain a sorting result;

[0139] According to the ranking result, target search information corresponding to the input information is determined.

[0140] In some possible embodiments, the information acquisition module 602 is specifically configured to:

[0141] Get the user's initial input information;

[0142] Determining a target encryption algorithm based on the data security level of the initial input information, and encrypting the initial input information based on the target encryption algorithm;

[0143] Performing sensitive information identification on the initial input information;

[0144] If sensitive information exists in the initial input information, anonymize the sensitive information, and determine the input information based on the anonymized sensitive information and the encrypted initial input information;

[0145] In a case where no sensitive information exists in the initial input information, the input information is determined based on the encrypted initial input information.

[0146] Based on the same technical concept, the embodiment of the present disclosure also provides a computer device. Figure 7 , which is a schematic diagram of the structure of a computer device 700 provided in an embodiment of the present disclosure, includes a processor 701, a memory 702, and a bus 703. The memory 702 is used to store execution instructions and includes a memory 7021 and an external memory 7022. The memory 7021, also referred to as internal memory, is used to temporarily store calculation data in the processor 701 and data exchanged with an external memory 7022, such as a hard disk. The processor 701 exchanges data with the external memory 7022 through the memory 7021.

[0147] In the embodiment of the present application, the memory 702 is specifically used to store application code for executing the solution of the present application, and the execution is controlled by the processor 701. That is, when the computer device 700 is running, the processor 701 communicates with the memory 702 via the bus 703, so that the processor 701 executes the application code stored in the memory 702, thereby performing the method described in any of the aforementioned embodiments.

[0148] Among them, the memory 702 can be, but is not limited to, random access memory (RAM), read only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), etc.

[0149] The processor 701 may be an integrated circuit chip with signal processing capabilities. The above-mentioned processor may be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it may also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The various methods, steps and logic block diagrams disclosed in the embodiments of the present invention can be implemented or executed. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0150] It should be understood that the structure illustrated in the embodiments of the present application does not constitute a specific limitation on the computer device 700. In other embodiments of the present application, the computer device 700 may include more or fewer components than shown, or may combine or separate certain components, or arrange the components differently. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.

[0151] The present disclosure also provides a computer-readable storage medium having a computer program stored thereon. When executed by a processor, the computer program executes the steps of the multimodal data search method described in the above method embodiment. The storage medium may be a volatile or non-volatile computer-readable storage medium.

[0152] The embodiments of the present disclosure also provide a computer program product, which carries program code. The instructions included in the program code can be used to execute the steps of the multimodal data search method described in the above method embodiment. For details, please refer to the above method embodiment and will not be repeated here.

[0153] The computer program product may be implemented in hardware, software, or a combination thereof. In one embodiment, the computer program product is implemented as a computer storage medium. In another embodiment, the computer program product is implemented as a software product, such as a software development kit (SDK).

[0154] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working process of the system and device described above can refer to the corresponding process in the aforementioned method embodiment, and will not be repeated here. In the several embodiments provided in the present disclosure, it should be understood that the disclosed system and method can be implemented in other ways. The device embodiments described above are merely schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some communication interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0155] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.

[0156] In addition, each functional unit in each embodiment of the present disclosure may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.

[0157] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a non-volatile computer-readable storage medium that is executable by a processor. Based on this understanding, the technical solution of the present disclosure, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present disclosure. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.

[0158] Finally, it should be noted that the above-described embodiments are only specific implementation methods of the present disclosure, which are used to illustrate the technical solutions of the present disclosure, rather than to limit them. The scope of protection of the present disclosure is not limited thereto. Although the present disclosure has been described in detail with reference to the above-described embodiments, those skilled in the art should understand that any person skilled in the art can modify or easily conceive of changes to the technical solutions described in the above-described embodiments within the technical scope disclosed in the present disclosure, or replace some of the technical features therein with equivalents. Such modifications, changes, or replacements do not deviate from the spirit and scope of the technical solutions of the embodiments of the present disclosure, and should be included in the scope of protection of the present disclosure. Therefore, the scope of protection of the present disclosure shall be subject to the scope of protection of the claims.

Claims

1. A multimodal data search method, characterized in that: include: Obtaining a knowledge base, and performing multimodal unified encoding on different modal data in the knowledge base to obtain a multimodal knowledge base; wherein the multimodal knowledge base includes feature vectors corresponding to different modal data in a unified semantic space; Obtaining user input information; wherein the user input information includes at least one of an image, text, audio, and video; Retrieving multimodal information from the multimodal knowledge base, the multimodal information having a similarity to the input information meeting the requirements; wherein the multimodal information includes feature vectors corresponding to different modal data having a similarity to the input information meeting the requirements; The feature vectors corresponding to the different modal data are sorted based on the user's historical behavior information, and the target search information corresponding to the input information is determined based on the sorting result.

2. The method according to claim 1, characterized in that The knowledge base includes a static knowledge database and a dynamic knowledge database, and the multimodal knowledge base includes a static knowledge base and a dynamic knowledge base; the acquisition of the knowledge base and performing multimodal unified encoding on different modal data in the knowledge base to obtain the multimodal knowledge base includes: Acquire the static knowledge database; wherein the static knowledge database includes image data, text data, audio data and video data; Performing multimodal unified encoding on the image data, the text data, the audio data, and the video data, respectively, and performing feature fusion processing and semantic alignment processing on the encoding results of each modal data to obtain the static knowledge base; wherein the static knowledge base includes feature vectors corresponding to each modal data in a unified semantic space; Establishing a real-time data acquisition channel; the real-time data acquisition channel supports access to multi-source and multi-modal real-time data; Acquiring the dynamic knowledge database based on the real-time data acquisition channel; wherein the dynamic knowledge database includes real-time image data, real-time text data, real-time audio data and real-time video data; The real-time image data, real-time text data, real-time audio data and real-time video data are respectively subjected to multimodal unified encoding, and the encoding results of each modal real-time data are subjected to feature fusion processing and semantic alignment processing to obtain the dynamic knowledge base; wherein, the dynamic knowledge base includes feature vectors corresponding to each modal real-time data in a unified semantic space.

3. The method according to claim 2, characterized in that The performing multimodal unified encoding on the different modal data in the knowledge base respectively includes: For image data, a convolutional neural network is used to extract visual features of the image data to obtain image visual features of the image data, and the image visual features are mapped to a unified semantic space to obtain an image feature vector; For text data, using a natural language processing method to perform semantic analysis on the text data, and generating a text feature vector in a unified semantic space based on the semantic analysis result; For audio data, extracting spectral features of the audio data based on an acoustic model, identifying and extracting semantic labels of the audio data, and encoding the spectral features and the semantic labels into audio feature vectors in a unified semantic space; For video data, visual features of the video frames in the video data are extracted based on a convolutional neural network, and features of the audio part of the video data are extracted in combination with audio processing technology. In addition, the visual features and audio features of the video data are fused and mapped to a unified semantic space to obtain a video feature vector.

4. The method according to claim 2, characterized in that The retrieving multimodal information having a similarity with the input information that meets the requirements from the multimodal knowledge base includes: Encoding the input information into a query feature vector in a unified semantic space; and determining a candidate data set matching the query feature vector based on the static knowledge base and the dynamic knowledge base; Calculating the similarity between the query feature vector and the feature vectors corresponding to each modality data in the candidate data set based on a vector similarity calculation method; According to a preset similarity threshold and the similarity between the query feature vector and the feature vectors corresponding to each modal data in the candidate data set, the feature vectors corresponding to the different modal data whose similarity with the input information meets the requirements are determined.

5. The method according to claim 4, characterized in that The step of encoding the input information into a query feature vector in a unified semantic space; and determining a candidate data set matching the query feature vector based on the static knowledge base and the dynamic knowledge base, comprises: Performing semantic parsing on the input information, and decomposing the input information into a plurality of query subtasks based on the semantic parsing results; wherein each query subtask corresponds to one or more data shards in the static knowledge base and / or the dynamic knowledge base; Allocating the multiple query subtasks to multiple computing nodes for parallel processing includes: For each query subtask, encoding the input information fragment corresponding to the query subtask into a query feature subvector in a unified semantic space; Using a vector similarity calculation method, searching, in the data shard corresponding to the query subtask, feature vectors corresponding to different modal data whose similarity to the query feature subvector meets a preset initial similarity threshold, to obtain an initial candidate data set corresponding to the computing node; The initial candidate data sets corresponding to the various computing nodes are aggregated to obtain the candidate data sets.

6. The method according to claim 4, characterized in that The step of sorting the feature vectors corresponding to the different modal data based on the user's historical behavior information and determining the target search information corresponding to the input information based on the sorting result includes: Determine a user interest feature vector based on user historical behavior information; wherein the user historical behavior information includes click history, stay time, and search history; Calculating the similarity between the user interest feature vector and the feature vectors corresponding to the different modal data; Based on the similarity calculation result, the feature vectors corresponding to the different modal data are sorted to obtain a sorting result; According to the ranking result, target search information corresponding to the input information is determined.

7. The method according to claim 1, characterized in that The obtaining of user input information includes: Get the user's initial input information; Determining a target encryption algorithm based on the data security level of the initial input information, and encrypting the initial input information based on the target encryption algorithm; Performing sensitive information identification on the initial input information; If sensitive information exists in the initial input information, anonymize the sensitive information, and determine the input information based on the anonymized sensitive information and the encrypted initial input information; In a case where no sensitive information exists in the initial input information, the input information is determined based on the encrypted initial input information.

8. A multimodal data search device, characterized in that: include: A knowledge base construction module is used to obtain a knowledge base and perform multimodal unified encoding on different modal data in the knowledge base to obtain a multimodal knowledge base; wherein the multimodal knowledge base includes feature vectors corresponding to different modal data in a unified semantic space; An information acquisition module, configured to acquire user input information; wherein the user input information includes at least one of images, text, audio, and video; an information retrieval module, configured to retrieve multimodal information from the multimodal knowledge base that meets the requirements for similarity with the input information; wherein the multimodal information includes feature vectors corresponding to different modal data that meet the requirements for similarity with the input information; An information determination module is used to sort the feature vectors corresponding to the different modal data based on the user's historical behavior information, and determine the target search information corresponding to the input information based on the sorting result.

9. A storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.

10. A computer device comprising a storage medium, a processor, and a computer program stored in the storage medium and executable on the processor, wherein: When the processor executes the computer program, the method according to any one of claims 1 to 7 is implemented.

Citation Information

Cited By

  • Digital library information retrieval method and system based on Internet data

    CN121233836A

  • A digital library information retrieval method and system based on Internet data

    CN121233836B