Multi-modal data retrieval method and system

By building a multimodal data processing and search engine and blockchain storage network, the problems of low efficiency and insufficient accuracy of multimodal data retrieval are solved, and efficient and secure multimodal data retrieval is achieved.

CN120492685AInactive Publication Date: 2025-08-15SHENZHEN HOLOGRAPHIC FIELD CULTURE TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510666793.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-22
Publication Date
2025-08-15
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

When existing data retrieval technology processes huge multimodal data sets, it has problems of low retrieval efficiency and insufficient accuracy, and it is impossible to effectively extract and fuse various data types such as images, text, and audio.

Method used

Using artificial intelligence algorithms to build a multimodal data processing and search engine, combining blockchain technology to build a multimodal data storage network, and through multimodal data processing models, feature fusion models and search map generation models, efficient processing and distributed storage of multimodal data are achieved.

Benefits of technology

It significantly improves the search speed and accuracy, effectively utilizes computing resources, prevents data tampering, and improves the security of data storage and the relevance of search results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120492685A_ABST
    Figure CN120492685A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of data retrieval, and discloses a multi-modal data retrieval method and system. The method comprises the following steps: constructing a multi-modal data processing and retrieval engine and a multi-modal data storage network; according to the real-time data information, a multi-modal data processing and retrieval engine is used for processing real-time multi-modal data, and a multi-modal data storage network is used for carrying out distributed storage on the real-time standard multi-modal data and a real-time retrieval atlas of the real-time standard multi-modal data; and according to the real-time query information, using a multi-modal data processing and retrieval engine to match a target real-time retrieval map in the multi-modal data storage network, and according to the target retrieval map, extracting target standard multi-modal data in the multi-modal data storage network. The problems of low retrieval efficiency and insufficient accuracy in the prior art are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of data retrieval, and in particular relates to a multimodal data retrieval method and system. Background Art

[0002] With the rapid development of information technology, information digitization has been widely used in enterprise management. However, as enterprises grow in size, the volume of enterprise data is growing exponentially, posing a severe challenge to data retrieval. Efficient and accurate retrieval in the ocean data center has become a key development direction and the foundation for scientific enterprise management.

[0003] Existing data retrieval technology has the following defects: 1) Low retrieval efficiency: Traditional retrieval methods often exhibit significant efficiency bottlenecks when processing large data sets, resulting in excessively long retrieval times and requiring a large amount of computing resources, resulting in retrieval efficiency that cannot meet requirements. 2) Insufficient accuracy: Enterprise data is often multimodal data with multiple types of data (such as images, text, audio, etc.), and each data type has its own unique characteristics. Existing technologies cannot fully extract and integrate these features, resulting in a loss of accuracy in retrieval results. Summary of the Invention

[0004] In order to solve the problems of low retrieval efficiency and insufficient accuracy in the prior art, the present invention aims to provide a multimodal data retrieval method and system.

[0005] The technical solution adopted in the present invention is: A multimodal data retrieval method comprises the following steps: Use artificial intelligence algorithms to build a multimodal data processing and retrieval engine, and use blockchain technology to build a multimodal data storage network, and connect the multimodal data storage network to the multimodal data processing and retrieval engine; Collect real-time multimodal data and its real-time data information, use the multimodal data processing and retrieval engine to process the real-time multimodal data based on the real-time data information, and use the multimodal data storage network to distribute the obtained real-time standard multimodal data and its real-time retrieval map; Collect users' real-time query information, use multimodal data processing and retrieval engines based on the real-time query information, match the target real-time retrieval graph in the multimodal data storage network, and extract target standard multimodal data in the multimodal data storage network based on the target retrieval graph.

[0006] Furthermore, the real-time multimodal data includes real-time image data, real-time text data, and real-time audio data.

[0007] Furthermore, the multimodal data processing and retrieval engine includes a multimodal data processing component and a multimodal data retrieval component; The multimodal data processing component includes a multimodal data processing model, a multimodal feature fusion model, and a retrieval graph generation model; The multimodal data retrieval component includes a keyword extraction model and a retrieval graph matching model; Multimodal data storage networks include graph databases and distributed storage networks.

[0008] Furthermore, the multimodal data processing model is constructed based on the MOGRPO-DPA-DALL-E-Wav2Vec-CCA-ASGA algorithm, and the multimodal data processing model includes a multimodal data preprocessing component constructed based on the MOGRPO-DPA algorithm and a multimodal data alignment component constructed based on the DALL-E-Wav2Vec-CCA-ASGA algorithm, which are connected in sequence. The multimodal data preprocessing component includes a preprocessing strategy generation module constructed based on the MOGRPO algorithm and a data preprocessing algorithm library constructed based on the DPA algorithm, which are connected in sequence. The multimodal data alignment component includes an image and text feature extraction module constructed based on the DALL-E algorithm, an audio feature extraction module constructed based on the Wav2Vec algorithm, a cross-modal association module constructed based on the CCA algorithm, and a multimodal data alignment module constructed based on the ASGA algorithm, which are connected in sequence. The multimodal feature fusion model is constructed based on FPN-LSTM-VGGish-MLP, and the multimodal feature fusion model includes a high-order image feature extraction module constructed based on the FPN algorithm, a deep text feature extraction module constructed based on the LSTM algorithm, a deep audio feature extraction module constructed based on the VGGish algorithm, and a multimodal feature fusion module constructed based on the MLP, which are connected in sequence. The multimodal feature fusion module is respectively connected to the high-order image feature extraction module, the deep text feature extraction module, and the deep audio feature extraction module; The retrieval graph generation model is constructed based on the GCN-AGN algorithm, and the retrieval graph generation model includes a retrieval graph generation module and an attention weight module connected in sequence.

[0009] Furthermore, the keyword extraction model is constructed based on the BERT-BiLSTM-CRF algorithm, and the keyword extraction model includes a word embedding module constructed based on the BERT algorithm, a query semantic feature extraction module constructed based on the BiLSTM algorithm, and a keyword annotation module constructed based on the CRF algorithm, which are sequentially connected; The retrieval graph matching model is constructed based on the LSTM-GCN-k-NN algorithm, and the retrieval graph matching model includes a keyword semantic feature extraction module constructed based on the LSTM algorithm, a retrieval graph feature extraction module constructed based on the GCN algorithm, and a retrieval graph matching module constructed based on the k-NN algorithm. The retrieval graph matching module is connected to the keyword semantic feature extraction module and the retrieval graph feature extraction module respectively.

[0010] Furthermore, real-time multimodal data and its real-time data information are collected, and based on the real-time data information, a multimodal data processing and retrieval engine is used to process the real-time multimodal data, and a multimodal data storage network is used to distributely store the obtained real-time standard multimodal data and its real-time retrieval graph, including the following steps: Use the multimodal data processing component of the multimodal data processing and retrieval engine to collect real-time multimodal data and its real-time data information; According to the real-time data information, the multimodal data processing model of the multimodal data processing component is used to process the real-time multimodal data to obtain real-time standard multimodal data; Use the multimodal feature fusion model to perform multimodal feature fusion on real-time standard multimodal data to obtain real-time multimodal fusion features; According to the real-time multimodal fusion features, a retrieval graph generation model is used to generate a retrieval graph to obtain a real-time retrieval graph; Using a distributed storage network of a multimodal data storage network, real-time standard multimodal data is distributedly stored and corresponding real-time storage addresses are obtained; A graph database of a multimodal data storage network is used to store real-time retrieval graphs and corresponding real-time storage addresses of real-time standard multimodal data.

[0011] Furthermore, based on the real-time data information, a multimodal data processing model is used to process the real-time multimodal data to obtain real-time standard multimodal data, including the following steps: According to the real-time data information, the multimodal data preprocessing component of the multimodal data processing model is used to preprocess the corresponding real-time multimodal data to obtain preprocessed real-time multimodal data; The multimodal data alignment component of the multimodal data processing model is used to perform multimodal data alignment on the preprocessed real-time multimodal data to obtain real-time standard multimodal data.

[0012] Furthermore, a multimodal feature fusion model is used to perform multimodal feature fusion on the real-time standard multimodal data to obtain real-time multimodal fusion features, including the following steps: Use the high-order image feature extraction module of the multimodal feature fusion model to extract real-time high-order image features of real-time standard multimodal data; Use the deep text feature extraction module of the multimodal feature fusion model to extract real-time deep text features from real-time standard multimodal data; A deep audio feature extraction module using a multimodal feature fusion model is used to extract real-time deep audio features from real-time standard multimodal data. The multimodal feature fusion module of the multimodal feature fusion model is used to fuse real-time high-order image features, real-time deep text features, and real-time deep audio features to obtain real-time multimodal fusion features.

[0013] Furthermore, real-time query information of users is collected, and based on the real-time query information, a multimodal data processing and retrieval engine is used to match a target real-time retrieval graph in a multimodal data storage network, and based on the target retrieval graph, target standard multimodal data is extracted from the multimodal data storage network, including the following steps: Use the multimodal data processing and retrieval engine's multimodal data retrieval component to collect users' real-time query information; Use the keyword extraction model of the multimodal data retrieval component to extract real-time keywords for real-time query information; Based on real-time keywords, the retrieval graph matching model of the multimodal data retrieval component is used to perform retrieval and matching in the graph database of the multimodal data storage network to obtain the target real-time retrieval graph; According to the target stored in the graph database, the target storage address corresponding to the graph is retrieved in real time, and the target standard multimodal data is extracted in the distributed storage network of the multimodal data storage network.

[0014] A multimodal data retrieval system is used to implement a multimodal data retrieval method. The system includes an engine and storage network construction unit, a data processing and distributed storage unit, and a query matching and data retrieval unit, which are connected in sequence.

[0015] The beneficial effects of the present invention are: The present invention provides a multimodal data retrieval method and system, which utilizes a multimodal data processing and retrieval engine constructed using an artificial intelligence algorithm. This method can efficiently process and analyze massive amounts of multimodal data. Combined with a distributed storage network based on blockchain technology, this method can achieve rapid access and parallel processing of data, significantly improving retrieval speed and overcoming the efficiency bottleneck of traditional methods in processing large-scale data. Furthermore, this method can more effectively utilize computing resources, achieve rational allocation and efficient utilization of computing resources, and avoid resource waste. Using artificial intelligence algorithms, various features (such as images, text, audio, etc.) in multimodal data can be extracted and integrated more comprehensively and deeply to construct a more accurate real-time retrieval graph, thereby improving the accuracy and relevance of retrieval results. By utilizing the decentralized and tamper-proof characteristics of blockchain technology, a multimodal data storage network can be constructed, which can effectively prevent data tampering and malicious attacks and improve the security of data storage.

[0016] Other beneficial effects of the present invention will be further described in the specific embodiments. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 It is a flowchart of the multimodal data retrieval method in the present invention.

[0018] Figure 2 It is a structural block diagram of the multimodal data retrieval system in the present invention. DETAILED DESCRIPTION

[0019] The present invention will be further explained below with reference to the accompanying drawings and specific embodiments.

[0020] Example 1: like Figure 1 As shown, this embodiment provides a multimodal data retrieval method, comprising the following steps: S1: Use artificial intelligence algorithms to build a multimodal data processing and retrieval engine, and use blockchain technology to build a multimodal data storage network, and connect the multimodal data storage network to the multimodal data processing and retrieval engine, including the following steps: S1-1: Collect a number of historical multimodal data and their historical data information, and use multiple artificial intelligence complex algorithms to build a multimodal data processing model based on the historical multimodal data and their historical data information, and obtain a number of historical standard multimodal data; The multimodal data processing model is based on Multi-Objective Group Relative Policy Optimization (MOGRPO) - Data Preprocessing Algorithm (DPA) - DALL-E-Wav2Vec - Cross-modal Correlation Algorithm (CCA) - Accurate Snow Geese Algorithm Algorithm (ASGA) algorithm, and the multimodal data processing model includes a multimodal data preprocessing component based on the MOGRPO-DPA algorithm and a multimodal data alignment component based on the DALL-E-Wav2Vec-CCA-ASGA algorithm, which are connected in sequence. The multimodal data preprocessing component includes a preprocessing strategy generation module based on the MOGRPO algorithm and a data preprocessing algorithm library based on the DPA algorithm, which are connected in sequence. The multimodal data alignment component includes an image and text feature extraction module based on the DALL-E algorithm, an audio feature extraction module based on the Wav2Vec algorithm, a cross-modal association module based on the CCA algorithm, and a multimodal data alignment module based on the ASGA algorithm, which are connected in sequence. The preprocessing strategy generation module includes an objective function set, a strategy network, an experience replay pool, and an intelligent agent. The intelligent agent is connected to the objective function set, the experience replay pool, and the strategy network respectively. The intelligent agent learns the historical preprocessing strategy through the experience replay pool and continuously optimizes its own strategy generation ability. The intelligent agent controls the strategy network based on the learned experience to generate a more effective preprocessing strategy. The design of the experience replay pool and the intelligent agent enables the model to continuously learn and optimize, thereby improving the quality of strategy generation. The preprocessing strategy generation module adopts a group exploration method, which can avoid falling into the local optimal solution to a certain extent. The strategy network outputs the distribution probability of the action under a given state. The preprocessing strategy generation module directly updates the strategy network through the gradient, eliminating the Critic model in traditional reinforcement learning, making the algorithm structure more concise. The data preprocessing algorithm library stores several DPA algorithm packages, including data According to the preprocessing strategy, a series of preprocessing algorithms such as cleaning algorithm, Gaussian denoising algorithm, text filling algorithm, image enhancement algorithm and normalization algorithm can be used to preprocess the data. The image and text feature extraction module adopts the Transformer network and learns the relationship between text and image through a large amount of training data. It performs well in image-text alignment and can generate text descriptions that are highly relevant to the image content for extracting image and text features. The audio feature extraction module is used to extract audio features. The cross-modal association module is used to integrate the extracted image features, text features and audio features to construct a cross-modal association matrix. In this embodiment, an attention mechanism (such as Co-Attention) is used to learn the association between different modalities. The multimodal data alignment module is used to iteratively search for the best modality alignment solution in the search space to achieve multimodal data alignment. S1-2: Based on several historical standard multimodal data, use multi-feature fusion and deep learning algorithms to construct multimodal feature fusion and obtain several historical multimodal fusion features; The multimodal feature fusion model is built based on the Feature Pyramid Network (FPN)-Long Short-Term Memory Network (LSTM)-VGGish-Multilayer Perceptron (MLP). The multimodal feature fusion model includes a high-order image feature extraction module based on the FPN algorithm, a deep text feature extraction module based on the LSTM algorithm, a deep audio feature extraction module based on the VGGish algorithm, and a multimodal feature fusion module based on the MLP, which are connected in sequence. The multimodal feature fusion module is connected to the high-order image feature extraction module, the deep text feature extraction module, and the deep audio feature extraction module respectively. The high-order image feature extraction module is used to extract high-order image features of the image. FPN forms a multi-scale feature pyramid by constructing a top-down path and a bottom-up path, which can effectively capture image features of different scales. Multi-scale feature extraction can better capture the details and contextual information in the image, thereby improving the representation ability of image features, which is very helpful for image understanding and subsequent multimodal fusion; the deep text feature extraction module is used to extract the deep features of the text. LSTM is a special recurrent neural network (RNN) that can effectively process sequence data and capture long-distance dependencies in the text; LSTM can better understand the semantics and contextual information of the text and generate high-quality text feature representations. This is very useful for tasks such as text classification and sentiment analysis, and can provide rich semantic information in multimodal fusion; the deep audio feature extraction module is used to extract deep features of audio. VGGish is an audio feature extraction method based on the VGG network, which can extract useful feature representations from audio signals; it captures key information in audio signals and generates high-quality audio feature representations, which can provide rich audio information in multimodal fusion; the multimodal feature fusion module is used to fuse features from the high-order image feature extraction module, the deep text feature extraction module and the deep audio feature extraction module; MLP fuses and combines features of different modalities through multiple fully connected layers and nonlinear activation functions; multimodal feature fusion can comprehensively utilize information from different modalities to improve the generalization ability and robustness of the model. Through the fusion of MLP, the model can better understand and process complex multimodal data, thereby achieving better performance in various tasks; S1-3: Based on several historical multimodal fusion features, a deep learning algorithm is used to build a retrieval graph generation model and obtain several historical retrieval graphs; The retrieval graph generation model is built based on the Graph Convolutional Network (GCN)-Attention Graph Network (AGN) algorithm, and includes a retrieval graph generation module and an attention weight module connected in sequence. The retrieval graph generation module is used to generate retrieval graphs. GCN is a deep learning model for graph data. It extracts features on the graph by learning the relationship between nodes and generates a graph representation. In the retrieval graph generation module, GCN can be used to analyze the associations between data such as documents, images, and audio, and construct a retrieval graph that reflects the relationships between the data. The retrieval graph generated by GCN can capture the complex relationships between the data and provide rich structured information, which helps improve the accuracy and efficiency of the retrieval system because the graph can reveal the potential connections between the data, thereby supporting more accurate retrieval and recommendation. The attention weight module is built based on the attention mechanism and is used to assign different weights to different parts of the retrieval graph. The attention mechanism is a mechanism that can dynamically adjust weights according to the importance of the input data, allowing the model to pay more attention to task-related information. In the attention weight module, the attention mechanism can be used to analyze the importance of different nodes and edges in the retrieval graph, thereby assigning appropriate weights to them. Through the attention weight module, the model can process the information in the retrieval graph more flexibly, highlighting the parts most relevant to the retrieval task. This helps improve the performance of the retrieval system because the model can focus more on the most useful information, thereby generating more accurate and relevant retrieval results. S1-4: Integrate the multimodal data processing model, the multimodal feature fusion model, and the retrieval graph generation model to obtain the multimodal data processing component; S1-5: Collect some historical query information, and use natural language processing and deep learning algorithms to build a keyword extraction model based on the historical query information to obtain some historical keywords; The keyword extraction model is based on the Bidirectional Encoder Representations from Transformers (BERT)-Bidirectional Long Short-Term Memory (BiLSTM)-Conditional Random Field (CRF) algorithm. The model consists of a word embedding module based on the BERT algorithm, a query semantic feature extraction module based on the BiLSTM algorithm, and a keyword tagging module based on the CRF algorithm. The word embedding module converts each word in the query information into a vector representation in a high-dimensional space. Utilizing BERT's pre-training capabilities, it captures the changes in the meaning of words in context. The word vectors generated by BERT are rich in semantic information, helping subsequent modules better understand the query information. The query semantic feature extraction module uses a bidirectional LSTM network to learn the long-term dependencies and contextual information of the input sequence. BiLSTM can capture the complex relationships between words, improving the model's ability to understand semantics, providing more accurate input features for the keyword annotation module, and improving the accuracy of keyword recognition. The keyword annotation module uses CRF to annotate sequences and identify keyword entities in the input text. CRF can use contextual information for constraints, reducing incorrect annotations and improving the accuracy of keyword recognition. S1-6: Based on several historical search graphs and several historical keywords, use deep learning algorithms to build a search graph matching model; The retrieval graph matching model is built based on the LSTM-GCN-k-nearest neighbor algorithm (k-NN), and includes a keyword semantic feature extraction module built based on the LSTM algorithm, a retrieval graph feature extraction module built based on the GCN algorithm, and a retrieval graph matching module built based on the k-NN algorithm. The retrieval graph matching module is connected to the keyword semantic feature extraction module and the retrieval graph feature extraction module respectively. The keyword semantic feature extraction module is used to extract the semantic features of keywords. By learning the long-term dependencies in keyword sequences, LSTM can capture the contextual information of keywords and generate rich semantic representations. The keyword semantic features generated by LSTM can capture the deep meaning and contextual information of keywords, which helps improve the accuracy of retrieval. By understanding the semantics of keywords, the model can better match documents or images related to user queries. The retrieval graph feature extraction module extracts features on the graph by learning the relationships between nodes. In this module, GCN is used to extract features of the retrieval graph. The retrieval graph can represent the association between documents, images, audio and other data. GCN can capture the structural information in the graph and generate the feature representation of the graph. The retrieval graph features extracted by GCN can capture the complex relationships between data and provide rich structured information, which helps to improve the accuracy and efficiency of retrieval because the graph can reveal the potential connections between data, thereby supporting more accurate retrieval and recommendation; the retrieval graph matching module finds the nearest k neighbors by calculating the distance between the keyword semantic features and several retrieval graph features. In this module, k-NN is used to match the keyword semantic features and the retrieval graph features. By calculating the similarity between the features, k-NN can find the retrieval graph that is most relevant to the keyword semantics; the k-NN algorithm can effectively match the keyword semantic features and the retrieval graph features to find the most relevant retrieval graph; by calculating the similarity between the features, k-NN can improve the accuracy and efficiency of retrieval and support more accurate and relevant retrieval results; S1-7: Integrate the keyword extraction model and the retrieval graph matching model to obtain a multimodal data retrieval component, and integrate the multimodal data retrieval component and the multimodal data processing component to obtain a multimodal data processing and retrieval engine; S1-8: Use the Apache Graph Execution (Apache AGE) architecture to build a graph database that supports the Cypher query language. Use blockchain technology to build a distributed storage network and connect the graph database to the distributed ledger of the distributed storage network. Apache AGE is an open source graph data processing framework built on the PostgreSQL database, designed to provide efficient graph query and graph algorithm capabilities; Apache AGE's architectural design allows users to perform graph operations in a familiar SQL environment while leveraging the advantages of graph databases to process complex graph data, which is beneficial for S1-9: Integrate the graph database and the distributed storage network to obtain a multimodal data storage network, and connect the multimodal data storage network to the multimodal data processing and retrieval engine; S2: Collect real-time multimodal data and its real-time data information, use a multimodal data processing and retrieval engine to process the real-time multimodal data based on the real-time data information, and use a multimodal data storage network to distributely store the obtained real-time standard multimodal data and its real-time retrieval graph, including the following steps: S2-1: Use the multimodal data processing component of the multimodal data processing and retrieval engine to collect real-time multimodal data and its real-time data information; Real-time multimodal data includes real-time image data, real-time text data, and real-time audio data; real-time data information includes data size, format, dimensions, timestamp, name, source, purpose, author, etc. S2-2: Based on the real-time data information, the multimodal data processing model of the multimodal data processing component is used to process the real-time multimodal data to obtain real-time standard multimodal data, including the following steps: S2-2-1: Based on the real-time data information, use the multimodal data preprocessing component of the multimodal data processing model to preprocess the corresponding real-time multimodal data to obtain preprocessed real-time multimodal data, including the following steps: S2-2-1-1: Based on the real-time data information, use the preprocessing strategy generation module of the multimodal data preprocessing component of the multimodal data processing model to generate a corresponding real-time preprocessing strategy; S2-2-1-2: Based on the real-time preprocessing strategy, call the corresponding DPA algorithm package in the data preprocessing algorithm library of the multimodal data preprocessing component; S2-2-1-3: Use the DPA algorithm to encapsulate and preprocess the corresponding real-time multimodal data to obtain preprocessed real-time multimodal data; S2-2-2: Use the multimodal data alignment component of the multimodal data processing model to perform multimodal data alignment on the preprocessed real-time multimodal data to obtain real-time standard multimodal data, including the following steps: S2-2-2-1: Using the image and text feature extraction module of the multimodal data alignment component of the multimodal data processing model, extracting real-time image features and real-time text features of the pre-processed real-time multimodal data; S2-2-2-2: Use the audio feature extraction module of the multimodal data alignment component to extract real-time audio features of the preprocessed real-time multimodal data; S2-2-2-3: Use the cross-modal association module of the multimodal data alignment component to generate a corresponding real-time cross-modal association matrix based on real-time image features, real-time text features, and real-time audio features; S2-2-2-4: Based on the real-time cross-modal correlation matrix, use the multimodal data alignment module of the multimodal data alignment component to generate the optimal real-time multimodal data alignment solution, including the following steps: S2-2-2-4-1: Encode the initial real-time multimodal data alignment scheme into individual vectors of the initial solution, and based on the individual vectors, use the Tent-Logistic-Cosine chaotic mapping sequence to initialize and generate several initial solutions, thus obtaining an initial ASGA population consisting of several initial ASGA individuals (initial solutions); The formula is:

[0021] Where, ASGA individuals (initial solutions) generated for the Tent-Logistic-Cosine chaotic mapping sequence; is a randomly generated ASGA individual; are preset parameters; i is the individual indicator of ASGA; the initial population is generated by the Tent-Logistic-Cosine chaotic mapping sequence. Compared with the randomly distributed population, the initial position distribution of the improved ASGA population is more uniform, which expands the search range of the ASGA population in space and increases the diversity of group positions. To a certain extent, it improves the defect that the algorithm is prone to falling into local extreme values, thereby improving the optimization efficiency of the algorithm; S2-2-2-4-2: Taking minimizing the prediction error as the optimization goal, based on the real-time cross-modal correlation matrix, set the fitness function of the ASGA algorithm, and set the ASGA population parameters and the maximum number of iterations; The formula is:

[0022] Where, is the fitness function; is the prediction error function; is the real-time cross-modal correlation matrix weight; For the The moment Real-time cross-modal correlation matrix for real-time multimodal data; It is the time indication quantity; It is the real-time multimodal data indicator; i is the ASGA indicator; S2-2-2-4-3: Use the fitness function to obtain the initial fitness value of each initial ASGA individual in the initial ASGA population, and take the initial ASGA individual with the lowest fitness value as the leader goose; S2-2-2-4-4: Entering the exploration phase, the leader goose rotation mechanism, the calling guidance mechanism, and the dynamic reverse mechanism are introduced to iteratively update the initial ASGA population to obtain an updated ASGA population and retain the optimal individual; The leader goose rotation mechanism selects a new leader goose in each iteration based on the fitness value of the ASGA individuals. This mechanism can prevent the leader goose from falling into the local optimum too early and enhance the global search capability of the algorithm. The formula is:

[0023] Where, Be the leader of an update; For the The third-to-last initial ASGA individual with the lowest fitness value in the initial ASGA population at the number of iterations; For the The fifth-to-last initial ASGA individual in the initial ASGA population with the highest fitness value at the iteration number; is the current iteration number; is the optimal individual; is the first weight factor; is a random number generation function; The calling guidance mechanism uses the sound wave propagation attenuation model to adjust the position update of the ASGA individuals according to their distance from the leader goose. The position updates of ASGA individuals that are closer are more influenced by the leader goose and can quickly approach the optimal solution. The position updates of ASGA individuals that are farther away are less influenced by the leader goose and can maintain a certain exploration ability. This mechanism can prevent the group from excessively gathering or dispersing, and improve the local search accuracy of the algorithm. The formula is:

[0024] Where, is an updated ASGA individual; For the The number of iterations of the initial ASGA individual; is the sound intensity received by the initial ASGA individual; is the sound intensity parameter; is the initial sound intensity; is the minimum acceptable sound intensity; is the convergence factor; is the initial ASGA individual with the farthest distance; is a random parameter; is the Brownian motion function; is the Brownian motion parameter; is the XOR processing symbol;

[0025] Where, is the convergence factor; tanh(.) is the hyperbolic tangent function; is the current iteration number; is the maximum number of iterations; a max 、 a min are the maximum and minimum values of the convergence factor respectively; λ is the deceleration rate parameter, is the decrement period parameter, λ =-2 π , = π ; In the early stages of iteration, When the value of is large, the leader goose’s call guidance effect on ASGA individuals is strong, which helps to explore more solution spaces. In the later stages of iteration, the leader goose’s call guidance effect on ASGA individuals is weak, which helps to conduct fine search in local areas. Dynamic reversal mechanism, which dynamically reverses the initial ASGA individuals to improve the diversity of exploration directions and avoid falling into local optimality; The formula is:

[0026] Where, is an updated reverse ASGA individual; γ is the decreasing inertia coefficient; L max 、 L min are the maximum and minimum values of the vector space respectively; The leader goose of one update, several ASGA individuals of one update, and several reverse ASGA individuals of one update are integrated to obtain an ASGA population of one update, and the ASGA individual with the lowest fitness value is retained as the optimal individual; S2-2-2-4-5: Enter the development stage, introduce the abnormal boundary strategy, perform a second update on the ASGA population that has been updated once, obtain the second updated ASGA population, and retain the optimal individual; The abnormal boundary strategy calculates the difference between the fitness value of each ASGA individual updated at a time and the average fitness value of the group. For ASGA individuals whose fitness value is much higher than the group average, their position update method will be adjusted, such as using a larger step size or a smaller step size. This mechanism can help individuals avoid falling into local optimality and improve the convergence speed and accuracy of the algorithm; The formula is:

[0027] Where, is the ASGA individual with secondary update; is an updated ASGA individual; is the fitness function; is the average fitness value of the group; The ASGA individual with the highest fitness value; are the second weight factor and the third weight factor; is the Levy flight strategy parameter; S2-2-2-4-6: If the number of iterations is greater than or equal to the iteration threshold or the fitness value of the optimal individual is less than the fitness threshold, the optimal individual is output as the optimal solution; S2-2-2-4-7: Decode the individual vectors of the optimal solution to obtain the optimal real-time multimodal data alignment solution; S2-2-2-5: Perform multimodal data alignment on the preprocessed real-time multimodal data according to the optimal real-time multimodal data alignment scheme to obtain real-time standard multimodal data; S2-3: Use the multimodal feature fusion model to perform multimodal feature fusion on the real-time standard multimodal data to obtain real-time multimodal fusion features, including the following steps: S2-3-1: Use the high-order image feature extraction module of the multimodal feature fusion model to extract real-time high-order image features of real-time standard multimodal data; S2-3-2: Use the deep text feature extraction module of the multimodal feature fusion model to extract real-time deep text features of real-time standard multimodal data; S2-3-3: Use the deep audio feature extraction module of the multimodal feature fusion model to extract real-time deep audio features of real-time standard multimodal data; S2-3-4: Use the multimodal feature fusion module of the multimodal feature fusion model to fuse real-time high-order image features, real-time deep text features, and real-time deep audio features to obtain real-time multimodal fusion features; S2-4: Based on the real-time multimodal fusion features, a retrieval graph generation model is used to generate a retrieval graph to obtain a real-time retrieval graph, including the following steps: S2-4-1: Based on the real-time multimodal fusion features, the attention weight module of the retrieval graph generation model is used to generate the corresponding real-time attention weights; S2-4-2: Based on the real-time attention weight and the real-time multimodal fusion features, the retrieval graph generation module of the retrieval graph generation model is used to generate a retrieval graph to obtain a real-time retrieval graph; S2-5: Using the distributed storage network of the multimodal data storage network, the real-time standard multimodal data is distributedly stored and the corresponding real-time storage address is obtained; S2-6: Using the graph database of the multimodal data storage network, store the real-time retrieval graph and the corresponding real-time storage address of the real-time standard multimodal data; S3: Collect the user's real-time query information. Based on the real-time query information, use the multimodal data processing and retrieval engine to match the target real-time retrieval graph in the multimodal data storage network. Then, based on the target retrieval graph, extract the target standard multimodal data from the multimodal data storage network. This includes the following steps: S3-1: Use the multimodal data processing and retrieval engine's multimodal data retrieval component to collect users' real-time query information; S3-2: Use the keyword extraction model of the multimodal data retrieval component to extract real-time keywords for real-time query information, including the following steps: S3-3-1: Use the word embedding module of the keyword extraction model of the multimodal data retrieval component to embed the real-time query information to obtain real-time word embedding query information; S3-3-2: Use the query semantic feature extraction module of the keyword extraction model to extract the real-time query semantic features of the real-time word embedding query information; S3-3-3: Based on the semantic features of the real-time query, use the keyword annotation module of the keyword extraction model to perform keyword annotation on the real-time word embedding query information to obtain the real-time keywords of the real-time query information; S3-3: Based on the real-time keywords, the retrieval graph matching model of the multimodal data retrieval component is used to perform retrieval matching in the graph database of the multimodal data storage network to obtain the target real-time retrieval graph, including the following steps: S3-3-1: Use the keyword semantic feature extraction module of the retrieval graph matching model of the multimodal data retrieval component to extract the real-time keyword semantic features of the real-time keywords; S3-3-2: Use the retrieval graph feature extraction module of the retrieval graph matching model to extract historical / real-time retrieval graph features of several historical / real-time retrieval graphs in the graph database of the multimodal data storage network; S3-3-3: Use the retrieval graph matching module of the retrieval graph matching model to obtain the real-time similarity between the real-time keyword semantic features and several historical / real-time retrieval graph features; S3-3-4: The historical / real-time search graph corresponding to the historical / real-time search graph feature with the highest real-time similarity is used as the target real-time search graph; S3-4: Retrieve the target storage address corresponding to the graph in real time according to the target stored in the graph database, and extract the target standard multimodal data in the distributed storage network of the multimodal data storage network.

[0028] Example 2: like Figure 2 As shown, this embodiment provides a multimodal data retrieval system for implementing a multimodal data retrieval method. The system includes an engine and storage network construction unit, a data processing and distributed storage unit, and a query matching and data retrieval unit connected in sequence; An engine and storage network construction unit, configured to use artificial intelligence algorithms to construct a multimodal data processing and retrieval engine, and to use blockchain technology to construct a multimodal data storage network, and to connect the multimodal data storage network to the multimodal data processing and retrieval engine; The data processing and distributed storage unit is used to collect real-time multimodal data and its real-time data information, process the real-time multimodal data using a multimodal data processing and retrieval engine based on the real-time data information, and use a multimodal data storage network to distributely store the obtained real-time standard multimodal data and its real-time retrieval graph; The query matching and data retrieval unit is used to collect the user's real-time query information, use the multimodal data processing and retrieval engine to match the target real-time retrieval graph in the multimodal data storage network based on the real-time query information, and extract the target standard multimodal data in the multimodal data storage network based on the target retrieval graph.

[0029] The present invention provides a multimodal data retrieval method and system, which utilizes a multimodal data processing and retrieval engine constructed using an artificial intelligence algorithm. This method can efficiently process and analyze massive amounts of multimodal data. Combined with a distributed storage network based on blockchain technology, this method can achieve rapid access and parallel processing of data, significantly improving retrieval speed and overcoming the efficiency bottleneck of traditional methods in processing large-scale data. Furthermore, this method can more effectively utilize computing resources, achieve rational allocation and efficient utilization of computing resources, and avoid resource waste. Using artificial intelligence algorithms, various features (such as images, text, audio, etc.) in multimodal data can be extracted and integrated more comprehensively and deeply to construct a more accurate real-time retrieval graph, thereby improving the accuracy and relevance of retrieval results. By utilizing the decentralized and tamper-proof characteristics of blockchain technology, a multimodal data storage network can be constructed, which can effectively prevent data tampering and malicious attacks and improve the security of data storage.

[0030] The present invention is not limited to the above optional embodiments. Anyone can derive various other forms of products based on the teachings of the present invention. The above specific embodiments should not be construed as limiting the scope of protection of the present invention. The scope of protection of the present invention shall be based on the scope defined in the claims, and the description can be used to interpret the claims.

Claims

1. A multimodal data retrieval method, characterized by: The steps include: Use artificial intelligence algorithms to build a multimodal data processing and retrieval engine, and use blockchain technology to build a multimodal data storage network, and connect the multimodal data storage network to the multimodal data processing and retrieval engine; Collect real-time multimodal data and its real-time data information, use the multimodal data processing and retrieval engine to process the real-time multimodal data based on the real-time data information, and use the multimodal data storage network to distribute the obtained real-time standard multimodal data and its real-time retrieval map; Collect users' real-time query information, use multimodal data processing and retrieval engines based on the real-time query information, match the target real-time retrieval graph in the multimodal data storage network, and extract target standard multimodal data in the multimodal data storage network based on the target retrieval graph.

2. A multimodal data retrieval method according to claim 1, characterized in that: The real-time multimodal data includes real-time image data, real-time text data and real-time audio data.

3. The multimodal data retrieval method according to claim 2, wherein: The multimodal data processing and retrieval engine includes a multimodal data processing component and a multimodal data retrieval component; The multimodal data processing component includes a multimodal data processing model, a multimodal feature fusion model and a retrieval graph generation model; The multimodal data retrieval component includes a keyword extraction model and a retrieval graph matching model; The multimodal data storage network includes a graph database and a distributed storage network.

4. The multimodal data retrieval method according to claim 3, wherein: The multimodal data processing model is constructed based on the MOGRPO-DPA-DALL-E-Wav2Vec-CCA-ASGA algorithm, and the multimodal data processing model includes a multimodal data preprocessing component constructed based on the MOGRPO-DPA algorithm and a multimodal data alignment component constructed based on the DALL-E-Wav2Vec-CCA-ASGA algorithm, which are connected in sequence. The multimodal data preprocessing component includes a preprocessing strategy generation module constructed based on the MOGRPO algorithm and a data preprocessing algorithm library constructed based on the DPA algorithm, which are connected in sequence. The multimodal data alignment component includes an image and text feature extraction module constructed based on the DALL-E algorithm, an audio feature extraction module constructed based on the Wav2Vec algorithm, a cross-modal association module constructed based on the CCA algorithm, and a multimodal data alignment module constructed based on the ASGA algorithm, which are connected in sequence. The multimodal feature fusion model is constructed based on FPN-LSTM-VGGish-MLP, and the multimodal feature fusion model includes a high-order image feature extraction module constructed based on the FPN algorithm, a deep text feature extraction module constructed based on the LSTM algorithm, a deep audio feature extraction module constructed based on the VGGish algorithm, and a multimodal feature fusion module constructed based on MLP, which are connected in sequence. The multimodal feature fusion module is respectively connected to the high-order image feature extraction module, the deep text feature extraction module, and the deep audio feature extraction module; The retrieval graph generation model is constructed based on the GCN-AGN algorithm, and the retrieval graph generation model includes a retrieval graph generation module and an attention weight module connected in sequence.

5. The multimodal data retrieval method according to claim 4, characterized in that: The keyword extraction model is constructed based on the BERT-BiLSTM-CRF algorithm, and the keyword extraction model includes a word embedding module constructed based on the BERT algorithm, a query semantic feature extraction module constructed based on the BiLSTM algorithm, and a keyword annotation module constructed based on the CRF algorithm, which are connected in sequence; The retrieval graph matching model is constructed based on the LSTM-GCN-k-NN algorithm, and the retrieval graph matching model includes a keyword semantic feature extraction module constructed based on the LSTM algorithm, a retrieval graph feature extraction module constructed based on the GCN algorithm, and a retrieval graph matching module constructed based on the k-NN algorithm. The retrieval graph matching module is respectively connected to the keyword semantic feature extraction module and the retrieval graph feature extraction module.

6. A multimodal data retrieval method according to claim 5, characterized in that: Collecting real-time multimodal data and its real-time data information, using a multimodal data processing and retrieval engine to process the real-time multimodal data based on the real-time data information, and using a multimodal data storage network to distributely store the obtained real-time standard multimodal data and its real-time retrieval graph, including the following steps: Use the multimodal data processing component of the multimodal data processing and retrieval engine to collect real-time multimodal data and its real-time data information; According to the real-time data information, the multimodal data processing model of the multimodal data processing component is used to process the real-time multimodal data to obtain real-time standard multimodal data; Use the multimodal feature fusion model to perform multimodal feature fusion on real-time standard multimodal data to obtain real-time multimodal fusion features; According to the real-time multimodal fusion features, a retrieval graph generation model is used to generate a retrieval graph to obtain a real-time retrieval graph; Using a distributed storage network of a multimodal data storage network, real-time standard multimodal data is distributedly stored and corresponding real-time storage addresses are obtained; A graph database of a multimodal data storage network is used to store real-time retrieval graphs and corresponding real-time storage addresses of real-time standard multimodal data.

7. The multimodal data retrieval method according to claim 6, characterized in that: According to the real-time data information, a multimodal data processing model is used to process the real-time multimodal data to obtain real-time standard multimodal data, including the following steps: According to the real-time data information, the multimodal data preprocessing component of the multimodal data processing model is used to preprocess the corresponding real-time multimodal data to obtain preprocessed real-time multimodal data; The multimodal data alignment component of the multimodal data processing model is used to perform multimodal data alignment on the preprocessed real-time multimodal data to obtain real-time standard multimodal data.

8. The multimodal data retrieval method according to claim 7, characterized in that: Using a multimodal feature fusion model, multimodal feature fusion is performed on real-time standard multimodal data to obtain real-time multimodal fusion features, including the following steps: Use the high-order image feature extraction module of the multimodal feature fusion model to extract real-time high-order image features of real-time standard multimodal data; Use the deep text feature extraction module of the multimodal feature fusion model to extract real-time deep text features from real-time standard multimodal data; A deep audio feature extraction module using a multimodal feature fusion model is used to extract real-time deep audio features from real-time standard multimodal data. The multimodal feature fusion module of the multimodal feature fusion model is used to fuse real-time high-order image features, real-time deep text features, and real-time deep audio features to obtain real-time multimodal fusion features.

9. The multimodal data retrieval method according to claim 8, characterized in that: Collecting real-time query information from users, using a multimodal data processing and retrieval engine based on the real-time query information, matching a target real-time retrieval graph in a multimodal data storage network, and extracting target standard multimodal data from the multimodal data storage network based on the target retrieval graph, including the following steps: Use the multimodal data processing and retrieval engine's multimodal data retrieval component to collect users' real-time query information; Use the keyword extraction model of the multimodal data retrieval component to extract real-time keywords for real-time query information; Based on real-time keywords, the retrieval graph matching model of the multimodal data retrieval component is used to perform retrieval and matching in the graph database of the multimodal data storage network to obtain the target real-time retrieval graph; According to the target stored in the graph database, the target storage address corresponding to the graph is retrieved in real time, and the target standard multimodal data is extracted in the distributed storage network of the multimodal data storage network.

10. A multimodal data retrieval system, configured to implement the multimodal data retrieval method according to any one of claims 1 to 9, characterized in that: The system comprises an engine and storage network construction unit, a data processing and distributed storage unit, and a query matching and data retrieval unit which are connected in sequence.