Multi-modal data extraction method based on deep learning network

Through the multimodal data extraction method based on deep learning network, the multimodal R&D experimental data management and query problems are solved, efficient classified storage and precise extraction are achieved, and the efficiency of scientific research is improved.

CN120145124APending Publication Date: 2025-06-13CHONGQING ACADEMY OF SCI & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510374940.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-27
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

The existing technology is difficult to effectively manage and utilize multimodal R&D experimental data, resulting in messy data, inefficient search efficiency, and difficult to meet complex query needs.

Method used

A multimodal data extraction method based on deep learning network is adopted to build a scientific research data set with clear structure through standardized processing, deep feature extraction, dynamic fusion and precise classification, and extract corresponding data based on user query needs.

Benefits of technology

It realizes efficient classification storage and precise extraction of R&D experimental data, improves the efficiency of data management and query, and meets the complex query needs of users.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120145124A_ABST
    Figure CN120145124A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-modal data extraction method based on a deep learning network. The method comprises the steps that S1, a multi-modal original data set is acquired and standardized; s2, performing feature extraction on each modal data in the standardized multi-modal data to obtain a depth feature of each modal; s3, performing dynamic fusion on the depth features of each modal to obtain a joint feature; s4, inputting the joint features of the standardized multi-modal data into the trained classification model to obtain corresponding scientific research data labels; s5, obtaining a scientific research data label of each piece of multi-modal data; constructing data samples based on the scientific research data labels and the corresponding multi-modal data to obtain a scientific research data set containing a plurality of data samples; and S6, matching the corresponding target data sample from the scientific research data set based on the query demand of the user, and extracting the target data from the matched target data sample. According to the invention, efficient classified storage and accurate extraction of research and development experiment data can be realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of big data and artificial intelligence, and particularly relates to a multi-modal data extraction method based on a deep learning network. Background Art

[0002] In the field of scientific and technological innovation research and development, with the rapid development of information technology and the increasing diversification of scientific research means, the scale and complexity of research and development experimental data have increased explosively. The research and development experimental data are no longer limited to a single text form, but cover multiple modalities such as images and videos. These multi-modal data record the whole process of research and development experiments and contain rich information, such as experimental scenarios, experimental conditions, experimental objects, scientific researchers, experimental time, experimental equipment, and reference documents, etc., which are of crucial significance for the promotion of scientific research work, the summary of achievements, and the inheritance of knowledge.

[0003] However, the management and utilization of current research and development experimental data face many challenges. On the one hand, due to the lack of an effective classification and storage method, a large amount of multi-modal research and development experimental data are often in a chaotic state. The data generated by different scientific research projects and different experimental stages are mixed with each other, making it difficult to carry out systematic sorting and classification. This not only causes scientific researchers to spend a lot of time and energy in searching for and using data, but also easily leads to the loss and damage of scientific research data, affecting the efficiency and quality of scientific research work. For example, when conducting a scientific research project involving multiple experimental stages and various experimental equipment, due to the lack of clear classification and storage, scientific researchers may need to spend several hours or even several days to search for relevant data of a specific experimental stage, which will undoubtedly seriously delay the scientific research progress.

[0004] On the other hand, users often need to quickly and accurately extract the required scientific research data according to their own query needs during the scientific research process. However, the existing data extraction methods mostly rely on manual screening and simple keyword searches, which are difficult to meet the complex query needs of users for multi-modal research and development experimental data. Manual screening is not only inefficient but also prone to omissions and errors; simple keyword searches cannot fully understand the associations and semantic information between multi-modal data, resulting in inaccurate and incomplete search results. For example, when a user wants to query the usage of a specific experimental equipment under different experimental conditions in a certain scientific research project, the existing search methods may not be able to accurately match the relevant image and video data, making it impossible for the user to obtain complete information.

[0005] In addition, the multi-modal R & D experimental data has characteristics such as heterogeneity, high dimensionality, and complexity. There are significant differences in format, structure, and semantics among data of different modalities. This makes it difficult to directly apply traditional data processing methods to the processing and analysis of multi-modal R & D experimental data. For example, text data usually exists in the form of character sequences, while image and video data have characteristics of spatial and temporal dimensions. Therefore, how to effectively fuse these different modalities of data and extract valuable information from them is a technical problem that urgently needs to be solved. Summary of the Invention

[0006] Aiming at the deficiencies of the above-mentioned prior art, the technical problem to be solved by the present invention is: how to provide a multi-modal data extraction method based on a deep learning network, which constructs a clearly structured and easy-to-query scientific research data set through standardizing multi-modal (R & D experimental) data, extracting deep features, dynamically fusing, and accurately classifying, and can extract corresponding scientific research data based on the user's query requirements, so as to realize the efficient classification storage and accurate extraction of R & D experimental data and provide strong support for the smooth development of scientific research work.

[0007] To solve the above technical problems, the present invention adopts the following technical solutions:

[0008] A multi-modal data extraction method based on a deep learning network, comprising:

[0009] S1: Obtain a multi-modal original data set and standardize the multi-modal data therein to obtain standardized multi-modal data;

[0010] S2: Based on the deep learning network, extract features from each modality data in the standardized multi-modal data to obtain deep features of each modality;

[0011] S3: Dynamically fuse the deep features of each modality in the standardized multi-modal data to obtain joint features;

[0012] S4: Input the joint features of the standardized multi-modal data into a trained classification model to obtain corresponding scientific research data labels;

[0013] S5: Repeat steps S1 to S4 to obtain scientific research data labels for each multi-modal data; construct data samples based on the scientific research data labels and the corresponding multi-modal data to obtain a scientific research data set containing several data samples;

[0014] S6: Match corresponding target data samples from the scientific research data set based on the user's query requirements, and extract target data from the matched target data samples.

[0015] Preferably, in step S1, the standardization processing steps include:

[0016] S101: The multimodal data includes text data, image data, and video data;

[0017] S102: Clean and standardize the text data to obtain standardized text data;

[0018] S103: Uniform the size, correct the color, and denoise the image data to obtain standardized image data;

[0019] S104: Split the video images and preprocess the frame images to obtain standardized video frame image data;

[0020] S105: Use the standardized text data, standardized image data, and standardized video frame image data as the standardized multimodal data of the multimodal data.

[0021] Preferably, in step S2, the processing steps of feature extraction include:

[0022] S201: The standardized multimodal data includes standardized text data, standardized image data, and standardized video frame image data;

[0023] S202: Input the standardized text data into the trained BERT model for deep encoding to obtain the deep features of the text data;

[0024] S203: Input the standardized image data into the trained convolutional neural network model to obtain the deep features of the image data;

[0025] S204: Input the standardized video frame image data into the trained convolutional neural network model to obtain the deep features of the video frame image data;

[0026] S205: The deep features of each modality include the deep features of text data, the deep features of image data, and the deep features of video frame image data.

[0027] Preferably, in step S3, the processing steps of feature fusion include:

[0028] S301: The deep features of each modality in the standardized multimodal data include the deep features f 1 of text data, the deep features f 2 of image data, and the deep features f 3 of video frame image data;

[0029] S302: Use a convolutional neural network to perform preliminary feature transformation on the deep features of each modality to obtain the transformed deep features {f 1 ′, f 2 ′, f 3 ′} of each modality;

[0030] S303: Calculate the internal attention weights of the depth features of each modality after transformation to obtain the first-level attention weights; perform weighted fusion on the depth features of each modality after transformation according to the first-level attention weights to obtain the internal fusion features of each modality;

[0031] S304: Calculate the inter-modal attention weights between the internal fusion features of each modality to obtain the second-level attention weights; perform weighted fusion on the internal fusion features of each modality according to the second-level attention weights to obtain the joint features.

[0032] Preferably, in step S303, the internal fusion features of each modality are calculated through the following steps:

[0033] S3031: For the depth feature vector f′ of the m-th modality m , divide it into N sub-regions {f′ m1 , f′ m2 ,..., f′ mN};

[0034] S3032: Calculate the attention weights of each sub-region;

[0035] The formula is expressed as:

[0036]

[0037] e mn =(W mn f′ mn +b mn ) T u m ;

[0038] In the formula: α mn represents the internal attention weight of the n-th sub-region of the m-th modality, and the internal attention weights of all sub-regions are the first-level attention weights; e mn represents the correlation score of the n-th sub-region of the m-th modality; W mn represents the weight matrix specific to the n-th sub-region of the m-th modality, which is used to map the sub-region f′ mn to the attention space; b mn represents the bias vector; u m represents the modality-specific attention vector;

[0039] S3033: Calculate the internal fusion features of the feature vector f′ m of the m-th modality according to the internal attention weights of each sub-region of the m-th modality

[0040] The formula is expressed as:

[0041]

[0042] Preferably, in step S304, the combined feature is calculated through the following steps:

[0043] S3041: Take the internal fusion feature of each modality as a new feature vector;

[0044] S3042: Calculate the inter-modal attention weights between the internal fusion features of each modality;

[0045] The formula is expressed as:

[0046]

[0047] In the formula: β m represents the inter-modal attention weight of the m-th modality; v m represents the internal fusion feature of the m-th modality;

[0048] S3043: Calculate the combined feature F according to the inter-modal attention weights;

[0049] The formula is expressed as:

[0050]

[0051] Preferably, in step S4, the data sample is constructed through the following steps:

[0052] S401: Assign a unique number ID to each multi-modal data;

[0053] S402: Construct a mapping table M map , which is used to record the number ID of each multi-modal data, its corresponding scientific research data label y i and the storage location of each modality data in the multi-modal original dataset;

[0054] The mapping table M map is expressed by the formula:

[0055] M map = {(id 1 , y 1 , loc1 t , loc1 i , loc1 v ), (id 2 , y 2 , loc2 t , loc2 i , loc2 v ),...};

[0056] In the formula: id1 Represents the serial number ID of the first multimodal data; y 1 Represents the scientific research data label of the first multimodal data; loc1 t , loc1 i , loc1 v Represents the storage locations of the text data, image data, and video data in the first multimodal data within the multimodal original dataset;

[0057] S403: Traverse the mapping table M map : For the i-th multimodal data, determine whether its scientific research data label y i belongs to the preset target scientific research data label, i.e., C target , if so, then according to the mapping table M map extract the corresponding text data te i , image data im i and video data ve i from the multimodal original dataset to construct the corresponding data sample (id i , y i , te i , im i , ve i ); otherwise, do not construct a data sample.

[0058] Preferably, in step S4, the scientific research dataset is constructed through the following steps:

[0059] S411: Store all the constructed data samples in the temporary data set D temp ;

[0060] The temporary data set D temp is represented by the formula:

[0061] D temp = {(id 1 , y 1 , te 1 , im 1 , ve 1 ), (id 2 , y 2 , te 2 , im 2 , ve 2 ),...}

[0062] In the formula: id 1 , y 1 , te 1 , im 1 , ve 1 respectively represent the multimodal data serial number ID, scientific research data label, text data, image data, and video data of the first data sample;

[0063] S412: Enhance the text data, image data, and video data in the temporary data set D temp to obtain an enhanced temporary data set D temp,avg ;

[0064] The enhanced temporary data set D temp,avg is represented by the formula:

[0065] D temp,avg ={(id 1 , y 1 , te 1,avg , im 1,avg , ve 1,avg ), (id 2 , y 2 , te 2,avg , im 2,avg , ve 2,avg ),...};

[0066] In the formula: te 1,avg , im 1,avg , ve 1,avg represent the enhanced text data, enhanced image data, and enhanced video data of the first data sample, respectively;

[0067] S413: Extract the text structured information, image features, and video features of each data sample from the enhanced temporary data set D temp,avg and integrate the extracted text structured information, image features, and video features to form the structured information of each data sample, obtaining the final scientific research data set D struct ;

[0068] The scientific research data set D struct is represented by the formula:

[0069] D struct ={(id 1 , y 1 , S 1 ), (id 2 , y 2 , S 2 ),...};

[0070] S 1 =(f t , f i , f v );

[0071] In the formula: S 1 represents the structured information of the first data sample; f t , f i , fv respectively represent the text structured information, image features, and video features.

[0072] Preferably, in step S6, the target data is extracted through the following steps:

[0073] S601: Obtain the user's query requirement Q raw and parse the query requirement Q raw to obtain the parsed query vector Q vec and the entity-relationship structure S QR ; where S QR contains all the entities and their relationships identified in the query requirement;

[0074] S602: Extract features from the structured information of each data sample in the scientific research dataset D struct to obtain the feature vector of each data sample and the corresponding set of numbers ID;

[0075] The feature vector set is obtained by concatenating the text feature vector and the multi-modal feature vector;

[0076] The text feature vector is obtained by converting the text structured information, i.e., each group of entity-relationship-entity triples, into a feature vector;

[0077] The multi-modal feature vector f m is obtained by fusing the image feature f i , the video key frame features {f f1 , f f2 ,...} and the video action feature f a using the attention mechanism; let the attention weights be a i , a f and a a , then the fused multi-modal feature vector f m can be expressed as:

[0078]

[0079] where the attention weights are calculated by a neural network based on the feature vectors:

[0080]

[0081] In the formula: w i , w f , w a represent trainable weight vectors;

[0082] S603: Calculate the query vector Q vecThe cosine similarity with the feature vector of each data sample, and put the ID numbers of the data samples whose cosine similarity exceeds the set cosine similarity threshold into the identification set of the preliminary matching samples;

[0083] S604: For each data sample in the identification set of the preliminary matching samples, extract its corresponding text structured information from the scientific research data set D struct to construct the entity-relationship structure S sample of the sample;

[0084] S605: Calculate the structural similarity sim QR between the entity-relationship structure S sample and the entity-relationship structure S struct , and retain the data samples whose structural similarity sim struct is greater than the set structural similarity threshold to obtain the matching sample identification set ID final ;

[0085] S606: Extract the corresponding target data samples from the scientific research data set D final according to the matching sample identification set ID struct ; Integrate the text structured information, image features, and video features of all target data samples into a structured matching result set R as the target data (the matching result set R needs to be converted into a format that is easy for users to understand).

[0086] Preferably, in step S605, the structural similarity sim struct is calculated through the following steps:

[0087]

[0088] In the formula: N represents the normalization constant; d represents the graph edit distance between the entity-relationship structure S QR and the entity-relationship structure S sample .

[0089] Compared with the prior art, the multi-modal data extraction method based on the deep learning network in the present invention has the following beneficial effects:

[0090] First, the present invention performs standardization processing on the obtained multi-modal original data set, converting each group of multi-modal data containing text data, image data, and video data during a certain period in the experimental process of scientific research projects into standardized multi-modal data, effectively eliminating the differences in format, scale, encoding method, etc. of different modal data, enabling the subsequent feature extraction and fusion processes to be carried out on a unified data benchmark. This not only improves the efficiency and accuracy of data processing but also avoids feature extraction biases caused by inconsistent data formats, laying a solid foundation for subsequent deep feature extraction and ensuring the improvement of the reliability and stability of data processing from the source.

[0091] Secondly, the present invention extracts features from each modal data in the standardized multi-modal data based on a deep learning network to obtain the deep features of each modality. The deep learning network has powerful non-linear fitting ability and automatic feature learning ability, and can mine high-level and abstract feature information hidden in complex text, image, and video data. These deep features comprehensively and accurately reflect the essential features of each modal data, providing rich and effective information for subsequent data fusion and classification.

[0092] Then, the present invention dynamically fuses the deep features of each modality in the standardized multi-modal data to obtain joint features. The dynamic fusion mechanism fully considers the correlation and complementarity between different modal data, adaptively adjusts the weights and fusion methods of each modal feature according to the characteristics of the data and actual needs. Through this fusion method, the advantages of each modal data can be fully utilized to make up for the deficiencies of single-modal data, so that the joint features contain more comprehensive and rich R & D experiment information.

[0093] Furthermore, the present invention inputs the joint features of the standardized multi-modal data into a trained classification model to obtain corresponding scientific research data labels, and then constructs data samples based on the scientific research data labels and the corresponding multi-modal data to obtain a scientific research data set containing several data samples. The construction method of the scientific research data set not only integrates multi-modal R & D experiment data but also systematically organizes and annotates it through labels, facilitating subsequent data management and query.

[0094] Finally, the present invention matches corresponding target data samples from the scientific research dataset based on the user's query requirements, and extracts target data from the matched target data samples. Since the scientific research dataset has undergone the previous standardization, feature extraction, fusion, and classification processes, the data samples have accurate labels and rich feature information. Therefore, it is possible to quickly and accurately match the target data samples that meet the user's query requirements, and the target data extracted from the target data samples has a high degree of accuracy and relevance, which can meet the personalized needs of users for R & D experimental data, improve the efficiency and accuracy of data acquisition, and provide strong support for scientific research decision-making, experimental analysis and other work. BRIEF DESCRIPTION OF THE DRAWINGS

[0095] In order to make the objectives, technical solutions, and advantages of the invention more clear, the present invention will be further described in detail below with reference to the accompanying drawings, where:

[0096] Figure 1 It is a logic block diagram of a multi-modal data extraction method based on a deep learning network. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0097] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Components of the embodiments of the present invention usually described and illustrated in the drawings here can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present invention provided in the drawings is not intended to limit the scope of the claimed invention, but merely represents selected embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.

[0098] It should be noted that similar reference numerals and letters denote similar items in the following figures. Therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures. In the description of the present invention, it should be noted that the orientation or positional relationship indicated by terms such as "center", "upper", "lower", "left", "right", "vertical", "horizontal", "inner", "outer", etc. is based on the orientation or positional relationship shown in the figures, or the orientation or positional relationship in which the inventive product is customarily placed during use. It is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be construed as a limitation of the present invention. In addition, terms such as "first", "second", "third", etc. are only used for distinguishing descriptions and cannot be understood as indicating or implying relative importance. In addition, terms such as "horizontal" and "vertical" do not mean that the components are required to be absolutely horizontal or hanging, but can be slightly inclined. For example, "horizontal" only means that its direction is more horizontal relative to "vertical", and does not mean that the structure must be completely horizontal, but can be slightly inclined. In the description of the present invention, it should also be noted that unless otherwise clearly specified and limited, the terms "set", "installed", "connected", "connected" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be directly connected, or indirectly connected through an intermediate medium, and can be the communication inside two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific situations.

[0099] The following is a more detailed description through specific embodiments:

[0100] Embodiment:

[0101] In this embodiment, a multi-modal data extraction method based on a deep learning network is disclosed.

[0102] As Figure 1 shown, a multi-modal data extraction method based on a deep learning network includes:

[0103] S1: Obtain a multi-modal original data set and standardize the multi-modal data therein to obtain standardized multi-modal data;

[0104] In this embodiment, the multi-modal original data set includes several groups of multi-modal data, and each group of multi-modal data is text data, image data, and video data during a certain period in the experimental process of a scientific research project.

[0105] S2: Based on the deep learning network, extract features from each modal data in the standardized multi-modal data to obtain the deep features of each modality;

[0106] S3: Dynamically fuse the deep features of each modality in the standardized multi-modal data to obtain joint features;

[0107] S4: Input the joint features of the standardized multi-modal data into the trained classification model to obtain the corresponding scientific research data labels;

[0108] In this embodiment, a convolutional neural network model, a support vector machine, a decision tree model, or a random forest model is selected as the classification model;

[0109] S5: Repeat steps S1 to S4 to obtain the scientific research data labels of each multi-modal data; construct data samples based on the scientific research data labels and the corresponding multi-modal data to obtain a scientific research data set containing several data samples;

[0110] In this embodiment, the scientific research data labels include labels such as the R & D experiment scenario, R & D experiment conditions, R & D experiment objects, scientific research personnel, R & D experiment time, R & D experiment equipment, and scientific research references of the scientific research project. Among them, the scientific research data labels of each data sample are associated with the corresponding scientific research project. For example, the scientific research data label of a certain data sample is: the R & D experiment scenario of the XX scientific research project, or the R & D experiment time of the XX scientific research project. That is, the format of the scientific research data label is XXXX of the XX scientific research project.

[0111] S6: Match the corresponding target data samples from the scientific research data set based on the user's query requirements, and extract the target data from the matched target data samples.

[0112] In this embodiment, the user's query requirements include query requirements for various information in the scientific research scenario. For example: query the scientific research personnel, R & D experiment time, and R & D experiment equipment of the XX scientific research project.

[0113] First of all, the present invention performs a standardization process on the obtained multi-modal original data set, and converts each group of multi-modal data containing text data, image data, and video data during a certain period in the scientific research project experiment process into standardized multi-modal data, effectively eliminating the differences in format, scale, coding method, etc. of different modal data, enabling the subsequent feature extraction and fusion processes to be carried out on a unified data basis, not only improving the efficiency and accuracy of data processing, but also avoiding feature extraction deviations caused by inconsistent data formats, laying a solid foundation for subsequent deep feature extraction, and ensuring the reliability and stability of data processing from the source.

[0114] Secondly, the present invention extracts features from each modality data in the standardized multi-modal data based on a deep learning network to obtain the deep features of each modality. The deep learning network has a powerful non-linear fitting ability and automatic feature learning ability, and can mine high-level and abstract feature information hidden in complex text, image, and video data. These deep features comprehensively and accurately reflect the essential features of each modality data, providing rich and effective information for subsequent data fusion and classification.

[0115] Then, the present invention dynamically fuses the deep features of each modality in the standardized multi-modal data to obtain joint features. The dynamic fusion mechanism fully considers the relevance and complementarity between different modality data, adaptively adjusts the weights and fusion methods of each modality feature according to the characteristics of the data and actual requirements. Through this fusion method, the advantages of each modality data can be fully exerted, and the deficiencies of single-modal data can be made up, so that the joint features contain more comprehensive and rich R & D experiment information.

[0116] Furthermore, the present invention inputs the joint features of the standardized multi-modal data into a trained classification model to obtain corresponding scientific research data labels, and then constructs data samples based on the scientific research data labels and the corresponding multi-modal data to obtain a scientific research data set containing several data samples. The construction method of the scientific research data set not only integrates multi-modal R & D experiment data, but also systematically organizes and annotates it through labels, facilitating subsequent data management and query.

[0117] Finally, the present invention matches the corresponding target data samples from the scientific research data set based on the user's query requirements, and extracts target data from the matched target data samples. Since the scientific research data set has undergone the previous standardization, feature extraction, fusion, and classification processing, the data samples have accurate labels and rich feature information. Therefore, it can quickly and accurately match the target data samples that meet the user's query requirements, and the target data extracted from the target data samples has high accuracy and relevance, which can meet the personalized needs of users for R & D experiment data, improve the efficiency and accuracy of data acquisition, and provide strong support for scientific research decision-making, experimental analysis and other work.

[0118] To better introduce the technical solution of the present invention, this embodiment is described through the following several parts.

[0119] I. Data Standardization

[0120] The processing steps of standardization include:

[0121] S101: Collect multi-modal data during the experiment of a scientific research project to construct a multi-modal original data set; the multi-modal data includes text data, image data, and video data;

[0122] In this embodiment, the text data sources include PDF / Word documents, scanned handwritten experimental records, and web texts. The image data sources include photos of experimental equipment, observation charts, and microscope imaging. The video data sources include experimental operation videos and monitoring video streams.

[0123] S102: Perform data cleaning and data standardization on the text data to obtain standardized text data;

[0124] In this embodiment, data cleaning includes processing such as removing non-standard characters, word segmentation, and removing stop words. Data standardization refers to unifying the text format, such as converting all text to lowercase and removing extra spaces.

[0125] S103: Perform size unification, color correction, and denoising processing on the image data to obtain standardized image data;

[0126] In this embodiment, size unification means using an image scaling algorithm (such as bilinear interpolation, nearest neighbor interpolation, etc.) to adjust the image to a unified size. Color correction means adjusting the image color according to a color correction algorithm to make it more realistic. Denoising processing means applying a denoising algorithm to remove noise in the image and improve the image quality.

[0127] S104: Perform frame splitting and preprocessing on the frame images of the video to obtain standardized video frame image data;

[0128] In this embodiment, the preprocessing of the frame images refers to performing processing such as size unification, color correction, and denoising processing on the frame images obtained by frame splitting.

[0129] S105: Use the standardized text data, standardized image data, and standardized video frame image data as the standardized multi-modal data of the multi-modal data.

[0130] II. Feature Extraction

[0131] The processing steps of feature extraction include:

[0132] S201: The standardized multi-modal data includes standardized text data, standardized image data, and standardized video frame image data;

[0133] S202: Input the standardized text data into the trained BERT (Bidirectional Encoder Representations from Transformers) model, and perform deep encoding on the standardized text data through a multi-layer Transformer structure to obtain the deep features of the text data;

[0134] S203: Input the standardized image data into the trained convolutional neural network model, extract features from the image data through the convolutional layer, pooling layer, and fully connected layer (to obtain a high-level feature map), and use the output before the last fully connected layer as the deep feature of the image data;

[0135] S204: Input the standardized video frame image data into the trained convolutional neural network model, extract features from the image data through the convolutional layer, pooling layer, and fully connected layer (to obtain a high-level feature map), and use the output before the last fully connected layer as the deep feature of the video frame image data;

[0136] S205: The deep features of each modality include the deep feature of text data, the deep feature of image data, and the deep feature of video frame image data.

[0137] III. Feature Fusion

[0138] The processing steps of feature fusion include:

[0139] S301: The deep features of each modality in the standardized multi-modal data include the deep feature f 1 of text data, the deep feature f 2 of image data, and the deep feature f 3 of video frame image data;

[0140] S302: Standardize the deep features {f 1 , f 2 , f 3} of each modality to eliminate the dimensional difference; use a convolutional neural network to perform a preliminary feature transformation on the deep feature of each modality to enhance its feature expression ability, and obtain the transformed deep features {f 1 ′, f 2 ′, f 3 ′} of each modality;

[0141] S303: Calculate the internal attention weight of the deep feature of each modality after transformation to obtain the first-level attention weight; perform weighted fusion on the deep feature of each modality after transformation according to the first-level attention weight to obtain the internal fusion feature of each modality;

[0142] In step S303, calculate the internal fusion feature of each modality through the following steps:

[0143] S3031: For the deep feature vector f′ m of the m-th modality, divide it into N sub-regions {f′ m1 , f′ m2 ,..., f′ mN};

[0144] In this embodiment, sub-regions are obtained by means of spatial division or feature dimension division.

[0145] S3032: Calculate the attention weights of each sub-region;

[0146] The formula is expressed as:

[0147]

[0148] e mn =(W mn f′ mn +b mn ) T u m ;

[0149] In the formula: α mn represents the internal attention weight of the nth sub-region of the mth modality, and the internal attention weights of all sub-regions are the first-level attention weights; e mn represents the correlation score of the nth sub-region of the mth modality; W mn represents the weight matrix specific to the nth sub-region of the mth modality, which is used to map the sub-region f m ′ n to the attention space; b mn represents the bias vector, which is used to adjust the mapped feature vector; u m represents the modality-specific attention vector, which is used to perform a dot product operation with the mapped and adjusted feature vector to obtain the correlation score;

[0150] S3033: Calculate the internal fusion feature of the feature vector f m ′ of the mth modality according to the internal attention weights of each sub-region of the mth modality

[0151] The formula is expressed as:

[0152]

[0153] S304: Calculate the inter-modal attention weights between the internal fusion features of each modality to obtain the second-level attention weights; perform weighted fusion on the internal fusion features of each modality according to the second-level attention weights to obtain the joint feature;

[0154] In this embodiment, it is also necessary to further activate the joint feature using a non-linear activation function to enhance its non-linear expression ability; perform normalization processing on the activated joint feature through batch normalization or layer normalization to improve the stability and convergence speed of the model.

[0155] In step S304, the joint feature is calculated through the following steps:

[0156] S3041: Take the internal fusion feature of each modality as a new feature vector;

[0157] S3042: Calculate the inter-modal attention weights between the internal fusion features of each modality;

[0158] The formula is expressed as:

[0159]

[0160] In the formula: β m represents the inter-modal attention weight of the m-th modality; v m represents the internal fusion feature of the m-th modality;

[0161] S3043: Calculate the joint feature F according to the inter-modal attention weights;

[0162] The formula is expressed as:

[0163]

[0164] IV. Construct data samples

[0165] Construct data samples through the following steps:

[0166] S401: Assign a unique number ID to each multi-modal data;

[0167] S402: Construct a mapping table M map , which is used to record the number ID of each multi-modal data, its corresponding scientific research data label y i and the storage location of each modality data in the multi-modal original dataset;

[0168] The mapping table M map is expressed by the formula:

[0169] M map = {(id 1 , y 1 , loc1 t , loc1 i , loc1 v ), (id 2 , y 2 , loc2 t , loc2 i , loc2 v ),...};

[0170] In the formula: id 1 represents the number ID of the first multi-modal data; y 1Research data label indicating the 1st multimodal data; loc1 t ,loc1 i ,loc1 v Indicates the storage locations of the text data, image data, and video data in the 1st multimodal data within the multimodal original dataset;

[0171] S403: Traverse the mapping table M map : For the i-th multimodal data, determine whether its research data label y i belongs to the preset target research data label, i.e., C target , if so, then according to the mapping table M map extract the corresponding text data te i , image data im i and video data ve i from the multimodal original dataset to construct the corresponding data sample (id i ,y i ,te i ,im i ,ve i ); otherwise, do not construct a data sample.

[0172] Construct a research dataset through the following steps:

[0173] S411: Store all the constructed data samples in the temporary data set D temp ;

[0174] The formula for the temporary data set D temp is expressed as:

[0175] D temp ={(id 1 ,y 1 ,te 1 ,im 1 ,ve 1 ),(id 2 ,y 2 ,te 2 ,im 2 ,ve 2 ),...}

[0176] In the formula: id 1 ,y 1 ,te 1 ,im 1 ,ve 1 respectively represent the multimodal data number ID, research data label, text data, image data, and video data of the 1st data sample;

[0177] S412: For the temporary data set D tempEnhance the text data, image data, and video data therein to obtain the enhanced temporary data set D temp,avg ;

[0178] In this embodiment, the text data is enhanced by means such as synonym replacement, random insertion, and random deletion; the image data is enhanced by using various image enhancement techniques, such as rotation, flipping, scaling, adding noise, etc.; the video data is enhanced by operations such as frame sampling, frame rate adjustment, and adding video effects (such as blurring, sharpening). Update the temporary data set D with the enhanced text data, image data, and video data temp , to obtain the enhanced temporary data set D temp,avg ;

[0179] The enhanced temporary data set D temp,avg is represented by the formula:

[0180] D temp,avg ={(id 1 , y 1 , te 1,avg , im 1,avg , ve 1,avg ), (id 2 , y 2 , te 2,avg , im 2,avg , ve 2,avg ),...};

[0181] In the formula: te 1,avg , im 1,avg , ve 1,avg respectively represent the enhanced text data, enhanced image data, and enhanced video data of the first data sample;

[0182] S413: Extract the text structured information, image features, and video features of each data sample from the enhanced temporary data set D temp,avg , and integrate the extracted text structured information, image features, and video features to form the structured information of each data sample, obtaining the final scientific research data set D struct ;

[0183] The scientific research data set D struct is represented by the formula:

[0184] D struct ={(id 1 , y 1 , S 1 ), (id 2 , y 2 , S 2 ),...};

[0185] S1 = (f t , f i , f v );

[0186] Where: S 1 represents the structured information of the first data sample; f t , f i , f v represent the text structured information, image features, and video features respectively;

[0187] In this embodiment, the text structured information is several groups of entity-relationship-entity triples that enhance the text data; the image features are the image features obtained by extracting features from the enhanced image data through a convolutional neural network; the video features include video key frame features and video action features, both of which are obtained by existing means.

[0188] V. User Requirement Matching

[0189] Extract the target data through the following steps:

[0190] S601: Obtain the user's query requirement Q raw and parse the query requirement Q raw to obtain the parsed query vector Q vec and the entity-relationship structure S QR ; where S QR contains all the entities and their relationships identified in the query requirement;

[0191] In this embodiment, use natural language processing tools to perform word segmentation and part-of-speech tagging on the user query Q raw to split the query into individual words and label the part of speech of each word (noun, verb, adjective, etc.). Then, use named entity recognition technology to identify the scientific research entities in the query. Finally, based on the identified entities and relationships, construct a query vector Q vec . Assuming that n entities and m relationships are identified, the query vector can be expressed as Q vec = [e 1 , e 2 , …, e 3 , r 1 , r 2 , …, r m , where e i represents the feature representation of the i-th entity, and r j represents the feature representation of the j-th relationship.

[0192] S602: Extract features from the structured information of each data sample in the scientific research dataset D struct to obtain the feature vector of each data sample and the corresponding set of numbered IDs;

[0193] In this embodiment, the feature vector set is obtained by concatenating the text feature vector and the multi-modal feature vector.

[0194] The text feature vector is obtained by converting the text structured information, i.e., each group of entity-relationship-entity triples, into a feature vector. The word vectors of the entities and relationships in each triple are averaged using the method of averaging word vectors to obtain the feature vector of the triple. Then, the feature vectors of all triples are concatenated to obtain the text feature vector f of the data sample. t 。

[0195] The multi-modal feature vector f m is obtained by fusing the image feature f i , the video key frame features {f f1 , f f2 ,...} and the video action feature f a using the attention mechanism. Let the attention weights be a i , a f and a a , respectively. Then, the fused multi-modal feature vector f m can be expressed as:

[0196]

[0197] Among them, the attention weights are calculated by a neural network based on the feature vectors:

[0198]

[0199] In the formula: w i , w f , w a represent trainable weight vectors.

[0200] S603: Calculate the cosine similarity between the query vector Q vec and the feature vector of each data sample, and put the ID numbers of the data samples whose cosine similarity exceeds the set cosine similarity threshold into the identification set of the preliminary matching samples;

[0201] S604: For each data sample in the identification set of the preliminary matching samples, extract its corresponding text structured information from the scientific research data set D struct and construct the entity-relationship structure S sample of the sample;

[0202] S605: Calculate the structural similarity sim QR between the entity-relationship structure S sample and the entity-relationship structure S struct , and use the structural similarity simstruct Data samples greater than the set structural similarity threshold are retained to obtain a set of matching sample identifiers ID final ;

[0203] In step S605, the structural similarity sim is calculated through the following steps struct :

[0204]

[0205] where: N represents the normalization constant; d represents the graph edit distance between the entity-relationship structure S QR and the entity-relationship structure S sample of the graph edit distance.

[0206] S606: According to the set of matching sample identifiers ID final extract the corresponding target data samples from the scientific research dataset D struct ; integrate the text structured information, image features, and video features of all target data samples into a structured matching result set R as the target data (the matching result set R needs to be converted into a format that is easy for users to understand).

[0207] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit the technical solutions. Those of ordinary skill in the art should understand that any modifications or equivalent replacements made to the technical solutions of the present invention without departing from the purpose and scope of the present technical solution shall be covered by the scope of the claims of the present invention.

Claims

1. A multimodal data extraction method based on deep learning network, characterized in that: include: S1: obtaining a multimodal original data set and standardizing the multimodal data therein to obtain standardized multimodal data; S2: Extract features of each modality in the standardized multimodal data based on a deep learning network to obtain deep features of each modality; S3: Dynamically fuse the deep features of each modality in the standardized multimodal data to obtain joint features; S4: Input the joint features of the standardized multimodal data into the trained classification model to obtain the corresponding scientific research data labels; S5: Repeat steps S1 to S4 to obtain a scientific research data label for each multimodal data; construct a data sample based on the scientific research data label and the corresponding multimodal data to obtain a scientific research data set containing a plurality of data samples; S6: Match the corresponding target data samples from the scientific research data set based on the user's query requirements, and extract the target data from the matched target data samples.

2. The multimodal data extraction method based on deep learning network according to claim 1, characterized in that: In step S1, the standardized processing steps include: S101: The multimodal data includes text data, image data and video data; S102: Perform data cleaning and data standardization on the text data to obtain standardized text data; S103: performing size unification, color correction and noise removal on the image data to obtain standardized image data; S104: performing frame de-framing and frame image pre-processing on the video image to obtain standardized video frame image data; S105: taking the standardized text data, the standardized image data and the standardized video frame image data as standardized multimodal data of the multimodal data.

3. The multimodal data extraction method based on deep learning network according to claim 1, characterized in that: In step S2, the feature extraction processing steps include: S201: Standardizing multimodal data including standardized text data, standardized image data, and standardized video frame image data; S202: Input the standardized text data into the trained BERT model for deep encoding to obtain deep features of the text data; S203: Input the standardized image data into the trained convolutional neural network model to obtain deep features of the image data; S204: inputting the standardized video frame image data into the trained convolutional neural network model to obtain the deep features of the video frame image data; S205: The depth features of each modality include text data depth features, image data depth features, and video frame image data depth features.

4. The multimodal data extraction method based on deep learning network according to claim 1, characterized in that: In step S3, the feature fusion processing steps include: S301: standardizing the depth features of each modality in the multimodal data including text data depth features f1, image data depth features f2, and video frame image data depth features f3; S302: using a convolutional neural network to perform preliminary feature transformation on the deep features of each modality to obtain transformed deep features {f′1, f′2, f′3} of each modality; S303: Calculate the internal attention weight of the transformed deep features of each modality to obtain the first-level attention weight; perform weighted fusion on the transformed deep features of each modality according to the first-level attention weight to obtain the internal fusion features of each modality; S304: Calculate the inter-modal attention weights between the internal fusion features of each modality to obtain the second-level attention weights; perform weighted fusion on the internal fusion features of each modality according to the second-level attention weights to obtain joint features.

5. The multimodal data extraction method based on deep learning network according to claim 4, characterized in that: In step S303, the internal fusion features of each modality are calculated by the following steps: S3031: For the deep feature vector f′ of the mth modality m , divide it into N sub-regions {f′ m1 ,f′ m2 ,...,f′ mN }; S3032: Calculate the attention weight of each sub-region; The formula is: have been mn =(W mn f′ mn +b mn ) T u m ; Where: α mn represents the internal attention weight of the nth sub-region of the mth modality, and the internal attention weights of all sub-regions are the first-level attention weights; e mn represents the correlation score of the nth sub-region of the mth mode; W mn represents the nth sub-region-specific weight matrix of the mth mode, which is used to transform the sub-region f′ mn Mapping to attention space; b mn represents the bias vector; u m represents the modality-specific attention vector; S3033: Calculate the feature vector f′ of the mth modality according to the internal attention weights of each sub-region of the mth modality m Internal fusion features The formula is:

6. The multimodal data extraction method based on deep learning network according to claim 4, characterized in that: In step S304, the joint features are calculated by the following steps: S3041: Fusion of internal features of each modality as the new feature vector; S3042: Calculate the inter-modal attention weights between the internal fusion features of each modality; The formula is: Where: β m represents the inter-modal attention weight of the m-th modality; v m represents the internal fusion features of the mth modality; S3043: Calculate the joint feature F according to the inter-modal attention weights; The formula is:

7. The multimodal data extraction method based on deep learning network according to claim 1, characterized in that: In step S4, a data sample is constructed by the following steps: S401: assigning a unique ID to each multimodal data; S402: Construct a mapping table M map , used to record the ID of each multimodal data and its corresponding scientific research data label y i and the storage location of each modality data in the multimodal original data set; Mapping Table M map The formula is expressed as: M map ={(id1,y1,loc1 t ,loc1 i ,loc1 v ),(id2,y2,loc2 t ,loc2 i ,loc2 v ),...}; Where: id1 represents the ID of the first multimodal data; y1 represents the scientific research data label of the first multimodal data; loc1 t ,loc1 i ,loc1 v Indicates the storage location of the text data, image data and video data in the first multimodal data in the multimodal original data set; S403: traverse the mapping table M map : For the i-th multimodal data, determine its scientific research data label y i Whether it belongs to the preset target scientific research data label, namely C target If so, then according to the mapping table M map Extract the corresponding text data te from the multimodal original dataset i 、Image dataim i and video data ve i To construct the corresponding data sample (id i ,y i ,te i ,im i ,ve i ); otherwise, no data sample is constructed.

8. The multimodal data extraction method based on deep learning network according to claim 7, characterized in that: In step S4, the scientific research dataset is constructed by the following steps: S411: Store all constructed data samples in a temporary data set D temp middle; Temporary data set D temp The formula is expressed as: D temp ={(id1,y1,te1,im1,ve1),(id2,y2,te2,im2,ve2),...} Where: id1, y1, te1, im1, ve1 represent the multimodal data ID, scientific research data label, text data, image data and video data of the first data sample respectively; S412: Temporary data set D temp The text data, image data and video data in the image data are enhanced to obtain the enhanced temporary data set D temp,avg ; Enhanced temporary data set D temp,avg The formula is expressed as: D temp,avg ={(id1,y1,te 1,avg ,im 1,avg ,ve 1,avg ),(id2,y2,te 2,avg ,im 2,avg ,ve 2,avg ),...}; In the formula: te 1,avg ,im 1,avg ,ve 1,avg Respectively represent the enhanced text data, enhanced image data and enhanced video data of the first data sample; S413: From the enhanced temporary data set D temp,avg The text structured information, image features and video features of each data sample are extracted, and the extracted text structured information, image features and video features are integrated to form the structured information of each data sample, and the final scientific research data set D is obtained. struct ; Scientific research dataset D struct The formula is expressed as: <h2 style=";text-align:left;direction:ltr">D<h2 style=";text-align:left;direction:ltr"> struct <h2 style=";text-align:left;direction:ltr"> ({(id1,y1,S1),(id2,y2,S2),...}) S1=(f t ,f i ,f v ); Where: S1 represents the structural information of the first data sample; f t ,f i ,f v They represent text structured information, image features, and video features respectively.

9. The multimodal data extraction method based on deep learning network according to claim 8, characterized in that: In step S6, the target data is extracted by the following steps: S601: Obtaining the user's query requirement Q raw And query demand Q raw Parse and obtain the parsed query vector Q vec and the entity-relationship structure S QR ; where S QR Contains all entities and their relationships identified in the query requirements; S602: Research dataset D struct Extract features from the structured information of each data sample in the data to obtain the feature vector and corresponding ID set of each data sample; The feature vector set is obtained by concatenating the text feature vector and the multimodal feature vector; The text feature vector is to convert the text structured information, i.e., each set of entity-relationship-entity triples, into a feature vector; Multimodal feature vector f m , is the image feature f i , video key frame features {f f1 ,f f2 ,...} and video motion features f a The attention mechanism is used for fusion; the attention weights are set to be a i 、a f and a a , then the fused multimodal feature vector f m It can be expressed as: Among them, the attention weight is calculated by the neural network based on the feature vector: Where: w i 、w f 、w a represents a trainable weight vector; S603: Calculate query vector Q vec The cosine similarity between the feature vectors of each data sample is calculated, and the ID of the data sample whose cosine similarity exceeds the set cosine similarity threshold is put into the identification set of the preliminary matching samples; S604: For each data sample in the identification set of the preliminary matching samples, struct Extract the corresponding text structured information and construct the entity-relationship structure S of the sample sample ; S605: Calculate entity-relationship structure S QR and the entity-relationship structure S sample The structural similarity between struct , the structural similarity sim struct Data samples with a structure similarity greater than the set threshold are retained to obtain the matching sample identification set ID. final ; S606: Identify the set ID according to the matching sample final From the scientific research dataset D struct The corresponding target data samples are extracted from the dataset; the text structured information, image features, and video features of all target data samples are integrated into a structured matching result set R as the target data (the matching result set R needs to be converted into a format that is easy for users to understand).

10. The multimodal data extraction method based on deep learning network according to claim 9, characterized in that: In step S605, the structural similarity sim is calculated by the following steps: struct : Where: N represents the normalization constant; d represents the entity-relationship structure S QR and the entity-relationship structure S sample The graph edit distance of .