Multi-modal fusion-based data processing method, construction method for multi-modal fusion-based model, and related device

By constructing a multimodal fusion data processing model and utilizing a multi-head attention mechanism to extract and fuse features from one-dimensional and two-dimensional training data, the inaccuracy of existing multimodal fusion data processing methods is solved, and efficient data analysis of sample detection is achieved.

WO2026002055A1PCT designated stage Publication Date: 2026-01-02FUDAN UNIVERSITY
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2025/103509
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-06-23
Filing Date
2025-06-25
Publication Date
2026-01-02

AI Technical Summary

Technical Problem

Existing data processing methods based on multimodal fusion still need improvement in terms of accuracy, especially in sample testing where efficient data analysis is difficult to achieve.

Method used

By acquiring a multimodal training dataset, including one-dimensional and two-dimensional training data, and conducting joint training, a multi-head attention mechanism is used for feature extraction and fusion processing to construct a data processing model based on multimodal fusion.

Benefits of technology

It improves the performance of the data processing model and the accuracy of sample analysis, and enhances the accuracy of prediction results for sample data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025103509_02012026_PF_FP_ABST
    Figure CN2025103509_02012026_PF_FP_ABST
Patent Text Reader

Abstract

A multi-modal fusion-based data processing method, a construction method for a multi-modal fusion-based model, and a related device. The construction method for a multi-modal fusion-based data processing model comprises: obtaining a multi-modal training data set, the multi-modal training data set comprising multi-modal training data; and using the multi-modal training data in the multi-modal training data set to perform joint training so as to obtain a multi-modal fusion-based data processing model. The technical solution of embodiments of the present invention enhances the performance of the construction method for a multi-modal fusion-based data processing model, thereby improving the accuracy of sample analysis.
Need to check novelty before this filing date? Find Prior Art

Description

Data processing and model construction method based on multi-modal fusion, and related equipment

[0001] The present application claims priority from patent applications with application numbers: 202410835224.8, filed on June 25, 2024, and titled "Data processing and model construction method based on multi-modal fusion, and related equipment", and 202510848886.3, filed on June 23, 2025, and titled "Data processing and model construction method based on multi-modal fusion, and related equipment" TECHNICAL FIELD

[0002] Embodiments of the present application relate to the technical field of biological information processing, and in particular to a data processing and model construction method based on multi-modal fusion, and related equipment. BACKGROUND

[0003] Sample detection technology has been widely used in medical, biological, environmental and food science fields, etc. By detecting samples, detection data with analytical significance is obtained.

[0004] However, as the data processing method based on multi-modal fusion develops towards a more complex direction, the quantity requirement of the detection data of the sample also increases accordingly. For example, with the development of artificial intelligence, machine learning and deep learning technologies have been gradually used to realize data analysis. However, the accuracy of the current data processing method based on multi-modal fusion still needs to be improved. TECHNICAL PROBLEM

[0005] The problem solved by embodiments of the present application is to provide a data processing and model construction method based on multi-modal fusion, and related equipment, which is beneficial to improve the performance of the construction method of the data processing model based on multi-modal fusion, thereby further improving the accuracy of sample analysis. TECHNICAL SOLUTION

[0006] To solve the above problems, embodiments of the present application provide a construction method of a data processing model based on multi-modal fusion, comprising:

[0007] Obtain a multi-modal training data set, wherein the multi-modal training data set comprises multi-modal training data;

[0008] Perform joint training using the multi-modal training data in the multi-modal training data set to obtain a data processing model based on multi-modal fusion.

[0009] Correspondingly, embodiments of the present application also provide a construction device of a data processing model based on multi-modal fusion, comprising:

[0010] The training data acquisition unit is adapted to acquire a multi-modal training data set, and the multi-modal training data set comprises multi-modal training data.

[0011] The joint training unit is adapted to perform joint training by using the multi-modal training data in the multi-modal training data set, and acquire a data processing model based on multi-modal fusion.

[0012] Correspondingly, the embodiment of the present application also provides a data processing method based on multi-modal fusion, comprising:

[0013] Acquiring multi-modal to-be-processed data;

[0014] Inputting the multi-modal to-be-processed data into the data processing model based on multi-modal fusion constructed by using the construction method of the data processing model based on multi-modal fusion according to any one of the above, and acquiring a corresponding data processing result.

[0015] Correspondingly, the embodiment of the present application also provides a data processing device based on multi-modal fusion, comprising:

[0016] The data acquisition unit is adapted to acquire multi-modal to-be-processed data;

[0017] The analysis and prediction unit is adapted to input the multi-modal to-be-processed data into the data processing model based on multi-modal fusion constructed by using the construction method of the data processing model based on multi-modal fusion according to any one of the above, and acquire a corresponding data processing result.

[0018] Correspondingly, the embodiment of the present application also provides a device, characterized in that comprising at least one memory and at least one processor, the memory stores one or more computer instructions, wherein the one or more computer instructions are executed by the processor to implement the construction method of the data processing model based on multi-modal fusion according to any one of the above or the data processing method based on multi-modal fusion according to the above.

[0019] Correspondingly, the embodiment of the present application also provides a storage medium, characterized in that the storage medium stores one or more computer instructions, and the one or more computer instructions are used to implement the construction method of the data processing model based on multi-modal fusion according to any one of the above or the data processing method based on multi-modal fusion according to the above. Advantages

[0020] Compared with the prior art, the technical scheme of the present application has the following advantages: in the construction method of the data processing model based on multi-modal fusion provided by the present application, the multi-modal training data set is obtained, the multi-modal training data set includes multi-modal training data, which can increase the information richness of the training data in the training data set, and the multi-modal training data in the multi-modal training data set is used for joint training subsequently to obtain a data processing model based on multi-modal fusion, which is helpful to improve the performance of the data processing model based on multi-modal fusion constructed, and then when the data processing model based on multi-modal fusion is used to detect the sample data to be processed, it is helpful to improve the accuracy of the prediction result of the sample data to be processed. BRIEF DESCRIPTION OF DRAWINGS

[0021] Fig. 1 is a flowchart of an embodiment of the construction method of the data processing model based on multi-modal fusion provided by the technical scheme of the present application;

[0022] Fig. 2 is a flowchart of an embodiment of the data processing model based on multi-modal fusion obtained by using one-dimensional training data in the first training data set and two-dimensional training data in the second training data set for joint training in the technical scheme of the present application;

[0023] Fig. 3 is a comparison diagram of the confusion matrix of the accuracy of the data processing model based on multi-modal fusion;

[0024] Fig. 4 is a schematic diagram of the framework structure of a construction device of a data processing model based on multi-modal fusion in an embodiment of the present application;

[0025] Fig. 5 is a flowchart of a data processing method based on multi-modal fusion in an embodiment of the present application;

[0026] Fig. 6 is a schematic diagram of the structure of a data processing device based on multi-modal fusion in an embodiment of the present application;

[0027] Fig. 7 is a flowchart of an embodiment of the mass spectrometry data processing method of the present application;

[0028] Fig. 8 is a schematic diagram of the original mass spectrometry data of an embodiment corresponding to step S10 in Fig. 7;

[0029] Fig. 9 is a schematic diagram of the Gaussian distribution of signal intensity in an embodiment of the mass spectrometry data processing method of the present application;

[0030] Fig. 10 is a schematic diagram of the newly added mass spectrometry data of an embodiment corresponding to step S20 in Fig. 7;

[0031] FIG. 11 is a t-SNE distribution diagram of an embodiment based on original mass spectrum data, and a t-SNE distribution diagram of an embodiment based on mass spectrum data obtained by processing the mass spectrum data according to the embodiment of the present application;

[0032] FIG. 12 is a comparison diagram of a confusion matrix of accuracy of a machine learning model;

[0033] FIG. 13 is a comparison diagram of a confusion matrix of accuracy of a deep learning model;

[0034] FIG. 14 is a structural diagram of an embodiment of a data processing system according to the present application;

[0035] FIG. 15 is a structural diagram of an apparatus according to an embodiment of the present application. Embodiments of the present application

[0036] As described in the background, the accuracy of the existing data processing method based on multi-modal fusion still needs to be improved.

[0037] To solve the above technical problems, the embodiments of the present application provide a method for constructing a data processing model based on multi-modal fusion, which comprises: obtaining a multi-modal training data set, wherein the multi-modal training data set comprises multi-modal training data; performing joint training on the multi-modal training data in the multi-modal training data set to obtain a data processing model based on multi-modal fusion.

[0038] In the method for constructing a data processing model based on multi-modal fusion provided by the embodiments of the present application, the multi-modal training data set is obtained, and the multi-modal training data set comprises multi-modal training data, which can increase the information richness of the training data in the training data set. Subsequently, joint training is performed on the multi-modal training data in the multi-modal training data set to obtain a data processing model based on multi-modal fusion, which is helpful to improve the performance of the constructed data processing model based on multi-modal fusion, and further improve the accuracy of the prediction result of the sample data to be processed when the data processing model based on multi-modal fusion is used to detect the sample data to be processed.

[0039] In order to make the above-mentioned objects, features and advantages of the present application more obvious and easy to understand, the specific embodiments of the present application will be described in detail below with reference to the accompanying drawings.

[0040] FIG. 1 shows a flow diagram of an embodiment of a method for constructing a data processing model based on multi-modal fusion according to the technical scheme of the present application. Referring to FIG. 1, a method for constructing a data processing model based on multi-modal fusion can specifically comprise the following steps:

[0041] Step S110: obtaining a multi-modal training data set, wherein the multi-modal training data set comprises multi-modal training data;

[0042] Step S120: jointly training by using the multi-modal training data in the multi-modal training data set to obtain a data processing model based on multi-modal fusion.

[0043] Please continue to refer to FIG. 1, and perform step S110 to obtain a multi-modal training data set, which includes multi-modal training data.

[0044] The multi-modal training data set is obtained to provide a basis for subsequent joint training by using the multi-modal training data in the multi-modal training data set to obtain a data processing model based on multi-modal fusion.

[0045] In some embodiments, the multi-modal training data set includes a first training data set and a second training data set. The first training data set includes a plurality of one-dimensional training data, and the second training data set includes a plurality of two-dimensional training data.

[0046] The one-dimensional training data is one-dimensional data. Specifically, one-dimensional data is composed of ordered or unordered data with a peer relationship.

[0047] In some embodiments, the one-dimensional training data in the first training data set has a plurality of first modalities, that is, the first training data set has multi-modal one-dimensional training data. It should be noted that the first modality is a general term for the modality of the one-dimensional training data in the first training data set, to distinguish from the modality of the two-dimensional training data in the second training data set. In the first training data set, at least part of the one-dimensional training data has a first modality different from that of the other part of the one-dimensional training data.

[0048] As an example, the one-dimensional training data includes Raman spectrum data, mass spectrum data, infrared spectrum data, and point-scanning fluorescence lifetime imaging (FLIM) data.

[0049] Mass spectrum data is commonly used in the fields of medicine, biology, environment, and food science, and is information about the mass-to-charge ratio (m / z) of molecules in a sample and the intensity at each mass-to-charge ratio obtained by detecting the sample with a mass spectrometer. Accordingly, mass spectrum data contains information about the intensity varying with the mass-to-charge ratio.

[0050] Raman spectrum is an analysis method that uses the Raman scattering effect of molecules to obtain information about the vibration and rotation of molecules. Specifically, Raman spectrum data is obtained by analyzing scattered light spectra different from the wavelength of excitation light to obtain information about the vibration and rotation of molecules.

[0051] Infrared spectroscopy is an analytical chemistry technique that uses the interaction of infrared radiation with matter to study the structure and chemical composition of a sample. Specifically, an infrared spectrometer is used to obtain the energy absorbed or scattered by the sample at different wavelengths, thereby revealing information about the molecular structure, functional groups, and chemical bonds of the sample.

[0052] Fluorescence lifetime imaging microscopy (FLIM) technology is a technique that uses fluorescent materials to stain a sample and then detects the fluorescence generated by the sample under light to obtain information about the shape of the sample at corresponding positions. The point scanning fluorescence lifetime imaging data is point-to-point time-correlated single-photon counting (TCSPC) data obtained by combining the fluorescence lifetime imaging technology with the laser point scanning microscopy technology.

[0053] In other embodiments, the one-dimensional training data can also be two or three of Raman spectroscopy data, mass spectrometry data, and point scanning fluorescence lifetime imaging data, or can also include other one-dimensional data of other modalities required for training the data processing model based on multi-modal fusion, which is not limited here.

[0054] The source of the one-dimensional training data with multiple first modalities in the first training data set can be selected according to actual needs.

[0055] As an example, the step of obtaining the first training data set includes detecting a sample to be measured to obtain first original training data, and performing first data augmentation processing on the first original training data to obtain corresponding first new training data, wherein the first new training data and the first original training data constitute the first training data set.

[0056] As an example, the first original training data is first copied, and the signal intensity in the copied first original training data is adjusted according to the distribution rule of the signal intensity in the first original training data to obtain corresponding first new training data.

[0057] The number of one-dimensional training data in the first training data set can be set according to the training needs of the data processing model based on multi-modal fusion, which is not limited here.

[0058] The two-dimensional training data is two-dimensional data. The two-dimensional data is a combination of one-dimensional data, which is composed of multiple one-dimensional data.

[0059] The two-dimensional training data includes at least one of surface scanning optical imaging data, confocal imaging data, and surface scanning fluorescence lifetime imaging data.

[0060] The surface scanning optical imaging is a method for obtaining images of a sample surface using optical technology, which generally uses an optical lens and a detector to capture optical signals of the sample surface and convert them into digital images. The surface scanning optical imaging can generally be used to observe microstructures, detect surface defects, measure sizes and shapes, etc.

[0061] The surface scanning fluorescence lifetime imaging data, which is one of the surface scanning optical imaging data, is obtained by scanning the entire sample surface to obtain the fluorescence intensity and lifetime information of each point, so as to reveal the distribution and lifetime changes of fluorescent substances in the sample, thereby providing information about the structure and properties of the sample. By analyzing the surface scanning fluorescence lifetime imaging data, the fluorescence lifetime imaging image of the sample and the related quantitative information can be obtained.

[0062] The confocal imaging is an optical imaging method that uses point-by-point illumination and spatial pinhole modulation to remove scattered light from the non-focal plane of the sample, which can improve the optical resolution and visual contrast. Specifically, the confocal imaging data includes the obtained fluorescence intensity and position information, and the related information such as the three-dimensional structure and fluorescence lifetime of the sample obtained therefrom.

[0063] In other embodiments, the two-dimensional training data can also be both the surface scanning optical imaging data and the confocal imaging data, or can also include other two-dimensional imaging data of the sample, which is not limited herein.

[0064] The source of the two-dimensional training data in the second training data set can be selected according to actual needs. As an example, the step of obtaining the second training data set includes: detecting the sample to be tested to obtain second original training data; and performing second data augmentation processing on the second original training data to obtain corresponding second new training data, wherein the second new training data and the second original training data constitute the second training data set.

[0065] As an example, the step of performing second data augmentation processing on the second original training data to obtain corresponding second new training data includes: copying the second original training data to obtain copied second original training data; and adjusting the signal intensity in the copied second original training data according to the signal intensity distribution law in the second original training data to obtain corresponding second new training data.

[0066] The number of two-dimensional training data in the second training data set can be set according to the model training requirements, which is not limited herein.

[0067] Please continue to refer to FIG. 1, and perform step S120 to jointly train based on the multi-modal training data in the multi-modal training data set, so as to obtain a data processing model based on multi-modal fusion.

[0068] The multi-modal training data in the multi-modal training data set is jointly trained to obtain a data processing model based on multi-modal fusion, and subsequent use of the data processing model based on multi-modal fusion can realize data analysis such as data classification.

[0069] In some embodiments, the multi-modal training data set includes a first training data set and a second training data set. The first training data set includes a plurality of one-dimensional training data, and the second training data set includes a plurality of two-dimensional training data. Accordingly, the step of jointly training based on the multi-modal training data in the multi-modal training data set to obtain a data processing model based on multi-modal fusion includes jointly training based on the one-dimensional training data in the first training data set and the two-dimensional training data in the second training data set to obtain a data processing model based on multi-modal fusion.

[0070] FIG. 2 shows a flowchart of an embodiment of the present application in which the first one-dimensional training data in the first training data set and the two-dimensional training data in the second training data set are jointly trained to obtain a data processing model based on multi-modal fusion. Referring to FIG. 2, taking the first one-dimensional training data with multiple first modalities in the first training data set and the second one-dimensional training data with one or more second modalities in the second training data set as an example, the step of jointly training based on the first one-dimensional training data in the first training data set and the two-dimensional training data in the second training data set to obtain a data processing model based on multi-modal fusion can include:

[0071] Step S1201: performing first fusion processing on the first one-dimensional training data with multiple first modalities in the first training data set to obtain a plurality of first fusion data corresponding to the first fusion data set;

[0072] Step S1202: performing first directional attention feature extraction processing on the first fusion data in the first fusion data set to obtain first directional attention feature data corresponding to each first modality;

[0073] Step S1203: performing first operation processing on each first directional attention feature data corresponding to each first modality and the one-dimensional training data corresponding to other first modalities to obtain a plurality of first feature data corresponding to the first feature data set;

[0074] Step S1204: performing second attention feature extraction processing on the two-dimensional training data in the second training data set respectively to obtain corresponding second attention feature data;

[0075] Step S1205: performing second operation processing on the second attention feature data respectively and one-dimensional training data of a corresponding first modality to obtain a plurality of second feature data respectively and generate a second feature data set;

[0076] Step S1206: performing first aggregation processing on the first feature data in the first feature data set and corresponding second feature data in the second feature data set to obtain a plurality of first aggregated feature data respectively and generate a first aggregated feature data set;

[0077] Step S1207: performing first classification processing on the first aggregated feature data in the first aggregated feature data set to obtain a corresponding first prediction result;

[0078] Step S1208: calculating a first loss value of the data processing model based on multi-modal fusion by using a preset first loss function according to the obtained first prediction result, and adjusting the weight of the data processing model based on multi-modal fusion according to the calculated first loss value until the first loss value of the data processing model based on multi-modal fusion converges.

[0079] Please continue to refer to FIG. 2, and perform step S1201 to perform first fusion processing on one-dimensional training data with a plurality of first modalities in the first training data set to obtain a plurality of first fusion data respectively and generate the first fusion data set.

[0080] The first fusion processing on the one-dimensional training data with a plurality of first modalities in the first training data set to obtain a plurality of first fusion data respectively and generate the first fusion data set provides a basis for subsequent first directional attention feature extraction processing on the first fusion data in the first fusion data set to obtain first directional attention feature data corresponding to each first modality.

[0081] The first fusion processing on the one-dimensional training data with a plurality of first modalities in the first training data set means that a plurality of one-dimensional training data with different first modalities are fused. In other words, the plurality of one-dimensional training data fused have different first modalities.

[0082] As an example, the first training data set includes Raman spectrum data and mass spectrum data, and a piece of Raman spectrum data and a piece of mass spectrum data are respectively acquired from the first training data set, and the acquired piece of Raman spectrum data and the acquired piece of mass spectrum data are subjected to first fusion processing, and a corresponding piece of first fusion data is acquired.

[0083] According to actual needs, a preset data fusion algorithm can be used to fuse the one-dimensional training data with multiple first modalities in the first training data set. The data fusion algorithm can be selected by a person skilled in the art according to actual needs, and is not limited herein.

[0084] The one-dimensional training data with multiple first modalities in the first training data set is subjected to first fusion processing, so that the generated corresponding first fusion data has the characteristic information of the one-dimensional training data with multiple first modalities, so that the generated first fusion data has more rich and comprehensive characteristic information, which is helpful to improve the prediction accuracy and generalization ability of the obtained multi-modal fusion-based data processing model.

[0085] Please continue to refer to FIG. 2. In step S1202, the first fusion data in the first fusion data set is subjected to first directional attention feature extraction processing respectively, and first directional attention feature data corresponding to each first modality is acquired.

[0086] The first fusion data in the first fusion data set is subjected to first directional attention feature extraction processing respectively, and first directional attention feature data corresponding to each first modality is acquired, which provides a basis for subsequent first operation processing of each first modality corresponding first directional attention feature data and other one-dimensional training data of corresponding first modalities to acquire a plurality of first feature data and generate a first feature data set.

[0087] The first fusion data in the first fusion data set is subjected to first directional attention feature extraction processing, so that the acquired first directional attention feature data can pay more attention to the detailed information related to the corresponding directional features, which is helpful to improve the expression ability of the first directional attention feature and to speed up the processing efficiency of the information.

[0088] The first fusion data in the first fusion data set is subjected to first directional attention feature extraction processing, which means that the first fusion data in the first fusion data set is subjected to directional attention feature extraction processing. The corresponding direction corresponds to the one-dimensional training data of the multiple first modalities of the first fusion data.

[0089] As an example, the first training data set includes Raman spectrum data and mass spectrum data, and the step of respectively performing first directional attention feature extraction processing on the first fusion data in the first fusion data set to obtain the first directional attention feature data corresponding to each first modality includes: respectively performing attention feature extraction processing in the Raman direction and the mass spectrum direction on the first fusion data in the first fusion data set to obtain the first directional attention feature data in the Raman direction and the first directional attention feature data in the mass spectrum direction.

[0090] As an example, the first directional attention feature data corresponding to each first modality is obtained by respectively performing first directional attention feature extraction processing on the first fusion data in the first fusion data set using a first multi-head attention mechanism.

[0091] Specifically, the first fusion data is multiplied by a corresponding plurality of groups of trainable parameter matrices WQ, WK, and WV respectively to obtain a plurality of groups of query matrices (Q), key matrices (K), and value (V) matrices corresponding to the parameter matrices WQ, WK, and WV respectively, and the plurality of groups of query matrices, key matrices, and value matrices are subjected to splicing operations and vector dot product operations to obtain corresponding attention matrices. Wherein, the corresponding plurality of groups of trainable parameter matrices WQ, WK, and WV are respectively corresponding to the plurality of groups of trainable parameter matrices WQ, WK, and WV of each first modality, in other words, each first modality corresponds to a group of trainable parameter matrices WQ, WK, and WV.

[0092] Please continue to refer to FIG. 2, and perform step S1203 to respectively perform first operation processing on the first directional attention feature data corresponding to each first modality and one-dimensional training data of a corresponding other first modality to obtain a plurality of first feature data corresponding to each first modality, and generate a first feature data set.

[0093] Respectively performing first operation processing on the first directional attention feature data corresponding to each first modality and one-dimensional training data of a corresponding other first modality to obtain a plurality of first feature data corresponding to each first modality, and generate a first feature data set, in preparation for subsequent first aggregation processing on the first feature data in the first feature data set and the corresponding second feature data in the second feature data set to obtain a plurality of first aggregation feature data corresponding to each first modality, and generate a first aggregation feature data set.

[0094] As an example, the first operation processing is convolution processing.

[0095] As an example, the first training data set includes Raman spectrum data and mass spectrum data, and correspondingly, the step of performing first operation processing on the first directional attention feature data corresponding to each first modality and one-dimensional training data of other corresponding first modalities respectively includes: fusing the first directional attention feature data in the Raman direction with the mass spectrum data, and performing convolution processing on the first directional attention feature data in the mass spectrum direction and the Raman data.

[0096] It can be understood that the first operation processing can also include other types of processing methods in addition to convolution processing, such as addition, division, or multiplication, and the like, which can be selected by those skilled in the art according to actual needs, and are not limited herein.

[0097] In obtaining the first directional attention feature data corresponding to each first modality, the one-dimensional training data with multiple first modalities is first fused to obtain first fusion data, which corresponds to the one-dimensional training data with multiple first modalities before the fusion processing, and then the first fusion data is subjected to first directional attention feature extraction processing to obtain the first directional attention feature data corresponding to each first modality. The first directional attention feature data and the one-dimensional training data with multiple first modalities corresponding to the first fusion data are one-to-one corresponding.

[0098] Correspondingly, in the step of performing first operation processing on the first directional attention feature data corresponding to each first modality and one-dimensional training data of other corresponding first modalities respectively, the one-dimensional training data of other corresponding first modalities is the one-dimensional training data of other first modalities corresponding to the first fusion data, except for the one-dimensional training data of the first modality corresponding to the first directional attention feature data.

[0099] For example, in the case that the one-dimensional training data subjected to the first fusion processing includes one-dimensional input data 1, one-dimensional input data 2, and one-dimensional input data 3, the first directional attention feature data obtained by performing first directional attention feature extraction processing on the one-dimensional input data 1 is subjected to first operation processing on the one-dimensional input data 2 and the one-dimensional input data 3 respectively, the first directional attention feature data obtained by performing first directional attention feature extraction processing on the one-dimensional input data 2 is subjected to first operation processing on the one-dimensional input data 1 and the one-dimensional input data 3 respectively, and the first directional attention feature data obtained by performing first directional attention feature extraction processing on the one-dimensional input data 3 is subjected to first operation processing on the one-dimensional input data 2 and the one-dimensional input data 1 respectively.

[0100] The first directional attention feature data corresponding to each first modality is respectively subjected to first operation processing with one-dimensional training data of other corresponding first modalities, so that the obtained first feature data can contain associated feature information between one-dimensional training data of two different first modalities, which is beneficial to further increase the richness of feature information in the first feature data.

[0101] Please continue to refer to FIG. 2, and perform step S1204 to respectively perform second attention feature extraction processing on two-dimensional training data in the second training data set to obtain corresponding second attention feature data.

[0102] The two-dimensional training data in the second training data set are respectively subjected to second attention feature extraction processing to obtain corresponding second attention feature data, which provides a basis for subsequent second operation processing of the second attention feature data with one-dimensional training data of corresponding first modalities to obtain corresponding multiple second feature data and generate a second feature data set.

[0103] The two-dimensional training data in the second training data set are respectively subjected to second attention feature extraction processing, so that the obtained second attention feature data can pay more attention to detailed information related to features in corresponding directions.

[0104] As an example, the second training data set includes confocal imaging data and face-scanning fluorescence lifetime imaging training data. Accordingly, the step of respectively performing second attention feature extraction processing on two-dimensional training data in the second training data set to obtain corresponding second attention feature data includes: performing self-attention feature extraction processing in the confocal imaging direction on the confocal imaging data in the second training data set to obtain second attention feature data in the confocal imaging direction, and performing self-attention feature extraction processing in the face-scanning fluorescence lifetime imaging direction on the face-scanning fluorescence lifetime imaging training data in the second training data set to obtain second attention feature data in the face-scanning fluorescence lifetime imaging direction.

[0105] In some embodiments, the second attention feature extraction processing on the two-dimensional training data in the second training data set is performed by using a second multi-head attention mechanism to obtain corresponding multiple second attention feature data. The second multi-head attention mechanism can be the same as or different from the first multi-head attention mechanism.

[0106] Please continue to refer to FIG. 2, and perform step S1205 to respectively perform second operation processing on the second attention feature data with one-dimensional training data of corresponding first modalities to obtain corresponding multiple second feature data and generate a second feature data set.

[0107] The second attention feature data of each second modality is respectively subjected to second operation processing with the one-dimensional training data of the corresponding first modality, to obtain a plurality of second feature data, and generate a second feature data set, to prepare for subsequent first aggregation processing of the first feature data in the first feature data set and the corresponding second feature data in the second feature data set, to obtain a plurality of first aggregation feature data, and generate a first aggregation feature data set.

[0108] In the step of subjecting the second attention feature data to second operation processing with the one-dimensional training data of the corresponding first modality, the one-dimensional training data of the corresponding first modality is the one-dimensional training data with multiple first modalities used when the corresponding first fusion data is obtained by the first fusion processing.

[0109] As an example, the first training data set includes Raman spectrum data and mass spectrum data, and the second training data set includes confocal imaging data. Correspondingly, the step of subjecting the second attention feature data to second operation processing with the one-dimensional training data of the corresponding first modality includes: subjecting the second attention feature data in the confocal imaging direction to second convolution processing with the corresponding Raman spectrum data and mass spectrum data. As another example, the first training data set includes infrared spectrum data, and the second training data set includes face scanning optical imaging data. Correspondingly, the step of subjecting the second attention feature data of each second modality to second operation processing with the one-dimensional training data of the corresponding first modality includes: subjecting the second attention feature data in the face scanning optical imaging direction to second operation processing with the infrared spectrum data.

[0110] The second attention feature data is respectively subjected to second operation processing with the one-dimensional training data of the corresponding first modality, so that the obtained second feature data can contain the associated feature information between the one-dimensional training data and the two-dimensional training data, which is beneficial to further increase the richness of information in the second feature data.

[0111] The second operation processing is performed according to the description of the first operation processing, and will not be repeated here.

[0112] Please continue to refer to FIG. 2, and perform step S1206 to aggregate the first feature data in the first feature data set and the corresponding second feature data in the second feature data set, to obtain a plurality of first aggregation feature data, and generate a first aggregation feature data set.

[0113] The first feature data in the first feature data set and the corresponding second feature data in the second feature data set are subjected to first aggregation processing, a plurality of first aggregated feature data corresponding to the first aggregated feature data are obtained, a first aggregated feature data set is generated, and the first aggregated feature data in the first aggregated feature data set is subjected to subsequent classification processing to obtain a corresponding first prediction result.

[0114] The step of performing first aggregation processing on the first feature data in the first feature data set and the corresponding second feature data in the second feature data set includes splicing the first feature data in the first feature data set and the second feature data in the second feature data set to obtain a plurality of aggregated feature data.

[0115] In the step of performing aggregation processing on the first feature data in the first feature data set and the corresponding second feature data in the second feature data set, the first operation processing is performed to obtain the first feature data and the corresponding second feature data by using the same one-dimensional training data with multiple first modalities as operation cores for first operation processing and second operation processing.

[0116] For example, in the case that the one-dimensional training data subjected to the first fusion processing includes one-dimensional input data 1, one-dimensional input data 2 and one-dimensional input data 3, correspondingly, in the step of performing first operation processing on the first directional attention feature data of each first modality and the one-dimensional training data of the corresponding other first modality in step S1203, the first directional attention feature data obtained by performing first directional attention feature extraction processing on the one-dimensional input data 1 is subjected to first operation processing on the one-dimensional input data 2 and the one-dimensional input data 3, the first directional attention feature data obtained by performing first directional attention feature extraction processing on the one-dimensional input data 2 is subjected to first operation processing on the one-dimensional input data 1 and the one-dimensional input data 3, and the first directional attention feature data obtained by performing first directional attention feature extraction processing on the one-dimensional input data 3 is subjected to first operation processing on the one-dimensional input data 1 and the one-dimensional input data 2.

[0117] Meanwhile, in the case that the two-dimensional training data in the second training data set includes two-dimensional input data 1 and two-dimensional input data 2, correspondingly, in the step of performing the second operation processing on the second attention feature data and the one-dimensional training data of the corresponding first modality in step S1205, the second directional attention feature data obtained by performing the second attention feature extraction processing on the two-dimensional input data 1 is respectively subjected to the second operation processing with the one-dimensional input data 1, the one-dimensional input data 2 and the one-dimensional input data 3, and the second directional attention feature data obtained by performing the second attention feature extraction processing on the two-dimensional input data 2 is respectively subjected to the second operation processing with the one-dimensional input data 1, the one-dimensional input data 2 and the one-dimensional input data 3.

[0118] Correspondingly, in the step of performing the aggregation processing on the first feature data in the first feature data set and the corresponding second feature data in the second feature data set in step S1206, the first directional attention feature data corresponding to the one-dimensional input data 1 is respectively subjected to the first operation processing with the one-dimensional input data 2 and the one-dimensional input data 3 to obtain the first feature data, the first directional attention feature data corresponding to the one-dimensional input data 2 is respectively subjected to the first operation processing with the one-dimensional input data 1 and the one-dimensional input data 3 to obtain the first feature data, the first directional attention feature data corresponding to the one-dimensional input data 3 is respectively subjected to the first operation processing with the one-dimensional input data 1 and the one-dimensional input data 2 to obtain the first feature data, and the second directional attention feature data corresponding to the two-dimensional input data 1 is respectively subjected to the second operation processing with the one-dimensional input data 1, the one-dimensional input data 2 and the one-dimensional input data 3 to obtain the second feature data, and the second directional attention feature data corresponding to the two-dimensional input data 2 is respectively subjected to the second operation processing with the one-dimensional input data 1, the one-dimensional input data 2 and the one-dimensional input data 3 to obtain the second feature data, are subjected to the aggregation processing.

[0119] Correspondingly, the plurality of aggregated features includes first feature data and second feature data, the first feature data is obtained by one-dimensional training data with a plurality of first modalities, and the second feature data is obtained by two-dimensional training data with one or more second modalities, so that the aggregated features have more rich and more comprehensive sample feature information.

[0120] In some embodiments, after the first directional attention feature data corresponding to each first modality is respectively subjected to first operation processing on one-dimensional training data of the corresponding other first modality to obtain a plurality of pieces of first feature data and a first feature data set is generated, and before the first feature data in the first feature data set and the corresponding second feature data in the second feature data set are subjected to first aggregation processing to obtain a plurality of pieces of first aggregated feature data and a first aggregated feature data set is generated, the one-dimensional training data in the first training data set and the two-dimensional training data in the second training data set are used for joint training to obtain a data processing model based on multi-modal fusion, and the method further comprises: performing first dimension reduction processing on the first feature data in the first feature data set.

[0121] The first dimension reduction processing on the first feature data in the first feature data set can realize dimension reduction of the first feature data and the second feature data, which is helpful to reduce the subsequent operation amount.

[0122] As an example, the step of performing first dimension reduction processing on the first feature data in the first feature data set comprises: performing first convolution dimension reduction processing on the first feature data in the first feature data set.

[0123] In other embodiments, other dimension reduction methods can also be used to perform first dimension reduction processing on the first feature data in the first feature data set, which can be selected by those skilled in the art according to actual needs, and is not limited herein.

[0124] After the second attention feature data is respectively subjected to second operation processing on one-dimensional training data of the corresponding first modality to obtain a plurality of pieces of second feature data and a second feature data set is generated, and before the first feature data in the first feature data set and the corresponding second feature data in the second feature data set are subjected to first aggregation processing to obtain a plurality of pieces of first aggregated feature data and a first aggregated feature data set is generated, the one-dimensional training data in the first training data set and the two-dimensional training data in the second training data set are used for joint training to obtain a data processing model based on multi-modal fusion, and the method further comprises: performing second dimension reduction processing on the second feature data in the second feature data set.

[0125] The second dimension reduction processing on the second feature data in the second feature data set can realize dimension reduction of the second feature data, which is helpful to reduce the subsequent operation amount.

[0126] As an example, the step of performing second dimension reduction processing on the second feature data in the second feature data set comprises: performing second operation dimension reduction processing on the second feature data in the second feature data set.

[0127] In other embodiments, other dimension reduction methods can also be used for the second dimension reduction processing of the second feature data in the second feature data set, and those skilled in the art can select according to actual needs, which are not limited here.

[0128] Correspondingly, the first feature data in the first feature data set and the corresponding second feature data in the second feature data set are subjected to first aggregation processing to obtain a plurality of corresponding first aggregated feature data and generate a first aggregated feature data set. The step of generating the first aggregated feature data set includes: subjecting the first feature data subjected to the first dimension reduction processing to first aggregation processing with the corresponding second feature data subjected to the second dimension reduction processing, obtaining a plurality of corresponding first aggregated feature data, and generating a first aggregated feature data set.

[0129] Please continue to refer to FIG. 2, and perform step S1207 to respectively subject the first aggregated feature data in the first aggregated feature data set to first classification processing to obtain corresponding first prediction results.

[0130] The first aggregated feature data in the first aggregated feature data set is respectively subjected to first classification processing to obtain corresponding first prediction results, which provides a basis for subsequently calculating the first loss value of the data processing model based on multi-modal fusion using a preset first loss function according to the obtained first prediction results.

[0131] In some embodiments, a first linear classifier is used to respectively subject the first aggregated feature data in the first aggregated feature data set to first linear classification processing to obtain corresponding first prediction results. The corresponding first prediction results are information of probability values of the first aggregated feature data corresponding one-dimensional training data and two-dimensional training data belonging to each sample category.

[0132] Please continue to refer to FIG. 2, and perform step S1208 to calculate the first loss value of the data processing model based on multi-modal fusion using a preset first loss function according to the obtained first prediction results, and adjust the weight of the data processing model based on multi-modal fusion according to the calculated first loss value until the first loss value of the data processing model based on multi-modal fusion converges.

[0133] The first loss value of the data processing model based on multi-modal fusion is calculated using a preset first loss function according to the obtained first prediction results, and the weight of the data processing model based on multi-modal fusion is adjusted according to the calculated first loss value until the first loss value of the data processing model based on multi-modal fusion converges, thereby obtaining a final data processing model based on multi-modal fusion.

[0134] In some embodiments, according to the obtained first prediction result, a preset first loss function is used to calculate a first loss value of the data processing model based on multi-modal fusion, and the weight of the data processing model based on multi-modal fusion is adjusted according to the calculated first loss value, thereby completing one iteration training of the data processing model based on multi-modal fusion.

[0135] Specifically, the process of each iteration training includes: using a preset number (such as batch size) of aggregated feature data to respectively train the data processing model based on multi-modal fusion to be trained, and obtaining a preset number of first prediction results; according to the difference between the first prediction result and the true result, a preset first loss function is used to calculate the corresponding first loss value; according to the first gradient value obtained by back propagation derivation, the weight of the data processing model based on multi-modal fusion is adjusted once.

[0136] Correspondingly, multiple iteration trainings are performed, and the weight of the data processing model based on multi-modal fusion is iteratively updated until the first loss value of the data processing model based on multi-modal fusion on the preset verification set converges. Wherein, the first loss value of the data processing model based on multi-modal fusion on the preset first verification set converges, means that the first loss value of the data processing model based on multi-modal fusion on the preset first verification set reaches a minimum value.

[0137] For more detailed content of the multiple iteration training process of the data processing model based on multi-modal fusion, please refer to the iteration training process of the neural network model in the prior art, which will not be repeated here.

[0138] In some embodiments, the data processing model based on multi-modal fusion is used for classifying pathogenic bacteria. In other embodiments, the data processing model based on multi-modal fusion can also include other types of samples such as particulate matter for classification, which is not limited here.

[0139] The above takes the example that the first training data set includes one-dimensional training data with multiple first modalities, and the second training data set includes two-dimensional training data with one or more second modalities, and introduces how to use one-dimensional training data in the first training data set and two-dimensional training data in the second training data set for joint training to obtain the data processing model based on multi-modal fusion. The present application is not limited to this.

[0140] FIG. 2 describes the construction method of the data processing model based on multi-modal fusion in the embodiment of the present application by taking the example of jointly training the data processing model based on multi-modal fusion by using the one-dimensional training data in the first training data set and the two-dimensional training data in the second training data set. It can be understood that the first training data set can also include one-dimensional training data with a first modality, and the second training data set includes two-dimensional training data with one or more second modalities

[0141] Correspondingly, in the case that the first training data set includes one-dimensional training data with a first modality, and the second training data set includes two-dimensional training data with one or more second modalities, the step of jointly training the data processing model based on multi-modal fusion by using the one-dimensional training data in the first training data set and the two-dimensional training data in the second training data set includes: performing first self-attention feature extraction processing on the one-dimensional training data in the first training data set respectively to obtain corresponding multiple pieces of first attention feature data, forming a first attention feature data set; performing second self-attention feature extraction processing on the two-dimensional training data in the second training data set respectively to obtain corresponding multiple pieces of second attention feature data, generating a second attention feature data set; performing second aggregation processing on the first attention feature data in the first attention feature data set and the second attention feature data in the second attention feature data set to obtain corresponding multiple pieces of second aggregation feature data, generating a second aggregation feature data set; performing second classification processing on the second aggregation feature data in the second aggregation feature data set respectively to obtain corresponding second prediction results; calculating a second loss value of the data processing model based on multi-modal fusion by using a preset second loss function according to the obtained second prediction results, and adjusting the weight of the data processing model based on multi-modal fusion according to the calculated second loss value until the second loss value of the data processing model based on multi-modal fusion converges.

[0142] For the step of jointly training the data processing model based on multi-modal fusion by using the one-dimensional training data in the first training data set and the two-dimensional training data in the second training data set in the case that the first training data set includes one-dimensional training data with a first modality, and the second training data set includes two-dimensional training data with one or more second modalities, please refer to the corresponding part of the aforementioned step of jointly training the data processing model based on multi-modal fusion by using the one-dimensional training data in the first training data set and the two-dimensional training data in the second training data set in the case that the first training data set includes one-dimensional training data with multiple first modalities, and the second training data set includes two-dimensional training data with one or more second modalities, and the description is not repeated here.

[0143] In another embodiment, the data processing model based on multi-modal fusion can also be trained by using one-dimensional training data with multiple first modalities in the first training data set.

[0144] Specifically, the step of training the data processing model based on multi-modal fusion by using one-dimensional training data with multiple first modalities in the first training data set comprises: performing first fusion processing on the one-dimensional training data with multiple first modalities in the first training data set to obtain corresponding multiple pieces of first fusion data and generate the first fusion data set; performing first directional attention feature extraction processing on the first fusion data in the first fusion data set respectively to obtain first directional attention feature data corresponding to each first modality; performing first operation processing on the first directional attention feature data corresponding to each first modality and one-dimensional training data of other corresponding first modalities respectively to obtain corresponding multiple pieces of first feature data and generate a first feature data set; performing first aggregation processing on the first feature data in the first feature data set to obtain corresponding multiple pieces of first aggregation feature data and generate a first aggregation feature data set; performing first classification processing on the aggregation feature data in the first aggregation feature data set respectively to obtain corresponding first prediction results; calculating a first loss value of the data processing model based on multi-modal fusion by using a preset first loss function according to the obtained first prediction results, and adjusting the weight of the data processing model based on multi-modal fusion according to the calculated first loss value until the first loss value of the data processing model based on multi-modal fusion converges.

[0145] For the content of training the data processing model based on multi-modal fusion by using one-dimensional training data with multiple first modalities in the first training data set, please refer to the content of the corresponding part shown in FIG. 2, which will not be repeated here.

[0146] Referring to FIG. 3, a comparative schematic diagram of a prediction accuracy confusion matrix of a data processing model based on multi-modal fusion is shown, and the darker the color of the grid where the number is located, the higher the corresponding prediction accuracy.

[0147] Among them, FIG. 3(a) shows the prediction accuracy confusion matrix of the data processing model based on multi-modal fusion trained by using only Raman spectrum training data, FIG. 3(b) shows the prediction accuracy confusion matrix of the data processing model based on multi-modal fusion trained by using only mass spectrum training data, and FIG. 3(c) shows the prediction accuracy confusion matrix of the data processing model based on multi-modal fusion trained by using the first training data set formed by the Raman spectrum training data and the mass spectrum training data obtained from the pathogenic bacteria samples and the second training data set formed by the pathogenic bacteria image data obtained from the pathogenic bacteria samples.

[0148] As an example, the sample model is used for pathogenic bacteria analysis, and in the confusion matrix, True represents the true result, Predicted represents the predicted result, and the type of pathogenic bacteria includes different species or subspecies of Enterobacter cloacae, such as E. bugandensis, E. hormaechei, E. chengduensis, E. cloacae, E. dissolvens, E. asburiae, E. kobei, and E. ludwigii. Taking E. hormaechei as an example, referring to FIG. 3(a), when the sample is analyzed by the data processing model for hospital clinical identification of pathogenic bacteria based on multi-modal fusion trained by using the Raman spectrum training data, the probability of predicting the accurate result is 0.49 (i.e., 49%), the probability of predicting E. chengduensis is 0.46 (i.e., 46%), and the probability of predicting E. cloacae is 0.05 (i.e., 5%); referring to FIG. 3(b), when the sample is analyzed by the data processing model for hospital clinical identification of pathogenic bacteria based on multi-modal fusion trained by using the mass spectrum training data, the probability of predicting the accurate result is 0.3 (i.e., 30%), the probability of predicting E. bugandensis is 0.55 (i.e., 55%), and the probability of predicting E. chengduensis is 0.15 (i.e., 15%); referring to FIG. 3(c), when the sample is analyzed by the data processing model based on multi-modal fusion in the embodiment, the probability of predicting the accurate result is 1 (i.e., 100%).

[0149] As can be seen from FIG. 3(a), the overall prediction accuracy (acc) of the data processing model based on multi-modal fusion trained by using only the Raman spectrum training data for pathogenic bacteria classification is 71.72%, as can be seen from FIG. 3(b), the overall prediction accuracy (acc) of the data processing model based on multi-modal fusion trained by using only the mass spectrum training data for pathogenic bacteria classification is 60.71%, and as can be seen from FIG. 3(c), the prediction accuracy (acc) of the data processing model based on multi-modal fusion in the embodiment for pathogenic bacteria classification is 98.91%.

[0150] FIG. 3(d) shows a schematic diagram of the prediction accuracy confusion matrix of the particle sample generated by the data processing model based on multi-modal fusion constructed by using the method in the embodiment of the application.

[0151] Referring to FIG. 3(d), the multi-modal fusion-based data processing model generated by the method for constructing a multi-modal fusion-based data processing model in the embodiment of the present application can be used to analyze various particulate samples, and the prediction accuracy can reach 100%. Taking polypropylene as an example, the multi-modal fusion-based data processing model generated by the method for constructing a multi-modal fusion-based data processing model in the embodiment of the present application can be used to analyze polypropylene particulate samples, and the probability of predicting accurate results is 1 (i.e. 100%).

[0152] As can be seen from the above description, the multi-modal fusion-based data processing model generated by the method for generating a multi-modal fusion-based data processing model in the embodiment of the present application has high prediction accuracy.

[0153] It is worth noting that for biological samples that need to be cultured for a period of time, the multi-modal fusion-based data processing model constructed by the method for constructing a multi-modal fusion-based data processing model in the embodiment of the present application can effectively integrate information of different scales of samples within a short period of time, shorten the culture period of biological samples required to obtain corresponding sample data, and improve the identification speed of actual clinical samples or other biological samples.

[0154] For example, when using the current multi-modal fusion-based data processing model to analyze and identify biological samples, the biological samples usually need to be cultured for at least 48 hours to obtain mass spectrometry data with resolved peaks, so that the biological samples can be accurately analyzed and identified. However, the multi-modal fusion-based data processing model constructed by the method for constructing a multi-modal fusion-based data processing model in the embodiment of the present application can be used to analyze and identify mass spectrometry data obtained from biological samples cultured for only 24 hours, effectively shortening the culture period of the biological samples.

[0155] Correspondingly, the embodiment of the present application also provides a construction device of a multi-modal fusion-based data processing model.

[0156] FIG. 4 shows a schematic diagram of the framework structure of a construction device of a multi-modal fusion-based data processing model in the embodiment of the present application. Referring to FIG. 4, a construction device 400 of a multi-modal fusion-based data processing model includes a training data acquisition unit 401 adapted to acquire a multi-modal training data set including multi-modal training data, and a joint training unit 402 adapted to perform joint training using the multi-modal training data in the multi-modal training data set to obtain a multi-modal fusion-based data processing model.

[0157] The construction apparatus of the data processing model based on multi-modal fusion in the embodiment can be used to execute the construction method of the data processing model based on multi-modal fusion, or can also execute the construction method of the data processing model based on multi-modal fusion by using other functional structures. The construction apparatus of the data processing model based on multi-modal fusion is described in the foregoing construction method of the data processing model based on multi-modal fusion, and will not be described herein again.

[0158] Correspondingly, the embodiment of the application further provides a data processing method based on multi-modal fusion.

[0159] FIG. 5 shows a flow diagram of a data processing method based on multi-modal fusion in the embodiment of the application. Referring to FIG. 5, a data processing method based on multi-modal fusion can include:

[0160] Step S510: acquiring multi-modal data to be processed;

[0161] Step S520: inputting the multi-modal data to be processed into the data processing model based on multi-modal fusion constructed by the construction method of the data processing model based on multi-modal fusion, and acquiring a corresponding data processing result.

[0162] In some embodiments, the multi-modal data to be processed includes one-dimensional data to be processed and two-dimensional data to be processed.

[0163] Correspondingly, the data processing model based on multi-modal fusion constructed by the construction method of the data processing model based on multi-modal fusion is used to analyze and process the multi-modal data to be processed, so as to acquire a prediction result of the multi-modal data to be processed, such as a data classification result of the multi-modal data to be processed. The construction method of the data processing model based on multi-modal fusion is described in the foregoing description, and will not be described herein again.

[0164] Correspondingly, the embodiment of the application further provides a data processing apparatus based on multi-modal fusion.

[0165] FIG. 6 shows a structural diagram of a data processing apparatus based on multi-modal fusion in the embodiment of the application. Referring to FIG. 6, a data processing apparatus 600 based on multi-modal fusion can include: a data acquisition unit 601 adapted to acquire data to be processed; and an analysis and prediction unit 602 adapted to input the data to be processed into a data processing model based on multi-modal fusion constructed by the construction method of the data processing model based on multi-modal fusion, and acquire a corresponding data processing result.

[0166] The data processing apparatus based on multi-modal fusion in the embodiment can be used to execute the data processing method based on multi-modal fusion as described above, or can also execute the data processing method based on multi-modal fusion as described above by using other functional structures. For the data processing apparatus based on multi-modal fusion, please refer to the content of the data processing method based on multi-modal fusion as described above, which will not be repeated here.

[0167] In addition, in order to improve the training effect of the model, the quantity of data is required to be higher and higher. However, due to the limitation of the number of actual detection of the to-be-detected sample, it is difficult to obtain diverse data (for example, the quantity of mass spectrum data obtained by detecting a single strain in a hospital is usually a single digit), so that the quantity of the actually collected data is difficult to meet the demand of the model training.

[0168] To solve the technical problem, the embodiment of the present application further provides a data processing method. Referring to FIG. 7, a flowchart of an embodiment of the data processing method of the present application is shown.

[0169] In the embodiment, the data processing method comprises the following basic steps:

[0170] Step S10: obtaining original mass spectrum data, wherein the original mass spectrum data comprises signal intensity varying with mass-to-charge ratio;

[0171] Step S20: obtaining Gaussian distribution of the signal intensity;

[0172] Step S30: data sampling is performed on the Gaussian distribution of the signal intensity, and based on the sampling result, the distribution of the signal intensity of the original mass spectrum data is adjusted to different degrees to obtain a plurality of different new mass spectrum data corresponding to the original mass spectrum data.

[0173] The embodiment of the present application utilizes the Gaussian distribution of the signal intensity, and through the sampling manner, based on the sampling result, the distribution of the signal intensity of the original mass spectrum data is adjusted to different degrees to obtain a plurality of different new mass spectrum data corresponding to the original mass spectrum data. Since the Gaussian distribution of the signal intensity can represent the normal fluctuation rule of the signal intensity, the signal intensity is sampled based on the Gaussian distribution of the signal intensity, so that the reliability of the new mass spectrum data is higher, and the data expansion of the original mass spectrum data is realized, so that the quantity of the data is increased while the diversity of the data is improved.

[0174] In order to make the above-mentioned purposes, features and advantages of the present application more obvious and easy to understand, the specific embodiments of the present application will be described in detail below with reference to the accompanying drawings.

[0175] With reference to FIG. 7 and FIG. 8, FIG. 8 is a schematic diagram of an embodiment of the original mass spectrum data, and the original mass spectrum data comprises signal intensity varying with mass-to-charge ratio, and the step S10 is performed to obtain the original mass spectrum data.

[0176] Subsequently, the original mass spectrum data is adjusted to different degrees to obtain adjusted signal intensity, and new mass spectrum data different from the original mass spectrum data is obtained based on the adjusted signal intensity, so that more new mass spectrum data is obtained based on the original mass spectrum data, and data augmentation is realized.

[0177] It should be noted that in the mass spectrum data, the horizontal coordinate of the mass spectrum data is the mass-to-charge ratio (m / z), and the vertical coordinate is the signal intensity.

[0178] The mass spectrum data is: the sample is dissociated, and different fragmentation mass spectrum characteristics are formed according to the mechanism that the flight time is different according to the different molecular sizes.

[0179] Therefore, the mass spectrum data is augmented by adjusting the distribution of the signal intensity to different degrees. Since the test of the mass spectrum data usually has certain volatility, the data processing method of the embodiment can generate fitting mass spectrum data caused by signal intensity volatility.

[0180] It should be further noted that the original mass spectrum data is obtained by detecting a preset number of measured samples.

[0181] As shown in FIG. 8, the horizontal coordinate in FIG. 8 represents the mass-to-charge ratio, and the vertical coordinate represents the signal intensity. In the embodiment, the signal intensity of the mass spectrum data is the peak intensity. For example, the measured sample can be a pathogenic bacterium, and the mass spectrum data is obtained by detecting the pathogenic bacterium.

[0182] Mass spectrometry is a technology for analyzing the chemical composition of molecules in a sample to be tested. By detecting the sample to be tested by a mass spectrometer, the mass-to-charge ratio (m / z) of the molecule and the intensity at each mass-to-charge ratio can be obtained. Therefore, the mass spectrum data comprises intensity varying with mass-to-charge ratio, and the mass spectrum data is commonly used in the fields of medicine, biology, environment and food science.

[0183] Correspondingly, the mass-to-charge ratio of the mass spectrum data is the position of the mass-to-charge ratio corresponding to each molecule, and the signal intensity is the intensity at the mass-to-charge ratio.

[0184] To further improve the reliability of sample analysis, a way of fusing or splicing mass spectrum data with data of other modalities (for example, Raman data) is gradually adopted. For example, when a machine learning model or a deep learning model is used for sample analysis, the mass spectrum data and the Raman data of the sample to be tested can be input into the model, and the sample analysis can be realized through the fused data or spliced data of the two, thereby improving the accuracy of sample analysis.

[0185] Specifically, the original mass spectrum data is obtained by detecting a sample to be tested by a detection device. For example, the detection device can be a mass spectrometer.

[0186] It should be noted that the number of original mass spectrum data can be one or multiple.

[0187] With reference back to FIG. 1, step S20 is performed to obtain a Gaussian distribution of the signal intensity.

[0188] The Gaussian distribution of the signal intensity can represent the normal fluctuation rule of the signal intensity, and therefore, the Gaussian distribution of the signal intensity is obtained first, so that the signal intensity can be sampled based on the Gaussian distribution subsequently, thereby generating simulated mass spectrum data as new mass spectrum data.

[0189] Moreover, sampling the signal intensity based on the Gaussian distribution of the signal intensity makes the reliability of the new mass spectrum data higher.

[0190] In this embodiment, in the step of obtaining the Gaussian distribution of the signal intensity, the Gaussian distribution of the signal intensity that can be detected by a detection device is obtained, and the detection device is used to obtain the original mass spectrum data.

[0191] When the detection device detects a sample to be tested, the detected signal intensity usually fluctuates to a certain extent (for example, the fluctuation of the detection result of the detection device can be caused by human operation, sample preparation, or the detection device itself), and the random distribution of the signal intensity of the statistical mass spectrum data satisfies the Gaussian distribution.

[0192] Moreover, the fluctuation of the detection result of the detection device usually has an approximate or same influence between different mass-to-charge ratios of the same sample to be tested, that is, it can be considered that the same Gaussian distribution is applicable to different mass-to-charge ratios, or it has an approximate or same influence between different samples to be tested, that is, it can be considered that the same Gaussian distribution is applicable to different samples to be tested, and therefore, selecting the Gaussian distribution of the signal intensity that can be detected by the detection device is beneficial to improve the universality of the Gaussian distribution, and accordingly, compared with a scheme of setting an independent Gaussian distribution for each mass-to-charge ratio, this embodiment is beneficial to reduce the complexity of the data processing method.

[0193] In addition, when the model is trained by using the newly added mass spectrum data, feature extraction is usually performed, and the Gaussian distribution of signal intensity used to represent the stability of the detection result of the detection equipment indicates that the fluctuation cannot reflect the characteristics of the measured sample itself, and the fluctuation is caused by the detection equipment, which is beneficial to reducing the attention to the fluctuation caused by the detection equipment during model training. Therefore, the model trained by using the newly added mass spectrum data obtained by using the data processing method can increase the training samples while reducing the probability of negatively affecting the training effect, so that the trained model is more accurate.

[0194] Specifically, in combination with reference to FIG. 9, FIG. 9 is a schematic diagram of the Gaussian distribution of signal intensity according to an embodiment of the present application. In this embodiment, the Gaussian distribution is the Gaussian distribution of signal intensity offset percentage, and the signal intensity offset percentage refers to the offset percentage of signal intensity relative to the signal intensity of the original mass spectrum data.

[0195] As shown in FIG. 9, the abscissa in FIG. 9 represents the signal intensity offset percentage, and the ordinate represents the number. The higher the number of the ordinate is, the higher the probability corresponding to the offset percentage is.

[0196] The offset percentage is used to represent the relative intensity, so that the same Gaussian distribution data (i.e., signal intensity offset percentage) can be shared by different mass-to-charge ratios or different measured samples during subsequent sampling, thereby reducing the complexity of the data processing.

[0197] It should be noted that before the distribution of the signal intensity of the original mass spectrum data is adjusted to different degrees, step S25 is performed to perform first data augmentation on the original mass spectrum data by copying the data to obtain a plurality of copied data corresponding to the original mass spectrum data.

[0198] The copied data is obtained by copying the original mass spectrum data, that is, the copied data is the same as the original mass spectrum data corresponding thereto, so that the distribution of the signal intensity of the copied data is adjusted to different degrees subsequently, which is equivalent to adjusting the distribution of the signal intensity of the original mass spectrum data to different degrees.

[0199] It should be noted that by copying the data, the plurality of sampling results can be used to adjust the respective copied data simultaneously after the data sampling, thereby improving the efficiency of the data processing.

[0200] It can be understood that in other embodiments, the step of copying the data can not be performed, and the distribution of the signal intensity of the original mass spectrum data can be adjusted by using different sampling results, and the adjusted data can be stored as newly added mass spectrum data.

[0201] It should be further explained that in the case of multiple original mass spectrum data, the number of replicated data corresponding to each original mass spectrum data can be determined according to data requirements.

[0202] In the embodiment, in the first data amplification of the original mass spectrum data by the replication data, the higher the quality of the original mass spectrum data is, the more the number of the corresponding replicated data is.

[0203] The higher the quality of the original mass spectrum data is, the higher the accuracy of the original mass spectrum data is, and accordingly, the more the features embodied are, and the greater the effect on data analysis is (for example, it is beneficial to realize data calculation or classification). Therefore, the more the number of data amplification is for the original mass spectrum data with higher quality, the more beneficial it is to improve the reliability of the result of data analysis, and accordingly, when training the model, increasing the number of reliable training samples is beneficial to improve the training effect of the model, so that the trained model is more accurate.

[0204] As an example, the number of replicated data corresponding to each original mass spectrum data includes: setting a weight corresponding to each original mass spectrum data based on the quality of the original mass spectrum data, and the higher the quality of the original mass spectrum data is, the greater the corresponding weight is; and obtaining the number of replicated data corresponding to each original mass spectrum data based on the target total number of mass spectrum data and the weight of each original mass spectrum data.

[0205] The weight is set based on the quality, so as to increase the proportion of high-quality mass spectrum data in the whole data, that is, to improve the effectiveness of data information.

[0206] In the embodiment, the way of judging the quality of the original mass spectrum data includes: comparing the original mass spectrum data with standard mass spectrum data, and the higher the matching degree is, the higher the quality of the original mass spectrum data is.

[0207] The matching degree with the standard mass spectrum data is used as the evaluation standard, which reduces the complexity of quality judgment.

[0208] Specifically, the way of comparing the original mass spectrum data with the standard mass spectrum data includes: comparing the position distribution of the mass-to-charge ratio of the original mass spectrum data with the position distribution of the mass-to-charge ratio of the standard mass spectrum data, and the higher the matching degree of the position distribution is, the higher the quality of the original mass spectrum data is.

[0209] The signal intensity obtained by testing different samples is often different, and the signal intensity has little significance in evaluating the quality. In a single original mass spectrum, the position of each mass-to-charge ratio should be fixed in theory, and therefore, the quality of the original mass spectrum data is comprehensively evaluated by the position distribution of the mass-to-charge ratio to obtain the overall quality of the single original mass spectrum.

[0210] It should be noted that in other embodiments, other quality evaluation methods can also be selected according to the specific type of the original mass spectrum data and other conditions or actual needs.

[0211] Referring to FIG. 7, step S30 is performed to sample the Gaussian distribution of the signal intensity, and based on the sampling result, different degrees of adjustment are implemented on the distribution of the signal intensity of the original mass spectrum data, thereby obtaining a plurality of different new mass spectrum data corresponding to the original mass spectrum data.

[0212] Since the Gaussian distribution of the signal intensity can represent the normal fluctuation rule of the signal intensity, sampling the signal intensity based on the Gaussian distribution of the signal intensity makes the reliability of the new mass spectrum data higher, and realizes data augmentation of the original mass spectrum data, thereby increasing the quantity of data while improving the diversity of data. Accordingly, model training based on more diverse data is beneficial to improving the accuracy of the model.

[0213] In addition, by augmenting the original mass spectrum data, when the original mass spectrum data needs to be fused or spliced with data of other modalities in actual operation, the data quantity of the mass spectrum data of the embodiment and the data of other modalities can be matched, thereby being compatible with the data of other modalities.

[0214] Since the original mass spectrum data is obtained by copying data first, sampling the Gaussian distribution of the signal intensity and implementing different degrees of adjustment on the distribution of the signal intensity of the original mass spectrum data based on the sampling result includes: adjusting the signal intensity of the plurality of copied data based on different sampling results, respectively, to realize second data augmentation, thereby obtaining a plurality of different new mass spectrum data.

[0215] That is, for a single original mass spectrum data, a plurality of new mass spectrum data different from the original mass spectrum data are obtained through second data augmentation, and the new mass spectrum data are also different from each other.

[0216] In the embodiment, in the process of sampling the Gaussian distribution of the signal intensity and implementing different degrees of adjustment on the distribution of the signal intensity of the original mass spectrum data based on the sampling result, an offset percentage is selected for each mass-to-charge ratio from the Gaussian distribution of the offset percentage, and the signal intensity corresponding to the mass-to-charge ratio in the original mass spectrum data is adjusted based on the selected offset percentage.

[0217] In the case that the quantity of the original mass spectrum data is multiple, since each original mass spectrum data shares the same Gaussian distribution data, and the Gaussian distribution is the Gaussian distribution of signal intensity offset percentage, for each mass-to-charge ratio in each original mass spectrum data, the original value of the signal intensity of the mass-to-charge ratio can be adjusted based on the selected offset percentage of the sampling.

[0218] For example, for each mass-to-charge ratio, the signal intensity can be adjusted by formula I2=I1*(1+m%);wherein I2 represents the signal intensity of the mass-to-charge ratio in the new mass spectrum data, I1 represents the signal intensity of the mass-to-charge ratio in the original mass spectrum data, and m% represents the offset percentage of the signal intensity, which can be zero, positive or negative. That is, for each mass-to-charge ratio, after determining the signal intensity offset percentage, the original signal intensity of the mass-to-charge ratio is adjusted to obtain the new signal intensity.

[0219] It can be seen that although the signal intensities of different mass-to-charge ratios can be different, and the signal intensity distributions of different measured samples can also be different, by using the Gaussian distribution of relative intensity, the original values of the signal intensities of each mass-to-charge ratio can be adjusted by the same calculation method for different original mass spectrum data and different mass-to-charge ratios, thereby reducing the complexity of data expansion.

[0220] In this embodiment, in the same new mass spectrum data, the adjustment degrees of the signal intensities of each mass-to-charge ratio relative to the signal intensities of the mass-to-charge ratio in the original mass spectrum data are different.

[0221] Specifically, since the original mass spectrum data is obtained by copying data, the signal intensities of the several copied data are adjusted based on different sampling results, and the signal intensities of each mass-to-charge ratio in the same copied data are adjusted to different degrees.

[0222] By making the change degrees of the signal intensities of each mass-to-charge ratio in the same new mass spectrum data relative to the original values different, it is beneficial to obtain more diverse data. Moreover, when the data obtained by the method of this embodiment is used for data analysis, even if the data is normalized, the diversity of the normalized data can still be maintained.

[0223] For example, for a certain mass-to-charge ratio, the offset of the signal intensity in the new mass spectrum data relative to the signal intensity in the original mass spectrum data is +3%(i.e. the increase is 3%), and for another mass-to-charge ratio, the offset of the signal intensity in the new mass spectrum data relative to the signal intensity in the original mass spectrum data is -2%(i.e. the decrease is 2%).

[0224] Referring to FIG. 10, FIG. 10 is a schematic diagram of the newly added mass spectrum data according to an embodiment of the present application, and in order to show the difference between the original mass spectrum data and the newly added mass spectrum data, in FIG. 4, the newly added mass spectrum data and the original mass spectrum data in FIG. 2 are shown in the same coordinate.

[0225] In FIG. 10, the data with darker color represents the original mass spectrum data, and the data with lighter color represents the newly added mass spectrum data. Therefore, by performing data augmentation, the newly added mass spectrum data different from the original mass spectrum data can be obtained.

[0226] Referring to FIG. 11, FIG. 11 shows t-SNE distribution diagrams obtained based on mass spectrum data, wherein FIG. 11(a) shows a t-SNE distribution diagram of an embodiment obtained based on original mass spectrum data, and FIG. 11(b) shows a t-SNE distribution diagram of an embodiment obtained based on mass spectrum data obtained by the method according to the embodiment.

[0227] t-SNE (t-Distributed Stochastic Neighbor Embedding) is an unsupervised nonlinear technique, which is mainly used to realize dimension reduction of data, and make the data in the low-dimensional space still carry the information carried in the high-dimensional space.

[0228] As shown in FIG. 11, taking the analysis of pathogenic bacteria as an example, each point represents a data, and the data points in the same circle, oval, square or trapezoidal circle belong to the same pathogenic bacteria category, for example, the pathogenic bacteria include M. abscessus, M. fortuitum, M. ulcerans, M. peregrinum, M. phlei and M. chelonae.

[0229] It should be noted that, in order to facilitate illustration, in FIG. 11, the points in the circle represent M. chelonae, the points in the square circle represent M. fortuitum, the points in the trapezoidal circle represent M. abscessus, the points in the solid line oval circle represent M. ulcerans, the points in the dashed line oval circle represent M. peregrinum, and the points in the double-dot chain line oval circle represent M. phlei.

[0230] As can be seen from FIG. 11(a), the t-SNE distribution diagram obtained by analyzing the sample through the original mass spectrum data shows that the number of data points in the same pathogen category is small. As can be seen from FIG. 11(b), the sample analysis through the mass spectrum data obtained by the method described in the embodiment shows that the number of data points in the same pathogen category increases and is stably distributed, the distribution of the data points does not appear to be severely deformed, and the classification can still be well distinguished.

[0231] Referring to FIG. 12, FIG. 12 is a comparison diagram of accuracy confusion matrices obtained based on machine learning models, and the deeper the color of the grid where the number is located, the higher the accuracy.

[0232] FIG. 12(a) shows an accuracy confusion matrix of an embodiment of a machine learning model trained based on original mass spectrum data, and FIG. 12(b) shows an accuracy confusion matrix of an embodiment of a machine learning model trained based on mass spectrum data obtained by the data processing method described in the embodiment.

[0233] As an example, the model is used for pathogen analysis, and in the confusion matrix, True represents the true result, and Predicted represents the predicted result. The types of pathogenic bacteria include M. abscessus, M. fortuitum, M. ulcerans, M. peregrinum, M. phlei, and M. chelonae.

[0234] For example, taking M. chelonae as True, when the machine learning model trained based on the original mass spectrum data is used for analysis, the probability of predicting the accurate result is 0.83 (i.e., 83%), and the probability of predicting M. ulcerans is 0.17 (i.e., 17%).

[0235] As shown in FIG. 12(a), after training based on the original mass spectrum data, the overall prediction accuracy of the machine learning model for the pathogen category is 69.40%, and as shown in FIG. 12(b), after training based on the mass spectrum data obtained by the data processing method described in the embodiment, the overall prediction accuracy of the machine learning model for the pathogen category is 78.43%, which can obtain a higher prediction accuracy.

[0236] Referring to FIG. 13, FIG. 13 is a comparison diagram of accuracy confusion matrices obtained based on deep learning models.

[0237] Fig. 13(a) shows an accuracy confusion matrix of an embodiment of the deep learning model trained based on original mass spectrum data, and Fig. 13(b) shows an accuracy confusion matrix of an embodiment of the deep learning model trained based on mass spectrum data obtained by the data processing method.

[0238] As an example, the model is used for pathogenic bacteria analysis, and in the confusion matrix, True represents the true result, and Predicted represents the predicted result.

[0239] As shown in Fig. 13(a), after training based on original mass spectrum data, the overall prediction accuracy of the deep learning model for pathogenic bacteria categories is 75.38%, and as shown in Fig. 13(b), after training based on mass spectrum data obtained by the data processing method, the overall prediction accuracy of the deep learning model for pathogenic bacteria categories is 90.33%, which can obtain higher prediction accuracy.

[0240] Correspondingly, the present application also provides a data processing system. Fig. 14 is a structural schematic diagram of an embodiment of the data processing system of the present application.

[0241] Referring to Fig. 14, and in combination with Figs. 8-13, the data processing system comprises: an original data acquisition module 10 configured to acquire original mass spectrum data, the original mass spectrum data comprising signal intensity varying with mass-to-charge ratio; a Gaussian distribution acquisition module 20 configured to acquire Gaussian distribution of the signal intensity; and a data augmentation module 30 configured to sample the Gaussian distribution of the signal intensity, and based on the sampling result, to adjust the distribution of the signal intensity of the original mass spectrum data to different degrees to obtain a plurality of different new mass spectrum data corresponding to the original mass spectrum data.

[0242] The data processing system adjusts the distribution of the signal intensity of the original mass spectrum data to different degrees to obtain adjusted signal intensity, and accordingly obtains new mass spectrum data different from the original mass spectrum data, thereby obtaining more new mass spectrum data based on the original mass spectrum data, and realizing data augmentation.

[0243] Since the Gaussian distribution of the signal intensity can represent the normal fluctuation rule of the signal intensity, sampling the signal intensity based on the Gaussian distribution of the signal intensity makes the reliability of the new mass spectrum data higher, and in addition, realizes data augmentation of the original mass spectrum data, thereby increasing the number of data while improving the diversity of the data; correspondingly, model training based on more diverse data is beneficial to improving the accuracy of the model.

[0244] In addition, by data augmentation on the original mass spectrum data, in actual operation, when the original mass spectrum data needs to be fused or spliced with data of other modalities, the data amount of the mass spectrum data of the embodiment can be matched with the data of other modalities, so as to be compatible with the data of other modalities.

[0245] It should be noted that in the mass spectrum data, the abscissa of the mass spectrum data is the mass-to-charge ratio, and the ordinate is the signal intensity.

[0246] It should also be noted that the original mass spectrum data is obtained by detecting a preset number of measured samples.

[0247] As shown in FIG. 8, the abscissa in FIG. 8 represents the position of the mass-to-charge ratio, and the ordinate represents the signal intensity. In the embodiment, the signal intensity of the mass spectrum data is the peak intensity. For example, the measured sample can be a pathogenic bacterium, and the mass spectrum data is obtained by detecting the pathogenic bacterium.

[0248] Mass spectrometry is a technique for analyzing the chemical composition of molecules in a sample to be tested. By detecting the measured sample by a mass spectrometer, the mass-to-charge ratio (m / z) of the molecule and the intensity at each mass-to-charge ratio can be obtained. Therefore, the mass spectrum data contains the intensity varying with the mass-to-charge ratio, and the mass spectrum data is commonly used in the fields of medicine, biology, environment and food science.

[0249] Correspondingly, the mass-to-charge ratio of the mass spectrum data is the position of the mass-to-charge ratio corresponding to each molecule, and the signal intensity is the intensity at the mass-to-charge ratio.

[0250] In order to further improve the reliability of sample analysis, the mass spectrum data is gradually fused or spliced with data of other modalities (for example, Raman data). For example, when a machine learning model or a deep learning model is used for sample analysis, the mass spectrum data and the Raman data of the sample to be tested can be input into the model, and the sample analysis can be realized by the fused data or the spliced data of the two, thereby improving the accuracy of sample analysis.

[0251] Specifically, the original mass spectrum data is obtained by detecting the measured sample by a detection device. For example, the detection device can be a mass spectrometer.

[0252] The Gaussian distribution of the signal intensity can represent the normal fluctuation rule of the signal intensity, so as to obtain the Gaussian distribution of the signal intensity, so that the data augmentation module 30 samples the signal intensity based on the Gaussian distribution, thereby generating simulated mass spectrum data as new mass spectrum data.

[0253] Moreover, sampling the signal intensity based on the Gaussian distribution of the signal intensity makes the reliability of the new mass spectrum data higher.

[0254] In this embodiment, the Gaussian distribution obtaining module 20 is configured to obtain a Gaussian distribution of signal intensity that can be detected by the detection device used to obtain the original mass spectrum data.

[0255] When the detection device detects the sample to be detected, the detected signal intensity usually fluctuates (for example, the detection result of the detection device fluctuates, which can be caused by human operation, sample preparation, or the detection device itself), and the random distribution of the signal intensity of the mass spectrum data usually satisfies a Gaussian distribution.

[0256] Moreover, the fluctuation of the detection result of the detection device usually has an approximate or same effect on different mass-to-charge ratios of the same sample to be detected, that is, it can be considered that the same Gaussian distribution is applicable to different mass-to-charge ratios, or has an approximate or same effect on different samples to be detected, that is, it can be considered that the same Gaussian distribution is applicable to different samples to be detected. Therefore, selecting the Gaussian distribution of the signal intensity that can be detected by the detection device is beneficial to improve the universality of the Gaussian distribution, and accordingly, compared with the scheme of setting an independent Gaussian distribution for each mass-to-charge ratio, the embodiment is beneficial to reduce the complexity of the data processing method.

[0257] In addition, when the newly added mass spectrum data is used for model training, feature extraction is usually performed, and the Gaussian distribution of the signal intensity used to represent the stability of the detection result of the detection device indicates that the fluctuation cannot reflect the characteristics of the sample to be detected, and the fluctuation is caused by the detection device. This is beneficial to reduce the attention to the fluctuation caused by the detection device during model training. Therefore, the newly added mass spectrum data obtained by using the data processing method of the embodiment can increase the training samples while reducing the probability of negatively affecting the training effect, so that the trained model is more accurate.

[0258] In combination with reference to FIG. 9, in this embodiment, the Gaussian distribution is a Gaussian distribution of signal intensity offset percentage, and the signal intensity offset percentage refers to the offset percentage of the signal intensity relative to the signal intensity of the original mass spectrum data. As shown in FIG. 9, the abscissa in FIG. 9 represents the signal intensity offset percentage, the ordinate represents the number, and the higher the number of the ordinate, the higher the probability corresponding to the offset percentage.

[0259] The offset percentage is used to represent the relative intensity, so that the same Gaussian distribution data (that is, the offset percentage of the signal intensity) can be shared by different mass-to-charge ratios or different samples to be detected during subsequent sampling, thereby reducing the complexity of the data processing.

[0260] It should be noted that the data processing system further comprises a data replication module 25 configured to replicate the original mass spectrum data to obtain a plurality of replicated data corresponding to the original mass spectrum data before adjusting the distribution of the signal intensity of the original mass spectrum data to different degrees.

[0261] The replicated data is obtained by replicating the original mass spectrum data, that is, the replicated data is the same as the original mass spectrum data, and thus adjusting the distribution of the signal intensity of the replicated data to different degrees is equivalent to adjusting the distribution of the signal intensity of the original mass spectrum data to different degrees.

[0262] It should be noted that by replicating the data, after subsequent data sampling, it is convenient to adjust each replicated data using a plurality of sampling results, thereby improving the efficiency of data processing.

[0263] It can be understood that in other embodiments, the data replication module can be omitted, and the data expansion module 30 adjusts the distribution of the signal intensity of the original mass spectrum data to different degrees using different sampling results, and stores the adjusted data as new mass spectrum data.

[0264] It should be further noted that in the case where the number of original mass spectrum data is multiple, the number of replicated data corresponding to each original mass spectrum data can be determined according to data requirements.

[0265] In this embodiment, the higher the quality of the original mass spectrum data, the more the number of corresponding replicated data.

[0266] The higher the quality of the original mass spectrum data, the higher the accuracy of the original mass spectrum data, and accordingly, the more features it embodies, and the greater the effect on data analysis (for example, it is beneficial to realize data calculation or classification). Therefore, the more the number of data expansion for the original mass spectrum data with higher quality, the more beneficial to improve the reliability of data analysis results, and accordingly, the number of reliable training samples can be increased when training the model, which is beneficial to improve the training effect of the model, so that the trained model is more accurate.

[0267] As an example, the data replication module 25 comprises a weight setting unit configured to set a weight corresponding to each original mass spectrum data based on the quality of the original mass spectrum data, and the higher the quality of the original mass spectrum data, the greater the corresponding weight; and a number setting unit configured to obtain the number of replicated data corresponding to each original mass spectrum data based on the target total number of mass spectrum data and the weight of each original mass spectrum data.

[0268] The weight is set based on the quality, so as to increase the proportion of high-quality mass spectrum data in the whole data, that is, to improve the effectiveness of data information.

[0269] In the embodiment, the way of judging the quality of the original mass spectrum data includes: comparing the original mass spectrum data with standard mass spectrum data, and the higher the matching degree is, the higher the quality of the original mass spectrum data is.

[0270] The matching degree with the standard mass spectrum data is used as the evaluation standard, so as to reduce the complexity of quality judgment.

[0271] Specifically, the way of comparing the original mass spectrum data with the standard mass spectrum data includes: comparing the position distribution of the mass-to-charge ratio of the original mass spectrum data with the position distribution of the mass-to-charge ratio of the standard mass spectrum data, and the higher the matching degree of the position distribution is, the higher the quality of the original mass spectrum data is.

[0272] The signal intensity obtained by testing different samples is often different, and the signal intensity has little significance for the quality evaluation. In a single original mass spectrum, the position of each mass-to-charge ratio should be fixed in theory. Therefore, the quality of the original mass spectrum data is comprehensively evaluated by the position distribution of the mass-to-charge ratio, so as to obtain the overall quality of a single original mass spectrum.

[0273] It should be noted that in other embodiments, according to the specific type of the original mass spectrum data and other conditions, or actual needs, the quality of the original mass spectrum data can also be obtained by other quality evaluation methods.

[0274] In the embodiment, since the several copied data corresponding to the original mass spectrum data are obtained by copying the data, the data expansion module 30 adjusts the signal intensity of the several copied data based on different sampling results, so as to realize the second data expansion and obtain a plurality of different new mass spectrum data.

[0275] That is, for a single original mass spectrum data, a plurality of new mass spectrum data different from the original mass spectrum data are obtained by the second data expansion, and the new mass spectrum data are also different from each other.

[0276] In the case where the number of original mass spectrum data is multiple, since each original mass spectrum data shares the same Gaussian distribution data, and the Gaussian distribution is the Gaussian distribution of the signal intensity offset percentage, for each mass-to-charge ratio in each original mass spectrum data, the original value of the signal intensity of the mass-to-charge ratio can be adjusted based on the selected offset percentage of the sampled data.

[0277] For example, for each mass-to-charge ratio, the data amplification module 30 adjusts the signal intensity based on the formula I2=I1*(1+m%) according to the formula I2=I1*(1+m%), wherein I2 represents the signal intensity of the mass-to-charge ratio in the new mass spectrum data, I1 represents the signal intensity of the mass-to-charge ratio in the original mass spectrum data, and m% represents the offset percentage of the signal intensity, which can be zero, positive or negative. That is, for each mass-to-charge ratio, after determining the signal intensity offset percentage, the original signal intensity of the mass-to-charge ratio is adjusted to obtain the new signal intensity.

[0278] As can be seen, although the signal intensities of different mass-to-charge ratios can be different, and the signal intensity distributions of different samples to be measured can also be different, by using the Gaussian distribution of the relative intensity, the original values of the signal intensities of the mass-to-charge ratios can be adjusted in the same way for different original mass spectrum data and different mass-to-charge ratios, thereby reducing the complexity of data amplification.

[0279] In this embodiment, in the same new mass spectrum data, the signal intensities of the mass-to-charge ratios are adjusted to different degrees relative to the signal intensities of the mass-to-charge ratios in the original mass spectrum data.

[0280] Specifically, since the original mass spectrum data is obtained by copying data, the data amplification module 30 adjusts the signal intensities of the mass-to-charge ratios in the same copy data to different degrees.

[0281] By making the signal intensities of the mass-to-charge ratios in the same new mass spectrum data change to different degrees relative to the original values, more diverse data can be obtained. Moreover, when the data obtained by the method of this embodiment is subjected to data analysis, even if the data is subjected to normalization processing, the diversity of the data after normalization processing can still be maintained.

[0282] For example, for a certain mass-to-charge ratio, the offset of the signal intensity in the new mass spectrum data relative to the signal intensity in the original mass spectrum data is +3% (i.e., the increase is 3%), and for another mass-to-charge ratio, the offset of the signal intensity in the new mass spectrum data relative to the signal intensity in the original mass spectrum data is -2% (i.e., the decrease is 2%).

[0283] Referring to FIG. 10, FIG. 10 is a schematic diagram of new mass spectrum data according to an embodiment of the present application, and in order to show the difference between the original mass spectrum data and the new mass spectrum data, in FIG. 4, the new mass spectrum data and the original mass spectrum data in FIG. 2 are shown in the same coordinate.

[0284] In Figure 10, the darker-colored data represents the original mass spectrometry data, and the lighter-colored data represents newly added mass spectrometry data. Therefore, by performing data amplification, new mass spectrometry data different from the original mass spectrometry data can be obtained.

[0285] Referring to Figure 11, Figure 11 shows a t-SNE distribution map obtained based on mass spectrometry data. Figure 11(a) shows a t-SNE distribution map of an embodiment obtained based on the original mass spectrometry data, and Figure 11(b) shows a t-SNE distribution map of an embodiment obtained based on the mass spectrometry data obtained by the method described in this embodiment.

[0286] As shown in Figure 11, taking the analysis of pathogens as an example, each point represents a data point. Data points in the same circle, ellipse, square, or trapezoid belong to the same pathogen category. For example, pathogens include: Mycobacterium abscessus, Mycobacterium fortuitum, Mycobacterium ulcerans, Mycobacterium peregrinum, Mycobacterium phlei, and Mycobacterium chelonae.

[0287] It should be noted that, for ease of illustration, in Figure 11, the dots in the circles represent Mycobacterium chelonae, the dots in the square circles represent Mycobacterium fortuitum, the dots in the trapezoidal circles represent Mycobacterium abscessus, the dots in the solid-lined elliptical circles represent Mycobacterium ulcerans, the dots in the dashed-lined elliptical circles represent Mycobacterium peregrinum, and the dots in the double-dotted-lined elliptical circles represent Mycobacterium phlei.

[0288] As can be seen from Figure 11(a), in the t-SNE distribution map obtained by analyzing samples using only the original mass spectrometry data, the number of data points under the same pathogen category is relatively small. As can be seen from Figure 11(b), when the mass spectrometry data obtained by the method described in this embodiment is used for sample analysis, the number of data points under the same pathogen category increases and the distribution is stable. Moreover, the distribution of data points does not show serious distortion and can still achieve good distinction between categories.

[0289] Referring to Figure 12, which is a comparative diagram of the accuracy confusion matrix obtained based on the machine learning model, the darker the color of the grid containing the number, the higher the accuracy.

[0290] Fig. 12(a) is an accuracy confusion matrix of an embodiment of the machine learning model trained based on the original mass spectrum data, and Fig. 12(b) is an accuracy confusion matrix of an embodiment of the machine learning model trained based on the mass spectrum data obtained by the data processing method.

[0291] As an example, the model is used for pathogenic bacteria analysis, and in the confusion matrix, True represents the true result, and Predicted represents the predicted result. The types of pathogenic bacteria include M. abscessus, M. fortuitum, M. ulcerans, M. peregrinum, M. phlei, and M. chelonae.

[0292] For example, taking M. chelonae as True, when the machine learning model trained based on the original mass spectrum data is used for analysis, the probability of predicting the correct result is 0.83 (i.e., 83%), and the probability of predicting M. ulcerans is 0.17 (i.e., 17%).

[0293] As shown in Fig. 12(a), after training based on the original mass spectrum data, the overall prediction accuracy of the machine learning model for pathogenic bacteria is 69.40%, and as shown in Fig. 12(b), after training based on the mass spectrum data obtained by the data processing method, the overall prediction accuracy of the machine learning model for pathogenic bacteria is 78.43%, and a higher prediction accuracy can be obtained.

[0294] Referring to Fig. 13, Fig. 13 is a comparison diagram of accuracy confusion matrices based on deep learning models.

[0295] Fig. 13(a) is an accuracy confusion matrix of an embodiment of the deep learning model trained based on the original mass spectrum data, and Fig. 13(b) is an accuracy confusion matrix of an embodiment of the deep learning model trained based on the mass spectrum data obtained by the data processing method.

[0296] As an example, the model is used for pathogenic bacteria analysis, and in the confusion matrix, True represents the true result, and Predicted represents the predicted result.

[0297] As shown in FIG. 13(a), after training based on the original mass spectrum data, the overall prediction accuracy of the deep learning model for the pathogenic bacteria class is 75.38%, as shown in FIG. 13(b), after training based on the mass spectrum data obtained by the data processing method described in the embodiment, the overall prediction accuracy of the deep learning model for the pathogenic bacteria class is 90.33%, and a higher prediction accuracy can be obtained.

[0298] It should be noted that in the present embodiment, the data processing system is used to implement the data processing method described in the foregoing embodiments, and the specific description of the data processing system can be combined with the related description in the foregoing embodiments.

[0299] Correspondingly, the embodiment of the present application also provides a device comprising at least one memory and at least one processor, the memory stores one or more computer instructions, wherein the one or more computer instructions are executed by the processor to implement the construction method of the data processing model based on multi-modal fusion or the data processing method based on multi-modal fusion. Wherein, the construction method of the data processing model based on multi-modal fusion or the data processing method based on multi-modal fusion is described in the foregoing part, and will not be repeated here.

[0300] Correspondingly, the embodiment of the present application also provides a storage medium, the storage medium stores one or more computer instructions, the one or more computer instructions are used to implement the construction method of the data processing model based on multi-modal fusion or the data processing method based on multi-modal fusion. Wherein, the construction method of the data processing model based on multi-modal fusion or the data processing method based on multi-modal fusion or the data processing method is described in the foregoing part, and will not be repeated here.

[0301] Correspondingly, the embodiment of the present application also provides a device, which can load the construction method of the data processing model based on multi-modal fusion or the data processing method based on multi-modal fusion in the form of program to implement the construction method of the data processing model based on multi-modal fusion or the data processing method based on multi-modal fusion provided by the embodiment of the present application.

[0302] Referring to FIG. 15, a hardware structure diagram of the device provided by an embodiment of the present application is shown. The device of the present embodiment comprises at least one processor 01, at least one communication interface 02, at least one memory 03 and at least one communication bus 04.

[0303] In some embodiments, the number of the processor 01, the communication interface 02, the memory 03 and the communication bus 04 is at least one, and the processor 01, the communication interface 02 and the memory 03 complete the communication with each other through the communication bus 04.

[0304] The communication interface 02 can be an interface of a communication module for network communication, for example, an interface of a GSM module.

[0305] The processor 01 can be a central processing unit CPU, or an application specific integrated circuit ASIC, or one or more integrated circuits configured to implement the methods described in the embodiments.

[0306] The memory 03 can include a high-speed RAM memory, and can also include a non-volatile memory, for example, at least one disk memory. The memory 03 stores one or more computer instructions, and the one or more computer instructions are executed by the processor 01 to implement the construction method of the multi-modal fusion based data processing model or the multi-modal fusion based data processing method or the data processing method provided in the foregoing embodiments.

[0307] It should be noted that the implementation device described above can also include other devices (not shown) that can not be necessary for the disclosure of the embodiments of the application; since these other devices can not be necessary for understanding the disclosure of the embodiments of the application, the embodiments of the application do not introduce them one by one.

[0308] The embodiments of the application also provide a storage medium, which stores one or more computer instructions, and the one or more computer instructions are used to implement the construction method of the multi-modal fusion based data processing model or the multi-modal fusion based data processing method or the data processing method provided in the foregoing embodiments.

[0309] The above-described embodiments of the application are combinations of elements and features of the application. Unless otherwise mentioned, elements or features can be considered selective. Each element or feature can be practiced without being combined with other elements or features. In addition, embodiments of the application can be constructed by combining some elements and / or features. The order of the operations described in the embodiments of the application can be rearranged. Some configurations of any embodiment can be included in another embodiment, and can be replaced with corresponding configurations of another embodiment. It is obvious to those skilled in the art that the claims in the appended claims, which have no explicit reference relationship with each other, can be combined into embodiments of the application, or can be included as new claims in modifications after the submission of the application.

[0310] Embodiments of the present application can be implemented in various forms, for example, hardware, firmware, software, or a combination thereof. In a hardware configuration, the method according to exemplary embodiments of the present application can be implemented by one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), processors, controllers, micro-controllers, microprocessors, or the like.

[0311] In a firmware or software configuration, embodiments of the present application can be implemented in the form of modules, procedures, functions, and the like. Software code can be stored in a memory unit and executed by a processor. The memory unit is located at the interior or exterior of the processor and can deliver data to and receive data from the processor via various known means.

[0312] The above description of disclosed embodiments is merely intended to illustrate the principles of the present application. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein can be applied to other embodiments without departing from the spirit or scope of the present application. Thus, the present application is not intended to be limited to the embodiments shown herein but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

[0313] While the present application has been disclosed in connection with the embodiments presented, it should be understood that the application as claimed can be practiced otherwise than as specifically described. Any conceivable modification, variation or alteration to the disclosed embodiments will be apparent to those skilled in the art and the general principles defined herein can be applied to other embodiments without departing from the spirit or scope of the application. Thus, the present application is not intended to be limited to the embodiments shown herein but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A method for constructing a data processing model based on multimodal fusion, characterized in that, include: Obtain a multimodal training dataset, which includes multimodal training data; The multimodal training data in the multimodal training dataset is used for joint training to obtain a data processing model based on multimodal fusion.

2. The method for constructing a data processing model based on multimodal fusion as described in claim 1, characterized in that, The multimodal training dataset includes a first training dataset, which includes one-dimensional training data for multiple first modalities. The method involves jointly training the multimodal training data in the multimodal training dataset to obtain a data processing model based on multimodal fusion. This includes: performing a first fusion process on one-dimensional training data with multiple first modalities in the first training dataset to obtain multiple corresponding first fusion data points, generating the first fusion dataset; performing a first directional attention feature extraction process on the first fusion data in the first fusion dataset to obtain first directional attention feature data corresponding to each first modality; performing a first operation process on the first directional attention feature data corresponding to each first modality and one-dimensional training data of other corresponding first modalities to obtain multiple corresponding first feature data points, generating a first feature dataset; performing a first aggregation process on the first feature data in the first feature dataset to obtain multiple corresponding first aggregated feature data points, generating a first aggregated feature dataset; performing a first classification process on the aggregated feature data in the first aggregated feature dataset to obtain corresponding first prediction results; calculating the first loss value of the data processing model based on multimodal fusion using a preset first loss function based on the obtained first prediction results, and adjusting the weights of the data processing model based on multimodal fusion based on the calculated first loss value until the first loss value of the data processing model based on multimodal fusion converges.

3. The method for constructing a data processing model based on multimodal fusion as described in claim 2, wherein the multimodal training dataset further includes a second training dataset, the second training dataset including two-dimensional training data having one or more second modalities; The method further includes: performing joint training on the multimodal training data in the multimodal training dataset to obtain a data processing model based on multimodal fusion; performing second attention feature extraction processing on the two-dimensional training data in the second training dataset to obtain corresponding second attention feature data; and performing second operation processing on the second attention feature data and the corresponding one-dimensional training data of the first modality to obtain multiple corresponding second feature data and generate a second feature dataset. Performing a first aggregation process on the first feature data in the first feature dataset to obtain multiple corresponding first aggregated feature data and generate a first aggregated feature dataset includes: performing a first aggregation process on the first feature data in the first feature dataset and the corresponding second feature data in the second feature dataset to obtain multiple corresponding first aggregated feature data and generate a first aggregated feature dataset.

4. The method for constructing a data processing model based on multimodal fusion as described in claim 3, characterized in that, After performing a first operation on the first directional attention feature data corresponding to each first modality and the corresponding one-dimensional training data of other first modalities to obtain multiple first feature data and generate a first feature dataset, and performing a first aggregation on the first feature data in the first feature dataset and the corresponding second feature data in the second feature dataset to obtain multiple first aggregated feature data, and before generating the first aggregated feature dataset, performing joint training using the one-dimensional training data in the first training dataset and the two-dimensional training data in the second training dataset to obtain a data processing model based on multimodal fusion, the method further includes: performing a first dimensionality reduction on the first feature data in the first feature dataset; After performing a second operation on the second attention feature data and the corresponding one-dimensional training data of the first modality to obtain multiple corresponding second feature data and generate a second feature dataset, and performing a first aggregation operation on the first feature data in the first feature dataset and the corresponding second feature data in the second feature dataset to obtain multiple corresponding first aggregated feature data, and before generating the first aggregated feature dataset, performing joint training using the one-dimensional training data in the first training dataset and the two-dimensional training data in the second training dataset to obtain a data processing model based on multimodal fusion, the method further includes: performing a second dimensionality reduction operation on the second feature data in the second feature dataset to obtain multiple corresponding second dimensionality reduction feature data; Performing a first aggregation process on the first feature data in the first feature dataset and the corresponding second feature data in the second feature dataset to obtain multiple corresponding first aggregated feature data and generate a first aggregated feature dataset includes: performing a first aggregation process on the first feature data after first dimensionality reduction and the corresponding second feature data after second dimensionality reduction to obtain multiple corresponding first aggregated feature data and generate a first aggregated feature dataset.

5. The method for constructing a data processing model based on multimodal fusion as described in claim 4, characterized in that, The first dimensionality reduction process includes a first convolutional dimensionality reduction process, and the second dimensionality reduction process includes a second convolutional dimensionality reduction process.

6. The method for constructing a data processing model based on multimodal fusion as described in claim 3, characterized in that, The one-dimensional training data includes at least two of the following: Raman spectroscopy data, mass spectrometry data, infrared spectroscopy data, and point scan fluorescence lifetime imaging data. The two-dimensional training data includes at least one of surface scan imaging data and confocal imaging data.

7. The method for constructing a data processing model based on multimodal fusion as described in claim 3, characterized in that, The multimodal training dataset includes a first training dataset and a second training dataset. The first training dataset includes one-dimensional training data with a first modality, and the second training dataset includes two-dimensional training data with one or more second modalities. The data processing model based on multimodal fusion is obtained by jointly training one-dimensional training data from the first training dataset and two-dimensional training data from the second training dataset, including: Perform first self-attention feature extraction processing on the one-dimensional training data in the first training dataset to obtain multiple corresponding first attention feature data, forming a first attention feature dataset; The second self-attention feature extraction process is performed on the two-dimensional training data in the second training dataset to obtain multiple corresponding second attention feature data and generate the second attention feature dataset. A second aggregation process is performed on the first attention feature data in the first attention feature dataset and the second attention feature data in the second attention feature dataset to obtain multiple corresponding second aggregated feature data and generate a second aggregated feature dataset. The second aggregated feature data in the second aggregated feature dataset are subjected to a second classification process to obtain the corresponding second prediction results; Based on the obtained second prediction result, the second loss value of the data processing model based on multimodal fusion is calculated using a preset second loss function, and the weights of the data processing model based on multimodal fusion are adjusted according to the calculated second loss value until the second loss value of the data processing model based on multimodal fusion converges.

8. The method for constructing a data processing model based on multimodal fusion as described in claim 7, characterized in that, After performing a first self-attention feature extraction process on the one-dimensional training data in the first training dataset to obtain multiple corresponding first attention feature data and forming a first attention feature dataset, and performing a second aggregation process on the first attention feature data in the first attention feature dataset and the second attention feature data in the second attention feature dataset to obtain multiple corresponding second aggregated feature data, and before generating the second aggregated feature dataset, performing joint training using the one-dimensional training data in the first training dataset and the two-dimensional training data in the second training dataset to obtain a data processing model based on multimodal fusion, the method further includes: performing a third dimensionality reduction process on the first attention feature data in the first attention feature dataset. The process further includes: performing second self-attention feature extraction on the two-dimensional training data in the second training dataset to obtain multiple corresponding second attention feature data, generating a second attention feature dataset, and performing second aggregation on the first attention feature data in the first attention feature dataset and the second attention feature data in the second attention feature dataset to obtain multiple corresponding second aggregated feature data. Before generating the second aggregated feature dataset, the process further includes: performing joint training on the one-dimensional training data in the first training dataset and the two-dimensional training data in the second training dataset to obtain a data processing model based on multimodal fusion. The process of performing a second aggregation process on the first attention feature data in the first attention feature dataset and the second attention feature data in the second attention feature dataset to obtain multiple corresponding second aggregated feature data and generate a second aggregated feature dataset includes: performing a second aggregation process on the first attention feature data after a third dimensionality reduction process and the corresponding second attention feature data after a fourth dimensionality reduction process to obtain multiple corresponding second aggregated feature data and generate a second aggregated feature dataset.

9. The method for constructing a data processing model based on multimodal fusion as described in claim 8, characterized in that, The third dimensionality reduction process includes a third convolutional dimensionality reduction process, and the fourth dimensionality reduction process includes a fourth convolutional dimensionality reduction process.

10. The method for constructing a data processing model based on multimodal fusion as described in claim 8, characterized in that, The one-dimensional training data includes Raman spectroscopy data, mass spectrometry data, infrared spectroscopy data, or point scan fluorescence lifetime imaging data. The two-dimensional training data includes at least one of surface scan imaging data and confocal imaging data.

11. The method for constructing a data processing model based on multimodal fusion as described in claim 1, characterized in that, The data processing model based on multimodal fusion is used to classify pathogens or particulate matter.

12. A device for constructing a data processing model based on multimodal fusion, characterized in that, include: The training data acquisition unit is adapted to acquire a multimodal training dataset, wherein the multimodal training dataset includes multimodal training data; The joint training unit is adapted to perform joint training using multimodal training data in the multimodal training dataset to obtain a data processing model based on multimodal fusion.

13. A data processing method based on multimodal fusion, characterized in that, include: Acquire multimodal data to be processed; The multimodal data to be processed is input into a multimodal fusion-based data processing model constructed using the method described in any one of claims 1-11, and the corresponding data processing results are obtained.

14. A data processing device based on multimodal fusion, characterized in that, include: The data acquisition unit is suitable for acquiring multimodal data to be processed. The analysis and prediction unit is adapted to input the multimodal data to be processed into a multimodal fusion-based data processing model constructed using the construction method of the multimodal fusion-based data processing model as described in any one of claims 1-11, and obtain the corresponding data processing results.

15. A device, characterized in that, It includes at least one memory and at least one processor, the memory storing one or more computer instructions, wherein the one or more computer instructions are executed by the processor to implement the method for constructing a data processing model based on multimodal fusion as described in any one of claims 1-11 or the data processing method based on multimodal fusion as described in claim 13.

16. A storage medium, characterized in that, The storage medium stores one or more computer instructions, which are used to implement the method for constructing a data processing model based on multimodal fusion as described in any one of claims 1-11 or the data processing method based on multimodal fusion as described in claim 13.

Citation Information

Patent Citations

  • Multi-modal information fusion method and device and electronic equipment

    CN111563551A

  • Video classification method and device, model training method and device, medium and electronic equipment

    CN115311599A

  • Method for establishing multi-modal data fusion model based on multi-attention mechanism

    CN116644381A

  • Multi-modal pre-training model training method and device and multi-modal data processing method and device

    CN116861995A

  • Data processing and model construction method based on multi-modal fusion, and related equipment

    CN118708969A