Method for data classification and model construction based on raman and mass spectrometry, related devices
By fusing Raman spectroscopy and mass spectrometry training data and employing self-attention feature extraction and convolution processing, a more accurate data classification model was constructed. This addresses the problem of insufficient accuracy in existing Raman and mass spectrometry data classification methods and improves the accuracy of sample analysis.
Patent Information
- Application Number
- CN202410836234.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-06-25
- Publication Date
- 2026-02-06
- Estimated Expiration
- 2044-06-25
AI Technical Summary
Existing data classification methods based on Raman and mass spectrometry still need improvement in accuracy.
By acquiring a first training dataset and a second training dataset, which contain Raman spectroscopy training data and mass spectrometry training data respectively, and fusing them to form a fused training dataset, a data classification model based on Raman and mass spectrometry is obtained through training using the fused training data. Self-attention feature extraction and convolution processing are used to improve the richness of feature information, and finally, the model weights are adjusted through a loss function to improve accuracy.
It improves the performance and prediction accuracy of the data classification model, and enhances the analytical capabilities for sample testing.
Smart Images

Figure CN118760980B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] Embodiments of the present application relate to the technical field of bio-information processing, and in particular to a method for data classification and model construction based on Raman and mass spectrometry, and related equipment. BACKGROUND
[0002] Sample detection technology has been widely used in the fields of medicine, biology, environment and food science, etc. Through detection of samples, detection data with analytical significance is obtained.
[0003] However, with the development of data classification methods based on Raman and mass spectrometry towards more complex directions, the quantity requirement of detection data of samples also increases accordingly. For example, with the development of artificial intelligence, machine learning and deep learning technologies have been gradually used to realize data analysis. However, the accuracy of the current data classification method based on Raman and mass spectrometry still needs to be improved. SUMMARY
[0004] The problem solved by embodiments of the present application is to provide a method for data classification and model construction based on Raman and mass spectrometry, and related equipment, which is beneficial to improve the performance of the data classification model construction method based on Raman and mass spectrometry, thereby further improving the accuracy of sample analysis.
[0005] To solve the above problems, embodiments of the present application provide a method for constructing a data classification model based on Raman and mass spectrometry, comprising:
[0006] obtaining a first training data set and a second training data set, wherein the first training data set comprises a plurality of Raman spectrum training data, and the second training data set comprises a plurality of mass spectrometry training data;
[0007] fusing the Raman spectrum training data in the first training data set with the mass spectrometry training data in the second training data set respectively to obtain a plurality of fusion training data corresponding thereto, thereby forming a fusion training data set;
[0008] training using the fusion training data in the fusion training data set to obtain a data classification model based on Raman and mass spectrometry.
[0009] Optionally, the training using the fusion training data in the fusion training data set to obtain a data classification model based on Raman and mass spectrometry comprises:
[0010] The fusion training data in the fusion training data set are respectively subjected to self-attention feature extraction processing in the Raman spectrum direction, a plurality of self-attention feature data in the Raman spectrum direction corresponding are obtained, and a first directional attention feature data set is formed; the self-attention feature data in the Raman spectrum direction in the first directional attention feature data set are respectively subjected to first convolution processing with corresponding mass spectrum training data, a plurality of first feature data corresponding are obtained, and a first feature data set is generated;
[0011] The fusion training data in the fusion training data set are respectively subjected to self-attention feature extraction processing in the Raman spectrum direction, a plurality of self-attention feature data in the Raman spectrum direction corresponding are obtained, and a first directional attention feature data set is formed; the self-attention feature data in the Raman spectrum direction in the first directional attention feature data set are respectively subjected to first convolution processing with corresponding mass spectrum training data, a plurality of first feature data corresponding are obtained, and a first feature data set is generated;
[0012] The first feature data in the first feature data set and the corresponding second feature data in the second feature data set are subjected to aggregation processing, a plurality of aggregation feature data corresponding are obtained, and an aggregation feature data set is generated;
[0013] The aggregation feature data in the aggregation feature data set is subjected to classification processing, and a prediction result corresponding is obtained;
[0014] According to the obtained prediction result, a preset loss function is used to calculate the loss value of the data classification model based on Raman and mass spectrum, and the weight of the data classification model based on Raman and mass spectrum is adjusted according to the calculated loss value, until the loss value of the data classification model based on Raman and mass spectrum converges.
[0015] Optionally, after the self-attention feature data in the Raman spectrum direction in the first directional attention feature data set and the corresponding mass spectrum training data are subjected to first convolution processing to obtain a plurality of first feature data corresponding, and a first feature data set is generated, and before the first feature data in the first feature data set and the corresponding second feature data in the second feature data set are subjected to aggregation processing to obtain a plurality of aggregation feature data corresponding, and an aggregation feature data set is generated, the fusion training data in the fusion training data set are used for training to obtain a data classification model based on Raman and mass spectrum, and the first feature data in the first feature data set is subjected to first dimension reduction processing.
[0016] The self-attention feature data of the mass spectrum direction in the second directional attention feature data set is respectively subjected to a second convolutional processing with corresponding Raman spectrum training data, a plurality of second feature data corresponding to the second feature data set are obtained, and the first feature data in the first feature data set is aggregated with the corresponding second feature data in the second feature data set to obtain a plurality of aggregated feature data corresponding to the aggregated feature data set before the first feature data in the first feature data set is aggregated with the corresponding second feature data in the second feature data set to obtain a plurality of aggregated feature data corresponding to the aggregated feature data set, and the fusion training data in the fusion training data set is used for training to obtain a Raman and mass spectrum based data classification model, and the second feature data in the second feature data set is subjected to a second dimension reduction processing.
[0017] The first feature data in the first feature data set is aggregated with the corresponding second feature data in the second feature data set to obtain a plurality of aggregated feature data corresponding to the aggregated feature data set, including: the first feature data subjected to the first dimension reduction processing is aggregated with the corresponding second feature data subjected to the second dimension reduction processing to obtain a plurality of aggregated feature data corresponding to the aggregated feature data set.
[0018] Optionally, the first dimension reduction processing includes first convolutional dimension reduction processing, and the second dimension reduction processing includes second convolutional dimension reduction processing.
[0019] Optionally, the self-attention feature extraction processing of the fusion training data in the fusion training data set in the Raman spectrum direction respectively to obtain a plurality of self-attention feature data in the Raman spectrum direction to form the first directional attention feature data set includes: the first multi-head self-attention mechanism is used to perform the self-attention feature extraction processing of the fusion training data in the fusion training data set in the Raman spectrum direction respectively to obtain a plurality of self-attention feature data in the Raman spectrum direction to form the first directional attention feature data set.
[0020] The self-attention feature extraction processing of the fusion training data in the fusion training data set in the mass spectrum direction respectively to obtain a plurality of self-attention feature data in the mass spectrum direction to form the second directional attention feature data set includes: the second multi-head self-attention mechanism is used to perform the self-attention feature extraction processing of the fusion training data in the fusion training data set in the mass spectrum direction respectively to obtain a plurality of self-attention feature data in the mass spectrum direction to form the second directional attention feature data set.
[0021] Optionally, the Raman and mass spectrum based data classification model is used for classifying pathogenic bacteria.
[0022] Correspondingly, the embodiment of the application also provides a construction device of a Raman and mass spectrum based data classification model, including:
[0023] The first obtaining unit is adapted to obtain a first training data set and a second training data set, the first training data set comprising a plurality of Raman spectrum training data, and the second training data set comprising a plurality of mass spectrum training data;
[0024] The fusion processing unit is adapted to fuse the Raman spectrum training data in the first training data set with the mass spectrum training data in the second training data set respectively, to obtain a plurality of corresponding fusion training data, and to form a fusion training data set.
[0025] The training unit is adapted to train based on the fusion training data in the fusion training data set, to obtain a Raman and mass spectrum based data classification model.
[0026] Correspondingly, the embodiment of the present application also provides a Raman and mass spectrum based data classification method, comprising:
[0027] Obtaining Raman spectrum data and mass spectrum data of a sample to be analyzed;
[0028] Fusing the Raman spectrum data and the mass spectrum data of the sample to be analyzed, and inputting the fused data into a Raman and mass spectrum based data classification model constructed by the construction method of the Raman and mass spectrum based data classification model according to any one of the above, to obtain a corresponding data classification result.
[0029] Correspondingly, the embodiment of the present application also provides a Raman and mass spectrum based data classification device, comprising:
[0030] The second obtaining unit is adapted to obtain Raman spectrum data and mass spectrum data of a sample to be analyzed;
[0031] The analysis and prediction unit is adapted to input the Raman spectrum data and the mass spectrum data of the sample to be analyzed into a Raman and mass spectrum based data classification model constructed by the construction method of the Raman and mass spectrum based data classification model according to any one of the above, to obtain a corresponding data classification result.
[0032] Correspondingly, the embodiment of the present application also provides a device comprising at least one memory and at least one processor, the memory storing one or more computer instructions, wherein the one or more computer instructions are executed by the processor to implement the construction method of the Raman and mass spectrum based data classification model according to any one of the above or the Raman and mass spectrum based data classification method according to the above.
[0033] Correspondingly, the embodiment of the present application also provides a storage medium storing one or more computer instructions, the one or more computer instructions being used to implement the construction method of the Raman and mass spectrum based data classification model according to any one of the above or the Raman and mass spectrum based data classification method according to the above.
[0034] Compared with the prior art, the technical scheme of the embodiment of the present application has the following advantages:
[0035] The embodiment of the present application provides a method for constructing a data classification model based on Raman and mass spectrometry, comprising: obtaining a first training data set and a second training data set, wherein the first training data set comprises a plurality of Raman spectrum training data, and the second training data set comprises a plurality of mass spectrometry training data; performing fusion processing on the Raman spectrum training data in the first training data set and the mass spectrometry training data in the second training data set respectively to obtain a plurality of corresponding fusion training data, and form a fusion training data set; and training the fusion training data in the fusion training data set to obtain a data classification model based on Raman and mass spectrometry.
[0036] In the method for constructing a data classification model based on Raman and mass spectrometry provided by the embodiment of the present application, the Raman spectrum training data in the first training data set and the mass spectrometry training data in the second training data set are first fused to obtain a plurality of corresponding fusion training data, which can improve the feature information richness of the fusion training data in the fusion training data set, and then the fusion training data in the fusion training data set is trained to obtain a data classification model based on Raman and mass spectrometry, which is helpful to improve the performance of the data classification model based on Raman and mass spectrometry constructed, and further improve the prediction accuracy of the data classification model based on Raman and mass spectrometry when the data classification model based on Raman and mass spectrometry is used to detect sample data subsequently. BRIEF DESCRIPTION OF DRAWINGS
[0037] Figure 1 FIG. 1 is a flowchart of an embodiment of the method for constructing a data classification model based on Raman and mass spectrometry provided by the technical scheme of the present application;
[0038] Figure 2 FIG. 2 is a schematic diagram of a Raman spectrum training data;
[0039] Figure 3 FIG. 3 is a schematic diagram of a mass spectrometry training data;
[0040] Figure 4 FIG. 4 is a schematic diagram of fusion training data obtained by fusing a Raman spectrum training data and a mass spectrometry training data;
[0041] Figure 5 FIG. 5 is a flowchart of an embodiment of the method for training the fusion training data in the fusion training data set to obtain a data classification model based on Raman and mass spectrometry in the technical scheme of the present application;
[0042] FIG. 6 is a comparative schematic diagram of a confusion matrix of the accuracy rate of the data classification model based on Raman and mass spectrometry;
[0043] Figure 7 A framework structure schematic diagram of a device for constructing a data classification model based on Raman and mass spectrum in an embodiment of the present application;
[0044] Figure 8 A flowchart schematic diagram of a data classification method based on Raman and mass spectrum in an embodiment of the present application;
[0045] Figure 9 A structure schematic diagram of a data classification device based on Raman and mass spectrum in an embodiment of the present application;
[0046] Figure 10 A structure schematic diagram of a device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0047] As described in the background, the accuracy of the existing data classification method based on Raman and mass spectrum still needs to be improved.
[0048] In order to solve the above technical problems, the present application provides a method for constructing a data classification model based on Raman and mass spectrum, comprising: obtaining a first training data set and a second training data set, wherein the first training data set comprises a plurality of Raman spectrum training data, and the second training data set comprises a plurality of mass spectrum training data; performing fusion processing on the Raman spectrum training data in the first training data set and the mass spectrum training data in the second training data set respectively, obtaining a plurality of corresponding fusion training data, and forming a fusion training data set; training using the fusion training data in the fusion training data set, and obtaining a data classification model based on Raman and mass spectrum.
[0049] In the method for constructing a data classification model based on Raman and mass spectrum provided in the present application, the Raman spectrum training data in the first training data set is first fused with the mass spectrum training data in the second training data set, and a plurality of corresponding fusion training data is obtained, which can improve the feature information richness of the fusion training data in the fusion training data set. Then, the fusion training data in the fusion training data set is trained to obtain a data classification model based on Raman and mass spectrum, which is helpful to improve the performance of the constructed data classification model based on Raman and mass spectrum, and further improve the prediction accuracy of the data classification model based on Raman and mass spectrum when detecting the sample data to be processed.
[0050] In order to make the above-mentioned objects, features and advantages of the present application more obvious and easy to understand, the specific embodiments of the present application will be described in detail below with reference to the accompanying drawings.
[0051] Figure 1A flowchart of an embodiment of the method for constructing a data classification model based on Raman and mass spectrometry provided by the present application is shown. Referring to Figure 1 A method for constructing a data classification model based on Raman and mass spectrometry can specifically include the following steps:
[0052] Step S110: Obtain a first training data set and a second training data set, wherein the first training data set includes a plurality of Raman spectrum training data, and the second training data set includes a plurality of mass spectrum training data.
[0053] Step S120: Fuse the Raman spectrum training data in the first training data set with the mass spectrum training data in the second training data set respectively to obtain a plurality of corresponding fusion training data, and form a fusion training data set.
[0054] Step S130: Train using the fusion training data in the fusion training data set to obtain a data classification model based on Raman and mass spectrometry.
[0055] Please continue to refer to Figure 1 Step S110 is performed to obtain a first training data set and a second training data set, wherein the first training data set includes a plurality of Raman spectrum training data, and the second training data set includes a plurality of mass spectrum training data.
[0056] The first training data set and the second training data set are obtained to provide a basis for subsequent fusion of the Raman spectrum training data in the first training data set with the mass spectrum training data in the second training data set to obtain a plurality of corresponding fusion training data and form a fusion training data set.
[0057] Raman spectrum (Raman) is an analysis method that uses the Raman scattering effect of molecules to obtain information about the vibration and rotation of molecules. Specifically, Raman spectrum data is obtained by analyzing scattered light with a wavelength different from the excitation light to obtain information about the vibration and rotation of molecules.
[0058] Referring to Figure 2 wherein the abscissa represents the difference between the relative frequency of the sample and the laser light source, also known as the Raman shift or frequency shift, that is, the relative wavenumber, and the ordinate represents the intensity or signal strength of the sample scattered light, or the Raman intensity.
[0059] Mass spectrometry data is commonly used in the fields of medicine, biology, environment, and food science, and is information about the mass-to-charge ratio (m / z) of molecules in a sample and the intensity at each mass-to-charge ratio obtained by detecting the sample using a mass spectrometer. Accordingly, mass spectrometry data contains information about the intensity as a function of mass-to-charge ratio.
[0060] Referring toFigure 3 , Figure 3 The horizontal axis represents the mass-to-charge ratio of the mass spectrometry data, and the vertical axis represents the intensity at that mass-to-charge ratio. For example, the sample being tested can be a pathogen, and the mass spectrometry data is obtained by detecting the pathogen using a mass spectrometer.
[0061] Depending on the actual needs, the Raman spectroscopy training data in the first training dataset can come from a variety of sources.
[0062] As an example, the steps for obtaining the first training dataset include: using a Raman spectrometer to detect the sample to obtain raw Raman spectral training data; performing a first data augmentation process on the raw Raman spectral training data to obtain corresponding first newly added training data, wherein the first newly added training data and the raw Raman spectral training data constitute the first training dataset.
[0063] As an example, the original Raman spectrum training data is first copied, and the Raman frequency shift in the copied original Raman spectrum training data is adjusted according to the distribution law of the Raman frequency shift in the original Raman spectrum training data to obtain the corresponding first new training data.
[0064] The number of Raman spectroscopy training data in the first training dataset can be set according to the training requirements of the data classification model based on Raman and mass spectrometry, and is not limited here.
[0065] Depending on the actual needs, the sources of mass spectrometry training data in the second training dataset can be diverse.
[0066] As an example, the steps for obtaining the second training dataset include: using a mass spectrometer to detect the sample to obtain original mass spectrometry training data; performing a second data augmentation process on the original mass spectrometry training data to obtain corresponding second newly added training data, wherein the second newly added training data and the original mass spectrometry training data constitute the second training dataset.
[0067] As an example, the steps of performing a second data augmentation process on the original mass spectrometry training data to obtain the corresponding second new training data include: copying the original mass spectrometry training data to obtain the corresponding copied training data; adjusting the signal intensity in the copied training data according to the distribution pattern of signal intensity in the original mass spectrometry training data to obtain the corresponding second new training data.
[0068] The number of mass spectrometry training data in the second training dataset can be set according to the training requirements of the Raman and mass spectrometry-based data classification model, and there is no limit here.
[0069] Figure 4A schematic diagram of fusion training data obtained by fusing Raman spectrum training data and mass spectrum training data is shown. Please refer to Figures 1 to 4 In step S120, the Raman spectrum training data in the first training data set are fused with the mass spectrum training data in the second training data set respectively to obtain a plurality of corresponding fusion training data, thereby forming a fusion training data set.
[0070] The Raman spectrum training data in the first training data set are fused with the mass spectrum training data in the second training data set respectively to obtain a plurality of corresponding fusion training data, thereby forming a fusion training data set, which provides a basis for subsequent training using the fusion training data in the fusion training data set to obtain a data classification model based on Raman and mass spectrum.
[0071] As an example, the step of fusing the Raman spectrum training data in the first training data set with the mass spectrum training data in the second training data set respectively includes: traversing the Raman spectrum training data in the first training data set to obtain a current Raman spectrum training data; fusing the current Raman spectrum training data with each mass spectrum training data in the second training data set respectively to obtain a plurality of corresponding fusion training data; obtaining a next Raman spectrum training data from the first training data set as the current Raman spectrum training data, and repeating the step of fusing the current Raman spectrum training data with each mass spectrum training data in the second training data set respectively to obtain a plurality of corresponding fusion training data until the Raman spectrum training data in the first training data set are completely traversed.
[0072] According to actual needs, a suitable data fusion algorithm can be used to fuse the Raman spectrum training data in the first training data set with the mass spectrum training data in the second training data set respectively. The data fusion algorithm can be selected by those skilled in the art according to actual needs, which is not limited herein.
[0073] The Raman spectrum training data in the first training data set are fused with the mass spectrum training data in the second training data set respectively, so that the generated corresponding fusion training data includes the feature information of the Raman spectrum training data, the feature information of the mass spectrum training data, and the association feature information between the Raman spectrum training data and the mass spectrum training data, so that the generated fusion training data has more rich and comprehensive feature information.
[0074] Figure 5A flowchart of an embodiment of the present application is shown. The embodiment is a process for obtaining a Raman and mass spectrum based data classification model by training using fusion training data in a fusion training data set. Figure 5 The step of obtaining a Raman and mass spectrum based data classification model by training using fusion training data in a fusion training data set can specifically include:
[0075] Step S1301: Perform self-attention feature extraction processing on the fusion training data in the fusion training data set in the Raman spectrum direction, respectively, to obtain corresponding multiple pieces of self-attention feature data in the Raman spectrum direction, to form a first directional attention feature data set.
[0076] Step S1302: Perform first convolution processing on the self-attention feature data in the Raman spectrum direction in the first directional attention feature data set and the corresponding mass spectrum training data, respectively, to obtain corresponding multiple pieces of first feature data, to generate a first feature data set.
[0077] Step S1303: Perform self-attention feature extraction processing on the fusion training data in the fusion training data set in the mass spectrum direction, respectively, to obtain corresponding multiple pieces of self-attention feature data in the mass spectrum direction, to form a second directional attention feature data set.
[0078] Step S1304: Perform second convolution processing on the self-attention feature data in the mass spectrum direction in the second directional attention feature data set and the corresponding Raman spectrum training data, respectively, to obtain corresponding multiple pieces of second feature data, to generate a second feature data set.
[0079] Step S1305: Perform aggregation processing on the first feature data in the first feature data set and the corresponding second feature data in the second feature data set, to obtain corresponding multiple pieces of aggregated feature data, to generate an aggregated feature data set.
[0080] Step S1306: Perform classification processing on the aggregated feature data in the aggregated feature data set, to obtain a corresponding prediction result.
[0081] Step S1307: According to the obtained prediction result, calculate a loss value of the Raman and mass spectrum based data classification model using a preset loss function, and adjust the weight of the Raman and mass spectrum based data classification model according to the calculated loss value, until the loss value of the Raman and mass spectrum based data classification model converges.
[0082] Please continue to refer to Figures 1 to 4, execute step S1301, respectively perform self-attention feature extraction processing of the Raman spectrum direction on the fusion training data in the fusion training data set, obtain a plurality of pieces of self-attention feature data of the Raman spectrum direction corresponding thereto, and form a first directional attention feature data set.
[0083] Respectively perform self-attention feature extraction processing of the Raman spectrum direction on the fusion training data in the fusion training data set, obtain a plurality of pieces of self-attention feature data of the Raman spectrum direction corresponding thereto, and form a first directional attention feature data set, so as to provide a basis for subsequently performing first convolution processing on the self-attention feature data of the Raman spectrum direction in the first directional attention feature data set and the corresponding mass spectrum training data respectively, obtaining a plurality of pieces of first feature data corresponding thereto, and generating a first feature data set.
[0084] Respectively perform self-attention feature extraction processing of the Raman spectrum direction on the fusion training data in the fusion training data set, so that the obtained self-attention feature data of the Raman spectrum direction can pay more attention to the detailed information related to the features of the Raman spectrum direction, which is helpful to improve the expression ability of the self-attention feature data of the Raman spectrum direction and to speed up the processing efficiency of information.
[0085] In some embodiments, the first multi-head self-attention mechanism is used to perform self-attention feature extraction processing of the Raman spectrum direction on the fusion training data in the fusion training data set, to obtain a plurality of pieces of self-attention feature data of the Raman spectrum direction corresponding thereto, and to form a first directional attention feature data set.
[0086] Specifically, when the first multi-head self-attention mechanism is used to perform self-attention feature extraction processing of the Raman spectrum direction on the fusion training data in the fusion training data set, the fusion training data in the fusion training data set are respectively multiplied by a plurality of groups of trainable parameter matrices W Q , W K , and W V to obtain a plurality of groups of query matrices (Q), key matrices (K), and value (V) matrices corresponding to the parameter matrices W Q , W K , and W V respectively, and the plurality of groups of query matrices, key matrices, and value matrices are subjected to splicing operation and vector dot product operation to obtain a corresponding attention matrix.
[0087] Please continue to refer to Figures 1 to 4 , execute step S1302, respectively perform first convolution processing on the self-attention feature data of the Raman spectrum direction in the first directional attention feature data set and the corresponding mass spectrum training data, obtain a plurality of pieces of first feature data corresponding thereto, and generate a first feature data set.
[0088] The self-attention feature data of the Raman spectrum direction in the first directional attention feature data set are respectively subjected to first convolution processing with corresponding mass spectrum training data, a plurality of first feature data corresponding to the first feature data set are obtained, and the first feature data set is generated, so as to prepare for subsequent aggregation processing of the first feature data in the first feature data set and corresponding second feature data in the second feature data set, and obtain a plurality of aggregation feature data to generate an aggregation feature data set.
[0089] In some embodiments, in the step of performing first convolution processing on the self-attention feature data of the Raman spectrum direction in the first directional attention feature data set with corresponding mass spectrum training data, the mass spectrum training data is the mass spectrum data used when the corresponding fusion data is obtained by fusion processing. Wherein, the fusion data used when the self-attention feature data of the Raman spectrum direction is obtained by performing self-attention feature extraction processing on the Raman spectrum direction.
[0090] The self-attention feature data of the Raman spectrum direction in the first directional attention feature data set are respectively subjected to first convolution processing with corresponding mass spectrum training data, so that the obtained first feature data can contain the association feature information between the corresponding Raman spectrum training data and the mass spectrum training data, which is beneficial to further increase the information richness of the first feature data.
[0091] Please continue to refer to Figure 2 , perform step S1303, and perform self-attention feature extraction processing on the fusion training data in the fusion training data set in the mass spectrum direction to obtain a plurality of self-attention feature data of the mass spectrum direction to form a second directional attention feature data set.
[0092] The self-attention feature extraction processing is performed on the fusion training data in the fusion training data set in the mass spectrum direction to obtain a plurality of self-attention feature data of the mass spectrum direction to form a second directional attention feature data set, which provides a basis for subsequent second convolution processing of the self-attention feature data of the mass spectrum direction in the second directional attention feature data set with corresponding Raman spectrum training data to obtain a plurality of second feature data to generate a second feature data set.
[0093] The self-attention feature extraction processing is performed on the fusion training data in the fusion training data set in the mass spectrum direction, so that the obtained self-attention feature data of the mass spectrum direction can pay more attention to the detailed information related to the features of the mass spectrum direction, which helps to improve the expression ability of the self-attention feature data of the mass spectrum direction and helps to speed up the processing efficiency of information.
[0094] In some embodiments, the second multi-head self-attention mechanism is used to perform self-attention feature extraction processing in the mass spectrum direction on the fusion training data in the fusion training data set respectively, to obtain a plurality of pieces of self-attention feature data in the mass spectrum direction corresponding to the self-attention feature extraction processing, and to form a second directional attention feature data set.
[0095] The second multi-head self-attention mechanism used for performing self-attention feature extraction processing in the mass spectrum direction on the fusion training data in the fusion training data set will not be described here again, and the first multi-head self-attention mechanism used for performing self-attention feature extraction processing in the Raman spectrum direction on the fusion training data in the fusion training data set is referred to the foregoing description.
[0096] Please continue to refer to Figures 1 to 4 , and perform step S1304 to perform second convolution processing on the self-attention feature data in the mass spectrum direction in the second directional attention feature data set and the corresponding Raman spectrum training data respectively, to obtain a plurality of pieces of second feature data corresponding to the second convolution processing, and to generate a second feature data set.
[0097] The second convolution processing on the self-attention feature data in the mass spectrum direction in the second directional attention feature data set and the corresponding Raman spectrum training data respectively is to prepare for subsequent aggregation processing of the first feature data in the first feature data set and the corresponding second feature data in the second feature data set, to obtain a plurality of pieces of aggregated feature data corresponding to the aggregation processing, and to generate an aggregated feature data set.
[0098] In some embodiments, in the step of performing second convolution processing on the self-attention feature data in the mass spectrum direction in the second directional attention feature data set and the corresponding Raman spectrum training data respectively, the corresponding Raman spectrum training data is the Raman spectrum training data used when the fusion processing is performed to obtain the corresponding fusion data. The fusion data used when the self-attention feature extraction processing in the mass spectrum direction is performed to obtain the self-attention feature data in the mass spectrum direction.
[0099] The second convolution processing on the self-attention feature data in the mass spectrum direction in the second directional attention feature data set and the corresponding Raman spectrum training data respectively is to make the obtained second feature data contain the associated feature information between the corresponding Raman spectrum training data and the mass spectrum training data, which is beneficial to further increase the information richness of the second feature data.
[0100] Please continue to refer to Figures 1 to 4, performing step S1305, aggregating the first feature data in the first feature data set and the corresponding second feature data in the second feature data set, obtaining a plurality of corresponding aggregated feature data, and generating an aggregated feature data set.
[0101] The first feature data in the first feature data set and the corresponding second feature data in the second feature data set are aggregated to obtain a plurality of corresponding aggregated feature data, and an aggregated feature data set is generated, which provides a basis for subsequent classification processing of the aggregated feature data in the aggregated feature data set to obtain a corresponding prediction result.
[0102] The step of aggregating the first feature data in the first feature data set and the corresponding second feature data in the second feature data set includes respectively splicing the first feature data in the first feature data set and the second feature data in the second feature data set to obtain a plurality of corresponding aggregated feature data.
[0103] In some embodiments, in the step of splicing the first feature data in the first feature data set and the corresponding second feature data in the second feature data set to obtain a plurality of corresponding aggregated feature data, the Raman spectrum direction self-attention feature data used when performing first convolution processing to obtain the first feature data corresponds to the mass spectrum direction self-attention feature data used when performing second convolution processing to obtain the second feature data, which is obtained by performing Raman spectrum direction self-attention feature extraction processing and mass spectrum direction self-attention feature extraction processing on the same fusion training data.
[0104] Correspondingly, the first feature data in the first feature data set and the corresponding second feature data in the second feature data set are aggregated to obtain a plurality of corresponding aggregated feature data, which has more rich and comprehensive sample feature information.
[0105] In this embodiment, before the first feature data in the first feature data set and the corresponding second feature data in the second feature data set are aggregated to obtain a plurality of corresponding aggregated feature data and generate an aggregated feature data set, the fusion training data in the fusion training data set is used for training to obtain a data classification model based on Raman and mass spectrum.
[0106] The first feature data in the first feature data set is processed by first dimension reduction, which can realize dimension reduction of the first feature data and the second feature data, and correspondingly help to reduce the subsequent operation amount.
[0107] As an example, the step of performing first dimension reduction processing on the first feature data in the first feature data set comprises performing first convolution dimension reduction processing on the first feature data in the first feature data set.
[0108] In other embodiments, other dimension reduction methods can also be used to perform first dimension reduction processing on the first feature data in the first feature data set, which can be selected by those skilled in the art according to actual needs, and is not limited herein.
[0109] In this embodiment, before the first feature data in the first feature data set and the corresponding second feature data in the second feature data set are aggregated to obtain a plurality of aggregated feature data and generate an aggregated feature data set, the step of obtaining a data classification model based on Raman and mass spectrometry by training using the fusion training data in the fusion training data set further comprises performing second dimension reduction processing on the second feature data in the second feature data set.
[0110] Performing second dimension reduction processing on the second feature data in the second feature data set can achieve dimension reduction of the second feature data, which is helpful to reduce the subsequent computational complexity.
[0111] As an example, the step of performing second dimension reduction processing on the second feature data in the second feature data set comprises performing second convolution dimension reduction processing on the second feature data in the second feature data set.
[0112] In other embodiments, other dimension reduction methods can also be used to perform second dimension reduction processing on the second feature data in the second feature data set, which can be selected by those skilled in the art according to actual needs, and is not limited herein.
[0113] Correspondingly, the step of aggregating the first feature data in the first feature data set and the corresponding second feature data in the second feature data set comprises first aggregating the first feature data after first dimension reduction processing and the corresponding second feature data after second dimension reduction processing to obtain a plurality of first aggregated feature data and generate a first aggregated feature data set.
[0114] Please continue to refer to Figures 1 to 4 , and perform step S1306 to classify the aggregated feature data in the aggregated feature data set to obtain a corresponding prediction result.
[0115] Classifying the aggregated feature data in the aggregated feature data set to obtain a corresponding prediction result provides a basis for subsequently calculating a loss value of the data classification model based on Raman and mass spectrometry using a preset loss function according to the obtained prediction result.
[0116] In some embodiments, the linear classifier is used to perform linear classification on the aggregated feature data in the aggregated feature data set respectively, to obtain a prediction result corresponding to the aggregated feature data. The prediction result corresponds to a probability value of the sample corresponding to the aggregated feature data belonging to each sample category.
[0117] Please continue to refer to Figures 1 to 4 According to the obtained prediction result, a preset loss function is used to calculate the loss value of the Raman and mass spectrum based data classification model, and the weight of the Raman and mass spectrum based data classification model is adjusted according to the calculated loss value, until the loss value of the Raman and mass spectrum based data classification model converges.
[0118] According to the obtained prediction result, a preset loss function is used to calculate the loss value of the Raman and mass spectrum based data classification model, and the weight of the Raman and mass spectrum based data classification model is adjusted according to the calculated loss value, until the loss value of the Raman and mass spectrum based data classification model converges.
[0119] In some embodiments, according to the obtained prediction result, a preset loss function is used to calculate the loss value of the Raman and mass spectrum based data classification model, and the weight of the Raman and mass spectrum based data classification model is adjusted according to the calculated loss value, to complete one iteration training of the Raman and mass spectrum based data classification model.
[0120] Specifically, the process of each iteration training includes: using a preset number (such as batch size) of aggregated feature data to train the Raman and mass spectrum based data classification model to be trained respectively, to obtain a preset number of prediction results; according to the difference between the training result and the true result, a preset loss function is used to calculate the corresponding loss value; according to the gradient value obtained by back propagation derivation, the weight of the Raman and mass spectrum based data classification model is adjusted once.
[0121] Accordingly, multiple iteration training is performed, and the weight of the Raman and mass spectrum based data classification model is iteratively updated until the loss value of the Raman and mass spectrum based data classification model on the preset verification set converges. Wherein, the loss value of the Raman and mass spectrum based data classification model on the preset verification set converges, means that the loss value of the Raman and mass spectrum based data classification model on the preset verification set reaches a minimum value.
[0122] For more details about the multiple iteration training process of the Raman and mass spectrum based data classification model, please refer to the iteration training process of the neural network model in the prior art, which will not be repeated here.
[0123] In some embodiments, the Raman and mass spectrum based data classification model is used for classifying pathogenic bacteria. In other embodiments, the Raman and mass spectrum based data classification model can also be used for classifying other types of samples, which are not limited here.
[0124] Referring to FIG. 6, FIG. 6 is a comparative diagram of the prediction accuracy confusion matrix of the Raman and mass spectrum based data classification model, and the darker the color of the grid where the number is located, the higher the accuracy.
[0125] Wherein, FIG. 6(a) represents the prediction accuracy confusion matrix of an embodiment of the Raman and mass spectrum based data classification model trained by using only Raman spectrum training data, FIG. 6(b) represents the prediction accuracy confusion matrix of an embodiment of the Raman and mass spectrum based data classification model trained by using only mass spectrum training data, and FIG. 6(c) represents the prediction accuracy confusion matrix of an embodiment of the Raman and mass spectrum based data classification model trained by using the fusion training data obtained by fusing the Raman spectrum training data and the mass spectrum training data in the embodiment of the present application.
[0126] As an example, the sample model is used for analyzing pathogenic bacteria, and in the confusion matrix, True represents the true result, and Predicted represents the predicted result. The types of pathogenic bacteria include M. chelonae, M. Huston, M. abscessus, M. mucogenicum, M. ulcerans, and M. peregrinum.
[0127] For example, taking M. chelonae as an example, referring to Fig. 6(a), when the Raman and mass spectrum based data classification model trained by using the Raman training data is used for analysis, the probability of predicting the accurate result is 0.85 (i.e. 85%), the probability of predicting M. Huston is 0.06 (i.e. 6%), the probability of predicting M. abscessus is 0.04 (i.e. 4%), and the probability of predicting M. mucogenicum is 0.05 (i.e. 5%); referring to Fig. 6(b), when the sample classification model trained by using the mass spectrum training data is used for analysis, the probability of predicting the accurate result is 0.46 (i.e. 46%), the probability of predicting M. ulcerans is 0.27 (i.e. 27%), and the probability of predicting M. peregrinum is 0.27 (i.e. 27%); referring to Fig. 6(c), when the Raman and mass spectrum based data classification model trained by using the fusion training data in some embodiments is used for analysis, the probability of predicting the accurate result is 0.99 (i.e. 99%), and the probability of predicting M. ulcerans is 0.01 (i.e. 1%).
[0128] As can be seen from Fig. 6(a), the overall prediction accuracy (acc) of the Raman and mass spectrum based data classification model trained by using only the Raman spectrum training data for pathogenic bacteria classification is 72.89%, as can be seen from Fig. 6(b), the overall prediction accuracy (acc) of the Raman and mass spectrum based data classification model trained by using only the mass spectrum training data for pathogenic bacteria classification is 63.22%, and as can be seen from Fig. 6(c), the prediction accuracy (acc) of the Raman and mass spectrum based data classification model trained by using the fusion training data obtained by fusing the Raman spectrum training data and the mass spectrum training data in the embodiments of the present application for pathogenic bacteria classification is 97.06%. It can be known that the prediction accuracy of the Raman and mass spectrum based data classification model constructed by the construction method of the Raman and mass spectrum based data classification model in the embodiments of the present application is significantly improved.
[0129] It is worth noting that for biological samples that need a culture period, the Raman and mass spectrum based data classification model constructed by the construction method of the Raman and mass spectrum based data classification model in the embodiments of the present application can effectively integrate different scale sample information in a short time, shorten the culture period of biological samples required for obtaining corresponding sample data, and improve the identification speed of actual clinical samples or other biological samples.
[0130] For example, in the analysis and identification of biological samples by using the current Raman and mass spectrum-based data classification model, it usually takes at least 48 hours of biological sample culture to obtain mass spectrum data with resolved peaks, so that the biological sample can be accurately analyzed and identified. However, by using the Raman and mass spectrum-based data classification model constructed by the construction method of the Raman and mass spectrum-based data classification model in the embodiment of the present application, the mass spectrum data obtained from the biological sample cultured for only 24 hours can be analyzed and identified, effectively shortening the culture period of the biological sample.
[0131] Correspondingly, the embodiment of the present application also provides a construction device of a Raman and mass spectrum-based data classification model.
[0132] Figure 7 The framework structure schematic diagram of a construction device of a Raman and mass spectrum-based data classification model in the embodiment of the present application is shown. Referring to Figure 7 A construction device 700 of a Raman and mass spectrum-based data classification model, comprising: a first acquisition unit 701 adapted to acquire a first training data set and a second training data set, wherein the first training data set comprises a plurality of Raman spectrum training data, and the second training data set comprises a plurality of mass spectrum training data; a fusion processing unit 702 adapted to fuse the Raman spectrum training data in the first training data set with the mass spectrum training data in the second training data set respectively, acquire a plurality of corresponding fusion training data, and form a fusion training data set; and a training unit 703 adapted to train by using the fusion training data in the fusion training data set, and acquire a Raman and mass spectrum-based data classification model.
[0133] The construction device of the Raman and mass spectrum-based data classification model in some embodiments can be used to execute the foregoing construction method of the Raman and mass spectrum-based data classification model, or other functional structures can also be used to execute the foregoing construction method of the Raman and mass spectrum-based data classification model. For the construction device of the Raman and mass spectrum-based data classification model, please refer to the foregoing content of the construction method of the Raman and mass spectrum-based data classification model, which will not be repeated here.
[0134] Correspondingly, the embodiment of the present application also provides a Raman and mass spectrum-based data classification method.
[0135] Figure 8 The flowchart of a Raman and mass spectrum-based data classification method in the embodiment of the present application is shown. Referring to FIG. 6, a Raman and mass spectrum-based data classification method can comprise:
[0136] Step S810: acquiring Raman spectrum data and mass spectrum data of a sample to be analyzed;
[0137] Step S820: inputting the Raman spectrum data and mass spectrum data of the sample to be analyzed into the Raman and mass spectrum based data classification model constructed by the method for constructing a Raman and mass spectrum based data classification model to obtain a corresponding data classification result.
[0138] Correspondingly, the Raman and mass spectrum based data classification model constructed by the method for constructing a Raman and mass spectrum based data classification model is used for detecting the sample data to be analyzed, and the prediction result of the sample data to be analyzed. The method for constructing a Raman and mass spectrum based data classification model is described in the foregoing part, and will not be described here.
[0139] Correspondingly, the embodiment of the present application further provides a Raman and mass spectrum based data classification device.
[0140] Figure 9 The structure of a Raman and mass spectrum based data classification device in the embodiment of the present application is shown. Referring to Figure 9 A Raman and mass spectrum based data classification device 900 can comprise: a second acquisition unit 901 adapted to acquire Raman spectrum data and mass spectrum data of a sample to be analyzed; and a data classification unit 902 adapted to input the Raman spectrum data and mass spectrum data of the sample to be analyzed into a Raman and mass spectrum based data classification model constructed by the method for constructing a Raman and mass spectrum based data classification model to obtain a corresponding data classification result.
[0141] The Raman and mass spectrum based data classification device in some embodiments can be used to execute the foregoing Raman and mass spectrum based data classification method, or can also be used to execute the foregoing Raman and mass spectrum based data classification method with other functional structures. The Raman and mass spectrum based data classification device is described in the foregoing part of the Raman and mass spectrum based data classification method, and will not be described here.
[0142] Correspondingly, the embodiment of the present application further provides a device comprising at least one memory and at least one processor, wherein the memory stores one or more computer instructions, and the one or more computer instructions are executed by the processor to implement the method for constructing a Raman and mass spectrum based data classification model or the Raman and mass spectrum based data classification method. The method for constructing a Raman and mass spectrum based data classification model or the Raman and mass spectrum based data classification method is described in the foregoing part, and will not be described here.
[0143] Accordingly, embodiments of the present invention also provide a storage medium storing one or more computer instructions, which are used to implement the method for constructing the data classification model based on Raman and mass spectrometry or the data classification method based on Raman and mass spectrometry as described above. The method for constructing the data classification model based on Raman and mass spectrometry or the data classification method based on Raman and mass spectrometry are described in the foregoing sections and will not be repeated here.
[0144] Accordingly, embodiments of the present invention also provide an apparatus that can implement the method for constructing a data classification model based on Raman and mass spectrometry or the data classification method based on Raman and mass spectrometry provided in the embodiments of the present invention by loading the above-described method for constructing a data classification model based on Raman and mass spectrometry in the form of a program.
[0145] refer to Figure 10 The diagram illustrates the hardware structure of a device according to an embodiment of the present invention. The device of this embodiment includes: at least one processor 01, at least one communication interface 02, at least one memory 03, and at least one communication bus 04.
[0146] In some embodiments, the number of processor 01, communication interface 02, memory 03 and communication bus 04 is at least one, and processor 01, communication interface 02 and memory 03 communicate with each other through communication bus 04.
[0147] Communication interface 02 can be an interface for a communication module used for network communication, such as the interface for a GSM module.
[0148] Processor 01 may be a central processing unit (CPU), an application-specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the methods described in this embodiment.
[0149] Memory 03 may include high-speed RAM, and may also include non-volatile memory, such as at least one disk storage device. Memory 03 stores one or more computer instructions, which are executed by processor 01 to implement the Raman and mass spectrometry-based data classification model construction method or the Raman and mass spectrometry-based data classification method provided in the foregoing embodiments.
[0150] It should be noted that the implementation device described above can further include other devices (not shown) that can not be necessary for the disclosure of the embodiments of the present application; since these other devices can not be necessary for understanding the disclosure of the embodiments of the present application, the embodiments of the present application do not introduce them one by one.
[0151] The above-described embodiments of the present application are combinations of elements and features of the present application. The elements or features can be considered selective unless otherwise mentioned. Each of the elements or features can be practiced without the other elements or features being present. Also, some constructions of any one of the embodiments can be included in the other embodiments, and an arbitrary combination of a part of elements and / or features of the embodiments can constitute another embodiment. The order of operations described in the embodiments of the present application can be re-arranged. Some construction of any one of the embodiments can be included in another embodiment and can be substituted for the corresponding construction of another embodiment. It is evident that the scope of the present application is not limited by the detailed description of the embodiments of the present application, and those skilled in the art within the scope of the present application can implement the present application by appropriate modifications for the embodiments disclosed in the present application without deviating from the scope of the present application. Accordingly, the scope of the present application should be ascertained by the technical scope of the claims rather than the above description, and all differences within equivalent scope of the claims should be construed as being included in the scope of the present application.
[0152] Embodiments of the present application can be implemented in various forms, for example, hardware, firmware, software, or a combination thereof. In a hardware configuration, the method according to the exemplary embodiments of the present application can be implemented by one or more Application Specific Integrated Circuits (ASICs), Digital Signal Processors (DSPs), Digital Signal Processing Devices (DSPDs), Programmable Logic Devices (PLDs), Field Programmable Gate Arrays (FPGAs), processors, controllers, micro-controllers, microprocessors, etc.
[0153] In a firmware or software configuration, the embodiments of the present application can be implemented in the form of modules, procedures, functions, and the like. Software code can be stored in a memory unit and executed by a processor. The memory unit is located at the interior or exterior of the processor and can deliver data to and receive data from the processor via various means.
[0154] The above description of disclosed embodiments is merely intended to teach a person skilled in the art how to implement or use the present application. Many modifications to these embodiments will be readily apparent to those skilled in the art, and the generic principles defined herein can be applied to other embodiments without departing from the spirit or scope of the application. Accordingly, the present application is not intended to be limited to the embodiments shown herein but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
[0155] While the present application has been disclosed as above, the present application is not limited to the above. Any person skilled in the art can make various changes and modifications without departing from the spirit and scope of the present application, and the scope of protection of the present application should be construed on the basis of the scope of the claims.
Claims
1. A method for constructing a data classification model based on Raman and mass spectrometry, characterized in that, include: A first training dataset and a second training dataset are obtained. The first training dataset includes multiple Raman spectral training data, and the second training dataset includes multiple mass spectrometry training data. The Raman spectral training data is Raman spectral data acquired by collecting samples from the test sample using a Raman spectrometer, and the mass spectrometry training data is mass spectrometry data acquired by collecting samples from the test sample using a mass spectrometer. The Raman spectroscopy training data in the first training dataset is fused with the mass spectrometry training data in the second training dataset to obtain multiple fused training data, forming a fused training dataset. The multiple fused training data each include feature information of the Raman spectroscopy training data, feature information of the mass spectrometry training data, and correlation feature information between the Raman spectroscopy training data and the mass spectrometry training data. The method involves training the fused training data in the fused training dataset to obtain a data classification model based on Raman and mass spectrometry. This includes: performing self-attention feature extraction processing on the fused training data in the fused training dataset for Raman spectral directions to obtain multiple corresponding self-attention feature data in Raman spectral directions, forming a first directional attention feature dataset; performing a first convolution processing on the self-attention feature data in the Raman spectral directions of the first directional attention feature dataset with the corresponding mass spectrometry training data to obtain multiple corresponding first feature data, generating a first feature dataset; and performing self-attention feature extraction processing on the fused training data in the fused training dataset for mass spectrometry directions to obtain multiple corresponding self-attention feature data in mass spectrometry directions, forming a second directional attention feature dataset. The self-attention feature data in the mass spectrometry direction of the second directional attention feature dataset are respectively convolved with the corresponding Raman spectroscopy training data to obtain multiple corresponding second feature data, generating a second feature dataset; the first feature data in the first feature dataset is aggregated with the corresponding second feature data in the second feature dataset to obtain multiple corresponding aggregated feature data, generating an aggregated feature dataset; the aggregated feature data in the aggregated feature dataset is classified to obtain corresponding prediction results; based on the obtained prediction results, the loss value of the Raman and mass spectrometry-based data classification model is calculated using a preset loss function, and the weights of the Raman and mass spectrometry-based data classification model are adjusted according to the calculated loss value until the loss value of the Raman and mass spectrometry-based data classification model converges.
2. The method for constructing a data classification model based on Raman and mass spectrometry as described in claim 1, characterized in that, The self-attention feature data in the Raman spectral direction of the first directional attention feature dataset is convolved with the corresponding mass spectrometry training data to obtain multiple first feature data and generate a first feature dataset. Then, the first feature data in the first feature dataset is aggregated with the corresponding second feature data in the second feature dataset to obtain multiple aggregated feature data. Before generating the aggregated feature dataset, the fusion training data in the fusion training dataset is used for training to obtain a data classification model based on Raman and mass spectrometry. The method also includes: performing a first dimensionality reduction process on the first feature data in the first feature dataset. The process involves performing a second convolution on the self-attention feature data in the mass spectrometry direction of the second directional attention feature dataset and the corresponding Raman spectroscopy training data to obtain multiple corresponding second feature data. After generating the second feature dataset, the process also involves aggregating the first feature data in the first feature dataset with the corresponding second feature data in the second feature dataset to obtain multiple corresponding aggregated feature data. Before generating the aggregated feature dataset, the process involves training the data using the fusion training dataset to obtain a data classification model based on Raman and mass spectrometry. The process further includes performing a second dimensionality reduction on the second feature data in the second feature dataset. The first feature data in the first feature dataset and the corresponding second feature data in the second feature dataset are aggregated to obtain multiple aggregated feature data and generate an aggregated feature dataset. This includes: aggregating the first feature data after first dimensionality reduction with the corresponding second feature data after second dimensionality reduction to obtain multiple aggregated feature data and generate an aggregated feature dataset.
3. The method for constructing a data classification model based on Raman and mass spectrometry as described in claim 2, characterized in that, The first dimensionality reduction process includes a first convolutional dimensionality reduction process, and the second dimensionality reduction process includes a second convolutional dimensionality reduction process.
4. The method for constructing a data classification model based on Raman and mass spectrometry as described in claim 1, characterized in that, The step of performing self-attention feature extraction processing on the fused training data in the fused training dataset according to the Raman spectral direction to obtain corresponding self-attention feature data in multiple Raman spectral directions and forming a first directional attention feature dataset includes: using a first multi-head self-attention mechanism to perform self-attention feature extraction processing on the fused training data in the fused training dataset according to the Raman spectral direction to obtain corresponding self-attention feature data in multiple Raman spectral directions and forming a first directional attention feature dataset; The step of performing self-attention feature extraction processing in the mass spectrometry direction on the fused training data in the fused training dataset to obtain corresponding self-attention feature data in multiple mass spectrometry directions and forming a second directional attention feature dataset includes: using a second multi-head self-attention mechanism to perform self-attention feature extraction processing in the mass spectrometry direction on the fused training data in the fused training dataset to obtain corresponding self-attention feature data in multiple mass spectrometry directions and forming a second directional attention feature dataset.
5. The method for constructing a data classification model based on Raman and mass spectrometry as described in claim 1, characterized in that, The Raman and mass spectrometry-based data classification model is used to classify pathogens.
6. A device for constructing a data classification model based on Raman and mass spectrometry, characterized in that, include: The first acquisition unit is adapted to acquire a first training dataset and a second training dataset. The first training dataset includes multiple Raman spectral training data, and the second training dataset includes multiple mass spectrometry training data. The Raman spectral training data is Raman spectral data acquired by collecting samples from a Raman spectrometer, and the mass spectrometry training data is mass spectrometry data acquired by collecting samples from a mass spectrometer. The fusion processing unit is adapted to fuse the Raman spectroscopy training data in the first training dataset with the mass spectrometry training data in the second training dataset to obtain multiple fused training data to form a fused training dataset. The multiple fused training data include feature information of the Raman spectroscopy training data, feature information of the mass spectrometry training data, and correlation feature information between the Raman spectroscopy training data and the mass spectrometry training data. The training unit is adapted to train using the fused training data in the fused training dataset to obtain a data classification model based on Raman and mass spectrometry, including: performing self-attention feature extraction processing on the fused training data in the fused training dataset according to the Raman spectral direction to obtain multiple self-attention feature data in the corresponding Raman spectral direction, forming a first directional attention feature dataset; performing a first convolution processing on the self-attention feature data in the Raman spectral direction of the first directional attention feature dataset with the corresponding mass spectrometry training data to obtain multiple first feature data, generating a first feature dataset; and performing self-attention feature extraction processing on the fused training data in the fused training dataset according to the mass spectrometry direction to obtain multiple self-attention feature data in the corresponding mass spectrometry direction, forming a second directional attention feature dataset. The self-attention feature data in the mass spectrometry direction of the second directional attention feature dataset are respectively convolved with the corresponding Raman spectroscopy training data to obtain multiple corresponding second feature data, generating a second feature dataset; the first feature data in the first feature dataset is aggregated with the corresponding second feature data in the second feature dataset to obtain multiple corresponding aggregated feature data, generating an aggregated feature dataset; the aggregated feature data in the aggregated feature dataset is classified to obtain corresponding prediction results; based on the obtained prediction results, the loss value of the Raman and mass spectrometry-based data classification model is calculated using a preset loss function, and the weights of the Raman and mass spectrometry-based data classification model are adjusted according to the calculated loss value until the loss value of the Raman and mass spectrometry-based data classification model converges.
7. A data classification method based on Raman and mass spectrometry, characterized in that, include: Acquire Raman spectral data and mass spectrometry data of the sample to be analyzed; The Raman spectral data and mass spectrometry data of the sample to be analyzed are fused and input into the Raman and mass spectrometry-based data classification model constructed by the method of constructing a Raman and mass spectrometry-based data classification model as described in any one of claims 1-5, and the corresponding data classification results are obtained.
8. A data classification device based on Raman and mass spectrometry, characterized in that, include: The second acquisition unit is suitable for acquiring Raman spectral data and mass spectrometry data of the sample to be analyzed; The analysis and prediction unit is adapted to input the Raman spectral data and mass spectrometry data of the sample to be analyzed into the Raman and mass spectrometry-based data classification model constructed by the method for constructing a data classification model based on Raman and mass spectrometry as described in any one of claims 1-5, and obtain the corresponding data classification results.
9. A device, characterized in that, It includes at least one memory and at least one processor, the memory storing one or more computer instructions, wherein the one or more computer instructions are executed by the processor to implement the method for constructing a data classification model based on Raman and mass spectrometry as described in any one of claims 1-5 or the data classification method based on Raman and mass spectrometry as described in claim 7.
10. A storage medium, characterized in that, The storage medium stores one or more computer instructions, which are used to implement the method for constructing a data classification model based on Raman and mass spectrometry as described in any one of claims 1-5 or the data classification method based on Raman and mass spectrometry as described in claim 7.
Citation Information
Patent Citations
A liquor vintage recognition method based on a fusion technology of ion mobility spectrometry / mass spectrometry / Raman spectroscopy
CN103293141A
Raman spectrum classification method, species blood semen and species classification method
CN117349741A
Data processing and model construction method based on multi-modal fusion, and related equipment
CN118708969A