Bacterial drug resistance prediction method and system based on deep fusion of multi-modal Raman spectrum data
Through the deep fusion method of multimodal Raman spectroscopy data, the changes in bacterial drug resistance are monitored in real time, and the problem of inability to effectively integrate static and dynamic data in the existing technology is solved, efficient and accurate drug resistance prediction is achieved, and accurate adjustment of clinical drug regimens is supported.
Patent Information
- Application Number
- CN202510649472.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-20
- Publication Date
- 2025-08-29
- Estimated Expiration
- 2045-05-20
AI Technical Summary
The existing technology cannot monitor the dynamic changes in bacterial drug resistance in real time, resulting in a lack of accuracy in clinical antibiotic drug regimens, and the existing Raman spectroscopy technology cannot effectively integrate static and dynamic data, and lacks the spatial and temporal dimensional correlation of multimodal data.
The multimodal Raman spectral data deep fusion method is adopted, and the static and dynamic Raman spectral data is obtained, embedded encoding and feature extraction is performed, and deep cross-fusion is performed with multi-layer perceptrons. The spatial and temporal dual-flow attention mechanism and the Transformer encoder are used to extract features, and a cross-modal cross-fusion module is built, combining reinforcement learning and transfer learning optimization model.
Real-time monitoring of dynamic changes in bacterial drug resistance is achieved, detection time is shortened, and efficient and accurate drug resistance monitoring solutions are provided, supporting clinical precise adjustment of drug dosage and treatment cycle.
Smart Images

Figure CN120564889A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of bacterial resistance prediction, and specifically relates to a bacterial resistance prediction method and system based on deep fusion of multimodal Raman spectroscopy data. Background Art
[0002] Traditional bacterial drug sensitivity detection technology mainly uses culture method and phenotypic analysis. This method requires steps such as bacterial isolation, amplification, drug exposure and phenotypic observation. The whole process usually takes 18-24 hours. The detection time of some special pathogens such as Mycobacterium tuberculosis can even reach several weeks. This time-consuming process directly leads to the difficulty of quickly obtaining drug resistance data in the early stages of infection in the clinic, thereby delaying the window period for precise antibiotic treatment. In addition, existing technologies can only obtain static drug resistance phenotypic results and cannot monitor the dynamic response of bacteria under the action of antibiotics (such as changes in membrane fluidity, metabolic reprogramming, etc.), resulting in a lag in the evaluation of the current trend of drug resistance development.
[0003] In the field of drug resistance mechanism research, although gene sequencing technology can analyze the mutation sites of drug-resistant genes, it cannot directly link gene expression with real-time drug resistance phenotypes; although mass spectrometry technology can detect bacterial metabolites, it is not sensitive enough to low-abundance molecules and has difficulty capturing dynamic changes. In contrast, Raman spectroscopy technology, with its advantages such as non-invasive detection, high specificity and molecular fingerprint recognition capabilities, is gradually realizing its potential in the field of drug resistance monitoring. However, static Raman spectroscopy can only obtain static molecular information and cannot capture the real-time dynamic response of bacteria under the action of antibiotics. In addition, existing methods still lack multimodal data fusion capabilities. For example, the simple splicing of static and dynamic data cannot effectively mine high-order correlations in the spatiotemporal dimension. Existing feature fusion methods have problems such as insufficient versatility in cross-modal association modeling.
[0004] In clinical practice, doctors urgently need tools that can monitor the dynamic changes in bacterial resistance in real time to precisely adjust antibiotic dosage and treatment duration. Currently, clinical practice primarily relies on empirical medication regimens, which can easily lead to problems such as undertreatment or overdose. To address this issue, this paper proposes a bacterial resistance prediction method and system based on deep fusion of multimodal Raman spectroscopy. Summary of the Invention
[0005] In response to the above-mentioned problems in the prior art, the present invention provides a bacterial resistance prediction method and system based on deep fusion of multimodal Raman spectroscopy data.
[0006] Bacterial resistance prediction method based on deep fusion of multimodal Raman spectroscopy data, including:
[0007] Obtain static Raman spectral data of target bacteria and target antibiotics, as well as dynamic Raman spectral data of target bacteria under the action of target antibiotics;
[0008] Performing embedded coding and feature extraction on the static Raman spectrum data and the dynamic Raman spectrum data respectively to obtain static spectrum features and dynamic spectrum features;
[0009] The static spectral features and the dynamic spectral features are deeply cross-fused and combined with a multi-layer perceptron to obtain the drug resistance prediction results of the target bacteria.
[0010] Preferably, the method further includes preprocessing the acquired dynamic Raman spectroscopy data, wherein the preprocessing method includes:
[0011] Using discrete wavelet transform, multi-scale analysis is performed on the dynamic Raman spectroscopy data to obtain high-frequency coefficients;
[0012] Performing threshold processing on the high-frequency coefficients, combining low-pass filter coefficients and high-pass filter coefficients, reconstructing the signal, and obtaining denoised dynamic Raman spectroscopy data;
[0013] The baseline correction of the denoised dynamic Raman spectroscopy data was performed using the asymmetric least squares method.
[0014] The baseline-corrected dynamic Raman spectral data are subjected to area normalization and standard normal transformation, and the transformed dynamic Raman spectral data are subjected to time correction using a dynamic time warping algorithm to obtain the pre-processed dynamic Raman spectral data.
[0015] Preferably, the method for embedding the static Raman spectrum data includes:
[0016] Performing a matrix multiplication operation on the static Raman spectrum data using a preset weight matrix to complete input embedding of the static Raman spectrum data;
[0017] Position encoding is performed on the static Raman spectrum data that has been embedded, thereby completing the embedded encoding of the static Raman spectrum data.
[0018] Preferably, the method for embedding the pre-processed dynamic Raman spectroscopy data includes:
[0019] The pre-processed dynamic Raman spectroscopy data is discretized according to time series and spatial distribution using a token embedding method to obtain a token embedding representation;
[0020] Position embedding is added to the token embedding representation to obtain the token embedding representation with position information, completing the embedding encoding of the preprocessed dynamic Raman spectroscopy data.
[0021] Preferably, a Transformer encoder with a preset number of layers is used to extract features from the static Raman spectrum data after embedding coding, wherein the Transformer encoder includes a self-attention mechanism module, a pre-feedback layer, and a second-layer normalization; the self-attention mechanism module includes a multi-head attention layer and a first-layer normalization; the method for extracting features from the static Raman spectrum data after embedding coding using the Transformer encoder includes:
[0022] Perform linear transformation on the static Raman spectrum data after embedding coding to obtain a linear matrix;
[0023] Introducing the bias of the antibiotic-bacteria affinity into the multi-head attention layer, processing the linear matrix, and obtaining the multi-head attention layer output;
[0024] Perform first-layer normalization on the output of the multi-head attention layer to obtain the first-layer normalized output;
[0025] The output of the self-attention mechanism module is obtained by adding the normalized output of the first layer to the static Raman spectrum data after embedding encoding;
[0026] Inputting the output of the self-attention mechanism module into the pre-feedback layer, and adding the pre-feedback output and the output of the self-attention mechanism module to perform a second-layer normalization process to obtain a Transformer encoder output;
[0027] The output of the previous layer of Transformer encoder is used as the input of the next layer of Transformer encoder until the output of the last layer of Transformer encoder is obtained to obtain the static spectral features.
[0028] Preferably, a spatiotemporal dual-stream attention mechanism encoder with a preset number of layers is combined with a second multi-layer perceptron to perform feature extraction on the dynamic Raman spectrum data after embedding encoding, wherein the spatiotemporal dual-stream attention mechanism encoder includes a spatiotemporal parallel cross self-attention mechanism module, a fourth layer of normalization, and a first multi-layer perceptron; the spatiotemporal parallel cross self-attention mechanism module includes a third layer of normalization and a spatiotemporal parallel multi-head attention layer;
[0029] The method for extracting features from dynamic Raman spectroscopy data after embedding coding by using a spatiotemporal dual-stream attention mechanism encoder combined with a second multi-layer perceptron includes:
[0030] Reshape the embedded encoded dynamic Raman spectroscopy data to obtain spatial dimension input representation and time dimension input representation;
[0031] Performing third-layer normalization processing on the spatial dimension input representation and the time dimension input representation to obtain a third-layer normalized output;
[0032] Performing linear transformation on the normalized output of the third layer to obtain a spatial dimension correlation matrix and a temporal dimension correlation matrix;
[0033] Inputting the spatial dimension correlation matrix and the temporal dimension correlation matrix into the spatiotemporal parallel multi-head attention layer to perform multi-head attention calculation, and obtaining the spatiotemporal parallel multi-head attention layer output;
[0034] Adding the output of the spatiotemporal parallel multi-head attention layer to the embedded encoded dynamic Raman spectrum data to obtain the output of the spatiotemporal parallel cross self-attention mechanism module;
[0035] Performing fourth-layer normalization processing and first multi-layer perceptron processing on the output of the spatiotemporal parallel cross self-attention mechanism module to obtain a first multi-layer perceptron output;
[0036] Adding the output of the first multi-layer perceptron to the output of the spatiotemporal parallel cross self-attention mechanism module to obtain the spatiotemporal dual-stream attention mechanism encoder output;
[0037] The spatiotemporal dual-stream attention mechanism encoder output is processed by a second multi-layer perceptron to obtain dynamic spectral features.
[0038] Preferably, the method for deeply cross-fusing the static spectral features with the dynamic spectral features includes:
[0039] Performing matrix multiplication on the static spectral feature and the dynamic spectral feature to obtain a result matrix;
[0040] The sum of squared differences of the result matrices is calculated to obtain a deep cross-fused spectral feature vector; wherein the sum of squared differences is used to measure the degree of difference between the result matrices.
[0041] The present invention also provides a bacterial resistance prediction system based on deep fusion of multimodal Raman spectroscopy data, which is used to implement the method, including:
[0042] A data acquisition module is used to acquire static Raman spectral data of target bacteria and target antibiotics, as well as dynamic Raman spectral data of target bacteria under the action of target antibiotics;
[0043] a feature extraction module, configured to perform embedding coding and feature extraction on the static Raman spectrum data and the dynamic Raman spectrum data, respectively, to obtain static spectrum features and dynamic spectrum features;
[0044] The drug resistance prediction module is used to deeply cross-fuse the static spectral features with the dynamic spectral features, and combine them with a multi-layer perceptron to obtain the drug resistance prediction results of the target bacteria.
[0045] Compared with existing technologies, the present invention offers the following advantages: a method and system for predicting bacterial resistance based on deep fusion of multimodal Raman spectral data is proposed. This method integrates static Raman spectral molecular information of bacteria / antibiotics, dynamic Raman spectral information from long-term continuous monitoring of single bacteria under antibiotic action, and information on drug resistance genes to perform multimodal deep feature extraction and interaction. This method exploits the correlations between multimodal data, predicts the dynamic development trend of bacterial resistance, and shortens the time required for clinical resistance detection. By combining time-resolved CARS microscopy (with a time resolution of 100ms / frame) with microfluidics, the system captures the molecular changes of bacteria under antibiotic action in real time and collects dynamic Raman spectral data. An innovative spatiotemporal dual-stream attention mechanism is employed to simultaneously mine real-time features of dynamic data from both spatial and temporal dimensions. A cross-modal cross-fusion module is constructed to effectively integrate interactive dynamic and static spectral features, and reinforcement learning and transfer learning are combined to optimize model performance. This technology overcomes the limitations of traditional drug susceptibility testing, which is time-consuming and provides limited information, providing a new solution for efficient, accurate, and interpretable drug resistance monitoring in the clinic. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] In order to more clearly illustrate the technical solution of the present invention, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0047] Figure 1 This is a flow chart of a method for predicting bacterial resistance based on deep fusion of multimodal Raman spectroscopy data according to an embodiment of the present invention;
[0048] Figure 2 This is a schematic diagram of uniform frame sampling according to an embodiment of the present invention;
[0049] Figure 3 This is a schematic diagram of the Transformer model structure for extracting static spectral features of bacteria / antibiotics according to an embodiment of the present invention;
[0050] Figure 4 A schematic diagram of the structure of a self-attention module with added bias according to an embodiment of the present invention;
[0051] Figure 5 This is a schematic diagram of the structure of the spatiotemporal dual-stream attention mechanism encoder according to an embodiment of the present invention;
[0052] Figure 6 Schematic diagram of the spatiotemporal parallel multi-head attention layer structure of an embodiment of the present invention. DETAILED DESCRIPTION
[0053] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0054] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the present invention is further described in detail below with reference to the accompanying drawings and specific embodiments.
[0055] Example 1
[0056] like Figure 1 As shown in the figure, the bacterial resistance prediction method based on deep fusion of multimodal Raman spectroscopy data includes:
[0057] S1: Obtain static Raman spectral data of target bacteria and target antibiotics, as well as dynamic Raman spectral data of target bacteria under the action of target antibiotics. The static Raman spectral data includes static Raman spectral data of bacteria and static Raman spectral data of antibiotics.
[0058] In this embodiment, the process of acquiring static Raman spectrum data includes:
[0059] (1) Determine target data
[0060] Bacterial data: Identify the bacterial species that need to be studied (such as Staphylococcus aureus and Escherichia coli) and their resistance phenotypes.
[0061] Antibiotic data: Select antibiotics related to target bacteria and determine their characteristic peak ranges.
[0062] (2) Select the database
[0063] Download static Raman spectroscopy data of common clinical bacteria and antibiotics from the following public databases:
[0064] 1) Bacterial Raman spectroscopy data: http: / / www.sers-cbp.net /
[0065] 2) Antibiotic Raman spectroscopy data: http: / / sers.test.bniu.net /
[0066] (3) Data retrieval and screening
[0067] When retrieving target data, a specific keyword combination must be entered. Bacterial data must include strain number, culture conditions (such as culture medium type, culture time) and drug resistance information (such as MIC value, drug resistance genotype), etc. Antibiotic data must include molecular structure and characteristic peak position.
[0068] The process of acquiring static Raman spectroscopy data includes:
[0069] (1) Data Acquisition
[0070] Dynamic Raman spectroscopy data can monitor the molecular changes of bacteria under the action of antibiotics in real time. This technology is usually used in combination with a microfluidic system and a time-resolved Raman microscope. The experiment uses a time-resolved CARS microscope or a confocal Raman microscope, combined with a microfluidic chip to precisely control the contact environment between bacteria and antibiotics. The laser wavelength is usually selected to be 785nm or 532nm to avoid fluorescence interference, the power is controlled at 5-20mW to prevent sample damage, the time resolution is set to 100ms-1s / frame to capture fast dynamic responses, and the flow rate of the microfluidic chip is maintained at 0.5-1μL / min to ensure that the bacteria are evenly exposed to the gradient concentration of antibiotics. The specific experimental steps are as follows:
[0071] 1) Sample preparation: The target bacteria were inoculated into the culture channel of the microfluidic chip and maintained in LB medium.
[0072] 2) Antibiotic introduction: Inject antibiotics through the side channel of the chip, and the concentration gradient is determined according to the preliminary experiment. Data acquisition: Start the time series acquisition, and the system will automatically collect spectral data at a fixed time interval of 5 minutes, covering the fingerprint area of 600-1800cm-1 and 2800-3200cm-1. -1 The CH vibration region was monitored continuously for 8 hours, ultimately generating a spectral dataset totaling 96 data sets. The raw data is a time series of Raman spectra recorded as a multidimensional array. Each frame contains spatial coordinates (XY) and a spectral intensity matrix (wavenumber-intensity), stored as a multidimensional array (time × space × wavenumber) or in HDF5 format.
[0073] A further embodiment includes preprocessing the acquired dynamic Raman spectral data. Dynamic Raman spectral data records the molecular response characteristics of bacteria under the action of antibiotics. However, due to interference from factors such as instrument noise, environmental interference, and the complexity of biological samples, the raw data often contains significant noise and baseline drift. To eliminate these interference factors, the system preprocesses the raw spectral data to improve data quality and ensure the accuracy of subsequent analysis.
[0074] Wavelet transform has multi-resolution analysis characteristics and can effectively separate signals and noise. Therefore, the present invention decomposes the signal into different scales by selecting appropriate wavelet basis functions and effectively removes the noise distributed in the high-frequency detail coefficients through threshold processing.
[0075] Preprocessing methods include:
[0076] Using discrete wavelet transform (DWT ), multi-scale analysis is performed on the dynamic Raman spectroscopy data to obtain high-frequency coefficients; the decomposition process is shown in formulas (1) and (2):
[0077] cA j+1 [n]=∑ k h[k-2n]cA j [k] (1)
[0078] cD j+1 [n]=∑ k g[k-2n]cA j [k] (2)
[0079] Among them, cA j+1 [n] and cD j+1 [n] represents the approximate coefficient obtained by decomposition at the j+1th layer, h[k-2n] and h[k-2n] are the low-pass filter coefficients, cA j [k] is the approximation coefficient of the jth layer.
[0080] The high-frequency coefficients are subjected to threshold processing, and the signal is reconstructed by combining the low-pass filter coefficients and the high-pass filter coefficients to obtain the denoised dynamic Raman spectrum data. Specifically, the soft threshold function is shown in formula (3).
[0081]
[0082] where sign(cD j [n]) is the sign function used to obtain cD j [n] is the symbol, T is the set threshold. (cD j [n]|-T) + Indicates when |cD j [n]|-When T>0, take cD j [n]|-T, otherwise 0.
[0083] Finally, the signal is reconstructed. The reconstruction process is shown in formula (4), where h[n] and g[n] are the low-pass and high-pass filter coefficients, respectively, and cA and cD are the approximation and detail coefficients.
[0084] cA j [n]=∑ k h[n-2k]cA j+1 [k]+∑ k g[n-2k]cD j+1 [k] (4)
[0085] The asymmetric least squares method is used to perform baseline correction on the denoised dynamic Raman spectroscopy data; after signal denoising, in order to eliminate the influence of sample autofluorescence and instrument system errors, the baseline drift problem in the signal needs to be further processed. During the dynamic monitoring process, the baseline drift manifests as a low-frequency background signal that changes with time, and its intensity is often 1-2 orders of magnitude higher than the Raman characteristic peak. This drift will seriously interfere with the identification and quantitative analysis of characteristic peaks. However, the traditional baseline correction method is not effective when processing dynamic data, so the present invention adopts an improved asymmetric least squares method, the objective function is shown in formula (5), and the weight update is shown in formula (6), where λ controls the smoothness and p controls the asymmetry:
[0086] min b ∑ i w i (y i -b i ) 2 +λ∑ i (Δ 2 b i ) 2 (5)
[0087]
[0088] where w i is the weight coefficient, y i is the original signal value, b i is the baseline estimate.
[0089] The dynamic Raman spectral data after baseline correction are subjected to area normalization and standard normal transformation, and the transformed dynamic Raman spectral data are time-corrected by the dynamic time warping algorithm to obtain the pre-processed dynamic Raman spectral data. Specifically, after completing the baseline correction, the spectral intensity collected at different time points may produce systematic differences due to factors such as laser power fluctuations and sample concentration changes, and normalization processing is required to eliminate systematic deviations. Area normalization eliminates the overall intensity change by dividing by the total area of the spectrum, retains the relative peak intensity information, and studies the change in the characteristic peak intensity ratio. The area normalization formula is shown in (7), where Δx j is the wave number interval:
[0090]
[0091] The standard normal transformation performs mean centering and standard deviation scaling on the intensity of each wave number point, which is suitable for multi-sample integrated analysis, where μ and σ are the mean and standard deviation of all spectra at the wave number, respectively. The standard normal transformation is shown in formula (8):
[0092]
[0093] During the dynamic data collection process of microfluidic experiments, unstable flow rate will cause the collection time point to shift, and time correction is required to ensure the accuracy of dynamic analysis. DTW (Dynamic Time Warping) aligns time series by finding the minimum cumulative distance path, allowing nonlinear time expansion and contraction to handle dynamic data with unstable flow rate. The distance matrix is shown in formula (9), where D(i, j) represents the minimum cumulative distance from the starting point to (i, j), represents the value of the first time series at the i-th moment, represents the value of the second time series at the jth moment, i represents the time index of the first series, and j represents the time index of the second series. The cumulative distance is shown in formula (10), where γ(i, j) represents the total distance of the optimal path from the starting point to (i, j), W(i, j) represents the local distance of the current point, and min{} represents the minimum cumulative distance of the previous step:
[0094]
[0095] The preprocessing described above ultimately yielded a temporal-spatial spectral dataset with a high signal-to-noise ratio and a stable baseline. This process effectively eliminated instrumental errors and environmental interference, significantly improving data quality and enabling the clear visualization of weak Raman peaks.
[0096] This example also collects drug resistance data and antibiotic-bacteria affinity data:
[0097] Acquisition of drug resistance data:
[0098] CARD (https: / / card.mcmaster.ca / ) is a comprehensive antibiotic resistance gene database that includes: sequence, functional, and classification information of antibiotic resistance genes (AROs), resistance mechanisms (such as enzymatic degradation and efflux pumps), and data related to clinical phenotypes (such as MIC values). CARD provides multiple data files, with the core data package containing all resistance genes, protein sequences, and Ontology (AROs). The files include the resistance gene ontology (structured classification), resistance protein sequences (FASTA format), resistance gene nucleic acid sequences, and resistance gene prediction rules.
[0099] (1) Search method
[0100] Keyword search (suitable for quick location): Enter a gene name, antibiotic, or species. For example, searching for blaCTX-M will retrieve annotations for all CTX-M type β-lactamase genes.
[0101] Advanced filtering (for batch acquisition): In the "Advanced Search" module, select genes, resistance mechanisms, associated pathogens and other filtering conditions to filter out the required data.
[0102] (2) Data download
[0103] Download a single record: Click the target gene on the search results page and enter the details page to download: FASTA sequence (gene / protein), annotation information (JSON / TSV format, including resistance rules and references).
[0104] Batch download: Get the complete dataset through the "Download" page.
[0105] (3) Screening of valid data:
[0106] Duplicate or incompletely annotated records, such as those lacking descriptions of resistance mechanisms, were eliminated.
[0107] Genes with experimental verification were retained and marked as “Experimental Evidence” in CARD.
[0108] Acquisition of antibiotic-bacteria affinity data:
[0109] The PubChem database (https: / / pubchem.ncbi.nlm.nih.gov / ) contains over 110 million compounds (including antibiotics, small molecule drugs, etc.) and provides biological activity data (such as binding affinity, inhibition constant, MIC value, etc.), including compound-target interaction information (such as binding data of antibiotics and bacterial proteins).
[0110] Enter the PubChem BioAssay library and enter keywords such as antibiotic names or bacterial names in the search bar. Alternatively, use the advanced search function to limit the search to experimental types, species, and other criteria. This will screen for biochemical experimental data related to antibiotic-bacteria interactions, looking for results that directly measure affinity or related activities. Once you've found entries related to your target antibiotic or bacteria, review the "Related Data" section and use literature citations, patent information, and other resources to identify related studies and obtain clues to affinity data.
[0111] S2: Static Raman spectral data and dynamic Raman spectral data are embedded and feature extracted to obtain static and dynamic spectral features. Addressing the inability of existing technologies to monitor the dynamic changes in drug resistance in real time, this model predicts the dynamic development trend of bacterial drug resistance through the fusion of multimodal data, significantly shortening clinical resistance detection time and assisting doctors in accurately determining medication dosage and timing.
[0112] However, there are many difficulties in extracting features from dynamic data. In terms of spatiotemporal characteristics, the molecular changes in bacteria under the action of antibiotics are continuous and rapid in the temporal dimension, and vary in the spatial dimension depending on the different regions within the cell. Therefore, accurately capturing the key features of spatiotemporal changes has become an important issue that needs to be addressed. In addition, dynamic data features are highly heterogeneous. The dynamic response patterns produced by different bacterial species and different antibiotics vary greatly, making it difficult to use unified standards and methods for feature extraction and analysis. This further increases the difficulty of building a universal feature extraction model.
[0113] Traditional Transformers struggle to effectively extract both temporal and spatial features simultaneously. Their self-attention mechanism primarily captures sequential dependencies within a sequence, without explicitly distinguishing between the temporal and spatial dimensions. When processing dynamic Raman spectroscopy data for bacterial resistance monitoring, traditional Transformers struggle to effectively address the complex spatiotemporal variations of this data. Furthermore, instrument noise, sample complexity, environmental interference, and the high heterogeneity of data features make it difficult for traditional Transformers to accurately extract effective spatiotemporal features. Furthermore, they cannot use a unified approach to process complex and diverse spatiotemporal features, making it impossible to efficiently extract both temporal and spatial features simultaneously.
[0114] A further embodiment is that the method for embedding encoding static Raman spectrum data includes:
[0115] Static Raman spectral data is embedded by performing matrix multiplication on the input using a preset weight matrix. Specifically, the raw Raman spectral data is matrix multiplied using a learnable weight matrix. The raw Raman spectral data is defined as x, with dimension n0, and mapped to a low-dimensional space y with dimension m0. A weight matrix W0 is defined with dimensions m0 × n0. The low-dimensional vector y is obtained by performing the operation y = W0x. This process effectively reduces the data dimensionality while preserving key features.
[0116] The embedded static Raman spectrum data is position-encoded to complete the embedded coding of the static Raman spectrum data. Specifically, the embedded static Raman spectrum data is position-encoded. For a dt-dimensional vector y, its i0th element is The position encoding calculation method is shown in formulas (11) and (12):
[0117]
[0118]
[0119] Where pos represents the position of the data in the sequence, i0 represents the index (dimension) of the vector element, and dt represents the dimension of the embedding vector. In this way, position encoding information is added to each data point, and the final embedded encoded static data is obtained for subsequent feature extraction.
[0120] A further embodiment is that the method for embedding the pre-processed dynamic Raman spectroscopy data includes:
[0121] A token embedding method is used to discretize the preprocessed dynamic Raman spectral data according to its time series and spatial distribution, obtaining a token embedded representation. In this embodiment, the token embedding method is used to discretize the dynamic Raman spectral data according to its time series and spatial distribution, considering its changing characteristics in the temporal and spatial dimensions, thereby converting it into a token sequence suitable for model processing. For example, bacterial Raman spectral data at different time points is divided into multiple time slices, and the data within each time slice is then spatially partitioned and processed, and each small block of data is converted into a token.
[0122] Therefore, the present invention divides the dynamic data collected from each strain at 5 minutes into nt=96 frames in time order, i.e., nt images. Each frame of image is spatially divided into non-overlapping nw blocks, i.e., x n1 …x nw Each block can be regarded as a token. All tokens of nt images are connected in sequence to form an input sequence, which contains N = nt × nw tokens in total. The feature dimension is set to d, and the final embedding expression of the token is defined as like Figure 2 shown.
[0123] Add position embedding to the token embedding representation to obtain a token embedding representation with position information, and complete the embedding encoding of the preprocessed dynamic Raman spectroscopy data. Specifically, then add position embedding to these N non-overlapping image blocks. Its encoding method is the same as the position encoding method of Transformer, which can be referred to formulas (11) and (12). After token embedding and position embedding, a token embedding representation with position information is obtained, which is defined as z, as shown in formula (13). z is composed of the learnable embedding vector z cls and each image block embedding vector E k etc., plus the position code p0. cls To classify tokens, the projection of E is equivalent to a two-dimensional convolution operation. In order to preserve the position information, a learnable position embedding p0∈R is added to z N×d, giving each token its position information in the time and space sequence so that the model can perceive the order and position relationship of the data.
[0124] z=[z cls ,Ex1,Ex2,…,Ex N ]+p0 (13)
[0125] In the process of building a high-precision bacterial resistance prediction model, feature extraction is a crucial core link. Its accuracy and efficiency play a decisive role in the overall performance of the model. The feature extraction of static data and dynamic data each has its own unique process and significance.
[0126] First, the embedded coded bacterial Raman data and antibiotic Raman data are input into independent Transformer modules respectively. After being processed by the S-layer Transformer mechanism, the features of bacterial Raman data and antibiotic Raman data are output respectively. The specific value of S is learned after model training, such as Figure 3 As shown. Each layer of the Transformer mechanism is mainly composed of a self-attention mechanism, which uses the self-attention mechanism to extract the Raman spectral data features of bacteria and antibiotics and mine the potential features and patterns in the data. It is worth noting that the present invention adds a bias amount about the affinity of antibiotics and bacteria to the multi-head attention layer based on the basic self-attention mechanism, so that the model can achieve local attention to the correlation between antibiotics and bacteria, as shown in Figure 2. Figure 4 As shown, the bias introduces a scaled dot-product attention mechanism layer in the self-attention module.
[0127] A further implementation method is to use a Transformer encoder with a preset number of layers (S layers) to extract features from the static Raman spectral data after embedding coding, wherein the Transformer encoder includes a self-attention mechanism module, a pre-feedback layer, and a second-layer normalization; the self-attention mechanism module includes a multi-head attention layer and a first-layer normalization. The processing process of bacteria and antibiotics is the same. Taking bacteria as an example here, the method of using the Transformer encoder to extract features from the static Raman spectral data after embedding coding includes:
[0128] For the static Raman spectrum data X after embedding coding B Perform linear transformation to obtain the linear matrix (query) Q B 、(key)K B and (value)V B , and Q B =K B =V B .
[0129] Introduce the deviation of antibiotic-bacteria affinity into the multi-head attention layer, process the linear matrix, and obtain the output of the multi-head attention layer; specifically, input these three matrices into the multi-head attention layer respectively, and B and K B Perform the dot product calculation and divide the result by And add the bias value, and then use softmax to get V B The weight of the multi-head attention layer is finally obtained, Attention(Q B ,K B ,V B ) can be expressed as:
[0130]
[0131] Among them, d k It's Q B , K B and V B dimension, bias is the deviation value calculated based on the antibiotic-bacteria affinity data.
[0132] Each attention mechanism consists of h parallel multi-head attention layers, 3h+1 linear layers, and 1 connection operation, where h is the number of heads and can be learned. Define the hth i The output of the multi-head attention layer is head b , b=1,2,…,h, then head a It can be expressed as:
[0133]
[0134] Here, the input Q=K=V of the linear layer is the input X of the Transformer module. B . is the linear projection matrix.
[0135] Finally, the output of the multi-head attention layer is connected and passed to the linear layer to obtain the output of the attention mechanism layer. MultiHead(Q,K,V) can be expressed as
[0136] MultiHead(Q,K,V)=Concat(head1,...,head h )W O (16)
[0137] Among them, W O is the linear projection matrix.
[0138] Perform first-layer normalization on the output of the multi-head attention layer to obtain the first-layer normalized output;
[0139] After adding the first layer normalized output to the embedded encoded static Raman spectrum data, the output of the self-attention mechanism module is obtained; specifically, the final output Y B , which can be expressed as:
[0140] Y B =LN(MultiHead(Q,K,V))+X B (17)
[0141] Among them, LN() represents the layer normalization operation.
[0142] The output of the self-attention mechanism module is input to the pre-feedback layer, and the pre-feedback output is added to the output of the self-attention mechanism module for second-layer normalization to obtain the Transformer encoder output;
[0143] The output of the previous layer of Transformer encoder is used as the input of the next layer of Transformer encoder until the output of the last layer of Transformer encoder is obtained to obtain the static spectral features. B The input is fed into the pre-feedback mechanism, added to itself, and then normalized once to obtain the final output, which can be expressed as:
[0144] Y outB =LN(Y B +FeedForword(Y B )) (18)
[0145] Among them, Y outB is the output of the first-layer Transformer module, which is then fed into the second-layer Transformer module as input. After processing in this way, the output of the S-th layer Transformer module is obtained. That is, the feature extraction result obtained by the bacteria after the Transformer encoder is processed. It will be sent to the cross-fusion module as input in the future, and deeply interact with the dynamic data features and antibiotic static data feature extraction results. The antibiotic static data feature extraction process is the same as the bacterial static data feature extraction, so it will not be repeated here. The feature extraction result is expressed as Where S is the number of layers of the Transformer encoder used to extract static data features of antibiotics, which can be obtained through learning.
[0146] The embedded, encoded dynamic data is fed into the spatiotemporal dual-stream attention mechanism. The spatiotemporal dual-stream attention encoder, also leveraging self-attention, comprehensively extracts features from the data across both time and space. The self-attention mechanism automatically calculates the association weights between each token and all other tokens, capturing the complex dependencies of the data across time and space, successfully extracting high-level features from the dynamic data.
[0147] A further implementation method is to use a preset number of layers of spatiotemporal dual-stream attention mechanism encoder combined with a second multi-layer perceptron to extract features from the embedded encoded dynamic Raman spectroscopy data, wherein the spatiotemporal dual-stream attention mechanism encoder includes a spatiotemporal parallel cross self-attention mechanism module, a fourth normalization layer, and a first multi-layer perceptron; the spatiotemporal parallel cross self-attention mechanism module includes a third normalization layer and a spatiotemporal parallel multi-head attention layer;
[0148] The method for extracting features from dynamic Raman spectroscopy data after embedding coding by using a spatiotemporal dual-stream attention mechanism encoder combined with a second multi-layer perceptron includes:
[0149] The dynamic Raman spectral data (dynamic Raman matrix z) after embedding encoding is reshaped to obtain spatial dimension input representation and time dimension input representation; specifically, the dynamic Raman matrix z after token embedding and position embedding is input into the spatiotemporal dual-stream attention module. After being processed by L layers of spatiotemporal dual-stream attention mechanism encoder, the bacterial dynamic Raman spectral feature extraction result is output. The specific value of L is learned after model training. Each layer of spatiotemporal dual-stream attention mechanism encoder consists of a spatiotemporal parallel cross self-attention mechanism, layer normalization (LN) and a multi-layer perceptron (MLP), as shown in Figure 2. Figure 5 shown.
[0150] The spatial dimension input representation and the time dimension input representation are normalized in the third layer to obtain the third layer normalized output; then, the input of the lth (l=1,2,…,L) layer is defined as z l , where the input z of the first layer is 1 =z. l Reshape to obtain input expressions of spatial and temporal dimensions respectively and Then perform layer normalization on the result and It is sent to the spatiotemporal parallel multi-head attention layer for parallel multi-head self-attention operations in the time dimension and space dimension, such as Figure 6 As shown, the spatiotemporal parallel multi-head attention layer includes a spatial head and a temporal head, and both the spatial head and the temporal head include a scaled dot product attention mechanism layer.
[0151] Perform linear transformation on the normalized output of the third layer to obtain the spatial dimension correlation matrix and the temporal dimension correlation matrix. Here, the multi-head attention operation of the spatial dimension and the temporal dimension is represented by MSA(), which is similar to the multi-head attention mechanism of Transformer. and The results are linearly transformed to obtain the spatial dimension correlation matrix Q s , K s and V s , and the time dimension correlation matrix Q t , K t and V t, Get the label Y for the spatial dimension s and the time dimension is labeled Y t Among them, d k1 It's Q s , K s and V s The dimension, d k2 It's Q t , K t and V t dimension.
[0152]
[0153] The spatial dimension correlation matrix and the temporal dimension correlation matrix are input into the spatiotemporal parallel multi-head attention layer for multi-head attention calculation to obtain the spatiotemporal parallel multi-head attention layer output; specifically, the spatial dimension and temporal dimension multi-head attention layers are respectively composed of h s and h t It consists of multiple attention layers running in parallel, where h s and h t The multi-head attention mechanism in spatial and temporal dimensions is similar to that of Transformer. You can refer to formulas (15) and (16). s and Y t Calculate the output of the multi-head attention module in spatial and temporal dimensions, respectively, as and After concatenating and linearly processing the results, we finally get the output Y of the spatiotemporal parallel multi-head attention module. l , which is the output of the MSA() operation. The calculation process is shown in formula (21), where W b is a linear weight.
[0154]
[0155] The output of the spatiotemporal parallel multi-head attention layer is added to the dynamic Raman spectrum data after embedding encoding to obtain the output of the spatiotemporal parallel cross-self-attention mechanism module; specifically, the output of the spatiotemporal parallel cross-self-attention module is defined as y l , and its calculation method is shown in formula (22).
[0156]
[0157] The output of the spatiotemporal parallel cross self-attention mechanism module is normalized in the fourth layer and processed by the first multi-layer perceptron to obtain the output of the first multi-layer perceptron; next, the output y l The layer normalization processing LN() and the multi-layer perceptron processing MLP() are performed in sequence, and finally the output of the spatiotemporal dual-stream attention mechanism encoder of the lth layer is obtained, which is also the input of the spatiotemporal dual-stream attention mechanism encoder of the next layer, defined as z l+1 , calculated as shown in formula (23).
[0158] z l+1 =MLP(LN(y l ))+y l (twenty three)
[0159] Add the output of the first multi-layer perceptron to the output of the spatiotemporal parallel cross self-attention mechanism module to obtain the spatiotemporal dual-stream attention mechanism encoder output;
[0160] The output of the spatiotemporal dual-stream attention mechanism encoder is processed by the second multi-layer perceptron to obtain dynamic spectral features. Specifically, z l+1 As input, it is sent to the next layer (the l+1 layer) of the spatiotemporal dual-stream attention mechanism module, and after processing in the same way, the output of the L-th layer spatiotemporal dual-stream attention mechanism module is obtained. This is the output of the spatiotemporal dual-stream attention mechanism encoder. It is processed by the multi-layer perceptron to obtain the feature extraction result of the dynamic spectral data, which is expressed as Y outD This is then fed into the cross-fusion module as input, interacting deeply with the feature extraction results of the bacterial and antibiotic static data. This innovative spatiotemporal parallel cross-self-attention strategy not only retains the global modeling capabilities of the traditional self-attention mechanism, but also significantly improves the model's accuracy in analyzing microbial metabolic dynamics by decoupling features in the spatial and temporal dimensions.
[0161] S3: Deeply cross-fuse static spectral features with dynamic spectral features and combine them with a multi-layer perceptron to obtain the drug resistance prediction results of the target bacteria.
[0162] A further embodiment is that the method for deeply cross-fusing static spectral features with dynamic spectral features includes:
[0163] Perform matrix multiplication on the static spectral features and the dynamic spectral features to obtain a result matrix;
[0164] The sum of squared differences of the result matrices is calculated to obtain the spectral feature vector of deep cross fusion; the sum of squared differences is used to measure the degree of difference between the result matrices.
[0165] In the field of bacterial resistance prediction research, different features originate from different information channels. Integrating these features can construct a comprehensive multimodal description and achieve a more accurate modeling process. To achieve accurate prediction of bacterial resistance, the dynamic data features extracted by the spatiotemporal dual-stream attention mechanism encoder and the static data features of bacteria and antibiotics extracted by the Transformer need to be sent to the cross-fusion module for deep cross-fusion. Traditional cross-fusion methods, such as feature concatenation, simply connect features from different sources in sequence without considering the intrinsic correlation and importance differences between features. This will obscure key features and affect the accuracy of bacterial resistance prediction. Although weighted summation takes certain weights into account, since the weights are usually pre-set or simply learned, it is difficult to accurately capture the true relationship between different feature types in complex scenarios.
[0166] In clinical testing, bacterial resistance analysis faces the challenge of integrating multi-source, dynamic, and complex features. This is especially true when test samples contain multiple pathogens with overlapping Raman spectral signatures. Traditional methods often struggle to effectively distinguish and integrate this information, limiting model performance. Therefore, to accurately predict bacterial resistance, this paper employs a novel cross-fusion approach.
[0167] The dynamic feature extraction result of the spatiotemporal dual-stream attention mechanism encoder is represented as H d =Y outD , the static feature extraction results of bacteria and antibiotics through Transformer are represented as H b and H a , and The matrix obtained by matrix multiplication of these three eigenvectors is used to calculate the loss function and assist in model optimization, as shown in formulas (24), (25) and (26):
[0168] r1=H d T H b (twenty four)
[0169] r2=H d T H a (25)
[0170] r3=H a TH b (26)
[0171] Next, the formula for calculating the sum of squared differences is used to measure the degree of difference between the two sets of values by calculating the sum of squared differences. The smaller the difference, the smaller the loss value and the better the fusion effect, as shown in formulas (27) and (28):
[0172] L1=∑(r1-r3) 2 (27)
[0173] L2=∑(r2-r3) 2 (28)
[0174] L1 and L2 reflect the degree of difference between these eigenvectors. They integrate the information of two eigenvectors into a scalar value, which measures the similarity or difference between the two eigenvectors. To a certain extent, they fuse their feature information and prepare for subsequent analysis.
[0175] Then the results L1 and L2 obtained through the cross-fusion and loss calculation steps are fed into the multi-layer perceptron (MLP). MLP further extracts key information from features through multiple layers of linear and nonlinear transformations and explores more complex high-order relationships between features. Its output is recorded as h out Compared with the previous feature fusion stage, the feature representation after MLP processing is more abstract and refined, effectively removing redundant information and highlighting the feature dimensions that are more valuable for model decision-making.
[0176] Finally, the feature output h obtained after MLP processing is out Combined with the parameter weights determined by the previously downloaded drug resistance data. Suppose the weight matrix obtained after preprocessing the downloaded drug resistance data is (n1 is the number of drug resistance categories, m1 is the dimension of the MLP output feature vector), the bias vector is b r ∈R n . The feature vector h output by MLP out With the weight matrix W r Perform matrix multiplication and add the bias vector b r , we get the predicted score vector s0, as shown in formula (29):
[0177] s0=W r h out +b r (29)
[0178] The predicted score vector s is then processed by the Softmax function and converted into a probability distribution, as shown in formula (30):
[0179]
[0180] Among them, each element in p0 Indicates that bacteria show nth c The probability of drug resistance is determined based on the probability value. A probability threshold τ is set. If max(p0)>τ, the bacteria are judged to be resistant to the antibiotic, and the category corresponding to the maximum value is the specific resistance type; if max(p0)≤τ, the bacteria are judged to be sensitive to the antibiotic.
[0181] In this embodiment, a specific implementation process of model training and optimization is provided:
[0182] (1) Dataset division
[0183] To ensure effective training and accurate evaluation of the model, this paper scientifically divides the dynamic dataset:
[0184] Training set (5 sets, 0-8 hours of data): Contains complete dynamic Raman spectroscopy data (96 frames, collected every 5 minutes), covering the complete dynamic response process of bacteria under the action of antibiotics.
[0185] Test set (1 set, 0-2 hours data + prediction for the next 6 hours): Only the first 2 hours of data are provided, and the model is required to predict the resistance development trend for the next 6 hours.
[0186] This division simulates the need to predict the dynamic changes of bacterial resistance based on early data in clinical scenarios.
[0187] (2) Training and testing
[0188] First, the model is trained end-to-end using the training set (0-8 hours of data) to optimize the model parameters. The loss function uses a combination of mean square error (MSE) and cross-entropy to optimize the dynamic feature fitting and drug resistance classification tasks, respectively. Next, the first 2 hours of data from the test set are input into the model, and the prediction results for the next 6 hours (such as the probability of drug resistance, the trend of changes in key metabolic peaks, etc.) are output. Finally, the predicted development trend data for the next 6 hours is compared with the 6 hours of dynamic data corresponding to the actual 2-8 hours of the test set.
[0189] (3) Error analysis and optimization
[0190] The mean square error (MSE) is used to measure the average square error between the predicted value and the true value. MSE can be expressed as:
[0191]
[0192] where n 0is the number of test set samples, It is the p f The predicted value of the sample, It is the p f The true value of the samples.
[0193] In addition, the mean absolute error (MAE) can be calculated to more intuitively reflect the average absolute deviation between the predicted value and the true value. MAE can be expressed as:
[0194]
[0195] Finally, a reasonable error threshold is set based on clinical needs. This threshold is used to determine the acceptability of the model's prediction results. For example, if the error in bacterial resistance prediction needs to be kept within a certain range in clinical applications to ensure the accuracy of medication decisions, a corresponding MSE or MAE threshold can be set based on this requirement. Next, the calculated error index is compared with the set error threshold. If the error of the test set is greater than the threshold, it indicates that the model performs poorly in predicting the test set data and requires further training and optimization.
[0196] Through the above "training-evaluation-optimization" cycle, the performance of the model is continuously improved until the error of the prediction rate is less than the set threshold, meets the clinical requirements, and is finally used for actual bacterial resistance detection.
[0197] Example 2
[0198] The present invention also provides a bacterial resistance prediction system based on deep fusion of multimodal Raman spectroscopy data, which is used to implement the method, including:
[0199] A data acquisition module is used to acquire static Raman spectral data of target bacteria and target antibiotics, as well as dynamic Raman spectral data of target bacteria under the action of target antibiotics;
[0200] A feature extraction module is used to perform embedding coding and feature extraction on static Raman spectrum data and dynamic Raman spectrum data respectively to obtain static spectrum features and dynamic spectrum features;
[0201] The drug resistance prediction module is used to deeply cross-fuse static spectral features with dynamic spectral features and combine them with a multi-layer perceptron to obtain the drug resistance prediction results of the target bacteria.
[0202] The embodiments described above are merely descriptions of preferred embodiments of the present invention and are not intended to limit the scope of the present invention. Without departing from the spirit of the present invention, various modifications and improvements made to the technical solutions of the present invention by persons skilled in the art should fall within the scope of protection defined by the claims of the present invention.
Claims
1. A bacterial resistance prediction method based on deep fusion of multimodal Raman spectroscopy data, characterized by: include: Obtain static Raman spectral data of target bacteria and target antibiotics, as well as dynamic Raman spectral data of target bacteria under the action of target antibiotics; Performing embedded coding and feature extraction on the static Raman spectrum data and the dynamic Raman spectrum data respectively to obtain static spectrum features and dynamic spectrum features; The static spectral features and the dynamic spectral features are deeply cross-fused and combined with a multi-layer perceptron to obtain the drug resistance prediction results of the target bacteria.
2. The method according to claim 1, characterized in that The method further includes preprocessing the acquired dynamic Raman spectrum data, wherein the preprocessing method includes: Using discrete wavelet transform, multi-scale analysis is performed on the dynamic Raman spectroscopy data to obtain high-frequency coefficients; Performing threshold processing on the high-frequency coefficients, combining low-pass filter coefficients and high-pass filter coefficients, reconstructing the signal, and obtaining denoised dynamic Raman spectroscopy data; The baseline correction of the denoised dynamic Raman spectroscopy data was performed using the asymmetric least squares method. The baseline-corrected dynamic Raman spectral data are subjected to area normalization and standard normal transformation, and the transformed dynamic Raman spectral data are subjected to time correction using a dynamic time warping algorithm to obtain the pre-processed dynamic Raman spectral data.
3. The method according to claim 1, characterized in that The method for embedding and encoding the static Raman spectrum data includes: Performing a matrix multiplication operation on the static Raman spectrum data using a preset weight matrix to complete input embedding of the static Raman spectrum data; Position encoding is performed on the static Raman spectrum data that has been embedded, thereby completing the embedded encoding of the static Raman spectrum data.
4. The method according to claim 2, characterized in that The method for embedding and encoding the pre-processed dynamic Raman spectroscopy data includes: The pre-processed dynamic Raman spectroscopy data is discretized according to time series and spatial distribution using a token embedding method to obtain a token embedding representation; Position embedding is added to the token embedding representation to obtain the token embedding representation with position information, completing the embedding encoding of the preprocessed dynamic Raman spectroscopy data.
5. The method according to claim 1, characterized in that A Transformer encoder with a preset number of layers is used to extract features from static Raman spectral data after embedding coding, wherein the Transformer encoder includes a self-attention mechanism module, a pre-feedback layer, and a second-layer normalization; the self-attention mechanism module includes a multi-head attention layer and a first-layer normalization; the method for extracting features from static Raman spectral data after embedding coding using the Transformer encoder includes: Perform linear transformation on the static Raman spectrum data after embedding coding to obtain a linear matrix; Introducing the bias of the antibiotic-bacteria affinity into the multi-head attention layer, processing the linear matrix, and obtaining the multi-head attention layer output; Perform first-layer normalization on the output of the multi-head attention layer to obtain the first-layer normalized output; The output of the self-attention mechanism module is obtained by adding the normalized output of the first layer to the static Raman spectrum data after embedding encoding; Inputting the output of the self-attention mechanism module into the pre-feedback layer, and adding the pre-feedback output and the output of the self-attention mechanism module to perform a second-layer normalization process to obtain a Transformer encoder output; The output of the previous layer of Transformer encoder is used as the input of the next layer of Transformer encoder until the output of the last layer of Transformer encoder is obtained to obtain the static spectral features.
6. The method according to claim 1, characterized in that A spatiotemporal dual-stream attention mechanism encoder with a preset number of layers is combined with a second multi-layer perceptron to perform feature extraction on the dynamic Raman spectral data after embedding encoding, wherein the spatiotemporal dual-stream attention mechanism encoder includes a spatiotemporal parallel cross self-attention mechanism module, a fourth normalization layer, and a first multi-layer perceptron; the spatiotemporal parallel cross self-attention mechanism module includes a third normalization layer and a spatiotemporal parallel multi-head attention layer; The method for extracting features from dynamic Raman spectroscopy data after embedding coding by using a spatiotemporal dual-stream attention mechanism encoder combined with a second multi-layer perceptron includes: Reshape the embedded encoded dynamic Raman spectroscopy data to obtain spatial dimension input representation and time dimension input representation; Performing third-layer normalization processing on the spatial dimension input representation and the time dimension input representation to obtain a third-layer normalized output; Performing linear transformation on the normalized output of the third layer to obtain a spatial dimension correlation matrix and a temporal dimension correlation matrix; Inputting the spatial dimension correlation matrix and the temporal dimension correlation matrix into the spatiotemporal parallel multi-head attention layer to perform multi-head attention calculation, and obtaining the spatiotemporal parallel multi-head attention layer output; Adding the output of the spatiotemporal parallel multi-head attention layer to the embedded encoded dynamic Raman spectrum data to obtain the output of the spatiotemporal parallel cross self-attention mechanism module; Performing fourth-layer normalization processing and first multi-layer perceptron processing on the output of the spatiotemporal parallel cross self-attention mechanism module to obtain a first multi-layer perceptron output; Adding the output of the first multi-layer perceptron to the output of the spatiotemporal parallel cross self-attention mechanism module to obtain the spatiotemporal dual-stream attention mechanism encoder output; The spatiotemporal dual-stream attention mechanism encoder output is processed by a second multi-layer perceptron to obtain dynamic spectral features.
7. The method according to claim 1, characterized in that The method for deeply cross-fusing the static spectral features with the dynamic spectral features includes: Performing matrix multiplication on the static spectral feature and the dynamic spectral feature to obtain a result matrix; The sum of squared differences of the result matrices is calculated to obtain a deep cross-fused spectral feature vector; wherein the sum of squared differences is used to measure the degree of difference between the result matrices.
8. A bacterial resistance prediction system based on deep fusion of multimodal Raman spectroscopy data, used to implement the method according to any one of claims 1 to 7, characterized in that: include: A data acquisition module is used to acquire static Raman spectral data of target bacteria and target antibiotics, as well as dynamic Raman spectral data of target bacteria under the action of target antibiotics; a feature extraction module, configured to perform embedding coding and feature extraction on the static Raman spectrum data and the dynamic Raman spectrum data, respectively, to obtain static spectrum features and dynamic spectrum features; The drug resistance prediction module is used to deeply cross-fuse the static spectral features with the dynamic spectral features, and combine them with a multi-layer perceptron to obtain the drug resistance prediction results of the target bacteria.
Citation Information
Patent Citations
Scale-adaptive microbial Raman spectrum detection method and system
CN112986210A
Fuel oil Raman spectrum determination method and system
CN116008248A
Classification and identification method of antibiotic drugs based on deep learning and Raman spectrum
CN116028863A
High-temporal-spatial-resolution multiplexing wide-field Raman imaging system and method based on flat-top illumination
CN116577316A
Method for identifying microorganism or detecting its morphology alteration using surface enhanced ramam scattering (SERS)
EP2063253A2
Cited By
Raman spectrum pathogen rapid identification and detection system
CN121997192A
Surface enhanced Raman detection method and system for monitoring dynamic evolution of bacterial drug resistance, storage medium and electronic equipment
CN122171522A