A bone infection pathogenic bacteria identification model, a construction method thereof and an application thereof

CN122780948APending Publication Date: 2026-09-18NANFANG HOSPITAL OF SOUTHERN MEDICAL UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202611007237.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-07
Publication Date
2026-09-18

AI Technical Summary

Technical Problem

但该技术落地面临关键瓶颈:骨感染病灶标本获取需通过有创操作,临床样本量本就稀缺,且致病菌纯化过程中易出现污染导致有效光谱数据量进一步减少,难以满足深度学习模型对大规模训练数据的需求,导致SERS与深度学习的融合应用始终停留在实验室阶段,无法向临床转化

Benefits of technology

本发明针对骨感染致病菌样本稀缺场景,设计基于生成对抗网络的样本数据增强整合策略,实现数据扩充,进一步进行表面增强拉曼光谱和 SpecGAN-LSTM的融合,开发骨感染致病菌识别模型及方法等方案,实现骨感染致病菌的快速精准识别,检测效率显著提升,且经准确率、精确率等指标验证,识别性能稳定可靠,准确率可达96%以上。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122780948A_ABST
    Figure CN122780948A_ABST
Patent Text Reader

Abstract

This invention relates to a model for identifying pathogenic bacteria causing bone infections, its construction method, and its application. The construction method includes: acquiring surface-enhanced Raman spectral data of pathogenic bacteria from bone infection lesion samples and preprocessing it; dividing the preprocessed spectral data into training and testing sets; using a generative adversarial network (GAN) to enhance and integrate the training set data; and using the enhanced and integrated training set to train a SpecGAN-LSTM model to obtain the model for identifying pathogenic bacteria causing bone infections. This invention addresses the scarcity of pathogenic bacteria samples causing bone infections by designing a sample data enhancement and integration strategy based on a GAN to expand the data. Further, it integrates surface-enhanced Raman spectroscopy and SpecGAN-LSTM to develop a model and method for identifying pathogenic bacteria causing bone infections. This achieves rapid and accurate identification of pathogenic bacteria causing bone infections, significantly improving detection efficiency. Furthermore, the identification performance has been verified by indicators such as accuracy and precision, demonstrating stable and reliable performance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of biomedical detection and machine learning technology, and relates to a bone infection pathogen identification model, its construction method and application. Background Technology

[0002] Bone infection is a common and refractory infectious disease in orthopedic clinics, mainly including acute suppurative osteomyelitis, chronic osteomyelitis, and periprosthetic infection. This disease has an insidious onset and rapid progression; if the causative pathogen is not identified and targeted treatment is not initiated in a timely manner, it can easily lead to bone necrosis, limb dysfunction, and even amputation, seriously threatening the patient's life and health. In clinical practice, accurate diagnosis and treatment of bone infection highly depend on etiological diagnostic results; therefore, developing efficient and accurate pathogen detection technologies is crucial to improving the treatment outcomes of bone infection.

[0003] Currently, the mainstream clinical methods for detecting pathogens causing bone infections are mainly based on traditional biological testing. The core process is "specimen collection - in vitro culture - bacterial identification - drug sensitivity testing". This method obtains pure colonies of pathogens through in vitro culture and then completes identification by combining biochemical reactions or mass spectrometry. Although it has the advantage of high reliability, it has significant shortcomings: First, the detection cycle is long. The culture and identification of common aerobic bacteria (such as Staphylococcus aureus and Staphylococcus epidermidis) takes 3 to 4 days, and if anaerobic bacteria or drug-resistant bacteria are involved, it takes more than 7 days. This is far from meeting the urgent clinical need for "precise medication within 48 hours". As a result, clinicians often rely on empirical use of broad-spectrum antibiotics, which increases the risk of drug-resistant bacteria and may delay treatment due to inappropriate medication. Second, the operation process is cumbersome, requiring multiple rounds of purification culture, which places strict requirements on laboratory conditions and operator skills, making it difficult to adapt to the rapid testing needs of primary healthcare institutions.

[0004] Surface-enhanced Raman spectroscopy (SERS) technology amplifies the Raman signal of pathogens by using a nano-enhanced substrate, significantly improving detection sensitivity and giving it the potential to identify pathogens. In recent years, with the development of deep learning technology, spectral data analysis methods based on machine learning models have further improved the accuracy of SERS detection due to their ability to capture the sequence correlation characteristics of characteristic peaks. For example, CN117491339A discloses a rapid SERS detection method for identifying pathogenic bacteria, including: (1) collecting the SERS spectrum of pathogenic bacteria as a calibration set sample; (2) performing spectral preprocessing on the SERS spectrum and classifying the data; (3) establishing different types of machine learning classification models and selecting the best classification algorithm based on the generalization ability of the model; (4) collecting the SERS spectrum of the test sample; (5) inputting the SERS spectrum of the test sample into the established best classification model to achieve rapid identification of the test bacteria. However, the application of this technology faces a key bottleneck: obtaining bone infection lesion specimens requires invasive procedures, clinical sample sizes are already scarce, and contamination during the purification of pathogenic bacteria can further reduce the amount of effective spectral data, making it difficult to meet the needs of deep learning models for large-scale training data. As a result, the integration and application of SERS and deep learning has remained in the laboratory stage and cannot be translated into clinical practice.

[0005] Existing attempts based on SERS and deep learning largely fail to address the core contradiction of the unique characteristics of spectral data and the scarcity of samples. On the one hand, some models directly use general deep learning architectures to process spectral data, failing to effectively screen for specific characteristic peaks among pathogens. This leads to the model being interfered with by irrelevant noise, limiting recognition accuracy. On the other hand, the lack of effective data augmentation methods, relying solely on training with a small number of real samples, results in insufficient model generalization ability, making them prone to misclassification and missed detection in complex clinical samples. Furthermore, traditional data augmentation methods (such as simple interpolation and noise addition) cannot simulate the unique molecular fingerprint feature distribution of SERS spectra, resulting in data lacking biological significance and failing to fundamentally solve the problem of insufficient sample size. Simultaneously, existing fusion models do not fully consider the subtle spectral differences among common bone infection pathogens (such as Staphylococcus aureus and Staphylococcus epidermidis), failing to enhance the capture of key discriminative features through network structure optimization, further restricting the improvement of detection accuracy.

[0006] In conclusion, there is an urgent need to develop a technical solution that can accurately adapt to scenarios with scarce samples, so as to achieve efficient utilization and in-depth mining of scarce spectral data. Summary of the Invention

[0007] To address the shortcomings of existing technologies and practical needs, this invention provides a model for identifying pathogenic bacteria causing bone infections, its construction method, and its application, aiming to achieve rapid and accurate identification of pathogenic bacteria causing bone infections.

[0008] To achieve the above objectives, the present invention adopts the following technical solution: In a first aspect, the present invention provides a method for constructing a model for identifying pathogenic bacteria causing bone infections, the method comprising the following steps: S1. Collect and preprocess bone infection lesion samples; S2. Obtain surface-enhanced Raman spectroscopy data of pathogenic bacteria from bone infection lesion samples; S3. Preprocess the surface-enhanced Raman spectroscopy data obtained in step S2; S4. Divide the surface-enhanced Raman spectroscopy data after preprocessing in step S3 into a training set and a test set; S5. Use a generative adversarial network to enhance and integrate the data from the training set in step S4; S6. Use the enhanced and integrated training set from step S5 to train the SpecGAN-LSTM model to obtain a bone infection pathogen identification model.

[0009] This invention utilizes the distributed learning capabilities of Generative Adversarial Networks (GANs) to generate virtual data consistent with real spectral features, thus compensating for sample size deficiencies at the source. Simultaneously, combining the molecular fingerprint characteristics of spectral data, a deep learning model incorporating feature peak screening and attention mechanisms is designed to focus on specific spectral differences among pathogens, enhancing the targeting and effectiveness of feature extraction. A high-performance model for identifying bone infection pathogens is constructed, combining the rapid detection advantages of SERS technology with the precise identification capabilities of deep learning. This overcomes the time-consuming bottlenecks of traditional detection methods and the sample adaptation limitations of existing fusion technologies, achieving rapid and accurate identification of bone infection pathogens. This promotes the transformation of optical detection technology from "auxiliary localization" to "core diagnosis," providing reliable technical support for timely intervention and precise medication in clinical bone infections.

[0010] Optionally, the preprocessing in step S1 includes: S1-1. Mix the collected bone infection lesion samples with buffer solution, centrifuge, and discard the supernatant; S1-2. Resuspend the precipitate in buffer solution and filter under negative pressure using a microporous membrane to separate and retain pathogenic bacteria; S1-3. Elute the retained bacterial cells to obtain a purified bacterial solution; wherein the pathogenic bacteria include at least Staphylococcus aureus and Staphylococcus epidermidis.

[0011] Optionally, the buffer solution described in step S1-1 or S1-2 includes a phosphate buffer solution.

[0012] Optionally, in step S1-1, the volume ratio of the bone infection lesion sample to the buffer solution is 1:(1~3).

[0013] Optionally, the centrifugation conditions are 3000~8000 rpm for 10~5 min.

[0014] Optionally, the eluent used for elution includes physiological saline.

[0015] Optionally, step S2 specifically includes: S2-1. The purified bacterial solution is mixed with the Raman-enhanced substrate material to obtain a mixture; S2-2. Use a capillary glass tube to draw up the mixture and place it on the optical acquisition platform; S2-3. Randomly select at least 10 different sites on each sample spot, and collect samples from 400 to 1800 cm⁻¹. -1 The spectrum within the range.

[0016] Optionally, the Raman-enhanced substrate material includes any one of silver nanosphere sol-reinforcement (20 nm-50 nm), gold nanosphere sol, or conventionally shaped gold / silver nanospheres such as square or star.

[0017] Optionally, the preprocessing in step S3 includes: S3-1. Perform baseline correction on the surface-enhanced Raman spectroscopy data obtained in step S2; S3-2. Smoothing and denoising the baseline-corrected surface-enhanced Raman spectroscopy data; S3-3. Normalize the surface-enhanced Raman spectral data after smoothing and denoising to map the surface-enhanced Raman spectral intensity values ​​to the standard dimension range.

[0018] Optionally, step S5 specifically includes: S5-1. Construct a generative adversarial network with the same structure for each target pathogen in the training set. The generator of each generative adversarial network is responsible for learning the spectral distribution characteristics of the corresponding bacterial species and generating its virtual spectrum. S5-2. Input the real surface-enhanced Raman spectrum dataset of each target pathogen into its corresponding generative adversarial network for independent training; during training, the Fréchet distance is used to evaluate the generation quality of each generative adversarial network. The Fréchet distance calculation formula is as follows: ; in, and These are the mean and covariance of the true spectral characteristics, respectively. and These represent the mean and covariance of the generated spectral features, respectively. The trace of the matrix is ​​represented; training aims to minimize the Fréchet distance. S5-3. Using the generative adversarial network obtained from training, generate a specified number of virtual surface-enhanced Raman spectra of target pathogens; integrate the virtual surface-enhanced Raman spectra of each target pathogen with their corresponding original training sets to form an amplified training set covering the target pathogens.

[0019] Optionally, the training described in step S6 includes: S6-1. Perform significance analysis on the spectral data of the target pathogens in the training set after the enhancement and integration in step S5. Use statistical methods such as analysis of variance or t-test to identify the wavenumber intervals corresponding to the characteristic peaks that have significant differences among the target pathogens. Construct a key feature wavenumber mask and use the key feature wavenumber mask to screen out the key feature peaks that can best distinguish the bacterial species from the full-band spectrum. S6-2. Construct the SpecGAN-LSTM model, the structure of which includes: Input layer: Receives key feature peaks obtained after being filtered by the key feature wavenumber mask; Long Short-Term Memory (LSTM) network layers: used for deep learning and memorizing the sequence dependencies, intensity ratios, and combination patterns among the key feature peaks; Attention mechanism layer: Automatically evaluates and weights the contribution of different key feature peak time steps to the final classification decision, focusing on the most discriminative features; Fully connected output layer: outputs the probability that each sample belongs to the target pathogen category; S6-3. Input the key feature peaks obtained by the key feature wavenumber mask screening into the SpecGAN-LSTM model for training to obtain a bone infection pathogen identification model.

[0020] Optionally, the training described in step S6-3 uses the Adam optimizer, with a learning rate of 0.001 to 0.0001, a batch size of 16, 32, 64 or 128, and a training cycle of 50 to 100.

[0021] Optionally, in the SpecGAN-LSTM model, the gating units in the long short-term memory network layer use the Sigmoid function, the attention mechanism layer uses the Softmax function, and the fully connected output layer uses the Softmax function.

[0022] Optionally, step S6 may also include a step of testing the recognition performance of the bone infection pathogen identification model on a test set.

[0023] Optionally, the evaluation metrics for the test include accuracy, precision, recall, and F1 score.

[0024] Secondly, the present invention provides a bone infection pathogen identification model, which is obtained by the bone infection pathogen identification model construction method described in the first aspect.

[0025] Thirdly, the present invention provides a method for identifying bone infection pathogens not for disease diagnosis and / or treatment purposes, the method comprising: The surface-enhanced Raman spectroscopy data of the sample to be tested is input into the bone infection pathogen identification model described in the second aspect, and the category probability of the target pathogen is output. The judgment is made based on the category probability.

[0026] The bone infection pathogen identification method of the present invention can also be used for non-disease diagnosis, such as detecting pathogens in non-human samples for pathogen monitoring and control.

[0027] Fourthly, the present invention provides a bone infection pathogen identification system, the system being used to perform the steps in the bone infection pathogen identification method described in the third aspect, including a data acquisition unit and a determination unit; The data acquisition unit is used to perform the following: acquiring surface-enhanced Raman spectral data of the sample to be tested; the determination unit is used to perform the following: inputting the surface-enhanced Raman spectral data of the sample to be tested into the bone infection pathogen identification model described in the second aspect, outputting the category probability of the target pathogen, making a determination based on the category probability, and outputting the result.

[0028] Optionally, the criteria for determining whether a sample belongs to a specific pathogenic bacterium include (taking Staphylococcus aureus and Staphylococcus epidermidis as examples): (1) When the probability of Staphylococcus aureus type output by the model is ≥80%, and the difference between the probability of Staphylococcus aureus type and the probability of Staphylococcus epidermidis type is ≥20%, the sample to be tested is determined to be Staphylococcus aureus; (2) When the probability of Staphylococcus epidermidis type output by the model is ≥80%, and the difference between the probability of Staphylococcus epidermidis type and the probability of Staphylococcus aureus type is ≥20%, the sample to be tested is determined to be Staphylococcus epidermidis. (3) When the probability of Staphylococcus aureus and the probability of Staphylococcus epidermidis are both less than 80%, or the difference between the two probabilities is less than 20%, it is determined that the pathogenic bacteria category of the sample cannot be clearly identified.

[0029] Fifthly, the present invention provides an electronic device comprising one or more processors and a memory for storing executable instructions, wherein the one or more processors are configured to invoke the executable instructions stored in the memory to implement the function of the bone infection pathogen identification system described in the fourth aspect.

[0030] In a sixth aspect, the present invention provides a computer-readable storage medium having stored thereon computer program instructions, which, when executed by a processor, implement the function of the bone infection pathogen identification system described in the fourth aspect.

[0031] Compared with the prior art, the present invention has at least the following beneficial effects: This invention addresses the scarcity of bone infection pathogen samples by designing a sample data enhancement and integration strategy based on generative adversarial networks to expand the data. Further, it integrates surface-enhanced Raman spectroscopy and SpecGAN-LSTM to develop a bone infection pathogen identification model and method. This enables rapid and accurate identification of bone infection pathogens, significantly improving detection efficiency. Furthermore, the accuracy and precision indicators have verified that the identification performance is stable and reliable, with an accuracy rate exceeding 96%. Attached Figure Description

[0032] Figure 1 A schematic diagram of the process for constructing a model to identify pathogens causing bone infections.

[0033] Figure 2 This is a schematic diagram of the structure of an electronic device according to an embodiment of the present invention, wherein 01 is a processor, 02 is a memory, and 03 is a computer program. Detailed Implementation

[0034] The technical solution of the present invention will be further described below with reference to the accompanying drawings and specific embodiments. However, the following examples are merely simplified examples of the present invention and do not represent or limit the scope of protection of the present invention. The scope of protection of the present invention is determined by the claims.

[0035] Where specific techniques or conditions are not specified in the examples, they shall be performed in accordance with the techniques or conditions described in the literature in this field, or in accordance with the product instructions. Reagents or instruments whose manufacturers are not specified are all conventional products that can be purchased from legitimate channels.

[0036] Unless otherwise defined, scientific and technical terms and their abbreviations used in conjunction with this invention shall have the meanings commonly understood by one of ordinary skill in the art to which this invention pertains. Some of the terms and abbreviations used in this invention are listed below.

[0037] SpecGAN-LSTM is a hybrid model that combines SpecGAN (a generative adversarial network for spectral data, often used for audio generation, which directly models and generates audio spectrograms) with LSTM (Long Short-Term Memory network, which is good at processing time-series data).

[0038] Fréchet distance: Fréchet Inception Distance, also known as "Earth travel distance", measures the similarity between two multidimensional probability distributions by calculating the average of the shortest distances between corresponding points in the distributions.

[0039] Generative Adversarial Network (GAN): A type of unsupervised deep learning model consisting of two networks, a generator and a discriminator, playing against each other.

[0040] The Sigmoid function is a classic sigmoid monotonic activation function, and its mathematical expression is: .

[0041] The Softmax function is a normalized activation function used in multi-class classification tasks. It maps any real vector to a probability distribution (where each element takes values ​​in the interval [0, 1] and sums to 1). Its mathematical expression is: .

[0042] Accuracy (ACC): Measures the overall prediction accuracy of a model, i.e., the proportion of all correctly predicted samples out of the total sample. It is effective when the data is class-balanced, but can be misleading on imbalanced data. Calculation formula: Accuracy = (TP + TN) / (TP + TN + FP + FN).

[0043] Precision (Pre): Measures the reliability of a model's predictions of positive examples. It represents the proportion of samples where the model correctly predicts a positive example. High precision means fewer false positives. The formula is: Precision = TP / (TP + FP).

[0044] Specificity: Measures a model's ability to identify negative examples. It's the proportion of samples that the model correctly excludes from all true negative samples. High specificity means a low false positive rate. Calculation formula: Specificity = TN / (TN + FP).

[0045] Sensitivity / Recall: Measures a model's ability to detect positive examples. It's the proportion of true positive examples that the model correctly predicts. High sensitivity means a low false negative rate. Calculation formula: Recall = TP / (TP + FN).

[0046] F1 score: The harmonic mean of precision and recall, used to comprehensively evaluate these two metrics. It is particularly effective in reflecting the true performance of a model than accuracy, especially when there is an imbalance of positive and negative examples. Calculation formula: F1 = 2 × (Precision × Recall) / (Precision + Recall).

[0047] TP: True positive, meaning that the actual case is positive and the model also predicts it to be positive.

[0048] FP: False positive, meaning that the actual case is negative, but the model incorrectly predicts it as a positive case.

[0049] TN: True negative, meaning that the actual case is negative and the model also predicts it to be negative.

[0050] FN: False negative, meaning the actual case is positive, but the model incorrectly predicts it as negative.

[0051] This invention develops a rapid identification method for bone infection pathogens in scenarios where samples are scarce. A bone infection pathogen identification model is constructed based on surface-enhanced Raman spectroscopy and SpecGAN-LSTM. The flowchart is shown in the figure, and includes the following steps: S1. Collect and preprocess bone infection lesion samples; S2. Collect surface-enhanced Raman spectroscopy (SERS) data of pathogenic bacteria from bone infection lesion samples; S3. Preprocess the spectral data of pathogens causing bone infections; S4. Divide the spectral dataset of bone infection pathogens after preprocessing in step S3 into a training set and a test set; S5. Use Generative Adversarial Networks (GANs) to enhance and integrate the spectral data of the training set of bone infection pathogens; S6. Use the enhanced and integrated training set from step S5 to train the SpecGAN-LSTM model to obtain the final identification model for bone infection pathogens.

[0052] Specifically, the preprocessing in step S1 includes: S1-1. Mix the collected bone infection lesion samples with buffer (such as sterile phosphate buffer, the volume ratio can be 1:(1~1.5)), disperse them thoroughly by vortexing, and then centrifuge (such as centrifuging at 3000 rpm for 10 min), discard the supernatant to remove blood and soft tissue impurities. S1-2. Resuspend the precipitate in fresh sterile phosphate buffer and filter under negative pressure using a microporous membrane (e.g., a 0.22 μm membrane) to separate and retain pathogenic bacteria; S1-3. Rinse the filter membrane with an appropriate amount of physiological saline and elute the retained bacteria to obtain a purified bacterial solution; wherein the pathogenic bacteria include at least Staphylococcus aureus and Staphylococcus epidermidis.

[0053] Specifically, the SERS data acquisition method in step S2 includes: S2-1. Mix the purified bacterial solution with the Raman-enhanced substrate material (such as silver nanoparticle sol enhancer); use a micropipette to mix thoroughly and incubate at room temperature in the dark to ensure that the pathogenic bacteria and nanoparticles interact fully; S2-2. Using the same transfer volume, draw up the mixture with a capillary glass tube and place it on the optical acquisition platform; S2-3. Using the automated Raman spectroscopy mapping (imaging) function, randomly select no fewer than 10 different sites on each sample spot and collect data from 400 to 1800 cm⁻¹. -1 The spectrum is within the range (standardized parameters can be obtained by using a 785 nm laser at a power of 8-10 mW and an integration time of 3-20 seconds, with 20-3 accumulations); wherein the pathogenic bacteria are Staphylococcus aureus and Staphylococcus epidermidis.

[0054] Specifically, step S3 involves preprocessing the spectral data of bone infection pathogens to eliminate instrument errors and sample concentration differences, including the following steps: S3-1. Baseline correction is performed on the acquired raw surface-enhanced Raman spectroscopy data to eliminate non-target signals and baseline drift introduced by factors such as laser fluorescence, sample background and detector dark current; S3-2. Perform smoothing and denoising processing on the baseline-corrected spectral data to filter out high-frequency random noise while retaining the characteristic peak shapes that represent molecular fingerprints in the spectrum; S3-3. Normalize the spectral data after smoothing and denoising to map the spectral intensity values ​​to a standard dimension range, so as to eliminate the difference in absolute spectral intensity caused by fluctuations in bacterial concentration or laser power.

[0055] Specifically, in step S4, the preprocessed spectral dataset is randomly divided into a training set and a test set according to a ratio of (5~7):(3~5) (e.g., 7:3).

[0056] Specifically, in step S5, for each pathogenic bacterium (such as Staphylococcus aureus and Staphylococcus epidermidis), a generative adversarial network is independently trained for data augmentation, which includes the following steps: S5-1. Network Construction: For each target pathogen (such as Staphylococcus aureus and Staphylococcus epidermidis), construct a corresponding generative adversarial network (GAN) with the same structure; the generator of each GAN is responsible for learning the spectral distribution characteristics unique to the corresponding bacterial species and generating its virtual spectrum; S5-2. Independent Training and Evaluation: The real spectral dataset of each target pathogen is input into its corresponding GAN for independent training. During training, the Fréchet distance is used to evaluate the generation quality of each GAN, which measures the difference in distribution between the generated spectrum and the real spectrum in the feature space. FD The calculation formula is as follows: ; in, and These are the mean and covariance of the true spectral characteristics, respectively. and These represent the mean and covariance of the generated spectral features, respectively. The trace of the matrix is ​​represented; training aims to minimize the Fréchet distance. S5-3. Data Integration: Using the trained generative adversarial network, generate a specified number of virtual spectra of the target pathogens; finally, integrate the virtual spectral data of each target pathogen with their corresponding original training sets to form an amplified training set that covers the target pathogens.

[0057] Specifically, step S6 uses the enhanced and integrated training set from step S5 to train the SpecGAN-LSTM model, which includes the following steps: S6-1. Key Feature Peak Interval Screening: Significance analysis is performed on the spectral data of target pathogens in the enhanced and integrated training set. Statistical methods such as analysis of variance or t-test are used to identify the wavenumber intervals corresponding to the feature peaks that show significant differences among the target pathogens. Based on this, a key feature wavenumber mask is constructed to screen out the key feature peaks that best distinguish bacterial species from the full-band spectrum, thereby significantly reducing the input dimensionality of subsequent models. S6-2. Constructing the SpecGAN-LSTM classification model: This model is a classifier specifically designed for spectral sequences, and its network structure includes the following components: Input layer: Receives key feature peaks obtained after being filtered by a key feature wavenumber mask; Long Short-Term Memory (LSTM) layer: used for deep learning and memorizing the sequence dependencies, intensity ratios, and combination patterns among the key feature peaks; Attention mechanism layer: Automatically evaluates and weights the contribution of different feature peak time steps to the final classification decision, focusing on the most discriminative features; Fully connected output layer: The final output layer contains the probability that each sample belongs to the target pathogen category. S6-3. Model Training and Validation: After applying key feature wavenumber masks to the expanded training set data, the data is input into the SpecGAN-LSTM model for training (the Adam optimizer can be used during training, with a learning rate of 0.001, a batch size of 64, and a training cycle of 100; the cross-entropy loss function can be optimized to enable the model to accurately learn and distinguish the unique spectral fingerprint of the target pathogen), ultimately obtaining an efficient bone infection pathogen identification model.

[0058] Specifically, the activation functions used in each layer of the SpecGAN-LSTM model in training step S6 are as follows: The gating units in the Long Short-Term Memory (LSTM) network layer use the Sigmoid function to calculate the on / off state of the input, forget, and output gates; the candidate cell states use the hyperbolic tangent function; the attention mechanism layer uses the Softmax function to assign weights to the hidden states of all time steps output by the LTM network layer; and the fully connected output layer uses the Softmax function to convert the final output value into the probability that the sample belongs to the target pathogen category.

[0059] Specifically, step S6 also includes testing the model's recognition performance on the test set. The evaluation metrics used include accuracy, precision, recall, and F1 score. For multi-class tasks, precision, recall, and F1 score are calculated using macro-average or micro-average to comprehensively evaluate the model's recognition performance for different types of bone infection pathogens.

[0060] The specific embodiments of the present invention use samples containing Staphylococcus aureus and Staphylococcus epidermidis as examples to verify the technical solution of the present invention.

[0061] Example 1 This embodiment constructs a model for identifying pathogenic bacteria causing bone infections based on surface-enhanced Raman spectroscopy and SpecGAN-LSTM.

[0062] The enhancing substrate materials for surface-enhanced Raman spectroscopy include silver nanosol and capillary glass tubes. In use, the sample to be tested is mixed with silver nanosol and incubated. The mixture is then drawn up with a capillary glass tube and placed on an optical acquisition platform. The surface-enhanced Raman spectrum of the sample to be tested on the Raman-enhanced substrate is then acquired using a data acquisition module.

[0063] Silver nanosols were prepared by the following method: 1) Dissolve 34 mg of silver nitrate in 200 mL of ultrapure water. Then add trisodium citrate (1%, 4 mL) to the boiled silver nitrate solution and stir continuously for 30 min until the mixture turns grayish-green. Cool to room temperature, make up to 200 mL, mix well, and store at 4°C protected from light. 2) Centrifuge 500 mL of Ag colloidal solution (9000 rpm, 10 min) and immediately remove the supernatant to obtain silver nanosol for later use.

[0064] The data acquisition module is a DXR3 laser confocal Raman (Thermo Scientific™, USA).

[0065] The methods for constructing recognition models include: S1. Collect and preprocess bone infection lesion samples. S1-1. The collected bone infection lesion samples were mixed with sterile phosphate buffer at a volume ratio of 1:1, and thoroughly dispersed by vortexing. Then, the mixture was centrifuged at 3000 rpm for 5 minutes, and the supernatant was discarded to remove blood and soft tissue impurities. S1-2. Resuspend the precipitate in fresh sterile phosphate buffer and filter under negative pressure using a 0.22 μm microporous membrane to separate and retain pathogenic bacteria; S1-3. Rinse the filter membrane with an appropriate amount of physiological saline and elute the retained bacteria to obtain a purified bacterial solution; wherein the pathogenic bacteria include Staphylococcus aureus and Staphylococcus epidermidis; S2. Collect surface-enhanced Raman spectroscopy (SERS) data of pathogenic bacteria from bone infection lesion samples: S2-1. Accurately transfer 5 μL of purified bacterial solution and 5 μL of silver nano-sol enhancer, and mix them by pipetting using a micropipette; S2-2. Using the same transfer volume, quantitatively add the mixture to a specific area of ​​the marked silicon wafer substrate, and place it in a controlled desiccator to complete the drying process under constant temperature of 25°C to ensure the consistency and reproducibility of sample preparation; S2-3. Using the automated Raman spectroscopy mapping function, at least 10 different sites were randomly selected on each sample, and each site was tested 3 times. Data were collected from 400 to 1800 cm⁻¹ using a 785 nm laser at standardized parameters of 10 mW power and 3-second integration time. -1 The spectrum within the range; wherein the pathogenic bacteria are Staphylococcus aureus and Staphylococcus epidermidis.

[0066] S3. Preprocess the spectral data of pathogens causing bone infections: The Raman spectra of each bacterial species were plotted and processed. S3-1. Baseline correction is performed on the acquired raw surface-enhanced Raman spectra to eliminate non-target signals and baseline drift introduced by factors such as laser fluorescence, sample background and detector dark current; S3-2. The spectrum after baseline correction is smoothed and denoised to filter out high-frequency random noise while retaining the characteristic peaks that represent the molecular fingerprint in the spectrum. S3-3. Normalize the smoothed and denoised spectral data to map the spectral intensity values ​​to a standard dimension range, thereby eliminating differences in absolute spectral intensity caused by fluctuations in bacterial concentration or laser power. S4: Divide the original spectral dataset of bone infection pathogens into training and testing sets in a 7:3 ratio.

[0067] S5: Augmentation of the training set spectral data of bone infection pathogens using SpecGAN-LSTM: S5-1. Network Construction: Two identical generative adversarial networks (GANs) are constructed for the target pathogens, Staphylococcus aureus and Staphylococcus epidermidis. The generator of each GAN is responsible for learning the spectral distribution characteristics specific to the corresponding bacterial species and generating its virtual spectrum. S5-2. Independent Training and Evaluation: The real spectral datasets of Staphylococcus aureus and Staphylococcus epidermidis are input into their respective generative adversarial networks for independent training; during the training process, the Fréchet distance is used to evaluate the generation quality of each GAN.

[0068] S5-3. Data Integration: Using two trained generators, generate a specified number (1024) virtual spectra of Staphylococcus aureus and Staphylococcus epidermidis. Finally, integrate the virtual spectral data of the two types of bacteria with their corresponding original training sets to form an amplified training set covering the two target pathogens.

[0069] S6. Train the LSTM model using the enhanced and integrated training set to identify pathogens causing bone infections: S6-1. Key Feature Peak Range Screening: Significance analysis was performed on the spectral data of Staphylococcus aureus and Staphylococcus epidermidis in the enhanced integrated training set. Using analysis of variance (ANOVA), wavenumber ranges corresponding to feature peaks showing significant differences between the two pathogens were identified. The screening dimensions included not only the intensity differences of individual peaks but also comprehensive information such as the position, shape, and full width at half maximum (FWHM) of all peaks. This information collectively constitutes the "fingerprint" distinguishing the two bacteria. Based on this, a key feature wavenumber mask was constructed to screen the subset of spectral features that best distinguishes the bacterial species from the full-band spectrum, thereby significantly reducing the input dimensionality of subsequent models. S6-2. Constructing the SpecGAN-LSTM classification model: This model is a classifier specifically designed for spectral sequences, and its network structure includes the following components: Input layer: Receives spectral data filtered through a key feature wavenumber mask; Two stacked long short-term memory network layers: used for deep learning and memorizing the sequence dependencies, intensity ratios, and combination patterns among the key feature peaks; Attention mechanism layer: Automatically evaluates and weights the contribution of different feature peak time steps to the final classification decision, focusing on the most discriminative features; Fully connected output layer: The final output sample is the probability of belonging to the Staphylococcus aureus and Staphylococcus epidermidis categories, respectively; S6-3. Model Training and Validation: The expanded training set data was applied with a key feature wavenumber mask and then input into the SpecGAN-LSTM model for training. The Adam optimizer was used during training, with a learning rate of 0.001, a batch size of 64, and a training period of 100. By optimizing the cross-entropy loss function, the model was able to accurately learn and distinguish the unique spectral fingerprints of the two target pathogens, ultimately obtaining an efficient bone infection pathogen identification model.

[0070] The activation functions used in each layer of the SpecGAN-LSTM model are as follows: 1) The gating units in the Long Short-Term Memory network layer use the Sigmoid function to calculate the degree of opening and closing of the input gate, forget gate and output gate; the candidate cell states use the hyperbolic tangent function.

[0071] 2) The attention mechanism layer uses the Softmax function to assign weights to the hidden states of all time steps output by the LSTM layer.

[0072] 3) The fully connected output layer uses the Softmax function to convert the final output value into the probability that the sample belongs to the category of Staphylococcus aureus or Staphylococcus epidermidis.

[0073] Example 2 This embodiment tests the recognition model constructed in the previous embodiment.

[0074] The model's recognition performance was tested on a test set. Data from the test set was input into the recognition model, and the output was the probability that a sample belonged to either Staphylococcus aureus or Staphylococcus epidermidis. For the binary classification recognition scenario of Staphylococcus aureus and Staphylococcus epidermidis, and considering the requirements of clinical testing for recognition accuracy and reliability, a dual judgment standard of a single-class probability threshold and the difference between the two class probabilities was adopted. The specific judgment rules are as follows: (1) When the probability of Staphylococcus aureus type output by the model is ≥80%, and the difference between the probability of Staphylococcus aureus type and the probability of Staphylococcus epidermidis type is ≥20%, the sample to be tested is determined to be Staphylococcus aureus; (2) When the probability of Staphylococcus epidermidis type output by the model is ≥80%, and the difference between the probability of Staphylococcus epidermidis type and the probability of Staphylococcus aureus type is ≥20%, the sample to be tested is determined to be Staphylococcus epidermidis. (3) When the probability of Staphylococcus aureus and the probability of Staphylococcus epidermidis are both less than 80%, or the difference between the two probabilities is less than 20%, it is determined that the pathogenic bacteria type of the sample cannot be clearly identified. It is recommended to collect the sample again and obtain its surface-enhanced Raman spectroscopy data before retesting.

[0075] Based on the above judgment rules, the class probabilities of the test samples and the judgment results are shown in Table 1; based on the judgment results, the accuracy, precision, specificity, recall and F1 score of the recognition model are further calculated, and the results are shown in Table 2.

[0076] Table 1 Table 2 As shown in Table 2, the recognition model constructed in this invention has good recognition performance.

[0077] Example 3 This embodiment provides a system for identifying pathogenic bacteria causing bone infections. The system includes a data acquisition unit and a determination unit. The data acquisition unit is used to perform the following: acquiring the SERS spectrum of the sample to be tested. The determination unit is used to perform the following: inputting the SERS spectrum of the sample to be tested into an established identification model, outputting the category probability of the target pathogen, making a determination based on the category probability, and outputting the result.

[0078] Example 4 This embodiment provides an electronic device, such as... Figure 2 As shown, the electronic device includes a memory 02 and a processor 01. The memory 02 stores a computer program 03 that can run on the processor 01. When the processor 01 executes the computer program 03, it implements the function of the bone infection pathogen identification system as described in Example 3.

[0079] Those skilled in the art will understand that the device of the present invention can be obtained using various forms of hardware, software, firmware, dedicated processors, or combinations thereof.

[0080] Example 5 This embodiment provides a computer-readable storage medium storing a computer program. The computer program has program code, which, when run in a corresponding processor, controller, computing device, or terminal, implements the function of the bone infection pathogen identification system as described in Embodiment 3.

[0081] In summary, this invention addresses the scarcity of bone infection pathogen samples by designing a sample data enhancement and integration strategy based on generative adversarial networks to expand the data. Furthermore, it integrates surface-enhanced Raman spectroscopy and SpecGAN-LSTM to develop a bone infection pathogen identification model and method. This approach enables rapid and accurate identification of bone infection pathogens, significantly improving detection efficiency. Furthermore, the identification performance has been verified by indicators such as accuracy and precision, demonstrating stable and reliable performance.

[0082] The applicant declares that the above description is only a specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Those skilled in the art should understand that any changes or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention fall within the protection and disclosure scope of the present invention.

Claims

1. A method for constructing a model for identifying pathogenic bacteria causing bone infections, characterized in that, The construction method includes the following steps: S1. Collect and preprocess bone infection lesion samples; S2. Obtain surface-enhanced Raman spectroscopy data of pathogenic bacteria from bone infection lesion samples; S3. Preprocess the surface-enhanced Raman spectroscopy data obtained in step S2; S4. Divide the surface-enhanced Raman spectroscopy data after preprocessing in step S3 into a training set and a test set; S5. Use a generative adversarial network to enhance and integrate the data from the training set in step S4; S6. Use the enhanced and integrated training set from step S5 to train the SpecGAN-LSTM model to obtain a bone infection pathogen identification model.

2. The method for constructing a bone infection pathogen identification model according to claim 1, characterized in that, The preprocessing described in step S1 includes: S1-1. Mix the collected bone infection lesion samples with buffer solution, centrifuge, and discard the supernatant; S1-2. Resuspend the precipitate in buffer solution and filter under negative pressure using a microporous membrane to separate and retain pathogenic bacteria; S1-3. Elute the retained bacterial cells to obtain a purified bacterial solution; wherein the pathogenic bacteria include at least Staphylococcus aureus and Staphylococcus epidermidis; The buffer solution mentioned in step S1-1 or S1-2 includes a phosphate buffer solution; In step S1-1, the volume ratio of the bone infection lesion sample to the buffer solution is 1:(1~3); The centrifugation conditions are 3000-8000 rpm for 10-5 min; The elution solution includes physiological saline.

3. The method for constructing a bone infection pathogen identification model according to claim 2, characterized in that, Step S2 specifically includes: S2-1. The purified bacterial solution is mixed with the Raman-enhanced substrate material to obtain a mixture; S2-2. Use a capillary glass tube to draw up the mixture and place it on the optical acquisition platform; S2-3. Randomly select at least 10 different sites on each sample spot, and collect samples from 400 to 1800 cm⁻¹. -1 Spectrum within the range; The Raman-enhanced substrate material includes any one of silver nanosphere sol reinforcing agent, gold nanosphere sol, or square or star-shaped gold / silver nanomaterials.

4. The method for constructing a bone infection pathogen identification model according to claim 1, characterized in that, The preprocessing described in step S3 includes: S3-1. Perform baseline correction on the surface-enhanced Raman spectroscopy data obtained in step S2; S3-2. Smoothing and denoising the baseline-corrected surface-enhanced Raman spectroscopy data; S3-3. Normalize the surface-enhanced Raman spectral data after smoothing and denoising to map the surface-enhanced Raman spectral intensity values ​​to the standard dimension range. Step S5 specifically includes: S5-1. Construct a generative adversarial network with the same structure for each target pathogen in the training set. The generator of each generative adversarial network is responsible for learning the spectral distribution characteristics of the corresponding bacterial species and generating its virtual spectrum. S5-2. Input the real surface-enhanced Raman spectrum dataset of each target pathogen into its corresponding generative adversarial network for independent training; during training, the Fréchet distance is used to evaluate the generation quality of each generative adversarial network. The Fréchet distance calculation formula is as follows: ; in, and These are the mean and covariance of the true spectral characteristics, respectively. and These represent the mean and covariance of the generated spectral features, respectively. The trace of the matrix is ​​represented; training aims to minimize the Fréchet distance. S5-3. Using the generative adversarial network obtained from training, generate a specified number of virtual surface-enhanced Raman spectra of target pathogens; integrate the virtual surface-enhanced Raman spectra of each target pathogen with their corresponding original training sets to form an amplified training set covering the target pathogens.

5. The method for constructing a bone infection pathogen identification model according to claim 1, characterized in that, The training described in step S6 includes: S6-1. Perform significance analysis on the spectral data of the target pathogens in the training set after the enhancement and integration in step S5. Use statistical methods such as analysis of variance or t-test to identify the wavenumber intervals corresponding to the characteristic peaks that have significant differences among the target pathogens. Construct a key feature wavenumber mask and use the key feature wavenumber mask to screen out the key feature peaks that can best distinguish the bacterial species from the full-band spectrum. S6-2. Construct the SpecGAN-LSTM model, the structure of which includes: Input layer: Receives key feature peaks obtained after being filtered by the key feature wavenumber mask; Long Short-Term Memory (LSTM) network layers: used for deep learning and memorizing the sequence dependencies, intensity ratios, and combination patterns among the key feature peaks; Attention mechanism layer: Automatically evaluates and weights the contribution of different key feature peak time steps to the final classification decision, focusing on the most discriminative features; Fully connected output layer: outputs the probability that each sample belongs to the target pathogen category; S6-3. Input the key feature peaks obtained by the key feature wavenumber mask screening into the SpecGAN-LSTM model for training to obtain a bone infection pathogen identification model; The training described in step S6-3 uses the Adam optimizer, with a learning rate of 0.001 to 0.0001, a batch size of 16, 32, 64 or 128, and a training cycle of 50 to 100. In the SpecGAN-LSTM model, the gating units in the long short-term memory network layer use the Sigmoid function, the attention mechanism layer uses the Softmax function, and the fully connected output layer uses the Softmax function. Step S6 also includes testing the recognition performance of the bone infection pathogen identification model on the test set.

6. A model for identifying pathogenic bacteria causing bone infections, characterized in that, The bone infection pathogen identification model is obtained by the method for constructing the bone infection pathogen identification model according to any one of claims 1-5.

7. A method for identifying pathogenic bacteria causing bone infections for purposes other than disease diagnosis and / or treatment, characterized in that, The method for identifying pathogens causing bone infections includes: The surface-enhanced Raman spectroscopy data of the sample to be tested is input into the bone infection pathogen identification model described in claim 6, and the category probability of the target pathogen is output. The judgment is made based on the category probability.

8. A system for identifying pathogenic bacteria causing bone infections, characterized in that, The system is used to perform the steps in the bone infection pathogen identification method of claim 7, including a data acquisition unit and a determination unit; The data acquisition unit is used to perform the following: acquiring surface-enhanced Raman spectral data of the sample to be tested; the determination unit is used to perform the following: inputting the surface-enhanced Raman spectral data of the sample to be tested into the bone infection pathogen identification model of claim 6, outputting the category probability of the target pathogen, making a determination based on the category probability and outputting the result.

9. An electronic device comprising one or more processors and a memory for storing executable instructions, characterized in that, The one or more processors are configured to invoke executable instructions stored in the memory to implement the function of the bone infection pathogen identification system of claim 8.

10. A computer-readable storage medium having computer program instructions stored thereon, characterized in that, When the computer program instructions are executed by the processor, they implement the function of the bone infection pathogen identification system of claim 8.

Citation Information

Patent Citations

  • SERS (Surface Enhanced Raman Scattering) detection method and system for rapidly identifying types of pathogenic bacteria

    CN117491339A