Apparatus and method for improving accuracy of microorganism identification

CN122531485APending Publication Date: 2026-08-07ST TERESA (CHONGQING) BIOMEDICAL TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
ST TERESA (CHONGQING) BIOMEDICAL TECHNOLOGY CO LTD
Filing Date
2026-05-18
Publication Date
2026-08-07

AI Technical Summary

Technical Problem

[0003]为了解决现有方法在对微生物进行鉴别时存在的鉴别结果精确性较低的问题,本发明的目的在于提供一种提升微生物鉴别精确性的装置和方法,所采用的技术方案具体如下:

Benefits of technology

本发明获取了待测微生物样本的离子强度及对应的时空信息,有效捕捉了待测微生物样本在制备过程中由内源性生化差异驱动的动态演化信息;基于离子强度和采样时刻筛选起始态组分与终态组分,并根据各组分的离子强度及采样位置计算差异因子,构建位置校正成本矩阵,通过对该矩阵求解全局最优匹配,并基于最优匹配路径上的成本数值的统计特征构建动态演化特征向量,能够将传统方法中丢失的时空耦合演化信息转化为具有物种特异性的可量化特征;最终依据该动态演化特征向量确定待测微生物物种类别,提升了在遗传背景高度相似的微生物之间的鉴别精确性,克服了现有静态终态产物分析方法中特征谱图高度重叠、难以有效区分的技术难题。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122531485A_ABST
    Figure CN122531485A_ABST
Patent Text Reader

Abstract

The application relates to the technical field of microorganism identification, and particularly relates to a device and method for improving the accuracy of microorganism identification. The method comprises the following steps: acquiring ion intensity and space-time information of a microorganism sample to be detected; screening initial state components and final state components based on the ion intensity and sampling time; obtaining difference factors of each component pair according to the ion intensity and sampling position corresponding to each component pair, wherein one component pair is composed of one initial state component and one final state component; constructing a position correction cost matrix according to the space-time information and difference factors corresponding to each component pair; solving the global optimal matching of the position correction cost matrix; constructing a dynamic evolution feature vector based on the statistical characteristics of the cost values on the optimal matching path, and further determining the species category of the microorganism sample to be detected. The application improves the accuracy of the microorganism identification result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of microbial identification technology, and more specifically to a device and method for improving the accuracy of microbial identification. Background Technology

[0002] Currently, matrix-assisted laser desorption / ionization time-of-flight mass spectrometry (MALDI-TOF MS) has become the mainstream method in clinical microbiological testing. Its basic principle is to compare the characteristic fingerprint spectra of microbial proteins with a database. However, this technology has significant limitations in distinguishing microorganisms with highly similar genetic backgrounds (e.g., Escherichia coli and Shigella). This limitation stems from the fact that conventional microbial mass spectrometry methods only focus on the static final products formed after sample drying and crystallization, ignoring the crucial dynamic information contained during sample preparation. In fact, after microbial cells are mixed with a chemical matrix, during solvent evaporation and co-crystallization, protein components with different physicochemical properties exhibit varying dissolution rates, precipitation sequences, and spatial non-uniformity at the microscopic scale of the target plate. Conventional methods, through full-spot scanning averaging, discard this spatiotemporal coupling evolutionary information driven by endogenous biochemical differences, resulting in highly overlapping characteristic spectra of closely related species. This makes effective separation within the existing characteristic space impossible, leading to low accuracy in microbial identification. Summary of the Invention

[0003] To address the problem of low accuracy in the identification of microorganisms using existing methods, the present invention aims to provide an apparatus and method for improving the accuracy of microbial identification. The specific technical solution adopted is as follows: In a first aspect, the present invention provides a method for improving the accuracy of microbial identification, the method comprising the following steps: The ionic strength and corresponding spatiotemporal information of the microbial sample to be tested are obtained, wherein the spatiotemporal information includes the sampling time and sampling location; Based on all the stated ion intensities and sampling times, the initial and final state components are screened; the difference factor of each component pair is obtained according to the ion intensities and sampling locations corresponding to each component pair; wherein, a component pair consists of an initial state component and a final state component; Based on the spatiotemporal information corresponding to each component pair and the difference factor, a position correction cost matrix is ​​constructed; the global optimal matching is solved for the position correction cost matrix, and a dynamic evolution feature vector is constructed based on the statistical characteristics of the cost values ​​on the optimal matching path; The species category of the microbial sample to be tested is determined based on the dynamic evolution feature vector.

[0004] Preferably, the acquisition of each component includes: Obtain the mass-to-charge ratio of mass spectral peaks with a signal-to-noise ratio higher than a preset signal-to-noise ratio threshold at different acquisition times; All mass spectrometry peaks with a signal-to-noise ratio higher than a preset signal-to-noise ratio threshold are arranged in descending order of mass-to-charge ratio to obtain a mass spectrometry peak sequence. In the mass spectrometry peak sequence, the absolute value of the difference between the mass-to-charge ratios of each pair of mass spectrometry peaks is calculated, and the absolute value is taken as the difference between the mass-to-charge ratios of the corresponding two mass spectrometry peaks. Mass spectrometry peaks in the mass spectrometry peak sequence whose difference is less than or equal to the preset difference threshold and which are adjacent are grouped into the same category, and each category is taken as a component.

[0005] Preferably, the screening of the initial state component and the final state component based on all the said ion intensities and sampling times includes: Define the start and end time windows based on the sampling time; Calculate the first average signal strength of each component within the initial time window and the second average signal strength within the final time window; Arrange all the first average signal indices in descending order to obtain a first average signal intensity sequence; determine the components corresponding to the first preset number of first average signal indices in the first average signal intensity sequence as the initial state components; Arrange all the second average signal indices in descending order to obtain a second average signal intensity sequence; determine the components corresponding to the first preset number of second average signal indices in the second average signal intensity sequence as the final state components.

[0006] Preferably, obtaining the difference factor for each component pair based on the ionic intensity and sampling location of each component pair includes: The DTW distance between the ion intensity sequences of the two components in the candidate component pair is calculated, and the DTW distance is used as the difference factor of the candidate component pair; wherein, the ion intensity sequence of the initial state component in the candidate component pair is obtained by arranging all the ion intensities of the corresponding component in chronological order from front to back, and the ion intensity sequence of the final state component in the candidate component pair is obtained by arranging all the ion intensities of the final state component in chronological order from back to front. The candidate component pair can be any component pair.

[0007] Preferably, the step of constructing the location correction cost matrix based on the spatiotemporal information corresponding to each component pair and the difference factor includes: For any component in any component pair: the spatial centroid coordinates of any component are determined by weighting the position coordinates of the corresponding component using the ionic strength of the corresponding component. Based on the Euclidean distance between the spatial centroid coordinates of each component pair, the Gaussian kernel spatial weight of each component pair is obtained; combined with the Gaussian kernel spatial weight and the difference factor of each component pair, the position correction association cost of each component pair is obtained. Based on the location correction associated costs of all component pairs, a location correction cost matrix is ​​constructed.

[0008] Preferably, the step of solving the global optimal matching of the position correction cost matrix includes: using a bipartite graph matching algorithm to solve the position correction cost matrix to obtain an allocation scheme that minimizes the total association cost.

[0009] Preferably, the construction of the dynamic evolution feature vector based on the statistical characteristics of cost values ​​on the optimal matching path includes: Obtain the total associated cost corresponding to the allocation scheme and the cost values ​​of each item on the optimal matching path; Calculate the mean, variance, skewness, and kurtosis of each cost value on the optimal matching path; The total associated cost, mean, variance, skewness, and kurtosis constitute a dynamic evolution feature vector.

[0010] Preferably, determining the species category of the microbial sample to be tested based on the dynamic evolution feature vector includes: The dynamic evolution feature vector of the sample to be tested is input into the trained reference microbial classification model to obtain the species category of the microbial sample to be tested.

[0011] Preferably, the step of acquiring the ionic strength and corresponding spatiotemporal information of the microbial sample to be tested includes: The laser beam of the mass spectrometer is controlled to continuously scan and acquire the microbial sample to be tested along a preset path to obtain spatiotemporal information and raw mass spectrometry signal; the ion intensity is obtained based on the raw mass spectrometry signal.

[0012] In a second aspect, the present invention provides an apparatus for improving the accuracy of microbial identification, the apparatus being used to implement the method of the first aspect, the apparatus comprising: The acquisition module is used to acquire the ionic strength and corresponding spatiotemporal information of the microbial sample to be tested, wherein the spatiotemporal information includes the sampling time and sampling location; The analysis module is used to screen the initial state components and final state components based on all the ion intensities and sampling times; and to obtain the difference factor of each component pair according to the ion intensities corresponding to each component pair and the sampling location; wherein, a component pair consists of an initial state component and a final state component; The vector construction module is used to construct a position correction cost matrix based on the spatiotemporal information corresponding to each component pair and the difference factor; solve the global optimal matching for the position correction cost matrix; and construct a dynamic evolution feature vector based on the statistical characteristics of the cost values ​​on the optimal matching path. The identification module is used to determine the species category of the microbial sample to be tested based on the dynamic evolution feature vector.

[0013] The present invention has at least the following beneficial effects: This invention acquires the ionic intensity and corresponding spatiotemporal information of the microbial sample to be tested, effectively capturing the dynamic evolution information driven by endogenous biochemical differences during the preparation process. Based on ionic intensity and sampling time, it screens initial and final state components, calculates difference factors based on the ionic intensity and sampling location of each component, constructs a location correction cost matrix, solves for the global optimal match of this matrix, and constructs a dynamic evolution feature vector based on the statistical characteristics of the cost values ​​on the optimal matching path. This transforms the spatiotemporal coupling evolution information lost in traditional methods into quantifiable features specific to the species. Finally, based on this dynamic evolution feature vector, it determines the species category of the microorganism to be tested, improving the accuracy of identification between microorganisms with highly similar genetic backgrounds and overcoming the technical challenge of highly overlapping feature spectra and difficulty in effective differentiation in existing static final-state product analysis methods. Attached Figure Description

[0014] To more clearly illustrate the technical solutions and advantages in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0015] Figure 1 A flowchart illustrating a method for improving the accuracy of microbial identification provided in an embodiment of the present invention; Figure 2 This is a structural block diagram of a device for improving the accuracy of microbial identification provided in an embodiment of the present invention. Detailed Implementation

[0016] To further illustrate the technical means and effects adopted by the present invention to achieve the intended purpose, the following detailed description of an apparatus and method for improving the accuracy of microbial identification, in conjunction with the accompanying drawings and preferred embodiments, is provided below.

[0017] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.

[0018] The following description, in conjunction with the accompanying drawings, details a specific embodiment of the apparatus and method for improving the accuracy of microbial identification provided by the present invention.

[0019] An example of a method to improve the accuracy of microbial identification: This embodiment proposes a method to improve the accuracy of microbial identification, such as... Figure 1 As shown, a method for improving the accuracy of microbial identification in this embodiment includes the following steps: Step S1: Obtain the ionic strength and corresponding spatiotemporal information of the microbial sample to be tested, wherein the spatiotemporal information includes the sampling time and sampling location.

[0020] Considering the dynamic evolution of microbial samples under specific chemical microenvironmental disturbances, including solvent evaporation and crystallization processes, and recognizing the differentiated cell wall structures and hydrophobic protein properties of different microbial species, when mixed with specific chemical probes and matrix solutions, various protein components exhibit specific differences in release rates and spatial non-uniformity during the drying and crystallization process. To transform this species-specific dynamic process into computable data entities, this embodiment performs controlled temporal scanning acquisition and structured processing on the microbial samples to be tested.

[0021] Specifically, a clean stainless steel target plate for a MALDI-TOF mass spectrometer was selected as the reaction carrier; at the sample spotting locations on the target plate, [the following substances were added sequentially]. The microbial sample to be tested (bacterial suspension, such as cultured Escherichia coli or Shigella). concentration lactose solution as a chemical probe, and of A cyano-4-hydroxycinnamic acid (CHCA) matrix solution was used. The mixture was initially mixed at the microscale using a pipette, and then sent into the vacuum chamber of the mass spectrometer for detection. An automated time-series scanning sequence was set using the mass spectrometer control software. The laser beam was programmed to start from the geometric center of the sample point and move along an Archimedean spiral trajectory towards the edge. This scanning strategy covered different crystalline regions from the center to the edge of the sample point. The laser pulse frequency was set to 20 Hz, and the accumulated signal within each 0.5 seconds was treated as an independent acquisition unit. The scanning process lasted 90 seconds, with a total of 180 independent acquisition operations. During this process, the raw mass spectrometry signal was acquired, and the sampling time and the two-dimensional coordinates of the laser beam on the target plate plane (sampling position) were recorded simultaneously. The sampling time and two-dimensional coordinates were used as spatiotemporal information.

[0022] Considering that the raw mass spectrometry signal contains instrument background noise, baseline drift, and slight deviations in the mass axis due to target plate flatness, directly using the raw mass spectrometry signal would introduce calculation errors. Therefore, it is necessary to standardize and clean the raw mass spectrometry signal obtained from T acquisitions and establish a unified sampling time index. Where t represents the sampling time index, and T is the total number of samples.

[0023] For each acquisition, the top-hat transform algorithm is first used to remove the low-frequency baseline background, and then the Savitzky-Golay filter is used to smooth and suppress high-frequency random noise. The original mass spectrometry signal after the above processing is recorded as the mass spectrometry signal, and the ion intensity of the mass spectrometry signal is obtained.

[0024] Step S2: Screen the initial state components and final state components based on all the ion intensities and sampling times; obtain the difference factor of each component pair according to the ion intensities and sampling positions corresponding to each component pair; wherein, a component pair consists of an initial state component and a final state component.

[0025] A continuous wavelet transform (CWT) algorithm is used to extract all mass spectral peaks with a signal-to-noise ratio (SNR) higher than a preset SNR threshold from the mass spectrometry signals identified at different acquisition times. The mass-to-charge ratio (MMR) and ion intensity of these peaks are recorded. The preset SNR threshold can be 3. Further, all extracted mass spectral peaks with SNRs higher than the preset SNR threshold are arranged in descending order of MMR to obtain a mass spectral peak sequence. In this sequence, the absolute value of the difference between the MMRs of each pair of peaks is calculated. This absolute value is taken as the difference in MMR between the corresponding two peaks. Adjacent peaks in the mass spectral peak sequence with a difference less than or equal to a preset difference threshold are grouped into the same category, with each category forming a component. The preset difference threshold can be 0.5 Da.

[0026] To distinguish between the initial and final states of system evolution, it is necessary to screen out the key components that play a dominant role at both ends of the reaction from all the data.

[0027] Specifically, a start time window and an end time window are defined based on the sampling time; as one specific implementation, the time window will be calculated from the first sampling time to the last sampling time. The time interval between the sampling times is used as the starting time window, and the first sampling time is used as the starting time window. The time period between the first sampling time and the Tth sampling time is used as the termination time window, where: The symbol indicates rounding up.

[0028] Calculate the average value of all ion intensities for each component in the initial time window, and record this average value as the first average signal intensity; calculate the average value of all ion intensities for each component in the final time window, and record this average value as the second average signal intensity; each component has a corresponding first average signal intensity and a corresponding second average signal intensity.

[0029] Arrange all first average signal strengths in descending order of their first average signal strength to obtain a first average signal strength sequence. The components corresponding to the first preset number of first average signal strengths in the first average signal strength sequence are determined as initial state components; in this embodiment, the preset number is 20, but in specific applications, the implementer can set it according to specific circumstances. Simultaneously, arrange all second average signal strengths in descending order of their second average signal strength to obtain a second average signal strength sequence; the components corresponding to the first preset number of second average signal strengths in the second average signal strength sequence are determined as final state components.

[0030] Using the above method, multiple initial state components and multiple final state components can be obtained. The initial state components and final state components are combined in pairs to obtain component pairs, that is, each component pair contains one initial state component and one final state component.

[0031] Given that related components (such as precursors and products) in biochemical reactions usually exhibit causally related abundance trends over time, it is necessary to evaluate the similarity between initial and final state components in order to quantify this morphological association.

[0032] The following embodiment uses one component pair as an example for illustration. Other component pairs can be processed using the method provided in this embodiment.

[0033] Specifically, any pair of components is designated as a candidate component pair. For the initial-state component in the candidate component pair, all ion intensities of that component are arranged in chronological order to obtain the ion intensity sequence of that component. For the final-state component in the candidate component pair, all ion intensities of the final-state component are arranged in chronological order to obtain the ion intensity sequence of that component. The ion intensity sequences of the initial-state and final-state components are normalized using a maximum-minimum normalization method. It should be noted that all ion intensity sequences mentioned later are normalized sequences, and the maximum and minimum values ​​used in the data normalization can be obtained from statistical historical data. The Dynamic Time Warping Distance (DTW) between the ion intensity sequences of the two components in the candidate component pair is calculated, and this distance is used as the difference factor of the candidate component pair. The smaller the difference factor, the more similar the initial-state and final-state components in the candidate component pair are. It should be noted that when calculating the DTW distance between the ion intensity sequences of two components, the absolute physical time corresponding to the data values ​​in the two ion intensity sequences is ignored.

[0034] The difference factor for each component pair can be obtained through the above methods.

[0035] Step S3: Construct a position correction cost matrix based on the spatiotemporal information corresponding to each component pair and the difference factor; solve the global optimal matching for the position correction cost matrix; and construct a dynamic evolution feature vector based on the statistical characteristics of the cost values ​​on the optimal matching path.

[0036] During solvent evaporation, molecules with similar interactions or physicochemical properties tend to deposit in close spatial locations. Based on this characteristic, this embodiment will next evaluate the degree of co-localization of components on the target plate plane.

[0037] Specifically, for any component in any component pair: the spatial centroid coordinates of the component are determined by weighting the position coordinates of the component using the ionic strength corresponding to that component.

[0038] As a specific implementation method, a specific formula for calculating the spatial centroid coordinates is given. The spatial centroid coordinates of the i-th component can be expressed as: in, This represents the spatial centroid coordinates of the i-th component. This represents the total number of sampling times for the i-th component. Indicates the i-th component. The ion intensity corresponding to each sampling time. Indicates the first The two-dimensional coordinates corresponding to each sampling time.

[0039] Using the above method, the spatial centroid coordinates of each component in each component pair can be obtained. Next, the Gaussian kernel spatial weights of each component pair are determined by measuring the proximity of the spatial centroid coordinates of the two components in each pair.

[0040] Specifically, the spatial distribution overlap of each component pair is evaluated based on the Euclidean distance between the spatial centroid coordinates of each component pair, and the Gaussian kernel spatial weights of each component pair are obtained.

[0041] As a specific implementation method, a specific formula for calculating the Gaussian kernel space weights is given. The Gaussian kernel space weights of the component pairs formed by the u-th initial state component and the v-th final state component can be expressed as: in, Let represent the Gaussian kernel space weights of the component pairs consisting of the u-th initial state component and the v-th final state component. This represents the spatial centroid coordinates of the u-th initial state component. This represents the spatial centroid coordinates of the v-th final state component. Let represent the Euclidean distance between the spatial centroid coordinates of the component pair consisting of the u-th initial state component and the v-th final state component. Indicates the adaptive scaling parameter. This represents an exponential function with the natural constant as its base.

[0042] Adaptive scaling parameters The median of the Euclidean distance between the spatial centroid coordinates of all component pairs is set to the sum of the median and a preset minimum spatial smoothing factor. In this embodiment, the preset minimum spatial smoothing factor is 0.0001. This smoothing factor is introduced to prevent the median from being 0 in the case of an extremely homogeneous distribution of sample crystals, thereby ensuring the adaptive scaling parameter. The Gaussian kernel is always greater than 0, thus avoiding a crash exception where the denominator is zero during internal calculations of the exponential term. Furthermore, based on the mathematical properties of the natural exponential function, the Gaussian kernel space weights... Always being positive ensures that subsequent calculations of location correction associated costs are performed. The denominator is not zero.

[0043] The Gaussian kernel spatial weight is used to characterize the spatial overlap between the initial and final state components. The larger the value, the more the spatial distributions of the initial and final state components overlap.

[0044] To identify real biochemical evolutionary pathways that are both kinetically related and spatially proximate, it is necessary to fuse kinetic differences with spatial weights. This embodiment quantifies the location-corrected correlation cost of each component pair by combining the Gaussian kernel spatial weights and difference factors.

[0045] As a specific implementation method, a specific formula for calculating the position correction association cost is given. The position correction association cost of the component pair consisting of the u-th initial state component and the v-th final state component can be expressed as: in, Let $\frac{u}{v}$ represent the position correction association cost of the component pair consisting of the $u$-th initial state component and the $v$-th final state component. This represents the difference factor between the component pair consisting of the u-th initial state component and the v-th final state component. The Gaussian kernel space weight represents the component pair consisting of the u-th initial state component and the v-th final state component.

[0046] In the formula for calculating the location correction association cost, the Gaussian kernel space weight acts as a penalty factor. The closer the Gaussian kernel space weight of the component pair consisting of the u-th initial state component and the v-th final state component is to 0, the more the association cost will be amplified.

[0047] Using the above method, the location correction association cost for each component pair can be obtained. Based on the location correction association costs of all component pairs, a location correction cost matrix is ​​constructed. The number of dimensions of the location correction cost matrix is... The element in the u-th row and v-th column of the position correction cost matrix represents the position correction associated cost of the component pair consisting of the u-th initial state component and the v-th final state component. ,in, This indicates the preset quantity, that is, the quantity of the initial state component or the quantity of the final state component.

[0048] Furthermore, the Hungarian Algorithm is used to solve the minimum weight perfect matching problem of a bipartite graph on the position correction cost matrix. This algorithm outputs an allocation scheme that minimizes the total association cost. Based on this allocation scheme, the total association cost corresponding to the scheme and the cost values ​​of each cost value on the optimal matching path are extracted. The mean, variance, skewness, and kurtosis of each cost value on the optimal matching path are calculated. The total association cost corresponding to the allocation scheme with the minimum total association cost, along with the mean, variance, skewness, and kurtosis of each cost value on the optimal matching path, constitute a dynamically evolving feature vector. This dynamically evolving feature vector can be represented as: ,in, This represents the dynamic evolution feature vector of the microbial sample to be tested. This represents the total associated cost corresponding to the allocation scheme that minimizes total associated costs. This represents the average cost value along the optimal matching path. This represents the variance of the cost values ​​along the optimal matching path. This represents the skewness of each cost value on the optimal matching path. This represents the kurtosis of each cost value on the optimal matching path. It should be noted that when calculating skewness and kurtosis, if the standard deviation is 0, the sum of the standard deviation and the preset minimum value is used as the denominator in the skewness and kurtosis calculation formulas. The preset minimum value can be 0.0001.

[0049] The dynamic evolution feature vector eliminates random noise and retains only the spatiotemporal coupled evolutionary pattern determined by the species’ genetic background. Microorganisms of different species exhibit significant geometric separability in this new feature space.

[0050] Step S4: Determine the species category of the microbial sample to be tested based on the dynamic evolution feature vector.

[0051] In this embodiment, the dynamic evolution feature vector of the microbial sample to be tested is obtained through the above steps. Next, the species category of the microbial sample to be tested will be determined based on the dynamic evolution feature vector of the microbial sample to be tested.

[0052] Multiple standard strains were selected from historical databases to confirm species identity using gold standards (such as 16S rRNA sequencing), including Escherichia coli and Shigella strains for target differentiation. For each standard strain, the dynamic evolutionary feature vector acquisition method described in this embodiment was used to obtain the dynamic evolutionary feature vector for each standard strain. In this embodiment, the number of standard strains is 1000; in specific applications, the implementer can set this number according to specific circumstances.

[0053] Next, Support Vector Machine (SVM) is used as the classification algorithm. SVM aims to find a decision hyperplane that can separate samples of different classes with the maximum margin, which aligns with the geometrically separable property of the dynamically evolving feature vectors in this embodiment. The dynamically evolving feature vectors of the standard strain are used as the training dataset to train the SVM, solving for the parameters (normal vector and intercept) of the optimal classification hyperplane. The trained model is then designated as the reference strain classification model. The training of SVM is a prior art technique and will not be elaborated upon further here.

[0054] Furthermore, the dynamic evolution feature vector of the test sample is input into the trained reference bacterial classification model. The model calculates the geometric position of the dynamic evolution feature vector of the test sample in the feature space according to the decision hyperplane defined inside it. Specifically, it calculates the signed Euclidean distance from the data point to the decision hyperplane and outputs the identification result based on the distance: (1) Species determination: The sign of the distance determines which side of the hyperplane the sample is located on, thus directly corresponding to the species category in the binary classification problem (e.g., a positive value corresponds to Escherichia coli, and a negative value corresponds to Shigella); (2) Confidence assessment: The absolute value of the distance reflects the degree to which the feature vector deviates from the decision boundary. The farther the distance, the more typical the evolution pattern of the sample is, and the more reliable the classification result is. The geometric distance is transformed into a geometric distance by methods such as Platt scaling. The posterior probability value within the interval is used as the classification confidence output. For example, the final output is: "Identification result: Shigella; Confidence: " This indicates that the component evolution characteristics of the sample under test are highly consistent with those of Shigella and are far from the region of confusion with Escherichia coli in the feature space.

[0055] Thus, by using the method provided in this embodiment, accurate identification of microbial species has been achieved.

[0056] This embodiment acquires the ionic intensity and corresponding spatiotemporal information of the microbial sample to be tested, effectively capturing the dynamic evolution information driven by endogenous biochemical differences during the preparation process of the microbial sample to be tested. Based on the ionic intensity and sampling time, the initial and final state components are screened, and the difference factor is calculated according to the ionic intensity and sampling position of each component to construct a position correction cost matrix. By solving the global optimal match of this matrix, and constructing a dynamic evolution feature vector based on the statistical characteristics of the cost values ​​on the optimal matching path, the spatiotemporal coupling evolution information lost in traditional methods can be transformed into quantifiable features with species specificity. Finally, the species category of the microorganism to be tested is determined based on this dynamic evolution feature vector, which improves the identification accuracy between microorganisms with highly similar genetic backgrounds and overcomes the technical problem of highly overlapping feature spectra and difficulty in effective differentiation in existing static final product analysis methods.

[0057] An embodiment of a device for improving the accuracy of microbial identification: See Figure 2 The diagram illustrates a structural block diagram of an apparatus for improving the accuracy of microbial identification according to an embodiment of the present invention. The apparatus may include an acquisition module, an analysis module, a vector construction module, and an identification module.

[0058] The acquisition module is used to acquire the ionic strength and corresponding spatiotemporal information of the microbial sample to be tested, wherein the spatiotemporal information includes the sampling time and sampling location; The analysis module is used to screen the initial state components and final state components based on all the ion intensities and sampling times; and to obtain the difference factor of each component pair according to the ion intensities corresponding to each component pair and the sampling location; wherein, a component pair consists of an initial state component and a final state component; The vector construction module is used to construct a position correction cost matrix based on the spatiotemporal information corresponding to each component pair and the difference factor; solve the global optimal matching for the position correction cost matrix; and construct a dynamic evolution feature vector based on the statistical characteristics of the cost values ​​on the optimal matching path. The identification module is used to determine the species category of the microbial sample to be tested based on the dynamic evolution feature vector.

[0059] It should be understood that Figure 2 The structural block diagram and modules of the device for improving the accuracy of microbial identification shown can be implemented in various ways. For example, in some embodiments, the device and modules can be implemented by hardware, software, or a combination of software and hardware. The hardware portion can be implemented using dedicated logic; the software portion can be stored in memory and executed by a suitable instruction execution device, such as a microprocessor or dedicated hardware. Those skilled in the art will understand that the above-described methods and devices can be implemented using computer-executable instructions and / or included in processor control code, for example, such code provided on a carrier medium such as a disk, CD, or DVD-ROM, a programmable memory such as read-only memory (firmware), or a data carrier such as an optical or electronic signal carrier. The devices and modules of this specification can be implemented not only by hardware circuits such as very large-scale integrated circuits or gate arrays, semiconductors such as logic chips, transistors, or programmable hardware devices such as field-programmable gate arrays, programmable logic devices, etc., but also by software executed by various types of processors, or by a combination of the above-described hardware circuits and software (e.g., firmware).

[0060] For more details about the above modules, please refer to other parts of this manual; they will not be repeated here.

[0061] The provided device is used to execute the corresponding method provided above. Therefore, the beneficial effects it can achieve can be referred to the beneficial effects of the corresponding method provided above, and will not be repeated here.

[0062] It should be noted that the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A method for improving the accuracy of microbial identification, characterized in that, The method includes the following steps: The ionic strength and corresponding spatiotemporal information of the microbial sample to be tested are obtained, wherein the spatiotemporal information includes the sampling time and sampling location; Based on all the stated ion intensities and sampling times, the initial and final state components are screened; the difference factor of each component pair is obtained according to the ion intensities and sampling locations corresponding to each component pair; wherein, a component pair consists of an initial state component and a final state component; Based on the spatiotemporal information corresponding to each component pair and the difference factor, a position correction cost matrix is ​​constructed; the global optimal matching is solved for the position correction cost matrix, and a dynamic evolution feature vector is constructed based on the statistical characteristics of the cost values ​​on the optimal matching path; The species category of the microbial sample to be tested is determined based on the dynamic evolution feature vector.

2. The method for improving the accuracy of microbial identification according to claim 1, characterized in that, The acquisition of each component includes: Obtain the mass-to-charge ratio of mass spectral peaks with a signal-to-noise ratio higher than a preset signal-to-noise ratio threshold at different acquisition times; All mass spectrometry peaks with a signal-to-noise ratio higher than a preset signal-to-noise ratio threshold are arranged in descending order of mass-to-charge ratio to obtain a mass spectrometry peak sequence. In the mass spectrometry peak sequence, the absolute value of the difference between the mass-to-charge ratios of each pair of mass spectrometry peaks is calculated, and the absolute value is taken as the difference between the mass-to-charge ratios of the corresponding two mass spectrometry peaks. Mass spectrometry peaks in the mass spectrometry peak sequence whose difference is less than or equal to the preset difference threshold and which are adjacent are grouped into the same category, and each category is taken as a component.

3. The method for improving the accuracy of microbial identification according to claim 2, characterized in that, The screening of initial and final state components based on all the stated ion intensities and sampling times includes: Define the start and end time windows based on the sampling time; Calculate the first average signal strength of each component within the initial time window and the second average signal strength within the final time window; Arrange all the first average signal indices in descending order to obtain a first average signal intensity sequence; determine the components corresponding to the first preset number of first average signal indices in the first average signal intensity sequence as the initial state components; Arrange all the second average signal indices in descending order to obtain a second average signal intensity sequence; determine the components corresponding to the first preset number of second average signal indices in the second average signal intensity sequence as the final state components.

4. The method for improving the accuracy of microbial identification according to claim 2, characterized in that, The step of obtaining the difference factor for each component pair based on the ion intensity and sampling location corresponding to each component pair includes: The DTW distance between the ion intensity sequences of the two components in the candidate component pair is calculated, and the DTW distance is used as the difference factor of the candidate component pair; wherein, the ion intensity sequence of the initial state component in the candidate component pair is obtained by arranging all the ion intensities of the obtained initial state component in chronological order from front to back, and the ion intensity sequence of the final state component in the candidate component pair is obtained by arranging all the ion intensities of the obtained final state component in chronological order from back to front. The candidate component pair can be any component pair.

5. The method for improving the accuracy of microbial identification according to claim 2, characterized in that, The step of constructing a location correction cost matrix based on the spatiotemporal information corresponding to each component pair and the difference factor includes: For any component in any component pair: the spatial centroid coordinates of any component are determined by weighting the position coordinates of the corresponding component using the ionic strength of the corresponding component. Based on the Euclidean distance between the spatial centroid coordinates of each component pair, the Gaussian kernel spatial weight of each component pair is obtained; combined with the Gaussian kernel spatial weight and the difference factor of each component pair, the position correction association cost of each component pair is obtained. Based on the location correction associated costs of all component pairs, a location correction cost matrix is ​​constructed.

6. The method for improving the accuracy of microbial identification according to claim 1, characterized in that, The step of solving the global optimal matching of the position correction cost matrix includes: using a bipartite graph matching algorithm to solve the position correction cost matrix to obtain an allocation scheme that minimizes the total association cost.

7. The method for improving the accuracy of microbial identification according to claim 6, characterized in that, The construction of a dynamically evolving feature vector based on the statistical characteristics of cost values ​​on the optimal matching path includes: Obtain the total associated cost corresponding to the allocation scheme and the cost values ​​of each item on the optimal matching path; Calculate the mean, variance, skewness, and kurtosis of each cost value on the optimal matching path; The total associated cost, mean, variance, skewness, and kurtosis constitute a dynamic evolution feature vector.

8. The method for improving the accuracy of microbial identification according to claim 1, characterized in that, The process of determining the species category of the microbial sample to be tested based on the dynamic evolution feature vector includes: The dynamic evolution feature vector of the sample to be tested is input into the trained reference microbial classification model to obtain the species category of the microbial sample to be tested.

9. The method for improving the accuracy of microbial identification according to claim 1, characterized in that, The acquisition of the ionic strength and corresponding spatiotemporal information of the microbial sample to be tested includes: The laser beam of the mass spectrometer is controlled to continuously scan and acquire the microbial sample to be tested along a preset path to obtain spatiotemporal information and raw mass spectrometry signal; the ion intensity is obtained based on the raw mass spectrometry signal.

10. An apparatus for improving the accuracy of microbial identification, the apparatus being used to implement the method of claim 1, characterized in that, The device includes: The acquisition module is used to acquire the ionic strength and corresponding spatiotemporal information of the microbial sample to be tested, wherein the spatiotemporal information includes the sampling time and sampling location; The analysis module is used to screen the initial state components and final state components based on all the ion intensities and sampling times; and to obtain the difference factor of each component pair according to the ion intensities corresponding to each component pair and the sampling location; wherein, a component pair consists of an initial state component and a final state component; The vector construction module is used to construct a position correction cost matrix based on the spatiotemporal information corresponding to each component pair and the difference factor; solve the global optimal matching for the position correction cost matrix; and construct a dynamic evolution feature vector based on the statistical characteristics of the cost values ​​on the optimal matching path. The identification module is used to determine the species category of the microbial sample to be tested based on the dynamic evolution feature vector.