Target pathogenic bacterium immune antibody detection method based on DNA detection
By combining DNA sequence and immune antibody data analysis, high-frequency mutation regions and antibody response characteristics are identified, and detection thresholds are dynamically adjusted. This solves the problem of identifying and warning of the risk of immune escape from target pathogens, and improves the accuracy of detection and early warning capabilities.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-22
- Publication Date
- 2026-03-20
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing detection methods cannot effectively identify the risk of immune escape caused by gene mutations in target pathogens, and the fixed antibody affinity detection threshold leads to misjudgment and missed detection. Furthermore, the lack of synergistic analysis of changes in bacterial copy number and antibody expression levels makes it impossible to provide timely warnings of infection recurrence or disease aggravation.
By acquiring DNA sequence data of target pathogenic bacteria and expression data of immune antibody proteins, the boundaries of high-frequency mutation regions are identified, the sequence mutation degree and antibody response intensity indicators are calculated, and the historical correlation between the two is combined to construct an immune escape risk coefficient and an antibody inhibition trend factor. The antibody affinity detection threshold is dynamically adjusted to output an effective set of neutralizing antibodies.
It enables precise capture of genotype changes in target pathogenic bacteria, dynamically reflects the strength of the host's immune response, timely identifies the risk of immune escape, accurately assesses antibody inhibition, provides reliable basis for treatment plans, and improves detection accuracy and early warning capabilities.
Smart Images

Figure CN121709029A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of biological detection, and particularly to a target pathogenic bacteria immune antibody detection method based on DNA detection. BACKGROUND
[0002] In the diagnosis and prevention process of diseases related to target pathogenic bacteria infection, effective detection of target pathogenic bacteria immune antibodies is an important means to monitor the infection state and evaluate the treatment effect. Traditional target pathogenic bacteria immune antibody detection methods mostly rely on direct detection of antigen-antibody reactions. Although such methods are relatively simple to operate, they can only detect known bacterial antigens and are difficult to cope with antigen changes caused by bacterial gene mutations, which may lead to missed detection or false detection.
[0003] With the development of molecular biology technology, detection methods based on target pathogenic bacteria DNA sequence analysis have been gradually applied in the field of bacterial detection. By determining and analyzing the DNA sequence of target pathogenic bacteria, the genotypic characteristics of bacteria can be identified to assist in judging the species and variation of bacteria. However, existing methods based on target pathogenic bacteria DNA detection often only focus on the sequence changes of bacteria themselves and fail to establish a correlation between bacterial gene mutations and immune antibody responses produced by the host immune system, making it impossible to fully reflect the infection dynamics and immune escape potential of bacteria in the host.
[0004] In practical applications, target pathogenic bacteria are prone to gene mutations during proliferation, especially in the gene regions related to antigen epitopes. High-frequency mutations may lead to changes in the antigenicity of target pathogenic bacteria, thereby evading the recognition and attack of the host immune system, i.e., the immune escape phenomenon. At this time, even if the host immune system produces a large number of immune antibodies, the antibodies may not be able to effectively neutralize the target pathogenic bacteria due to the decreased affinity between the antibodies and the mutated target pathogenic bacteria antigens, resulting in detection results that cannot accurately reflect the actual immune protection effect. Traditional detection methods and single DNA sequence analysis methods cannot effectively solve this problem, making it difficult to accurately assess the current immune status of the host in terms of the immune escape risk of target pathogenic bacteria.
[0005] Current detection methods often employ fixed thresholds for antibody affinity detection, failing to adjust for dynamic changes in antibody response intensity based on the genetic mutations of the target pathogen. When significant genetic mutations occur in the target pathogen or antibody response intensity fluctuates considerably, fixed thresholds can easily lead to misjudgments of effective neutralizing antibodies. This may result in either identifying low-affinity, ineffective antibodies as effective ones or missing some effective antibodies with certain affinity, impacting the reliability of the test results and the formulation of subsequent treatment plans. Furthermore, existing methods lack analysis of the synergistic relationship between changes in the target pathogen's copy number and antibody expression levels, making it difficult to promptly detect the proliferative trend of the target pathogen due to weakened antibody inhibition, and hindering early warning of the risk of infection recurrence or disease exacerbation. Summary of the Invention
[0006] The purpose of this invention is to provide a method for detecting target pathogenic bacteria immune antibodies based on DNA detection, so as to solve the problems mentioned in the background art.
[0007] To achieve the above objectives, the present invention provides a method for detecting target pathogenic bacteria immune antibodies based on DNA detection, the method comprising:
[0008] Obtain the target pathogen bacterial DNA sequence data and immune antibody protein expression level data of the biological sample to be tested;
[0009] Mutation site scanning was performed on the DNA sequence data of the target pathogenic bacteria to identify the boundaries of high-frequency mutation regions;
[0010] Based on the distribution density and base substitution frequency of the high-frequency mutation region boundary, the sequence mutation degree of the target pathogenic bacterial DNA sequence is calculated.
[0011] The periodic fluctuation characteristics of the immune antibody protein expression data are extracted simultaneously, and an antibody response intensity index is generated based on the fluctuation amplitude change rate and phase shift.
[0012] By analyzing the historical correlation between the sequence mutation rate and the antibody response intensity index, and combining the synergy of their changing trends, the immune escape risk coefficient at the current moment is determined.
[0013] The antibody inhibition trend factor was constructed by detecting the recent decline slope of the expression level of the immune antibody protein and the increase rate of the copy number of the target pathogen bacterial DNA sequence.
[0014] By combining the immune escape risk coefficient and the antibody inhibition trend factor, an adaptive evolution index of the target pathogen bacteria is generated;
[0015] The antibody affinity detection threshold is dynamically adjusted based on the gradient of the adaptive evolution index of the target pathogenic bacteria and the preset detection sensitivity threshold.
[0016] Based on the adjusted antibody affinity detection threshold, the expression data of the immune antibody proteins are screened for affinity, and an effective set of neutralizing antibodies is output.
[0017] Preferably, the identification of the high-frequency mutation region boundary further includes:
[0018] Divide the target pathogenic bacterial DNA sequence data into a preset length sliding window and count the frequency of non-conserved bases in each window;
[0019] Regions where the frequency of three or more consecutive windows exceeds the mutation threshold are marked as candidate mutation regions;
[0020] Calculate the boundary entropy value of the candidate mutation region, merge adjacent regions whose boundary entropy value difference is less than the tolerance factor, and output the boundary of the merged high-frequency mutation region.
[0021] Preferably, calculating the sequence mutation degree of the target pathogenic bacterial DNA sequence further includes:
[0022] Extract the length of the boundary and the coordinates of the center position of each high-frequency mutation region;
[0023] The deviation of the base substitution type from the standard reference sequence in each region was statistically analyzed;
[0024] The weighted product of the length, the center position dispersion, and the deviation is used as the local mutation contribution value of each region;
[0025] The sequence mutation degree is generated by normalizing and summing all local mutation contribution values.
[0026] Preferably, the antibody response strength indicator further includes:
[0027] The expression data of the immune antibody proteins were segmented in the time domain, and the peak and trough positions in each time period were extracted.
[0028] Calculate the average amplitude difference between adjacent peaks and troughs as the rate of change of fluctuation amplitude;
[0029] The absolute value of the phase deviation between the actual peak position and the theoretical immune cycle is measured to generate the phase offset.
[0030] The inverse of the rate of change of the fluctuation amplitude is linearly combined with the phase offset to output an antibody response intensity index.
[0031] Preferably, determining the immune escape risk coefficient at the current moment further includes:
[0032] Establish a sliding correlation coefficient matrix between the sequence mutation rate and the antibody response intensity index within a historical time window;
[0033] Calculate the eigenvalue variance of the correlation coefficient matrix as the historical correlation decay rate;
[0034] Extract the first-order difference of sequence mutation degree and the first-order difference of antibody response intensity index from consecutive time points before the current time.
[0035] The sum of the products of two differences at the same time point is used as the degree of synergy in the trend of change.
[0036] The geometric mean of the historical correlation decay and the correlation of the change trend is used as the immune escape risk coefficient.
[0037] Preferably, the construction of the antibody inhibition trend factor further includes:
[0038] Extract the expression level data of immune antibody proteins at the most recent preset time and fit the first derivative of its logarithmic change curve;
[0039] The mean of the smallest negative segment of the first derivative is taken as the recent descent slope;
[0040] The logarithmic increment of the target pathogen bacterial DNA sequence copy number per unit time is calculated simultaneously and used as the copy number increase rate.
[0041] The antibody inhibition trend factor is generated by taking the logarithm of the ratio of the absolute value of the recent decline slope to the copy number increase rate.
[0042] Preferably, the generation of the target pathogenic bacteria adaptive evolution index further includes:
[0043] The immune escape risk coefficient is transformed by the Sigmoid function to obtain a standardized risk value;
[0044] Multiply the antibody inhibition trend factor by a preset evolution acceleration weighting coefficient;
[0045] The standardized risk value is added to the weighted antibody inhibition trend factor to output the adaptive evolution index of the target pathogen bacteria.
[0046] Preferably, the dynamic adjustment of the antibody affinity detection threshold further includes:
[0047] Calculate the difference in the adaptive evolution index of the target pathogen bacteria between the current time and the previous detection time;
[0048] When the difference exceeds the evolutionary gradient threshold, the antibody affinity detection threshold is reduced by a preset step size.
[0049] When the difference is lower than the negative gradient threshold twice in a row, the antibody affinity detection threshold is increased by a fixed ratio.
[0050] Preferably, the output effective neutralizing antibody set further includes:
[0051] Under the adjusted antibody affinity detection threshold, records in the immune antibody protein expression data whose affinity parameters exceed the threshold are screened.
[0052] The antibody-encoding gene identifiers corresponding to the records are aggregated to generate an effective set of neutralizing antibodies.
[0053] Preferably, the method further includes:
[0054] Map the boundaries of the high-frequency mutation regions to the target pathogen bacterial antigen protein coding regions;
[0055] Cross-align the binding site coordinates of the effective neutralizing antibody set with the boundary of the mutation region;
[0056] Output a subset of neutralizing antibodies covering key mutation sites and their predicted binding free energy.
[0057] Compared with the prior art, the beneficial effects of the present invention are:
[0058] By simultaneously acquiring target pathogen bacterial DNA sequence data and immune antibody protein expression data, and performing multi-dimensional correlation analysis between the two, this method effectively overcomes the shortcomings of traditional detection methods that focus only on a single indicator and cannot establish a correlation between target pathogen bacterial variation and immune response. In terms of target pathogen bacterial DNA sequence analysis, by scanning mutation sites to identify the boundaries of high-frequency mutation regions, and calculating the sequence mutation degree based on the distribution density and base substitution frequency of these boundaries, the method can accurately capture the variation characteristics of the target pathogen bacterial gene sequence, especially high-frequency mutation regions related to antigenicity. This comprehensively reflects the dynamic changes in the target pathogen bacterial genotype, providing detailed sequence-level evidence for subsequent analysis of the target pathogen bacterial immune escape potential, and avoiding detection biases caused by traditional methods neglecting target pathogen bacterial gene mutations.
[0059] In the analysis of immune antibody protein expression levels, this method extracts the periodic fluctuation characteristics of antibody expression data and combines the rate of change of fluctuation amplitude and phase shift to generate an antibody response intensity index. This can dynamically reflect the changes in the host's immune system response intensity to the target pathogen infection. Compared with traditional methods that only detect the absolute value of antibody expression, this method can better reflect the dynamic trend of antibody response and effectively identify the stages of antibody response enhancement or weakening, providing richer information for judging the host's immune status.
[0060] By analyzing the synergy between the historical correlation between sequence mutation rate and antibody response intensity, the immune escape risk coefficient can be determined. This approach closely links the gene mutation status of the target pathogen with the host's immune response status, clearly revealing the potential for immune escape due to gene mutations. When the number of high-frequency mutation regions and the frequency of base substitution in the target pathogen increase, and the synergy between the trend of antibody response intensity and the trend of sequence mutation rate decreases, an increased risk of immune escape can be identified in a timely manner. This correlation analysis avoids the problem that analyzing DNA sequence or antibody expression levels alone cannot assess the risk of immune escape, and helps to predict the possibility of the target pathogen escaping host immune attack in advance.
[0061] In constructing the antibody inhibition trend factor, this method simultaneously considers the recent decline slope of immune antibody protein expression and the increase rate of target pathogen bacterial DNA sequence copy number, which can intuitively reflect the changes in the inhibitory effect of antibodies on target pathogen bacteria. When antibody expression decreases and the target pathogen bacterial copy number increases, the antibody inhibition trend factor can accurately reflect the weakening of antibody inhibition, promptly detect the potential proliferation trend of target pathogen bacteria due to weakened inhibition, and provide an important reference for early warning of infection recurrence or disease aggravation. This overcomes the shortcomings of existing methods that lack analysis of the synergistic changes between the two and cannot promptly detect the risk of target pathogen bacterial proliferation.
[0062] The adaptive evolutionary index of target pathogens, generated by integrating immune escape risk coefficients and antibody inhibition trend factors, comprehensively reflects the evolutionary trend of the target pathogens' adaptability within the host. It encompasses both the immune escape potential resulting from gene mutations and the proliferative capacity of the target pathogens under antibody inhibition, providing a comprehensive indicator for assessing the threat level posed by the target pathogens to the host. Based on the gradient of this index's changes and a preset detection sensitivity threshold, the antibody affinity detection threshold is dynamically adjusted. This allows the threshold setting to match the mutation status of the target pathogens and the antibody response status in real time, avoiding the problem of false positives for effective antibodies caused by fixed thresholds and significantly improving the accuracy of antibody affinity screening.
[0063] Based on the adjusted antibody affinity detection threshold, the data on the expression levels of immune antibody proteins are screened, and the output set of effective neutralizing antibodies is more in line with the actual needs of immune protection. This can provide a reliable basis for clinical development of targeted treatment plans, evaluation of treatment effects, and monitoring of infection status. At the same time, it provides more comprehensive and accurate technical support for the diagnosis and prevention of diseases related to target pathogen bacterial infections, and helps to promote the development of target pathogen bacterial immune detection technology towards a more efficient direction and more in line with the needs of practical applications. Attached Figure Description
[0064] Figure 1 This is a schematic diagram illustrating the working principle of the target pathogen bacterial immune antibody detection method based on DNA detection described in this invention.
[0065] Figure 2 A flowchart for calculating sequence mutation degree;
[0066] Figure 3 Flowchart for determining the risk factor of immune escape;
[0067] Figure 4 Flowchart for constructing antibody inhibition trend factors;
[0068] Figure 5 This is a flowchart for the analysis of neutralizing antibodies at key mutation sites. Detailed Implementation
[0069] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0070] Please see Figure 1 This invention provides a method for detecting target pathogenic bacteria immune antibodies based on DNA detection, the method comprising:
[0071] This process involves acquiring the target pathogen's bacterial DNA sequence data and the expression levels of immune antibody proteins from the biological samples to be tested; scanning for mutation sites in the target pathogen's bacterial DNA sequence data to identify the boundaries of high-frequency mutation regions; calculating the sequence mutation degree of the target pathogen's bacterial DNA sequence based on the distribution density and base substitution frequency of the high-frequency mutation region boundaries; simultaneously extracting the periodic fluctuation characteristics of the immune antibody protein expression data, and generating an antibody response intensity index based on the rate of change of fluctuation amplitude and phase shift; determining the immune escape risk coefficient at the current moment by analyzing the historical correlation between the sequence mutation degree and the antibody response intensity index, and combining the synergy of their changing trends; constructing an antibody inhibition trend factor by detecting the recent decline slope of the immune antibody protein expression data and the copy number increase rate of the target pathogen's bacterial DNA sequence; fusing the immune escape risk coefficient and the antibody inhibition trend factor to generate the target pathogen's bacterial adaptive evolution index; dynamically adjusting the antibody affinity detection threshold based on the gradient of the target pathogen's bacterial adaptive evolution index and the preset detection sensitivity threshold; and performing affinity screening on the immune antibody protein expression data based on the adjusted antibody affinity detection threshold to output an effective set of neutralizing antibodies.
[0072] Example 1: See Figure 2In the implementation process, for the processing of DNA sequence data of the target pathogen bacteria, a fixed-length analysis window is first set. This window moves stepwise along the DNA sequence, with a fixed step size each time. At each window position, the number of all bases within that window that do not match the standard reference sequence is counted in detail. The standard reference sequence is usually a known original sequence or consensus sequence of the target pathogen bacterial strain. After the statistics are completed, the number of non-matching bases counted in each window is compared with the total number of bases in that window, and the proportion is calculated.
[0073] A mutation frequency threshold is set to determine whether a window possesses significant mutational characteristics. When the mutation frequency of multiple consecutive windows exceeds this threshold, the DNA regions covered by these consecutive windows are initially marked as candidate mutation regions. The number of consecutive windows is required to ensure the coherence and significance of the regions. Subsequently, the stability of the boundary regions of each candidate mutation region is evaluated. The boundary region is defined as a fixed-length sequence fragment upstream of the start position and downstream of the end position of the candidate mutation region. Sequence complexity is quantitatively analyzed on these boundary sequence fragments, and their information entropy values are calculated. The information entropy value reflects the randomness or uncertainty of the sequence; the lower the entropy value, the more conserved or ordered the sequence.
[0074] The boundary entropy values of adjacent candidate mutation regions are compared. If the difference in boundary entropy values between adjacent regions is below a preset tolerance range, the boundary stability of the two regions is considered similar, and they are merged into a larger high-frequency mutation region. This merging operation helps reduce regional fragmentation caused by the division of the analysis window, more accurately reflecting the actual continuous mutation hotspot regions. The final output is the boundary information of the merged high-frequency mutation region, including the start position, end position, and length of each region.
[0075] After obtaining the boundaries of high-frequency mutation regions, the sequence mutation rate is calculated. For each identified high-frequency mutation region, its physical length is recorded, and the coordinates of its center point on the DNA sequence are calculated. The center point coordinates are obtained by averaging the start and end positions. Simultaneously, the specific base substitution types occurring within this region are analyzed in detail. The types and frequencies of base changes occurring at each site within the region are statistically analyzed and compared with the base composition at the corresponding positions in the standard reference sequence. This comparison aims to quantify the degree of difference between the mutation pattern of the region and the background distribution of the reference sequence, i.e., the deviation.
[0076] To comprehensively assess the impact of each high-frequency mutation region on overall sequence variation, the concept of a local mutation contribution value is introduced. This value is determined by three key parameters: the physical length of the region, the dispersion of the region's center point relative to the average center point of all identified high-frequency mutation regions, and the deviation of base substitutions within the region. These three parameters measure the size of the mutation region, its spatial distribution characteristics, and the differences in mutation properties, respectively. Each parameter is assigned a weighting coefficient, and the three weighted parameter values are multiplied to obtain the local mutation contribution value for that region. The weighting coefficients are used to adjust the influence of different factors on the final contribution value.
[0077] After calculating the local mutation contribution values of all high-frequency mutation regions, normalization is performed. Normalization eliminates differences in numerical scale, facilitating comparisons between different regions and overall summarization. The maximum value among all local mutation contributions is used as the normalization benchmark; each contribution value is divided by this maximum value to obtain the normalized contribution value. Finally, the normalized contribution values of all high-frequency mutation regions are summed, and the sum represents the sequence mutation degree of the entire target pathogen bacterial DNA sequence. This sequence mutation degree is a comprehensive indicator that quantifies the overall variation intensity and distribution characteristics of high-frequency mutation regions in the sequence. For example, when analyzing a specific sample, the analysis window length is set to 100 base pairs, and the window movement step size is 50 base pairs. The mutation frequency threshold is set to 15%, meaning that a window with more than 15% of non-conserved bases is considered a high-mutation window. Only windows meeting the threshold are marked as candidate mutation regions. The boundary entropy analysis range is set to 10 base pairs before and after the boundary position. The tolerance factor is set to 0.2, meaning that adjacent candidate regions with a boundary entropy difference of less than 0.2 are merged. When calculating the local mutation contribution value, the length weight is set to 0.3, the center position dispersion weight to 0.2, and the deviation weight to 0.5. The center position dispersion is obtained by calculating the absolute distance between the center coordinates of the region and the average center coordinates of all regions. The deviation calculation involves the frequency of various base substitutions such as A->G and C->T within the statistical region, and comparing it with the expected background frequency of base substitutions at the corresponding position in the reference sequence (possibly based on sequence context or evolutionary models), calculating its KL divergence. The sum of the normalized contribution values of all regions is the final sequence mutation degree. This process is executed automatically by a computer program, with the input being target pathogen bacterial DNA sequencing data (FASTA format) and a standard reference sequence, and the output being the sequence mutation degree value and a detailed list of mutation region coordinates.
[0078] Example 2: See Figure 3For processing data on the expression levels of immune antibody proteins, the continuous antibody concentration time series is first divided into time periods according to a fixed cycle. Each time period corresponds to a complete immune response cycle, and its duration is set according to the organism's immune rhythm. Within each divided time period, a specific algorithm is used to identify the peak and trough locations of antibody concentration changes. This algorithm determines the local highest and lowest points of the concentration curve by comparing the numerical relationships between adjacent data points and combining numerical comparisons within a local range.
[0079] After identifying the peak and trough locations for each time period, the concentration difference between adjacent peaks and troughs is calculated. This difference reflects the amplitude of antibody concentration change within a fluctuation cycle. The average of these differences over multiple consecutive cycles is taken as an indicator of the intensity of antibody concentration fluctuations, i.e., the rate of change of fluctuation amplitude. The larger this value, the greater the oscillation amplitude of antibody concentration. Simultaneously, the deviation between the time point of antibody concentration peak occurrence and the expected time point is analyzed. The expected time point is determined based on a pre-defined theoretical immune cycle model, which defines the time interval for the immune response to reach its peak under ideal conditions. The absolute value of the time difference between the actual detected peak time point and the theoretical time point is calculated; this value is the phase shift. The phase shift reflects the degree of deviation of the actual immune response from the theoretical model in the time dimension.
[0080] The reciprocal of the rate of change of fluctuation amplitude is combined with the phase offset. The reciprocal of the rate of change of fluctuation amplitude serves as a reverse indicator of fluctuation intensity, while the phase offset directly reflects the time deviation. The two are linearly weighted and summed according to a preset proportional coefficient to generate a comprehensive antibody response intensity index. Changes in this index value can reflect the overall response status of the body's immune system to the stimulation of target pathogenic bacteria.
[0081] In another part of the implementation, it is necessary to determine the immune escape risk coefficient. This process first establishes a dynamic correlation model between sequence mutation rate and antibody response intensity index over a historical period. A fixed-length historical time window is selected, and within this window, the statistical correlation between sequence mutation rate and antibody response intensity index is calculated step by step using smaller sliding windows as units. Each time a time unit is slid, the correlation value corresponding to that unit of time is calculated. This process is repeated until the entire historical time window is covered, forming a series of correlation values arranged in chronological order, i.e., the sliding correlation coefficient matrix. The characteristics of this correlation coefficient matrix are analyzed. The main eigenvectors of the matrix are extracted through mathematical transformations, and the dispersion of the distribution of these eigenvalues, i.e., the eigenvalue variance, is calculated. This variance is defined as the historical correlation decay, which quantifies the stability of the historical correlation between sequence mutation rate and antibody response intensity index. The higher the decay, the greater the fluctuation in historical correlation and the lower the stability.
[0082] The trends of sequence mutation rate and antibody response intensity indices were analyzed at the most recent consecutive time points. The numerical change at each time point relative to the previous time point was calculated, i.e., the first-order difference value. The sequence mutation rate difference value at the same time point was multiplied by the antibody response intensity index difference value. The products from multiple consecutive time points were summed to obtain the trend coherence. This value reflects the degree of consistency between the two indices in the short-term direction of change. The historical correlation decay and trend coherence were combined. The geometric mean method was used, i.e., the square root of the product of the two, as the final immune escape risk coefficient. This coefficient integrates the stability of long-term historical correlation and the coherence of short-term trends to assess the potential risk level of immune escape from the target pathogen at the current moment.
[0083] When analyzing a specific sample, antibody concentration time-series data were obtained from hourly monitoring records of the patient over seven consecutive days. The time domain was segmented into seven complete cycles, each lasting 24 hours. Within each 24-hour cycle, a local extremum-based algorithm was used to identify peak and trough locations, with a comparison window of five data points (five hours). The concentration difference between adjacent peaks and troughs was calculated, and the arithmetic mean of these differences over the seven cycles was taken as the rate of change of fluctuation amplitude. The theoretical immune cycle was set at the 24-hour mark, meaning a peak was expected to occur at midnight each day. The actual peak time was determined through detection, and the absolute value of the time difference between it and midnight (in hours) was calculated as the phase shift. The reciprocal of the rate of change of fluctuation amplitude was multiplied by a coefficient of 0.6, and the phase shift was multiplied by a coefficient of 0.4; the sum of these two values yielded the antibody response intensity index.
[0084] For calculating the immune escape risk coefficient, the historical time window is set to the most recent 30 days. A sliding window of 7 days with a step size of 1 day is used. The Pearson correlation coefficient between the sequence mutation rate and the antibody response intensity index within the corresponding 7-day window is calculated, forming a 30-dimensional correlation coefficient sequence (matrix). The eigenvalues of the covariance matrix of this sequence are calculated, and the variance of the eigenvalues is used as the historical correlation decay. Simultaneously, the sequence mutation rate and antibody response intensity index values for the most recent five consecutive time points (one point per day) are extracted, and the first-order difference for each day (i.e., the current day's value minus the previous day's value) is calculated. The two difference values for the same day are multiplied, and the sum of the products over five days is used to obtain the trend coherence. Finally, the immune escape risk coefficient is the geometric mean of the historical correlation decay and the trend coherence (i.e., the square root of their product). This process is automatically executed by a computer program, with time series data and preset parameters as inputs, and antibody response intensity index and immune escape risk coefficient values as outputs.
[0085] Example 3: See Figure 4To analyze recent trends in immune antibody protein expression levels, we first extracted concentration monitoring data over a recent fixed time period. This period is sufficient to cover a typical immune response cycle. Logarithmic transformation was applied to the antibody concentration values during this time period to linearize the trend and reduce variations in numerical range. Subsequently, curve fitting was performed on the transformed logarithmic concentration values to smooth random fluctuations and reveal their intrinsic patterns. Based on the fitted curve, its first derivative was calculated, reflecting the instantaneous rate of change of logarithmic concentration over time. From these derivative values, the longest-lasting or lowest-valued negative intervals were selected, and the average value of the derivatives within these intervals was calculated as the recent decline slope of antibody concentration. This slope quantifies the rate of decline of antibody concentration in recent times.
[0086] The amplification of the target pathogen's DNA sequence was analyzed simultaneously, and the change in the target pathogen's DNA copy number per unit time was statistically analyzed. This change was also logarithmically transformed to standardize the comparison scale. The difference between the logarithmic copy number at the current time point and a previous fixed time point was calculated. This difference was divided by the time interval to obtain the copy number increase rate. This rate reflects the dynamic intensity of the target pathogen's replication within the host. The absolute value of the recent decline slope was compared with the copy number increase rate. Since both are logarithmically scaled, their ratio eliminates the influence of dimensions and more accurately reflects the relative relationship between antibody inhibition and target pathogen proliferation. Taking the logarithm of this ratio further compressed the numerical range and enhanced the sensitivity to changing trends, ultimately generating an antibody inhibition trend factor. An increase in the positive value of this factor indicates a relative weakening of antibody inhibition and an increase in target pathogen proliferation. The previously calculated immune escape risk coefficient needs to be standardized. It is mapped to a fixed range using a mathematical function to eliminate the skewness of its original numerical distribution and enhance the comparability between different samples. This function has an S-shaped curve characteristic, which can smoothly compress the input value to between zero and one.
[0087] The antibody inhibition trend factor is multiplied by a pre-defined weighting coefficient. This weighting coefficient is adjusted based on the general evolutionary rate of the target pathogen and the response characteristics of the immune system to regulate the contribution of the trend factor in the final index. The standardized risk value is then added to the weighted antibody inhibition trend factor to generate the target pathogen's adaptive evolutionary index. This index integrates risk assessment based on historical data and trend analysis based on recent dynamics, providing a comprehensive indicator to quantify the current adaptability and evolutionary potential of the target pathogen.
[0088] The standardization process in this index employs the following transformations:
[0089]
[0090] in: Represents the standardized risk value. Represents the original risk coefficient of immune escape. It is a scaling factor that adjusts the steepness of the transform curve. It is the offset parameter that determines the position of the curve center. It is the base of the natural logarithm.
[0091] When analyzing a specific sample, hourly concentration monitoring data of immune antibody proteins over the most recent 72 hours were extracted. First, the natural logarithm of each concentration value was calculated. The logarithmic concentration data was then smoothed using a Savitzky-Golay filter (7-hour window, polynomial order 2). The first derivative (numerical difference) of the smoothed curve was calculated. The period with the lowest consecutive negative values in the derivative sequence was identified. A segment with a duration of at least 5 hours and the smallest average derivative was selected, and the arithmetic mean of all derivative values within this segment was calculated as the recent downward slope. (Unit: per hour). Analyze the DNA copy number data of the target pathogen bacteria. Read the copy number detection values at the current time point and 24 hours ago, and take the natural logarithm for each. and Calculate the rate of increase in copy number. (Unit is also per hour). Calculate the antibody inhibition trend factor. .
[0092] Known risk factor for immune escape (For example, a value of 2.1) Standardize. Set parameters. , .calculate Antibody inhibition trend factor (For example, a value of 0.8) multiplied by the preset evolution acceleration weight coefficient ,get Finally, the adaptive evolution index of the target pathogen bacteria was calculated. This process is executed automatically by a computer program. The inputs are time series data, copy number data, and preset parameters. The outputs are antibody inhibition trend factor and the adaptive evolution index of the target pathogen bacteria.
[0093] Example 4: A continuous time-point recording mechanism was established for dynamic monitoring of the adaptive evolution index of target pathogenic bacteria. The current index value was calculated at each detection time and directly compared with the index value at the previous detection time to calculate the numerical difference. This difference reflects the direction and magnitude of change in the adaptiveness of the target pathogenic bacteria within adjacent detection intervals.
[0094] When the current index difference exceeds the threshold, it indicates that the target pathogen has recently exhibited strong adaptive evolution, significantly increasing its potential risk of immune escape. At this point, the system automatically triggers a mechanism to lower the antibody affinity detection threshold. The downward adjustment is based on a preset step size parameter, typically set as a fixed percentage of the current threshold level to ensure gradual and controllable adjustment. Simultaneously, a negative threshold is set. When the index difference remains below this negative threshold for multiple consecutive detection cycles, it indicates that the adaptive evolutionary pressure on the target pathogen has eased, and the risk of immune escape has relatively decreased. At this point, the system automatically triggers a mechanism to raise the antibody affinity detection threshold. The raising operation is performed at a fixed percentage of the current threshold to gradually restore it to the baseline level. Under the adjusted antibody affinity detection threshold, the immune antibody protein expression database is scanned and screened. The screening condition is that the affinity parameter value of each antibody record must be greater than or equal to the currently effective detection threshold. Affinity parameters are typically derived from experimental data, such as the value of the dissociation constant measured by surface plasmon resonance (SPR) after negative logarithmic transformation. Identifiers are extracted from all antibody records that meet the threshold criteria. Identifiers are typically unique identifiers of the antibody-encoding genes, such as specific hash values of the heavy and light chain variable region gene sequences. All extracted identifiers are aggregated to form a non-repeating set, which represents the set of antibodies that still maintain effective neutralizing potential under the current immune stress environment.
[0095] Table 1: Record of dynamic adjustment of antibody affinity detection threshold.
[0096] Detection time point Current adaptive evolution index Index difference value Affinity threshold before adjustment Adjustment operation Affinity threshold after adjustment D1-08:00 1.25 - 6.80 Initial state 6.80 D2-08:00 1.52 +0.27 6.80 Reduce 5% 6.46 D3-08:00 1.78 +0.26 6.46 Reduce 5% 6.14 D4-08:00 1.65 -0.13 6.14 No operation 6.14 D5-08:00 1.59 -0.06 6.14 No operation 6.14 D6-08:00 1.48 -0.11 6.14 Raise 3% 6.32
[0097] Referring to Table 1, in a continuous monitoring case, the system calculates the adaptive evolution index of the target pathogen bacteria at fixed time points each day. The initial antibody affinity detection threshold is set at 6.80 (corresponding to a dissociation constant KD value of approximately 10^(-6.8)M). On the second day, the calculated index difference is positive 0.27, exceeding the set positive evolutionary gradient threshold of 0.15. Based on preset rules, the system lowers the affinity detection threshold by 5%, resulting in a new threshold of 6.46. On the third day, the index difference again shows a positive value and exceeds the threshold, so the threshold is further lowered to 6.14. On the fourth and fifth days, although the index difference is negative, it does not meet the condition of being below the negative threshold (set at -0.1) twice consecutively, so the threshold remains unchanged. Until the sixth day, the index difference is negative again, marking the second consecutive time it has been below the negative threshold. The system triggers the threshold adjustment mechanism, raising the threshold by 3%, adjusting it to 6.32.
[0098] After each threshold adjustment, the system immediately scans the immune antibody database. Assuming that after the threshold is lowered to 6.14, antibodies with affinity parameters between 6.14 and 6.46 in the original database records are reinstated as valid antibodies. The system extracts the corresponding genetic identifiers for these antibodies (e.g., the MD5 hash of antibody A: VH3-2301-VK1-3901; the MD5 hash of antibody B: VH1-6901-VL2-1401) and merges them with the identifiers of the existing high-affinity antibodies, generating an updated set of valid neutralizing antibodies. This set dynamically reflects all antibody clones expected to maintain neutralizing potency under the current evolutionary pressure of the target pathogen bacteria. The entire process is executed automatically, ensuring the consistency and timeliness of monitoring, decision-making, and output.
[0099] Example 5: See Figure 5 This process integrates the boundary information of previously identified high-frequency mutation regions in the target pathogen's DNA with the target pathogen's genome annotation information. The genome annotation information precisely identifies the start and end positions of genes encoding various functional proteins (especially antigen proteins) on the DNA sequence. Through bioinformatics alignment, the physical coordinates of each high-frequency mutation region are mapped to its corresponding coding region. If a high-frequency mutation region falls entirely or partially within the coding region of an antigen protein, that region is marked as being located within the antigen protein coding region. This mapping process identifies which amino acid residues might be affected by DNA-level mutations, potentially altering the epitope structure of the antigen protein. Antigen-binding site information corresponding to each antibody record is extracted from the effective neutralizing antibody set. This information, derived from structural biology databases or experimental assays, details the position coordinates of specific amino acid residues on the antigen protein sequence that the antibody's complementarity-determining region contacts when interacting with the target pathogen's antigen protein. These coordinates are typically represented by residue numbers or positional identifiers in a three-dimensional structure.
[0100] The coordinates of each antibody's binding site were compared one by one with the boundaries of all high-frequency mutated regions mapped to the antigen's protein coding region. This alignment aimed to identify whether the antibody binding site spatially or sequentially overlapped with any mutated region. The overlap criterion was that if the amino acid residue number targeted by an antibody binding site fell within the start and end residue numbers of a high-frequency mutated region, the binding site was considered covered by that mutated region. This means that the epitope originally recognized by the antibody may have changed due to mutations in that region, potentially affecting its binding ability and neutralizing potency. Based on the alignment results, a subset was selected from the set of effective neutralizing antibodies. This subset contained all antibodies with at least one binding site overlapping with any key mutated region. These antibodies were considered candidates most directly affected by the current evolutionary mutations of the target pathogen, and their functional effectiveness needed to be prioritized for evaluation.
[0101] For each antibody in this subset, its binding free energy with the current variant antigen protein (integrating identified mutations) is further predicted using computational simulation methods. The prediction process typically employs molecular mechanics combined with Poisson-Boltzmann surface area calculations or similar methods. This calculation is based on three-dimensional structural models of the antibody and antigen, estimating the thermodynamic stability of their binding through an energy function. The binding free energy value is negative; the more negative the value (the larger the absolute value), the stronger and more stable the binding. The subset of neutralizing antibodies is output, along with the corresponding predicted binding free energy value for each antibody record in the subset. This output provides quantitative predictive information on which neutralizing antibodies are likely to remain effective against currently prevalent target pathogen bacterial variants, and which may have diminished or lost their effectiveness.
[0102] In one analysis, three high-frequency mutation regions were identified, whose boundaries, after mapping, corresponded to amino acid residues 120-135, 255-270, and 310-330 of the antigen protein VP1, respectively. The effective neutralizing antibody set comprised 15 different antibody clones. By querying structural databases, detailed binding site information for eight of these antibodies to the standard antigen protein was obtained. For example, the binding site of antibody mAb-12 was recorded as targeting residues 128, 129, 131, and 265. Comparing this to the mutation regions: residues 128, 129, and 131 fall within the first mutation region (120-135); residue 265 falls within the second mutation region (255-270). Therefore, the binding site of antibody mAb-12 was determined to be covered by a mutation region and included in the subset of antibodies targeting key mutations.
[0103] Binding free energy prediction was performed on antibodies in this subset. Taking antibody mAb-12 as an example, its binding free energy with a variant antigen protein model integrating all recognized mutations was calculated. The results showed that its predicted binding free energy value differed specifically from the reference value for the original strain. The final output subset list included antibody identifiers such as mAb-12, mAb-07, and mAb-15, along with their respective calculated predicted values. The entire process was executed through an automated analysis workflow, combining sequence variation information with antibody function predictions, providing detailed data support for assessing immune escape risk and guiding antibody selection.
[0104] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus.
[0105] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.
Claims
1. A method for detecting immune antibodies against target pathogenic bacteria based on DNA detection, characterized in that, Includes the following steps: Obtain the target pathogen bacterial DNA sequence data and immune antibody protein expression level data of the biological sample to be tested; Mutation site scanning was performed on the DNA sequence data of the target pathogenic bacteria to identify the boundaries of high-frequency mutation regions; Based on the distribution density and base substitution frequency of the high-frequency mutation region boundary, the sequence mutation degree of the target pathogenic bacterial DNA sequence is calculated; The periodic fluctuation characteristics of the immune antibody protein expression data are extracted simultaneously, and an antibody response intensity index is generated based on the rate of change of fluctuation amplitude and phase shift. By analyzing the historical correlation between the sequence mutation rate and the antibody response intensity index, and combining the synergy of their changing trends, the immune escape risk coefficient at the current moment is determined. The antibody inhibition trend factor was constructed by detecting the recent decline slope of the expression level of the immune antibody protein and the increase rate of the copy number of the target pathogen bacterial DNA sequence. By combining the immune escape risk coefficient and the antibody inhibition trend factor, an adaptive evolution index of the target pathogen bacteria is generated; The antibody affinity detection threshold is dynamically adjusted based on the gradient of the adaptive evolution index of the target pathogenic bacteria and the preset detection sensitivity threshold. Based on the adjusted antibody affinity detection threshold, the expression data of the immune antibody proteins are screened for affinity, and an effective set of neutralizing antibodies is output.
2. The method for detecting target pathogenic bacteria immune antibodies based on DNA detection as described in claim 1, characterized in that, The identification of the high-frequency mutation region boundary further includes: Divide the target pathogenic bacterial DNA sequence data into a preset length sliding window and count the frequency of non-conserved bases in each window; Regions where the frequency of three or more consecutive windows exceeds the mutation threshold are marked as candidate mutation regions; Calculate the boundary entropy value of the candidate mutation region, merge adjacent regions whose boundary entropy value difference is less than the tolerance factor, and output the boundary of the merged high-frequency mutation region.
3. The method for detecting target pathogenic bacteria immune antibodies based on DNA detection as described in claim 2, characterized in that, The calculation of the sequence mutation degree of the target pathogenic bacterial DNA sequence further includes: Extract the length of the boundary and the coordinates of the center position of each high-frequency mutation region; The deviation of the base substitution type from the standard reference sequence in each region was statistically analyzed; The weighted product of the length, the center position dispersion, and the deviation is used as the local mutation contribution value of each region; The sequence mutation degree is generated by normalizing and summing all local mutation contribution values.
4. The method for detecting target pathogenic bacteria immune antibodies based on DNA detection as described in claim 1, characterized in that, The antibody response strength indicator further includes: The expression data of the immune antibody proteins were segmented in the time domain, and the peak and trough positions in each time period were extracted. Calculate the average amplitude difference between adjacent peaks and troughs as the rate of change of fluctuation amplitude; The absolute value of the phase deviation between the actual peak position and the theoretical immune cycle is measured to generate the phase offset. The inverse of the rate of change of the fluctuation amplitude is linearly combined with the phase offset to output an antibody response intensity index.
5. The method for detecting target pathogenic bacteria immune antibodies based on DNA detection as described in claim 4, characterized in that, The determination of the immune escape risk coefficient at the current moment further includes: Establish a sliding correlation coefficient matrix between the sequence mutation rate and the antibody response intensity index within a historical time window; Calculate the eigenvalue variance of the correlation coefficient matrix as the historical correlation decay rate; Extract the first-order difference of sequence mutation degree and the first-order difference of antibody response intensity index from consecutive time points before the current time. The sum of the products of two differences at the same time point is used as the degree of synergy in the trend of change. The geometric mean of the historical correlation decay and the correlation of the change trend is used as the immune escape risk coefficient.
6. The method for detecting target pathogenic bacteria immune antibodies based on DNA detection as described in claim 1, characterized in that, The constructed antibody inhibition trend factor further includes: Extract the expression level data of immune antibody proteins at the most recent preset time and fit the first derivative of its logarithmic change curve; The mean of the smallest negative segment of the first derivative is taken as the recent descent slope; The logarithmic increment of the target pathogen bacterial DNA sequence copy number per unit time is calculated simultaneously and used as the copy number increase rate. The antibody inhibition trend factor is generated by taking the logarithm of the ratio of the absolute value of the recent decline slope to the copy number increase rate.
7. The method for detecting target pathogenic bacteria immune antibodies based on DNA detection as described in claim 6, characterized in that, The generated target pathogenic bacteria adaptive evolution index further includes: The immune escape risk coefficient is transformed by the Sigmoid function to obtain a standardized risk value; Multiply the antibody inhibition trend factor by a preset evolution acceleration weighting coefficient; The standardized risk value is added to the weighted antibody inhibition trend factor to output the adaptive evolution index of the target pathogen bacteria.
8. The method for detecting target pathogenic bacteria immune antibodies based on DNA detection as described in claim 7, characterized in that, The dynamic adjustment of the antibody affinity detection threshold further includes: Calculate the difference in the adaptive evolution index of the target pathogen bacteria between the current time and the previous detection time; When the difference exceeds the evolutionary gradient threshold, the antibody affinity detection threshold is reduced by a preset step size. When the difference is lower than the negative gradient threshold twice in a row, the antibody affinity detection threshold is increased by a fixed ratio.
9. The method for detecting target pathogenic bacteria immune antibodies based on DNA detection as described in claim 8, characterized in that, The output set of effective neutralizing antibodies further includes: Under the adjusted antibody affinity detection threshold, records in the immune antibody protein expression data whose affinity parameters exceed the threshold are screened. The antibody-encoding gene identifiers corresponding to the records are aggregated to generate an effective set of neutralizing antibodies.
10. The method for detecting target pathogenic bacteria immune antibodies based on DNA detection as described in claim 9, characterized in that, Also includes: Map the boundaries of the high-frequency mutation regions to the target pathogen bacterial antigen protein coding regions; Cross-align the binding site coordinates of the effective neutralizing antibody set with the boundary of the mutation region; Output a subset of neutralizing antibodies covering key mutation sites and their predicted binding free energy.