Method and system for mining therapeutic targets of pathogenic bacteria and coronavirus infectious diseases

By extracting protein and nucleic acid expression reads from bacterial and coronavirus infection samples, locating molecular sites, and constructing a target interaction diagram, the limitations of target selection and low accuracy in existing technologies are solved. This enables target identification and screening in complex infection environments, improving the accuracy and adaptability of therapeutic targets.

CN121459939APending Publication Date: 2026-02-03BEIJING INST OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511610873.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-05
Publication Date
2026-02-03

AI Technical Summary

Technical Problem

Existing technologies face challenges in identifying therapeutic targets for pathogenic bacteria and coronavirus infections, including limited target selection, low accuracy, difficulty in capturing the correspondence between sequence expression and tissue structure, and a lack of reliable interaction information comparison support in complex infection environments. This results in unstable target responses and makes it difficult to support the universal applicability of downstream drug targeted development or vaccine construction.

Method used

By extracting protein and nucleic acid expression reads from bacterial and coronavirus infection samples, locating molecular sites, removing redundant signals, and grouping them with sample information, a set of infection molecular features is generated. Based on expression intensity and stability grading, repetitive and consistent fragments are identified, a target interaction diagram is constructed, the overlap and response patterns between targets are analyzed, shared binding regions in differentiated infection types are identified, and core fragments with stable responses are extracted.

Benefits of technology

It improves the spatial resolution of feature recognition, enhances the stability and discriminative power of fragment screening, improves the accuracy and adaptability of interaction screening, and forms a set of therapeutic targets that cover multiple infection states and have consistent responses.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121459939A_ABST
    Figure CN121459939A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of anti-infection mining, in particular to a pathogenic bacterium and coronavirus infection disease treatment target mining method and system.The method comprises the following steps of extracting expression read segments and positioning molecular sites, grouping, comparing and screening stable sequences, analyzing the fragment overlapping and co-occurrence relation, verifying a binding area and eliminating accidental interaction, and identifying the shared high-frequency binding region and the stable response fragment to obtain an infection treatment target mining result. According to the method, the spatial resolution of feature recognition is improved by expressing accurate positioning of read segments in a dyeing area and a sequencing sequence, the stability and the distinction degree of segment screening are enhanced by multi-dimensional screening of signal intensity and peak shape features, and the connectivity relation between the sequences is constructed by structural integration of segment co-occurrence frequency and interval distribution; by combining the combination position statistics under the multi-infection condition, the accuracy and adaptability of interaction screening are improved, and finally a therapeutic target set which can cover the multi-infection state and is consistent in response is formed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of anti-infection discovery technology, and in particular to methods and systems for discovering therapeutic targets for pathogenic bacteria and coronavirus infections. Background Technology

[0002] The field of anti-infective drug discovery technology primarily involves using scientific methods to study the interactions between pathogenic microorganisms (such as bacteria, viruses, and fungi) and host organisms, thereby discovering and developing targeted treatment methods or drugs. The core content of this field includes research on the pathogenic mechanisms of pathogens, screening for drug targets, development of anti-infective drugs, and design of novel vaccines. The technology encompasses the interdisciplinary application of genomics, proteomics, metabolomics, and other disciplines, providing innovative solutions for anti-infective therapy by accurately identifying key molecular targets in pathogens and hosts. Traditional methods for discovering therapeutic targets for pathogenic bacteria and coronavirus infections involve analyzing the molecular mechanisms of interactions between pathogenic microorganisms (such as pathogens and coronaviruses) and host organisms to discover potential therapeutic targets, thereby developing anti-infective drugs or treatments. Traditional target discovery methods typically employ genomics and proteomics, using sequencing, data comparison, and molecular biology experiments to identify key genes and proteins in pathogens and analyze their functions and roles in the infection process. However, these methods often face limitations such as limited target selection and low accuracy. Therefore, the technical content proposed in the patent mainly provides new target screening ideas and methods by combining innovative target discovery strategies with the characteristics of pathogens and coronaviruses.

[0003] Existing technologies rely on single-dimensional gene or protein sequencing data to construct target identification pathways. In actual samples, it is difficult to capture the correspondence between sequence expression and tissue structure, resulting in a disconnect between data sources and biological background. This leads to one-sided feature extraction, neglecting the co-occurrence and structural association between expressed signals during the signal screening stage. It is difficult to determine the functional connectivity and intervention pathways between fragments. In complex infection environments, there is no integration mechanism for cross-infection signals. When facing complex infections or synergistic changes of multiple pathogens, there is a lack of reliable interaction information comparison support, causing the selected targets to respond unstablely in dynamic pathological states. This makes it difficult to support the universal applicability of downstream drug targeted development or vaccine construction. Summary of the Invention

[0004] To achieve the above objectives, the present invention employs the following technical solution: a method for identifying therapeutic targets for pathogenic bacteria and coronavirus infections, comprising the following steps: S1: Extract protein and nucleic acid expression reads from bacterial and coronavirus infection samples, locate molecular sites, remove redundant signals, and group them in combination with sample information to generate a set of infection molecular features; S2: Based on the set of infectious molecular features, the feature set is graded by expression intensity and stability, repeating and uniformly distributed fragments are identified, abnormal signals are excluded, and sequences appearing in multiple batches are extracted to form a stable set of infectious signals. S3: Call the stable infection signal set, compare the boundary overlap and co-occurrence frequency of stable signal segments, analyze the connection relationship, merge adjacent segments with regular co-occurrence, and generate an infection-associated structural sequence group; S4: Compare the infection-associated structural sequence group with the public database, analyze the repetition of the binding regions, exclude occasional interactions, retain repeated binding pairs, and construct a target interaction diagram. S5: Based on the target interaction diagram, analyze the overlap and response patterns between targets, identify the shared binding regions in differentiated infection types, extract the core fragments of stable responses, and obtain the results of infection treatment target mining.

[0005] As a further aspect of the present invention, the set of infectious molecular features includes expression intensity information, staining site markers, sequencing sequence localization, and sample source labels; the set of stable infection signals includes high-frequency fragment sequences, distribution consistency identifiers, cross-batch stability parameters, and abnormal signal filtering records; the cluster of infection-associated structural sequences includes boundary overlapping regions, co-occurrence frequency matrices, fragment spacing parameters, and continuous structural units; the target interaction diagram includes a binding position lookup table, repeated contact statistics, cross-conditional interaction pairs, and interaction screening results; and the results of infection treatment target mining include shared binding regions, stable response fragments, dual infection marker sites, and high-frequency interacting targets.

[0006] As a further aspect of the present invention, the merging of adjacent segments with regular co-occurrence will merge stable signal segments that frequently appear adjacently and simultaneously in the sample into a continuous structural unit based on boundary overlap and co-occurrence patterns.

[0007] As a further aspect of the present invention, the core fragment of the stable response refers to the key functional sequence that exhibits expression and consistent response patterns under differentiated infection conditions.

[0008] As a further aspect of the present invention, the specific steps of S1 are as follows: S101: Based on the expression reads in bacterial and coronavirus infection samples, protein and nucleic acid sequences are extracted, and the positions of molecular sites in the reads in the stained regions and sequencing sequences are retrieved to establish a molecular site localization matrix; S102: Call the molecular site localization matrix, perform joint sorting and remove duplicate and low-frequency sites based on the occurrence frequency and peak width of the molecular sites, and obtain a set of non-redundant molecular sites; S103: Based on the non-redundant molecular site set, the samples are grouped and summarized in combination with the sample source identifier. The sites in the same group are aggregated and duplicate signals are removed to generate an infected molecular feature set.

[0009] As a further aspect of the present invention, the specific steps of S2 are as follows: S201: Call the sequence entries in the infected molecular feature set, and perform value sorting and fluctuation range comparison in the same type of samples according to the expression intensity and signal stability parameters of each sequence, and filter the sequence entries that simultaneously meet the expression and fluctuation range requirements to generate a set of expression stable entries; S202: Based on the expressed stable entry set, analyze the distribution information of the sequence in the differential source samples, calculate the source repetition rate between sequences, identify entries with a repetition rate lower than the repetition threshold, mark and remove all abnormal signal entries, and obtain the source repetition entry set; S203: Based on the sequence content in the source repeating entry set, retrieve the occurrence records in multiple batches of samples, filter the sequence entries that can continuously appear in the batches of samples, establish the corresponding index sequence, and generate a stable infection signal set.

[0010] As a further aspect of the present invention, the specific steps of S3 are as follows: S301: Call all segment sequences in the stable infection signal set, extract the start and end boundary position information of each group of segments, count the overlapping intervals between the boundaries of the differentiated segments, and accumulate the co-occurrence times of segments with overlapping intervals to generate a segment overlap co-occurrence matrix. S302: Based on the fragment pair information recorded in the fragment overlap co-occurrence matrix, retrieve the boundary distance value of each fragment pair and compare it with a fixed interval distance threshold. If the distance is less than the fixed interval distance threshold and the number of co-occurrences is greater than the co-occurrence frequency threshold, then mark the fragment pair as connectable and obtain the fragment connection index set. S303: Based on all connectable fragment pairs marked in the fragment connection index set, aggregate adjacent fragment sequences in index order, remove fragments that do not meet the connection conditions, construct a list of continuous fragment combination sequences, and generate an infection-associated structure sequence group.

[0011] As a further aspect of the present invention, the specific steps of S4 are as follows: S401: Call all fragment combination information in the infection-associated structural sequence group, combine with the fragment combination position index recorded in the public interaction data, extract the start and end position information of the sequence within the fragment combination and perform a pairwise cross-alignment with the interaction position, calculate the alignment hit value and generate the sequence interaction hit matrix. S402: Based on the number of combinations recorded in the sequence interaction hit matrix, retrieve the cumulative hit records of each pair of sequence combination regions, and remove combination pairs with fewer hits than the interaction statistical benchmark value, retaining only the combination information that meets the record frequency requirement, to obtain a set of multi-condition co-occurrence combination pairs; S403: Based on the connection relationship of all binding pairs in the multi-condition co-occurrence binding pair set, construct an undirected graph structure and set the sequence index as a node and the binding pair as an edge, remove independent nodes without connection pairs, and generate a target interaction graph.

[0012] As a further aspect of the present invention, the specific steps of S5 are as follows: S501: Call the binding pair position index recorded in the target interaction diagram, extract the infection target position segment in the pathogen and coronavirus association data, perform coordinate overlap comparison for each group of infection types, construct the corresponding overlap matrix according to the intersection ratio, and generate the target position overlap distribution table. S502: Based on the overlap ratio values ​​in the target location overlap distribution table, select the binding regions that exist in both differential infection types and have an overlap ratio greater than the target intersection benchmark threshold, and remove regions with an overlap ratio lower than the target intersection benchmark threshold to obtain a set of shared high-frequency binding regions. S503: Based on the sequence segments identified by the shared high-frequency combined region, the response frequency in the dual infection signal records is counted, and segments with stable response frequencies in all infection samples are selected to construct a list of locatable target segments and obtain the results of infection treatment target mining.

[0013] A system for identifying therapeutic targets for pathogenic bacteria and coronavirus infections, including: The infection feature extraction module is used to achieve S1: extracting protein and nucleic acid expression reads from bacterial and coronavirus infection samples, locating molecular sites, removing redundant signals, and grouping them in combination with sample information to generate a set of infection molecular features; The stable signal screening module is used to implement S2: based on the set of infectious molecular features, the feature set is graded by expression intensity and stability, repeating and uniformly distributed fragments are identified, abnormal signals are excluded, and sequences appearing in multiple batches are extracted to form a stable set of infectious signals; The structural sequence construction module is used to implement S3: call the stable infection signal set, compare the boundary overlap and co-occurrence frequency of stable signal segments, analyze the connection relationship, merge adjacent segments with regular co-occurrence, and generate an infection-associated structural sequence group; The target interaction screening module is used to implement S4: compare the infection-related structural sequence group with the public database, analyze the repetition of the binding region, exclude occasional interactions, retain repeated binding pairs, and construct a target interaction connection graph. The core target identification module is used to implement S5: based on the target interaction diagram, analyze the overlap and response patterns between targets, identify the shared binding regions in differentiated infection types, extract the core fragments of stable responses, and obtain the results of infection treatment target mining.

[0014] Compared with the prior art, the advantages and positive effects of the present invention are as follows: In this invention, the spatial resolution of feature recognition is improved by accurately locating the expressed reads in the stained region and the sequencing sequence. The stability and discriminativeness of fragment screening are enhanced by multidimensional screening of signal intensity and peak shape features. The connectivity between sequences is constructed by structural integration of fragment co-occurrence frequency and spacing distribution. The accuracy and adaptability of interaction screening are improved by combining binding position statistics under multiple infection conditions. Finally, a set of therapeutic targets that can cover multiple infection states and have consistent responses is formed. Attached Figure Description

[0015] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0016] Figure 1 This is a schematic diagram of the steps of the present invention; Figure 2 This is a detailed schematic diagram of S1 of the present invention; Figure 3 This is a detailed schematic diagram of S2 of the present invention; Figure 4 This is a detailed schematic diagram of S3 of the present invention; Figure 5 This is a detailed schematic diagram of S4 of the present invention; Figure 6 This is a detailed schematic diagram of S5 of the present invention; Figure 7 This is a system module diagram of the present invention. Detailed Implementation

[0017] The technical solution of the present invention will now be described with reference to the accompanying drawings.

[0018] In embodiments of the present invention, words such as "exemplarily," "for example," etc., are used to indicate that something is an example, illustration, or description. Any embodiment or design described as "exemplary" in the present invention should not be construed as being more preferred or advantageous than other embodiments or designs. Specifically, the use of the word "exemplary" is intended to present the concept in a concrete manner. Furthermore, in embodiments of the present invention, the meaning expressed by "and / or" can be both, or either one.

[0019] In the embodiments of this invention, the terms "image" and "picture" may sometimes be used interchangeably. It should be noted that, without emphasizing the distinction between them, they convey the same meaning. Similarly, the terms "of," "corresponding (relevant)," and "corresponding" may sometimes be used interchangeably. It should be noted that, without emphasizing the distinction between them, they convey the same meaning.

[0020] In this embodiment of the invention, sometimes a subscript such as W1 may be written in a non-subscript form such as W1. When the difference is not emphasized, the meaning they express is the same.

[0021] To make the technical problems, technical solutions and advantages of the present invention clearer, a detailed description will be given below in conjunction with the accompanying drawings and specific embodiments.

[0022] Please see Figure 1 This invention provides a method for identifying therapeutic targets for pathogenic bacteria and coronavirus infections, comprising the following steps: S1: Extract expression reads of proteins and nucleic acids from bacterial and coronavirus infection samples, record the positions of molecular sites in the stained regions and sequencing sequences, compare them according to the frequency of occurrence and peak width and remove redundant signals, combine the sample source information to group and summarize, and generate a set of infection molecular features after integrating repeated signals; S2: Call the set of infectious molecular features, perform hierarchical comparison based on expression intensity and signal stability, identify items that appear multiple times and have the same distribution range, compare the repetition rate between sources and exclude abnormal signals, select fragment sequences that appear stably in multiple batches of data, and generate a stable set of infectious signals; S3: Call the stable infection signal set, compare the boundary overlap range of the segments and count the number of co-occurrences, calculate the distance between adjacent segments and determine the connection relationship, merge segments with stable intervals and frequent co-occurrences into a continuous segment combination to form an infection-associated structural sequence group; S4: Call the infection-associated structural sequence group, retrieve public interaction data, perform cross-alignment of the binding positions between fragments, filter by the number of repeated records in the contact area, delete occasional interaction signals and retain binding pairs that appear under multiple infection conditions, and generate a target interaction relationship map. S5: Call the target interaction diagram to perform overlap analysis on the target locations in the data on the association between pathogens and coronaviruses, screen the high-frequency binding regions shared in the differentiated infection types, extract the core fragments that maintain a stable response in dual infection signals, and obtain the results of infection treatment target mining.

[0023] The set of infection molecular features includes expression intensity information, staining site markers, sequencing sequence localization, and sample source labels. The set of stable infection signals includes high-frequency fragment sequences, distribution consistency markers, cross-batch stability parameters, and abnormal signal filtering records. The infection-associated structural sequence cluster includes boundary overlapping regions, co-occurrence frequency matrix, fragment spacing parameters, and continuous structural units. The target interaction relationship map includes a binding position comparison table, repeated contact statistics, cross-condition interaction pairs, and interaction screening results. The results of infection treatment target mining include shared binding regions, stable response fragments, dual infection marker sites, and high-frequency interacting targets.

[0024] Please see Figure 2 The specific steps of S1 are as follows: S101: Based on the expression reads in bacterial and coronavirus infection samples, protein and nucleic acid sequences are extracted, and the positions of molecular sites in the reads in the stained regions and sequencing sequences are retrieved to establish a molecular site localization matrix; First, the base sequence information should be extracted from the original sequencing file, and its quality value should be statistically analyzed. Base segments with a quality value greater than 30 in each sequencing read should be selected for retention. This quality threshold is based on the general understanding in the literature that bases with a quality value of Q30 or higher (representing an error rate of 0.1%) are high-quality bases. Then, adapter sequence identification and removal should be performed, and the reads should be spliced ​​into the longest continuous usable fragment. The high-quality reads should be aligned one by one with the reference genome using an alignment tool. During the alignment process, the chromosome number and start and end positions corresponding to each read should be extracted, and the region it falls into should be marked as the staining region number. For example, if read A aligns to chr7 between 1521300 and 1521580, the staining region is the chr7_1521 region. Next, the reference base sequence should be extracted from this staining region, and the presence of protein-coding information in this region should be determined. Whether to continue translation, if the region is a coding region, the base sequence is divided into triplets according to the start codon, and translated segment by segment into the corresponding amino acids to form a protein sequence. During this process, the specific base position index of each mutation site within the read is recorded. For each aligned read, the mutation site is identified by comparing whether the corresponding base position in the reference genome is the same. For example, if the base in the 43rd read is C and the reference genome is T, it is identified as a mutation type of T>C. After aligning and locating all reads, a molecular site localization matrix containing all variants of all samples is constructed. The matrix is ​​arranged with sample number as the row and site number as the column. The element value is 0 or 1, indicating whether the sample has a corresponding variant or expression at that site. For example, if sample_001 has a mutation site at chr7_1521_43, then that position in the matrix is ​​assigned a value of 1.

[0025] S102: Call the molecular site localization matrix, perform joint sorting and remove duplicate and low-frequency sites based on the occurrence frequency and peak width of the molecular sites, and obtain a set of non-redundant molecular sites; First, statistics are performed on each column of the matrix to obtain the frequency of each locus in all samples. Frequency is the number of elements with a value of 1 in that column. After determining the frequency, the peak width of the locus in the alignment data is further extracted. By analyzing the base distribution in the coverage variation region, it is determined whether the signal peak of the locus is continuous and representative. The peak width is calculated by setting a baseline coverage threshold. For example, the average coverage depth plus the standard deviation is set as the peak bottom threshold. That is, if the average coverage of a certain segment is 80 and the standard deviation is 10, the peak bottom threshold is set to 90. The range where the coverage of consecutive base positions in this segment exceeds 90 is defined as the peak width. The peak width is the number of bases in this continuous range. After obtaining the frequency and peak width, a joint sorting operation is performed, that is, sorting first by the descending order of the locus frequency, and then sorting the loci with the same frequency by... The peak widths are sorted in descending order. The priority of frequency sorting depends on the threshold setting, which is generally set to 10%. That is, if the total number of samples is 200, the sites with a frequency below 20 are considered low-frequency sites and will be removed. This setting comes from the need to control the proportion of noise variation in biological samples. For sites with redundancy, such as two sites less than 10 bp apart and with the same mutation type, it is determined whether they are on the same expression read and the site with the higher frequency and larger peak width is retained as the representative, while duplicate sites are removed. For example, sites A and B are only 5 bp apart and both are represented as C>T mutations in the sample, but A has a frequency of 60 and a peak width of 35, while B has a frequency of 58 and a peak width of 30. In this case, A is retained and B is removed, and finally a non-redundant molecular site set composed of high-frequency, high-peak-width, and non-redundant representative molecular sites is obtained.

[0026] S103: Based on the non-redundant molecular site set, the sample source identifier is combined with the grouping and summarizing, the sites of the same group are aggregated and duplicate signals are removed to generate a set of infectious molecular features. First, all samples are grouped according to the source labels marked in the sample information. For example, the pathogen infection group, coronavirus infection group, and healthy control group are treated as three separate groups. For each group, the row vector of the expression matrix of all samples belonging to that group at non-redundant sites is extracted. A logical OR operation is performed on the values ​​of all samples within the group at the same site; that is, if any sample is 1 at that site, then the group is also marked as 1 at that site. This method automatically removes duplicate expression signals within the group when constructing the feature vector of each group. In the example, if the coronavirus group has 5 samples, their expression at site chr5_2030_89... If the expression is [0, 1, 0, 0, 1], then the site is marked as 1 in the group expression. Then, the number of sites expressing 1 in each group sample is counted, and according to the expression difference of the site in different groups, the subsequent cross-analysis is prepared. Finally, a binary vector set is formed, where each vector corresponds to a sample group and the expression status of all non-redundant molecular sites. Combining the example, assuming there are 1500 non-redundant sites, each group gets a 1500-dimensional feature vector, where the elements are 0 or 1, indicating whether the group has an expression signal at that site. This feature set is the infectious molecule feature set.

[0027] Please see Figure 3 The specific steps of S2 are as follows: S201: Call the sequence entries in the infected molecular feature set, and perform value sorting and fluctuation range comparison in the same type of samples according to the expression intensity and signal stability parameters of each sequence. Select the sequence entries that meet both the expression and fluctuation range requirements and generate a set of expression stable entries. First, the expression intensity and corresponding signal stability parameter values ​​of each sequence in various samples of the same type are read. The expression intensity is represented by the average coverage depth of the sequence at the corresponding non-redundant site in the sample. If a sequence has a coverage depth of 110 in sample A, 95 in sample B, and 102 in sample C, then the average expression intensity of the sequence is (110 + 95 + 102) ÷ 3, which is 102.33. The signal stability parameter is obtained by the ratio of the standard deviation to the mean of the sample expression intensity, indicating the degree of expression fluctuation of the sequence within the sample group. The lower the value, the more stable the expression. Then, all sequences are divided according to the sample category, for example, bacterial infection samples are in one group, and coronavirus infection samples are in another group. For the sequence entries within each sample group, they are first sorted in descending order according to the average expression intensity value. The top 20% of the sorted results are extracted as strong expression candidates. The threshold for this set is set based on observations of the expression value distribution histogram of actual samples. Sequences representing epistatic expression characteristics are selected as the starting point, and then signal volatility screening is performed on these candidate sequences. The expression variation amplitude within the group is calculated for each sequence, i.e., the ratio of the standard deviation to the mean is used as the volatility value. A volatility range threshold of 0.2 is set, meaning sequences with a CV value less than 0.2 are considered to have low volatility. In practice, if an entry in a coronavirus infection sample has an average expression of 120 and a standard deviation of 15, then the CV is 0.125, which is less than the 0.2 threshold, and therefore retained. If another sequence has an average expression of 100 and a standard deviation of 30, then the CV is 0.3, and it is excluded. Finally, only those sequences that belong to the top 20% of expression intensity in the sample group and also satisfy the CV requirement of less than 0.2 are considered as expression-stable entries, forming the expression-stable entry set.

[0028] S202: Based on the expression stable entry set, analyze the distribution information of the sequence in the differential source samples, calculate the source repetition rate between sequences, identify entries with a repetition rate lower than the repetition threshold, mark and remove all abnormal signal entries, and obtain the source repetition entry set; The distribution of each sequence across different source sample groups was extracted. Cross-group stability was assessed by statistically analyzing the presence of each sequence in each group. A source distribution matrix was constructed for each sequence, where rows represent the sample source group, columns represent sequence numbers, and cells are 0 or 1, indicating whether the sequence is expressed in that source group. The source repetition count was obtained by summing each column of the matrix and then calculating the ratio to the total number of groups to obtain the source repetition rate. For example, if a sequence is present in the pathogen and coronavirus groups but not expressed in the healthy group, its source repetition rate is 2 / 3, approximately 66.7%. To ensure the representativeness of the selected sequences across multiple sources, a repetition rate threshold of 50% was set, meaning the sequence must appear in at least half of the source groups to be retained. This value was found through simulation testing across multiple datasets to balance representativeness and specificity. The balance point is that if a sequence is expressed only in a single group, such as only in the coronavirus group, its repetition rate is 1 / 3, which is less than 50%, and it is considered an entry with excessive source selectivity. Based on this, anomalous signal entries are further identified. The criteria for anomaly judgment are set as follows: the expression intensity is more than twice the mean of the sequence in the group, or its standard deviation is more than three times the mean standard deviation of the sequence in the group, or its fluctuation coefficient (CV) is greater than 0.4. If any one of the three judgment conditions is met, it is judged as an anomalous signal entry. For example, if the mean expression of a sequence in the pathogen infection group is 80, and the expression of the sequence in a certain sample is 180, it is an anomalous sample and the sequence is removed. By performing labeling and removal operations on all low repetition rate and anomalous signal entries, the sequence entries with a source repetition rate ≥ 50% and that do not meet the anomalous conditions are finally retained to form the source repetition entry set.

[0029] S203: Based on the sequence content of the source duplicate entry set, retrieve the occurrence records in multiple batches of samples, filter the sequence entries that can continuously appear in batches of samples, establish corresponding index sequences, and generate a stable infection signal set; A search was performed on sample data from multiple batches to determine the presence of each sequence in each batch. Batch division was based on sample collection date, experimental operation number, or sequencing device number, dividing the samples into different experimental batches. Each batch contained at least 10 samples to ensure statistical representativeness. Within each batch, the presence of a particular sequence was determined by the presence of at least 30% of the samples within that batch with an expression value greater than 10. Setting 30% as the presence threshold considered the unavoidable presence of a small number of low-expression samples and the tolerance range for intra-sample variability. If a sequence had an expression value greater than 10 in 4 out of 10 samples in batch A, it was considered to have appeared validly. If only 1 sample in batch B had an expression value greater than 10, it was considered to have appeared validly. This expression is not considered valid. The number of batches in which each sequence appears validly is recorded across all batches, and then the ratio is calculated with the total number of batches to obtain the sustained occurrence rate. If this rate is higher than a set threshold of 70%, the entry is considered a stable expression entry. This threshold is set to take into account the impact of changes in experimental conditions across batches on expression stability. If a sequence appears validly in 5 out of 7 batches, the occurrence rate is 71.4%, which can be considered a stable entry. If it appears validly in only 2 batches, it is excluded as an unstable entry. Finally, the retained stable expression entries are numbered and their corresponding batch, sample location information, expression intensity, and other index items are recorded to construct a complete set of indexed sequences. The entries in this set are the stable infection signal set.

[0030] Please see Figure 4 The specific steps of S3 are as follows: S301: Call all segment sequences in the stable infection signal set, extract the start and end boundary position information of each group of segments, count the overlapping intervals between the boundaries of the differential segments, and accumulate the co-occurrence times of segments with overlapping intervals to generate a segment overlap co-occurrence matrix. First, the start and end positions of each fragment in the chromosome or reference genome are extracted. A boundary index list is constructed for each group of fragments, with each fragment recorded as an interval information item consisting of start and end coordinates. For example, fragment A is [12540, 12760], fragment B is [12650, 12820], and fragment C is [13000, 13200]. By traversing the boundary intervals between any two fragments, it is determined whether the two fragments overlap. Specifically, if the start position of fragment B is less than the end position of fragment A, and the end position of fragment B is greater than the start position of fragment A, then the two fragments are considered to have an overlapping interval. If A and B satisfy the above conditions... The condition is to determine if there is overlap. Then, the number of times each pair of segments with an overlapping relationship co-occurs in different samples is counted. The number of times two segments appear together in the same batch or the same sample. For example, if segment A and segment B co-occur in 47 out of 100 samples, the number of times they co-occurs is 47. A two-dimensional table with segment numbers as rows and columns is constructed to record this co-occurrence information. The value of each cell in this matrix is ​​the number of times the corresponding two segments co-occur between samples. If there is no co-occurrence, 0 is filled in. Through all the above segment boundary judgments and co-occurrence statistics, a segment overlap co-occurrence matrix is ​​finally obtained, which includes the overlap relationship judgment results and sample co-occurrence statistics, with each segment pair as the unit.

[0031] S302: Based on the fragment pair information recorded in the fragment overlap co-occurrence matrix, retrieve the boundary distance value of each fragment pair and compare it with the fixed interval distance threshold. If the distance is less than the fixed interval distance threshold and the number of co-occurrences is greater than the co-occurrence frequency threshold, then mark the fragment pair as connectable and obtain the fragment connection index set. For each pair of co-occurring fragments, their boundary positions are further obtained, and the boundary distance between the two fragments is calculated. The boundary distance is defined as the absolute distance between the nearest boundaries of the two fragments. If the termination position of fragment A is 12760 and the start position of fragment B is 12800, then the boundary distance is 40. This distance value is compared with a preset fixed interval distance threshold, which is set to 100. The value is set based on the statistical distribution of the average fragment length, and the maximum tolerable length that can represent the actual inter-gene interval is selected as the connection criterion. When the boundary distance value is less than 100, it is determined that the connection distance condition is met. At the same time, the co-occurrence of the fragment pair in the fragment overlap co-occurrence matrix is ​​also considered. The frequency of occurrence is extracted and compared with a set co-occurrence frequency threshold of 20. This means that at least 20 samples in the same fragment pair must co-occur to be considered as having actual relevance. If the boundary distance between fragments A and B is 40 and the co-occurrence frequency is 47, then both conditions are met and they are marked as a connectable fragment pair. If the boundary distance is 90 but the co-occurrence frequency is 12, then they are excluded. This co-occurrence frequency threshold is set by performing frequency histogram analysis on the number of co-occurring fragments in multiple sets of samples to ensure the statistical stability of the connection relationship. In this way, all fragment pairs that meet the dual threshold conditions of distance and frequency are selected to generate a fragment connection index set composed of fragment pairs.

[0032] S303: Based on all connectable fragment pairs marked in the fragment connection index set, aggregate adjacent fragment sequences in index order, remove fragments that do not meet the connection conditions, construct a list of continuous fragment combination sequences, and generate an infection-associated structure sequence group; Aggregation is performed according to the index order of the fragments in the original sequence. Each group of fragments is checked sequentially to see if a valid label relationship exists in the connection index set. If it does, it is retained and sequence splicing is performed. During splicing, the starting boundary of the smaller fragment is taken as the start point of the combination, and the ending boundary of the larger fragment is taken as the end point. The intermediate sequence is extracted into continuous combined segments. If fragment C starts at position 13000, fragment D starts at position 13080, and ends at position 13220, and their connection relationship is valid, then the range of the combined fragments is from 13000 to 13220, and this range is taken as... New sequences are recorded. Segments that do not appear in the connection index set or do not meet the connection conditions are not spliced ​​and are directly excluded. For example, the boundary distance between segment E and segment F is 200, which does not meet the condition of setting a maximum distance of 100, so they are considered to be unable to be connected. Even if their co-occurrence frequency is met, they are not combined. By traversing all segments and connecting them in index order, a sequence set composed of multiple adjacent connected segments is finally obtained. Each item is a continuous, uninterrupted, and statistically supported structural combination. All aggregation results form an infection-associated structural sequence group.

[0033] Please see Figure 5The specific steps of S4 are as follows: S401: Call all fragment combination information in the infection-associated structure sequence group, combine with the fragment combination position index recorded in the public interaction data, extract the start and end position information of the sequence within the fragment combination, perform pairwise cross-alignment with the interaction position, calculate the alignment hit value and generate the sequence interaction hit matrix. First, each fragment combination is broken down into its constituent individual fragments, and the start and end positions of each fragment in the reference genome are extracted and recorded as a set of boundary coordinates. Then, all fragments within each combination are numbered, and their combination information is recorded. Next, fragment binding position information from an external public interaction database is imported. The binding position coordinates of each interaction record are extracted and cross-referenced one-to-one with the start and end positions in the fragment combination. Specifically, the cross-reference operation determines whether a given interaction coordinate falls within the start and end interval of a fragment combination. If the interaction coordinate value is greater than or equal to the start position of a fragment and less than or equal to the end position, then the interaction record and the fragment are considered to have a match. For example, the interval for fragment P1 is [15400, 1...].

[5620] If a combination position in a public interaction record is 15530, it is considered a hit. This hit result is recorded in a two-dimensional matrix, where the rows represent fragment combination numbers and the columns represent interaction record numbers. If a hit occurs, the corresponding cell is marked as 1; otherwise, it is marked as 0. For multiple interaction records hitting the same combination, the hit count is accumulated. After comparing all fragment combinations with all interaction sites, the hit records of each fragment combination are statistically analyzed by summing their row vectors to obtain the cumulative hit count of each fragment combination in the interaction data. For a certain interaction site hitting multiple fragment combinations, the column vectors can also be summed to obtain its overall activity level. Finally, a complete sequence interaction hit matrix is ​​generated for subsequent screening and analysis.

[0034] S402: Based on the number of combinations recorded in the sequence interaction hit matrix, retrieve the cumulative hit records of each pair of sequence combination regions, and remove combination pairs with fewer hits than the interaction statistical benchmark value, retaining only the combination information that meets the record frequency requirement, to obtain the multi-condition co-occurrence combination pair set; First, extract the cross-hit values ​​between each pair of fragment combinations from the matrix. That is, for each row and column cell with an interaction hit, mark the value of that cell as a combination record. Then, for the location of the combination in each pair of combinations, number and accumulate them according to the specific coordinates provided in the interaction data. Sum the hit counts of all corresponding combination events recorded in all samples and batches to obtain the cumulative hit value of that combination region. For example, if combination A and B appear 3 times in the interaction data, corresponding to the coordinates in samples X, Y, and Z respectively, then its cumulative hit count is 3. After performing this cumulative operation on all combination pairs, compare with the pre-... A statistical baseline value for the interaction is set to determine whether to retain the binding pair. The baseline value is set to 2, meaning that only binding pairs with a cumulative hit count greater than or equal to 2 are retained. This threshold is calculated based on the average repetition of homologous binding events in 10 batches of data. It is set to avoid misjudgment caused by occasional hits. If a binding pair hits only once in one sample, the cumulative hit count is 1, which is less than 2. It is directly removed and not counted as a valid binding pair. Binding pairs with a hit frequency of 2 or more are retained as reliable events. The set of all binding pairs that meet this multi-condition judgment criterion is constructed, which is the multi-condition co-occurrence binding pair set.

[0035] S403: Based on the connection relationship of all binding pairs in the multi-condition co-occurrence binding pair set, construct an undirected graph structure and set the sequence index as a node and the binding pair as an edge, remove independent nodes without connection pairs, and generate a target interaction graph. Based on the index number of the fragment combination in the infection-associated structure sequence group, all combination numbers are used as nodes in the graph structure. Each combination pair is regarded as an edge connecting two nodes, thus constructing an undirected graph structure. Specifically, the fragment numbers at both ends of the multi-condition co-occurrence combination pair are read one by one and used as a pair of connecting points in the graph for edge addition. If the two end numbers are combination C and combination D, then nodes C and D and the connecting edge CD are added to the graph. All combination pairs that meet the conditions are added in this way. After all additions are completed, all nodes in the graph are traversed to determine whether there are connecting edges. If a node does not appear in any edge, that is, the fragment combination has not formed a combination hit relationship with any other combination, then the node is removed from the graph, excluding fragment combinations that exist independently but have no actual connection. After deleting nodes without connecting pairs, the remaining graph structure is the connection network formed between all fragment combinations with interaction relationships. Each node represents a specific fragment combination, and each edge represents a combination event between two combinations that has a hit in the public interaction data and meets the frequency requirement. Finally, a complete target interaction relationship graph is generated.

[0036] Please see Figure 6 The specific steps of S5 are as follows: S501: Call the binding pair position index recorded in the target interaction diagram, extract the infection target position segment in the pathogen and coronavirus association data, perform coordinate overlap comparison for each group of infection types, construct the corresponding overlap matrix according to the intersection ratio, and generate the target position overlap distribution table. First, the chromosome numbers and boundary segment information of the two segments in each binding pair are extracted, and a binding pair position list is constructed. Each item in this list records the start and end coordinate ranges of the binding pair. For example, binding pair AB contains segment A: [15000, 15200] and segment B: [15210, 15450], and the combined binding pair position is [15000, 15450]. Then, the infection target location sets independently labeled in the pathogenic bacteria and coronavirus infection data are introduced, and the coordinate range of each target region is extracted from them. For example, pathogen target X is [15100, 15300], and coronavirus target Y is [15050, 15500]. For each infection type, a target set is established, and cross-comparison is performed with the binding pair coordinates. The comparison process determines the coordinate range of the two segments. If there is an intersection between the segments, and the start and end intervals of the two segments overlap (i.e., there is an intersecting area between the two start and end coordinates), then calculate the length of the intersection and perform a ratio operation between this value and the length of the binding pair segment. The calculated value is the overlap ratio between the binding pair and the target point of the corresponding infection type. For example, if the length of binding pair AB is 450bp and the overlap segment with target point X is 100bp, then the overlap ratio is 100 / 450≈22.2%. Record this ratio value in the overlap matrix, with binding pairs as rows, infection types as columns, and matrix cells as the corresponding overlap ratios. Repeat the above steps to complete the cross-comparison and overlap ratio calculation of all binding pairs with each type of infection target area. After the construction is completed, output the statistical values ​​of the overlap ratio of all binding pairs under different infection types, and finally generate the target location overlap distribution table.

[0037] S502: Based on the overlap ratio values ​​in the target location overlap distribution table, select the binding regions that exist in both differential infection types and have an overlap ratio greater than the target intersection benchmark threshold, and remove regions with an overlap ratio lower than the target intersection benchmark threshold to obtain a set of shared high-frequency binding regions. First, the overlap data of all binding pairs in the distribution table under the two infection types are read. The overlap ratios under bacterial infection and coronavirus infection are extracted separately. Binding pairs with non-zero overlap in both types are selected, that is, binding pairs whose overlap values ​​in both the bacterial and viral columns of the distribution table are not 0. If either is 0, they are excluded. Next, the binding pairs that meet the above conditions are further evaluated by averaging their overlap ratios under the two infection types. For example, if the overlap ratio of a binding pair is 36.5% under bacterial infection and 41.7% under coronavirus infection, the average overlap is 39.1%. This average value is compared with a preset target intersection benchmark threshold to determine whether to retain it. The threshold for point intersection is set at 35%. This value is determined by analyzing statistical charts of typical pathogen intersection segments. The median level, which represents the coverage of the co-occurring area, is selected as the standard. If the average overlap ratio is greater than the threshold, the binding pair is retained; if it is lower than the threshold, it is removed. For example, if the average of the above binding pairs is 39.1%, which is greater than 35%, it is considered an effective shared binding region. Conversely, if the two overlap ratios of another binding pair are 22.4% and 27.6%, with an average of 25%, which is less than 35%, the binding pair is removed. Finally, all binding pairs that simultaneously meet the criteria of having an intersection under two infection types and having an average overlap ratio greater than 35% are selected to construct a set of shared high-frequency binding regions.

[0038] S503: Based on the sequence segments of shared high frequency combined with regional centralized identification, the response frequency in dual infection signal records is statistically analyzed, and fragments with stable response frequencies in all infection samples are screened to construct a list of locatable target fragments and obtain the results of infection treatment target mining. The expression matrix corresponding to each sample in the dual-infection signal record is extracted. For each shared region, an expression response is statistically analyzed across all samples according to the sample number. The criterion is whether the signal reading in a sample within a given region exceeds a set response threshold (set to 10). To ensure the signal in that region is significantly higher than the background noise, if a region has signal readings of 12, 15, and 9 in samples S1, S2, and S3 respectively, it is considered to have responded in S1 and S2, but not in S3. The response frequency of each segment is recorded across all samples, i.e., the number of samples in which that region is judged to have responded out of the total number of samples. For example, if 32 out of 40 samples respond to that region, the response frequency is... 32 ÷ 40 = 80%. The response frequency values ​​of all shared regions are summarized and compared with the stable response frequency threshold, which is set to 70%. That is, a fragment is considered stable only if it responds in more than 70% of the samples. This ratio is set with reference to the minimum ratio requirement for stable expression of the same site in multiple batches of repeated experiments. If a shared region responds 42 times in 60 samples, the ratio is 70%, which meets the setting. The fragment is retained as a localizable target fragment. If another fragment has a response rate of only 55%, it is removed. Finally, all regions with a response frequency of 70% or more in dual-infection samples are retained to construct a list of localizable target fragments, forming the final result of infection treatment target mining.

[0039] Please see Figure 7 A system for identifying therapeutic targets for pathogenic bacteria and coronavirus infections, including: The infection feature extraction module is used to achieve S1: extracting protein and nucleic acid expression reads from bacterial and coronavirus infection samples, locating molecular sites, removing redundant signals, and grouping them in combination with sample information to generate a set of infection molecular features; The stable signal screening module is used to implement S2: based on the set of infected molecular features, the feature set is graded by expression intensity and stability, repetitive and uniformly distributed fragments are identified, abnormal signals are excluded, and sequences appearing in multiple batches are extracted to form a stable set of infected signals. The structural sequence construction module is used to implement S3: call the stable infection signal set, compare the boundary overlap and co-occurrence frequency of stable signal segments, analyze the connection relationship, merge adjacent segments with regular co-occurrence, and generate an infection-associated structural sequence group; The target interaction screening module is used to implement S4: compare infection-related structural sequence groups with public databases, analyze the repetition of binding regions, exclude occasional interactions, retain repeated binding pairs, and construct a target interaction connection graph. The core target identification module is used to implement S5: based on the target interaction diagram, it analyzes the overlap and response patterns between targets, identifies the shared binding regions in differentiated infection types, extracts the core fragments of stable responses, and obtains the results of infection treatment target mining.

[0040] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.

Claims

1. A method for identifying therapeutic targets for pathogenic bacteria and coronavirus infections, characterized in that, Includes the following steps: S1: Extract protein and nucleic acid expression reads from bacterial and coronavirus infection samples, locate molecular sites, remove redundant signals, and group them in combination with sample information to generate a set of infection molecular features; S2: Based on the set of infectious molecular features, the feature set is graded by expression intensity and stability, repeating and uniformly distributed fragments are identified, abnormal signals are excluded, and sequences appearing in multiple batches are extracted to form a stable set of infectious signals. S3: Call the stable infection signal set, compare the boundary overlap and co-occurrence frequency of stable signal segments, analyze the connection relationship, merge adjacent segments with regular co-occurrence, and generate an infection-associated structural sequence group; S4: Compare the infection-associated structural sequence group with the public database, analyze the repetition of the binding regions, exclude occasional interactions, retain repeated binding pairs, and construct a target interaction diagram. S5: Based on the target interaction diagram, analyze the overlap and response patterns between targets, identify the shared binding regions in differentiated infection types, extract the core fragments of stable responses, and obtain the results of infection treatment target mining.

2. The method for identifying therapeutic targets for pathogenic bacteria and coronavirus infections according to claim 1, characterized in that, The set of infectious molecular features includes expression intensity information, staining site markers, sequencing sequence localization, and sample source labels. The set of stable infection signals includes high-frequency fragment sequences, distribution consistency markers, cross-batch stability parameters, and abnormal signal filtering records. The cluster of infection-associated structural sequences includes boundary overlapping regions, co-occurrence frequency matrices, fragment spacing parameters, and continuous structural units. The target interaction diagram includes a binding position lookup table, repeated contact statistics, cross-conditional interaction pairs, and interaction screening results. The results of infection treatment target mining include shared binding regions, stable response fragments, dual infection marker sites, and high-frequency interacting targets.

3. The method for identifying therapeutic targets for pathogenic bacteria and coronavirus infections according to claim 1, characterized in that, The merging of adjacent segments that co-occur regularly will merge stable signal segments that frequently appear adjacently and simultaneously in the sample into a continuous structural unit based on boundary overlap and co-occurrence patterns.

4. The method for identifying therapeutic targets for pathogenic bacteria and coronavirus infections according to claim 1, characterized in that, The core fragments of the stable response refer to key functional sequences that exhibit expression and consistent response patterns under differentiated infection conditions.

5. The method for identifying therapeutic targets for pathogenic bacteria and coronavirus infections according to claim 1, characterized in that, The specific steps of S1 are as follows: S101: Based on the expression reads in bacterial and coronavirus infection samples, protein and nucleic acid sequences are extracted, and the positions of molecular sites in the reads in the stained regions and sequencing sequences are retrieved to establish a molecular site localization matrix; S102: Call the molecular site localization matrix, perform joint sorting and remove duplicate and low-frequency sites based on the occurrence frequency and peak width of the molecular sites, and obtain a set of non-redundant molecular sites; S103: Based on the non-redundant molecular site set, the samples are grouped and summarized in combination with the sample source identifier. The sites in the same group are aggregated and duplicate signals are removed to generate an infected molecular feature set.

6. The method for identifying therapeutic targets for pathogenic bacteria and coronavirus infections according to claim 1, characterized in that, The specific steps of S2 are as follows: S201: Call the sequence entries in the infected molecular feature set, and perform value sorting and fluctuation range comparison in the same type of samples according to the expression intensity and signal stability parameters of each sequence, and filter the sequence entries that simultaneously meet the expression and fluctuation range requirements to generate a set of expression stable entries; S202: Based on the expressed stable entry set, analyze the distribution information of the sequence in the differential source samples, calculate the source repetition rate between sequences, identify entries with a repetition rate lower than the repetition threshold, mark and remove all abnormal signal entries, and obtain the source repetition entry set; S203: Based on the sequence content in the source repeating entry set, retrieve the occurrence records in multiple batches of samples, filter the sequence entries that can continuously appear in the batches of samples, establish the corresponding index sequence, and generate a stable infection signal set.

7. The method for identifying therapeutic targets for pathogenic bacteria and coronavirus infections according to claim 1, characterized in that, The specific steps for S3 are as follows: S301: Call all segment sequences in the stable infection signal set, extract the start and end boundary position information of each group of segments, count the overlapping intervals between the boundaries of the differentiated segments, and accumulate the co-occurrence times of segments with overlapping intervals to generate a segment overlap co-occurrence matrix. S302: Based on the fragment pair information recorded in the fragment overlap co-occurrence matrix, retrieve the boundary distance value of each fragment pair and compare it with a fixed interval distance threshold. If the distance is less than the fixed interval distance threshold and the number of co-occurrences is greater than the co-occurrence frequency threshold, then mark the fragment pair as connectable and obtain the fragment connection index set. S303: Based on all connectable fragment pairs marked in the fragment connection index set, aggregate adjacent fragment sequences in index order, remove fragments that do not meet the connection conditions, construct a list of continuous fragment combination sequences, and generate an infection-associated structure sequence group.

8. The method for identifying therapeutic targets for pathogenic bacteria and coronavirus infections according to claim 1, characterized in that, The specific steps of S4 are as follows: S401: Call all fragment combination information in the infection-associated structural sequence group, combine with the fragment combination position index recorded in the public interaction data, extract the start and end position information of the sequence within the fragment combination and perform a pairwise cross-alignment with the interaction position, calculate the alignment hit value and generate the sequence interaction hit matrix. S402: Based on the number of combinations recorded in the sequence interaction hit matrix, retrieve the cumulative hit records of each pair of sequence combination regions, and remove combination pairs with fewer hits than the interaction statistical benchmark value, retaining only the combination information that meets the record frequency requirement, to obtain a set of multi-condition co-occurrence combination pairs; S403: Based on the connection relationship of all binding pairs in the multi-condition co-occurrence binding pair set, construct an undirected graph structure and set the sequence index as a node and the binding pair as an edge, remove independent nodes without connection pairs, and generate a target interaction graph.

9. The method for identifying therapeutic targets for pathogenic bacteria and coronavirus infections according to claim 1, characterized in that, The specific steps of S5 are as follows: S501: Call the binding pair position index recorded in the target interaction diagram, extract the infection target position segment in the pathogen and coronavirus association data, perform coordinate overlap comparison for each group of infection types, construct the corresponding overlap matrix according to the intersection ratio, and generate the target position overlap distribution table. S502: Based on the overlap ratio values ​​in the target location overlap distribution table, select the binding regions that exist in both differential infection types and have an overlap ratio greater than the target intersection benchmark threshold, and remove regions with an overlap ratio lower than the target intersection benchmark threshold to obtain a set of shared high-frequency binding regions. S503: Based on the sequence segments identified by the shared high-frequency combined region, the response frequency in the dual infection signal records is counted, and segments with stable response frequencies in all infection samples are selected to construct a list of locatable target segments and obtain the results of infection treatment target mining.

10. A system for identifying therapeutic targets for pathogenic bacteria and coronavirus infections, characterized in that, The system is used to implement the method for identifying therapeutic targets for pathogenic bacteria and coronavirus infections as described in any one of claims 1-9, the system comprising: The infection feature extraction module is used to achieve S1: extracting protein and nucleic acid expression reads from bacterial and coronavirus infection samples, locating molecular sites, removing redundant signals, and grouping them in combination with sample information to generate a set of infection molecular features; The stable signal screening module is used to implement S2: based on the set of infectious molecular features, the feature set is graded by expression intensity and stability, repeating and uniformly distributed fragments are identified, abnormal signals are excluded, and sequences appearing in multiple batches are extracted to form a stable set of infectious signals; The structural sequence construction module is used to implement S3: call the stable infection signal set, compare the boundary overlap and co-occurrence frequency of stable signal segments, analyze the connection relationship, merge adjacent segments with regular co-occurrence, and generate an infection-associated structural sequence group; The target interaction screening module is used to implement S4: compare the infection-related structural sequence group with the public database, analyze the repetition of the binding region, exclude occasional interactions, retain repeated binding pairs, and construct a target interaction connection graph. The core target identification module is used to implement S5: based on the target interaction diagram, analyze the overlap and response patterns between targets, identify the shared binding regions in the differentiated infection types, extract the core fragments of stable responses, and obtain the results of infection treatment target mining.