Genome comparative analysis system for multi-drug-resistant staphylococcus aureus, prophage and split phage

By designing a genome comparison analysis system for multidrug-resistant Staphylococcus aureus, prophages, and schistosomes, this study solves the problem of the difficulty in systematically analyzing genome changes in existing technologies. It achieves a clear presentation of dynamic genome changes and reveals the role of phages, supporting drug development and drug resistance research.

CN120895089APending Publication Date: 2025-11-04SHIHEZI UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511054228.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-30
Publication Date
2025-11-04

AI Technical Summary

Technical Problem

Existing technologies struggle to fully integrate genomic data from multidrug-resistant Staphylococcus aureus, prophages, and schistophages. They also lack systematic analytical tools, making it difficult to deeply reveal changes in gene function, dynamic changes in drug resistance, and the impact of phages on the host genome.

Method used

A genome comparison analysis system for multidrug-resistant Staphylococcus aureus, prophages, and schistophages was designed. Through an initial gene module, a first gene module, a second gene module, a gene mutation module, and a gene synthesis module, genomic sequence data at different time periods were acquired and analyzed. Initial, dynamic changes, and mutation information were integrated to construct a complete analysis chain.

Benefits of technology

It enables comprehensive analysis of the genomes of multidrug-resistant Staphylococcus aureus, prophages, and schistophages, clearly presenting the dynamic changes in genes, revealing the interaction between phages and the host, and providing a basis for decision-making in drug development and research on drug resistance mechanisms.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120895089A_ABST
    Figure CN120895089A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of genome comparison, and particularly relates to a genome comparative analysis system for multi-drug-resistant staphylococcus aureus, prophage and split phage. According to the present invention, the genome sequence analysis of the multi-drug-resistant staphylococcus aureus, the prophage and the split phage is integrated, the initial and dynamic change and mutation information is covered, the complete analysis chain is constructed, the gene data analysis in different time periods can clearly present the gene dynamic change process, and the analysis result is accurate. The invention discloses a response mechanism of a microbial genome to the influence of an external environment or bacteriophage, discloses the interaction between the bacteriophage and the multi-drug-resistant staphylococcus aureus, particularly the action in a drug-resistant gene transfer or anti-bacteriophage mechanism, can efficiently recognize important gene changes, and can be used for preparing the multi-drug-resistant staphylococcus aureus. And a powerful decision basis is provided for drug research and development and drug resistance mechanism research.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of genome comparison technology, specifically relating to a genome comparison analysis system for multidrug-resistant Staphylococcus aureus, prophages, and schistosomes. Background Technology

[0002] Staphylococcus aureus is a Gram-positive, coagulase-positive opportunistic pathogen, and the most frequently detected Gram-positive pathogen globally. Multidrug-resistant Staphylococcus aureus (MRSA), in particular, poses a significant challenge to clinical treatment due to its drug resistance and pathogenicity. In recent years, the incidence of MRSA has increased in all regions of the world, imposing a substantial public health burden. Therefore, in-depth research into the genomics of MRSA to understand its drug resistance mechanisms and pathogenicity is crucial for developing new treatment strategies.

[0003] MRSA exhibits resistance to multiple antibiotics, primarily due to the presence of numerous resistance genes. These resistance genes can spread between different strains via horizontal gene transfer, leading to the continuous development of MRSA resistance. Furthermore, MRSA possesses potent virulence genes, enabling systemic infections and resulting in high mortality rates.

[0004] At the genomic level, MRSA genomes exhibit high genetic diversity and evolutionary relationships. Comparative genomic analysis can reveal the phylogenetic relationships, evolutionary pathways, and genetic basis of drug resistance and pathogenicity among MRSA strains. This is of great significance for understanding the transmission mechanisms of MRSA, predicting its epidemic trends, and developing targeted prevention and control measures.

[0005] Prophages are dormant lysogenic bacteriophages whose genetic material is integrated into or embedded in the host bacterium's chromosome. Under certain environmental conditions, prophages can transition from dormancy to an active state, entering the lysis cycle or lysis phase and releasing progeny phages. The presence of prophages not only affects the biological characteristics of the host bacterium but may also spread drug resistance and virulence genes through horizontal gene transfer.

[0006] Lytic phages, also known as virulent phages, do not integrate their genetic material with the host bacterium's chromosome. Instead, they directly replicate, assemble, lyse the host bacterium, and release progeny phages. Compared to lysogenic phages, strictly lysogenic phages do not contain any harmful genes such as lysogens, integrases, resistance genes, or virulence genes. During the lysis of the host bacterium, lytic phages may introduce fragments of the host bacterium's genome (including drug resistance and virulence genes) into new host bacteria, thereby accelerating the spread of these genes.

[0007] Currently, genomic analysis technology is widely used in studying microbial resistance and the relationship between bacteriophages and hosts. However, systematic comparative genomic analyses of multidrug-resistant Staphylococcus aureus, prophages, and schizophages still have shortcomings. Most current studies focus only on genomic data at a single time point, making it difficult to reveal the dynamic changes in the genome over different time periods. There is a lack of a system that can comprehensively integrate initial genes, dynamically changing genes, and mutation information, making it difficult to achieve in-depth interpretation of multidrug resistance gene transfer and the impact of bacteriophage genomes. Existing analytical tools often focus on single target genes and cannot simultaneously analyze gene function changes, dynamic changes in drug resistance, and the impact of bacteriophages on the host genome. Summary of the Invention

[0008] The purpose of this invention is to provide a genome comparison analysis system for multidrug-resistant Staphylococcus aureus, prophages, and schistosomes, which can comprehensively reveal the genome response mechanism of microorganisms and the mode of action of bacteriophages by starting from the initial genome sequence and combining dynamic changes and mutation information at different time periods.

[0009] The specific technical solution adopted by this invention is as follows: A genome comparison analysis system for multidrug-resistant Staphylococcus aureus, prophages, and schistosomes includes an initial gene module, a first gene module, a second gene module, a gene mutation module, and a gene synthesis module. The initial gene module is used to obtain the initial genome sequence data of multidrug-resistant Staphylococcus aureus, prophage, and schistophage; The first gene module is used to construct the first time period, obtain the first genome sequence data of multidrug-resistant Staphylococcus aureus, prophage and cleavage phage within the first time period, and obtain the first gene information based on the initial genome sequence data and the first genome sequence data; The second gene module is used to construct the second time period, obtain the second genome sequence data of multidrug-resistant Staphylococcus aureus, prophage and cleavage phage within the second time period, and obtain the second gene information based on the initial genome sequence data and the second genome sequence data. The gene mutation module is used to construct the third time period, obtain the third genome sequence data of multidrug-resistant Staphylococcus aureus, prophage and schizophage within the third time period, and obtain gene mutation information of multidrug-resistant Staphylococcus aureus, prophage and schizophage based on the third genome sequence data; The gene synthesis module is used to obtain gene analysis information of multidrug-resistant Staphylococcus aureus, prophages, and schistophages based on initial genome sequence data, first gene information, second gene information, and gene mutation information, and to obtain gene synthesis information based on the gene analysis information.

[0010] In a preferred embodiment, the initial gene module includes a first gene unit, a second gene unit, a third gene unit, and a genome unit; The first gene unit is used to obtain genomic data of multidrug-resistant Staphylococcus aureus; The second gene unit is used to obtain the genomic data of the prophage; The third gene unit is used to obtain the genomic data of the schizophage; The genomic unit is used to summarize the genomic data of multidrug-resistant Staphylococcus aureus, the genomic data of prophages, and the genomic data of schizophages, and to label the summarized results as the initial genomic sequence data of multidrug-resistant Staphylococcus aureus, prophages, and schizophages.

[0011] In a preferred embodiment, the first gene module includes a first time period unit, a first sequence unit, a first initial vector unit, an evolutionary pressure unit, a first vector unit, and a first information unit; The first time period unit is used to construct the first time period. The first sequence unit is used to obtain the first genome sequence data of multidrug-resistant Staphylococcus aureus, prophage, and schistosome within the first time period; The first initial vector unit is used to obtain multiple initial genome sequence feature vectors based on the initial genome sequence data. Evolutionary pressure unit, used to obtain evolutionary pressure information of multiple genes based on first genome sequence data; The first vector unit is used to obtain multiple first genome sequence feature vectors corresponding to each initial genome sequence feature vector based on gene evolution pressure information; The first information unit is used to obtain the first gene value based on multiple initial genome sequence feature vectors and multiple first genome sequence feature vectors corresponding to each initial genome sequence feature vector, and to mark the first gene value as the first gene information.

[0012] In a preferred embodiment, the first time period unit includes a first start unit, a standard duration unit, a first end unit, and a first time period construction unit; The first start unit is used to mark the time node for acquiring the initial genome sequence data as the start time of the first time period; The first standard duration unit is used to obtain the standard duration; The first ending unit is used to obtain the ending time of the first time period based on the brick marking time and the start time of the first time period. The first time period construction unit is used to obtain the first time period based on the start time and end time of the first time period.

[0013] In a preferred embodiment, the second gene module includes a second time period unit, a second sequence unit, a first proportion unit, a second proportion unit, a third proportion unit, a second initial vector unit, and a second information unit; The second time period unit is used to construct the second time period; The second sequence unit is used to obtain the second genome sequence data of multidrug-resistant Staphylococcus aureus, prophage, and schistophage during the second time period; The first proportion unit is used to obtain the proportion of conserved sequences of multidrug-resistant Staphylococcus aureus based on the second genome sequence data and mark it as the first conserved sequence proportion. The second proportion unit is used to obtain the proportion of conserved sequences in the prophage based on the second genome sequence data and mark it as the proportion of the second conserved sequence. The third proportion unit is used to obtain the proportion of conserved sequences in the splitting phage based on the second genome sequence data and mark it as the third conserved sequence proportion. The second initial vector unit is used to obtain multiple initial genome sequence feature vectors based on the initial genome sequence data. The second information unit is used to identify the second gene value based on the characteristic vectors of multiple initial genome sequences, the proportion of the first conserved sequence, the proportion of the second conserved sequence, and the proportion of the third conserved sequence, and to label it as the second gene information.

[0014] In a preferred embodiment, the second time period unit includes a second start unit, a first acquisition unit, a second standard duration unit, a standard value unit, a second duration unit, a second end unit, and a second time period construction unit; The second start unit is used to obtain the end time of the first time period and mark the start time of the second time period; The first acquisition unit is used to acquire the corresponding first gene value based on the first gene information; The second standard duration unit is used to obtain the standard duration of the first time period; Standard value unit, used to obtain standard gene values; The second duration unit is used to obtain the second duration based on the first gene value, the standard duration, and the standard gene value. The second end unit is used to obtain the end time of the second time period based on the second duration and the start time of the second time period; The second time period construction unit is used to obtain the second time period based on the start time and end time of the second time period.

[0015] In a preferred embodiment, the gene mutation module includes a third time segment unit, a third sequence unit, a third vector unit, and a gene mutation unit. The third time period unit is used to construct the third time period; The third sequence unit is used to obtain the third genome sequence data of multidrug-resistant Staphylococcus aureus, prophage, and schistosome within the third time period; The third vector unit is used to obtain multiple corresponding third genome sequence feature vectors based on the third genome sequence data; Gene mutation unit is used to obtain gene mutation values ​​based on the characteristic vectors of multiple third genome sequences and to mark them as gene mutation information.

[0016] In a preferred embodiment, the third time period unit includes a third start unit, a third end unit, and a third time period construction unit; The third start unit is used to obtain the start time of the first time period and mark it as the start time of the third time period; The third end unit is used to obtain the end time of the second time period and mark it as the end time of the third time period; The third time period construction unit is used to obtain the third time period based on the start time and end time of the third time period.

[0017] In a preferred embodiment, the gene integration module includes an integration vector unit, an analysis value unit, an integration table unit, a target interval unit, and an integration gene unit; The integrated vector unit is used to obtain multiple initial genome sequence feature vectors based on the initial genome sequence data. The analysis value unit is used to obtain gene analysis values ​​for multidrug-resistant Staphylococcus aureus, prophages, and schistophages based on initial genome sequence data, first gene information, second gene information, and gene mutation information. The comprehensive table unit is used to obtain a comprehensive table, which includes multiple gene analysis interval values ​​and gene comparison gene comprehensive information corresponding to each gene analysis interval value; The target interval unit is used to obtain the target gene analysis interval value based on the gene analysis value; The integrated gene unit is used to obtain the corresponding gene integrated information from the integrated table based on the target gene analysis interval value.

[0018] And, a genome comparison analysis terminal for multidrug-resistant Staphylococcus aureus, prophages, and schistophages, including: One or more processors; A storage device on which one or more programs are stored; When one or more programs are executed by one or more processors, the one or more processors enable a genome comparison analysis system for multidrug-resistant Staphylococcus aureus, prophages, and schistophages.

[0019] The technical effects achieved by this invention are as follows: This invention integrates genomic sequence analysis of multidrug-resistant Staphylococcus aureus, prophages, and schistosomes, covering initial, dynamic changes, and mutation information, constructing a complete analytical chain. Genomic data analysis at different time points clearly presents the dynamic changes in genes, revealing the response mechanisms of microbial genomes to external environmental or phage influences, and elucidating the interactions between phages and multidrug-resistant Staphylococcus aureus, particularly their role in drug resistance gene transfer or antiphage mechanisms. It can efficiently identify important gene changes and provide strong decision-making support for drug development and research on drug resistance mechanisms. Attached Figure Description

[0020] Figure 1 This is a system module diagram provided by the present invention; Figure 2 This is a flowchart provided by the present invention. Detailed Implementation

[0021] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings.

[0022] Many specific details are set forth in the following description in order to provide a full understanding of the invention. However, the invention may also be practiced in other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the spirit of the invention. Therefore, the invention is not limited to the specific embodiments disclosed below.

[0023] Secondly, the term "an embodiment" or "embodiment" as used herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The phrase "in a preferred embodiment" appearing in different places throughout this specification does not necessarily refer to the same embodiment, nor is it a single or selective embodiment that mutually excludes other embodiments.

[0024] Furthermore, the present invention will be described in detail with reference to the schematic diagrams. When describing the embodiments of the present invention in detail, the schematic diagrams are merely examples for ease of explanation and should not limit the scope of protection of the present invention.

[0025] Please see the appendix Figure 1As shown, a genome comparison analysis system for multidrug-resistant Staphylococcus aureus, prophages, and schistosomes is provided, including an initial gene module, a first gene module, a second gene module, a gene mutation module, and a gene synthesis module; The initial gene module is used to obtain the initial genome sequence data of multidrug-resistant Staphylococcus aureus, prophage, and schistophage; The first gene module is used to construct the first time period, obtain the first genome sequence data of multidrug-resistant Staphylococcus aureus, prophage and cleavage phage within the first time period, and obtain the first gene information based on the initial genome sequence data and the first genome sequence data; The second gene module is used to construct the second time period, obtain the second genome sequence data of multidrug-resistant Staphylococcus aureus, prophage and cleavage phage within the second time period, and obtain the second gene information based on the initial genome sequence data and the second genome sequence data. The gene mutation module is used to construct the third time period, obtain the third genome sequence data of multidrug-resistant Staphylococcus aureus, prophage and schizophage within the third time period, and obtain gene mutation information of multidrug-resistant Staphylococcus aureus, prophage and schizophage based on the third genome sequence data; The gene synthesis module is used to obtain gene analysis information of multidrug-resistant Staphylococcus aureus, prophages, and schistophages based on initial genome sequence data, first gene information, second gene information, and gene mutation information, and to obtain gene synthesis information based on the gene analysis information.

[0026] As described above, the initial gene module is responsible for acquiring the initial genomic sequence data of multidrug-resistant Staphylococcus aureus, prophages, and schizophages. It performs high-precision analysis and stores the genomes of the target microorganisms using sequencing technology, providing raw data support for subsequent modules. The first gene module constructs the first time period and acquires the genomic sequence data within that period, as well as the differences in the initial genomic sequence data, extracting gene change information (first gene information) within the first time period. The second gene module is similar to the first gene module, but it targets genomic changes within the second time period, extracting second gene information. The gene mutation module acquires data for the third time period, extracting gene mutation information for multidrug-resistant Staphylococcus aureus, prophages, and schizophages. The gene integration module integrates the initial genomic sequence data, first gene information, second gene information, and gene mutation information to form a complete gene sequence. Based on gene analysis information, this system generates comprehensive gene information, including gene function annotation, dynamic changes in drug resistance-related genes, and the impact of bacteriophages on the host genome. It provides comprehensive genome comparison analysis conclusions, supporting scientific research and clinical applications. The system systematically integrates genome sequence analysis of multidrug-resistant Staphylococcus aureus, prophages, and schistosomes, covering initial, dynamic, and variant information, constructing a complete analytical chain. Gene data analysis at different time points clearly presents the dynamic changes in genes, revealing the response mechanisms of microbial genomes to external environmental or bacteriophage influences. It also reveals the interaction between bacteriophages and multidrug-resistant Staphylococcus aureus, particularly their role in drug resistance gene transfer or antiphage mechanisms. The system can efficiently identify important gene changes and provide strong decision-making support for drug development and research into drug resistance mechanisms.

[0027] In a preferred embodiment, the initial gene module includes a first gene unit, a second gene unit, a third gene unit, and a genome unit; The first gene unit is used to obtain genomic data of multidrug-resistant Staphylococcus aureus; The second gene unit is used to obtain the genomic data of the prophage; The third gene unit is used to obtain the genomic data of the schizophage; The genomic unit is used to summarize the genomic data of multidrug-resistant Staphylococcus aureus, the genomic data of prophages, and the genomic data of schizophages, and to label the summarized results as the initial genomic sequence data of multidrug-resistant Staphylococcus aureus, prophages, and schizophages.

[0028] As described above, the first gene unit acquires genomic sequence data of multidrug-resistant Staphylococcus aureus (MRSA) using high-throughput sequencing technology. The data includes gene information related to drug resistance, virulence factors, and metabolic pathways. The second gene unit extracts genomic data of prophages, mainly including prophage sequences embedded in the host genome. The third gene unit acquires genomic data of schizophages, mainly from schizophages independent of the host, including information on genes related to their infection cycle, especially lysis function and possible drug resistance gene fragments. The genomic unit summarizes the data generated by the above three units, integrating the genomic data of MRSA, prophages, and schizophages into an initial genomic sequence dataset. This ensures that the genomic data from MRSA, prophages, and schizophages are collected comprehensively and systematically, providing a unified input source for genomic comparative analysis.

[0029] In a preferred embodiment, the first gene module includes a first time period unit, a first sequence unit, a first initial vector unit, an evolutionary pressure unit, a first vector unit, and a first information unit; The first time period unit is used to construct the first time period. The first sequence unit is used to obtain the first genome sequence data of multidrug-resistant Staphylococcus aureus, prophage, and schistosome within the first time period; The first initial vector unit is used to obtain multiple initial genome sequence feature vectors based on the initial genome sequence data. Evolutionary pressure unit, used to obtain evolutionary pressure information of multiple genes based on first genome sequence data; The first vector unit is used to obtain multiple first genome sequence feature vectors corresponding to each initial genome sequence feature vector based on gene evolution pressure information; The first information unit is used to obtain the first gene value based on multiple initial genome sequence feature vectors and multiple first genome sequence feature vectors corresponding to each initial genome sequence feature vector, and to mark the first gene value as the first gene information.

[0030] The time dimension of the first time-period unit construction system analysis, defined as "first time period," refers to a specific observation or experimental time range. By clearly defining the time frame, the range of genomic sequence data to be acquired within that time period is determined. The first sequence unit extracts genomic sequence data of multidrug-resistant Staphylococcus aureus, prophages, and schistosomes within the first time period, ensuring data coverage of all analytical objects and forming a complete first-time-period genomic dataset. The first initial vector unit converts the initial genomic sequence data into vector representations to capture the characteristics of each genome (e.g., base composition, GC content, functional gene region distribution, etc.). The initial vector represents the state of the initial genome through mathematical modeling. The evolutionary pressure unit calculates the evolutionary pressure information of genes under environmental selection pressure based on the genomic sequence data of the first time period. The first vector unit adjusts the initial genomic sequence characteristic vector according to the gene evolutionary pressure information to generate the genomic sequence characteristic vector of the first time period. The first information unit calculates the first gene value by combining the initial genomic sequence characteristic vector with the genomic sequence characteristic vector of the first time period. The formula for calculating the first gene value is as follows: In the formula, Let i represent the first gene value, and i represent the index of the multiple initial genome sequence feature vectors, i = 1, 2, 3…n. Let be the feature vector of the i-th initial genome sequence, and h represent the numbers of the multiple first genome sequence feature vectors corresponding to each initial genome sequence feature vector, h = 1, 2, 3…t. This represents the h-th first genome sequence feature vector corresponding to each initial genome sequence feature vector. This enables the system to dynamically monitor genome changes over time periods, helping to reveal the adaptive evolutionary process of the genome at different times. It allows the system to keenly capture minute changes in genes within the first time period, providing a powerful tool for studying dynamic changes in the genome.

[0031] In a preferred embodiment, the first time period unit includes a first start unit, a standard duration unit, a first end unit, and a first time period construction unit; The first start unit is used to mark the time node for acquiring the initial genome sequence data as the start time of the first time period; The first standard duration unit is used to obtain the standard duration; The first ending unit is used to obtain the ending time of the first time period based on the brick marking time and the start time of the first time period. The first time period construction unit is used to obtain the first time period based on the start time and end time of the first time period.

[0032] As described above, the first start unit records and marks the time node when the initial genome sequence data is acquired as the start time of the "first time period". This time point serves as the time baseline for the entire analysis. The first standard duration unit acquires the standard duration set by the system to determine the time span of each time period. The standard duration can be set according to research needs, such as using hours, days, weeks, or other time units as the cycle, which facilitates the adjustment of the system's adaptability. The first end unit calculates the end time based on the start time and standard duration of the first time period, generating the end point of the time range of the "first time period". After determining the end time, a clear time window is formed. The first time period construction unit combines the start time of the first start unit and the end time of the first end unit to construct the complete time range of the first time period. This allows the system to clearly distinguish the time range to which the data belongs, avoids data errors caused by time ambiguity, and provides flexibility. The duration can be adjusted according to research needs, thereby adapting to the needs of genome data analysis at different time scales.

[0033] In a preferred embodiment, the second gene module includes a second time period unit, a second sequence unit, a first proportion unit, a second proportion unit, a third proportion unit, a second initial vector unit, and a second information unit; The second time period unit is used to construct the second time period; The second sequence unit is used to obtain the second genome sequence data of multidrug-resistant Staphylococcus aureus, prophage, and schistophage during the second time period; The first proportion unit is used to obtain the proportion of conserved sequences of multidrug-resistant Staphylococcus aureus based on the second genome sequence data and mark it as the first conserved sequence proportion. The second proportion unit is used to obtain the proportion of conserved sequences in the prophage based on the second genome sequence data and mark it as the proportion of the second conserved sequence. The third proportion unit is used to obtain the proportion of conserved sequences in the splitting phage based on the second genome sequence data and mark it as the third conserved sequence proportion. The second initial vector unit is used to obtain multiple initial genome sequence feature vectors based on the initial genome sequence data. The second information unit is used to identify the second gene value based on the characteristic vectors of multiple initial genome sequences, the proportion of the first conserved sequence, the proportion of the second conserved sequence, and the proportion of the third conserved sequence, and to label it as the second gene information.

[0034] The second time period unit constructs the second time period, determining the time window for analyzing the genome data of multidrug-resistant Staphylococcus aureus (MRSA), prophage, and schizophage. The second sequence unit extracts the genome sequence data of MRSA, prophage, and schizophage within the second time period. The first proportion unit calculates the proportion of conserved sequences of MRSA in the second genome sequence and marks it as the first conserved sequence proportion. The second proportion unit calculates the proportion of conserved sequences of prophage and marks it as the second conserved sequence proportion. The third proportion unit calculates the proportion of conserved sequences of schizophage and marks it as the third conserved sequence proportion. The second initial vector unit generates multiple initial genome sequence feature vectors based on the initial genome sequence data. The second information unit integrates the initial genome feature vectors and the proportions of the first, second, and third conserved sequences, and calculates and generates a second gene value, which is marked as the second gene information. The formula for calculating the second gene value is as follows: , This is represented as the second gene value, where i represents the number of the feature vectors of multiple initial genome sequences, i = 1, 2, 3…n. Represented as the feature vector of the i-th initial genome sequence, This is represented by the proportion of the first conservative sequence. This is represented by the proportion of the second conservative sequence. Represented as the proportion of the third conserved sequence, it can dynamically monitor the genomic evolution of multidrug-resistant Staphylococcus aureus and its related bacteriophages, provide data support for studying gene mutation patterns, quantify the degree of influence of evolutionary pressure on the genome within a specific time period, and help understand gene adaptive changes.

[0035] In a preferred embodiment, the second time period unit includes a second start unit, a first acquisition unit, a second standard duration unit, a standard value unit, a second duration unit, a second end unit, and a second time period construction unit; The second start unit is used to obtain the end time of the first time period and mark the start time of the second time period; The first acquisition unit is used to acquire the corresponding first gene value based on the first gene information; The second standard duration unit is used to obtain the standard duration of the first time period; Standard value unit, used to obtain standard gene values; The second duration unit is used to obtain the second duration based on the first gene value, the standard duration, and the standard gene value. The second end unit is used to obtain the end time of the second time period based on the second duration and the start time of the second time period; The second time period construction unit is used to obtain the second time period based on the start time and end time of the second time period.

[0036] As described above, the second start unit obtains the end time of the first time period and marks it as the start time of the second time period, ensuring continuity between time periods, establishing correlation, and providing a time node basis for multi-time period gene analysis. The first acquisition unit extracts the first gene value (i.e., the comprehensive gene characteristics within the first time period) based on the first gene information. The second standard duration unit obtains the standard duration from the first time period as a reference basis for constructing the second time period. The standard value unit extracts the standard gene value, representing the expected value of gene characteristics or evolutionary indicators, which is used as reference data to measure the gene change trend in the second time period and compared with the actual gene value. The second duration unit comprehensively considers the first gene value, the standard duration, and the standard gene value to calculate the second duration. The formula for calculating the second duration is as follows: In the formula, A represents the second duration and B represents the first duration. Represented as the first gene value, The standard gene value is represented by the second end unit. Based on the start time and the second duration of the second time period, the end time of the second time period is determined. The second time period construction unit combines the start time and end time of the second time period to formally construct the second time period, providing a clear time range. Based on the first gene value, the standard gene value, and the standard duration, the duration of the second time period is dynamically adjusted, which can adapt to the actual rate of genomic change and improve the flexibility and scientific nature of time period division.

[0037] In a preferred embodiment, the gene mutation module includes a third time period unit, a third sequence unit, a third vector unit, and a gene mutation unit; The third time period unit is used to construct the third time period; The third sequence unit is used to obtain the third genome sequence data of multidrug-resistant Staphylococcus aureus, prophage, and schistosome within the third time period; The third vector unit is used to obtain multiple corresponding third genome sequence feature vectors based on the third genome sequence data; Gene mutation unit is used to obtain gene mutation values ​​based on the characteristic vectors of multiple third genome sequences and to mark them as gene mutation information.

[0038] As described above, the third time period unit constructs a third time period based on analysis needs, determining the time range for gene mutation analysis. The third sequence unit extracts genomic sequence data of multidrug-resistant Staphylococcus aureus, prophages, and schistophages within the third time period. The third vector unit transforms the third genomic sequence data into multiple feature vectors, each representing a specific feature of the genomic sequence (such as nucleotide frequency, sequence length, structural features, etc.). The gene mutation unit compares the differences between the feature vectors of the third genomic sequence and calculates the gene mutation value. The formula for calculating the gene mutation value is as follows: In the formula, The value is represented as the gene mutation value, and f represents the number of multiple third-genome sequence characteristic vectors, f=1,2,3…q. Represented as the feature vector of the f-th third genome sequence, it transforms genomic changes into quantifiable gene variation values, making complex gene difference information more intuitive and easier to analyze.

[0039] In a preferred embodiment, the third time period unit includes a third start unit, a third end unit, and a third time period construction unit; The third start unit is used to obtain the start time of the first time period and mark it as the start time of the third time period; The third end unit is used to obtain the end time of the second time period and mark it as the end time of the third time period; The third time period construction unit is used to obtain the third time period based on the start time and end time of the third time period.

[0040] As described above, the third start unit extracts a time node from the start time of the first time period and marks it as the start time of the third time period. The third end unit extracts a time node from the end time of the second time period and marks it as the end time of the third time period, ensuring that the time range of the third time period can cover the first two time periods, which facilitates the global judgment of gene changes. The third time period construction unit combines the start and end times of the third time period to construct the third time period and marks it as the last time period in the analysis scope, ensuring seamless connection of analysis time periods and avoiding the omission of possible important gene changes. Through automated start and end time marking, human errors in the time period division process are avoided, improving the standardization and reliability of the analysis.

[0041] In a preferred embodiment, the gene synthesis module includes a synthesis vector unit, an analysis value unit, a synthesis table unit, a target interval unit, and a synthesis gene unit. The integrated vector unit is used to obtain multiple initial genome sequence feature vectors based on the initial genome sequence data. The analysis value unit is used to obtain gene analysis values ​​for multidrug-resistant Staphylococcus aureus, prophages, and schistophages based on initial genome sequence data, first gene information, second gene information, and gene mutation information. The comprehensive table unit is used to obtain a comprehensive table, which includes multiple gene analysis interval values ​​and gene comparison gene comprehensive information corresponding to each gene analysis interval value; The target interval unit is used to obtain the target gene analysis interval value based on the gene analysis value; The integrated gene unit is used to obtain the corresponding gene integrated information from the integrated table based on the target gene analysis interval value.

[0042] The above-mentioned comprehensive vector unit extracts multiple initial genome sequence feature vectors from the initial genome sequence data, serving as the basic vector set for analysis. These vectors cover the characteristics of multidrug-resistant Staphylococcus aureus, prophage, and schistophage genome sequences. The analysis value unit integrates the initial genome sequence data, first gene information, second gene information, and gene mutation information to calculate the gene analysis value. The formula for calculating the gene analysis value is as follows: In the formula, F represents the gene analysis value, and i represents the number of the feature vectors of multiple initial genome sequences, i=1,2,3…n. Represented as the feature vector of the i-th initial genome sequence, Represented as the first gene value, This is represented as the second gene value. Represented as gene mutation values, the comprehensive table unit creates a comprehensive table, dividing gene analysis values ​​into multiple intervals (gene analysis interval values), and assigning corresponding gene comparison and comprehensive information to each interval. This table serves as a mapping table for gene analysis values, facilitating rapid lookup and classification. The target interval unit matches the calculated gene analysis values ​​with the corresponding target gene analysis interval values. The matching process follows a preset interval range to ensure accurate classification of gene analysis values. The comprehensive gene unit uses the target gene analysis interval values ​​to retrieve the corresponding comprehensive gene information from the comprehensive table. This information integrates gene variations, conserved sequences, and evolutionary trends of multidrug-resistant Staphylococcus aureus, prophages, and schistosomes. It integrates the characteristics of the initial gene sequence, gene information changes and mutation information within a time period, making the analysis results more scientific and comprehensive. It quickly locates the comprehensive information corresponding to gene analysis values, reduces computational complexity, improves analysis efficiency, and can be dynamically optimized according to new data or new requirements, enhancing the system's adaptability.

[0043] And, a genome comparison analysis terminal for multidrug-resistant Staphylococcus aureus, prophages, and schistophages, including: One or more processors; A storage device on which one or more programs are stored; When one or more programs are executed by one or more processors, the one or more processors enable a genome comparison analysis system for multidrug-resistant Staphylococcus aureus, prophages, and schistophages.

[0044] The above description is merely a preferred embodiment of the present invention. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principles of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention. Structures, devices, and operating methods not specifically described or explained in this invention are implemented according to conventional methods in the art unless otherwise specified or limited.

Claims

1. A genome comparison analysis system for multidrug-resistant Staphylococcus aureus, prophages, and schistophages, characterized in that, It includes the initial gene module, the first gene module, the second gene module, the gene mutation module, and the gene synthesis module; The initial gene module is used to obtain the initial genome sequence data of multidrug-resistant Staphylococcus aureus, prophage, and schistophage; The first gene module is used to construct the first time period, obtain the first genome sequence data of multidrug-resistant Staphylococcus aureus, prophage and cleavage phage within the first time period, and obtain the first gene information based on the initial genome sequence data and the first genome sequence data; The second gene module is used to construct the second time period, obtain the second genome sequence data of multidrug-resistant Staphylococcus aureus, prophage and cleavage phage within the second time period, and obtain the second gene information based on the initial genome sequence data and the second genome sequence data. The gene mutation module is used to construct the third time period, obtain the third genome sequence data of multidrug-resistant Staphylococcus aureus, prophage and schizophage within the third time period, and obtain gene mutation information of multidrug-resistant Staphylococcus aureus, prophage and schizophage based on the third genome sequence data; The gene synthesis module is used to obtain gene analysis information of multidrug-resistant Staphylococcus aureus, prophages, and schistophages based on initial genome sequence data, first gene information, second gene information, and gene mutation information, and to obtain gene synthesis information based on the gene analysis information.

2. The genome comparison analysis system for multidrug-resistant Staphylococcus aureus, prophages, and schistophages according to claim 1, characterized in that, The initial gene module includes a first gene unit, a second gene unit, a third gene unit, and a genome unit; The first gene unit is used to obtain genomic data of multidrug-resistant Staphylococcus aureus; The second gene unit is used to obtain the genomic data of the prophage; The third gene unit is used to obtain the genomic data of the schizophage; The genomic unit is used to summarize the genomic data of multidrug-resistant Staphylococcus aureus, the genomic data of prophages, and the genomic data of schizophages, and to label the summarized results as the initial genomic sequence data of multidrug-resistant Staphylococcus aureus, prophages, and schizophages.

3. The genome comparison analysis system for multidrug-resistant Staphylococcus aureus, prophages, and schistophages according to claim 1, characterized in that, The first gene module includes a first time period unit, a first sequence unit, a first initial vector unit, an evolutionary pressure unit, a first vector unit, and a first information unit; The first time period unit is used to construct the first time period. The first sequence unit is used to obtain the first genome sequence data of multidrug-resistant Staphylococcus aureus, prophage, and schistosome within the first time period; The first initial vector unit is used to obtain multiple initial genome sequence feature vectors based on the initial genome sequence data. Evolutionary pressure unit, used to obtain evolutionary pressure information of multiple genes based on first genome sequence data; The first vector unit is used to obtain multiple first genome sequence feature vectors corresponding to each initial genome sequence feature vector based on gene evolution pressure information; The first information unit is used to obtain the first gene value based on multiple initial genome sequence feature vectors and multiple first genome sequence feature vectors corresponding to each initial genome sequence feature vector, and to mark the first gene value as the first gene information.

4. The genome comparison analysis system for multidrug-resistant Staphylococcus aureus, prophages, and schistophages according to claim 2, characterized in that, The first time period unit includes the first start unit, the standard duration unit, the first end unit, and the first time period construction unit; The first start unit is used to mark the time node for acquiring the initial genome sequence data as the start time of the first time period; The first standard duration unit is used to obtain the standard duration; The first ending unit is used to obtain the ending time of the first time period based on the brick marking time and the start time of the first time period. The first time period construction unit is used to obtain the first time period based on the start time and end time of the first time period.

5. The genome comparison analysis system for multidrug-resistant Staphylococcus aureus, prophages, and schistophages according to claim 1, characterized in that, The second gene module includes a second time period unit, a second sequence unit, a first proportion unit, a second proportion unit, a third proportion unit, a second initial vector unit, and a second information unit; The second time period unit is used to construct the second time period; The second sequence unit is used to obtain the second genome sequence data of multidrug-resistant Staphylococcus aureus, prophage, and schistophage during the second time period; The first proportion unit is used to obtain the proportion of conserved sequences of multidrug-resistant Staphylococcus aureus based on the second genome sequence data and mark it as the first conserved sequence proportion. The second proportion unit is used to obtain the proportion of conserved sequences in the prophage based on the second genome sequence data and mark it as the proportion of the second conserved sequence. The third proportion unit is used to obtain the proportion of conserved sequences in the splitting phage based on the second genome sequence data and mark it as the third conserved sequence proportion. The second initial vector unit is used to obtain multiple initial genome sequence feature vectors based on the initial genome sequence data. The second information unit is used to identify the second gene value based on the characteristic vectors of multiple initial genome sequences, the proportion of the first conserved sequence, the proportion of the second conserved sequence, and the proportion of the third conserved sequence, and to label it as the second gene information.

6. The genome comparison analysis system for multidrug-resistant Staphylococcus aureus, prophages, and schistophages according to claim 5, characterized in that, The second time period unit includes a second start unit, a first acquisition unit, a second standard duration unit, a standard value unit, a second duration unit, a second end unit, and a second time period construction unit; The second start unit is used to obtain the end time of the first time period and mark the start time of the second time period; The first acquisition unit is used to acquire the corresponding first gene value based on the first gene information; The second standard duration unit is used to obtain the standard duration of the first time period; Standard value unit, used to obtain standard gene values; The second duration unit is used to obtain the second duration based on the first gene value, the standard duration, and the standard gene value. The second end unit is used to obtain the end time of the second time period based on the second duration and the start time of the second time period; The second time period construction unit is used to obtain the second time period based on the start time and end time of the second time period.

7. The genome comparison analysis system for multidrug-resistant Staphylococcus aureus, prophages, and schistophages according to claim 1, characterized in that, The gene mutation module includes a third time period unit, a third sequence unit, a third vector unit, and a gene mutation unit; The third time period unit is used to construct the third time period; The third sequence unit is used to obtain the third genome sequence data of multidrug-resistant Staphylococcus aureus, prophage, and schistosome within the third time period; The third vector unit is used to obtain multiple corresponding third genome sequence feature vectors based on the third genome sequence data; Gene mutation unit is used to obtain gene mutation values ​​based on the characteristic vectors of multiple third genome sequences and to mark them as gene mutation information.

8. The genome comparison analysis system for multidrug-resistant Staphylococcus aureus, prophages, and schistophages according to claim 7, characterized in that, The third time period unit includes the third start unit, the third end unit, and the third time period construction unit; The third start unit is used to obtain the start time of the first time period and mark it as the start time of the third time period; The third end unit is used to obtain the end time of the second time period and mark it as the end time of the third time period; The third time period construction unit is used to obtain the third time period based on the start time and end time of the third time period.

9. The genome comparison analysis system for multidrug-resistant Staphylococcus aureus, prophages, and schistophages according to claim 1, characterized in that, The gene integration module includes an integration vector unit, an analysis value unit, an integration table unit, a target interval unit, and an integration gene unit; The integrated vector unit is used to obtain multiple initial genome sequence feature vectors based on the initial genome sequence data. The analysis value unit is used to obtain gene analysis values ​​for multidrug-resistant Staphylococcus aureus, prophages, and schistophages based on initial genome sequence data, first gene information, second gene information, and gene mutation information. The comprehensive table unit is used to obtain a comprehensive table, which includes multiple gene analysis interval values ​​and gene comparison gene comprehensive information corresponding to each gene analysis interval value; The target interval unit is used to obtain the target gene analysis interval value based on the gene analysis value; The integrated gene unit is used to obtain the corresponding gene integrated information from the integrated table based on the target gene analysis interval value.

10. A genome comparison analysis terminal for multidrug-resistant Staphylococcus aureus, prophages, and schistophages, characterized in that, include: One or more processors; A storage device on which one or more programs are stored; When one or more programs are executed by one or more processors, the one or more processors implement the genome comparison analysis system for multidrug-resistant Staphylococcus aureus, prophages and schistosomes as described in any one of claims 1 to 9.