Method and system for analyzing genotype and key mutation site of chikungunya virus

By constructing high-coverage and low-coverage datasets and combining evolutionary diversity analysis and machine learning models, the problems of low resolution and insufficient efficiency in CHIKV genotyping were solved, enabling rapid and accurate genotyping of CHIKV genotypes and identification of key mutation sites, thus improving the reliability and robustness of genotyping results.

CN121565249AInactive Publication Date: 2026-02-24XIAN MEDICAL UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511745671.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-24
Publication Date
2026-02-24
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

In existing technologies, CHIKV genotyping and key mutation site analysis methods have limited genotyping resolution, making it difficult to accurately distinguish between IOL and sublineages such as AUL and AAL. The analysis efficiency is low, failing to meet the needs of refined genotyping, and the model robustness is insufficient, with genotyping accuracy dropping significantly when faced with low coverage data.

Method used

By acquiring CHIKV genome sequence data, high-coverage and low-coverage datasets were constructed. Non-conserved sites were identified by combining evolutionary diversity analysis. A position weight matrix model was constructed and genotyping was performed by combining machine learning models with phylogenetic trees. The contribution of amino acid sites to genotype classification was quantified using the SHAP interpretable machine learning framework, and key mutation sites were identified.

Benefits of technology

It enables rapid and accurate genotyping of CHIKV, improves genotyping resolution and efficiency, ensures high reliability and robustness of genotyping results, maintains high accuracy even with low coverage data, and supports virus tracing and vaccine development.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121565249A_ABST
    Figure CN121565249A_ABST
Patent Text Reader

Abstract

The invention relates to a chikungunya virus genotype and key mutation site analysis method and system.The chikungunya virus genotype and key mutation site analysis method comprises the steps that CHIKV genome sequence data is obtained, a high-coverage-rate data set and a low-coverage-rate data set are obtained through preprocessing, and evolutionary diversity analysis is conducted on the high-coverage-rate data set to determine non-conservative sites; constructing a position weight matrix model according to the high-coverage-rate data set and the non-conservative sites, and performing typing labeling on unlabeled gene sequences in the high / low-coverage-rate data set in combination with a machine learning model and a phylogenetic tree; constructing a CHIKV genotyping model through data set division, feature screening and model training optimization on the basis of the labeled high-coverage data set; based on a CHIKV genotyping model, key mutation sites are identified by quantifying the contribution degree of amino acid sites to genotype classification through an SHAP interpretable machine learning framework. According to the application, high-precision typing and key mutation site recognition of the chikungunya virus can be realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the interdisciplinary field of bioinformatics and artificial intelligence, and in particular to a method and system for analyzing the genotype and key mutation sites of chikungunya virus. Background Technology

[0002] Chikungunya fever is a global mosquito-borne infectious disease caused by the chikungunya virus (CHIKV). Clinical manifestations include acute fever and joint pain, which can develop into chronic polyarthritis and arthritis in severe cases, posing a persistent threat to global public health. As an RNA virus, CHIKV's RNA-dependent RNA polymerase lacks proofreading function, leading to a high mutation rate during transmission and the continuous evolution of new circulating strains. Based on its genetic evolutionary characteristics, CHIKV has been classified into several genotypes, including the East / Central / South African genotype (ECSA), the West African genotype (WA), the Asian genotype, and the Indian Ocean lineage (IOL) derived from ECSA. Adaptive mutations in genotypes such as IOL significantly enhance the virus's transmissibility and host adaptability, directly triggering multiple global outbreaks. Therefore, accurate genotyping and identification of key mutation sites in CHIKV are crucial for virus tracing, real-time epidemic monitoring, prevention and control strategy development, and vaccine development.

[0003] In the field of CHIKV genotyping and key mutation site analysis, some technical solutions exist, mainly including traditional phylogenetic tree construction methods and machine learning-based classification models. Traditional phylogenetic tree methods construct evolutionary relationships through sequence alignment to achieve genotyping, while some machine learning models construct classifiers by extracting sequence features to assist in genotyping. However, the genotyping resolution of existing technologies is limited. Most methods can only identify three major genotypes: ECSA, WA, and Asian, and it is difficult to accurately distinguish between IOL and sub-lineages such as AUL and AAL, failing to meet the needs of refined genotyping. Furthermore, the analysis efficiency is low. Traditional phylogenetic tree methods have high computational complexity and are time-consuming, making them unsuitable for the rapid analysis of large-scale genomic data and unable to support real-time epidemic monitoring. At the same time, the models lack robustness. When faced with incomplete sequence data such as low coverage, the genotyping accuracy drops significantly, making it difficult to adapt to situations with inconsistent data quality in practical applications.

[0004] Therefore, it is necessary to improve one or more of the problems existing in the above-mentioned related technical solutions.

[0005] It should be noted that the information disclosed in the background section above is only used to enhance the understanding of the background of this disclosure, and therefore may include information that does not constitute prior art known to those skilled in the art. Summary of the Invention

[0006] The purpose of this disclosure is to provide a method and system for analyzing the genotype and key mutation sites of chikungunya virus, thereby overcoming, to at least some extent, one or more problems caused by the limitations and defects of related technologies.

[0007] In a first aspect, this application provides a method for analyzing the genotype and key mutation sites of chikungunya virus, including: CHIKV genome sequence data were acquired, and high-coverage and low-coverage datasets were obtained through preprocessing. Evolutionary diversity analysis was then performed on the high-coverage datasets to identify non-conserved sites. Based on the high-coverage dataset and the non-conserved sites, a position weight matrix model is constructed, and combined with a machine learning model and a phylogenetic tree, the unlabeled gene sequences in the high / low-coverage dataset are genotyped and labeled. Based on the labeled high-coverage dataset, a CHIKV genotyping model was constructed through dataset partitioning, feature selection, and model training optimization. Based on the CHIKV genotyping model, the contribution of amino acid sites to genotype classification was quantified using the SHAP interpretable machine learning framework to identify key mutation sites.

[0008] In one possible implementation, the steps of acquiring CHIKV genome sequence data, obtaining high-coverage and low-coverage datasets through preprocessing, and performing evolutionary diversity analysis on the high-coverage dataset to determine non-conserved sites include: CHIKV genome sequence data were obtained from a gene database; some sequences in the CHIKV genome sequence data contain genotyping tags. Using the complete genome sequence as a reference, multiple sequence alignment was performed on the CHIKV genome sequence data to ensure consistent site coordinates; The 5' and 3' untranslated regions of the CHIKV genome sequence with consistent site coordinates were pruned to retain gene sequences containing complete coding regions, and the datasets were divided into high-coverage and low-coverage datasets. By introducing a site diversity index to quantify the degree of variation of each site in the high-coverage dataset, non-conservative sites are obtained.

[0009] In one possible implementation, after trimming the 5' and 3' untranslated regions of the CHIKV genome sequence with consistent site coordinates, retaining the gene sequence containing the complete coding region, the method further includes: Samples with less than 80% genome coverage were removed, and duplicate data and data with fewer than 5 entries for the same genotype were also removed.

[0010] In one possible implementation, the formula for calculating the site diversity index is: in, This is the diversity index for this site. For nucleotide / amino acid sites, It is a nucleotide / amino acid type. , Nucleotide / amino acid type The corresponding frequency, For site Nucleotide / Amino Acid Type The number of times it appears, This represents the total number of valid sequences at that site.

[0011] In one possible implementation, the step of constructing a position weight matrix model based on the high-coverage dataset and the non-conserved sites, and combining it with a machine learning model and a phylogenetic tree to perform genotyping annotation on unlabeled gene sequences within the high / low coverage dataset, includes: The high-coverage dataset is downsampled using a greedy algorithm based on maximizing sequence differences, so that each genotype in the high-coverage dataset retains at most 20 samples. The reference sequence set is obtained by iteratively selecting the sample with the largest average difference from the selected samples. Based on the non-conserved sites, the occurrence frequency of nucleotides in the high coverage dataset at the sites is calculated, and the occurrence frequency is converted into a log ratio relative to the background genome frequency. The position weight matrix is ​​then calculated to score the nucleotides at each non-conserved site, thus obtaining the position weight matrix model. The total matching score between the sequence to be tested and the position weight matrix of each genotype is calculated according to the position weight matrix model. The total matching score is converted into a probability distribution to obtain the first classification confidence score, which is the difference between the highest probability and the second highest probability. A random forest binary classification model is trained based on the high coverage dataset, and the parameters are optimized through 5-fold cross-validation. The predicted probability of the random forest binary classification model is used as the confidence score of the second classification. The unlabeled gene sequences in the high / low coverage dataset are genotyped and labeled using the first classification confidence score and the second classification confidence score, combined with the phylogenetic tree results.

[0012] In one possible implementation, the formula for calculating the frequency of occurrence is: in, Indicates site Nucleotide type Frequency of occurrence Indicates the target genotype. Indicates the nucleotide type, Indicates nucleotide site, Nucleotides At the site The count, This is a pseudo-count. Genotype The number of sequences in a high-coverage dataset; The formula for calculating the score of the nucleotide at each non-conservative site in the position weight matrix is ​​as follows: in, Indicates non-conserved sites Above, genotype Position weight matrix for nucleotides The score, Nucleotides The background frequency, i.e., the average probability of its occurrence in all sequences; The position weight matrix obtained for each genotype is one. The score matrix for a test sequence The above is related to genotype. The matching score is calculated using the following formula: in, Genotype The corresponding total matching score, The sequence length is given.

[0013] In one possible implementation, the step of typing and labeling unlabeled gene sequences in the high / low coverage dataset using the first classification confidence score and the second classification confidence score includes: If the confidence level of the first classification is higher than the first preset threshold, and the identification result is a high-discrimination genotype, the genotyping result is adopted; When the identification result is a low-discrimination genotype, the genotyping result is adopted if the confidence of the first category is higher than the first preset threshold and the confidence of the second category is greater than the second preset threshold. When neither of the above two conditions is met, i.e. the typing result cannot be adopted, a phylogenetic tree is constructed based on the maximum likelihood method and the adjacency method, and the phylogenetic tree constructed by the maximum likelihood method and the adjacency method is used. The typing result is adopted when the typing result of the phylogenetic tree constructed by the maximum likelihood method and the adjacency method is consistent; otherwise, the typing is marked as undetermined.

[0014] In one possible implementation, the steps of constructing a CHIKV genotyping model based on a labeled high-coverage dataset through dataset partitioning, feature selection, and model training optimization include: The labeled high-coverage dataset is stratified by genotype and randomly divided into a training set and a first independent test set, while the low-coverage dataset is used as the second independent test set. Three machine learning algorithms, Random Forest, XGBoost, and LightGBM, were used to construct preliminary classification models on the training set. The importance weight of each site was obtained, and sites with importance weights greater than zero in all three machine learning algorithms were selected as feature sites to obtain a feature set. The feature set includes two independent feature sets: a key nucleotide feature set and a key amino acid feature set. Based on the key nucleotide feature set and the key amino acid feature set respectively, the final classification models of three machine learning algorithms, Random Forest, XGBoost and LightGBM, are constructed through the training set; The three final genotyping models were tested using the first and second independent test sets to determine the CHIKV genotyping models for nucleotide and amino acid sequences.

[0015] In one possible implementation, the step of identifying key mutation sites by quantifying the contribution of amino acid sites to genotype classification using the SHAP interpretable machine learning framework based on the CHIKV genotyping model includes: The multi-genotype classification problem is transformed into a binary classification problem. For each target genotype, the genotype labels of the dataset are converted into two categories: the target genotype and other genotypes. By using key amino acid feature sets, an independent machine learning classifier is trained for each binary classification problem to obtain a classification model; The SHAP value of each sample in the training set at each amino acid site is calculated using a classification model. The absolute mean of all SHAP values ​​for each amino acid site is calculated and sorted in descending order to obtain the key mutation sites.

[0016] Secondly, this application provides an analysis system for chikungunya virus genotypes and key mutation sites, comprising: The data processing module acquires CHIKV genome sequence data, obtains high-coverage and low-coverage datasets through preprocessing, and performs evolutionary diversity analysis on the high-coverage datasets to identify non-conserved sites. The genotyping annotation module constructs a position weight matrix model based on the high-coverage dataset and the non-conserved sites, and combines a machine learning model with a phylogenetic tree to perform genotyping annotation on the unlabeled gene sequences in the high / low-coverage dataset. The genotyping module, based on a labeled high-coverage dataset, constructs the CHIKV genotyping model through dataset partitioning, feature selection, and model training optimization. The site identification module, based on the CHIKV genotyping model, uses the SHAP interpretable machine learning framework to quantify the contribution of amino acid sites to genotype classification and identify key mutation sites.

[0017] The technical solution provided in this application may include the following beneficial effects: This application presents a method and system for analyzing chikungunya virus genotypes and key mutation sites. The system's data preprocessing workflow filters high-quality sequences and divides datasets into high / low coverage categories. Combined with evolutionary diversity analysis, it identifies non-conserved sites, laying a data foundation for subsequent typing and site identification. Furthermore, by constructing a position weight matrix model, it achieves rapid initial screening of unlabeled samples. This is followed by a combination of a random forest binary classification model and phylogenetic analysis to complete fine-grained typing annotation, solving the problems of low typing resolution and long processing times associated with traditional methods. Simultaneously, based on the annotated dataset, a multi-algorithm high-precision typing model is constructed. The SHAP interpretability framework is used to quantify the contribution of amino acid sites to classification, accurately identifying key mutation sites and ensuring high reliability and robustness of the typing results.

[0018] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this disclosure. Attached Figure Description

[0019] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure. It is obvious that the drawings described below are merely some embodiments of this disclosure, and those skilled in the art can obtain other drawings based on these drawings without any inventive effort.

[0020] Figure 1 A flowchart illustrating the method for analyzing the genotype and key mutation sites of the chikungunya virus in an exemplary embodiment of this disclosure is shown. Figure 2 A detailed flowchart of step S100 of the method for analyzing the genotype and key mutation sites of Chikungunya virus in an exemplary embodiment of this disclosure is shown. Figure 3 A detailed flowchart of step S200 of the method for analyzing the genotype and key mutation sites of Chikungunya virus in an exemplary embodiment of this disclosure is shown. Figure 4 A detailed flowchart of step S300 of the method for analyzing the genotype and key mutation sites of Chikungunya virus in an exemplary embodiment of this disclosure is shown. Figure 5 A detailed flowchart of step S400 of the method for analyzing the genotype and key mutation sites of Chikungunya virus in an exemplary embodiment of this disclosure is shown. Figure 6 The diagram shows the distribution of CHIKV nucleotide site diversity using the analytical method for analyzing chikungunya virus genotypes and key mutation sites in an exemplary embodiment of this disclosure. Figure 7 The diagram shows the distribution of CHIKV amino acid site diversity using the analytical method for analyzing chikungunya virus genotypes and key mutation sites in an exemplary embodiment of this disclosure. Figure 8 The optimal model for analyzing chikungunya virus genotypes and key mutation sites in exemplary embodiments of this disclosure is based on a high-coverage test set (…). Confusion matrix diagram of test results; Figure 9 The optimal model for analyzing chikungunya virus genotypes and key mutation sites in exemplary embodiments of this disclosure is based on a low-coverage test set (…). Confusion matrix diagram of test results; Figure 10 A SHAP summary diagram of the top 30 most important features of the analysis method for chikungunya virus genotypes and key mutation sites in an exemplary embodiment of this disclosure is shown. Figure 11 This diagram illustrates a partial amino acid site SHAP feature dependency plot of the method for analyzing chikungunya virus genotypes and key mutation sites in an exemplary embodiment of this disclosure. Figure 12 This diagram illustrates the structure of a system for analyzing chikungunya virus genotypes and key mutation sites in an exemplary embodiment of this disclosure. Detailed Implementation

[0021] Exemplary embodiments will now be described more fully with reference to the accompanying drawings. However, these exemplary embodiments can be implemented in many forms and should not be construed as limited to the examples set forth herein; rather, they are provided so that this disclosure will be more comprehensive and complete, and will fully convey the concept of the exemplary embodiments to those skilled in the art. The described features, structures, or characteristics may be combined in any suitable manner in one or more embodiments.

[0022] Furthermore, the accompanying drawings are merely illustrative of this disclosure and are not necessarily drawn to scale. The same reference numerals in the drawings denote the same or similar parts, and therefore repeated descriptions of them will be omitted. Some block diagrams shown in the drawings are functional entities and do not necessarily correspond to physically or logically independent entities. These functional entities may be implemented in software, in one or more hardware modules or integrated circuits, or in different network and / or processor devices and / or microcontroller devices.

[0023] This example implementation first provides a method for analyzing the genotype and key mutation sites of chikungunya virus. This method can be applied to a terminal device, such as a mobile terminal like a mobile phone, desktop computer, personal digital assistant, laptop, tablet computer, or smartwatch. (Reference) Figure 1 As shown, the method may include the following steps: Step S100: Obtain CHIKV genome sequence data, obtain high-coverage dataset and low-coverage dataset through preprocessing, and perform evolutionary diversity analysis on the high-coverage dataset to determine non-conserved sites.

[0024] Step S200: Construct a position weight matrix model based on the high coverage dataset and the non-conserved sites, and combine it with a machine learning model and a phylogenetic tree to perform genotyping and labeling of unlabeled gene sequences in the high / low coverage dataset.

[0025] Step S300: Based on the labeled high-coverage dataset, construct the CHIKV genotyping model through dataset partitioning, feature selection, and model training optimization.

[0026] Step S400: Based on the CHIKV genotyping model, key mutation sites are identified by quantifying the contribution of amino acid sites to genotype classification using the SHAP interpretable machine learning framework.

[0027] The aforementioned method can construct high-coverage and low-coverage datasets through standardized data preprocessing, and identify non-conservative sites by combining evolutionary diversity analysis, laying a high-quality data foundation for subsequent typing and site identification. Through a hierarchical typing strategy of "initial screening using a positional weight matrix model + fine discrimination using a random forest binary classification model + phylogenetic tree verification," it achieves rapid and accurate annotation of eight CHIKV genotypes, overcoming the shortcomings of traditional methods such as low typing resolution and long processing times. By fusing multiple algorithms to screen key features and construct typing models, the RF model, validated on high and low coverage test sets, achieves an accuracy of 99.54% in high-coverage scenarios and 97.97% in low-coverage scenarios, demonstrating strong robustness and generalization ability. The SHAP interpretability framework quantifies the contribution of amino acid sites to typing, accurately identifies key mutation sites and visualizes their impact patterns, revealing the molecular genetic basis of genotyping and providing key technical support for virus tracing and vaccine development.

[0028] Below, we will refer to Figures 2 to 5 The steps of the method described above in this example embodiment will be explained in more detail.

[0029] In step S100, CHIKV genome sequence data is acquired, and high-coverage datasets and low-coverage datasets are obtained through preprocessing. Evolutionary diversity analysis is then performed on the high-coverage datasets to identify non-conserved sites.

[0030] In one embodiment, such as Figure 2 As shown, step S100 may include the following sub-steps: In step S110, CHIKV genome sequence data is obtained from a gene database; some sequences in the CHIKV genome sequence data contain genotyping tags.

[0031] For example, the complete genome sequence of Chikungunya virus (CHIKV) can be obtained from authoritative international databases, with a total of 1799 sequences acquired from the National Center for Biotechnology Information (NCBI) database. Of these, 1099 sequences can have their finely defined genotypes determined through phylogenetic analysis obtained via the Nextstrain platform, covering nine genotypes: MAL, WA, SAL, EAL, IOL, AAL, AUL, AUL-Am, and sECSA. 5087 sequences were also obtained from the Global Initiative on Sharing All Influenza Data (GISAID) database, which annotates three main genotypes: ECSA, WA, and Asian.

[0032] In step S120, a complete genome sequence is selected as a reference, and multiple sequence alignment is performed on the CHIKV genome sequence data to ensure consistent site coordinates.

[0033] It should be noted that after obtaining the raw sequences, the system's preprocessing workflow is executed to construct a high-quality dataset suitable for machine learning models. First, using MAFFT multiple sequence alignment software, with the complete genome sequence numbered MN974211.1 in the NCBI database as a reference, multiple sequence alignment is performed on all samples to ensure that the site coordinates of all sequences are consistent and comparable.

[0034] In step S130, the 5' and 3' untranslated regions of the CHIKV genome sequence with consistent site coordinates are pruned to retain the gene sequence containing the complete coding region, and the high-coverage dataset and the low-coverage dataset are divided.

[0035] Understandably, the untranslated regions at the 5' and 3' ends of the aligned sequence are trimmed, leaving the sequence containing the complete coding region.

[0036] Among them, after retaining gene sequences containing complete coding regions, samples with genome coverage of less than 80% were removed, and duplicate data and data with fewer than 5 data points of the same genotype were removed.

[0037] It should be noted that the methods for processing sequences containing complete coding regions include removing samples with less than 80% genome coverage to ensure the integrity of sequence information; removing duplicate samples based on sequence consistency to ensure sample independence; and removing genotypes with fewer than 5 samples (such as sECSA genotyping, which contains only 1 sample) to avoid model overfitting due to insufficient sample size.

[0038] In addition, to evaluate the robustness of the model in handling incomplete data in real-world applications, samples with coverage between 80% and 90% (n=111) were reserved as a separate low-coverage test set.

[0039] After the above preprocessing steps, two core datasets were finally constructed: High-coverage dataset: Composed of 1091 high-coverage (≥90%) samples from NCBI and 1923 high-coverage samples from GISAID, used for model training, optimization, testing, and interpretability analysis.

[0040] Low-coverage dataset: Contains 111 sequences, used to verify the model's generalization ability under conditions of poor data quality.

[0041] In step S140, the degree of variation of each site in the high coverage dataset is quantified by introducing a site diversity index to obtain non-conservative sites.

[0042] Understandably, to understand the genetic variation characteristics of CHIKV and provide mutation site feature inputs for subsequent machine learning model construction, evolutionary diversity analysis at the nucleotide and amino acid levels was performed on the preprocessed high-coverage dataset. Taking amino acid sequence diversity as an example: To quantify the degree of variation at each amino acid site, a site diversity index was introduced. The index is calculated as follows: for a specific amino acid site... ,set up The type of amino acid at this site The number of times it appears, This represents the total number of valid sequences at this site (excluding missing sites).

[0043] The formula for calculating the site diversity index is as follows: in, This is the diversity index for this site. For nucleotide / amino acid sites, It is a nucleotide / amino acid type. , Nucleotide / amino acid type The corresponding frequency, For site Nucleotide / Amino Acid Type The number of times it appears, This represents the total number of valid sequences at that site.

[0044] In step S200, a position weight matrix model is constructed based on the high coverage dataset and the non-conserved sites, and the unlabeled gene sequences in the high / low coverage dataset are genotyped and labeled by combining the machine learning model and the phylogenetic tree.

[0045] Understandably, in order to achieve rapid and accurate typing of a large number of unlabeled GISAID samples, a hierarchical classification process was designed, which integrates a probability matching model based on sequence patterns with a machine learning classifier, and uses phylogenetic analysis to finally determine the samples with low confidence, thus balancing efficiency and accuracy.

[0046] In one embodiment, such as Figure 3 As shown, step S200 may include the following sub-steps: In step S210, the high-coverage dataset is downsampled using a greedy algorithm based on maximizing sequence differences, so that each genotype in the high-coverage dataset retains at most 20 samples. The reference sequence set is obtained by iteratively selecting the sample with the largest average difference from the selected samples.

[0047] Understandably, to ensure the reliability of the genotyping criteria, sequences with a coverage rate greater than 95% were first screened from the NCBI high-coverage (≥90%) samples to construct a high-quality, highly diverse representative reference sequence set for the eight known genotypes. To avoid model bias caused by an excessive number of samples within a particular genotype, a greedy algorithm based on maximizing sequence differences is proposed to downsample genotypes with large sample sizes (keeping a maximum of 20 samples per genotype). The core of this algorithm is to iteratively select the sample with the largest average difference from the selected sample set, thereby ensuring... It can cover the genetic diversity within each genotype to a great extent.

[0048] In step S220, based on the non-conservative sites, the occurrence frequency of nucleotides in the high coverage dataset at the sites is calculated, and the occurrence frequency is converted into a logarithmic ratio relative to the background genome frequency. The position weight matrix is ​​then calculated to score the nucleotides at each non-conservative site, thus obtaining the position weight matrix model.

[0049] Understandably, based on NCBI's high-coverage data, position weight matrices (PWMs) were constructed for each of the eight genotypes. The PWMs were designed to broadly cover the sequence characteristics of each genotype, and their core principle is to quantify sequence characteristics by statistically analyzing the nucleotide distribution patterns of each genotype at multiple sites. The PWM construction process is as follows: Suppose a certain genotype have The aligned sequences were first selected based on the sequence evolutionary diversity analysis results, and non-conserved sites were screened from the full-length aligned sequences. >0), the length of the filtered sequence is denoted as For the site (1≤ ≤ The occurrence frequency of four nucleotides (A, C, G, T) was counted, ignoring deletions. A pseudo-count was introduced. (Set to 1) To avoid the zero probability problem, then nucleotide X∈{A,C,G,T} at the site frequency of occurrence The calculation formula is: ,in, .

[0050] The frequencies are converted to log-odds form, which is the log ratio relative to the background genome frequencies. Background frequencies By calculating nucleotides in the entire dataset The proportion is obtained. Then the genotype is... PWM at position Nucleotides Score for: This score quantifies the performance of a particular genotype. In the middle, site Nucleotides appear The relative probability. Performing the above calculations on each non-conserved locus of this genotype yields the size corresponding to that genotype. PWM matrix.

[0051] The formula for calculating the frequency of occurrence is as follows: in, Indicates site Nucleotide type Frequency of occurrence Indicates the target genotype. Indicates the nucleotide type, Indicates nucleotide site, Nucleotides At the site The count, It is a pseudo-count and is set to 1. Genotype The number of sequences in a high-coverage dataset; The formula for calculating the score of the nucleotide at each non-conservative site in the position weight matrix is ​​as follows: in, Indicates non-conserved sites Above, genotype Position weight matrix for nucleotides The score, Nucleotides The background frequency, i.e., the average probability of its occurrence in all sequences.

[0052] In step S230, the total matching score between the sequence to be tested and the position weight matrix of each genotype is calculated according to the position weight matrix model. The total matching score is converted into a probability distribution to obtain the first classification confidence score, which is the difference between the highest probability and the second highest probability.

[0053] Understandably, the process of using PWM for classification involves considering a test sequence. (Aligned to the same length as the reference sequence), treat it as a nucleotide vector. Calculate the relationship between this sequence and a given genotype. PWM matching total score : ,in, Indicates the location of the sequence to be tested. Nucleotides. If If the character is missing or ambiguous, the contribution of that locus is zero. Match scores between the test sequence and all eight genotype PWMs are calculated and normalized using a softmax function to convert the scores into probability distributions. Classification confidence. Defined as the difference between the highest and second-highest probabilities after normalization. A larger value indicates a more reliable classification result. By setting a confidence threshold, high-confidence classification results can be effectively distinguished from ambiguous cases requiring further verification.

[0054] The position weight matrix obtained for each genotype is a The score matrix for a test sequence The above is related to genotype. The matching score is calculated using the following formula: in, Genotype The corresponding total matching score, The sequence length is given.

[0055] In step S240, a random forest binary classification model is trained based on the high coverage dataset, and the parameters are optimized through 5-fold cross-validation. The predicted probability of the random forest binary classification model is used as the confidence level of the second classification.

[0056] Understandably, the analysis revealed that the PWM model had limited ability to distinguish between three pairs of genotypes with similar genetic backgrounds: AUL and AUL-Am, IOL and EAL, and AAL and MAL. Therefore, for each pair of easily confused genotypes, three independent random forest binary classification models were trained using corresponding samples from the NCBI high-coverage dataset. Parameters were optimized through 5-fold cross-validation, achieving 100% accuracy in internal validation. The classification confidence of the machine learning model (…) Then its predicted probability output is used.

[0057] In step S250, the unlabeled gene sequences in the high / low coverage dataset are genotyped and labeled using the first classification confidence and the second classification confidence, combined with the phylogenetic tree results.

[0058] Understandably, the following classification process is performed on unlabeled samples: Step 1: Perform initial screening using a PWM model. If the classification confidence level... Higher than the preset threshold ( If the result is a highly discriminative genotype such as WA or SAL, then the result is directly adopted.

[0059] Step 2: If the PWM model identifies one of the three pairs of easily confused genotypes mentioned above, then the corresponding random forest binary classification model is activated for fine-tuning. When the PWM confidence level... > And the confidence level of random forest > At that time, the results of machine learning models are adopted.

[0060] Step 3: For fuzzy samples whose confidence level did not reach the threshold in the above two steps, use MEGA11 software, based on the maximum likelihood method and the neighbor method, to compare them with the reference set. A phylogenetic tree is constructed simultaneously. The genotype of a sample is finally determined only if the genotyping conclusions of the phylogenetic trees constructed by the two methods are consistent; if the conclusions are inconsistent, the sample is marked as undetermined.

[0061] This hierarchical strategy enables rapid and automatic typing of most samples, while ensuring the reliability of typing of fuzzy samples through phylogenetic analysis, providing a high-quality dataset for subsequent research.

[0062] Furthermore, the method for genotyping and labeling unlabeled gene sequences within the high / low coverage dataset is as follows: If the confidence level of the first classification is higher than the first preset threshold, and the identification result is a high-discrimination genotype, the genotyping result is adopted; When the identification result is a low-discrimination genotype, the genotyping result is adopted if the confidence of the first category is higher than the first preset threshold and the confidence of the second category is greater than the second preset threshold. When neither of the above two conditions is met, i.e. the typing result cannot be adopted, a phylogenetic tree is constructed based on the maximum likelihood method and the adjacency method, and the phylogenetic tree constructed by the maximum likelihood method and the adjacency method is used. The typing result is adopted when the typing result of the phylogenetic tree constructed by the maximum likelihood method and the adjacency method is consistent; otherwise, the typing is marked as undetermined.

[0063] It is understandable that the two situations mentioned in "when the above two conditions are not met" refer to "if the confidence level of the first classification is higher than the first preset threshold and the identification result is a high-discrimination genotype, the genotyping result is adopted" and "when the identification result is a low-discrimination genotype, the confidence level of the first classification is higher than the first preset threshold and the confidence level of the second classification is greater than the second preset threshold, the genotyping result is adopted". When the genotyping result cannot be adopted in accordance with these two conditions, this method is used to perform genotyping or mark the genotyping as undetermined.

[0064] Among them, high-discrimination genotypes can be understood as genotypes that are easy to distinguish, that is, genotypes that are evolutionarily distant from other genotypes, specifically referring to the WA and SAL genotypes; Low-discrimination genotypes can be understood as genotypes that are not easy to distinguish, specifically referring to genotypes other than WA and SAL.

[0065] In step S300, based on the labeled high-coverage dataset, the CHIKV genotyping model is constructed through dataset partitioning, feature selection, and model training optimization.

[0066] Understandably, after obtaining a high-quality dataset with accurate genotype labels, this application employs various machine learning algorithms to construct a high-precision CHIKV genotyping model and systematically evaluates its performance.

[0067] In one embodiment, such as Figure 4 As shown, step S300 may include the following sub-steps: In step S310, the labeled high-coverage dataset is stratified by genotype and randomly divided into a training set and a first independent test set, and the low-coverage dataset is used as the second independent test set.

[0068] It should be noted that the high-coverage (≥90%) datasets from NCBI and GISAID were merged, and duplicate samples were removed to form the initial high-coverage sample set. To eliminate the potential bias caused by the imbalance of sample sizes among different genotypes to model training, a downsampling strategy based on maximizing sequence differences was adopted, setting the upper limit of the number of samples for each genotype to 200, resulting in a balanced high-coverage dataset. Subsequently, this balanced high-coverage dataset was stratified by genotype and randomly divided into the training set in a 6:4 ratio. ) and independent test set ( Training set Used for model training and parameter optimization. This is used to initially evaluate the model's generalization ability on high-coverage data. Secondly, the low-coverage (80%–90%) datasets from NCBI and GISAID are merged, deduplicated, and then a separate low-coverage test set is created. This low-coverage test set is used to test the robustness of the model when faced with incomplete sequence data.

[0069] In step S320, three machine learning algorithms, Random Forest, XGBoost, and LightGBM, are used to construct a preliminary classification model on the training set to obtain the importance weight of each site. Sites with importance weights greater than zero in all three machine learning algorithms are selected as feature sites to obtain a feature set. The feature set includes two independent sets: a key nucleotide feature set and a key amino acid feature set.

[0070] It should be noted that, in order to reduce feature dimensionality and select the features most discriminative for typing, model-based feature importance analysis was performed. The initial feature sets consisted of non-conserved nucleotide sites (for typing models based on nucleotide site features) and post-translational non-conserved amino acid sites (for typing models based on amino acid site features).

[0071] Three machine learning algorithms, Random Forest (RF), XGBoost (XGB), and LightGBM (LGB), were used on the training set. A preliminary classification model was constructed. After optimizing the core parameters of each model through k-fold cross-validation (k=5), a training model was built using the entire training set to obtain the importance weight of each feature (i.e., each locus). Feature importance reflects the degree to which the locus contributes to correct classification.

[0072] To integrate the advantages of different algorithms, a conservative strategy was adopted: the union of sites whose importance weights are all greater than zero in the RF, XGB, and LGB models was selected as the final feature subset. This process screened 1197 feature sites from the original nucleotide sites and 591 feature sites from the amino acid sites, significantly reducing feature dimensionality while maximizing feature diversity.

[0073] In step S330, based on the key nucleotide feature set and the key amino acid feature set, the final classification models of three machine learning algorithms, Random Forest, XGBoost and LightGBM, are constructed through the training set.

[0074] It should be noted that, based on the key nucleotide feature set (n=1197) and key amino acid feature set (n=591) selected above, the following methods were used: The training set was used to reconstruct the three final subtyping models: RF, XGB, and LGB. Each model underwent hyperparameter tuning again using k-fold cross-validation to ensure optimal performance.

[0075] In step S340, the three final genotyping models are tested using the first independent test set and the second independent test set to determine the CHIKV genotyping model for nucleotide and amino acid sequences.

[0076] It should be noted that the independently constructed test suite was used. (High coverage) and (Low Coverage) A systematic evaluation was conducted on all final models. To objectively measure the overall performance of the models in multi-class classification tasks and overcome the potential class imbalance in the test set, the weighted F1-score was selected as the core evaluation metric. This metric is the harmonic mean of precision and recall, and its weighted calculation takes into account the number of samples in each class, thus providing a fairer estimate of the overall performance of the model.

[0077] By comparing the weighted F1-score, accuracy, and confusion matrix of each model on two test sets, the optimal typing model for nucleotide sequences and amino acid sequences were determined, respectively. The latter was further applied to the model interpretability analysis.

[0078] In step S400, based on the CHIKV genotyping model, key mutation sites are identified by quantifying the contribution of amino acid sites to genotype classification using the SHAP interpretable machine learning framework.

[0079] It is understandable that the formation of viral evolutionary branches is closely related to amino acid mutations, which directly affect the functional properties of viral proteins. This application employs interpretable machine learning methods to identify amino acid mutation sites that contribute significantly to genotyping. Considering that amino acid mutations have a more significant and intuitive impact on viral phenotypes, this section is based on the aforementioned best-performing amino acid sequence genotyping model.

[0080] SHAP (SHapley Additive exPlanations) is an interpretable machine learning framework based on cooperative game theory. It decomposes the model's predictions into the sum of the contributions of each input feature, thereby quantifying the influence of each feature on a specific prediction. SHAP values ​​have a solid mathematical foundation and can simultaneously satisfy desirable properties such as local accuracy and consistency, providing reliable explanations for model decisions.

[0081] In one embodiment, such as Figure 5 As shown, step S400 may include the following sub-steps: In step S410, the multi-genotype classification problem is transformed into a binary classification problem. For each target genotype, the genotype labels of the dataset are converted into two categories: the target genotype and other genotypes.

[0082] Understandably, the 8-class classification problem is broken down into a series of binary classification problems. For each target genotype, the dataset labels are converted into two categories: target genotype and other genotypes.

[0083] In step S420, an independent machine learning classifier is trained for each binary classification problem using the key amino acid feature set to obtain a classification model.

[0084] Understandably, an independent machine learning classifier is trained for each binary classification problem using an amino acid feature training set.

[0085] In step S430, the SHAP value of each sample in the training set at each amino acid site is calculated using a classification model.

[0086] Understandably, the trained model is used to calculate the SHAP value for each sample in the training set at each amino acid site. The sign of the SHAP value indicates whether the amino acid mutation at that site promotes or inhibits the predicted target genotype, and the absolute value reflects the degree of influence.

[0087] The SHAP method interprets the model prediction value of a single sample as the sum of the contributions of all features, and its calculation method is as follows: in, The baseline value is the predicted average of all samples. Indicates the first Each feature for a sample The contribution value.

[0088] In step S440, the absolute mean of all SHAP values ​​for each amino acid site is calculated and sorted in descending order to obtain the key mutation sites.

[0089] Understandably, for each amino acid site, the absolute mean of the SHAP values ​​of all samples is calculated as an indicator of the importance of that site in distinguishing the target genotype, and the samples are arranged in descending order of this importance. The results can also be visualized: a SHAP summary plot comprehensively shows the overall impact of each amino acid site on the model output, intuitively presenting the distribution and contribution direction of important sites; a feature dependency plot provides an in-depth analysis of the relationship between different amino acid types at key sites and SHAP values, revealing the specific impact patterns of specific mutations on genotyping.

[0090] Furthermore, such as Figure 12 As shown, this application also provides an analysis system for chikungunya virus genotypes and key mutation sites, including: The data processing module acquires CHIKV genome sequence data, obtains high-coverage and low-coverage datasets through preprocessing, and performs evolutionary diversity analysis on the high-coverage datasets to identify non-conserved sites. The genotyping annotation module constructs a position weight matrix model based on the high-coverage dataset and the non-conserved sites, and combines a machine learning model with a phylogenetic tree to perform genotyping annotation on the unlabeled gene sequences in the high / low-coverage dataset. The genotyping module, based on a labeled high-coverage dataset, constructs the CHIKV genotyping model through dataset partitioning, feature selection, and model training optimization. The site identification module, based on the CHIKV genotyping model, uses the SHAP interpretable machine learning framework to quantify the contribution of amino acid sites to genotype classification and identify key mutation sites.

[0091] Furthermore, the analytical methods for the chikungunya virus genotype and key mutation sites of this application were tested and analyzed, including: sequence evolutionary diversity analysis, model testing, and identification of key mutation sites.

[0092] Results of sequence evolutionary diversity analysis: like Figure 6 As shown, this figure illustrates the distribution of genetic diversity at the CHIKV nucleotide level. The horizontal axis represents nucleotide sites, with the names of different gene coding regions labeled at the bottom. The vertical axis represents the nucleotide site diversity index (…). Highly diverse sites are highlighted and their location information is labeled. Furthermore, a violin plot overlaid on the diversity map shows the distribution of nucleotide site diversity values ​​within each protein region.

[0093] like Figure 7 As shown in the figure, this plot illustrates the genetic diversity distribution of CHIKV amino acid levels. The horizontal axis represents amino acid sites, with the names of different gene coding regions labeled at the bottom. The vertical axis represents the amino acid site diversity index. Highly diverse sites are highlighted and their location information is labeled. Furthermore, a violin plot overlaid on the diversity map shows the distribution of amino acid site diversity values ​​within each protein region.

[0094] Model test results: First, the model was tested on a high-coverage test set, and the confusion matrix of the test results is as follows: Figure 8As shown, the vertical axis of the confusion matrix represents the true genotype label of the sample, and the horizontal axis represents the predicted genotype label of the model. Each value in the matrix represents the proportion of samples whose corresponding true category was predicted to be in each category. The results show that the model has high classification accuracy for all genotypes, with only a very small number of misclassifications observed between closely related AUL and AUL-Am lineages.

[0095] The test results are shown in Table 1: Table 1 is based on a high-coverage test set ( Model test results As can be seen from Table 1, all three machine learning methods achieved high test accuracy, with F1-scores all above 99%. The random forest method achieved the highest test accuracy, with an F1-score of 99.52%.

[0096] Secondly, the model was tested on a low-coverage test set, and the confusion matrix of the optimal model test results is as follows: Figure 9 As shown, the vertical axis of the confusion matrix represents the true genotype label of the sample, and the horizontal axis represents the predicted genotype label of the model. Each value in the matrix represents the proportion of samples whose corresponding true category is predicted to be in each category. The model can correctly classify most low-coverage sequences, with only a portion (30%) of AUL-Am samples being misclassified as AUL.

[0097] The test results are shown in Table 2: Table 2 is based on the low coverage test set ( Model test results As can be seen from Table 2, consistent with the test results on high coverage data, the random forest classifier achieved the highest test accuracy, with an F1-score of 96.5%, indicating that it has good generalization ability for incomplete sequence data.

[0098] Key mutation site identification results: like Figure 10As shown, this figure is a SHAP summary diagram of the representative CHIKV genotypes IOL and AUL. The figure visualizes the top 30 features of the IOL and AUL genotypes, sorted by mean absolute SHAP value. Each point represents an independent sample, and the horizontal position of the point indicates the SHAP value, which quantifies the influence of the corresponding amino acid feature on the prediction result of the target genotype. Positive values ​​increase the probability of the sample being identified as the target genotype, while negative values ​​decrease this probability. The vertical axis displays amino acid sites arranged from highest to lowest feature importance. Analysis revealed that key sites that positively drive IOL genotype classification include nsP2_54, C_27, E2_264, and E1_211. For the AUL genotype, key discriminant sites include 6K_20, E2_368, nsP2_338, E3_19, and C_37.

[0099] like Figure 11 As shown in the figure, this plot illustrates the SHAP feature dependence of representative amino acid sites for the IOL and AUL genotypes. In the figure, the horizontal axis represents amino acid type, and the vertical axis displays the SHAP value (dashed lines indicate zero values). Each sample is colored according to its SHAP value contribution (blue: negative, red: positive), and the distribution of SHAP values ​​corresponding to each amino acid type is shown through overlaid box plots and violin plots. Taking E2_264 as an example, the 264th amino acid of the E2 protein, type A, has a positive effect on IOL genotyping, while the sequence prediction results for amino acid type V at this site tend to favor non-IOL genotyping.

[0100] In the description of this disclosure, it should be understood that the terms center, longitudinal, transverse, length, width, thickness, upper, lower, front, back, left, right, vertical, horizontal, top, bottom, inner, outer, clockwise, counterclockwise, etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. They are only for the convenience of describing this disclosure and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on this disclosure.

[0101] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of this disclosure, "multiple" means two or more, unless otherwise explicitly specified.

[0102] In this disclosure, unless otherwise expressly specified and limited, the terms installation, connection, linking, fixing, etc., should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral part; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; they can refer to the internal communication of two components or the interaction between two components. Those skilled in the art can understand the specific meaning of the above terms in this disclosure according to the specific circumstances.

[0103] In this disclosure, unless otherwise expressly specified and limited, "above or below" a first feature may include direct contact between the first and second features, or contact between the first and second features through another feature between them. Furthermore, "above," "over," and "on" a first feature includes the first feature directly above or diagonally above the second feature, or simply indicates that the first feature is at a higher horizontal level than the second feature. "Below," "below," and "under" a first feature includes the first feature directly below or diagonally below the second feature, or simply indicates that the first feature is at a lower horizontal level than the second feature.

[0104] In the description of this specification, the references to the terms "an embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of this disclosure. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. In addition, those skilled in the art can combine and integrate the different embodiments or examples described in this specification.

[0105] Other embodiments of this disclosure will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this disclosure are indicated by the appended claims.

Claims

1. A method for analyzing the genotype and key mutation sites of chikungunya virus, characterized in that, include: CHIKV genome sequence data were acquired, and high-coverage and low-coverage datasets were obtained through preprocessing. Evolutionary diversity analysis was then performed on the high-coverage datasets to identify non-conserved sites. Based on the high-coverage dataset and the non-conserved sites, a position weight matrix model is constructed, and combined with a machine learning model and a phylogenetic tree, the unlabeled gene sequences in the high / low-coverage dataset are genotyped and labeled. Based on the labeled high-coverage dataset, a CHIKV genotyping model was constructed through dataset partitioning, feature selection, and model training optimization. Based on the CHIKV genotyping model, the contribution of amino acid sites to genotype classification was quantified using the SHAP interpretable machine learning framework to identify key mutation sites.

2. The method for analyzing the genotype and key mutation sites of chikungunya virus according to claim 1, characterized in that, The steps of acquiring CHIKV genome sequence data, obtaining high-coverage and low-coverage datasets through preprocessing, and performing evolutionary diversity analysis on the high-coverage dataset to determine non-conserved sites include: CHIKV genome sequence data were obtained from a gene database; some sequences in the CHIKV genome sequence data contain genotyping tags. Using the complete genome sequence as a reference, multiple sequence alignment was performed on the CHIKV genome sequence data to ensure consistent site coordinates; The 5' and 3' untranslated regions of the CHIKV genome sequence with consistent site coordinates were pruned to retain gene sequences containing complete coding regions, and the datasets were divided into high-coverage and low-coverage datasets. By introducing a site diversity index to quantify the degree of variation of each site in the high-coverage dataset, non-conservative sites are obtained.

3. The method for analyzing the genotype and key mutation sites of chikungunya virus according to claim 2, characterized in that, After trimming the 5' and 3' untranslated regions of the CHIKV genome sequence with consistent site coordinates, retaining the gene sequence containing the complete coding region, the method further includes: Samples with less than 80% genome coverage were removed, and duplicate data and data with fewer than 5 entries for the same genotype were also removed.

4. The method for analyzing the genotype and key mutation sites of chikungunya virus according to claim 2, characterized in that, The formula for calculating the site diversity index is as follows: in, This is the diversity index for this site. For nucleotide / amino acid sites, It is a nucleotide / amino acid type. , Nucleotide / amino acid type The corresponding frequency, For site Nucleotide / Amino Acid Type The number of times it appears, This represents the total number of valid sequences at that site.

5. The method for analyzing the genotype and key mutation sites of chikungunya virus according to claim 1, characterized in that, The step of constructing a position weight matrix model based on the high-coverage dataset and the non-conserved sites, and combining it with a machine learning model and a phylogenetic tree to perform genotyping annotation on unlabeled gene sequences in the high / low coverage dataset, includes: The high-coverage dataset is downsampled using a greedy algorithm based on maximizing sequence differences, so that each genotype in the high-coverage dataset retains at most 20 samples. The reference sequence set is obtained by iteratively selecting the sample with the largest average difference from the selected samples. Based on the non-conserved sites, the occurrence frequency of nucleotides in the high coverage dataset at the sites is calculated, and the occurrence frequency is converted into a log ratio relative to the background genome frequency. The position weight matrix is ​​then calculated to score the nucleotides at each non-conserved site, thus obtaining the position weight matrix model. The total matching score between the sequence to be tested and the position weight matrix of each genotype is calculated according to the position weight matrix model. The total matching score is converted into a probability distribution to obtain the first classification confidence score, which is the difference between the highest probability and the second highest probability. A random forest binary classification model is trained based on the high coverage dataset, and the parameters are optimized through 5-fold cross-validation. The predicted probability of the random forest binary classification model is used as the confidence score of the second classification. The unlabeled gene sequences in the high / low coverage dataset are genotyped and labeled using the first classification confidence score and the second classification confidence score, combined with the phylogenetic tree results.

6. The method for analyzing the genotype and key mutation sites of chikungunya virus according to claim 5, characterized in that, The formula for calculating the frequency of occurrence is: in, Indicates site Nucleotide type Frequency of occurrence Indicates the target genotype. Indicates the nucleotide type, Indicates nucleotide site, Nucleotides At the site The count, This is a pseudo-count. Genotype The number of sequences in a high-coverage dataset; The formula for calculating the score of the nucleotide at each non-conservative site in the position weight matrix is ​​as follows: in, Indicates non-conserved sites Above, genotype Position weight matrix for nucleotides The score, Nucleotides The background frequency, i.e., the average probability of its occurrence in all sequences; The position weight matrix obtained for each genotype is one. The score matrix for a test sequence The above is related to genotype. The matching score is calculated using the following formula: in, Genotype The corresponding total matching score, The sequence length is given.

7. The method for analyzing the genotype and key mutation sites of chikungunya virus according to claim 5, characterized in that, The step of performing genotyping annotation on unlabeled gene sequences in the high / low coverage dataset using the first classification confidence score and the second classification confidence score includes: If the confidence level of the first classification is higher than the first preset threshold, and the identification result is a high-discrimination genotype, the genotyping result is adopted; When the identification result is a low-discrimination genotype, the genotyping result is adopted if the confidence of the first category is higher than the first preset threshold and the confidence of the second category is greater than the second preset threshold. When neither of the above two conditions is met, i.e. the typing result cannot be adopted, a phylogenetic tree is constructed based on the maximum likelihood method and the adjacency method, and the phylogenetic tree constructed by the maximum likelihood method and the adjacency method is used. The typing result is adopted when the typing result of the phylogenetic tree constructed by the maximum likelihood method and the adjacency method is consistent; otherwise, the typing is marked as undetermined.

8. The method for analyzing the genotype and key mutation sites of chikungunya virus according to claim 1, characterized in that, The steps for constructing the CHIKV genotyping model based on the labeled high-coverage dataset, through dataset partitioning, feature selection, and model training optimization, include: The labeled high-coverage dataset is stratified by genotype and randomly divided into a training set and a first independent test set, while the low-coverage dataset is used as the second independent test set. Three machine learning algorithms, Random Forest, XGBoost, and LightGBM, were used to construct preliminary classification models on the training set. The importance weight of each site was obtained, and sites with importance weights greater than zero in all three machine learning algorithms were selected as feature sites to obtain a feature set. The feature set includes two independent sets: a key nucleotide feature set and a key amino acid feature set. Based on the key nucleotide feature set and the key amino acid feature set respectively, the final classification models of three machine learning algorithms, RandomForest, XGBoost and LightGBM, are constructed through the training set; The three final genotyping models were tested using the first and second independent test sets to determine the CHIKV genotyping models for nucleotide and amino acid sequences.

9. The method for analyzing the genotype and key mutation sites of chikungunya virus according to claim 1, characterized in that, The steps for identifying key mutation sites based on the CHIKV genotyping model, using the SHAP interpretable machine learning framework to quantify the contribution of amino acid sites to genotype classification, include: The multi-genotype classification problem is transformed into a binary classification problem. For each target genotype, the genotype labels of the dataset are converted into two categories: the target genotype and other genotypes. By using key amino acid feature sets, an independent machine learning classifier is trained for each binary classification problem to obtain a classification model; The SHAP value of each sample in the training set at each amino acid site is calculated using a classification model. The absolute mean of all SHAP values ​​for each amino acid site is calculated and sorted in descending order to obtain the key mutation sites.

10. A system for analyzing the genotype and key mutation sites of chikungunya virus, characterized in that, include: The data processing module acquires CHIKV genome sequence data, obtains high-coverage and low-coverage datasets through preprocessing, and performs evolutionary diversity analysis on the high-coverage datasets to identify non-conserved sites. The genotyping annotation module constructs a position weight matrix model based on the high-coverage dataset and the non-conserved sites, and combines a machine learning model with a phylogenetic tree to perform genotyping annotation on the unlabeled gene sequences in the high / low-coverage dataset. The genotyping module, based on a labeled high-coverage dataset, constructs the CHIKV genotyping model through dataset partitioning, feature selection, and model training optimization. The site identification module, based on the CHIKV genotyping model, uses the SHAP interpretable machine learning framework to quantify the contribution of amino acid sites to genotype classification and identify key mutation sites.