Artificial Intelligence-Based Diagnostic Recognition System for Subtype Equine Encephalomyelitis Virus
Through the AI-based subtype equine encephalomyelitis virus diagnostic and identification system, the parallel processing of K-mer analysis and machine learning models is used to solve the problems of low data processing efficiency and improper model selection in the existing technology, and efficient and accurate virus subtype recognition is achieved.
Patent Information
- Application Number
- CN202510382848.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-28
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-03-28
AI Technical Summary
In the diagnosis and recognition of subtype equine encephalomyelitis viruses, genomic data and case data processing efficiency is inefficient, and improper selection of recognition models leads to poor recognition effect.
Using an artificial intelligence-based subtype equine encephalomyelitis virus diagnostic identification system, subset division module, feature extraction and aggregation module, classification module and evaluation module, K-mer analysis is used to extract mutation pattern feature data from genomic data, and clinical feature data is extracted from case data, and parallel processing and cluster evaluation are carried out to select the optimal machine training model.
The data processing efficiency is improved, the accuracy and efficiency of the identification system are ensured, and the optimal identification model is selected to meet the needs of large-scale data processing.
Smart Images

Figure CN119889735B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of virus recognition systems, and particularly to a diagnosis and recognition system for subtype equine encephalomyelitis virus based on artificial intelligence. Background Art
[0002] Subtype equine encephalomyelitis virus (abbreviated as EEV) is a disease caused by equine encephalomyelitis virus, belonging to the Flaviviridae family and being a type of Encephalitis-Virus. This virus is mainly transmitted by mosquitoes and is a pathogenic factor for horses and other animals, and can also pose a certain threat to humans. The recognition system is used to detect and identify different subtypes of equine encephalomyelitis virus, especially in the rapid response during the outbreak of the epidemic and in high-risk areas, which has important practical significance.
[0003] The existing technologies have the following deficiencies:
[0004] 1. In the diagnosis and recognition of virus subtypes, genomic data and case data are often huge and complex. Traditional data processing methods usually rely on single-machine computing and cannot efficiently process large-scale genomic data and case data, resulting in low data processing efficiency. Especially when faced with massive data, the computing bottleneck and memory limit will seriously affect the speed and accuracy of diagnosis and recognition;
[0005] 2. The existing recognition systems usually select the recognition model randomly or by simple index evaluation. Random selection, due to its strong randomness, cannot guarantee the optimal recognition efficiency and recognition effect of the recognition system. And for simple index evaluation, first, due to the different operating performances of different computing nodes, using simple index evaluation may lead to the lack of evaluation fairness, and second, the comprehensiveness of simple index evaluation is poor, resulting in the recognition system being unable to select the best recognition model for application.
[0006] Based on this, the present invention proposes a diagnosis and recognition system for subtype equine encephalomyelitis virus based on artificial intelligence, enabling each subset to be processed in parallel on independent computing nodes, effectively improving the data processing efficiency, and after clustering processing considering the performance of the computing nodes, evaluating the classification effect of the clusters to select the optimal machine training model for application. Summary of the Invention
[0007] The purpose of the present invention is to provide a diagnosis and recognition system for subtype equine encephalomyelitis virus based on artificial intelligence to solve the deficiencies in the background art.
[0008] To achieve the above purpose, the present invention provides the following technical solution: A diagnosis and recognition system for subtype equine encephalomyelitis virus based on artificial intelligence, including a subset division module, a feature extraction and aggregation module, a classification module, and an evaluation module;
[0009] Subset Partitioning Module: Collect genomic data and case data to construct a dataset. After preprocessing the dataset, divide it into multiple subsets;
[0010] Feature Extraction and Aggregation Module: Use K-mer analysis to extract mutant pattern feature data from genomic data, and extract clinical feature data from case data. Aggregate the mutant pattern feature data and clinical feature data into sample data;
[0011] Classification Module: Distribute multiple subset sample data to be classified to each computing node for parallel processing. Each computing node uses the trained machine learning model for classification to obtain the classification results of each sample data;
[0012] Evaluation Module: Perform clustering processing based on the running status of each computing node, evaluate the classification effects of each cluster, and then select the optimal machine training model for application.
[0013] In a preferred embodiment, the Feature Extraction and Aggregation Module uses the K-mer algorithm to cut each genomic sequence in the genomic data to generate K-mer subsequences, and analyzes the mutant characteristics in the genomic data based on the K-mer subsequences;
[0014] Collect case data information, perform standardization or normalization processing on numerical data, and convert categorical data into numerical data through One-Hot encoding or Label encoding;
[0015] Encode clinical symptoms, use text analysis technology to extract standardized symptom categories or severity levels from symptom descriptions, combine with clinical diagnosis results, extract disease progression or complication characteristics, and extract biomarkers or clinical trial results from laboratory test data;
[0016] Match and aggregate genomic data and case data based on patient identifiers or virus subtypes. After integrating the genomic data, clinical data, and symptom information of each patient, a sample dataset is obtained.
[0017] In a preferred embodiment, the Feature Extraction and Aggregation Module obtains the number of partitions of K-mer subsequences in the genomic sequence, and obtains the genotype difference value, Huffman coding length, and Cosine similarity of each K-mer subsequence;
[0018] Calculate and obtain the mutant factor of the genomic sequence based on the genotype difference value, Huffman coding length, and Cosine similarity;
[0019] Compare the obtained mutation factor with a preset mutation threshold. The mutation threshold is used to determine whether there is a gene mutation in the genomic sequence. If the mutation factor is less than or equal to the mutation threshold, it is determined that there is no gene mutation in the genomic sequence. If the mutation factor is greater than the mutation threshold, it is determined that there is a gene mutation in the genomic sequence.
[0020] In a preferred embodiment, the evaluation module obtains the real-time temperature, temperature rise rate, and voltage standard deviation of each computing node, performs normalization processing on the real-time temperature, temperature rise rate, and voltage standard deviation, maps the value ranges of the real-time temperature, temperature rise rate, and voltage standard deviation to between [0, 1], and obtains the operation index of the computing node by adding the normalized real-time temperature, temperature rise rate, and voltage standard deviation;
[0021] Suppose all computing nodes need to be divided into K clusters. First, randomly select the operation indices of K computing nodes as the cluster center values of the K clusters, and calculate the absolute values of the differences between the remaining computing nodes and the cluster center values of the K clusters. Then, assign the computing nodes to the cluster with the smallest absolute value of the difference. After all computing nodes are assigned, recalculate the average operation index of each cluster, use the average operation index as the cluster center value of the cluster, and repeat the clustering step. When the difference between the cluster center value obtained in any iteration and the cluster center value obtained in the previous iteration is less than 0.1, it is determined that the convergence condition is met, and K clusters are output;
[0022] Obtain the classification effect coefficient of each computing node in the cluster. The classification effect coefficient is obtained by adding the normalized classification speed and classification accuracy, and calculate the average value of the classification effect coefficients of all clusters. Select the cluster with the largest classification effect coefficient, and obtain the machine learning model used by each computing node in this cluster. Use the machine learning model with the largest usage quantity as the optimal machine training model for application.
[0023] In a preferred embodiment, the classification module assigns subsets to each computing node, and the computing node uses a pre-trained machine learning model to perform classification tasks. The machine learning model includes a decision tree model, a support vector machine model, a random forest model, and a neural network model;
[0024] After each computing node finishes processing the subset assigned to it, it outputs the classification result of the subset. The classification result of each subset is a label. After all computing nodes complete the classification tasks, the obtained classification results are summarized.
[0025] In a preferred embodiment, the calculation expression of the Huffman coding length is:
[0026] , where is the Huffman coding length, is the number of K-mer subsequences of the genomic sequence, represents the frequency of the i-th K-mer subsequence in the genomic data, represents the bit length occupied by the Huffman code of the i-th K-mer.
[0027] In a preferred embodiment, the calculation expression of the Cosine similarity is: , where, is the Cosine similarity, is the frequency vector of a K-mer subsequence in the genomic sequence, is the frequency vector of another K-mer subsequence in the genomic sequence, is the frequency vector and the frequency vector is the dot product of, are respectively the norms of the frequency vector and the norm of the frequency vector of.
[0028] In a preferred embodiment, the calculation expression of the genotype difference value is: , where, is the genotype difference value, is the number of K-mer subsequences of the genomic sequence, is the occurrence frequency of the sample data on the i-th K-mer subsequence, is the occurrence frequency of the reference genome on the i-th K-mer subsequence.
[0029] In a preferred embodiment, the subset partitioning module collects genomic data and case data, collects genomic data from different virus samples, each virus sample contains the gene sequence of the virus, obtains the genomic data through the genomic sequencing results of the virus or a public gene database, and the data format is FASTQ or FASTA genomic data format;
[0030] Collects case data related to patients, including patient information, clinical symptoms, diagnostic information, and laboratory test results, obtains case data through patient electronic health records, medical databases, or laboratory reports, and the data format is CSV, JSON, XML structured or semi-structured format;
[0031] The genomic data uses a standard format to save sequence information, the case data is unified into a table form, the numerical features are standardized or normalized, and the text data is encoded and processed into numerical data.
[0032] In a preferred embodiment, the subset partitioning module partitions the data set in ways including partitioning by virus subtype, partitioning by patient group, balancing data distribution, or parallel processing partitioning.
[0033] In the above technical solution, the technical effects and advantages provided by the present invention are as follows:
[0034] The present invention extracts mutation pattern feature data from genomic data through the feature extraction and aggregation module using K-mer analysis, extracts clinical feature data from case data, aggregates the mutation pattern feature data and the clinical feature data into sample data, the classification module distributes multiple subset sample data to be classified to each computing node for parallel processing, each computing node uses the trained machine learning model for classification, the evaluation module performs clustering processing according to the running state of each computing node, and after evaluating the classification effects of each cluster, selects the optimal machine training model for application. The recognition system enables each subset to be processed in parallel on an independent computing node, effectively improving the data processing efficiency, and after clustering processing considering the performance of the computing nodes, evaluating the classification effects of the clusters to select the optimal machine training model for application. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments recorded in the present invention. For those of ordinary skill in the art, other drawings can also be obtained based on these drawings.
[0036] Figure 1 It is a flowchart of the method of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0037] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0038] Embodiment 1: Please refer to Figure 1 As shown, the artificial intelligence-based subtype equine encephalomyelitis virus diagnosis and recognition system in this embodiment includes a subset partitioning module, a feature extraction and aggregation module, a classification module, and an evaluation module;
[0039] Subset Partitioning Module: Collect genomic data and case data to construct a dataset. After preprocessing the dataset (including data formatting and data cleaning), the dataset is divided into multiple subsets for parallel processing. The dataset can be segmented by virus subtype, patient group, etc. The subset partitioning results are sent to the Feature Extraction and Aggregation Module and the Classification Module;
[0040] Feature Extraction and Aggregation Module: Use K-mer analysis to extract mutation pattern feature data from genomic data and extract clinical feature data from case data. Aggregate the mutation pattern feature data and clinical feature data into sample data. For example, genomic data and case data may be stored separately, and they are aggregated according to virus subtype or patient identifier to ensure that all relevant information (genomic, clinical, symptoms, etc.) of each patient can be combined. The sample data of each subset is sent to the Classification Module;
[0041] Classification Module: Distribute multiple subset sample data to be classified to each computing node for parallel processing. Each computing node uses a trained machine learning model for classification to obtain the classification results of each sample data. The classification results are sent to the Evaluation Module;
[0042] Evaluation Module: Perform clustering processing based on the running status of each computing node, and evaluate the classification effects of each cluster, and then select the optimal machine training model for application.
[0043] In this application, the Feature Extraction and Aggregation Module uses K-mer analysis to extract mutation pattern feature data from genomic data and extract clinical feature data from case data, and aggregates the mutation pattern feature data and clinical feature data into sample data. The Classification Module distributes multiple subset sample data to be classified to each computing node for parallel processing. Each computing node uses a trained machine learning model for classification. The Evaluation Module performs clustering processing based on the running status of each computing node, and evaluates the classification effects of each cluster, and then selects the optimal machine training model for application. The recognition system enables each subset to be processed in parallel on an independent computing node, effectively improving data processing efficiency. And after considering the performance of the computing nodes for clustering processing, it evaluates the classification effects of the clusters to select the optimal machine training model for application.
[0044] The working process of the recognition system is as follows:
[0045] The recognition system collects genomic data and case data to construct a dataset. After preprocessing the dataset (including data formatting and data cleaning), the dataset is divided into multiple subsets for parallel processing. The dataset can be segmented by virus subtype, patient group, etc. K-mer analysis is used to extract mutant pattern feature data from the genomic data, and clinical feature data is extracted from the case data. The mutant pattern feature data and clinical feature data are aggregated into sample data. For example, the genomic data and case data may be stored separately, and they are aggregated according to the virus subtype or patient identifier to ensure that all relevant information (genomic, clinical, symptoms, etc.) of each patient can be combined. Then, multiple subset sample data to be classified are distributed to each computing node for parallel processing. Each computing node uses the trained machine learning model for classification to obtain the classification result of each sample data. After clustering processing based on the running status of each computing node and evaluating the classification effects of each cluster, the optimal machine training model is selected for application.
[0046] Example 2: The subset division module collects genomic data and case data to construct a dataset. After preprocessing the dataset (including data formatting and data cleaning), the dataset is divided into multiple subsets for parallel processing. The dataset can be segmented by virus subtype, patient group, etc.;
[0047] The subset division module collects genomic data and case data;
[0048] Genomic data collection: Collect genomic data from different virus samples. Each sample contains the gene sequence of the virus, which may involve large-scale genomic information (such as thousands of gene markers, gene data of multiple samples). The data is obtained through the genomic sequencing results of the virus or public gene databases. And the data format is genomic data formats such as FASTQ and FASTA.
[0049] Case data collection: Collect case data related to patients, including basic information, clinical symptoms, diagnostic information, laboratory test results, etc. of the patients. The data is obtained through patient electronic health records (EHRs), medical databases, or laboratory reports. And the data format is structured or semi-structured formats such as CSV, JSON, and XML.
[0050] Genomic data and case data are often stored separately and need to be integrated together through patient identifiers or virus subtype identifiers to form a complete sample dataset.
[0051] Unify the formats of data from different sources (genomic data and case data) for subsequent processing. For genomic data, a standard format (such as FASTA) can be used to save sequence information; for case data, it can be unified into a table form to ensure that the field formats of each record are consistent. Detect and handle missing values in the data. For example, some fields in the case data may be missing, and some sequence information in the genomic data may be incomplete. Impute or delete the missing values to ensure the integrity of the data. Delete duplicate data entries to avoid the same data participating in model training repeatedly. Identify and correct outliers in the data. For example, there are extreme values in the patient's age field, or abnormal mutation markers that do not conform to biological laws appear in the genomic data.
[0052] Standardize or normalize numerical features (such as sequencing intensity values in genomic data, laboratory index values in case data, etc.) so that data of different scales can be compared within the same range. Encode text data (such as symptom descriptions, diagnosis results, etc.) into numerical data for machine learning model processing. Generate different variants of the data as needed, such as simulating mutations for genomic data or adding noise to case data, to improve data diversity and enhance the generalization ability of the model.
[0053] The dataset division includes the following steps:
[0054] Divide by virus subtype: If the goal is to identify different virus subtypes, the dataset can be divided according to the data of different subtypes. For example, divide the data into subtype A, subtype B, subtype C, etc., and each subset only contains genomic data and case data of a specific subtype.
[0055] Divide by patient group: The dataset can be divided according to different groups of patients, such as by age group, gender, disease severity (mild, moderate, severe), etc. The purpose of doing this is to enable samples in different groups to be processed independently and enable the model to learn the differences between groups.
[0056] Balance the data distribution: When dividing the dataset, ensure that the data in each subset has a relatively balanced distribution. For example, within each subset, ensure that the number of samples of different virus subtypes, different patient groups, etc. is roughly the same, and avoid having too many or too few samples of a certain subtype or group, which may affect the subsequent training effect.
[0057] Parallel processing division: For each subset, they can be further divided into smaller task blocks according to computing resources for subsequent parallel processing on multiple computing nodes. For example, the genomic data subset may be very large and can be further divided into task blocks according to chromosomes or gene segments.
[0058] Feature Extraction and Aggregation Module: Use K-mer analysis to extract mutant pattern feature data from genomic data and extract clinical feature data from case data, and aggregate the mutant pattern feature data and clinical feature data into sample data. For example, genomic data and case data may be stored separately, and they are aggregated according to virus subtypes or patient identifiers to ensure that all relevant information (genomic, clinical, symptoms, etc.) of each patient can be combined;
[0059] K-mer refers to a substring of length K in a gene sequence (for example, when K = 3, all subsequences of length 3). K-mer analysis can help identify important mutation information or variation patterns in genomic sequences, especially the repeated and variable parts of short sequences.
[0060] The Feature Extraction and Aggregation Module uses the K-mer algorithm to cut each genomic sequence in the genomic data to generate all possible K-mer subsequences, and analyzes the mutation features in the genomic data based on the K-mer subsequences.
[0061] The Feature Extraction and Aggregation Module obtains the number of K-mer subsequence partitions in the genomic sequence, and obtains the genotype difference value, Huffman coding length, and Cosine similarity of each K-mer subsequence;
[0062] Calculate the mutation factor of the genomic sequence based on the genotype difference value, Huffman coding length, and Cosine similarity. The expression is:
[0063] , where is the mutation factor, is the genotype difference value, is the Huffman coding length, is the Cosine similarity, , , are adjustment coefficients, and the adjustment coefficients , , are all greater than 0;
[0064] The calculation expression of the genotype difference value is: , where is the genotype difference value, is the number of K-mer subsequences of the genomic sequence, is the occurrence frequency of the sample data on the i-th K-mer subsequence, is the occurrence frequency of the reference genome on the i-th K-mer subsequence;
[0065] The genotype difference value measures the frequency difference between the sample and the reference genome in the K-mer subsequences. A larger difference value indicates a significant variation between the sample genome and the reference genome at that position. K-mers with larger genotype differences usually correspond to mutation regions, where mutations may affect gene functions or lead to the emergence of different phenotypes. An increase in this value indicates an increase in the difference between genomes, usually meaning an accumulation of mutations.
[0066] The calculation expression for the Huffman coding length is:
[0067] , where is the Huffman coding length, is the number of K-mer subsequences in the genome sequence, represents the frequency of the i-th K-mer subsequence in the genomic data, represents the bit length occupied by the Huffman coding of the i-th K-mer;
[0068] The Huffman coding length can help detect mutation regions in the genome. If the coding length of a certain K-mer is significantly longer, it may mean that the frequency of this K-mer is low, which may represent a mutation or a genomic variation region. Conversely, a shorter coding length means that this K-mer is more common in the genome and the mutation incidence is low.
[0069] The calculation expression for Cosine similarity is: , where is the Cosine similarity, is the frequency vector of a K-mer subsequence in the genome sequence, is the frequency vector of another K-mer subsequence in the genome sequence, is the frequency vector and the frequency vector 's dot product, are the norms of the frequency vectors and the frequency vector 's norm respectively. For example, , the norm is calculated in the same way as the norm ;
[0070] The higher the Cosine similarity, the more consistent the frequency distributions of the two genomes in the K-mer subsequences are. A low Cosine similarity means a large difference between the genomes. K-mers with a low Cosine similarity indicate a large mutation difference between the genomes, which may be reflected in certain mutation regions or variation regions.
[0071] The greater the mutation factor of the genomic sequence, the greater the mutation difference between genomes of the genomic sequence. Compare the obtained mutation factor with a preset mutation threshold. The mutation threshold is used to determine whether there is a gene mutation in the genomic sequence. If the mutation factor is less than or equal to the mutation threshold, it is determined that there is no gene mutation in the genomic sequence. If the mutation factor is greater than the mutation threshold, it is determined that there is a gene mutation in the genomic sequence.
[0072] Collect key information in the case data, such as the patient's basic information (gender, age, weight, etc.), clinical symptoms (fever, cough, etc.), diagnostic information (virus type, infection status, etc.), laboratory test results (such as blood tests, imaging examinations, etc.), and treatment plans, etc.
[0073] Perform standardization or normalization on numerical data (such as age, body temperature, laboratory test values), and convert categorical data (such as symptom descriptions, virus types) into numerical data through One-Hot encoding or Label encoding.
[0074] Encode the clinical symptoms. Use text analysis techniques (such as natural language processing) to extract standardized symptom categories or severity levels from the symptom descriptions. Combine with clinical diagnosis results, such as virus infection type, the patient's immune response, etc., to extract disease progression or possible complication characteristics. Extract biomarkers or clinical trial results from laboratory test data, and these results can often reflect the patient's health status or immune response to the virus. For example, detect certain specific gene expressions, antibody levels, etc.
[0075] Match and aggregate the genomic data and case data based on the patient identifier or virus subtype. The genomic data, clinical data, symptom information, etc. of each patient should be integrated together to form a complete sample dataset, ensuring that all relevant data of each patient can be stored in the same data table to avoid data loss or redundancy. For each sample, design a unified data structure so that information such as genomic features, clinical features, and symptoms can be seamlessly combined. For example, a table or vector containing fields such as the patient's basic information, genomic features (such as K-mer frequency), clinical symptoms, and laboratory test results can be used to represent each patient. This step ensures that in subsequent machine learning models, the input data can cover all relevant information of the patient, enabling the model to comprehensively consider genomic variations, clinical manifestations, and other clinical information.
[0076] During the aggregation process, some patients may lack certain feature information (such as genomic data or certain clinical features), and missing values need to be filled or marked. Methods such as imputation, filling, median filling, or KNN imputation can be used to handle missing values in the data to ensure the integrity of the sample data. After data aggregation, the combined data needs to be standardized to ensure that the scales of all features are consistent. Numerical features (such as K-mer frequencies, age, etc.) can be standardized (such as subtracting the mean and dividing by the standard deviation) or normalized (such as compressing the values to between 0 and 1). Categorical features (such as symptom types, treatment regimens, etc.) are converted into numerical data through One-Hot encoding or Label encoding.
[0077] Integrate all features to form a high-dimensional feature vector or matrix. The genomic features (frequencies obtained through K-mer analysis) and clinical features (such as symptoms, signs, treatment responses, etc.) of each patient will be input into the model as a whole. For some features, dimensionality reduction processing (such as PCA, principal component analysis) may be required to reduce the dimensions and computational complexity. Output the sample data set after aggregation and processing, and prepare to send it to the subsequent classification module for training or inference. This sample data set can be stored and transmitted in formats such as CSV, JSON, Parquet, etc.
[0078] The feature extraction and aggregation module in this application can effectively extract useful features from genomic data and case data and aggregate them into unified sample data, providing accurate and comprehensive input for subsequent virus subtype identification. This feature extraction method integrating genomic and clinical information can significantly improve the diagnostic accuracy and generalization ability of the model.
[0079] The feature extraction and aggregation module collects case data information, standardizes or normalizes numerical data, and converts categorical data into numerical data through One-Hot encoding or Label encoding, including the following steps:
[0080] Collect relevant information of patients, including clinical data (such as age, gender, symptoms, etc.), laboratory test results (such as blood test, imaging test results, etc.), and other clinical variables (such as treatment regimens, drug use history, etc.). The data format usually includes electronic medical records (EMR), medical database records, or other structured / unstructured formats. Integrate case data from different sources to ensure the relevance between data (such as genomic data, clinical symptoms, test results corresponding to patient IDs, etc.). Process missing values, outliers, or missing entries to ensure the integrity and consistency of the data.
[0081] Standardize numerical data to make the data have the same dimension. After standardization, the data distribution has a zero mean and a unit standard deviation, enabling features of different scales to be compared on the same scale and avoiding the influence of large numerical values on model training. Normalize numerical data by scaling it to a specified interval (usually the [0, 1] interval). Normalization ensures that the numerical values of all features are within the same interval, which helps the convergence speed during algorithm training. Select the standardization or normalization method according to the characteristics of the data and the requirements of the subsequent model. Standardization is applicable to most machine learning algorithms, especially distance-based algorithms (such as KNN, SVM, etc.). Normalization is often used in cases such as deep learning and neural networks that require scale consistency processing of input data.
[0082] For unordered categorical data (such as gender, blood type, drug use history, etc.), use One-Hot encoding to convert categorical data into binary vectors. For example, gender (Male, Female) can be represented by One-Hot encoding as: Male → [1, 0], Female → [0, 1]. One-Hot encoding can be extended to variables with multiple categories, generating multiple binary feature columns, with each category corresponding to one feature.
[0083] For ordered categorical data (such as disease severity, age group, etc.), use Label encoding to convert categories into integers. The Label encoding method maps each category to a unique integer value, such as: Mild → 0, Moderate → 1, Severe → 2. This method can preserve the order information between categories and is applicable to categorical variables with a clear order.
[0084] For variables with a large number of categories (such as city names, drug names, etc.), using One-Hot encoding may lead to the curse of dimensionality (too high feature dimensionality). In this case, other methods such as frequency encoding or embedding representation (such as Word2Vec, embedding layer, etc.) can be used to reduce the dimension and avoid performance degradation.
[0085] The feature extraction and aggregation module encodes clinical symptoms, uses text analysis techniques to extract standardized symptom categories or severity levels from symptom descriptions, combines clinical diagnosis results, extracts disease progression or complication features, and extracts biomarkers or clinical trial results from laboratory test data, including the following steps:
[0086] Symptom Classification System: According to the symptom description text, apply a standardized symptom classification system (such as ICD-10 or other clinical standards) to classify symptoms. Each symptom can be mapped to a specific category. For example: headache → neurological symptoms, cough → respiratory symptoms. Based on the context of the symptom description, use natural language processing (NLP) techniques (such as entity recognition) to identify the specific symptom category.
[0087] Use text analysis techniques (such as sentiment analysis or sentiment classification) to analyze the severity of symptoms. For example, detect the difference between "severe headache" and "mild headache", and assign a severity value (such as a scale from 1 to 5) to each symptom according to the emotional intensity of the description. Determine the severity of symptoms by extracting keywords or phrases (such as "severe", "mild", "sudden") in the symptom description to form a standardized grading system. Based on the preprocessed symptom text and combined with the assessment of severity, generate a standardized symptom code. For each patient, the symptom data will be converted into numerical features in a unified format for subsequent analysis.
[0088] Extract information on the progression of diseases from electronic medical records or doctors' diagnostic records, such as the stage of the disease, course of the disease, clinical diagnosis results, etc. Standardize the diagnostic information using a medical coding system (such as ICD-10, SNOMED-CT, etc.) for subsequent analysis. If the diagnostic information has a time sequence, time series analysis methods can be used to extract disease progression features. For example, information such as the development of the patient's condition and treatment response can be converted into time series data to reflect the progression process of the disease. According to the patient's diagnostic record, extract the trend of changes in clinical indicators (such as body temperature, blood pressure, pulse, etc.) over time as a sign of disease progression. Extract possible complication information from the case, such as respiratory failure, organ failure, etc. Analyze the timing, frequency of complication occurrence and its relationship with the disease itself. Code the complications using a standardized coding system (such as ICD-10) and input these codes as one of the disease features into the subsequent analysis model.
[0089] Collect the patient's laboratory test results, such as blood tests, urine tests, imaging tests, etc. The data may include numerical results (such as blood glucose levels, white blood cell counts, etc.) and categorical results (such as virus test results, imaging scores, etc.). According to the known disease-related biomarkers, extract the biomarker information related to the disease from the laboratory test data. For example, certain blood biomarkers (such as C-reactive protein, white blood cell count) may be related to the severity of the disease or the inflammatory response. For each patient, record the levels of all relevant biomarkers for further analysis. If the patient participated in a clinical trial, extract the relevant clinical trial results (such as drug response, treatment effect, etc.). These results can serve as important features such as disease progression, treatment response, etc. Convert the clinical trial results into numerical or categorical features and combine them with the patient's medical history, treatment records, etc. to form complete clinical features.
[0090] The feature extraction and aggregation module matches and aggregates the genomic data and case data based on the patient identifier or virus subtype. After integrating the genomic data, clinical data, and symptom information of each patient, a sample data set is obtained, including the following steps:
[0091] Collect the patient's genomic sequence data from the genomic database or laboratory test results. The data includes information such as gene mutations, single nucleotide polymorphisms (SNPs), insertions and deletions (InDels), etc. Collect the patient's clinical information, including basic demographic data (such as age, gender), medical history, medication history, laboratory test results, etc. This data may be stored in an electronic medical record (EMR) system or other medical databases. Clean the genomic data and case data, including removing missing values, duplicate data, etc. Perform standardization or normalization on numerical data and One-Hot encoding or Label encoding on categorical data to convert them into a numerical form suitable for analysis.
[0092] Use the patient's unique identifier (such as medical record number or ID number) as the keyword to match the genomic data and case data. Ensure that the genomic information of each patient is associated with data such as their clinical and symptom information. For example, a patient's genomic data may contain mutation information, while the patient's case data contains clinical symptoms and laboratory test results. Combine these data through the patient identifier. If the virus subtype information is available, the data can be grouped according to the virus subtype. This helps to analyze the association between the genomic characteristics of different subtypes and the clinical manifestations. Combine the genomic data, clinical information, and symptom data of each patient and perform subset division based on the virus subtype.
[0093] Aggregate genomic data, clinical data, and symptom information. For example, for the same patient, genomic data and case data may be stored in different tables or databases respectively. The aggregation process combines these data into a complete dataset. The result of aggregation is to generate a multi-dimensional feature vector for each patient, which includes their genomic information, clinical characteristics, symptom descriptions, etc. If the case data and genomic data are stored separately, they are merged according to specific rules by virus subtype or patient identifier. Combine the processed genomic data, case data, and symptom information into a complete sample dataset. Each patient will correspond to a sample, and the sample includes the following features:
[0094] Genomic data: mutations, SNP information, etc.;
[0095] Clinical data: diagnosis results, laboratory test results, etc.;
[0096] Symptom information: symptom categories, severity levels, etc.;
[0097] The sample data of each patient is represented in the form of a feature vector, and each feature corresponds to a column of data. For example, the features of genomic data can be the presence or absence of gene mutations, the features of case data can be clinical diagnoses, laboratory indicators, etc., and the features of symptom data can be symptom categories and their severity levels. The final output dataset will contain the feature data of all patients, providing input for subsequent tasks such as machine learning model training, virus subtype identification, and clinical prediction.
[0098] Classification module: Distribute the sample data of multiple subsets to be classified to each computing node for parallel processing. Each computing node uses the trained machine learning model for classification to obtain the classification results of each sample data;
[0099] The classification module assigns subsets to each computing node. Each computing node is an independent processing unit, usually a working unit in a cluster. Each node runs a machine learning model to classify the data subset assigned to it. When performing the classification task, the computing node uses the pre-trained machine learning model for inference to generate the classification results of each subset. These models may be based on different machine learning algorithms, such as decision tree models, support vector machine (SVM) models, random forest models, neural network models, etc.;
[0100] The classification module assigns subsets to each computing node. After each computing node finishes processing the subset assigned to it, it outputs the classification result of that subset. The classification result of each subset can be a label (such as "subtype 1" or "subtype 2") or a probability distribution (such as "90% belongs to subtype 1", "10% belongs to subtype 2"). After all computing nodes complete the classification task, the obtained classification results are aggregated. The classification module integrates these results and passes them to subsequent modules, such as the evaluation module, for further analysis and evaluation of the model's performance.
[0101] The computing node processes the subset through a pre-trained neural network model. In this application, the pre-training of the neural network model includes the following steps:
[0102] Collect the dataset for pre-training, which usually includes genomic data, case data, clinical data, etc. These datasets should be representative and able to reflect the diversity of virus subtypes and classification requirements. Before the data is input into the neural network, data cleaning is performed to ensure that the data has no missing values, outliers, or duplicates. The data is standardized (for example, converting numerical data into a standard normal distribution with a mean of 0 and a variance of 1) or normalized (for example, scaling all data to the interval [0, 1]) to ensure the consistency of the data scale and avoid the excessive influence of certain features on model training. If the dataset is small, more training samples can be generated through data augmentation techniques (such as random perturbation, data transformation, etc.) to increase the generalization ability of the model.
[0103] The architecture of the fully connected neural network includes an input layer, hidden layers, and an output layer. Each layer consists of multiple neurons (nodes). The number of neurons in the input layer usually equals the number of features, and the number of neurons in the output layer equals the number of classes (in a classification task, usually the number of virus subtype classes).
[0104] Input layer: Each node corresponds to a feature (such as genomic mutations, clinical data, symptoms, etc.).
[0105] Hidden layers: Usually one to multiple hidden layers are selected to increase the complexity of the network and help capture the non-linear relationships in the data. The number of nodes in the hidden layers is usually between 100 and 1000 and is adjusted according to the scale of the dataset.
[0106] Output layer: The number of nodes in the output layer equals the number of classification categories. For virus subtype classification, the number of nodes in the output layer is the number of different virus subtype species.
[0107] Activation function: Each node in the hidden layer and output layer usually uses an activation function (such as ReLU (Rectified-Linear-Unit) or Sigmoid) to increase non-linear features.
[0108] For the hidden layer, the commonly used activation function is ReLU because it can effectively avoid the vanishing gradient problem and accelerate training.
[0109] For the output layer, if it is a multi-classification problem, the Softmax activation function is commonly used to output the probabilities of each class.
[0110] The dataset is divided into a training set, a validation set, and a test set, usually in a ratio of 7:2:1. The training set is used to train the model, the validation set is used to tune the hyperparameters of the model, and the test set is used to finally evaluate the performance of the model. According to the categories of virus subtypes, corresponding labels are created. For example, if it is a "subtype 1", "subtype 2" task, the labels can be integers such as 0, 1, etc.
[0111] The input data is passed through the network layer by layer, and weighted summation is performed at each layer and processed by the activation function. The final output layer will generate the probability distribution of each category, and the cross-entropy loss function (Cross-Entropy-Loss) is used to calculate the difference between the output and the actual labels. The cross-entropy loss function works well for multi-classification tasks and can measure the difference between the probability distribution predicted by the model and the actual labels. According to the results calculated by the loss function, the gradients of each weight and bias are calculated through the backpropagation algorithm. Backpropagation is a process that uses the chain rule. It reversely transmits the error from the output layer to the input layer and updates the parameters (weights and biases) of each layer. Optimization algorithms (such as Adam, SGD, etc.) are used to update the parameters in the network to minimize the loss function. The Adam optimizer performs excellently when dealing with sparse gradients and large-scale data, and it is usually selected for training.
[0112] The learning rate is a hyperparameter that controls the step size of each parameter update. Techniques such as learning rate decay can be used to avoid oscillations or overfitting during the training process. Select an appropriate batch size (such as 32, 64, etc.). Too small a batch size may lead to unstable training, while too large a batch may lead to insufficient memory. Define the number of training epochs. Usually, the early stopping method (Early-Stopping) is used, that is, stop training in advance when the performance on the validation set no longer improves to prevent overfitting. After training is completed, save the optimal model (including the network structure and weights). The API of the framework (such as TensorFlow, PyTorch) can be used to save and load the model. In the classification module, load the trained neural network model and use it to predict the subset data to be classified to obtain the classification results of each sample.
[0113] It should be noted that the pre-training process of other machine learning models belongs to the prior art and will not be elaborated in this application.
[0114] Subset Data: Suppose we have a dataset containing 1000 samples, which is divided into 5 subsets, with each subset containing 200 samples.
[0115] Distribute Data: Distribute each subset to 5 different computing nodes for parallel processing.
[0116] Node Computation:
[0117] Computing Node 1 uses the trained model to classify 200 samples, and the classification results are output: 100 samples are of "Subtype 1" and 100 samples are of "Subtype 2".
[0118] Computing Node 2 performs the same operation and outputs its classification results: 150 samples are of "Subtype 1" and 50 samples are of "Subtype 2".
[0119] Similarly, other computing nodes will also perform classification and obtain their respective results.
[0120] Result Aggregation: The classification module aggregates the results of the 5 computing nodes to obtain the final classification result.
[0121] The evaluation module performs clustering processing based on the operating status of each computing node, evaluates the classification effects of each cluster, and then selects the optimal machine to train the model for application;
[0122] Obtain the real-time temperature, temperature rising rate, and voltage standard deviation of each computing node, perform normalization processing on the real-time temperature, temperature rising rate, and voltage standard deviation, map the value ranges of the real-time temperature, temperature rising rate, and voltage standard deviation to between [0, 1], and obtain the operating index of the computing node by adding the normalized real-time temperature, temperature rising rate, and voltage standard deviation;
[0123] Suppose it is necessary to divide all computing nodes into K (in this application, K is set to 3) clusters. First, randomly select the operating indices of K computing nodes as the cluster center values of the K clusters, calculate the absolute values of the differences between the remaining computing nodes and the cluster center values of the K clusters, and then assign the computing nodes to the cluster with the smallest absolute value of the difference. After all computing nodes are assigned, recalculate the average value of the operating indices of each cluster, and use the average value of the operating indices as the cluster center value of the cluster. Repeat the clustering step. When the difference between the cluster center value obtained in any iteration and the cluster center value obtained in the previous iteration is less than 0.1, it is determined that the convergence condition is met, and K clusters are output;
[0124] Obtain the classification effect coefficient of each computing node in the cluster. The classification effect coefficient is obtained by adding the normalized classification speed and the classification accuracy, and calculate the average value of the classification effect coefficients of all clusters. Select the cluster with the largest classification effect coefficient, and obtain the machine learning model used by each computing node in this cluster. Use the machine learning model with the largest usage quantity as the optimal machine training model for application.
[0125] In this application, by clustering the computing nodes, it is possible to evaluate the classification effect of computing nodes with the same or similar performance, and the evaluation result is more accurate.
[0126] The above formulas are all dimensionless and take their numerical values for calculation. The formula is obtained by collecting a large amount of data for software simulation to obtain a formula closest to the actual situation. The preset parameters in the formula are set by those skilled in the art according to the actual situation.
[0127] It should be understood that the term "and / or" in this article is only a description of the association relationship between associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. Here, A and B can be singular or plural. In addition, the character " / " in this article generally represents an "or" relationship between the associated objects before and after, but it may also represent an "and / or" relationship. Specifically, it can be understood by referring to the context before and after.
[0128] It should be understood that in various embodiments of this application, the magnitude of the sequence numbers of the above processes does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of this application.
[0129] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed in this article can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but this implementation should not be considered to exceed the scope of this application. Those skilled in the art can clearly understand that for the convenience and simplicity of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0130] As described above, it is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of changes or substitutions, which should all be covered within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the protection scope of the said claims.
Claims
1. An artificial intelligence-based diagnosis and recognition system for subtype equine encephalomyelitis virus, characterized in that: It includes a subset division module, a feature extraction and aggregation module, a classification module, and an evaluation module; Subset division module: Collect genomic data and case data to construct a data set. After preprocessing the data set, divide it into multiple subsets; Feature extraction and aggregation module: Use K-mer analysis to extract mutation pattern feature data from genomic data, and extract clinical feature data from case data, and aggregate the mutation pattern feature data and clinical feature data into sample data; Classification module: Distribute multiple subset sample data to be classified to each computing node for parallel processing. Each computing node uses a trained machine learning model for classification to obtain the classification results of each sample data; Evaluation module: Perform clustering processing based on the running status of each computing node, evaluate the classification effects of each cluster, and then select the optimal machine training model for application; The evaluation module obtains the running index of each computing node, performs K-clustering processing on each computing node according to the running index, outputs K clusters, obtains the classification effect coefficient of each computing node in the cluster, and selects the machine learning model used by each computing node in the cluster with the largest average classification effect coefficient, and uses the machine learning model with the largest number of uses as the optimal machine training model for application.
2. The diagnosis and recognition system for equine encephalomyelitis virus subtypes based on artificial intelligence according to claim 1, characterized in that: The feature extraction and aggregation module uses the K-mer algorithm to cut each genomic sequence in the genomic data to generate K-mer subsequences, and analyzes the mutation characteristics in the genomic data based on the K-mer subsequences; Collect case data information, perform standardization or normalization processing on numerical data, and convert categorical data into numerical data through One-Hot encoding or Label encoding; Encode clinical symptoms, use text analysis techniques to extract standardized symptom categories or severity levels from symptom descriptions, combine clinical diagnosis results, extract disease progression or complication characteristics, and extract biomarkers or clinical trial results from laboratory test data; Match and aggregate genomic data and case data based on patient identifiers or virus subtypes. After integrating the genomic data, clinical data, and symptom information of each patient, a sample data set is obtained.
3. The artificial intelligence-based subtype equine encephalomyelitis virus diagnosis and recognition system according to claim 2, characterized in that: The feature extraction and aggregation module obtains the number of K-mer subsequence partitions in the genomic sequence, and obtains the genotype difference value, Huffman coding length, and Cosine similarity of each K-mer subsequence; Calculate and obtain the mutation factor of the genomic sequence based on the genotype difference value, Huffman coding length, and Cosine similarity; Compare the obtained mutation factor with a preset mutation threshold. The mutation threshold is used to judge whether there is a gene mutation in the genomic sequence. If the mutation factor is less than or equal to the mutation threshold, it is judged that there is no gene mutation in the genomic sequence. If the mutation factor is greater than the mutation threshold, it is judged that there is a gene mutation in the genomic sequence.
4. The artificial intelligence-based subtype equine encephalomyelitis virus diagnosis and recognition system according to claim 3, wherein: The evaluation module obtains the real-time temperature, temperature rise rate, and voltage standard deviation of each computing node, performs normalization processing on the real-time temperature, temperature rise rate, and voltage standard deviation, maps the value ranges of the real-time temperature, temperature rise rate, and voltage standard deviation to between [0, 1], and obtains the operation index of the computing node by adding the normalized real-time temperature, temperature rise rate, and voltage standard deviation; Suppose it is necessary to divide all computing nodes into K clusters. First, randomly select the operation indices of K computing nodes as the cluster center values of the K clusters, and calculate the absolute values of the differences between the remaining computing nodes and the cluster center values of the K clusters. Then, assign the computing nodes to the cluster with the smallest absolute value of the difference. After all computing nodes are assigned, recalculate the average operation index of each cluster, and use the average operation index as the cluster center value of the cluster. Repeat the clustering step. When the difference between the cluster center value obtained in any iteration and the cluster center value obtained in the previous iteration is less than 0.1, it is determined that the convergence condition is met, and K clusters are output; Obtain the classification effect coefficient of each computing node in the cluster. The classification effect coefficient is obtained by adding the normalized classification speed and classification accuracy, and calculate the average value of the classification effect coefficients of all clusters. Select the cluster with the largest classification effect coefficient, and obtain the machine learning model used by each computing node in this cluster. Use the machine learning model with the largest usage quantity as the optimal machine training model for application.
5. The diagnostic and recognition system for equine encephalomyelitis virus subtypes based on artificial intelligence according to claim 4, characterized in that: The classification module assigns subsets to each computing node, and the computing node uses a pre-trained machine learning model to perform classification tasks. The machine learning model includes a decision tree model, a support vector machine model, a random forest model, and a neural network model; After each computing node finishes processing the subset assigned to it, it outputs the classification result of the subset. The classification result of each subset is a label. After all computing nodes complete the classification tasks, the obtained classification results are summarized.
6. The diagnostic and recognition system for equine encephalomyelitis virus subtypes based on artificial intelligence according to claim 5, characterized in that: The calculation expression of the Huffman coding length is: , where is the Huffman coding length, is the number of K-mer subsequences of the genomic sequence, represents the frequency of occurrence of the i-th K-mer subsequence in the genomic data, represents the bit length occupied by the Huffman coding of the i-th K-mer.
7. The diagnosis and recognition system for equine encephalomyelitis virus subtypes based on artificial intelligence according to claim 6, characterized in that: The calculation expression of the Cosine similarity is as follows: , where is the Cosine similarity, is the frequency vector of a K-mer subsequence in the genomic sequence, is the frequency vector of another K-mer subsequence in the genomic sequence, is the frequency vector and the frequency vector is the dot product of, are respectively the norm of the frequency vector and the norm of the frequency vector .
8. The diagnostic and recognition system for equine encephalomyelitis virus subtypes based on artificial intelligence according to claim 7, characterized in that: The calculation expression of the genotype difference value is as follows: , where is the genotype difference value, is the number of K-mer subsequences of the genomic sequence, is the occurrence frequency of the sample data on the i-th K-mer subsequence, is the occurrence frequency of the reference genome on the i-th K-mer subsequence.
9. The artificial intelligence-based diagnosis and recognition system for subtype equine encephalomyelitis virus according to claim 8, characterized in that: The subset division module collects genomic data and case data, collects genomic data from different virus samples. Each virus sample contains the gene sequence of the virus, and the genomic data is obtained through the genomic sequencing results of the virus or a public gene database, and the data format is in FASTQ or FASTA genomic data format; Collect case data related to patients, including patient information, clinical symptoms, diagnostic information, and laboratory test results, and obtain case data through patient electronic health records, medical databases, or laboratory reports, and the data format is in CSV, JSON, XML structured or semi-structured format; The genomic data saves the sequence information in a standard format, and the case data is unified into a table form. The numerical features are standardized or normalized, and the text data is encoded and processed into numerical data.
10. The diagnosis and recognition system for equine encephalomyelitis virus subtypes based on artificial intelligence according to claim 9, characterized in that: The subset division module has the following ways to divide the dataset: dividing by virus subtype, dividing by patient group, balancing data distribution, or dividing by parallel processing.
Citation Information
Patent Citations
Virus gene classification method and device and electronic equipment
CN113299345A
Deep learning-based disease risk prediction model construction method and system
CN117831771A
Breast cancer risk assessment method and system combined with deep learning
CN117976185A