A method and system for constructing a neural genetic disease recognition model
By using inter-class difference analysis and multi-stage reinforcement learning, the problems of poor generalization performance and inter-class confusion in the neurogenetic disease identification model under small sample conditions were solved, and the adaptive reinforcement construction and recognition ability of the model were improved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- THE FIRST AFFILIATED HOSPITAL OF FUJIAN MEDICAL UNIV
- Filing Date
- 2026-04-30
- Publication Date
- 2026-05-29
AI Technical Summary
Under small sample conditions, neurogenetic disease identification models are prone to overfitting and have poor generalization performance. Traditional data synthesis methods tend to generate samples with blurred class boundaries, which exacerbates inter-class confusion, and lack adaptive synthesis strategies for inter-class differences.
By analyzing the inter-class differences, the difficulty of distinguishing features between different disease categories is quantified. Differentiated user feature synthesis strategies are configured to generate synthetic user feature sets and perform multi-stage reinforcement learning to optimize the recognition model.
It improves the generalization performance and diagnostic reliability of the neurogenetic disease identification model under small sample conditions, reduces the false positive rate of inter-class confusion, and enhances the model's ability to distinguish easily confused categories.
Smart Images

Figure CN122117342A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of neurogenetic disease identification technology, specifically to a method and system for constructing a neurogenetic disease identification model. Background Technology
[0002] Neurogenetic diseases are a class of nervous system diseases caused by gene mutations.
[0003] However, neurogenetic diseases fall into the category of rare diseases, resulting in a scarcity of clinical sample data, with a very limited number of available samples for each disease category. Recognition models trained under limited sample conditions are prone to overfitting, struggle to learn the true discriminative boundaries between categories, and exhibit poor generalization performance. To address the limited sample problem, existing techniques attempt to expand training samples through data synthesis. However, different categories of neurogenetic diseases share high similarities in clinical features and exhibit small inter-class differences. Samples randomly generated by traditional data synthesis methods tend to fall near or even cross class boundaries, exacerbating misclassification of easily confused categories. Furthermore, existing methods lack specific consideration for inter-class differences and cannot adaptively adjust the synthesis strategy based on the difficulty of distinguishing each category.
[0004] Therefore, there is an urgent need for a method that can synthesize targeted data based on inter-class differences and continuously optimize recognition capabilities in multi-stage testing. Summary of the Invention
[0005] This invention addresses the technical problems in existing technologies, such as poor generalization performance of models under small sample conditions, the tendency of traditional data synthesis methods to generate samples with blurred class boundaries leading to increased inter-class confusion, and the lack of adaptive synthesis strategies for inter-class differences. It provides a method and system for constructing a neurogenetic disease identification model.
[0006] The technical solution of the present invention to solve the above-mentioned technical problems is as follows:
[0007] In a first aspect, the present invention provides a method for constructing a model for identifying neurogenetic diseases, comprising:
[0008] Multiple identification results and multiple user feature sets for neurogenetic disease identification are obtained, and inter-class difference analysis is performed to obtain multiple inter-class differences.
[0009] Based on the multiple recognition results and multiple user feature sets, a first recognition model for the recognition of neurogenetic diseases is trained.
[0010] Based on the inter-class differences, a user feature synthesis strategy is configured to generate multiple synthesized user feature sets and multiple synthesized recognition results. Reinforcement learning of the first recognition model is then performed to obtain a second recognition model. The user feature synthesis strategy includes the number of synthesized features and the degree of synthesis deviation.
[0011] The second recognition model is tested for the combined recognition error rate of multiple recognition results to obtain multiple combined recognition error rates. Combined with the multiple inter-class differences, a user feature synthesis strategy is configured to continue generating multiple synthetic user feature sets and multiple synthetic recognition results. Reinforcement learning is then performed to obtain the third recognition model.
[0012] Secondly, the present invention provides a system for constructing a model for identifying neurogenetic diseases, comprising:
[0013] The difference analysis module is used to obtain multiple identification results and multiple user feature sets for neurogenetic disease identification, perform inter-class difference analysis, and obtain multiple inter-class differences.
[0014] The first training module is used to train a first recognition model for the recognition of neurogenetic diseases based on the multiple recognition results and multiple user feature sets.
[0015] The reinforcement learning module is used to configure user feature synthesis strategy based on multiple inter-class differences, generate multiple synthesized user feature sets and multiple synthesized recognition results, perform reinforcement learning on the first recognition model, and obtain a second recognition model. The user feature synthesis strategy includes the number of synthesis and the degree of synthesis deviation.
[0016] The iterative optimization module is used to test the combined recognition error rate of the second recognition model for multiple recognition results, obtain multiple combined recognition error rates, combine the multiple inter-class differences, configure the user feature synthesis strategy, continue to generate multiple synthetic user feature sets and multiple synthetic recognition results, perform reinforcement learning, and obtain the third recognition model.
[0017] The beneficial effects of this invention are:
[0018] Compared to existing technologies, this invention first quantifies the difficulty of distinguishing features between different categories of neurogenetic diseases through inter-class dissimilarity analysis, providing a quantitative basis for differential data synthesis. Secondly, it adaptively configures the number and deviation of synthesized samples based on inter-class dissimilarity, generating more and larger deviations for easily confused categories with small inter-class differences, and fewer and smaller deviations for categories with large inter-class differences, enabling the synthesized data to specifically enhance the model's discrimination ability on difficult categories. Thirdly, by testing the combined recognition error rate of the model, it identifies currently easily confused category pairs, and further configures the synthesis strategy based on inter-class dissimilarity, continuing reinforcement learning to form a multi-stage iterative optimization. This invention solves the problems of poor model generalization under small sample conditions and the exacerbation of inter-class confusion by traditional synthesis methods, achieving adaptive enhancement construction of neurogenetic disease identification models. Attached Figure Description
[0019] Figure 1A flowchart illustrating a method for constructing a neurogenetic disease identification model provided by the present invention;
[0020] Figure 2 This is a schematic diagram of the structure of a neurogenetic disease identification model construction system provided by the present invention.
[0021] In the attached diagram, the components represented by each number are as follows:
[0022] The module consists of: a difference analysis module 11, a first training module 12, a reinforcement learning module 13, and an iterative optimization module 14. Detailed Implementation
[0023] Example 1, as Figure 1 As shown, this embodiment of the invention provides a method for constructing a neurogenetic disease identification model, including:
[0024] S10: Obtain multiple identification results and multiple user feature sets for neurogenetic disease identification, perform inter-class difference analysis, and obtain multiple inter-class differences;
[0025] First, multiple identification results and multiple user feature sets are obtained for neurogenetic disease identification. Neurogenetic diseases are a class of nervous system disorders caused by gene mutations, including various different disease types, such as Huntington's disease, spinal muscular atrophy, and hereditary ataxia. Because different disease types differ in clinical symptoms, gene markers, and imaging features, multiple identification results need to be obtained, each representing a specific neurogenetic disease type. Simultaneously, multiple user feature sets are obtained, each containing multi-dimensional information such as clinical characteristics, gene testing data, and neuroimaging indicators from multiple patients within the corresponding disease type.
[0026] Furthermore, since neurogenetic diseases fall into the category of rare diseases, the number of available patient samples for each disease category is extremely limited, typically only a few dozen or even fewer, constituting a typical small sample dataset. Therefore, this step collects multiple identification results and multiple user feature sets to provide foundational data for subsequent inter-class difference analysis, quantifying the difficulty of feature differentiation between different disease types, thereby guiding the configuration of differentiated data synthesis strategies and improving the model's ability to identify easily confused categories.
[0027] Specifically, multiple identification results and multiple user feature sets for neurogenetic disease identification are obtained, and inter-class dissimilarity analysis is performed to obtain multiple inter-class dissimilarity sets, including:
[0028] Multiple identification results and multiple user feature sets for neurogenetic disease identification are obtained, where the multiple identification results and multiple user feature sets are small sample data;
[0029] For each user feature set of the recognition result, the difference magnitude between the user feature sets of other recognition results and the user feature sets of all other recognition results is calculated to obtain multiple inter-class differences.
[0030] First, multiple identification results and multiple user feature sets for neurogenetic disease identification are obtained, where the multiple identification results and multiple user feature sets are small sample data.
[0031] Specifically, neurogenetic diseases typically include various types such as Huntington's disease, spinal muscular atrophy, hereditary ataxia, and peroneal muscular atrophy. Each disease differs in its causative genes, clinical manifestations, and inheritance patterns, with the corresponding identification result being a specific disease name or category label. Each identification result is associated with a set of user features, which contains characteristic data of all patients under that disease category, such as variations at specific gene loci in genetic testing, age of onset, rate of disease progression, electromyography indicators, imaging features, and other multi-dimensional information.
[0032] Meanwhile, because neurogenetic diseases are rare, the number of patient samples in each category is very limited. For example, each category may only have complete clinical and genetic data from 10 to 30 patients, constituting a typical small sample dataset. Training a model based on small sample data is prone to overfitting, meaning the model over-memorizes noise and specific features from the training samples and cannot generalize to unseen patient data. Therefore, it is necessary to subsequently expand the model with synthetic data to improve its generalization ability.
[0033] However, since different categories of neurogenetic diseases may have high similarities in characteristics, blindly synthesizing data may exacerbate inter-class confusion. Therefore, it is necessary to conduct inter-class difference analysis first to provide a quantitative basis for differentiated synthesis strategies.
[0034] Specifically, for each user feature set of the identification result, the difference magnitude between the user feature sets of all other identification results is calculated to obtain multiple inter-class difference degrees.
[0035] The difference magnitude is used to quantify the average degree of difference between two categories in the feature space. For each disease category, the inter-class difference degree is obtained by calculating the difference magnitude between it and all other categories and taking the average. This inter-class difference degree characterizes the overall distinguishability of each disease category from all other disease categories in the feature space. The larger the inter-class difference degree, the lower the overlap of the feature distribution of this category with other categories, and the easier it is to identify. The smaller the inter-class difference degree, the more the feature distribution of this category overlaps with other categories or the boundaries are blurred, making it more difficult to identify. The model is prone to confusing it with other categories, requiring more attention and stronger synthesis strategies in subsequent data synthesis.
[0036] Specifically, for each user feature set of the identification result, the magnitude of the difference between the user feature sets of all other identification results is calculated to obtain multiple inter-class dissimilarity measures, including:
[0037] Select the first user feature set of the first recognition result, calculate the difference magnitude of each first user feature from each user feature in the other all user feature sets, and calculate the mean to obtain the inter-class difference degree;
[0038] Continue calculating the inter-class dissimilarity of the user feature set for each other recognition result to obtain multiple inter-class dissimilarity values.
[0039] It is important to note that before calculating inter-class dissimilarity, the multi-dimensional features in the user feature set need to be normalized and standardized feature vectors constructed. Because the user feature set contains various types of clinical features, gene testing data, and neuroimaging indicators, the dimensions and numerical ranges of different features vary significantly. For example, the age feature ranges from 0 to 100 years, the gene expression level feature ranges from 0 to 1, and the imaging feature ranges from hundreds to thousands. If the raw feature values are directly used to calculate the Euclidean distance, features with larger numerical ranges will dominate the distance calculation results, causing features with smaller numerical ranges but important discriminative significance to be ignored.
[0040] Specifically, a min-max normalization method is used to scale each feature dimension independently. For the j-th feature dimension, the minimum value min of that dimension is calculated across the entire user feature set. j and maximum value max j For any user feature, the original value x in this dimension. j The normalized value is (x j Subtract min j Divide by (max) j Subtract min j After normalization, the numerical range of all feature dimensions is compressed to the interval between 0 and 1, eliminating the influence of dimensional differences on distance calculation. After normalization, all normalized feature values for each user feature are concatenated into a fixed-dimensional one-dimensional feature vector according to a preset order, such as a fixed arrangement order of feature categories like age, gene expression level, imaging indicators, age of onset, disease progression rate, and electromyography indicators. The dimension of this feature vector is equal to the number of feature categories contained in the user feature set. For example, if the user features include 20 feature categories such as age, gene expression level, and imaging indicators, the dimension of the concatenated feature vector after normalization will be 20. Subsequent calculations of inter-class dissimilarity are all based on this standardized feature vector.
[0041] First, select the first user feature set of the first recognition result, calculate the difference magnitude of each first user feature from each user feature in the other all user feature sets, and calculate the mean to obtain the first inter-class difference degree.
[0042] The first identification result represents a specific type of neurogenetic disease. The first user feature set is the set of feature data for all patients under this disease type.
[0043] Specifically, suppose there are M disease categories. For the i-th category, its user feature set contains feature vectors of multiple patients. The inter-class dissimilarity of the i-th category is calculated as follows: calculate the distance between each patient feature vector in the i-th category and each patient feature vector in all other categories (i.e., all M-1 categories excluding the i-th category). Optionally, this dissimilarity can be calculated using Euclidean distance. Then, take the average distance between all cross-category patient pairs as the inter-class dissimilarity of the i-th category. The inter-class dissimilarity of each category is calculated sequentially using the same method, resulting in M inter-class dissimilarity values.
[0044] For example, consider three categories: Huntington's disease (Category A), spinal muscular atrophy (Category B), and hereditary ataxia (Category C). First, calculate the inter-class dissimilarity for category A. Specifically, take the feature vector of each patient in category A and calculate the Euclidean distance between it and the feature vectors of each patient in categories B and C. Sum the distances between all cross-category patient pairs and divide by the total number of comparisons to obtain the inter-class dissimilarity for category A. This value reflects the overall distinguishability of category A from categories B and C.
[0045] Similarly, to calculate the inter-class dissimilarity of class B: take the feature vector of each patient in class B, calculate the Euclidean distance with the feature vectors of each patient in classes A and C respectively, and take the average of all distances to obtain the inter-class dissimilarity of class B. This value reflects the degree of overall feature differentiation between class B and classes A and C.
[0046] Similarly, to calculate the inter-class dissimilarity of class C: take the feature vector of each patient in class C, calculate the Euclidean distance with the feature vectors of each patient in classes A and B respectively, and take the average of all distances to obtain the inter-class dissimilarity of class C. This value reflects the degree of overall feature differentiation between class C and classes A and B.
[0047] Specifically, if the inter-class difference of class A is small, it indicates that the feature distribution of class A is relatively similar to that of classes B and C, and class A is easily confused with other classes; if the inter-class difference of class A is large, it indicates that the feature distribution of class A is significantly different from that of classes B and C, and class A is easier to distinguish from other classes.
[0048] Through the above calculations, each category obtains a quantified inter-class dissimilarity level, which provides a basis for configuring subsequent differential synthesis strategies. The smaller the inter-class dissimilarity level, the more similar the category is to other categories, making identification more difficult and requiring greater attention during data synthesis. For example, if category B has the smallest inter-class dissimilarity level, it indicates that category B is quite similar to both categories A and C, making it the most easily confused category. Therefore, more boundary samples should be generated for category B during data synthesis.
[0049] S20: Based on the multiple recognition results and multiple user feature sets, train a first recognition model for the recognition of neurogenetic diseases;
[0050] Furthermore, based on the aforementioned multiple recognition results and multiple user feature sets, a first recognition model for identifying neurogenetic diseases is trained. This first recognition model serves as the initial foundational model for subsequent reinforcement learning, establishing preliminary recognition capabilities on the original small sample data and providing a starting point for subsequent synthetic data reinforcement learning.
[0051] Due to the limited number of original samples, the first recognition model may suffer from overfitting. It performs well on the training set but lacks generalization ability on the test set, especially for easily confused categories with small inter-class differences. Therefore, it needs to be further improved through subsequent synthetic data reinforcement learning.
[0052] Specifically, based on the multiple identification results and multiple user feature sets, a first identification model for identifying neurogenetic diseases is trained, including:
[0053] Based on machine learning, we construct a basic model architecture for the identification of neurogenetic diseases.
[0054] The basic model architecture is trained and validated using the multiple user feature sets as training and validation data, and multiple recognition results as supervision and validation labels. After the validation accuracy is qualified, the first recognition model is obtained.
[0055] First, a basic model architecture for the identification of neurogenetic diseases is constructed based on machine learning. Since neurogenetic disease data is characterized by high feature dimensionality and small sample size, and the input after normalization is a multi-dimensional feature vector, with each feature vector containing values for multiple feature categories, an algorithm that can naturally handle multi-dimensional feature vectors and has good robustness to high-dimensional, small-sample data needs to be selected. Random forest is an ensemble learning algorithm that classifies data by constructing multiple decision trees and combining their voting results. It can effectively handle high-dimensional feature data and is insensitive to non-linear relationships between features and feature scaling, making it suitable for this scenario.
[0056] The number of input layer nodes in a random forest model equals the dimension of the input feature vector, i.e., the number of feature categories contained in the user feature set. For example, if the user features include 20 feature categories such as age, gene expression levels, and imaging indicators, the dimension of the normalized and concatenated feature vector is 20, and the number of input nodes in the random forest model is 20. The main hyperparameters of the model include the number of decision trees and the maximum number of features. The number of decision trees controls the number of trees in the ensemble model; a larger number results in a more stable model but also higher computational cost. The maximum number of features controls the number of features randomly selected by each tree during splitting, used to reduce the correlation between trees. Relevant hyperparameters are optimized through grid search combined with cross-validation. For example, the search is performed within the range of 50, 100, and 200 for the number of decision trees and one-third or the maximum number of features for the maximum number of features, selecting the parameter combination that yields the highest accuracy on the validation set. The model output layer represents the disease category prediction results, and the number of output nodes equals the total number of disease categories, M.
[0057] Then, multiple user feature sets are used as training data and validation data, and multiple recognition results are used as supervision labels and validation labels to train and validate the basic model architecture. After the validation accuracy is qualified, the first recognition model is obtained.
[0058] Specifically, due to the limited number of original samples, hold-out or cross-validation methods can be used for training and validation. Taking five-fold cross-validation as an example, all samples are randomly divided into five equal parts. Each time, four parts are used as the training set and one part as the validation set, repeated five times, using a different validation set each time. On the training set, the user feature set is used as the input feature vector, and the corresponding recognition result is used as the supervision label. The model parameters are iteratively updated by minimizing the classification loss function through an optimization algorithm. After each training iteration, the classification accuracy is calculated on the validation set to evaluate the model's generalization performance. When the average accuracy of the five cross-validations reaches a preset passing threshold, for example, an accuracy of 85% or higher, the model is considered valid, the current model parameters are saved, and the first recognition model is obtained. If the average accuracy does not reach the passing threshold, the model hyperparameters are adjusted or feature engineering is performed, and training and validation are repeated.
[0059] Through the above training and verification process, it can be ensured that the first recognition model has basic recognition capabilities on the original small sample data, providing a reliable initial model for subsequent synthetic data augmentation learning.
[0060] S30: Based on the inter-class differences, configure the user feature synthesis strategy to generate multiple synthesized user feature sets and multiple synthesized recognition results, perform reinforcement learning on the first recognition model, and obtain the second recognition model. The user feature synthesis strategy includes the number of synthesized features and the degree of synthesis deviation.
[0061] Furthermore, based on the inter-class dissimilarity values obtained from the aforementioned analysis, user feature synthesis strategies are configured. Different disease categories exhibit different distribution characteristics in the feature space. Categories with high inter-class dissimilarity values have features that are clearly distinguishable from other categories, and the model already possesses good recognition capabilities, requiring only a small number of synthetic samples for mild enhancement. Categories with low inter-class dissimilarity values have features that overlap with other categories or have blurred boundaries, making the model prone to confusion. Therefore, more synthetic samples need to be generated, and the feature deviation of the synthetic samples should be relatively large to expand the feature coverage of the category and strengthen the model's recognition of its discrimination boundaries.
[0062] Specifically, the user feature synthesis strategy includes two core parameters: synthesis quantity and synthesis deviation. The synthesis quantity refers to the number of synthesized samples generated for each disease category; the smaller the inter-class variability, the larger the synthesis quantity. The synthesis deviation refers to the degree to which the features of the synthesized samples can be shifted relative to the original feature distribution; the smaller the inter-class variability, the larger the synthesis deviation.
[0063] Furthermore, based on the configured number of synthesized features and the degree of synthesization deviation, a corresponding synthesized user feature set is generated for each disease category. Simultaneously, a corresponding recognition result label, i.e., the disease category to which the synthesized feature belongs, is assigned to each synthesized user feature, forming a synthesized recognition result. The generated multiple synthesized user feature sets and multiple synthesized recognition results are used as augmented training data, merged with the original training data, and used to retrain the first recognition model to obtain the second recognition model.
[0064] Compared with the first recognition model, the second recognition model is trained based on differentially synthesized enhanced training data. While maintaining the ability to recognize easily distinguishable categories, it improves the recognition accuracy of easily confused categories. It can effectively reduce the misclassification rate of the model in disease categories with small inter-class differences and blurred feature boundaries, thereby enhancing the generalization performance and diagnostic reliability of the neurogenetic disease recognition model in real clinical small sample scenarios.
[0065] Specifically, based on the inter-class differences, a user feature synthesis strategy is configured to generate multiple synthesized user feature sets and multiple synthesized recognition results. Reinforcement learning is then performed on the first recognition model to obtain a second recognition model, including:
[0066] Calculate the mean of the inter-class dissimilarity to obtain the average inter-class dissimilarity.
[0067] Calculate the average number of user features within the multiple user feature sets and configure it as the baseline synthesis quantity;
[0068] Based on the ratio of each inter-class difference to the average inter-class difference, the baseline synthesis quantity is adjusted and calculated to obtain multiple synthesis quantities;
[0069] Calculate the reciprocal of the ratio of each inter-class difference to the average inter-class difference to obtain multiple difference correction coefficients, where the minimum difference correction coefficient is 1;
[0070] Multiple difference correction coefficients are used to correct multiple inter-class differences, resulting in multiple synthesis deviations. Based on the multiple synthesis deviations and multiple synthesis quantities, multiple synthetic user feature sets and multiple synthetic recognition results are generated.
[0071] By using multiple synthetic user feature sets and multiple synthetic recognition results, reinforcement learning is performed on the first recognition model to obtain the second recognition model.
[0072] First, calculate the mean of the inter-class dissimilarity scores to obtain the average inter-class dissimilarity score. The average inter-class dissimilarity score is the arithmetic mean of the inter-class dissimilarity scores of all disease categories, reflecting the average difficulty of distinguishing between the categories as a whole. For example, if there are three disease categories, namely Huntington's disease, spinal muscular atrophy, and hereditary ataxia, with inter-class dissimilarity scores of 0.85, 0.45, and 0.50 respectively, then the average inter-class dissimilarity score is 0.60.
[0073] Secondly, the average number of user features across multiple user feature sets is calculated and configured as the baseline synthesis quantity. The number of patient samples in each category of the original user feature sets may differ; therefore, the average number of samples across all categories is calculated as the baseline synthesis quantity. This baseline synthesis quantity serves as a basic reference value for the number of synthesized samples in each disease category. Subsequently, this baseline value is adjusted based on the inter-class variability among each category, ensuring that categories with lower inter-class variability receive more synthesized samples, and categories with higher inter-class variability receive fewer synthesized samples. For example, if the three categories have 8, 6, and 7 samples respectively, the average number is 7, and the baseline synthesis quantity is configured as 7.
[0074] Furthermore, based on the ratio of each inter-class dissimilarity to the average inter-class dissimilarity, the baseline number of synthesized samples is adjusted and calculated to obtain multiple synthesized samples. The ratio of inter-class dissimilarity to the average inter-class dissimilarity reflects the degree to which the difficulty of distinguishing a class deviates from the average level. A ratio less than 1 indicates that the inter-class dissimilarity of that class is lower than the average level, meaning that this class is more easily confused with other classes, and more synthesized samples need to be generated; a ratio greater than 1 indicates that the inter-class dissimilarity of that class is higher than the average level, meaning that this class is clearly distinguishable from other classes, and fewer synthesized samples can be generated.
[0075] Specifically, the formula for calculating the number of syntheses is: the number of syntheses equals the baseline number of syntheses multiplied by the average inter-class variability divided by the inter-class variability of that class. For example, the baseline number of syntheses is 7, and the average inter-class variability is 0.60. For the Huntington's disease category, with an inter-class variability of 0.85, the number of syntheses is calculated as 7 × 0.60 ÷ 0.85 ≈ 4.94, rounded down to 5; for the spinal muscular atrophy category, with an inter-class variability of 0.45, the number of syntheses is calculated as 7 × 0.60 ÷ 0.45 ≈ 9.33, rounded down to 9; for the hereditary ataxia category, with an inter-class variability of 0.50, the number of syntheses is calculated as 7 × 0.60 ÷ 0.50 = 8.4, rounded down to 8.
[0076] Furthermore, the reciprocal of the ratio of each inter-class difference to the average inter-class difference is calculated to obtain multiple difference correction coefficients, where the minimum difference correction coefficient is 1.
[0077] Specifically, the difference correction coefficient is equal to the average inter-class difference divided by the inter-class difference of that class. When the inter-class difference is less than the average inter-class difference, the ratio is less than 1, the reciprocal is greater than 1, and the difference correction coefficient is greater than 1. When the inter-class difference is greater than the average inter-class difference, the reciprocal is less than 1, but a minimum value of 1 is set to ensure that the difference correction coefficient is not lower than 1, thus avoiding reducing the composite deviation of classes with large inter-class differences.
[0078] For example, let the average inter-class variability be 0.60. For the Huntington's disease category, the inter-class variability is 0.85, and the variability correction coefficient is calculated as 0.60 ÷ 0.85 ≈ 0.71. Since the minimum value is set to 1, the value is taken as 1. For the spinal muscular atrophy category, the inter-class variability is 0.45, and the variability correction coefficient is calculated as 0.60 ÷ 0.45 ≈ 1.33. For the hereditary ataxia category, the inter-class variability is 0.50, and the variability correction coefficient is calculated as 0.60 ÷ 0.50 = 1.20.
[0079] Furthermore, multiple difference correction coefficients are used to correct multiple inter-class differences, resulting in multiple composite deviations. The composite deviation equals the inter-class difference for that class multiplied by the difference correction coefficient. Since the minimum value of the difference correction coefficient is 1, the composite deviation is not less than the original inter-class difference. For classes with smaller inter-class differences, multiplying by a difference correction coefficient greater than 1 significantly increases the composite deviation, requiring the generated composite sample to differ more from the features of other classes. This pushes the composite sample further away from the class boundary, enhancing the model's ability to discriminate that class.
[0080] For example, the inter-class variance for Huntington's disease is 0.85, the variance correction coefficient is 1, and the composite deviation is 0.85; the inter-class variance for spinal muscular atrophy is 0.45, the variance correction coefficient is 1.33, and the composite deviation is 0.60; and the inter-class variance for hereditary ataxia is 0.50, the variance correction coefficient is 1.20, and the composite deviation is 0.60.
[0081] Furthermore, based on multiple synthesis deviations and multiple synthesis quantities, multiple synthetic user feature sets and multiple synthetic recognition results are generated. For each disease category, according to its synthesis quantity, a specified number of candidate synthetic features are generated within the original feature range, and features whose differences from other category features meet the synthesis deviation requirements are selected as the final synthetic user features. Simultaneously, a corresponding disease category label is assigned to this synthetic feature as the synthetic recognition result. It should be noted that the feature difference refers to the distance between the overall feature vector of the candidate synthetic user feature and the feature vectors of other categories, rather than the difference in a single feature dimension. Taking the spinal muscular atrophy category as an example, a candidate synthetic user feature includes multiple feature values such as age, gene expression level, and imaging indicators. Its difference from the Huntington's disease category is obtained by calculating the Euclidean distance between the two complete feature vectors. This distance comprehensively reflects the degree of difference across all feature dimensions. When a difference of 0.60 or higher is required, it means that the overall performance of the candidate synthetic user feature across all feature dimensions is far removed from the feature distribution of other categories.
[0082] Specifically, based on multiple synthesis deviations and multiple synthesis quantities, multiple synthesis user feature sets and multiple synthesis recognition results are generated, including:
[0083] Multiple user feature ranges are constructed based on user feature endpoint values within multiple user feature sets.
[0084] Within multiple user feature ranges, multiple first synthetic user feature sets are randomly generated according to multiple synthesis quantities. The difference between the first synthetic user features in each first synthetic user feature set and other first synthetic user feature sets is calculated to obtain multiple first synthetic deviation sets.
[0085] Each time, it is determined whether the first synthetic deviation is greater than or equal to the corresponding synthetic deviation. If it is, the corresponding first synthetic user feature is retained. If not, the second synthetic user feature is regenerated until the difference between the first synthetic user feature set and the other first synthetic user feature set is greater than or equal to the synthetic deviation, thus obtaining multiple synthetic user feature sets.
[0086] Multiple synthetic user feature sets are labeled with recognition results to obtain multiple synthetic recognition results.
[0087] First, multiple user feature ranges are constructed based on the endpoint values of user features within multiple user feature sets. Specifically, for each disease category, its original user feature set contains feature vectors of multiple patients, with a minimum and maximum value in each feature dimension. Using the minimum and maximum values of all original user features for that category in each feature dimension as endpoints, the value range of that category in the feature space is constructed.
[0088] For example, for the spinal muscular atrophy category, if the minimum value of the age feature is 2 years and the maximum value is 15 years, then the range of the age dimension is 2 to 15; if the minimum value of the gene expression level feature is 0.3 and the maximum value is 0.8, then the range of the gene expression level dimension is 0.3 to 0.8. The value ranges of all feature dimensions together constitute the user feature range for this category.
[0089] Secondly, within multiple user feature ranges, multiple first synthetic user feature sets are randomly generated according to multiple synthesis quantities. Specifically, for each disease category, a specified number of candidate synthetic user features are generated within its user feature range based on the synthesis quantity configured for that category, forming the first synthetic user feature set. During the generation process, the physiological or clinical relevance constraints between each feature dimension must be followed. Since there may be intrinsic correlations between different feature dimensions—for example, for spinal muscular atrophy, there is a correlation between age and gene expression levels: patients with onset in infancy often have higher levels of abnormal gene expression, while patients with onset in adulthood have relatively lower levels of abnormal gene expression—ignoring this correlation and generating features randomly may result in invalid samples where the age is 50 years old but the gene expression level is at an infantile level, misleading the model to learn non-existent feature patterns.
[0090] To address this issue, a conditional sampling strategy is employed when generating synthetic user features. First, the constraints between relevant feature pairs are statistically analyzed from the original user feature set. For example, age is divided into infancy, childhood, and adulthood, and the mean and standard deviation of gene expression levels for each age range are calculated. During generation, a value for one feature is randomly generated, its corresponding range is determined, and then the value for another feature is generated based on the distribution parameters of that range. For feature pairs without a clear correlation, independent sampling is used, uniformly and randomly generating features within the range of values for each feature dimension. This constrained sampling ensures that the generated synthetic user features conform to real physiological and clinical patterns.
[0091] Each feature dimension value of the first synthetic user feature is generated within the feature range of its respective category according to the constrained sampling method described above. For example, if the number of synthesized features for the spinal muscular atrophy category is 9, then 9 candidate synthetic user features are generated within the feature range of this category according to the constrained sampling method, constituting the first synthetic user feature set.
[0092] Then, the degree of difference between each first synthetic user feature set and other first synthetic user feature sets is calculated to obtain multiple sets of first synthetic deviations. For each generated candidate synthetic user feature, it is necessary to evaluate the degree of difference between it and the original user features of other disease categories.
[0093] Specifically, the dissimilarity of the candidate synthetic user feature is calculated by comparing it with the original user feature sets of each other category, and the minimum or average value is taken as the dissimilarity of the candidate synthetic user feature. Optionally, the dissimilarity can be calculated using metrics such as Euclidean distance, Manhattan distance, or Mahalanobis distance. For example, for a candidate synthetic user feature in the spinal muscular atrophy category, the minimum Euclidean distance between it and the original user feature sets of the Huntington's disease category and the hereditary ataxia category is calculated, and the smaller value is taken as the dissimilarity of the candidate feature.
[0094] Further, it is determined whether the first synthetic deviation is greater than or equal to the corresponding synthetic deviation. For each candidate synthetic user feature, its calculated difference is compared with the synthetic deviation configured for that category. If the difference is greater than or equal to the synthetic deviation, it indicates that the candidate synthetic user feature is sufficiently far from features of other categories and will not fall into the ambiguous region of category boundaries, so the candidate synthetic user feature is retained as a valid synthetic user feature. If the difference is less than the synthetic deviation, it indicates that the candidate synthetic user feature is too close to features of other categories and may be located near the category boundary, which could easily lead to model confusion, so the candidate feature is rejected and a new candidate synthetic user feature is generated.
[0095] The generation and judgment process is repeated until the number of valid synthetic user features for that category reaches the configured requirement for the number of synthetic features. This deviation filtering mechanism ensures that all synthetic user features meet the difference requirements from other categories, thus avoiding the generation of samples with ambiguous boundaries.
[0096] Furthermore, the recognition results of multiple synthetic user feature sets are labeled to obtain multiple synthetic recognition results. Each valid synthetic user feature is generated based on the feature range of its category, so its recognition result label is the category itself. For example, a synthetic user feature generated within the feature range of the spinal muscular atrophy category is labeled as spinal muscular atrophy.
[0097] Synthetic user features and their corresponding recognition result labels for all categories are aggregated to form a synthetic user feature set and synthetic recognition results, which are then used for subsequent reinforcement learning training. Specifically, through the above generation and screening process, the synthetic samples maintain a feature distribution similar to the original samples, while the deviation constraint ensures that the synthetic samples between different categories have sufficient discriminative power, effectively avoiding the model confusion problem caused by samples with blurred boundaries in traditional synthesis methods.
[0098] Finally, multiple synthetic user feature sets and multiple synthetic recognition results are used to perform reinforcement learning on the first recognition model to obtain the second recognition model. Specifically, the generated synthetic samples are merged with the original samples to form an expanded training set, which is then used to retrain the first recognition model. Because the synthetic samples specifically strengthen easily confused categories with small inter-class differences, and the quality of the synthetic samples is ensured through synthetic deviation constraints, the second recognition model after reinforcement learning has clearer discrimination boundaries for easily confused categories, thus improving recognition accuracy.
[0099] Specifically, by using multiple synthetic user feature sets and multiple synthetic recognition results, reinforcement learning is performed on the first recognition model to obtain a second recognition model, including:
[0100] The multiple synthetic user feature sets and multiple synthetic recognition results are used as enhanced training data and enhanced supervision labels;
[0101] Using the enhanced training data and enhanced supervision labels, the first recognition model is subjected to reinforcement learning to obtain the second recognition model. If the accuracy decreases during training, the model network parameters are rolled back.
[0102] First, multiple synthetic user feature sets and multiple synthetic recognition results are used as augmented training data and augmented supervision labels. In the preceding steps, synthetic user features that meet the synthetic deviation requirements were generated for each disease category, and each synthetic user feature was labeled with a corresponding disease category label. The synthetic user feature sets are used as augmented training data, and the corresponding synthetic recognition results are used as augmented supervision labels to further train the first recognition model. The augmented training data is merged with the original training data to form the expanded training set.
[0103] Specifically, augmented training data and augmented supervision labels are used to perform reinforcement learning on the first recognition model to obtain the second recognition model. Reinforcement learning refers to the process of further optimizing model parameters using newly added training data based on an existing model. Specifically, the current parameters of the first recognition model are used as initial parameters, the expanded training set is input into the model, the loss value between the model's prediction results and the true labels is calculated, and the model parameters are iteratively updated through backpropagation and an optimizer. Because the synthetic samples specifically strengthen easily confused categories with small inter-class differences, and the quality of the synthetic samples is ensured through synthetic deviation constraints, the discrimination boundary of the model after reinforcement learning will be clearer for easily confused categories.
[0104] Furthermore, if accuracy decreases during training, the model network parameters should be rolled back. During reinforcement learning, it's crucial to continuously monitor changes in model accuracy on the validation set. The validation set typically uses original real samples to ensure objectivity in the evaluation. If, after a training round, the validation set accuracy decreases compared to the previous round, it indicates that the newly added synthetic samples are of poor quality or overfitting the distribution of the synthetic samples, negatively impacting the model's generalization performance. In this case, instead of saving the model parameters from the current round, the model parameters that achieved the highest accuracy in the previous round should be rolled back, and further training should be stopped or training should continue after reducing the learning rate.
[0105] This fallback mechanism ensures that the performance of the second recognition model after reinforcement learning is at least as good as that of the first recognition model, avoiding model degradation caused by fluctuations in the quality of synthetic samples. After reinforcement learning, a second recognition model is obtained, which maintains the ability to recognize easily distinguishable categories while improving the recognition accuracy of easily confused categories.
[0106] In summary, the aforementioned differentiated synthesis strategy achieves precise allocation of synthetic data resources. Categories with smaller inter-class differences receive more synthetic samples with greater deviations, effectively compensating for the insufficient ability to identify easily confused categories under small sample conditions. Simultaneously, the synthesis deviation screening mechanism ensures that synthetic samples are far from category boundary regions, avoiding interference from blurred-boundary samples in traditional synthesis methods. Finally, through reinforcement learning, a second recognition model with clearer boundaries and higher overall recognition accuracy is obtained for easily confused categories.
[0107] S40: Test the combined recognition error rate of the second recognition model for multiple recognition results, obtain multiple combined recognition error rates, combine the multiple inter-class differences, configure the user feature synthesis strategy, continue to generate multiple synthetic user feature sets and multiple synthetic recognition results, perform reinforcement learning, and obtain the third recognition model.
[0108] Finally, after completing the first round of reinforcement learning to obtain the second recognition model, it is necessary to further test whether there are still easily confused category pairs in the actual recognition process, and to carry out a second round of differential synthesis and reinforcement learning for these easily confused category pairs in order to continuously optimize the model performance.
[0109] Specifically, the second recognition model is tested for the combined recognition error rate of multiple recognition results to obtain multiple combined recognition error rates. Combined with the multiple inter-class differences, a user feature synthesis strategy is configured to continue generating multiple synthetic user feature sets and multiple synthetic recognition results. Reinforcement learning is then performed to obtain a third recognition model, including:
[0110] When the second recognition model identifies user features of the first recognition result, the error rate of identifying them as other recognition results is used as the first combined recognition error rate.
[0111] Continue testing the recognition error rate of multiple combinations of multiple recognition results;
[0112] Based on multiple combined recognition error rates and multiple inter-class differences, user feature synthesis strategy configuration is performed to obtain multiple secondary synthesis quantities and secondary synthesis deviations;
[0113] Based on the number of secondary synthesiss and the deviation of the secondary synthesis, new multiple synthetic user feature sets and multiple synthetic recognition results are generated. The second recognition model is then subjected to reinforcement learning to obtain the third recognition model.
[0114] First, the second recognition model is tested for the combined recognition error rate of multiple recognition results, resulting in multiple combined recognition error rates. The combined recognition error rate refers to the probability that the model incorrectly identifies a user feature of one category as another specific category. For example, for test samples in the spinal muscular atrophy category, the probability that the model incorrectly identifies it as Huntington's disease, the probability of incorrectly identifying it as hereditary ataxia, etc., are statistically analyzed.
[0115] For example, Huntington's disease was selected as the first identification result, and 10 user feature samples from its test set were input into the second identification model for identification. Statistics showed that one sample was incorrectly identified as spinal muscular atrophy, two samples were incorrectly identified as hereditary ataxia, and a total of three samples were incorrectly identified as other results. Therefore, the first combination identification error rate for Huntington's disease was 30%. This first combination identification error rate reflects the overall degree to which the model confuses Huntington's disease with other disease categories.
[0116] Continue testing multiple recognition results, iterating through all disease categories, and testing the combined recognition error rate of each category when it is misclassified as other categories, obtaining multiple combined recognition error rates. For example, for three disease categories, Huntington's disease has 10 test samples, of which 3 were misclassified as other diseases, so its combined recognition error rate is 30%; spinal muscular atrophy has 10 test samples, of which 4 were misclassified as other diseases, so its combined recognition error rate is 40%; hereditary ataxia has 10 test samples, of which 2 were misclassified as other diseases, so its combined recognition error rate is 20%.
[0117] By analyzing the combined recognition error rates described above, categories that are more difficult to identify using the current model can be identified. For example, the combined recognition error rate for spinal muscular atrophy is 40%, indicating that the recognition accuracy for this category is low and it should be the focus of optimization in the second round of reinforcement learning.
[0118] Furthermore, based on multiple combined recognition error rates and multiple inter-class dissimilarity levels, a user feature synthesis strategy is configured to obtain multiple secondary synthesis quantities and secondary synthesis deviations. The combined recognition error rate reflects the degree of confusion exposed by the second recognition model in the actual recognition process, and effectively supplements and corrects the original inter-class dissimilarity level. For class pairs with high combined recognition error rates, it indicates that the model's discriminative ability on that class pair is still insufficient, and it is necessary to further reduce the effective inter-class dissimilarity level of that class, thereby allocating more synthetic samples and a larger synthesis deviation in the second round of synthesis.
[0119] Specifically, based on multiple combined recognition error rates and multiple inter-class differences, a user feature synthesis strategy is configured to obtain multiple secondary synthesis quantities and secondary synthesis deviations, including:
[0120] Calculate the ratio of multiple combined recognition error rates to the mean of multiple combined recognition error rates, and perform correction calculations on multiple inter-class differences to obtain multiple corrected inter-class differences.
[0121] Based on the differences between multiple correction classes, calculate the number of secondary syntheses and the deviation of secondary syntheses.
[0122] First, the ratios of multiple combined recognition error rates to the mean of multiple combined recognition error rates are calculated separately. Corrected inter-class dissimilarity is then calculated to obtain multiple corrected inter-class dissimilarity. The combined recognition error rate reflects the accuracy of the second recognition model in identifying each category during the actual recognition process. For categories with high recognition error rates, it indicates that these categories are still easily confused with other categories under the current model, and their effective inter-class dissimilarity needs to be further reduced, thereby allocating more synthetic samples and a larger synthetic deviation in the second round of synthesis.
[0123] Specifically, the correction calculation is performed as follows: calculate the average error rate of all category combinations, divide the error rate of each category combination by this average, and obtain the error rate ratio for that category. The corrected inter-class dissimilarity is equal to the original inter-class dissimilarity divided by the error rate ratio.
[0124] If the combined recognition error rate of a certain category is higher than the average level, the error rate ratio is greater than 1, and the corrected inter-class variability is less than the original inter-class variability; if it is lower than the average level, the error rate ratio is less than 1, and the corrected inter-class variability is greater than the original inter-class variability.
[0125] For example, the original inter-class variances for the three disease categories are: Huntington's disease 0.85, spinal muscular atrophy 0.45, and hereditary ataxia 0.50. The combined identification error rates for the three categories are 30%, 40%, and 20%, respectively, with an average of 30%. The error rate ratio for Huntington's disease is 30% divided by 30%, which equals 1, and the corrected inter-class variance is equal to the original inter-class variance of 0.85 divided by 1, which equals 0.85; the error rate ratio for spinal muscular atrophy is 40% divided by 30%, which is approximately 1.33, and the corrected inter-class variance is 0.45 divided by 1.33, which equals approximately 0.34; the error rate ratio for hereditary ataxia is 20% divided by 30%, which is approximately 0.67, and the corrected inter-class variance is 0.50 divided by 0.67, which equals approximately 0.75.
[0126] Through the above corrections, the categories with high error rates are identified with smaller corrected inter-class differences, thereby obtaining more synthesis results and greater synthesis deviation in the second round of synthesis.
[0127] Furthermore, based on multiple corrected inter-class differences, multiple secondary synthesis quantities and secondary synthesis deviations are calculated and configured. The calculation methods for the secondary synthesis quantities and secondary synthesis deviations are the same as in the first round. First, the average value of the corrected inter-class differences and the average number of original user features are calculated as the baseline synthesis quantity. Then, based on the ratio of the corrected inter-class differences for each category to the average corrected inter-class differences, the baseline synthesis quantity is adjusted to obtain the secondary synthesis quantity: the smaller the corrected inter-class differences for a category, the larger the synthesis quantity. Simultaneously, the reciprocal of the ratio of the corrected inter-class differences to the average corrected inter-class differences is calculated to obtain the difference correction coefficient, which is multiplied by the corrected inter-class differences to obtain the secondary synthesis deviation: the smaller the corrected inter-class differences for a category, the larger the synthesis deviation.
[0128] For example, the corrected inter-class dissimilarity rates for the three disease categories are: Huntington's disease 0.85, spinal muscular atrophy 0.34, and hereditary ataxia 0.75. The mean of the corrected inter-class dissimilarity rate is calculated to be 0.65. The average number of original user features is 7, which serves as the baseline synthesis number.
[0129] For Huntington's disease, the corrected inter-class variability of 0.85 is greater than the mean of 0.65, the ratio is approximately 1.31, and the number of secondary composites is 7 divided by 1.31, which is approximately 5.34, rounded down to 5. The difference correction coefficient is the mean of 0.65 divided by the corrected inter-class variability of 0.85, which is approximately 0.76, rounded down to the minimum value of 1. The deviation of the secondary composite is 0.85 multiplied by 1, which equals 0.85.
[0130] For spinal muscular atrophy, the corrected inter-class variability of 0.34 is less than the mean of 0.65, the ratio is approximately 0.52, and the number of secondary composites is 7 divided by 0.52, which is approximately 13.46, rounded down to 13. The difference correction coefficient is 0.65 divided by 0.34, which is approximately 1.91, and the deviation of secondary composites is 0.34 multiplied by 1.91, which is approximately 0.65.
[0131] For hereditary ataxia, the corrected interclass difference of 0.75 is greater than the mean of 0.65, the ratio is approximately 1.15, and the number of secondary synthesis is 7 divided by 1.15, which is approximately 6.09, rounded down to 6. The difference correction coefficient is 0.65 divided by 0.75, which is approximately 0.87, and the minimum value of 1 is taken as 1. The deviation of secondary synthesis is 0.75 multiplied by 1, which is 0.75.
[0132] With the above configuration, the second round of synthesis will specifically enhance the easily confused categories that still have low recognition accuracy in the second recognition model, so that the model performance will continue to improve in iterative optimization.
[0133] Finally, based on the number of secondary syntheses and their deviation, multiple new synthetic user feature sets and multiple synthetic recognition results are generated. The second recognition model is then subjected to reinforcement learning to obtain the third recognition model. Specifically, following the same synthesis process as the first round, new synthetic user features are generated based on the number of secondary syntheses, and deviation is filtered based on the deviation of the secondary syntheses to ensure that the newly generated synthetic samples meet higher deviation requirements.
[0134] For example, the number of secondary synthesized features for the spinal muscular atrophy (SMA) category is 13, with a secondary synthesis deviation of 0.65; the number of secondary synthesized features for the Huntington's disease category is 5, with a secondary synthesis deviation of 0.85; and the number of secondary synthesized features for the hereditary ataxia category is 6, with a secondary synthesis deviation of 0.75. Based on these parameters, new synthetic user features are generated within each user feature range, and features whose differences from other category features meet the synthesis deviation requirement are selected to form the second round of synthetic user feature sets. For example, for the SMA category, 13 synthetic samples need to be generated, and the difference between each synthetic sample and other category features must reach at least 0.65.
[0135] The newly generated synthetic samples are merged with the original samples and the first round of synthetic samples to perform a second round of reinforcement learning on the second recognition model. Simultaneously, the accuracy on the validation set is monitored during training; if the accuracy decreases, the model parameters are rolled back. After the second round of reinforcement learning, the model's ability to distinguish between easily confused category pairs is further enhanced, resulting in a third recognition model. This third recognition model further strengthens its ability to distinguish between easily confused category pairs, and the recognition accuracy is continuously improved.
[0136] In summary, the embodiments of this application have at least the following technical effects:
[0137] This invention first analyzes the inter-class differences by acquiring multiple recognition results and multiple user feature sets, quantifying the difficulty of distinguishing features between different categories of neurogenetic diseases, and providing a quantitative basis for subsequent differential data synthesis. Second, a first recognition model is trained based on the original small sample data, establishing basic recognition capabilities. Third, a user feature synthesis strategy is adaptively configured according to the inter-class differences, generating more synthetic samples with greater deviations for easily confused categories with small inter-class differences, and generating fewer synthetic samples with smaller deviations for categories with large inter-class differences. This allows the synthesized data to specifically enhance the model's ability to distinguish easily confused categories, avoiding the boundary ambiguity problem caused by blind synthesis.
[0138] Finally, by testing the combined recognition error rate of the second recognition model, the class pairs that are still easily confused by the model are identified. The synthesis strategy is then reconfigured based on inter-class dissimilarity, and reinforcement learning continues, forming a multi-stage iterative optimization mechanism to continuously improve the model's recognition accuracy on difficult samples. This invention solves the problems of poor model generalization performance under small sample conditions and the exacerbation of inter-class confusion by traditional synthesis methods, achieving adaptive reinforcement construction of a neurogenetic disease recognition model.
[0139] Example 2, as Figure 2 As shown, based on the same inventive concept as the neurogenetic disease identification model construction method provided in Embodiment 1, this embodiment of the invention also provides a neurogenetic disease identification model construction system, including:
[0140] The difference analysis module 11 is used to acquire multiple identification results and multiple user feature sets for neurogenetic disease identification, perform inter-class difference analysis, and obtain multiple inter-class differences.
[0141] The first training module 12 is used to train a first recognition model for the recognition of neurogenetic diseases based on the multiple recognition results and multiple user feature sets.
[0142] The reinforcement learning module 13 is used to configure the user feature synthesis strategy based on multiple inter-class differences, generate multiple synthesized user feature sets and multiple synthesized recognition results, perform reinforcement learning on the first recognition model, and obtain a second recognition model. The user feature synthesis strategy includes the number of synthesis and the degree of synthesis deviation.
[0143] The iterative optimization module 14 is used to test the combined recognition error rate of the second recognition model for multiple recognition results, obtain multiple combined recognition error rates, combine the multiple inter-class differences, configure the user feature synthesis strategy, continue to generate multiple synthetic user feature sets and multiple synthetic recognition results, perform reinforcement learning, and obtain the third recognition model.
[0144] The difference analysis module 11 is specifically used for:
[0145] Specifically, multiple identification results and multiple user feature sets for neurogenetic disease identification are obtained, and inter-class dissimilarity analysis is performed to obtain multiple inter-class dissimilarity sets, including:
[0146] Multiple identification results and multiple user feature sets for neurogenetic disease identification are obtained, where the multiple identification results and multiple user feature sets are small sample data;
[0147] For each user feature set of the recognition result, the difference magnitude between the user feature sets of other recognition results and the user feature sets of all other recognition results is calculated to obtain multiple inter-class differences.
[0148] Furthermore, for each user feature set of the recognition result, the magnitude of the difference between the user feature sets of all other recognition results is calculated to obtain multiple inter-class dissimilarity measures, including:
[0149] Select the first user feature set of the first recognition result, calculate the difference magnitude of each first user feature from each user feature in the other all user feature sets, and calculate the mean to obtain the inter-class difference degree;
[0150] Continue calculating the inter-class dissimilarity of the user feature set for each other recognition result to obtain multiple inter-class dissimilarity values.
[0151] The first training module 12 is specifically used for:
[0152] Based on the multiple recognition results and multiple user feature sets, a first recognition model for identifying neurogenetic diseases is trained, including:
[0153] Based on machine learning, we construct a basic model architecture for the identification of neurogenetic diseases.
[0154] The basic model architecture is trained and validated using the multiple user feature sets as training and validation data, and multiple recognition results as supervision and validation labels. After the validation accuracy is qualified, the first recognition model is obtained.
[0155] Specifically, the reinforcement learning module 13 is used for:
[0156] Based on multiple inter-class differences, a user feature synthesis strategy is configured to generate multiple synthesized user feature sets and multiple synthesized recognition results. Reinforcement learning is then performed on the first recognition model to obtain a second recognition model, including:
[0157] Calculate the mean of the inter-class dissimilarity to obtain the average inter-class dissimilarity.
[0158] Calculate the average number of user features within the multiple user feature sets and configure it as the baseline synthesis quantity;
[0159] Based on the ratio of each inter-class difference to the average inter-class difference, the baseline synthesis quantity is adjusted and calculated to obtain multiple synthesis quantities;
[0160] Calculate the reciprocal of the ratio of each inter-class difference to the average inter-class difference to obtain multiple difference correction coefficients, where the minimum difference correction coefficient is 1;
[0161] Multiple difference correction coefficients are used to correct multiple inter-class differences, resulting in multiple synthesis deviations. Based on the multiple synthesis deviations and multiple synthesis quantities, multiple synthetic user feature sets and multiple synthetic recognition results are generated.
[0162] By using multiple synthetic user feature sets and multiple synthetic recognition results, reinforcement learning is performed on the first recognition model to obtain the second recognition model.
[0163] Specifically, based on multiple synthesis deviations and multiple synthesis quantities, multiple synthesis user feature sets and multiple synthesis recognition results are generated, including:
[0164] Multiple user feature ranges are constructed based on user feature endpoint values within multiple user feature sets.
[0165] Within multiple user feature ranges, multiple first synthetic user feature sets are randomly generated according to multiple synthesis quantities. The difference between the first synthetic user features in each first synthetic user feature set and other first synthetic user feature sets is calculated to obtain multiple first synthetic deviation sets.
[0166] Each time, it is determined whether the first synthetic deviation is greater than or equal to the corresponding synthetic deviation. If it is, the corresponding first synthetic user feature is retained. If not, the second synthetic user feature is regenerated until the difference between the first synthetic user feature set and the other first synthetic user feature set is greater than or equal to the synthetic deviation, thus obtaining multiple synthetic user feature sets.
[0167] Multiple synthetic user feature sets are labeled with recognition results to obtain multiple synthetic recognition results.
[0168] Specifically, by using multiple synthetic user feature sets and multiple synthetic recognition results, reinforcement learning is performed on the first recognition model to obtain a second recognition model, including:
[0169] The multiple synthetic user feature sets and multiple synthetic recognition results are used as enhanced training data and enhanced supervision labels;
[0170] Using the enhanced training data and enhanced supervision labels, the first recognition model is subjected to reinforcement learning to obtain the second recognition model. If the accuracy decreases during training, the model network parameters are rolled back.
[0171] Specifically, the iterative optimization module 14 is used for:
[0172] The second recognition model is tested for the combined recognition error rate of multiple recognition results to obtain multiple combined recognition error rates. Combined with the multiple inter-class differences, a user feature synthesis strategy is configured to continue generating multiple synthetic user feature sets and multiple synthetic recognition results. Reinforcement learning is then performed to obtain a third recognition model, including:
[0173] When the second recognition model identifies user features of the first recognition result, the error rate of identifying them as other recognition results is used as the first combined recognition error rate.
[0174] Continue testing the recognition error rate of multiple combinations of multiple recognition results;
[0175] Based on multiple combined recognition error rates and multiple inter-class differences, user feature synthesis strategy configuration is performed to obtain multiple secondary synthesis quantities and secondary synthesis deviations;
[0176] Based on the number of secondary synthesiss and the deviation of the secondary synthesis, new multiple synthetic user feature sets and multiple synthetic recognition results are generated. The second recognition model is then subjected to reinforcement learning to obtain the third recognition model.
[0177] Specifically, based on multiple combined recognition error rates and multiple inter-class differences, a user feature synthesis strategy is configured to obtain multiple secondary synthesis quantities and secondary synthesis deviations, including:
[0178] Calculate the ratio of multiple combined recognition error rates to the mean of multiple combined recognition error rates, and perform correction calculations on multiple inter-class differences to obtain multiple corrected inter-class differences.
[0179] Based on the differences between multiple correction classes, calculate the number of secondary syntheses and the deviation of secondary syntheses.
Claims
1. A method for constructing a model for identifying neurogenetic diseases, characterized in that, The method includes: Multiple identification results and multiple user feature sets for neurogenetic disease identification are obtained, and inter-class difference analysis is performed to obtain multiple inter-class differences. Based on the multiple recognition results and multiple user feature sets, a first recognition model for the recognition of neurogenetic diseases is trained. Based on the inter-class differences, a user feature synthesis strategy is configured to generate multiple synthesized user feature sets and multiple synthesized recognition results. Reinforcement learning of the first recognition model is then performed to obtain a second recognition model. The user feature synthesis strategy includes the number of synthesized features and the degree of synthesis deviation. The second recognition model is tested for the combined recognition error rate of multiple recognition results to obtain multiple combined recognition error rates. Combined with the multiple inter-class differences, a user feature synthesis strategy is configured to continue generating multiple synthetic user feature sets and multiple synthetic recognition results. Reinforcement learning is then performed to obtain the third recognition model.
2. The method for constructing a neurogenetic disease identification model according to claim 1, characterized in that, Multiple identification results and multiple user feature sets for neurogenetic disease identification were obtained, and inter-class dissimilarity analysis was performed to obtain multiple inter-class dissimilarity sets, including: Multiple identification results and multiple user feature sets for neurogenetic disease identification are obtained, where the multiple identification results and multiple user feature sets are small sample data; For each user feature set of the recognition result, the difference magnitude between the user feature sets of other recognition results and the user feature sets of all other recognition results is calculated to obtain multiple inter-class differences.
3. The method for constructing a neurogenetic disease identification model according to claim 2, characterized in that, For each user feature set of the recognition result, the magnitude of the difference between the user feature sets of all other recognition results is calculated to obtain multiple inter-class dissimilarity measures, including: Select the first user feature set of the first recognition result, calculate the difference magnitude of each first user feature from each user feature in the other all user feature sets, and calculate the mean to obtain the inter-class difference degree; Continue calculating the inter-class dissimilarity of the user feature set for each other recognition result to obtain multiple inter-class dissimilarity values.
4. The method for constructing a neurogenetic disease identification model according to claim 1, characterized in that, Based on the multiple recognition results and multiple user feature sets, a first recognition model for identifying neurogenetic diseases is trained, including: Based on machine learning, we construct a basic model architecture for the identification of neurogenetic diseases. The basic model architecture is trained and validated using the multiple user feature sets as training and validation data, and multiple recognition results as supervision and validation labels. After the validation accuracy is qualified, the first recognition model is obtained.
5. The method for constructing a neurogenetic disease identification model according to claim 1, characterized in that, Based on multiple inter-class differences, a user feature synthesis strategy is configured to generate multiple synthesized user feature sets and multiple synthesized recognition results. Reinforcement learning is then performed on the first recognition model to obtain a second recognition model, including: Calculate the mean of the inter-class dissimilarity to obtain the average inter-class dissimilarity. Calculate the average number of user features within the multiple user feature sets and configure it as the baseline synthesis quantity; Based on the ratio of each inter-class difference to the average inter-class difference, the baseline synthesis quantity is adjusted and calculated to obtain multiple synthesis quantities; Calculate the reciprocal of the ratio of each inter-class difference to the average inter-class difference to obtain multiple difference correction coefficients, where the minimum difference correction coefficient is 1; Multiple difference correction coefficients are used to correct multiple inter-class differences, resulting in multiple synthesis deviations. Based on the multiple synthesis deviations and multiple synthesis quantities, multiple synthetic user feature sets and multiple synthetic recognition results are generated. By using multiple synthetic user feature sets and multiple synthetic recognition results, reinforcement learning is performed on the first recognition model to obtain the second recognition model.
6. The method for constructing a neurogenetic disease identification model according to claim 5, characterized in that, Based on multiple synthesis deviations and multiple synthesis quantities, multiple synthesis user feature sets and multiple synthesis recognition results are generated, including: Multiple user feature ranges are constructed based on user feature endpoint values within multiple user feature sets. Within multiple user feature ranges, multiple first synthetic user feature sets are randomly generated according to multiple synthesis quantities. The difference between the first synthetic user features in each first synthetic user feature set and other first synthetic user feature sets is calculated to obtain multiple first synthetic deviation sets. Each time, it is determined whether the first synthetic deviation is greater than or equal to the corresponding synthetic deviation. If it is, the corresponding first synthetic user feature is retained. If not, the second synthetic user feature is regenerated until the difference between the first synthetic user feature set and the other first synthetic user feature set is greater than or equal to the synthetic deviation, thus obtaining multiple synthetic user feature sets. Multiple synthetic user feature sets are labeled with recognition results to obtain multiple synthetic recognition results.
7. The method for constructing a neurogenetic disease identification model according to claim 5, characterized in that, By employing multiple synthetic user feature sets and multiple synthetic recognition results, reinforcement learning is performed on the first recognition model to obtain a second recognition model, including: The multiple synthetic user feature sets and multiple synthetic recognition results are used as enhanced training data and enhanced supervision labels; Using the enhanced training data and enhanced supervision labels, the first recognition model is subjected to reinforcement learning to obtain the second recognition model. If the accuracy decreases during training, the model network parameters are rolled back.
8. The method for constructing a neurogenetic disease identification model according to claim 1, characterized in that, The second recognition model is tested for the combined recognition error rate of multiple recognition results to obtain multiple combined recognition error rates. Combined with the multiple inter-class differences, a user feature synthesis strategy is configured to continue generating multiple synthetic user feature sets and multiple synthetic recognition results. Reinforcement learning is then performed to obtain a third recognition model, including: When the second recognition model identifies user features of the first recognition result, the error rate of identifying them as other recognition results is used as the first combined recognition error rate. Continue testing the recognition error rate of multiple combinations of multiple recognition results; Based on multiple combined recognition error rates and multiple inter-class differences, user feature synthesis strategy configuration is performed to obtain multiple secondary synthesis quantities and secondary synthesis deviations; Based on the number of secondary synthesiss and the deviation of the secondary synthesis, new multiple synthetic user feature sets and multiple synthetic recognition results are generated. The second recognition model is then subjected to reinforcement learning to obtain the third recognition model.
9. The method for constructing a neurogenetic disease identification model according to claim 8, characterized in that, Based on multiple combined recognition error rates and multiple inter-class differences, a user feature synthesis strategy is configured to obtain multiple secondary synthesis quantities and secondary synthesis deviations, including: Calculate the ratio of multiple combined recognition error rates to the mean of multiple combined recognition error rates, and perform correction calculations on multiple inter-class differences to obtain multiple corrected inter-class differences. Based on the differences between multiple correction classes, calculate the number of secondary syntheses and the deviation of secondary syntheses.
10. A system for constructing a model for identifying neurogenetic diseases, characterized in that, A method for constructing a neurogenetic disease identification model according to any one of claims 1-9 includes: The difference analysis module is used to obtain multiple identification results and multiple user feature sets for neurogenetic disease identification, perform inter-class difference analysis, and obtain multiple inter-class differences. The first training module is used to train a first recognition model for the recognition of neurogenetic diseases based on the multiple recognition results and multiple user feature sets. The reinforcement learning module is used to configure user feature synthesis strategy based on multiple inter-class differences, generate multiple synthesized user feature sets and multiple synthesized recognition results, perform reinforcement learning on the first recognition model, and obtain a second recognition model. The user feature synthesis strategy includes the number of synthesis and the degree of synthesis deviation. The iterative optimization module is used to test the combined recognition error rate of the second recognition model for multiple recognition results, obtain multiple combined recognition error rates, combine the multiple inter-class differences, configure the user feature synthesis strategy, continue to generate multiple synthetic user feature sets and multiple synthetic recognition results, perform reinforcement learning, and obtain the third recognition model.