T2DM Test Data Association System Based on Adaptive Multivariate Calibration Model
Through the T2DM test data association system based on the adaptive multivariate correction model, a multivariate correction model and expression layer are constructed to generate characterization characteristics of patients and animals, solving the reliability of the test data in diabetes data analysis in the prior art, and achieving more accurate diagnosis and analysis.
Patent Information
- Application Number
- CN202510486903.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-18
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2045-04-18
AI Technical Summary
The prior art has the reliability of the test data in the analysis of diabetes data. A single test data is difficult to reflect the high correlation causal relationship, which affects the accuracy of other patients' investigation and diagnosis.
A T2DM test data association system based on adaptive multivariate correction model is adopted, and a multivariate correction model and expression layer is constructed through the model construction module, patient sample anchoring module, animal sample labeling module, evolutionary estimation module and model correction module to generate the characterization characteristics of patients and animals, and then data matching and result analysis are carried out.
It improves the reliability of the test data, can promptly detect abnormal test data caused by uncontrollable patient behavior, assists in diagnostic analysis, and improves the relevance and diagnostic accuracy of the data.
Smart Images

Figure CN120015353B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of diabetes data analysis, and more specifically, to a T2DM test data association system based on an adaptive multivariate calibration model. Background Art
[0002] With the increasing incidence of diabetes year by year, diabetic nephropathy (DN) has become the main cause of end-stage renal disease. Currently, it is believed that the occurrence of DN is the result of the long-term combined action of genetic factors and environmental factors. Just like other tissues and organs, the kidney will be in the stage of functional self-compensation for a long time before obvious structural changes occur under the action of pathogenic factors such as hyperglycemia, that is, the high filtration rate stage described by Mogensen staging. If the increase in urinary microalbuminuria (UMA) is used as a marker of clinical DN, then the above-mentioned high filtration rate stage, that is, the functional compensation stage when UMA has not increased, can be regarded as "subclinical" DN. There are many ways to conduct diabetes testing and research, such as the relatively common urine, serum, and feces. There are also various diabetes analysis programs. For example, the patent with publication number CN115029431B discloses a type 2 diabetes gene detection kit and a type 2 diabetes genetic risk assessment system, which can determine the genetic risk of diabetes by obtaining gene sequences. For another example, the patent with publication number CN109171694B discloses a method and system for evaluating the condition of diabetes based on pulse signals. By obtaining pulse signals, characteristic indicators corresponding to the pulse signals are obtained, and a model function of the corresponding relationship between the characteristic indicators of the pulse signals and the condition of diabetes is obtained; by obtaining pulse signals, calculating and based on the characteristic indicators of the pulse signals, through the model function, an evaluation result of the condition of diabetes is obtained. The system can also obtain multiple pulse signals, obtain multiple evaluation results of the condition of diabetes, record and analyze these results, and obtain an evaluation result of the development trend of the condition of diabetes for the daily condition monitoring and control of diabetic patients. For example, the patent with publication number CN117893836B discloses a method for predicting diabetic nephropathy based on fundus vascular geometric parameters. The method includes the following steps: obtaining a data set, which includes clinical data and retinal vascular imaging data; screening modeling indicators from the clinical data and fundus vascular geometric parameters, and establishing a training set; training the training set based on the method of logistic regression to obtain a prediction model; analyzing the data to be measured through the prediction model to predict diabetic nephropathy or non-diabetic nephropathy. Extracting fundus vascular geometric parameters from retinal vascular imaging data, training a prediction model using the data set, analyzing the data to be measured, and predicting whether the patient has diabetic nephropathy or non-diabetic nephropathy; realizing non-invasive and rapid prediction of diabetic nephropathy; strongly demonstrating the association between fundus vascular characteristics and diabetic nephropathy.The above data are collectively referred to as test data. At present, there are certain limitations in predicting diabetes from test data. Since the prediction of diabetes based on a single test data feedback is not a highly correlated causal relationship, it is very important to screen other conditions of the patient. For example, there are high requirements for the patient's eating and rest. Therefore, the test data of clinical patients are difficult to be directly used as the only judgment basis. If the patient does not perform the corresponding actions as required, the test data will be affected by other conditions, thus affecting the analysis results. Summary of the Invention
[0003] In view of this, the object of the present invention is to provide a T2DM test data association system based on an adaptive multivariate calibration model.
[0004] In order to solve the above technical problems, the technical solution of the present invention is:
[0005] A T2DM test data association system based on an adaptive multivariate calibration model, characterized by: a model construction module, a patient sample anchoring module, an animal sample marking module, an evolutionary estimation module, and a model calibration module;
[0006] The model construction module is used to construct a multivariate calibration model. The multivariate calibration model includes a number of diagnostic nodes, each diagnostic node corresponding to a diagnostic result item. Identification connection lines are formed between the diagnostic nodes, and the identification connection lines reflect the association relationship between the diagnostic nodes. The multivariate calibration model includes a number of expression layers, each expression layer corresponding to a test item setting, and patient characterization features are set corresponding to the diagnostic nodes in the expression layer;
[0007] The patient sample anchoring module includes a reliability screening unit, an anchoring clustering unit, and an anchoring matching unit. The reliability screening unit is configured with a reliability evaluation algorithm, and the reliability evaluation algorithm is used to calculate the reliability value of each patient sample and screen the patient samples according to a pre-generated patient reliability threshold. The anchoring clustering unit is configured with anchoring clustering conditions. When the clustered samples of the screened patient samples in the corresponding diagnostic nodes meet the corresponding anchoring clustering conditions, the corresponding patient samples are anchored to the corresponding diagnostic nodes. The anchoring matching unit is used to generate patient characterization features based on the test data in the patient samples corresponding to the anchored diagnostic nodes;
[0008] The animal sample marking module includes a marking clustering unit and a marking matching unit. The marking clustering unit is configured with marking clustering conditions. When the clustered samples of the animal samples in the corresponding diagnostic nodes meet the corresponding marking clustering conditions, the corresponding animal samples are marked on the corresponding diagnostic nodes. The marking matching unit is used to generate animal characterization features based on the test data of the animal samples corresponding to the marked diagnostic nodes;
[0009] The evolution estimation module includes an evolution configuration unit, an evolution training unit, and an evolution estimation unit; the evolution configuration unit is configured with an evolution estimation sub-model for different inspection items, the evolution training unit is used to obtain the patient characterization features and animal characterization features belonging to the same diagnostic node to establish an evolution sample, and train the corresponding evolution estimation sub-model according to the evolution sample, and the evolution estimation unit generates the corresponding evolution characterization features according to the animal characterization features of other diagnostic nodes by the evolution estimation sub-model as the patient characterization features of this diagnostic node;
[0010] The model correction module is configured with a reliability evaluation unit, a vector configuration unit, and a distance configuration unit. The reliability evaluation unit is used to generate dynamic reliability parameters for identifying connections, the vector configuration unit is used to generate vectors for identifying connections, and the distance configuration unit is used to generate the lengths of the identifying connections.
[0011] Furthermore: there is also a result analysis module, which includes a data acquisition unit, a data matching unit, a result positioning unit, and a result output unit. The data acquisition unit is used to acquire measured inspection information, and the measured inspection information includes inspection data corresponding to different inspection items. The data matching unit matches the corresponding inspection data according to the patient characterization features. The result positioning unit determines the estimation coordinates of the measured inspection information in the multivariate correction model according to the matching result of the inspection data, and the result output unit generates an analysis graph according to the estimation coordinates and outputs it.
[0012] Furthermore: the result positioning unit determines the closest diagnostic node coordinates in each expression layer as relative sub-coordinates by matching the patient characterization data with the inspection data, and calculates the estimation coordinates of the corresponding measured inspection information according to a preset coordinate mean formula, where , ,where, is the abscissa value of the estimation coordinate, is the ordinate value of the estimation coordinate, is the abscissa value of the relative sub-coordinate in the th expression layer, is the ordinate value of the relative sub-coordinate in the th expression layer, is the reliability weight of the inspection data in the th expression layer, is the reliability weight of the matched diagnostic node in the th expression layer, is the similarity weight of the inspection data and the corresponding patient characterization features in the th expression layer, is the total number of expression layers matched in the inspection information, and satisfies the constraint condition , , , , , , where is a preset reliable weight parameter, is a preset similarity weight parameter, is the reliable value of the test data in the th expression layer, is the sum of the reliable values of the test data in all expression layers, is the reliable value of the matched diagnostic node in the th expression layer, is the sum of the reliable values of the matched diagnostic nodes in all expression layers, is the similarity value between the test data and the corresponding patient characterization features in the th expression layer, is the sum of the similarity values between all test data and the corresponding patient characterization features.
[0013] Furthermore: It also includes a data processing module. The data processing module includes a quantization feature processing unit, an image feature processing unit, and a sequence feature processing unit. The expression layer includes a serum test expression layer, a fecal test expression layer, a urine test expression layer, an eye pattern recognition expression layer, a pulse test expression layer, and a gene test expression layer;
[0014] The quantization feature processing unit is used to process test data with a numerical data type. The quantization feature processing unit is pre-configured with test mapping functions for different test sub-items, and re-assigns the test data according to the test mapping functions. When the belonging clustering cluster meets the preset quantization feature extraction conditions, patient characterization features or animal characterization features are generated according to the mapping rules between the re-assigned test data of different test sub-items;
[0015] The image feature processing unit is used to process test data with an image data type. The image feature processing unit grayscales the image and generates a binarization threshold according to the corresponding gray mean value of the image to binarize the image. When the belonging clustering cluster meets the preset image feature extraction conditions, the element shape graphics in the test sub-items are extracted as patient characterization features or animal characterization features;
[0016] The sequence feature processing unit pre-stores several different gene recognition fragments. The sequence feature processing unit recognizes and marks the gene recognition fragments in the test data. When the clustering cluster meets the preset sequence feature extraction conditions, the patient characterization features or animal characterization features are generated according to the set of gene recognition fragments.
[0017] Furthermore: The reliability evaluation algorithm includes:
[0018] , where is for corresponding to the reliability of the test data, is the macroscopic reliability corresponding to the test data, and the macroscopic reliability is negatively correlated with the dispersion degree corresponding to the test sub-item, is the th reliable weight value corresponding to the influencing factor item related to the test data, and the reliable weight value reflects the stability of the influencing factor item itself, is the th controllable selection value corresponding to the influencing factor item related to the test data, and the controllable selection value reflects the controllability of the selection content corresponding to the influencing factor item, is the total number of influencing factor items related to the test data, is the th difference value corresponding to the microscopic difference item related to the test data, and the difference value reflects the dispersion degree of the category corresponding to the microscopic difference item, is the total number of microscopic difference items related to the test data, is the preset macroscopic reliability weight, is the preset influencing factor weight, is the preset microscopic difference weight, and there is .
[0019] Furthermore: The anchor clustering conditions include a similarity sub-condition, a reliability sub-condition, and a quantity sub-condition. The marker clustering conditions include a similarity sub-condition, a reliability sub-condition, and a quantity sub-condition. When the similarity sub-condition, the reliability sub-condition, and the quantity sub-condition are all satisfied, it is considered that the corresponding anchor clustering condition or marker clustering condition is satisfied.
[0020] The similarity sub-condition is that the average similarity between the test data of the clustering cluster is greater than the preset similarity benchmark;
[0021] The reliability sub-condition is that the average reliable value between the test data of the clustering cluster is greater than the preset reliable value benchmark;
[0022] The quantity sub-condition is that the number of test data of the clustering cluster is greater than the preset quantity benchmark.
[0023] Furthermore: The vector configuration unit includes obtaining the disease course migration value and the test migration value between two diagnostic nodes, and the vector expression of the identification connection line is , where represents the vector of the identification connection line between the diagnostic node numbered and the diagnostic node numbered , is the number of The disease course migration value between the diagnostic node numbered and the diagnostic node numbered is the test migration value between the diagnostic node numbered and the diagnostic node numbered . The disease course migration value reflects the disease course relationship between two diagnostic nodes, and the test migration value reflects the difference in characterization features between two diagnostic nodes. Each characterization feature is pre-associated with a characterization feature value, and the difference in characterization features is the difference in characterization feature values between diagnostic nodes.
[0024] Furthermore: The reliable evaluation unit is configured with a transfer function table, and the transfer function table stores a number of reliability adjustment parameters. Each reliability adjustment parameter is indexed by an evolution factor. When the evolution inference module generates an evolution characterization feature, it generates an evolution factor based on the original fitness and the target fitness of the evolution characterization feature. The original fitness reflects the degree of fitness between the generated evolution characterization feature and the animal characterization feature, and the target fitness reflects the degree of fitness between the evolution characterization feature and the patient sample corresponding to this diagnostic node. The dynamic reliability parameter is generated according to the reliability adjustment parameter, and the reliability value of the patient characterization feature corresponding to the diagnostic node pointed to by the identification connection is calculated according to the dynamic reliability parameter.
[0025] Furthermore: The model correction module further includes a similarity learning unit. The similarity learning unit is used to obtain new patient samples, match the corresponding diagnostic nodes according to the diagnostic results of the patient samples, extract the test data in the patient samples, compare the corresponding patient characterization features and test data to generate deviation correction information. The deviation correction information includes an association correction index, a characterization deviation index, and a characterization reliability index. The deviation correction information is substituted into the evolution inference module to correct the evolution inference sub-model.
[0026] Furthermore: The distance configuration unit performs an association analysis on all patient samples corresponding to two diagnostic nodes between the identification connections of each expression layer, and configures an association mapping function. The length of the corresponding identification connection is calculated through the association mapping function.
[0027] The technical effects of the present invention are mainly reflected in the following aspects: By setting like this, based on the experimental results of relatively stable animal samples as the basis for deduction, deduce the differences between animal test data and patient test data in the test data, and then obtain the corresponding prediction model through evolution by the evolution model. In this way, abnormal situations can be judged according to patient test data or auxiliary diagnostic analysis can be carried out, improving the reliability of the data and being able to timely detect abnormalities in test data caused by uncontrollable patient behaviors. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] Figure 1: System architecture schematic diagram of the T2DM test data correlation system based on the adaptive multivariate calibration model of the present invention;
[0029] Figure 2 : Schematic diagram of the patient sample anchoring module of the T2DM test data correlation system based on the adaptive multivariate calibration model of the present invention;
[0030] Figure 3 : Schematic diagram of the animal sample marking module of the T2DM test data correlation system based on the adaptive multivariate calibration model of the present invention;
[0031] Figure 4 : Schematic diagram of the evolution estimation module of the T2DM test data correlation system based on the adaptive multivariate calibration model of the present invention;
[0032] Figure 5 : Schematic diagram of the model calibration module of the T2DM test data correlation system based on the adaptive multivariate calibration model of the present invention;
[0033] Figure 6 : Schematic diagram of the result analysis of the T2DM test data correlation system based on the adaptive multivariate calibration model of the present invention;
[0034] Figure 7 : Schematic diagram of the data processing module of the T2DM test data correlation system based on the adaptive multivariate calibration model of the present invention.
[0035] Reference numerals: 100, model construction module; 200, patient sample anchoring module; 210, reliability screening unit; 220, anchoring clustering unit; 230, anchoring matching unit; 300, animal sample marking module; 310, marking clustering unit; 320, marking matching unit; 400, evolution estimation module; 410, evolution configuration unit; 420, evolution training unit; 430, evolution estimation unit; 500, model calibration module; 510, reliable evaluation unit; 520, vector configuration unit; 530, distance configuration unit; 540, similarity learning unit; 600, result analysis module; 610, data acquisition unit; 620, data matching unit; 630, result positioning unit; 640, result output unit; 700, data processing module; 710, quantitative feature processing unit; 720, image feature processing unit; 730, sequence feature processing unit. Detailed implementation manners
[0036] The following further details the specific implementation manners of the present invention in conjunction with the accompanying drawings, so that the technical solutions of the present invention are easier to understand and master.
[0037] A T2DM test data correlation system based on an adaptive multivariate correction model, comprising / model construction module 100, patient sample anchoring module 200, animal sample marking module 300, evolutionary estimation module 400, and model correction module 500; The purpose of the present invention is that in terms of detection, if the test data takes patients as an example, information such as drinking water, eating, exercise, or living status is difficult to control. Therefore, if the test data deviates due to these correlation variables, and it is difficult for testers to discover, it will lead to the same recognition of all test data, which will further lead to deviations in the test results and affect the reliability of the test data.
[0038] The model construction module 100 is used to construct a multivariate correction model. The multivariate correction model includes several diagnostic nodes, each diagnostic node corresponding to a diagnostic result item, and there are identification connection lines formed between the diagnostic nodes. The identification connection lines reflect the association relationship between the diagnostic nodes. The multivariate correction model includes several expression layers, each expression layer corresponding to a test item setting, and patient characterization features are set corresponding to the diagnostic nodes in the expression layer; First of all, this model can be understood as a network of diagnostic nodes composed of constellation diagrams and diagnostic nodes associated by identification connection lines. The diagnostic nodes reflect the disease conditions or diagnostic results, classify all diagnostic results regarding T2DM, and assign labels to form diagnostic nodes. Then, each sample can find the corresponding diagnostic node, and the relative positions of the diagnostic nodes in the expression layer can change. For example, the same diagnostic node may be in different positions in the first expression layer and the second expression layer. Each expression layer corresponds to a test item setting. For example, urine tests can be configured with different expression layers according to the test items, and serum tests and urine tests belong to different expression layers. Therefore, the corresponding test results corresponding to the characterization features of each diagnostic node are also different. For example, the diagnostic nodes in the test data can be: NC: Normal control group; SDN: Subclinical diabetic nephropathy group; EDN: Early diabetic nephropathy group; which can be used as diagnostic nodes, and the corresponding characterization features can be the characteristics of various aspects of the test data.
[0039] It further includes a data processing module 700. The data processing module 700 includes a quantitative feature processing unit 710, an image feature processing unit 720, and a sequence feature processing unit 730. The expression layer includes a serum test expression layer, a fecal test expression layer, a urine test expression layer, an eye pattern recognition expression layer, a pulse test expression layer, and a gene test expression layer; The purpose of the data processing module 700 is to pre-process the above data.
[0040] The quantization feature processing unit 710 is used to process the test data with a numerical data type. The quantization feature processing unit 710 is preconfigured with test mapping functions corresponding to different test sub-items, and reassigns the test data according to the test mapping functions. When the belonging clustering cluster meets the preset quantization feature extraction conditions, the patient characterization features or animal characterization features are generated according to the mapping rules between the test data reassigned according to different test sub-items. For example, taking urine test as an example: Scr: Serum creatinine; UMA: Urinary microalbuminuria; microglobubin; NAG: N-acetyl-D-glucosaminadase; GAL: galactosidase; RBP: Retinol binding protein; UAGT: Urinary angiotensinogen; The test results of different sub-items are different, and each item has a corresponding range. Therefore, if different situations need to be classified, the principle is the range value. However, since the range units exceeded by different test sub-items are different, corresponding test mapping functions are constructed, so that the corresponding numerical deviation can be quantified according to the abnormal situation, and the reaction can be more accurate. And if the corresponding features need to be extracted, for the digital features, after mapping, the mapping value ranges of different items can be used as the characterization features. At the same time, if the deviation between an actual data and the characterization features needs to be calculated, it can also be determined by the deviation from the central value of the range.
[0041] The image feature processing unit 720 is used to process the test data with an image data type. After graying the image, the image feature processing unit generates a binarization threshold according to the gray mean value corresponding to the image to binarize the image. When the belonging clustering cluster meets the preset image feature extraction conditions, the element shape graphics in the test sub-items are extracted as the patient characterization features or animal characterization features; By simplifying the image, the element shape can be extracted as the recognition basis, and by generating the binarization threshold through the gray mean value, the influence of the environment on the brightness value can be eliminated, so that the recognition result is more accurate and the data volume is smaller. After obtaining the element graphics, a reference feature graphic can be formed by combining the element graphics as the characterization feature. Subsequently, if the similarity between the test image and the reference feature graphic exceeds the threshold, it is considered that the feature matching is successful.
[0042] The sequence feature processing unit 730 prestores a number of different gene recognition fragments. The sequence feature processing unit 730 recognizes and marks the gene recognition fragments in the test data. When the clustering cluster meets the preset sequence feature extraction conditions, the patient characterization features or animal characterization features are generated according to the set of gene recognition fragments. By marking the gene recognition fragments to form corresponding characterization features, if there is a set of identical gene fragments, it is considered a successful match. In this way, the data extraction of test data in different formats can be completed. Then, as long as there is a sufficient amount of test data that meets the clustering conditions, the corresponding characterization features can be determined, and the matching can be completed based on the characterization features.
[0043] The patient sample anchoring module 200 includes a reliability screening unit 210, an anchoring clustering unit 220, and an anchoring matching unit 230. First, due to the reasons described above, if the diagnostic nodes are directly determined through the clustering analysis of patient samples, the features obtained in this way deviate greatly from the actual situation, and the test data cannot be associated with each other. Therefore, the first step is to find patient samples that can serve as the anchoring diagnostic nodes, that is, the patient samples with high reliability and less influence and interference in the collection are used as the anchor. The coordinates of the diagnostic nodes corresponding to the anchor samples in each expression layer of the model are the same. In this way, the non-anchored diagnostic nodes are located through the anchored diagnostic nodes, so as to associate the corresponding test data, and the reliability is higher. Specifically, first, the reliability analysis of all patient samples needs to be carried out to screen out the patient samples with high reliability. The specific method is as follows: The reliability screening unit 210 is configured with a reliability evaluation algorithm, and the reliability evaluation algorithm is used to calculate the reliability value of each patient sample and screen the patient samples according to the pre-generated patient reliability threshold.
[0044] The reliability evaluation algorithm includes:
[0045] , where, is the reliability of the corresponding test data, is the macroscopic reliability corresponding to this test data. The macroscopic reliability is negatively correlated with the dispersion degree corresponding to the test sub-items. That is, if the dispersion degree of the corresponding test results in the test sub-items is higher, the value of this macroscopic reliability is lower. is the th reliable weight value corresponding to the influencing factor item related to this test data. The reliable weight value reflects the stability of this influencing factor item itself. For example, age, gender, exercise situation, and eating situation all belong to the influencing factor items. Therefore, the influencing factor items have corresponding reliable weight values. is the Controllable selection values corresponding to the influencing factor items related to the test data. For each specific content of the influencing factor item actually collected, there are corresponding controllable selection values. Different age ranges have different controllable selection values. The controllable selection values reflect the controllability of the selection content corresponding to the influencing factor item. Is the total number of influencing factor items related to the test data. Is the Difference value corresponding to the micro-difference item related to the test data. The difference value reflects the dispersion degree of the category corresponding to the micro-difference item. In a test data, the micro-difference item itself has a certain degree of discreteness. For example, if the change stability of blood glucose data is low, the corresponding dispersion degree is high, so the corresponding difference value is large and the reliability is also high. Is the total number of micro-difference items related to the test data. Is the preset macro reliability weight. Is the preset influencing factor weight. Is the preset micro-difference weight, and there is .
[0046] The anchoring clustering unit 220 is configured with anchoring clustering conditions. When the clustering cluster of the selected patient sample at the corresponding diagnosis node meets the corresponding anchoring clustering conditions, the corresponding patient sample is anchored to the corresponding diagnosis node. The anchoring matching unit 230 is used to generate patient characterization features according to the test data in the patient samples corresponding to the anchored diagnosis nodes.
[0047] The anchoring clustering conditions include a similarity sub-condition, a reliability sub-condition, and a quantity sub-condition. The marking clustering conditions include a similarity sub-condition, a reliability sub-condition, and a quantity sub-condition. When the similarity sub-condition, the reliability sub-condition, and the quantity sub-condition are all met, it is considered that the corresponding anchoring clustering condition or marking clustering condition is met.
[0048] The similarity sub-condition is that the similarity mean value between the test data of the clustering cluster is greater than the preset similarity benchmark. The similarity mean value is also the difference value between the test data. For example, if the difference value between the test data of different patients is small, the similarity mean value is large.
[0049] The reliability sub-condition is that the reliability mean value between the test data of the clustering cluster is greater than the preset reliability benchmark. The reliability mean value reflects the reliability of the test data itself. The reliability of each test data is different, and the overall reliability can be judged by calculating the mean value.
[0050] The above-mentioned quantitative sub-condition is that the number of test data of the clustering cluster is greater than a preset quantitative benchmark. The quantitative benchmark requires that the number of test data is large enough to generate more accurate characteristic features. If the above three sub-conditions cannot be met, it means that this diagnostic node is not suitable as the diagnostic node to be anchored. After finding a suitable diagnostic node to be anchored, the characteristics of the patient sample can be evolved through the animal sample. There are differences between the animal sample and the patient sample, and the specific expression content of each value may be different, but generally speaking, they are regular and relevant. Therefore, when the reliability of the patient sample is low, through the collection of animal test data and based on the anchored diagnostic node, deduction can be achieved.
[0051] The animal sample marking module 300 includes a marking clustering unit 310 and a marking matching unit 320. The marking clustering unit 310 is configured with marking clustering conditions. When the clustering cluster of the animal sample at the corresponding diagnostic node meets the corresponding marking clustering conditions, the corresponding animal sample is marked at the corresponding diagnostic node. The marking matching unit 320 is used to generate animal characteristic features according to the test data of the animal sample corresponding to the marked diagnostic node; the method of clustering and extracting features has been described above and will not be elaborated here. The marking clustering unit 310 and the marking matching unit 320 are used to configure corresponding animal characteristic features for all diagnostic nodes.
[0052] The evolution deduction module 400 includes an evolution configuration unit 410, an evolution training unit 420, and an evolution deduction unit 430; the evolution configuration unit 410 is configured with evolution deduction sub-models corresponding to different test items.
[0053] The evolution training unit 420 is used to obtain the patient characteristic features and animal characteristic features belonging to the same diagnostic node to establish an evolution sample, and train the corresponding evolution deduction sub-model according to the evolution sample. The purpose of the evolution deduction sub-model is to discover the regular relationship between the patient characteristic features and animal characteristic features based on the same diagnostic node, so as to infer the corresponding result value through the animal characteristic features for the diagnostic node with unknown patient characteristic features. This evolution deduction sub-model first constructs a benchmark expression with variable parameters based on experience, then brings in the animal characteristic features of the anchored diagnostic node, corrects the variable parameters through deviation, and completes the construction of each deduction sub-model through machine learning.
[0054] The evolution deduction unit 430 generates corresponding evolution characteristic features according to the evolution deduction sub-model based on the animal characteristic features of other diagnostic nodes as the patient characteristic features of this diagnostic node; the evolution characteristic features are the speculated patient characteristic features. When the reliability of the patient sample is not high, deducing the patient characteristic features corresponding to the diagnostic node can correlate the subsequent test data.
[0055] The model calibration module 500 is configured with a reliability evaluation unit 510, a vector configuration unit 520, and a distance configuration unit 530.
[0056] The reliability evaluation unit 510 is used to generate dynamic reliability parameters for identifying connections. The reliability evaluation unit is configured with a transfer function table, which is pre-configured. Different evolution factors correspond to unreliability adjustment parameters. The transfer function table stores a number of reliability adjustment parameters. Each reliability adjustment parameter is indexed by an evolution factor. When the evolution inference module 400 generates evolution representation features, evolution factors are generated based on the original fitness and target fitness of the evolution representation features. The original fitness reflects the degree of fitness between the generated evolution representation features and the animal representation features, that is, the fitness relationship between the evolution representation features and the animal representation features. The target fitness reflects the degree of fitness between the evolution representation features and the patient samples corresponding to the diagnostic nodes, that is, the matching relationship between the evolution representation features and the patient samples with low reliability. The dynamic reliability parameters are generated based on the reliability adjustment parameters. The reliability value of the patient representation features corresponding to the diagnostic nodes pointed to by the identifying connections is calculated according to the dynamic reliability parameters. In this way, the reliability value of each diagnostic node corresponding to the patient representation features can be calculated through the identifying connections. Specifically, the reliability value of the starting diagnostic node is multiplied by the corresponding dynamic reliability parameter to transfer the reliability.
[0057] The vector configuration unit is used to generate vectors for identifying connections. The vector configuration unit includes obtaining the course migration value and the test migration value between two diagnostic nodes. The vector of the identifying connection is expressed as , where represents the vector of the identifying connection between the diagnostic node numbered and the diagnostic node numbered , is the course migration value between the diagnostic node numbered and the diagnostic node numbered , is the test migration value between the diagnostic node numbered and the diagnostic node numbered . The course migration value reflects the course relationship between two diagnostic nodes, that is, whether there is a correlation between the occurrences of two diagnostic nodes. For example, they occur at intervals or simultaneously. The test migration value reflects the difference in representation features between two diagnostic nodes. Each representation feature is pre-associated with a representation feature value. The representation feature value is a weighted value of the feature complexity and feature accuracy of the data. Theoretically, the more complex and accurate it is, the higher the correlation with other nodes may be. So in the vertical vectorization, the difference in representation features is the difference in representation feature values between diagnostic nodes.
[0058] The distance configuration unit 530 is used to generate the length of the identification line. The distance configuration unit 530 performs a correlation analysis on all patient samples corresponding to two diagnosis nodes between the identification lines of each expression layer, and is configured with a correlation mapping function, and the length of the corresponding identification line is calculated by the correlation mapping function. The higher the correlation, the shorter the length of the identification line calculated by the corresponding correlation mapping function.
[0059] The model correction module 500 also includes a similarity learning unit 540, which is used to obtain new patient samples, match the corresponding diagnostic nodes according to the diagnostic results of the patient samples, and extract the test data in the patient samples, and compare the corresponding patient characterization features and the test data to generate deviation correction information, the deviation correction information includes an association correction index, a characterization deviation index, and a characterization reliability index, and the deviation correction information is substituted into the evolution inference module 400 to correct the evolution inference sub-model. The deviation correction information is the deviation correction evolution inference sub-model of the actual results through the three values of the association correction index, the characterization deviation index, and the characterization reliability index, so that the characterization features obtained by the evolution inference sub-model are close to the patient samples, and abnormal test data can be excluded in time.
[0060] The result analysis module 600 includes a data acquisition unit 610, a data matching unit 620, a result positioning unit 630 and a result output unit 640.
[0061] The data acquisition unit 610 is used to acquire actual test information, which includes test data corresponding to different test items. If there is new test information, the corresponding test data is matched by acquiring the test information.
[0062] The data matching unit 620 matches the corresponding test data according to the patient's characterization features, so that the corresponding diagnosis node can be found.
[0063] The result positioning unit 630 determines the estimated coordinates of the measured test information in the multivariate correction model according to the matching results of the test data, and the result output unit 640 generates and outputs an analysis map according to the estimated coordinates. The result positioning unit 630 determines the closest diagnostic node coordinates as relative sub-coordinates at each expression layer by matching the test data with the patient representation data, and calculates the estimated coordinates of the corresponding measured test information according to the preset coordinate mean formula.
[0064] ,
[0065] ,in, is the abscissa value of the estimated coordinate, is the ordinate value of the estimated coordinate, For the The horizontal coordinate value of the relative sub-coordinate in the expression layer, For the The ordinate value of the relative sub-coordinate in the expression layer, For the The reliability weight of the test data in the expression layer, For the The reliability weight of the matched diagnosis node in the expression layer, For the The similarity weights of the test data and the corresponding patient representation features in the expression layer, is the total number of expression layers that match the test information and satisfy the constraints , , , , , ,in, is the preset reliable weight parameter, is the preset similarity weight parameter, For the The reliability value of the test data in the expression layer, is the sum of the reliability values of the test data of all expression layers, For the The reliability value of the diagnosis node matched by the expression layer, is the sum of the reliability values of the matched diagnosis nodes in all expression layers, For the The similarity value between the test data and the corresponding patient representation features in the expression layer, It is the sum of the similarity values of all test data and the corresponding patient characterization features. Because different test data will find different nodes, this solution can determine the comprehensive coordinates, determine the specific location, and then generate the judgment result based on the coordinates as the analysis result.
[0066] Of course, the above are only typical examples of the present invention. In addition, the present invention may also have many other specific implementations. All technical solutions formed by equivalent replacement or equivalent transformation fall within the scope of protection required by the present invention.
Claims
1. A T2DM test data association system based on an adaptive multivariate correction model, characterized in that: Model building module, patient sample anchoring module, animal sample labeling module, evolution inference module and model correction module; The model building module is used to build a multivariate correction model, the multivariate correction model includes a plurality of diagnosis nodes, each diagnosis node corresponds to a diagnosis result item, and an identification line is formed between the diagnosis nodes, and the identification line reflects the association relationship between the diagnosis nodes. The multivariate correction model includes a plurality of expression layers, each expression layer corresponds to a test item setting, and the expression layer is provided with a patient characterization feature corresponding to the diagnosis node; The patient sample anchoring module includes a reliability screening unit, an anchor clustering unit and an anchor matching unit. The reliability screening unit is configured with a reliability evaluation algorithm, which is used to calculate the reliability value of each patient sample and screen the patient samples according to the pre-generated patient reliability threshold. The anchor clustering unit is configured with an anchor clustering condition. When the cluster cluster of the screened patient sample at the corresponding diagnosis node meets the corresponding anchor clustering condition, the corresponding patient sample is anchored to the corresponding diagnosis node. The anchor matching unit is used to generate patient characterization features according to the test data in the patient sample corresponding to the anchored diagnosis node; The animal sample marking module includes a marking clustering unit and a marking matching unit. The marking clustering unit is configured with a marking clustering condition. When the cluster of the animal sample at the corresponding diagnosis node meets the corresponding marking clustering condition, the corresponding animal sample is marked at the corresponding diagnosis node. The marking matching unit is used to generate an animal characterization feature according to the inspection data of the animal sample corresponding to the marked diagnosis node; The evolutionary inference module includes an evolutionary configuration unit, an evolutionary training unit, and an evolutionary inference unit; the evolutionary configuration unit is configured with an evolutionary inference sub-model corresponding to different test items, the evolutionary training unit is used to obtain patient characterization features and animal characterization features belonging to the same diagnostic node to establish an evolutionary sample, and train the corresponding evolutionary inference sub-model according to the evolutionary sample, and the evolutionary inference unit generates corresponding evolutionary characterization features according to the animal characterization features of other diagnostic nodes according to the evolutionary inference sub-model as the patient characterization features of the diagnostic node; The model correction module is configured with a reliable evaluation unit, a vector configuration unit and a distance configuration unit. The reliable evaluation unit is used to generate dynamic reliability parameters of the identification line, the vector configuration unit is used to generate the vector of the identification line, and the distance configuration unit is used to generate the length of the identification line.
2. The T2DM test data association system based on the adaptive multivariate correction model according to claim 1, characterized in that: It also includes a result analysis module, which includes a data acquisition unit, a data matching unit, a result positioning unit and a result output unit. The data acquisition unit is used to acquire measured test information, and the measured test information includes test data corresponding to different test items. The data matching unit matches the corresponding test data according to the patient's characterization characteristics. The result positioning unit determines the estimated coordinates of the measured test information in the multivariate correction model according to the matching results of the test data. The result output unit generates an analysis map according to the estimated coordinates and outputs it.
3. The T2DM test data association system based on the adaptive multivariate correction model according to claim 2, characterized in that: The result positioning unit determines the closest diagnostic node coordinates as relative sub-coordinates at each expression layer by matching the patient representation data with the test data, and calculates the estimated coordinates corresponding to the measured test information according to the preset coordinate mean formula. , ,in, is the abscissa value of the estimated coordinate, is the ordinate value of the estimated coordinate, For the The horizontal coordinate value of the relative sub-coordinate in the expression layer, For the The ordinate value of the relative sub-coordinate in the expression layer, For the The reliability weight of the test data in the expression layer, For the The reliability weight of the matched diagnosis node in the expression layer, For the The similarity weights of the test data and the corresponding patient representation features in the expression layer, is the total number of expression layers that match the test information and satisfy the constraints , , , , , ,in, is the preset reliable weight parameter, is the preset similarity weight parameter, For the The reliability value of the test data in the expression layer, is the sum of the reliability values of the test data of all expression layers, For the The reliability value of the diagnosis node matched by the expression layer, is the sum of the reliability values of the matched diagnosis nodes in all expression layers, For the The similarity value between the test data and the corresponding patient representation features in the expression layer, It is the sum of the similarity values of all test data and the corresponding patient characterization features.
4. The T2DM test data association system based on the adaptive multivariate correction model according to claim 1, characterized in that: It also includes a data processing module, which includes a quantitative feature processing unit, an image feature processing unit and a sequence feature processing unit, and the expression layer includes a serum test expression layer, a stool test expression layer, a urine test expression layer, an eye pattern recognition expression layer, a pulse test expression layer and a gene test expression layer; The quantitative feature processing unit is used to process the test data whose data type is a numerical value. The quantitative feature processing unit is pre-configured with a test mapping function corresponding to different test sub-items, and re-assigns the test data according to the test mapping function. When the cluster to which it belongs meets the preset quantitative feature extraction condition, the patient characterization feature or animal characterization feature is generated according to the mapping rule between the test data after the different test sub-items are re-assigned; The image feature processing unit is used to process the inspection data whose data type is an image. After graying the image, the image feature processing unit generates a binarization threshold according to the grayscale mean value corresponding to the image to binarize the image. When the cluster to which it belongs meets the preset image feature extraction condition, the element shape graphic in the inspection sub-item is extracted as the patient characterization feature or the animal characterization feature; The sequence feature processing unit pre-stores a number of different gene identification fragments. The sequence feature processing unit identifies and marks the gene identification fragments in the test data. When the cluster meets the preset sequence feature extraction conditions, the patient characterization feature or animal characterization feature is generated according to the set of gene identification fragments.
5. The T2DM test data association system based on the adaptive multivariate correction model according to claim 1, characterized in that: The reliability evaluation algorithm includes ,in, To verify the reliability of the data, is the macro reliability corresponding to the test data, and the macro reliability is negatively correlated with the discrete degree corresponding to the test sub-item. For the a reliable weight corresponding to an influencing factor item related to the test data, wherein the reliable weight reflects the stability of the influencing factor item itself, For the a controllable selection value corresponding to an influencing factor item related to the inspection data, wherein the controllable selection value reflects the controllability of the selection content corresponding to the influencing factor item, is the total number of influencing factors related to the test data, For the a difference value corresponding to a micro-difference item related to the test data, wherein the difference value reflects the dispersion degree of the category corresponding to the micro-difference item, is the total number of micro-difference items related to the test data, is the preset macro reliability weight, is the preset influencing factor weight, is the preset micro-difference weight, .
6. The T2DM test data association system based on the adaptive multivariate correction model according to claim 1, characterized in that: The anchor clustering condition includes a similarity sub-condition, a reliability sub-condition and a quantity sub-condition, and the tag clustering condition includes a similarity sub-condition, a reliability sub-condition and a quantity sub-condition. When the similarity sub-condition, the reliability sub-condition and the quantity sub-condition are all satisfied, the corresponding anchor clustering condition or tag clustering condition is deemed to be satisfied; The similarity sub-condition is that the similarity mean between the test data of the clusters is greater than a preset similarity benchmark; The reliability sub-condition is that the mean of the reliability values between the test data of the clusters is greater than a preset reliability value benchmark; The quantity sub-condition is that the quantity of the inspection data of the cluster is greater than a preset quantity benchmark.
7. The T2DM test data association system based on the adaptive multivariate correction model according to claim 1, characterized in that: The vector configuration unit includes obtaining the disease course migration value and the test migration value between two diagnosis nodes. The vector expression of the identification connection line is: ,in, Indicates the number The diagnostic nodes and numbers are A vector of identification lines between diagnostic nodes, For the number The diagnostic nodes and numbers are The disease course migration value between the diagnosis nodes, For the number The diagnostic nodes and numbers are The test migration value between the diagnostic nodes, the disease course migration value reflects the disease course relationship between the two diagnostic nodes, the test migration value reflects the characterization feature difference between the two diagnostic nodes, each characterization feature is pre-associated with a characterization feature value, and the characterization feature difference is the difference between the characterization feature values between the diagnostic nodes.
8. The T2DM test data association system based on the adaptive multivariate correction model according to claim 1, characterized in that: The reliability evaluation unit is configured with a transfer function table, which stores a number of reliability adjustment parameters. Each reliability adjustment parameter is indexed by an evolution element. When the evolution inference module generates an evolution characterization feature, the evolution element is generated according to the original fit and target fit of the evolution characterization feature. The original fit reflects the fit between the generated evolution characterization feature and the animal characterization feature, and the target fit reflects the fit between the evolution characterization feature and the patient sample corresponding to the diagnostic node. The dynamic reliability parameter is generated according to the reliability adjustment parameter, and the reliability value of the patient characterization feature corresponding to the diagnostic node pointed to by the identification line is calculated according to the dynamic reliability parameter.
9. The T2DM test data association system based on the adaptive multivariate correction model according to claim 1, characterized in that: The model correction module also includes a similarity learning unit, which is used to obtain new patient samples, match corresponding diagnostic nodes according to the diagnostic results of the patient samples, and extract test data in the patient samples, and compare the corresponding patient characterization features and test data to generate deviation correction information. The deviation correction information includes an association correction index, a characterization deviation index, and a characterization reliability index. The deviation correction information is substituted into the evolutionary inference module to correct the evolutionary inference sub-model.
10. The T2DM test data association system based on the adaptive multivariate correction model according to claim 1, characterized in that: The distance configuration unit performs a correlation analysis on all patient samples corresponding to two diagnosis nodes between the identification lines of each expression layer, and is configured with a correlation mapping function, through which the length of the corresponding identification line is calculated.
Citation Information
Patent Citations
A method and system for assessing diabetes based on pulse signals
CN109171694B
A type 2 diabetes gene detection kit and a type 2 diabetes genetic risk assessment system
CN115029431B
Method and system for predicting diabetic nephropathy based on fundus vascular geometric parameters
CN117893836B
Adjuvant disease diagnosis method based on patient test results
CN107066791A
Intervertebral disc degenerative change analysis system based on multi-dimensional information
CN119601217A