Method, device, computer device and medium for detecting risk of diabetes

By identifying anomalous indicators and correlation factors in real-time detection data during diabetes risk assessment, and utilizing knowledge graph matching and similarity calculation, the accuracy and convenience issues of traditional screening methods are resolved, achieving efficient and accurate diabetes risk assessment.

CN121075656BActive Publication Date: 2026-02-17CENT SOUTH UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511575556.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-10-31
Publication Date
2026-02-17
Estimated Expiration
2045-10-31

AI Technical Summary

Technical Problem

Traditional diabetes risk screening methods suffer from insufficient accuracy and poor convenience, especially in prediabetes screening where the detection sensitivity is low. Furthermore, the traditional screening process is cumbersome, resulting in low user compliance and failing to meet the needs for immediate home screening.

Method used

By identifying abnormal data in the real-time detection data of the test subjects, core indicators are screened based on the correlation coefficient between abnormal indicators and diabetes. Candidate factors are then matched from the knowledge graph, similarity is calculated, target factors and their associated factors are determined, and a comprehensive risk value is generated.

Benefits of technology

It significantly improves the scientific rigor and accuracy of diabetes risk assessment, reduces interference from redundant indicators, enhances the efficiency and accuracy of screening, and meets the needs of home-based, on-the-spot screening.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121075656B_ABST
    Figure CN121075656B_ABST
Patent Text Reader

Abstract

The application relates to a diabetes risk detection method and device, computer equipment and a medium. The method comprises the following steps: S1, determining abnormal data in real-time detection data of a to-be-detected object; the abnormal data comprises an abnormal index and a corresponding index value; S2, determining a core index in each abnormal index based on a correlation coefficient between each abnormal index and diabetes; S3, screening a candidate factor matching a normalized core index from a knowledge graph; the knowledge graph comprises a plurality of factors, and each factor comprises an index and an index value; S4, calculating a similarity based on the index value of the candidate factor and the index value of the core index, and determining a target factor with a similarity greater than a similarity threshold value; and S5, determining an associated factor associated with the target factor in the knowledge graph, and determining a comprehensive risk value of the to-be-detected object based on the risk values of the target factor and the associated factor. The method can improve accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of diabetes risk detection technology, and in particular to a method, device, computer equipment, and medium for diabetes risk detection. Background Technology

[0002] Diabetes is a prevalent chronic metabolic disease worldwide, and its incidence is rising annually due to population aging, changes in dietary patterns, and the widespread prevalence of sedentary lifestyles. Because early symptoms of diabetes are often subtle, it can easily lead to serious complications such as cardiovascular disease, kidney disease, and neuropathy in later stages. Therefore, diabetes risk screening is a crucial aspect of disease prevention and control. However, traditional diabetes risk screening models have many limitations.

[0003] Regarding the accuracy of screening, traditional methods suffer from both "over-treatment" and "missed diagnoses." Early screening mainly relies on biochemical indicators such as fasting blood glucose and oral glucose tolerance tests, but these indicators are significantly affected by short-term factors such as diet, exercise, and emotions. For prediabetes, which is a stage where blood glucose is slightly elevated but has not yet reached the diagnostic criteria, the sensitivity of traditional indicators is only about 60%. This means that some prediabetic patients are not intervened in time because their indicators appear "normal," eventually developing into a confirmed diagnosis of diabetes.

[0004] In terms of screening convenience, traditional screening procedures are cumbersome, resulting in low user compliance. For example, the oral glucose tolerance test requires users to fast for 8 hours and have their blood drawn multiple times, taking over 3 hours in total, and the blood draw is an invasive procedure. For children and those with needle phobia, this method makes them resistant to screening, with actual compliance rates below 30%. Furthermore, traditional screening must be completed in a medical institution, failing to meet users' needs for home-based or immediate screening.

[0005] Based on the shortcomings of traditional diabetes risk screening models in terms of accuracy and convenience, new technological solutions are urgently needed to improve the efficiency, accuracy, and accessibility of diabetes risk screening. Summary of the Invention

[0006] Therefore, it is necessary to provide a method, device, computer equipment, and medium for detecting diabetes risk that can improve accuracy, in response to the above-mentioned technical problems.

[0007] A method for detecting the risk of diabetes, the method comprising:

[0008] S1. Identify abnormal data in the real-time detection data of the object to be tested; the abnormal data includes abnormal indicators and corresponding indicator values.

[0009] S2. Based on the correlation coefficient between each of the abnormal indicators and diabetes, determine the core indicators among the abnormal indicators;

[0010] S3. Select candidate factors from the knowledge graph that match the normalized core indicators; the knowledge graph includes multiple factors, and each factor includes an indicator and an indicator value.

[0011] S4. Based on the index values ​​of the candidate factors and the index values ​​of the core indicators, calculate the similarity, and determine the candidate factors with similarity greater than the similarity threshold as the target factors;

[0012] S5. Determine the associated factors in the knowledge graph that are related to the target factor, and determine the comprehensive risk value of the object to be tested based on the risk values ​​of the target factor and the associated factors.

[0013] In this embodiment, abnormal data in the real-time detection data of the test subject is identified. This abnormal data includes abnormal indicators and their corresponding values. Based on the correlation coefficient between each abnormal indicator and diabetes, the core indicators among the abnormal indicators are determined. This eliminates interference from redundant indicators and focuses on the core indicators that have the greatest impact on diabetes, significantly improving the scientific rigor and accuracy of risk assessment. Candidate factors matching the standardized core indicators are screened from a knowledge graph. The knowledge graph includes multiple factors, each containing indicators and values. Similarity is calculated based on the indicator values ​​of the candidate factors and the core indicators. Candidate factors with similarity greater than a similarity threshold are identified as target factors. This allows for reasoning based on the knowledge graph, improving the accuracy of associated factors in the knowledge graph related to the target factors. Ultimately, this improves the accuracy of the overall risk value of the test subject determined based on the risk values ​​of the target factors and associated factors.

[0014] In one embodiment, the knowledge graph generation process in step S2 includes:

[0015] Acquire historical detection data and corresponding historical detection results for multiple tested objects, and preprocess the historical detection data;

[0016] The preprocessed historical test data and historical test results are standardized; the historical test data includes multiple indicator data, the indicator data includes indicators and corresponding indicator values, and the types of indicators include physical examination indicators, lifestyle indicators and family medical history indicators;

[0017] Based on the historical detection data and the historical detection results, the risk impact relationship between each of the indicator data and the historical detection results is explored;

[0018] Based on the risk impact relationship of each indicator data and the type of indicator in the indicator data, the indicator data is labeled;

[0019] A knowledge graph is generated using the labeled indicator data as factors.

[0020] In this embodiment, historical detection data and corresponding historical detection results of multiple tested objects are acquired, and the historical detection data is preprocessed and standardized. The historical detection data includes multiple indicator data, each containing an indicator and its corresponding value. Indicator types include physical examination indicators, lifestyle habit indicators, and family medical history indicators. This enables the integration and standardization of multi-source heterogeneous data, improving data quality and providing a high-quality, computable structured data foundation for subsequent analysis. By mining the risk impact relationships between each indicator data and historical detection results based on the historical detection data and results, the nonlinear, multi-factor synergistic risk impact relationships between each indicator data and risk level can be revealed. Based on the risk impact relationships of each indicator data and the types of indicators within the data, the indicator data is labeled, generating an interpretable and reasonable knowledge graph.

[0021] In one embodiment, the normalization process for the preprocessed historical detection data and the historical detection results includes:

[0022] Risk levels are extracted from the preprocessed historical detection results, and indicators and corresponding indicator values ​​are extracted from the preprocessed historical detection data.

[0023] Based on a preset data dictionary and processing rule set, the extracted risk level, the indicator and the corresponding indicator value are standardized to obtain a standardized result.

[0024] In this embodiment, risk levels are extracted from preprocessed historical detection results, and indicators and corresponding indicator values ​​are extracted from preprocessed historical detection data. Based on a preset data dictionary and processing rule set, the extracted risk levels, indicators, and corresponding indicator values ​​are standardized to obtain standardized results. This eliminates differences in terminology, units, and formats, achieves semantic consistency, provides high-quality input for knowledge graphs, ensures the accuracy and consistency of the entire process, and allows data from different sources and in different forms to be effectively utilized within a unified framework.

[0025] In one embodiment, the risk impact relationship includes a first risk level, and the step of labeling the indicator data based on the risk impact relationship of each indicator data and the type of indicator in the indicator data includes:

[0026] When the indicator in the indicator data is the physical examination indicator or the lifestyle indicator, the indicator data is marked as a first marker pointing to the first risk level;

[0027] When the indicator in the indicator data is a family medical history indicator, the indicator data is labeled with a second label pointing from the first risk level to the indicator data.

[0028] In this embodiment, when the indicator in the indicator data is a physical examination indicator or a lifestyle indicator, the indicator data is marked as a first marker pointing from the indicator data to the first risk level. When the indicator in the indicator data is a family medical history indicator, the indicator data is marked as a second marker pointing from the first risk level to the indicator data. This can reflect the causal directionality in medical logic and enhance the semantic expression capability of the knowledge graph.

[0029] In one embodiment, the method further includes:

[0030] For the indicator data carrying the first mark, determine whether the first risk level pointed to by the indicator data is consistent with the risk level in the corresponding historical detection result. If they are consistent, then determine that the first mark is correct.

[0031] For the indicator data carrying the second mark, determine whether the indicator data pointed to by the first risk level matches the indicator data in the corresponding historical detection data. If they match, then determine that the second mark is correct.

[0032] In this embodiment, for indicator data carrying a first mark, it is determined whether the first risk level pointed to by the indicator data is consistent with the risk level in the corresponding historical detection results. If they are consistent, the first mark is determined to be correct. For indicator data carrying a second mark, it is determined whether the indicator data pointed to by the first risk level matches the indicator data in the corresponding historical detection data. If they match, the second mark is determined to be correct. In this way, the marks can be verified to ensure that the marks of each indicator data are correct.

[0033] In one embodiment, step S5 includes:

[0034] If the indicators in the target factor are physical examination indicators or lifestyle indicators, then the first label of the target factor is obtained, and the associated factors related to the target factor are determined from the knowledge graph along the direction indicated by the first label, starting from the target factor.

[0035] If the indicator in the target factor is a family medical history indicator, then the second label of the target factor is obtained, and the associated factors related to the target factor are determined from the knowledge graph along the direction indicated by the second label, starting from the risk level corresponding to the second label.

[0036] In this embodiment, when the indicators in the target factor are physical examination indicators or lifestyle indicators, a first label of the target factor is obtained. Starting from the target factor, associated factors are determined from the knowledge graph along the direction indicated by the first label. When the indicators in the target factor are family medical history indicators, a second label of the target factor is obtained. Starting from the risk level corresponding to the second label, associated factors are determined from the knowledge graph along the direction indicated by the second label. This avoids searching for a large number of weakly related or irrelevant factors, reduces false associations and noise interference, and improves the credibility of associated factors. At the same time, targeted searching along the direction indicated by the first / second label can reduce computational overhead.

[0037] In one embodiment, determining the comprehensive risk value of the object to be tested based on the risk values ​​of the target factor and the associated factor respectively includes:

[0038] Determine the second risk level corresponding to each of the target factor and the associated factor;

[0039] Based on the second risk level corresponding to each of the target factor and the associated factor, the risk value of each of the target factor and the associated factor is determined;

[0040] The risk values ​​of the target factor and the associated factor are weighted and summed to obtain the comprehensive risk value of the object to be tested.

[0041] In this embodiment, by determining the second risk level corresponding to each of the target factor and the associated factor, and based on the second risk level corresponding to each of the target factor and the associated factor, the risk value of each of the target factor and the associated factor is determined. The risk values ​​of each of the target factor and the associated factor are weighted and summed to obtain the comprehensive risk value of the object to be tested. This can avoid viewing a single indicator in isolation, truly reflect the overall risk level under the interaction of multiple factors, improve the comprehensiveness of the assessment, and obtain an accurate comprehensive risk value.

[0042] A diabetes risk detection device, the device comprising:

[0043] The data determination module is used to determine abnormal data in the real-time detection data of the object under test; the abnormal data includes abnormal indicators and corresponding indicator values.

[0044] The indicator screening module is used to determine the core indicators among the abnormal indicators based on the correlation coefficient between each abnormal indicator and diabetes.

[0045] The factor matching module is used to filter out candidate factors from the knowledge graph that match the normalized core indicators; the knowledge graph includes multiple factors, and each factor includes an indicator and an indicator value.

[0046] The similarity calculation module is used to calculate the similarity based on the index values ​​of the candidate factors and the index values ​​of the core indicators, and to determine the candidate factors whose similarity is greater than the similarity threshold as the target factors.

[0047] The risk value determination module is used to determine the associated factors in the knowledge graph that are related to the target factor, and to determine the comprehensive risk value of the object to be tested based on the risk values ​​of the target factor and the associated factors.

[0048] A computer device includes a memory and a processor, the memory storing a computer program, the processor executing the computer program to implement the steps of the method described above.

[0049] A computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of the above-described method.

[0050] The beneficial effects of the aforementioned diabetes risk detection device, computer equipment, and media are as follows: By identifying abnormal data in the real-time detection data of the test subject (including abnormal indicators and their corresponding values), and based on the correlation coefficient between each abnormal indicator and diabetes, the core indicators among the abnormal indicators are determined. This eliminates interference from redundant indicators and focuses on the core indicators that have the greatest impact on diabetes, significantly improving the scientific rigor and accuracy of risk assessment. Furthermore, by screening candidate factors that match the standardized core indicators from a knowledge graph (which includes multiple factors, each with indicators and values), similarity is calculated based on the indicator values ​​of the candidate factors and the core indicators. Candidate factors with similarity values ​​greater than a similarity threshold are identified as target factors. This allows for reasoning based on the knowledge graph, improving the accuracy of associated factors in the knowledge graph related to the target factors. Ultimately, this enhances the accuracy of the overall risk value of the test subject determined based on the risk values ​​of both the target factors and associated factors. Attached Figure Description

[0051] Figure 1 This is a diagram illustrating the application environment of a diabetes risk detection method in one embodiment.

[0052] Figure 2 This is a flowchart illustrating a method for detecting the risk of diabetes in one embodiment;

[0053] Figure 3 This is a structural block diagram of a diabetes risk detection device in one embodiment;

[0054] Figure 4 This is an internal structural diagram of a computer device in one embodiment. Detailed Implementation

[0055] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0056] The diabetes risk detection method provided in this application embodiment can be applied to, for example... Figure 1 In the application environment shown, terminal 102 interacts with server 104 via a wired / wireless channel. The data storage system can store the data that server 104 needs to process. The process includes: S1, identifying abnormal data in the real-time detection data of the target object; abnormal data includes abnormal indicators and their corresponding values; S2, determining the core indicators among the abnormal indicators based on the correlation coefficient between each abnormal indicator and diabetes; S3, filtering candidate factors that match the normalized core indicators from the knowledge graph; the knowledge graph includes multiple factors, each including indicators and indicator values; S4, calculating the similarity based on the indicator values ​​of the candidate factors and the core indicators, and identifying candidate factors with similarity greater than a similarity threshold as target factors; S5, determining the associated factors in the knowledge graph related to the target factors, and determining the comprehensive risk value of the target object based on the risk values ​​of the target factors and associated factors. Terminal 102 can be, but is not limited to, various personal computers, laptops, smartphones, tablets, IoT devices, etc. Server 104 can be a single server, a server cluster consisting of multiple servers, or a cloud computing center consisting of multiple servers.

[0057] In one embodiment, such as Figure 2 As shown, a method for detecting the risk of diabetes is provided, which can be applied to... Figure 1 Taking server 104 as an example, the following steps are included:

[0058] S1. Identify abnormal data in the real-time detection data of the object under test; abnormal data includes abnormal indicators and their corresponding indicator values.

[0059] The subjects being tested are those preparing for diabetes risk assessment. Real-time monitoring data includes data obtained from diabetes screening, lifestyle data, and family genetic data. Diabetes screening primarily measures data related to blood glucose levels. Real-time monitoring data from diabetes screening includes, but is not limited to, fasting blood glucose and glycated hemoglobin (HbA1c). Lifestyle data includes, but is not limited to, daily calorie, carbohydrate, and fat intake, and weekly exercise frequency, duration, and intensity. Family genetic data includes, but is not limited to, the generational ranking, number of people with diabetes, and duration of diabetes in the family.

[0060] Abnormal data refers to data in real-time monitoring that falls outside the normal range. For example, the normal range for fasting blood glucose is 3.9 mmol / L to 6.1 mmol / L. If the fasting blood glucose level in the real-time monitoring data is 7.5 mmol / L, then the fasting blood glucose level of 7.5 mmol / L is abnormal data, fasting blood glucose is the abnormal indicator, and 7.5 mmol / L is the corresponding indicator value for the abnormal indicator. Another example: the normal daily calorie intake should be 2000 kcal. If the subject consumes 3500 kcal daily, and high-sugar foods account for more than 50% of their intake, then the daily calorie intake of 3500 kcal is abnormal data, daily calorie intake is the abnormal indicator, and 3500 kcal is the corresponding indicator value for the abnormal indicator. If the subject only exercises for 1 hour per week, while the normal exercise duration is more than 5 hours, then exercising only for 1 hour per week is abnormal data, exercise duration is the abnormal indicator, and 1 hour is the corresponding indicator value for the abnormal indicator. For example, if the father of the subject is diagnosed with diabetes, then the father and the father generation with the disease are abnormal data. If both the grandfather and the father are ill, then the grandfather, father, father and the father generation with the disease are abnormal data.

[0061] S2. Based on the correlation coefficient between each abnormal indicator and diabetes, determine the core indicators among each abnormal indicator;

[0062] The correlation coefficients between each indicator and diabetes can be predetermined. Since abnormal indicators are also part of the indicators, their correlation coefficients can be selected from the overall correlation coefficients between each indicator and diabetes. The process of determining the correlation coefficients between each indicator and diabetes includes: using any indicator as the target indicator, determining the number of people with diabetes (n1) and the number of people without diabetes (n0) in multiple subjects; and calculating the standard deviation (s) of the target indicator values ​​for multiple subjects. x The average value of the test indicators of the patients who already have the disease The average value of the test indicators in non-disease-affected subjects Through formula Calculate the correlation coefficient r between the measured indicator and diabetes. pb, where n is the total number of objects.

[0063] Core indicators refer to those outlier indicators whose correlation coefficients are greater than a threshold. By filtering out outlier indicators with correlation coefficients greater than the threshold, the interference from redundant indicators can be reduced.

[0064] In some embodiments, core indicators can also be determined based on a preset indicator library. This library includes multiple indicators strongly correlated with diabetes. If an abnormal indicator matches any indicator in the preset library, that abnormal indicator is considered a core indicator. This saves time calculating correlation coefficients and improves efficiency.

[0065] S3. Select candidate factors from the knowledge graph that match the normalized core indicators; the knowledge graph includes multiple factors, and each factor includes an indicator and an indicator value.

[0066] The knowledge graph consists of multiple interconnected nodes. Different nodes represent different factors, risk levels of diabetes, and disease outcomes. Connections between nodes indicate the relationships between factors and between factors and risk levels. Correlation coefficients can be labeled on these connections to demonstrate the closeness of the relationships between factors and between factors and risk levels. The knowledge graph also displays the interactions between factors and the paths through which they affect risk levels. For example, the path of "elevated blood sugar → insulin resistance → medium-risk diabetes" is also represented in the graph, helping users intuitively understand the logical connections between various factors and diabetes risk levels, where elevated blood sugar and insulin resistance are factors, and medium risk is the risk level.

[0067] Knowledge graphs contain factors with consistent indicators, but they do not contain two factors with both consistent indicators and indicator values.

[0068] Normalization is the process of converting raw, unstructured, or semi-structured core indicators into standardized data with a unified format, standard units, clear structure, and machine readability. When normalizing core indicators, their values ​​are also normalized simultaneously. For example, fasting blood glucose data, described in natural language as "fasting blood glucose: 7.8 mmol / L," will be normalized to a format conforming to database field requirements, becoming "blood_glucose_fasting:7.8". Similarly, dietary calories may be expressed in different units such as "kilocalories" or "kilojoules," which will be uniformly converted to "kilocalories" during normalization, such as converting "5000 kilojoules" to "1195 kilocalories," ensuring unit consistency. Furthermore, core indicator values ​​described as "high" or "low" will be represented by specific codes or numerical ranges after normalization.

[0069] Before standardization, the core indicators and their values ​​were presented as follows: "Fasting blood glucose 7.8 mmol / L, daily dietary calorie intake 3200 kcal, exercise duration 1 hour per week (too low)." These were fragmented, naturally expressed raw data. After standardization, the core indicators and their values ​​are presented as follows: "blood_glucose_fasting: 7.8 (unit: mmol / L, reference range 3.9-6.1, abnormal); diet_calorie_daily: 3200 (unit: kcal, reference range 2000-2500, abnormal); exercise_duration_weekly: 1 (unit: hour, reference range 5-10, abnormal)."

[0070] In some embodiments, standardization is performed using a pre-defined factor database. Specifically, the factor database includes various standard indicators related to diabetes risk and their corresponding values, with each indicator having a standardized format. The core indicators and their corresponding values ​​are compared and matched with the content in the factor database; according to the format requirements of the fields, data types, etc., in the factor database, the core indicators and their values ​​are organized and transformed so that the transformed core indicators and their values ​​can be recognized and accepted by the factor database, thereby obtaining standardized core indicators and their values.

[0071] Candidate factors are those whose included indicators are consistent with the candidate indicators. For example, factor A includes fasting blood glucose, and the core indicator also includes fasting blood glucose, so factor A is a candidate factor that matches the core indicator.

[0072] In some embodiments, if there are no candidate factors in the knowledge graph that match the normalized core indicators, real-time detection data is supplemented or new real-time detection data is obtained through re-detection in order to obtain new core indicators, and candidate factors that match the new core indicators are screened from the knowledge graph.

[0073] In some embodiments, if there are no candidate factors in the knowledge graph that match the normalized core indicators, the core indicators that do not match the candidate factors are marked, and then the marked core indicators and their values ​​are evaluated in combination with expert experience assessment or risk model, so as to judge the diabetes risk of the subject based on the assessment results.

[0074] In some embodiments, if the number of core indicators that do not match candidate factors exceeds a threshold, the knowledge graph is updated. Specifically, if a large number of core indicators fail to match candidate factors, it indicates that the knowledge graph may have insufficient coverage, requiring an update to incorporate new factors and relationships to improve the success rate of subsequent matching.

[0075] S4. Based on the index values ​​of candidate factors and the index values ​​of core indicators, calculate the similarity and determine the candidate factors with similarity greater than the similarity threshold as the target factors.

[0076] Similarity refers to the similarity between the indicator values ​​of candidate factors and the indicator values ​​of core indicators. The calculation method for similarity differs depending on the type of indicator. For numerical indicators, the absolute difference method is used to calculate similarity. This involves calculating the absolute difference between the indicator value of the core indicator and the indicator value of the candidate factor, and then converting the difference into a similarity value by considering the normal range of the core indicator. Examples of numerical indicators include blood glucose, cholesterol, and dietary calorie intake. For instance, the normal range for the core indicator of fasting blood glucose is 3.9 mmol / L to 6.1 mmol / L. The fasting blood glucose value (core indicator) is 7.8 mmol / L, and the blood glucose value of the candidate factor is 8.0 mmol / L, resulting in a difference of 0.2 mmol / L. Based on the normal range of 6.1 - 3.9 = 2.2 mmol / L, the similarity is calculated as 1 - 0.2 / 2.2 = 0.91. For numerical indicators, the Euclidean distance method can be used to calculate similarity. The smaller the distance, the higher the similarity. The distance is then converted into a similarity value through a mapping relationship.

[0077] For categorization indicators, such as whether the affected relative is the father, mother, or other relative; and whether the diet is high-sugar, high-fat, or balanced, the similarity is 1 if the candidate factor's classification is completely consistent with the core indicator's; otherwise, the similarity is 0. For example, if the core indicator's affected relative is the father, and the candidate factor's affected relative is also the father, the similarity is 1; if the candidate factor's affected relative is the mother, the similarity is 0. The affected relative is an indicator in both the core indicator and the candidate factor, with father and mother being the indicator values.

[0078] For interval-based indicators, such as exercise duration, the general recommendation is more than 5 hours per week, categorized into insufficient and satisfactory intervals. Specifically, it's determined whether the core indicator's value and the candidate factor's value fall within the same interval. If they do, the similarity is 1; if they are adjacent intervals, partial similarity is calculated based on the interval's span; if they are not adjacent intervals, the similarity is 0. For example, if the general recommendation is more than 5 hours of exercise per week, and the input exercise duration (core indicator's value) is 4 hours per week, it falls into the insufficient interval. If the candidate factor's value is 3 hours per week, also falling into the insufficient interval, the similarity is 1. If the candidate factor's value is 6 hours per week, falling into the satisfactory interval, the similarity is 0.

[0079] The similarity threshold is set based on historical data and medical experience to determine whether the matching degree between the core indicator value and the candidate factor value is high enough. For example, the normal range for fasting blood glucose is 3.9 mmol / L to 6.1 mmol / L. The blood glucose value of the core indicator is 7.8 mmol / L, and the blood glucose value of the candidate factor is 8.0 mmol / L, with a difference of 0.2 mmol / L. Based on the range of the normal range (6.1 - 3.9 = 2.2 mmol / L), the similarity can be calculated as 1 − 0.2 / 2.2 = 0.91. Since the similarity threshold is 0.8, the candidate factor can be identified as the target factor.

[0080] In some embodiments, factors in the knowledge graph whose indicators are consistent with the core indicators and whose indicator values ​​are consistent with the core indicator values ​​can also be identified as target factors. For example, if the indicators in factor A are consistent with the core indicators and the indicator values ​​in factor A are also consistent with the core indicator values, then factor A is the target factor.

[0081] S5. Identify the associated factors in the knowledge graph that are related to the target factor, and determine the comprehensive risk value of the object to be tested based on the preset risk values ​​of the target factor and the associated factors.

[0082] Among them, the association factor is a factor in the knowledge graph that has a direct or indirect connection with the target factor.

[0083] In a knowledge graph, each factor is associated with a risk value. After determining the target factor and associated factors, the overall risk value of the subject can be determined based on the risk values ​​associated with each factor. This overall risk value is then used to determine the subject's risk level for diabetes. Specifically, a range of overall risk values ​​corresponding to each risk level is preset. The subject's risk level is determined according to the range of the overall risk value it falls within. For example, an overall risk value between 0 and 30 is considered low risk, 31 to 70 is medium risk, and 71 to 100 is high risk. Therefore, a subject with an overall risk value of 57 would be classified as having a medium risk level for diabetes.

[0084] The overall risk value can be determined by summing the risk values ​​of the target factor and the related factors, or by weighted summing of the risk values ​​of the target factor and the related factors.

[0085] The beneficial effects of the aforementioned diabetes risk detection method are as follows: By identifying anomalous data in the real-time monitoring data of the test subjects (including anomalous indicators and their corresponding values), and based on the correlation coefficient between each anomalous indicator and diabetes, the core indicators among the anomalous indicators are determined. This eliminates the interference of redundant indicators and focuses on the core indicators with the greatest impact on diabetes, significantly improving the scientific rigor and accuracy of risk assessment. Furthermore, by screening candidate factors that match the standardized core indicators from a knowledge graph (which includes multiple factors, each with indicators and values), similarity is calculated based on the indicator values ​​of the candidate factors and the core indicators. Candidate factors with similarity values ​​greater than a similarity threshold are identified as target factors. This allows for reasoning based on the knowledge graph, improving the accuracy of associated factors in the identified knowledge graph that are related to the target factors. Ultimately, this enhances the accuracy of the overall risk value of the test subjects determined based on the risk values ​​of the target factors and associated factors.

[0086] In one embodiment, the knowledge graph generation process in step S2 includes:

[0087] Acquire historical detection data and corresponding historical detection results for multiple tested objects, and preprocess the historical detection data;

[0088] The preprocessed historical test data and results are standardized. The historical test data includes multiple indicator data, which includes the indicator and its corresponding value. The types of indicators include physical examination indicators, lifestyle indicators and family medical history indicators.

[0089] Based on historical monitoring data and results, we explore the risk impact relationship between each indicator data and historical monitoring results;

[0090] Based on the risk impact relationships of each indicator data and the types of indicators in the indicator data, the indicator data is labeled;

[0091] A knowledge graph is generated using labeled indicator data as factors.

[0092] Historical testing data refers to the testing data obtained from diabetes testing of the tested subjects. Historical testing results represent a conclusion of low, medium, or high risk of diabetes after a diabetes risk assessment of the tested subjects. Each tested subject has a corresponding historical testing result.

[0093] Preprocessing includes, but is not limited to, cleaning historical test data and removing duplicate, erroneous, or severely missing data to ensure the accuracy and usability of historical test data.

[0094] Standardization is the process of converting raw, unstructured, or semi-structured historical test data and results into standardized data with a unified format, standard units, clear structure, and machine readability.

[0095] Physical examination indicators are objective indicators obtained through medical examinations that reflect the physiological or biochemical state of the human body. Lifestyle indicators refer to observable and recordable behavioral characteristics of the tested individual in daily life, reflecting their lifestyle choices and health-related behavioral patterns. Family medical history indicators are indicators characterizing the occurrence of specific diseases among the tested individual's immediate or extended family members. Physical examination indicators include, but are not limited to, fasting blood glucose and glycated hemoglobin. Lifestyle indicators include, but are not limited to, daily calorie intake and weekly exercise duration. Family medical history indicators include, but are not limited to, the incidence of diseases among immediate family members, the incidence of diseases among children, and the intergenerational transmission of diseases.

[0096] Risk-impact relationships are patterns or regularities that influence diabetes risk levels, extracted from historical testing data and results using data mining algorithms. They are primarily used to identify which indicators are key factors causing changes in risk levels and how they influence those levels. Data mining algorithms include, but are not limited to, decision trees and regression algorithms. For example, decision tree analysis might reveal that individuals with a fasting blood glucose level of 12 mmol / L and weekly exercise duration of less than 1 hour are more likely to be classified as high-risk.

[0097] Because the data used are historical detection data and corresponding historical detection results of multiple tested objects, multiple risk impact relationships are also discovered.

[0098] Labeling is used to indicate the risk level or data level associated with each indicator. For example, a risk-affect relationship might be that a fasting blood glucose level of 12 mmol / L and weekly exercise time of less than 1 hour are more likely to be classified as high-risk. Since fasting blood glucose is a physical examination indicator and exercise time is a lifestyle indicator, the labeling of the two indicators—fasting blood glucose of 12 mmol / L and exercise time of 1 hour—would be a high-risk label. As another example, if a risk-affect relationship exists where there is an intermediate risk of diabetes and a first-degree relative has the disease, then the labeling of the indicator "first-degree relative has the disease" would be a high-risk label.

[0099] Because the labels on indicator data have direction, a knowledge graph can be generated by connecting the indicator data with the objects (risk levels or indicator data) they point to, based on the risk levels or indicator data that each indicator data label points to. Indicator data and risk levels are the nodes in the knowledge graph, and risk impact relationships are the edges.

[0100] In some embodiments, statistical methods are used to calculate the correlation coefficient between each indicator data and the diabetes risk level, so as to quickly identify the indicator data that are closely related to the diabetes risk level through the correlation coefficient.

[0101] In some embodiments, all labeled indicator data and their risk impact relationships with risk levels are organized and visualized according to a certain logical structure; the knowledge graph displays the type and label of each indicator data, and how each indicator data affects the diabetes risk level. For example, the glycemic factor affects the risk level through the pathway of elevated blood glucose, insulin resistance, and increased diabetes risk; the family history factor is associated with the risk level through the presence of diabetes risk and tracing family genetic predisposition.

[0102] In this embodiment, historical detection data and corresponding historical detection results of multiple tested objects are acquired, and the historical detection data is preprocessed and standardized. The historical detection data includes multiple indicator data, each containing an indicator and its corresponding value. Indicator types include physical examination indicators, lifestyle habit indicators, and family medical history indicators. This enables the integration and standardization of multi-source heterogeneous data, improving data quality and providing a high-quality, computable structured data foundation for subsequent analysis. By mining the risk impact relationships between each indicator data and historical detection results based on the historical detection data and results, the nonlinear, multi-factor synergistic risk impact relationships between each indicator data and risk level can be revealed. Based on the risk impact relationships of each indicator data and the types of indicators within the data, the indicator data is labeled, generating an interpretable and reasonable knowledge graph.

[0103] In one embodiment, the preprocessed historical detection data and historical detection results are normalized, including:

[0104] Risk levels are extracted from preprocessed historical detection results, and indicators and corresponding indicator values ​​are extracted from preprocessed historical detection data.

[0105] Based on a pre-defined data dictionary and processing rule set, the extracted risk levels, indicators, and corresponding indicator values ​​are standardized to obtain standardized results.

[0106] The risk level refers to the risk information of the tested subjects regarding their risk of developing diabetes. Risk levels are categorized as high risk, medium risk, and low risk. High risk indicates a very high susceptibility to diabetes, low risk indicates a low susceptibility, and medium risk indicates a moderate probability of developing diabetes. Extracting the risk level clarifies the historical assessment of the tested subjects' risk of developing diabetes, which is crucial for subsequent risk comparisons and trend analyses. For example, it reveals the subsequent development of diabetes in individuals with different risk levels.

[0107] Extract indicators and corresponding values ​​from historical testing data. For example, extract blood glucose levels, dietary calories, exercise duration, and family medical history.

[0108] A data dictionary is a structured description of data fields, containing metadata information such as the name, meaning, type, unit, value range, and source of each field. It's essentially a "data instruction manual." The purpose of a data dictionary is to standardize terminology, avoid ambiguity, improve data readability, facilitate development, analysis, and maintenance, and support data exchange and integration between different systems.

[0109] The processing rule set is a set of predefined logical rules or algorithms used to transform, standardize, and manipulate the original risk levels, indicators, and corresponding indicator values. It is the key engine for realizing the transformation from "data to knowledge".

[0110] Standardizing the extracted risk levels, indicators, and corresponding indicator values ​​is the process of converting the raw, unstructured, or semi-structured risk levels, indicators, and corresponding indicator values ​​into standardized data with a unified format, standard units, clear structure, and machine readability.

[0111] In this embodiment, risk levels are extracted from preprocessed historical detection results, and indicators and corresponding indicator values ​​are extracted from preprocessed historical detection data. Based on a preset data dictionary and processing rule set, the extracted risk levels, indicators, and corresponding indicator values ​​are standardized to obtain standardized results. This eliminates differences in terminology, units, and formats, achieves semantic consistency, provides high-quality input for knowledge graphs, ensures the accuracy and consistency of the entire process, and allows data from different sources and in different forms to be effectively utilized within a unified framework.

[0112] In one embodiment, the risk impact relationship includes a first risk level. Based on the risk impact relationship of each indicator data and the type of indicator in the indicator data, the indicator data is labeled, including:

[0113] When the indicators in the indicator data are physical examination indicators or lifestyle indicators, the indicator data is marked as the first marker pointing to the first risk level;

[0114] When the indicator in the indicator data is a family medical history indicator, the indicator data is labeled with the second label pointing from the first risk level to the indicator data.

[0115] The first risk level refers to the risk information of the tested subjects having diabetes. The first risk level can be any one of high risk, medium risk, or low risk.

[0116] Since physical examination indicators and lifestyle indicators are usually considered factors that trigger risk outcomes—that is, the indicator data points to the first risk level—they are marked as the first label, representing the direction of influence from cause to effect. However, for family medical history indicators, when there is a risk of diabetes, the risk level can often be inversely correlated with the family medical history—that is, the first risk level points to the family medical history indicator—so they are marked as the second association label, representing the direction of influence from effect to cause. For example, a subject's blood glucose level at the time of the physical examination was 12 mmol / L, their daily calorie intake was 3000 kcal, which is about 2000 kcal higher than the healthy standard, and their weekly exercise time was 1 hour, which is less than the recommended 5 hours. Their historical test results indicate a high risk. After preprocessing and standardizing these historical test data, the generated indicator data includes a blood glucose level of 12 mmol / L, a daily calorie intake of 3000 kcal, and a weekly exercise time of 1 hour. Then, based on the identified risk-influence relationships, each indicator data is marked with a first label, indicating that the blood glucose level of 12 mmol / L, the daily calorie intake of 3000 kcal, and the weekly exercise time of 1 hour all point to a high risk.

[0117] The first risk level indicated by the indicator data refers to the first risk level included in the risk impact relationship to which the indicator data belongs. The indicator data indicated by the first risk level refers to the indicator data in the risk impact relationship to which that first risk level belongs.

[0118] In this embodiment, when the indicator in the indicator data is a physical examination indicator or a lifestyle indicator, the indicator data is marked as a first marker pointing from the indicator data to the first risk level. When the indicator in the indicator data is a family medical history indicator, the indicator data is marked as a second marker pointing from the first risk level to the indicator data. This can reflect the causal directionality in medical logic and enhance the semantic expression capability of the knowledge graph.

[0119] In one embodiment, the method further includes:

[0120] For indicator data carrying the first marker, determine whether the first risk level pointed to by the indicator data is consistent with the risk level in the corresponding historical detection results. If they are consistent, then the first marker is determined to be correct.

[0121] For indicator data carrying a second marker, determine whether the indicator data pointed to by the first risk level matches the indicator data in the corresponding historical monitoring data. If they match, then the second marker is confirmed to be correct.

[0122] In determining risk levels, the historical monitoring results and indicator data must belong to the same monitored object. For example, determine whether the first risk level indicated by the indicator data in the historical monitoring data of monitored object A is consistent with the risk level in the historical monitoring results of monitored object A.

[0123] When assessing indicator data, the first risk level and the historical monitoring data must belong to the same measured object. For example, assess whether the indicator data pointed to by the first risk level in the risk impact relationship of measured object A matches the indicator data in the historical monitoring data of measured object A. Whether the indicator data matches means whether the indicator and its corresponding indicator value are consistent.

[0124] In some embodiments, if the first risk level pointed to by the indicator data is inconsistent with the risk level in the historical detection results, or if the indicator data pointed to by the first risk level does not match the indicator data in the historical detection data, then the risk impact relationship between each indicator data and the historical detection results is re-mined, and the data is labeled based on the re-mined risk impact relationship.

[0125] In this embodiment, for indicator data carrying a first mark, it is determined whether the first risk level pointed to by the indicator data is consistent with the risk level in the corresponding historical detection results. If they are consistent, the first mark is determined to be correct. For indicator data carrying a second mark, it is determined whether the indicator data pointed to by the first risk level matches the indicator data in the corresponding historical detection data. If they match, the second mark is determined to be correct. In this way, the marks can be verified to ensure that the marks of each indicator data are correct.

[0126] In one embodiment, step S5 includes:

[0127] If the indicators in the target factor are physical examination indicators or lifestyle indicators, then the first label of the target factor is obtained. Starting from the target factor, the associated factors related to the target factor are determined from the knowledge graph along the direction indicated by the first label.

[0128] If the indicator in the target factor is a family medical history indicator, then the second label of the target factor is obtained. Starting from the risk level corresponding to the second label, the associated factors related to the target factor are determined from the knowledge graph along the direction indicated by the second label.

[0129] In this knowledge graph, each factor is labeled. Specifically, if the indicator in a factor is a physical examination indicator or a lifestyle indicator, the label on the factor is the first label pointing from the indicator data to the risk level; if the indicator in a factor is a family medical history indicator, the label on the factor is the second label pointing from the risk level to the indicator data, and the indicator data is the factor. Since the target factor is a factor in the knowledge graph, the first / second label of the target factor can be directly obtained. In the knowledge graph, the first label includes labels that directly point to the risk level and labels that indirectly point to the risk level. Directly pointing to the risk level means that the target factor of abnormal glucose metabolism directly points to high risk. Indirectly pointing means that the target factor of abnormal fasting blood glucose first points to the associated factor of abnormal glucose metabolism, and then further points from abnormal glucose metabolism to high risk.

[0130] Furthermore, if the next target factor points to a risk level, the overall risk value is determined directly based on the preset risk value of the target factor; if the next target factor points to a related factor, it is further determined whether to continue searching based on the object pointed to by the related factor. If the object pointed to by the related factor is still a factor, the search continues downward; if the object pointed to by the related factor is a risk level, the search stops, and the overall risk value is determined based on the risk values ​​of the target factor and the related factor.

[0131] The risk level corresponding to the second label is the risk level pointing to the target factor. For example, if medium risk points to the target factor, then the risk level corresponding to the second label of the target factor is medium risk.

[0132] In a specific application, when the indicators in the target factor are physical examination indicators or lifestyle indicators, the first marker of the target factor is obtained. The first marker indicates that the fasting blood glucose level of 7.8 mmol / L points to abnormal glucose metabolism, and further points to the risk of diabetes. At this time, starting from the fasting blood glucose level of 7.8 mmol / L, the associated factors are traced back from the knowledge graph along the direction indicated by the first marker. In the knowledge graph, the associated factor of abnormal glucose metabolism is first linked, and then from abnormal glucose metabolism are linked to other associated factors such as imbalance in insulin secretion regulation and decreased cellular sensitivity to insulin. When the indicator in the target factor is a family history indicator, the second marker of the target factor is obtained. The direction indicated by the second association marker is from the father having type 2 diabetes to genetic susceptibility, and then to the risk level of disease under the influence of offspring's lifestyle. Taking the result of the father having type 2 diabetes as the starting point, the association factors in the association mining graph are traced along the direction indicated by the second marker. In the knowledge graph, the father having type 2 diabetes is associated with the association factor of the presence of diabetes-related genetic locus mutations. Then, the presence of diabetes-related genetic locus mutations is associated with the association factors such as offspring being more susceptible to high-fat diets and the risk further increasing when offspring do not exercise enough.

[0133] In this embodiment, when the indicators in the target factor are physical examination indicators or lifestyle indicators, a first label of the target factor is obtained. Starting from the target factor, associated factors are determined from the knowledge graph along the direction indicated by the first label. When the indicators in the target factor are family medical history indicators, a second label of the target factor is obtained. Starting from the risk level corresponding to the second label, associated factors are determined from the knowledge graph along the direction indicated by the second label. This avoids searching for a large number of weakly related or irrelevant factors, reduces false associations and noise interference, and improves the credibility of associated factors. At the same time, targeted searching along the direction indicated by the first / second label can reduce computational overhead.

[0134] In one embodiment, the comprehensive risk value of the object to be measured is determined based on the risk values ​​of the target factor and the associated factors, including:

[0135] Determine the second risk level corresponding to each of the target factor and the associated factors;

[0136] Based on the second risk level corresponding to the target factor and the associated factor, the risk value of the target factor and the associated factor is determined.

[0137] The risk values ​​of the target factor and related factors are weighted and summed to obtain the comprehensive risk value of the object under test.

[0138] The second risk level can be a risk level directly or indirectly indicated by each factor in the knowledge graph. It can also be a risk level derived from historical testing data. For example, if analyzing a large amount of historical data reveals that a high percentage of people with fasting blood glucose levels between 7.0 and 8.0 mmol / L are classified as medium risk, then the second risk level for the factor of fasting blood glucose levels between 7.0 and 8.0 mmol / L is medium risk. Similarly, if a large percentage of people with daily calorie intake exceeding 2500 kcal are classified as medium risk, then the second risk level for the factor of daily calorie intake exceeding 2500 kcal is medium risk. Furthermore, if 50% of people have a family member with diabetes, then the second risk level for the factor of a family member with diabetes is medium risk.

[0139] In some embodiments, the second risk level can be determined by combining the risk levels directly or indirectly pointed to by each factor in the knowledge graph and the risk levels obtained based on historical detection data. Specifically, if the risk levels directly or indirectly pointed to by each factor in the knowledge graph are consistent with the risk levels obtained based on historical detection data, then the determined risk level is taken as the second risk level; if the risk levels directly or indirectly pointed to by each factor in the knowledge graph are inconsistent with the risk levels obtained based on historical detection data, then the risk levels directly or indirectly pointed to by each factor in the knowledge graph are taken as the second risk level.

[0140] Each second risk level has a pre-set corresponding risk value. Therefore, after determining the second risk levels for the target factor and related factors, the risk values ​​for each factor can be directly determined based on the second risk levels. For example, the risk value for low risk is 1 point, the risk value for medium risk is 3 points, and the risk value for high risk is 5 points. If the second risk level for the target factor is low risk, then the risk value for the target factor is 1 point.

[0141] The weights of the target factor and the associated factor can be determined based on the degree of influence of the factor on diabetes. A greater influence results in a larger weight, and a smaller influence results in a smaller weight. For example, fasting blood glucose levels have a significant impact on diabetes risk, so a weight of 0.4 is assigned; daily dietary calorie intake has a slightly less significant impact, so a weight of 0.3 is assigned; and the influence of a family member's diabetes condition has a slightly less significant impact, so a weight of 0.3 is assigned. Furthermore, the degree of influence can be determined based on the correlation coefficient between the factor and diabetes.

[0142] In this embodiment, by determining the second risk level corresponding to each of the target factor and the associated factor, and based on the second risk level corresponding to each of the target factor and the associated factor, the risk value of each of the target factor and the associated factor is determined. The risk values ​​of each of the target factor and the associated factor are weighted and summed to obtain the comprehensive risk value of the object to be tested. This can avoid viewing a single indicator in isolation, truly reflect the overall risk level under the interaction of multiple factors, improve the comprehensiveness of the assessment, and obtain an accurate comprehensive risk value.

[0143] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.

[0144] Based on the same inventive concept, this application also provides a diabetes risk detection device for implementing the aforementioned diabetes risk detection method. The solution provided by this device is similar to the implementation described in the above method; therefore, the specific limitations of one or more diabetes risk detection device embodiments provided below can be found in the limitations of the diabetes risk detection method described above, and will not be repeated here.

[0145] In one embodiment, such as Figure 3 As shown, a diabetes risk detection device is provided, comprising:

[0146] The data determination module is used to identify abnormal data in the real-time detection data of the object under test; abnormal data includes abnormal indicators and their corresponding indicator values.

[0147] The indicator screening module is used to determine the core indicators among the abnormal indicators based on the correlation coefficient between each abnormal indicator and diabetes.

[0148] The factor matching module is used to filter out candidate factors from the knowledge graph that match the normalized core indicators; the knowledge graph includes multiple factors, and each factor includes an indicator and an indicator value.

[0149] The similarity calculation module is used to calculate the similarity based on the index values ​​of candidate factors and the index values ​​of core indicators, and to identify candidate factors with similarity greater than the similarity threshold as target factors.

[0150] The risk value determination module is used to identify the associated factors in the knowledge graph that are related to the target factor, and to determine the comprehensive risk value of the object to be tested based on the risk values ​​of the target factor and the associated factors.

[0151] The various modules in the aforementioned diabetes risk monitoring device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in the processor of a computer device in hardware form or independent of it, or stored in the memory of a computer device in software form, so that the processor can call and execute the corresponding operations of each module.

[0152] In one embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 4 As shown, the computer device includes a processor, memory, and a network interface connected via a system bus. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and a database. The internal memory provides the environment for the operation of the operating system and computer programs in the non-volatile storage media. The database stores real-time detection data, anomaly data, correlation coefficients, core indicators, candidate factors, similarity, target factors, correlation factors, risk values ​​of the target factors and correlation factors, and the overall risk value. The network interface communicates with external terminals via a network connection. When executed by the processor, the computer program implements a method for detecting the risk of diabetes.

[0153] Those skilled in the art will understand that Figure 4 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0154] In one embodiment, a computer device is also provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps in the above method embodiments.

[0155] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon that, when executed by a processor, implements the steps in the above method embodiments.

[0156] In one embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps in the above method embodiments.

[0157] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments described above. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, etc., and are not limited to these.

[0158] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0159] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.

Claims

1. A method for detecting the risk of diabetes, characterized in that, The method includes: S1. Identify abnormal data in the real-time detection data of the object to be tested; the abnormal data includes abnormal indicators and corresponding indicator values. S2. Based on the correlation coefficient between each of the abnormal indicators and diabetes, determine the core indicators among the abnormal indicators; S3. Select candidate factors from the knowledge graph that match the normalized core indicators; the knowledge graph includes multiple factors, and each factor includes an indicator and an indicator value. S4. Based on the index values ​​of the candidate factors and the index values ​​of the core indicators, calculate the similarity, and determine the candidate factors with similarity greater than the similarity threshold as target factors; S5. Determine the associated factors in the knowledge graph that are related to the target factor, and determine the comprehensive risk value of the object to be tested based on the risk values ​​of the target factor and the associated factors respectively; The method further includes: When the candidate factor and the core indicator are numerical indicators, the absolute difference between the indicator value of the core indicator and the indicator value of the candidate factor is calculated, and the absolute difference is converted into similarity in combination with the normal range of the core indicator. When the candidate factor and the core indicator are categorical indicators, the similarity is determined according to the type to which the indicator value of the candidate factor and the indicator value of the core indicator belong respectively; When the candidate factor and the core indicator are interval-type indicators, the similarity is determined based on the interval to which the indicator value of the core indicator and the indicator value of the candidate factor belong.

2. The method according to claim 1, characterized in that, The knowledge graph generation process described in step S2 includes: Acquire historical detection data and corresponding historical detection results for multiple tested objects, and preprocess the historical detection data; The preprocessed historical test data and historical test results are standardized; the historical test data includes multiple indicator data, the indicator data includes indicators and corresponding indicator values, and the types of indicators include physical examination indicators, lifestyle indicators and family medical history indicators; Based on the historical detection data and the historical detection results, the risk impact relationship between each of the indicator data and the historical detection results is explored; Based on the risk impact relationship of each indicator data and the type of indicator in the indicator data, the indicator data is labeled; A knowledge graph is generated using the labeled indicator data as factors.

3. The method according to claim 2, characterized in that, The standardization process for the preprocessed historical detection data and the historical detection results includes: Risk levels are extracted from the preprocessed historical detection results, and indicators and corresponding indicator values ​​are extracted from the preprocessed historical detection data. Based on a preset data dictionary and processing rule set, the extracted risk level, the indicator and the corresponding indicator value are standardized to obtain a standardized result.

4. The method according to claim 2, characterized in that, The risk impact relationship includes a first risk level. The labeling of the indicator data based on the risk impact relationship of each indicator data and the type of indicator in the indicator data includes: When the indicator in the indicator data is the physical examination indicator or the lifestyle indicator, the indicator data is marked as a first marker pointing to the first risk level; When the indicator in the indicator data is a family medical history indicator, the indicator data is labeled with a second label pointing from the first risk level to the indicator data.

5. The method according to claim 4, characterized in that, The method further includes: For the indicator data carrying the first mark, determine whether the first risk level pointed to by the indicator data is consistent with the risk level in the corresponding historical detection result. If they are consistent, then determine that the first mark is correct. For the indicator data carrying the second mark, determine whether the indicator data pointed to by the first risk level matches the indicator data in the corresponding historical detection data. If they match, then determine that the second mark is correct.

6. The method according to claim 1, characterized in that, Step S5 includes: If the indicators in the target factor are physical examination indicators or lifestyle indicators, then the first label of the target factor is obtained, and the associated factors related to the target factor are determined from the knowledge graph along the direction indicated by the first label, starting from the target factor. If the indicator in the target factor is a family medical history indicator, then the second label of the target factor is obtained, and the associated factors related to the target factor are determined from the knowledge graph along the direction indicated by the second label, starting from the risk level corresponding to the second label.

7. The method according to claim 1, characterized in that, The process of determining the comprehensive risk value of the object under test based on the risk values ​​of the target factor and the associated factor includes: Determine the second risk level corresponding to each of the target factor and the associated factor; Based on the second risk level corresponding to each of the target factor and the associated factor, the risk value of each of the target factor and the associated factor is determined; The risk values ​​of the target factor and the associated factor are weighted and summed to obtain the comprehensive risk value of the object to be tested.

8. A diabetes risk detection device, characterized in that, The device includes: The data determination module is used to determine abnormal data in the real-time detection data of the object under test; the abnormal data includes abnormal indicators and corresponding indicator values. The indicator screening module is used to determine the core indicators among the abnormal indicators based on the correlation coefficient between each abnormal indicator and diabetes. The factor matching module is used to filter out candidate factors from the knowledge graph that match the normalized core indicators; the knowledge graph includes multiple factors, and each factor includes an indicator and an indicator value. The similarity calculation module is used to calculate the similarity based on the indicator values ​​of the candidate factors and the indicator values ​​of the core indicators, and to identify the candidate factors whose similarity is greater than a similarity threshold as target factors; when the candidate factors and the core indicators are numerical indicators, the module calculates the absolute difference between the indicator values ​​of the core indicators and the candidate factors, and converts the absolute difference into a similarity value by combining it with the normal range of the core indicators; when the candidate factors and the core indicators are categorical indicators, the module determines the similarity based on the type to which the indicator values ​​of the candidate factors and the core indicators belong; when the candidate factors and the core indicators are interval indicators, the module determines the similarity based on the interval to which the indicator values ​​of the core indicators and the candidate factors belong. The risk value determination module is used to determine the associated factors in the knowledge graph that are related to the target factor, and to determine the comprehensive risk value of the object to be tested based on the risk values ​​of the target factor and the associated factors.

9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, It stores a computer program that, when executed by a processor, implements the steps of the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Computer equipment, system and readable storage medium

    CN110289101A

  • DQN network-based sarcopenia risk factor analysis device

    CN116759084A