A method, system, device, medium and program product for predicting insulin resistance

By classifying and training the user's medical and questionnaire data, a model for insulin resistance prediction was generated, which solved the problem that high-cost detection methods in the existing technology could not capture the potential risks of insulin resistance, and achieved the effect of accurate prediction and reducing medical costs.

CN119650100BActive Publication Date: 2025-05-27湖南工商大学
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510150629.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-11
Publication Date
2025-05-27
Estimated Expiration
2045-02-11

AI Technical Summary

Technical Problem

The prior art cannot capture the potential risks of insulin resistance through high-cost insulin resistance detection methods, resulting in lower early signals and early screening capabilities for identifying insulin resistance.

Method used

By obtaining the user's medical data and questionnaire data, the original data set is characterized, and cost-effective and cost-free feature data are obtained. Based on these data, the pre-constructed model is trained, and the first prediction model and the second prediction model are generated for insulin resistance prediction for the predicted user.

Benefits of technology

Accurate prediction of insulin resistance is achieved, medical costs are reduced, and screening capabilities for insulin resistance risks are improved, and signals of insulin resistance can be identified early.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119650100B_ABST
    Figure CN119650100B_ABST
Patent Text Reader

Abstract

The present invention discloses a method, system, device, medium and program product for predicting insulin resistance. The method includes: classifying an original data set into cost feature data and non-cost feature data, where the original data set includes user medical data and user questionnaire data; obtaining a first prediction model and a second prediction model, the first prediction model being trained based on the cost feature data and the non-cost feature data, and the second prediction model being trained based on the non-cost feature data; determining a first type of user and a second type of user among the users to be predicted; inputting the first type of user data and the second type of user data into the first prediction model and the second prediction model respectively for insulin resistance prediction, so as to provide corresponding prediction models for different user groups, and realizing accurate prediction of insulin resistance by combining questionnaire data and medical data, without using high-cost insulin resistance detection, effectively reducing medical costs and improving the risk screening ability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data processing, and particularly to a method, system, device, medium and program product for predicting insulin resistance. Background Art

[0002] With the changes in lifestyle and the rising incidence of chronic diseases, the prevalence trend of insulin resistance has become increasingly severe. Insulin resistance (IR) has become an important challenge in the field of public health. Insulin resistance refers to the state in which the biological effects of insulin in the body are weakened, and it is also the core pathophysiological factor of various diseases such as type 2 diabetes, cardiovascular diseases and metabolic syndrome. In this context, the need for early identification and intervention of insulin resistance is becoming increasingly urgent.

[0003] However, traditional insulin resistance detection methods, such as Insulin Tolerance Test (ITT) and Euglycemic Hyperinsulinemic Clamp, are difficult to be applied on a large scale due to their complexity, invasiveness and high cost, so there are limitations. And due to the individual differences in physiology and environment among different subjects, the current high-cost insulin resistance detection methods cannot capture the potential risks of insulin resistance, thus identifying the early signals of insulin resistance, resulting in a low early screening ability for insulin resistance. Summary of the Invention

[0004] The main purpose of the present invention is to provide a method, system, device, medium and program product for predicting insulin resistance, aiming to solve the technical problem that the existing high-cost insulin resistance detection methods cannot capture the potential risks of insulin resistance, thus identifying the early signals of insulin resistance, resulting in a low early screening ability for insulin resistance.

[0005] To achieve the above purpose, the present invention provides a method for predicting insulin resistance, the method comprising the following steps:

[0006] Obtain an original data set, the original data set including sample data and label data corresponding to the sample data, the sample data including user medical data and user questionnaire data;

[0007] Perform feature classification on the original data set to obtain cost feature data and cost-free feature data;

[0008] Train a pre - constructed original model based on the cost - bearing feature data and the cost - free feature data to obtain a first prediction model and a second prediction model. The original model includes an ensemble tree model. The first prediction model is obtained by training based on the cost - bearing feature data and the cost - free feature data, and the second prediction model is obtained by training based on the cost - free feature data;

[0009] Classify the users to be predicted, determine the first - type users and the second - type users among the users to be predicted, and obtain the first - type user data and the second - type user data. The first - type user data includes the user questionnaire data and user medical data of the first - type users, and the second - type user data includes the user questionnaire data of the second - type users;

[0010] Input the first - type user data into the first prediction model for insulin resistance prediction, and input the second - type user data into the second prediction model for insulin resistance prediction to obtain the prediction results of the first - type users and the prediction results of the second - type users.

[0011] Optionally, the step of training a pre - constructed original model based on the cost - bearing feature data and the cost - free feature data to obtain a first prediction model and a second prediction model includes:

[0012] Train multiple pre - constructed original models based on the cost - bearing feature data and the cost - free feature data to obtain multiple first candidate models and multiple second candidate models. The first candidate models are obtained by training based on the cost - bearing feature data and the cost - free feature data, and the second candidate models are obtained by training based on the cost - free feature data;

[0013] Evaluate each of the first candidate models and each of the second candidate models to obtain candidate evaluation results;

[0014] Conduct a correlation analysis on the cost - bearing feature data and the cost - free feature data, and perform feature screening on the cost - bearing feature data and the cost - free feature data based on the correlation analysis results and the candidate evaluation results to obtain first target data and second target data;

[0015] Train the first candidate models and the second candidate models according to the first target data and the second target data to obtain a first prediction model and a second prediction model. The first prediction model is obtained by training based on the first target data and the second target data, and the second prediction model is obtained by training based on the second target data.

[0016] Optionally, the correlation analysis result includes a continuous feature analysis result and a discrete feature analysis result; the correlation analysis of the cost-bearing feature data and the cost-free feature data, and the feature screening of the cost-bearing feature data and the cost-free feature data based on the correlation analysis result and the candidate evaluation result to obtain first target data and second target data includes:

[0017] Determine the continuous feature data and discrete feature data in the cost-bearing feature data and the cost-free feature data;

[0018] Determine the linear correlation coefficients between the continuous features in the continuous feature data, and perform correlation analysis on the continuous feature data based on the linear correlation coefficients to obtain a continuous feature analysis result;

[0019] Perform a chi-square test on each discrete feature in the discrete feature data to obtain a discrete feature analysis result;

[0020] Perform correlation dimensionality reduction on the cost-bearing feature data and the cost-free feature data based on the continuous feature analysis result and the discrete feature analysis result;

[0021] Evaluate the importance of each feature in the cost-bearing feature data and the cost-free feature data according to the correlation dimensionality reduction result and the candidate evaluation result;

[0022] Perform feature screening on the cost-bearing feature data and the cost-free feature data based on the importance evaluation result to obtain first target data and second target data.

[0023] Optionally, the classification of the user to be predicted, determining the first type of users and the second type of users in the user to be predicted, and obtaining the first type of user data and the second type of user data includes:

[0024] Classify the user to be predicted to determine the first type of users and the second type of users in the user to be predicted;

[0025] Generate a first type of questionnaire text based on the first target data and the second target data, and generate a second type of questionnaire text based on the second target data;

[0026] Send the first type of questionnaire text to the first type of users, and send the second type of questionnaire text to the second type of users;

[0027] Obtain the first type of response text feedback by the first type of users based on the first type of questionnaire text and the second type of response text feedback by the second type of users based on the second type of questionnaire text;

[0028] Perform semantic analysis on the first type of response text to obtain the first type of user data;

[0029] Perform semantic analysis on the second type of response text to obtain the second type of user data.

[0030] Optionally, the obtaining of the original data set includes:

[0031] Collect user questionnaire data and user medical data;

[0032] Aggregate the user questionnaire data and the user medical data based on the user identity identifier of the user questionnaire data and the user medical data to obtain an initial multi-dimensional data set;

[0033] Perform data cleaning on the initial multi-dimensional data set according to the user identity identifier to obtain a candidate multi-dimensional data set;

[0034] Perform dependency analysis on the candidate multi-dimensional data set to obtain the dependency relationship information between the features in the candidate multi-dimensional data set;

[0035] Fill in the missing data in the candidate multi-dimensional data set based on the dependency relationship information;

[0036] Perform standardization processing on the candidate multi-dimensional data set with missing data filled to obtain the original data set.

[0037] Optionally, after inputting the first type of user data into the first prediction model for insulin resistance prediction and inputting the second type of user data into the second prediction model for insulin resistance prediction to obtain the prediction results of the first type of users and the prediction results of the second type of users, it includes:

[0038] Perform risk classification on the first type of users and the second type of users based on the prediction results of the first type of users and the prediction results of the second type of users to determine high-risk users and low-risk users;

[0039] Perform semantic analysis on the response texts of the high-risk users and the low-risk users to obtain user response information;

[0040] Generate a resistance question set according to a pre-constructed insulin resistance knowledge graph, where the insulin resistance knowledge graph includes medical knowledge information, symptom information, risk factor information, and diagnostic criterion information of insulin resistance;

[0041] Determine the text matching degree between the user response information of the high-risk users and the low-risk users and the resistance question set based on a deep interaction text matching model;

[0042] Judging insulin resistance of the high-risk users and the low-risk users according to the text matching degree.

[0043] In addition, to achieve the above object, the present invention also provides an insulin resistance prediction system, which includes:

[0044] A data acquisition module, configured to acquire an original data set, where the original data set includes sample data and label data corresponding to the sample data, and the sample data includes user medical data and user questionnaire data;

[0045] A feature classification module, configured to perform feature classification on the original data set to obtain cost-bearing feature data and cost-free feature data;

[0046] A model training module, configured to train a pre-constructed original model based on the cost-bearing feature data and the cost-free feature data to obtain a first prediction model and a second prediction model, where the original model includes an ensemble tree model, the first prediction model is obtained by training based on the cost-bearing feature data and the cost-free feature data, and the second prediction model is obtained by training based on the cost-free feature data;

[0047] A user classification module, configured to classify a user to be predicted, determine a first type of user and a second type of user in the user to be predicted, and acquire first type of user data and second type of user data, where the first type of user data includes the user questionnaire data and user medical data of the first type of user, and the second type of user data includes the user questionnaire data of the second type of user;

[0048] An insulin resistance prediction module, configured to input the first type of user data into the first prediction model for insulin resistance prediction, and input the second type of user data into the second prediction model for insulin resistance prediction to obtain a prediction result of the first type of user and a prediction result of the second type of user.

[0049] In addition, to achieve the above object, the present application also provides an insulin resistance prediction device, which includes: a memory, a processor, and a computer program stored on the memory and executable on the processor, where the computer program is configured to implement the steps of the insulin resistance prediction method as described above.

[0050] In addition, to achieve the above object, the present application also provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the steps of the insulin resistance prediction method as described above.

[0051] In addition, to achieve the above object, the present application further provides a computer program product, the computer program product includes a computer program, and when the computer program is executed by a processor, it implements the steps of the insulin resistance prediction method described above.

[0052] The present invention obtains an original data set, the original data set includes sample data and label data corresponding to the sample data, the sample data includes user medical data and user questionnaire data, classifies the features of the original data set to obtain cost-bearing feature data and cost-free feature data, trains a pre-constructed original model based on the cost-bearing feature data and the cost-free feature data to obtain a first prediction model and a second prediction model, the original model includes an ensemble tree model, the first prediction model is obtained by training based on the cost-bearing feature data and the cost-free feature data, the second prediction model is obtained by training based on the cost-free feature data, classifies the user to be predicted to determine the first type of users and the second type of users among the users to be predicted, and obtains the first type of user data and the second type of user data, the first type of user data includes the user questionnaire data and user medical data of the first type of users, the second type of user data includes the user questionnaire data of the second type of users, inputs the first type of user data into the first prediction model for insulin resistance prediction, and inputs the second type of user data into the second prediction model for insulin resistance prediction to obtain the prediction results of the first type of users and the prediction results of the second type of users; since the present invention classifies the original data set into cost-bearing feature data and cost-free feature data, corresponding prediction models are trained for different user groups, so as to provide appropriate prediction means for different user groups, effectively reduce computing resources while ensuring prediction accuracy, combine questionnaire data and medical data to achieve accurate prediction of insulin resistance, do not need to use high-cost insulin resistance detection means, effectively reduce medical costs, and predict insulin resistance through a prediction model based on user questionnaire data and user medical data, so as to accurately capture the potential risks of users, effectively identify the early signals of insulin resistance, greatly improve the screening ability of insulin resistance risk, and thus formulate intervention measures for users' health problems in advance. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] The drawings here are incorporated into the specification and form a part of this specification, showing embodiments consistent with the present application, and are used together with the specification to explain the principles of the present application.

[0054] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0055] Figure 1 It is a schematic structural diagram of an insulin resistance prediction device in the hardware operating environment related to the solution of the embodiment of the present invention;

[0056] Figure 2 It is a schematic flowchart of an embodiment of the insulin resistance prediction method of the present invention;

[0057] Figure 3 It is a schematic diagram of the data processing flow of an embodiment of the insulin resistance prediction method of the present invention;

[0058] Figure 4 It is a schematic diagram of the model training process of an embodiment of the insulin resistance prediction method of the present invention;

[0059] Figure 5 It is a schematic diagram of the data optimization process of an embodiment of the insulin resistance prediction method of the present invention;

[0060] Figure 6 It is a schematic diagram of the insulin resistance judgment process of an embodiment of the insulin resistance prediction method of the present invention;

[0061] Figure 7 It is a schematic diagram of the basic information questionnaire of an embodiment of the insulin resistance prediction method of the present invention;

[0062] Figure 8 It is a schematic diagram of the eating habit questionnaire of an embodiment of the insulin resistance prediction method of the present invention;

[0063] Figure 9 It is a schematic diagram of the health literacy questionnaire of an embodiment of the insulin resistance prediction method of the present invention;

[0064] Figure 10 It is a structural block diagram of an embodiment of the insulin resistance prediction system of the present invention.

[0065] The realization, functional characteristics and advantages of the object of the present invention will be further described with reference to the embodiments and the accompanying drawings. Detailed implementation manners

[0066] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0067] Refer to Figure 1 , Figure 1Schematic diagram of the insulin resistance prediction device for the hardware operating environment involved in the solution of the embodiment of the present invention.

[0068] As Figure 1 shown, the insulin resistance prediction device may include: a processor 1001, such as a central processing unit (CPU), a communication bus 1002, a user interface 1003, a network interface 1004, and a memory 1005. Among them, the communication bus 1002 is used to realize the connection and communication between these components. The user interface 1003 may include a display screen (Display) and an input unit such as a keyboard (Keyboard). Optionally, the user interface 1003 may further include a standard wired interface and a wireless interface. The network interface 1004 may optionally include a standard wired interface and a wireless interface (such as a wireless-fidelity (WI-FI) interface). The memory 1005 may be a high-speed random access memory (Random Access Memory, RAM), or a stable non-volatile memory (Non-Volatile Memory, NVM), such as a disk memory. Optionally, the memory 1005 may also be a storage system independent of the aforementioned processor 1001.

[0069] Those skilled in the art can understand that Figure 1 the structure shown in

[0070] does not constitute a limitation on the insulin resistance prediction device, and may include more or fewer components than shown in the figure, or combine certain components, or have different component arrangements. Figure 1 shown, as a computer-readable storage medium, the memory 1005 may include an operating system, a network communication module, a user interface module, and an insulin resistance prediction program.

[0071] In Figure 1 the insulin resistance prediction device shown, the network interface 1004 is mainly used for data communication with a network server; the user interface 1003 is mainly used for data interaction with a user; the processor 1001 and the memory 1005 in the insulin resistance prediction device of the present invention may be provided in the insulin resistance prediction device. The insulin resistance prediction device calls the insulin resistance prediction program stored in the memory 1005 through the processor 1001 and executes the insulin resistance prediction method provided by the embodiment of the present invention.

[0072] The embodiment of the present invention provides an insulin resistance prediction method. Refer to Figure 2 , Figure 2 which is a schematic flowchart of an embodiment of the insulin resistance prediction method of the present invention.

[0073] In this embodiment, the method for predicting insulin resistance comprises the following steps:

[0074] Step S10: Obtain the original data set.

[0075] It should be noted that the execution subject of this embodiment can be a computing service device with data processing, network communication and program running functions, such as a tablet computer, a personal computer, a mobile phone, etc., or a terminal electronic device capable of realizing the above functions, etc. The following takes an insulin resistance prediction device (referred to as the prediction device) as an example to illustrate this embodiment and the following embodiments.

[0076] It should be noted that the original data set includes sample data and label data corresponding to the sample data, and the sample data includes user medical data and user questionnaire data.

[0077] It should be noted that the questionnaire data can be the questionnaire data answered by the user collected by the prediction device through a pre-built questionnaire system. The above questionnaire system focuses on the comprehensive collection of multi-dimensional data, covering key areas such as physiological characteristics, lifestyle and psychological state, so as to reveal the potential risk factors of insulin resistance. The questionnaire content includes basic physiological information such as age, gender, body mass index, lifestyle information such as eating habits, exercise frequency, smoking and drinking history, as well as the assessment of psychological states such as mood swings and sleep quality.

[0078] In some embodiments, the questionnaire system can generate corresponding questionnaire questions for the user based on the user's initial answers, such as user identity, user type, etc., so as to adjust the content and order of subsequent questions, thereby optimizing data collection efficiency without increasing the burden on the user.

[0079] In some embodiments, users can fill out questionnaires online through a self-service platform, and the system will process the user's answers in real time and generate preliminary health risk assessment results after submission. This convenience not only lowers the threshold for filling out questionnaires, but also enables users to obtain health feedback anytime and anywhere, improving the overall user experience.

[0080] In some embodiments, in order to ensure the integrity and reliability of the questionnaire data, the system introduces multiple optimization strategies in the data processing link. First, for the missing data in the questionnaire, the system uses methods such as mean filling, mode filling or logical inference to reduce the impact of missing data on model performance. Secondly, standardize all collected data and unify the data into a standard format and range to facilitate the subsequent use of machine learning models. In addition, to further improve the efficiency of the model, the system will screen the questionnaire data and eliminate features that are irrelevant to or have low contribution to the prediction of insulin resistance, thereby optimizing the data input structure.

[0081] It should be noted that the above user medical data may include two key dimensions: blood biochemical indicators and physical examination data. Blood biochemical indicators include important biomarkers reflecting metabolism such as fasting blood glucose, triglycerides, high-density lipoprotein, and low-density lipoprotein, while physical examination data includes basic physiological information such as blood pressure, waist circumference, and BMI. The introduction of the above data aims to capture the complexity and diversity of insulin resistance, enabling the system to conduct a more detailed assessment of high-risk individuals.

[0082] In some embodiments, the prediction device can collect user medical data and user questionnaire data, match the user questionnaire data with the medical data using an individual identifier (such as an ID number), and integrate the data based on the matching result to form an overall view of multi-dimensional data fusion. The model reasonably classifies and optimizes data from different sources through hierarchical modeling, separately processes easily accessible and difficult-to-access data, and avoids potential problems caused by data inconsistency.

[0083] In some embodiments, the user questionnaire data can refer to Table 1 below, which is a table of questionnaire feature data; the user medical data can refer to Table 2, which is a table of medical feature data:

[0084]

[0085]

[0086] Further, referring to Figure 3 , Figure 3 is a schematic diagram of the data processing flow of an embodiment of the present invention. In order to improve the data quality of the original data set and thus improve the prediction efficiency, the above step S10 may include:

[0087] Step S101: Collect user questionnaire data and user medical data;

[0088] Step S102: Aggregate the user questionnaire data and the user medical data based on the user identity identifier of the user questionnaire data and the user medical data to obtain an initial multi-dimensional data set;

[0089] Step S103: Clean the data of the initial multi-dimensional data set according to the user identity identifier to obtain a candidate multi-dimensional data set;

[0090] Step S104: Conduct a dependency analysis on the candidate multi-dimensional data set to obtain dependency relationship information between the features in the candidate multi-dimensional data set;

[0091] Step S105: Fill in the missing data in the candidate multi-dimensional data set based on the dependency relationship information;

[0092] Step S106: Standardize the candidate multi-dimensional dataset with missing data filled to obtain the original dataset.

[0093] In some embodiments, the prediction device can collect the physiological and biochemical characteristic data of an individual through questionnaire data. After data collection, the questionnaire data and blood data collected at different times are merged and integrated into a unified dataset according to the physical examination number and ID number, and then the duplicate data is randomly de-duplicated according to the ID number, and some data that is difficult to code is deleted.

[0094] In some embodiments, for the missing data, the prediction device can view the missing values of each feature from the feature dimension and perform visualization processing. Features with the number of missing values exceeding 30% of the total data are deleted.

[0095] It should be noted that features with a subordinate relationship cannot be processed as ordinary missing values. For example, features with a subordinate relationship: the subordinate features of "Do you smoke?" [How many cigarettes do you smoke per day, the number of years of continuous smoking, how long have you quit smoking]; the subordinate features of "Do you drink alcohol?" [What kind of alcohol do you drink, how many times a week do you drink, how many taels do you drink each time, the number of years of continuous drinking, how long have you quit drinking]; the subordinate features of "Do you exercise?" [Exercise method, how many times a week do you exercise, how long do you exercise each time, how many years have you persisted in exercising]; the subordinate features of how is your sleep: [Main manifestations, reasons affecting sleep]. Among the features with a subordinate relationship, the number of missing values of the features [how long have you quit smoking, how long have you quit drinking] is too large, and they are directly deleted. For other subordinate relationship features with fewer missing values, the missing values are assigned 0.

[0096] In some embodiments, for features without a subordinate relationship and with the number of missing values less than 30%, continuous features are filled with the average value, and discrete features are filled with the mode; for features with a small number of missing values, the rows with missing values are directly deleted.

[0097] In some embodiments, considering the accuracy of medical data testing, the prediction device can perform average value inspection and replacement on the four features of [age, height, weight, waist circumference].

[0098] Step S20: Classify the features of the original dataset to obtain cost-bearing feature data and cost-free feature data.

[0099] It should be noted that the cost-bearing feature data can be the feature data that is difficult to obtain in the original dataset, such as biochemical index data, blood lipid index data, blood glucose index data, etc. The above cost-free obtained data can be the data that is easy to obtain in the original dataset, such as eating habits, smoking characteristics, drinking characteristics, etc.

[0100] In some embodiments, the prediction device may classify the features in the dataset into cost-free feature1 and cost-bearing feature2. The cost-free feature1 is further divided into physiological features, medical history, eating habits, smoking category, drinking category, exercise category, work category, mental state, sleep and health monitoring, health habits and prevention, and health indicators; the cost-bearing feature2 is divided into biochemical indicators, blood lipid indicators, blood glucose indicators, and protein tables. According to the ease of acquisition, the data is divided into two levels: easily acquired data: feature1, and easily acquired data + difficultly acquired data: feature1 + feature2.

[0101] Step S30: Train a pre-constructed original model based on the cost-bearing feature data and the cost-free feature data to obtain a first prediction model and a second prediction model.

[0102] It should be noted that the original model includes an ensemble tree model. The first prediction model is obtained by training based on the cost-bearing feature data and the cost-free feature data, and the second prediction model is obtained by training based on the cost-free feature data.

[0103] In some embodiments, the prediction device may construct multiple original ensemble tree models, generate a training set and a test set based on the cost-bearing feature data and the cost-free feature data, train each original ensemble data model based on the training set, test and evaluate each trained original ensemble tree model based on the test set, and select the model with the optimal performance from multiple original ensemble tree models as the prediction model based on the test and evaluation results.

[0104] Further, referring to Figure 4 , Figure 4 which is a schematic diagram of the model training process in an embodiment of the present invention. To improve the model performance, the above step S30 may include:

[0105] Step S301: Train a pre-constructed multiple original models based on the cost-bearing feature data and the cost-free feature data to obtain multiple first candidate models and multiple second candidate models;

[0106] Step S302: Evaluate each first candidate model and each second candidate model to obtain candidate evaluation results;

[0107] Step S303: Perform a correlation analysis on the cost-bearing feature data and the cost-free feature data, and perform feature screening on the cost-bearing feature data and the cost-free feature data based on the correlation analysis results and the candidate evaluation results to obtain first target data and second target data;

[0108] Step S304: Train the first candidate model and the second candidate model based on the first target data and the second target data to obtain a first prediction model and a second prediction model.

[0109] It should be noted that the first candidate model is obtained by training based on the cost-bearing feature data and the cost-free feature data, and the second candidate model is obtained by training based on the cost-free feature data. The first prediction model is obtained by training based on the first target data and the second target data, and the second prediction model is obtained by training based on the second target data.

[0110] In some embodiments, the prediction device may construct five ensemble tree models respectively, such as catboost, RandomForest (RF), lightgbm, xgboost, and GradientBoostingCart (GBC). The ensemble tree model is an ensemble learning method that improves the overall performance and generalization ability of the model by constructing multiple base decision trees and integrating their prediction results. The key advantage of the ensemble tree model is that it can reduce the overfitting problem that a single decision tree may encounter, and enhance the generalization ability of the model to new data by integrating the viewpoints of multiple models.

[0111] In some embodiments, the prediction device may use 80% of the dataset as the training set and 20% as the test set for model training and evaluation. Through grid search, the optimal hyperparameter combination in the machine learning model is determined. Calculate the F1 score, accuracy, recall, precision, and area under the curve (AUC) to evaluate the performance of these models. The overall performance of the models can be evaluated through the above metrics and their applicability in clinical decision-making can be determined.

[0112] In some embodiments, the prediction device may first train multiple pre-constructed original models based on the cost-free feature data to obtain an initial model, then evaluate the initial model, and train the initial model based on the initial evaluation result and the cost-bearing feature data to obtain a candidate model. Then, evaluate the candidate model, optimize and screen the cost-bearing feature data and the cost-free feature data based on the candidate evaluation result, remove the data with lower importance and relevance, and retain the data with higher importance and relevance to obtain the first target data and the second target data. Train the candidate model based on the first target data and the second target data, where the first target data is obtained by optimizing and screening the cost-bearing feature data, and the second target data is obtained by optimizing and screening the cost-free feature data.

[0113] In some embodiments, based on data integration, the system trains and optimizes multi-source data through advanced machine learning algorithms to ensure the efficiency and reliability of the prediction model. To address the complexity and diversity of non-linear feature relationships and high-dimensional data, the system uses multiple ensemble tree models (such as CatBoost and LightGBM) for modeling. Through a hierarchical data training strategy, the system models the easily obtainable questionnaire data and the difficult-to-obtain medical data separately and integrates them into a unified prediction framework to effectively reduce the impact of data differences on model performance while ensuring the scientificity and stability of the overall prediction results.

[0114] Next, through the Bayesian optimization method, the system automatically searches for the best combination of model hyperparameters, thus significantly improving the accuracy and efficiency of the model. In addition, the system conducts in-depth feature importance analysis on the fused data to identify the key features that contribute most to insulin resistance prediction, such as fasting blood glucose, BMI, and triglycerides, etc., providing a reliable scientific basis for further clinical analysis.

[0115] Further, referring to Figure 5 , Figure 5 which is a schematic diagram of the data optimization process in an embodiment of the present invention. To improve data quality and thus improve model performance, the correlation analysis results include continuous feature analysis results and discrete feature analysis results. The above step S303 may include:

[0116] Step S3031: Determine the continuous feature data and discrete feature data in the cost-bearing feature data and the cost-free feature data;

[0117] Step S3032: Determine the linear correlation coefficients between the continuous features in the continuous feature data, and conduct correlation analysis on the continuous feature data based on the linear correlation coefficients to obtain continuous feature analysis results;

[0118] Step S3033: Conduct chi-square tests on the discrete features in the discrete feature data to obtain discrete feature analysis results;

[0119] Step S3034: Conduct correlation dimensionality reduction on the cost-bearing feature data and the cost-free feature data based on the continuous feature analysis results and the discrete feature analysis results;

[0120] Step S3035: Conduct importance evaluation on each feature in the cost-bearing feature data and the cost-free feature data according to the correlation dimensionality reduction results and the candidate evaluation results;

[0121] Step S3036: Conduct feature screening on the cost-bearing feature data and the cost-free feature data based on the importance evaluation results to obtain the first target data and the second target data.

[0122] It should be noted that the prediction device can reduce the overfitting problem of the model through sampling training, balance the class distribution, improve the recognition ability of the model for minority classes, and the accuracy of the model after oversampling has been improved in the training set.

[0123] In some embodiments, the prediction device can classify the cost feature data and the non-cost feature data to determine the continuous feature data and the discrete feature data therein, perform correlation analysis on the continuous feature data using pearson, and use chi-square test for correlation analysis of discrete data. By observing the heat map, it can be seen that there are strongly correlated features in the modules of eating habits, smoking, exercise, drinking, biochemical indicators, blood lipid indicators, blood glucose indicators, and protein indicators.

[0124] For example, in the chi-square test, self-measured blood pressure and heart rate, eating animal offal, how many times of drinking per week, actively obtaining medical knowledge are significantly correlated with the training data, while depression, frustration, tension, difficulty in relaxation, anxiety, restlessness, and irritability are significantly uncorrelated with the training data.

[0125] For example, serum albumin, age, total bile acid, and low-density lipoprotein cholesterol are significantly uncorrelated with the training data, while creatinine, aspartate aminotransferase, and albumin / globulin ratio are significantly correlated with the training data. By observing the first four columns, it can be seen that the distributions of indicators such as age and serum albumin are relatively concentrated, and the distribution ranges of features such as creatinine and aspartate aminotransferase are relatively wide. From the group means and group standard deviations, it can be seen that as the training data increases, the means of creatinine, total bile acid, aspartate aminotransferase, total cholesterol, etc. also increase, that is, the training data is positively correlated with these features. At this time, the means of serum albumin, albumin / globulin ratio, direct bilirubin, total bilirubin, etc. decrease, indicating that the training data is negatively correlated with age.

[0126] In some embodiments, the prediction device can delete the features that fail the statistical test, delete the data with low correlation degree, and retain the data with high correlation degree. In some embodiments, the features that pass the correlation test are analyzed, and it can be observed that the performance of each model has been improved after removing the irrelevant features from the model.

[0127] In some embodiments, the prediction device can perform correlation dimensionality reduction on the cost feature data and the non-cost feature data. For example, for the three strongly correlated features in "physiological features": [body mass index BMI, height, weight, waist circumference], dimensionality reduction is performed, and body mass index BMI is selected. For the two strongly correlated features of [right upper arm systolic blood pressure, right upper arm diastolic blood pressure], the dimensionality reduction is reduced, and "right upper arm systolic blood pressure" is selected.

[0128] In some embodiments, the prediction device can perform importance analysis on the feature data, so as to quantify the importance of the feature data, screen the features based on the quantification results, and select the best-performing machine learning model according to the results of model evaluation. Perform feature importance analysis on the selected model to determine the impact of the features on the prediction, and the importance of each feature is calculated according to the internal mechanism of the model.

[0129] For example, the feature data retained in the cost-free feature data may include: Body Mass Index (BMI), systolic blood pressure of the right upper arm, age, physical intensity at work, sitting duration outside work, staple food structure, how often to have a physical examination, whether to eat fruits, whether to drink milk, sleep duration, self-measured blood pressure and heart rate, how many years of regular exercise, salt, gender, how much meat to eat per day on average, whether to drink coffee, observe defecation and urination, legumes and bean products, taste of diet, carry first-aid medicine when going out, have three meals on time, exercise time each time, drink sugary drinks, normal pulse, whether to eat fatty meat, how much vegetables to eat per day on average, how many times to exercise per week, seat belt, sun exposure.

[0130] The feature data retained in the cost-bearing feature data may include: fasting blood glucose (GLU), high-density lipoprotein cholesterol, triglycerides, creatinine, uric acid, alanine aminotransferase, total bilirubin, urea nitrogen, total bile acid, low-density lipoprotein cholesterol, serum albumin, aspartate aminotransferase, total serum protein, total cholesterol, albumin / globulin ratio, serum globulin, direct bilirubin.

[0131] In some embodiments, the prediction device can find the optimal model hyperparameters through Bayesian optimization.

[0132] In some embodiments, the prediction device can provide feature importance analysis inside the model and model interpretation using methods such as SHAP values. For example, the top five features in terms of model feature importance and SHAP analysis are BMI (Body Mass Index), SBP (Systolic Blood Pressure), Salt intake, Exercise type, Family history of medicine.

[0133] In some embodiments, the top five features in terms of model feature importance and SHAP analysis are BMI (Body Mass Index), FBG (Fasting Blood Glucose), HDL-C (High-Density Lipoprotein Cholesterol), TG (Triglycerides), Cr (Creatinine).

[0134] In some embodiments, for each prediction result, a local explanation is provided to help understand the logic behind a specific prediction. When the BMI value is constant and less than -0.3 and the gender is female, at this time, the contribution of BMI to the model is positive, and the insulin prediction value will increase accordingly. When the BMI value is constant and greater than -0.3, the SHAP value of BMI will gradually decrease, indicating that the negative effect of BMI on the model increases, and the insulin prediction value will decrease accordingly. When the SBP value is constant and less than -0.4, looking vertically, the SHAP value of SBP gradually increases, indicating that when people have the same SBP, as the BMI increases, the positive effect of SBP on the model gradually increases, and the prediction value of insulin will increase accordingly. When the SBP value is constant and greater than -0.4, the contribution of the SBP value to the model has a negative effect.

[0135] Step S40: Classify the user to be predicted, determine the first type of users and the second type of users among the users to be predicted, and obtain the first type of user data and the second type of user data.

[0136] It should be noted that the first type of user data includes the user questionnaire data and user medical data of the first type of users, and the second type of user data includes the user questionnaire data of the second type of users.

[0137] It should be noted that the first type of users can be users who have medical data such as fasting blood glucose and blood oxygen value and questionnaire data. For example, the first type of users can be medical staff. The second type of users are ordinary users who only have daily data such as height and weight and questionnaire data.

[0138] It can be understood that in this embodiment, the users to be predicted can be classified based on the identity type of the users and / or the data types owned by the users, and the user types of each user to be predicted are determined, so as to determine the first type of users and the second type of users among them.

[0139] It should be understood that since different types of users have different data types, in this embodiment, by classifying the users to be predicted, the user group is divided, and corresponding prediction methods are provided for different types of users, ensuring that while improving the prediction accuracy, the calculation load is reduced.

[0140] Further, in order to accurately collect the user data of the user to be predicted, the above step S40 may include:

[0141] Step S401: Classify the user to be predicted, determine the first type of users and the second type of users among the users to be predicted;

[0142] Step S402: Generate the first type of questionnaire text based on the first target data and the second target data, and generate the second type of questionnaire text based on the second target data;

[0143] Step S403: Send the first type of questionnaire text to the first type of users, and send the second type of questionnaire text to the second type of users;

[0144] Step S404: Obtain the first type of response text feedback by the first type of users based on the first type of questionnaire text and the second type of response text feedback by the second type of users based on the second type of questionnaire text;

[0145] Step S405: Conduct semantic analysis on the first type of response text to obtain the first type of user data;

[0146] Step S406: Conduct semantic analysis on the second type of response text to obtain the second type of user data.

[0147] It should be noted that the first type of users can be users who have medical data such as fasting blood glucose and blood oxygen values and questionnaire data. For example, the first type of users can be medical staff. The second type of users are ordinary users who only have daily data such as height and weight and questionnaire data.

[0148] It can be understood that this embodiment can classify the users to be predicted based on the user identity type and / or the data type owned by the users, determine the user types of each user to be predicted, so as to determine the first type of users and the second type of users among them, and generate corresponding questionnaire texts for different types of users.

[0149] In a specific implementation, the prediction device can send the first type of questionnaire text containing cost feature data and non-cost feature data to the first type of users, and send the second type of questionnaire text containing only non-cost feature data to the second type of users.

[0150] In some embodiments, the prediction device can generate questionnaire texts by combining the first target data and the second target data obtained after optimizing the training data during the model training process. Since the first target data and the second target data are obtained by screening and cleaning based on cost feature data and non-cost feature data, generating questionnaire texts based on the first target data and the second target data can improve the accuracy of the questionnaire texts and avoid generating unnecessary, unimportant features that are less helpful for the accuracy of the prediction results in the questionnaire texts.

[0151] For example, the cost-bearing features in the first type of questionnaire text may include: fasting blood glucose (GLU), high-density lipoprotein cholesterol, triglyceride, creatinine, blood uric acid, alanine aminotransferase, total bilirubin, urea nitrogen, total bile acid, low-density lipoprotein cholesterol, serum albumin, aspartate aminotransferase, total serum protein, total cholesterol, albumin / globulin ratio, serum globulin, direct bilirubin, etc.; the cost-free features in the first type of questionnaire text may include: body mass index (BMI), systolic blood pressure of the right upper arm, age, physical intensity at work, sitting duration outside work, staple food structure, frequency of physical examinations, fruit consumption, milk consumption, sleep duration, self-measured blood pressure and heart rate, etc.

[0152] Step S50: Input the first type of user data into the first prediction model for insulin resistance prediction, and input the second type of user data into the second prediction model for insulin resistance prediction, to obtain the prediction results of the first type of users and the prediction results of the second type of users.

[0153] It should be noted that the first prediction model is trained based on cost-bearing feature data and cost-free feature data, and the first type of users are the users to be predicted who have cost-free feature data and cost-bearing feature data. Therefore, input the user data of the first type of users into the first prediction model for insulin resistance prediction. The second prediction model is trained based on cost-free feature data, and the second type of users are the users to be predicted who have cost-free feature data. Therefore, input the user data of the second type of users into the second prediction model for insulin resistance prediction, so as to provide corresponding prediction means for different types of users, reduce the calculation amount, and improve the prediction accuracy at the same time.

[0154] It can be understood that in this embodiment, a deep learning model is trained based on user questionnaire data and user medical data, thereby effectively reducing the detection cost and the invasiveness of the detection, greatly improving the accessibility and efficiency of the screening. In this embodiment, by adopting non-invasive technologies and low-cost detection methods, the prediction of insulin resistance becomes more convenient and economical, thus providing a popularization basis for the early screening of insulin resistance.

[0155] In some embodiments, the prediction device can observe the risk of an individual developing diabetes under different insulin resistance states by analyzing and displaying the relationship between insulin resistance and the incidence of diabetes through cohort studies.

[0156] Further, referring to Figure 6 , Figure 6 which is a schematic diagram of the insulin resistance judgment process in an embodiment of the present invention. In order to accurately identify the insulin resistance risk, after the above step S50, the following may be included:

[0157] Step S501: Classify the first type of users and the second type of users according to the prediction results of the first type of users and the second type of users, and determine high-risk users and low-risk users;

[0158] Step S502: Perform semantic analysis on the response texts of the high-risk users and the low-risk users to obtain user response information;

[0159] Step S503: Generate a set of resistance questions based on the pre-constructed insulin resistance knowledge graph;

[0160] Step S504: Based on the deep interactive text matching model, determine the text matching degree between the user response information of the high-risk users and the low-risk users and the set of resistance questions;

[0161] Step S505: Judge insulin resistance for the high-risk users and the low-risk users according to the text matching degree.

[0162] It should be noted that the insulin resistance knowledge graph includes medical knowledge information, symptom information, risk factor information, and diagnostic criterion information of insulin resistance.

[0163] In some embodiment sets, the prediction device identifies key information and biomarkers related to insulin resistance by performing detailed semantic analysis on the user's response text.

[0164] Then, the prediction device can generate a set of resistance questions based on the pre-constructed insulin resistance knowledge graph, which includes medical knowledge, symptoms, risk factors, and diagnostic criteria related to insulin resistance, so as to generate a series of possible questions. These questions will be used to further confirm the user's health status.

[0165] In some embodiments, the prediction device can pre-construct a question similarity detection model, that is, a deep interactive text matching model, for calculating the similarity between the user's response and the questions generated by the system. This model can evaluate the matching degree between the user's response and the set of questions related to insulin resistance, so as to judge whether the user's response indicates the possibility of insulin resistance.

[0166] In some embodiments, if the prediction device determines that the user's response indicates the risk of insulin resistance, it will push relevant analysis and suggestions to help the user understand their health status, and may recommend further medical consultation or examination. If the user's response is not related to insulin resistance, the prediction device will provide corresponding feedback, pointing out that their response is not related to the insulin resistance question, so as to guide the user to ask more specific questions.

[0167] In some embodiments, the prediction device may construct an automatic analysis system, which is composed of two core subsystems: a problem generation and similarity detection system, and an analysis and push system. The problem generation and similarity detection system includes a problem generation model based on medical knowledge and a deep interactive text matching model; the analysis and push system includes an analysis model based on customer answers and a model for sorting analysis results according to matching scores. Through the collaborative work of these two subsystems, the system can provide efficient and accurate automatic analysis services for insulin resistance problems.

[0168] In some embodiments, in order to achieve early screening of insulin resistance risk, the present invention designs a set of efficient and scientific questionnaire systems to collect basic health data of individuals in a non-invasive, economical and convenient way. Through carefully designed questionnaire content, multi-dimensional data collection, and real-time processing technology, the system can quickly identify potential high-risk individuals, laying a solid foundation for subsequent medical data analysis. The questionnaire system takes the comprehensive collection of multi-dimensional data as the core, covering key areas such as physiological characteristics, lifestyle, and mental state, in order to reveal potential risk factors for insulin resistance. The questionnaire content includes basic physiological information such as age, gender, body mass index, lifestyle information such as eating habits, exercise frequency, smoking and drinking history, and the assessment of mental state such as mood swings and sleep quality, referring to Figure 7 , Figure 8 and Figure 9 , where Figure 7 is a schematic diagram of the basic information questionnaire, Figure 8 is a schematic diagram of the eating habits questionnaire, Figure 9 is a schematic diagram of the health literacy questionnaire.

[0169] In this embodiment, by obtaining an original data set, the original data set includes sample data and label data corresponding to the sample data, the sample data includes user medical data and user questionnaire data, classifying features of the original data set to obtain cost-bearing feature data and cost-free feature data, training a pre-constructed original model based on the cost-bearing feature data and the cost-free feature data to obtain a first prediction model and a second prediction model, the original model includes an ensemble tree model, the first prediction model is obtained by training based on the cost-bearing feature data and the cost-free feature data, the second prediction model is obtained by training based on the cost-free feature data, classifying a user to be predicted to determine a first type of user and a second type of user among the users to be predicted, and obtaining first type of user data and second type of user data, the first type of user data includes the user questionnaire data and user medical data of the first type of user, the second type of user data includes the user questionnaire data of the second type of user, inputting the first type of user data into the first prediction model for insulin resistance prediction, and inputting the second type of user data into the second prediction model for insulin resistance prediction to obtain a prediction result of the first type of user and a prediction result of the second type of user; Since in this embodiment, the original data set is classified into cost-bearing feature data and cost-free feature data, corresponding prediction models are trained for different user groups, so as to provide appropriate prediction means for different user groups, effectively reducing computing resources while ensuring prediction accuracy, realizing accurate prediction of insulin resistance by combining questionnaire data and medical data, without using high-cost insulin resistance detection means, effectively reducing medical costs, predicting insulin resistance through a prediction model based on user questionnaire data and user medical data, accurately capturing the potential risks of users, effectively identifying early signals of insulin resistance, greatly improving the screening ability of insulin resistance risk, and thus formulating intervention measures for users' health problems in advance.

[0170] In addition, an embodiment of the present invention further provides a computer-readable storage medium, on which an insulin resistance prediction program is stored, and when the insulin resistance prediction program is executed by a processor, the steps of the insulin resistance prediction method as described above are implemented.

[0171] The computer-readable storage medium provided by this application can be, for example, a USB flash drive, but is not limited to electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or components, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: electrical connections with one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM) or flash memory, optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above. In this embodiment, the computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, device, or component. The program code contained on the computer-readable storage medium can be transmitted using any appropriate medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.

[0172] The above computer-readable storage medium can be included in the insulin resistance prediction device; it can also exist independently without being assembled into the insulin resistance prediction device.

[0173] In addition, an embodiment of the present invention also provides a computer program product, including an insulin resistance prediction program. When the insulin resistance prediction program is executed by a processor, it implements the steps of the insulin resistance prediction method described above.

[0174] The specific implementation manner of the computer program product of the present invention is basically the same as that of each embodiment of the above insulin resistance prediction method, and will not be elaborated here.

[0175] Refer to Figure 10 , Figure 10 which is the structural block diagram of an embodiment of the insulin resistance prediction system of the present invention.

[0176] As Figure 10 shown, the insulin resistance prediction system proposed by the embodiment of the present invention includes:

[0177] A data acquisition module 10, configured to acquire an original data set, where the original data set includes sample data and label data corresponding to the sample data, and the sample data includes user medical data and user questionnaire data;

[0178] A feature classification module 20, configured to perform feature classification on the original data set to obtain cost feature data and non-cost feature data;

[0179] A model training module 30, configured to train a pre-constructed original model based on the cost feature data and the cost-free feature data to obtain a first prediction model and a second prediction model. The original model includes an ensemble tree model. The first prediction model is obtained by training based on the cost feature data and the cost-free feature data, and the second prediction model is obtained by training based on the cost-free feature data;

[0180] A user classification module 40, configured to classify a user to be predicted, determine a first type of user and a second type of user among the users to be predicted, and obtain first type of user data and second type of user data. The first type of user data includes user questionnaire data and user medical data of the first type of user, and the second type of user data includes user questionnaire data of the second type of user;

[0181] An insulin resistance prediction module 50, configured to input the first type of user data into the first prediction model for insulin resistance prediction, and input the second type of user data into the second prediction model for insulin resistance prediction, to obtain a prediction result of the first type of user and a prediction result of the second type of user.

[0182] In this embodiment, by obtaining an original data set, where the original data set includes sample data and label data corresponding to the sample data, and the sample data includes user medical data and user questionnaire data, feature classification is performed on the original data set to obtain cost-bearing feature data and cost-free feature data. Based on the cost-bearing feature data and the cost-free feature data, a pre-constructed original model is trained to obtain a first prediction model and a second prediction model. The original model includes an ensemble tree model. The first prediction model is obtained by training based on the cost-bearing feature data and the cost-free feature data, and the second prediction model is obtained by training based on the cost-free feature data. The users to be predicted are classified to determine the first-type users and the second-type users among the users to be predicted, and the first-type user data and the second-type user data are obtained. The first-type user data includes the user questionnaire data and user medical data of the first-type users, and the second-type user data includes the user questionnaire data of the second-type users. The first-type user data is input into the first prediction model for insulin resistance prediction, and the second-type user data is input into the second prediction model for insulin resistance prediction to obtain the prediction results of the first-type users and the prediction results of the second-type users. Since in this embodiment, the original data set is classified into cost-bearing feature data and cost-free feature data, corresponding prediction models are trained for different user groups, so as to provide appropriate prediction means for different user groups. While ensuring prediction accuracy, the computing resources are effectively reduced. By combining questionnaire data and medical data, accurate prediction of insulin resistance is achieved without using high-cost insulin resistance detection means, effectively reducing the medical cost. Insulin resistance prediction is performed through a prediction model based on user questionnaire data and user medical data, accurately capturing the potential risks of users, effectively identifying the early signals of insulin resistance, and greatly improving the screening ability of insulin resistance risks, so as to formulate intervention measures for users' health problems in advance.

[0183] The insulin resistance prediction system provided by this application adopts the insulin resistance prediction method in the above embodiment and can solve the technical problems of insulin resistance prediction. Compared with the prior art, the beneficial effects of the insulin resistance prediction system provided by this application are the same as those of the insulin resistance prediction method provided by the above embodiment, and the other technical features in the insulin resistance prediction system are the same as the features disclosed in the method of the above embodiment and will not be elaborated here.

[0184] It should be understood that the above is only an example and does not constitute any limitation to the technical solution of the present invention. In specific applications, those skilled in the art can set according to needs, and the present invention does not limit this.

[0185] It should be noted that the workflow described above is only illustrative and does not limit the scope of protection of the present invention. In actual applications, those skilled in the art can select some or all of them according to actual needs to achieve the purpose of the solution of this embodiment, and no limitation is made here.

[0186] In addition, for the technical details not described in detail in this embodiment, reference can be made to the insulin resistance prediction method provided in any embodiment of the present invention, and details will not be repeated here.

[0187] It should be noted that in this article, the term "including", "comprising" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or system including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or further includes elements inherent to such process, method, article or system. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of another identical element in the process, method, article or system including that element.

[0188] The serial numbers of the above embodiments of the present invention are only for description and do not represent the advantages or disadvantages of the embodiments.

[0189] Through the description of the above embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present invention, in essence, or the part that makes contributions to the prior art, can be embodied in the form of a software product. The computer software product is stored in a storage medium (such as a read-only memory / random access memory, magnetic disk, optical disk), and includes several instructions for causing a terminal device (which can be a mobile phone, a computer, a server, an air conditioner, or a network device, etc.) to execute the methods described in various embodiments of the present invention.

[0190] The above are only the preferred embodiments of the present invention, and do not limit the patent scope of the present invention accordingly. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of the present invention, or directly or indirectly applied to other related technical fields, shall be equally included in the patent protection scope of the present invention.

Claims

1. A method for predicting insulin resistance, characterized in that: The insulin resistance prediction method comprises: Acquire an original data set, wherein the original data set includes sample data and label data corresponding to the sample data, wherein the sample data includes user medical data and user questionnaire data, wherein the user questionnaire data includes lifestyle data, psychological state data, eating habit data, exercise frequency data, smoking and drinking history data, sleep quality data, and mood fluctuation data; Performing feature classification on the original data set to obtain cost feature data and cost-free feature data, wherein the cost feature data is feature data that is difficult to obtain in the original data set, and the cost-free feature data is feature data that is easy to obtain in the original data set; Training a pre-built original model based on the cost feature data and the cost-free feature data to obtain a first prediction model and a second prediction model, wherein the original model includes an integrated tree model, the first prediction model is obtained by training based on the cost feature data and the cost-free feature data, and the second prediction model is obtained by training based on the cost-free feature data; Classifying the users to be predicted, determining the first type of users and the second type of users among the users to be predicted, and acquiring the first type of user data and the second type of user data, wherein the first type of user data includes the user questionnaire data and the user medical data of the first type of users, and the second type of user data includes the user questionnaire data of the second type of users; Inputting the first type of user data into the first prediction model to predict insulin resistance, and inputting the second type of user data into the second prediction model to predict insulin resistance, to obtain prediction results for the first type of users and prediction results for the second type of users; The pre-built original model is trained based on the cost feature data and the cost-free feature data to obtain a first prediction model and a second prediction model, including: Training a plurality of pre-constructed original models based on the cost feature data and the cost-free feature data to obtain a plurality of first candidate models and a plurality of second candidate models, wherein the first candidate models are obtained by training based on the cost feature data and the cost-free feature data, and the second candidate models are obtained by training based on the cost-free feature data; Evaluate each first candidate model and each second candidate model to obtain a candidate evaluation result; Performing a correlation analysis on the cost feature data and the cost-free feature data, and performing feature screening on the cost feature data and the cost-free feature data based on the correlation analysis result and the candidate evaluation result to obtain first target data and second target data; The first candidate model and the second candidate model are trained according to the first target data and the second target data to obtain a first prediction model and a second prediction model, wherein the first prediction model is obtained by training based on the first target data and the second target data, and the second prediction model is obtained by training based on the second target data.

2. The method for predicting insulin resistance according to claim 1, wherein: The correlation analysis result includes a continuous feature analysis result and a discrete feature analysis result; performing correlation analysis on the cost feature data and the cost-free feature data, and performing feature screening on the cost feature data and the cost-free feature data based on the correlation analysis result and the candidate evaluation result to obtain the first target data and the second target data, including: Determining continuous feature data and discrete feature data in the cost feature data and the cost-free feature data; Determining a linear correlation coefficient between each continuous feature in the continuous feature data, and performing a correlation analysis on the continuous feature data based on the linear correlation coefficient to obtain a continuous feature analysis result; Performing a chi-square test on each discrete feature in the discrete feature data to obtain a discrete feature analysis result; Performing correlation dimension reduction on the cost feature data and the cost-free feature data based on the continuous feature analysis result and the discrete feature analysis result; Performing importance evaluation on each feature in the cost feature data and the cost-free feature data according to the correlation dimension reduction result and the candidate evaluation result; Based on the importance evaluation result, the cost feature data and the cost-free feature data are feature screened to obtain the first target data and the second target data.

3. The method for predicting insulin resistance according to claim 2, wherein: The classifying the users to be predicted, determining the first type of users and the second type of users among the users to be predicted, and acquiring the first type of user data and the second type of user data, includes: Classifying the users to be predicted, and determining a first type of users and a second type of users among the users to be predicted; Generate a first type of questionnaire text based on the first target data and the second target data, and generate a second type of questionnaire text based on the second target data; Sending the first type of questionnaire text to the first type of users, and sending the second type of questionnaire text to the second type of users; Acquire a first type of answer text fed back by the first type of users based on the first type of questionnaire text and a second type of answer text fed back by the second type of users based on the second type of questionnaire text; Performing semantic analysis on the first type of answer text to obtain first type of user data; Perform semantic analysis on the second type of answer text to obtain second type of user data.

4. The method for predicting insulin resistance according to any one of claims 1 to 3, characterized in that: The obtaining of the original data set comprises: Collect user questionnaire data and user medical data; Aggregating the user questionnaire data and the user medical data based on user identities of the user questionnaire data and the user medical data to obtain an initial multidimensional data set; Performing data cleaning on the initial multidimensional data set according to the user identity identifier to obtain a candidate multidimensional data set; Performing dependency analysis on the candidate multidimensional data set to obtain dependency relationship information between features in the candidate multidimensional data set; Filling missing data in the candidate multidimensional data set based on the dependency information; The candidate multidimensional datasets filled with missing data are standardized to obtain the original dataset.

5. The method for predicting insulin resistance according to any one of claims 1 to 3, characterized in that: After inputting the first type of user data into the first prediction model to predict insulin resistance, and inputting the second type of user data into the second prediction model to predict insulin resistance, and obtaining the prediction results of the first type of users and the prediction results of the second type of users, the method comprises: Based on the prediction results of the first type of users and the prediction results of the second type of users, the first type of users and the second type of users are classified by risk to determine high-risk users and low-risk users; Performing semantic analysis on the answer texts of the high-risk user and the low-risk user to obtain user answer information; Generate a resistance question set according to a pre-constructed insulin resistance knowledge graph, wherein the insulin resistance knowledge graph includes medical knowledge information, symptom information, risk factor information, and diagnostic criteria information of insulin resistance; Determining the degree of text matching between the user answer information of the high-risk user and the low-risk user and the resistance question set based on a deep interactive text matching model; Insulin resistance is determined for the high-risk user and the low-risk user according to the text matching degree.

6. An insulin resistance prediction system, characterized in that: The insulin resistance prediction system comprises: A data acquisition module is used to acquire an original data set, wherein the original data set includes sample data and label data corresponding to the sample data, wherein the sample data includes user medical data and user questionnaire data, wherein the user questionnaire data includes lifestyle data, psychological state data, eating habit data, exercise frequency data, smoking and drinking history data, sleep quality data, and mood fluctuation data; A feature classification module is used to perform feature classification on the original data set to obtain cost feature data and cost-free feature data, wherein the cost feature data is feature data that is difficult to obtain in the original data set, and the cost-free feature data is feature data that is easy to obtain in the original data set; A model training module, used for training a pre-built original model based on the cost feature data and the cost-free feature data to obtain a first prediction model and a second prediction model, wherein the original model includes an integrated tree model, the first prediction model is obtained by training based on the cost feature data and the cost-free feature data, and the second prediction model is obtained by training based on the cost-free feature data; A user classification module, used to classify the users to be predicted, determine the first type of users and the second type of users among the users to be predicted, and obtain the first type of user data and the second type of user data, wherein the first type of user data includes the user questionnaire data and user medical data of the first type of users, and the second type of user data includes the user questionnaire data of the second type of users; An insulin resistance prediction module, used for inputting the first type of user data into the first prediction model to predict insulin resistance, and inputting the second type of user data into the second prediction model to predict insulin resistance, to obtain prediction results for the first type of users and prediction results for the second type of users; The model training module is also used to train multiple pre-constructed original models based on the cost feature data and the cost-free feature data to obtain multiple first candidate models and multiple second candidate models, wherein the first candidate models are obtained by training based on the cost feature data and the cost-free feature data, and the second candidate models are obtained by training based on the cost-free feature data; evaluate each first candidate model and each second candidate model to obtain a candidate evaluation result; perform correlation analysis on the cost feature data and the cost-free feature data, and perform feature screening on the cost feature data and the cost-free feature data based on the correlation analysis result and the candidate evaluation result to obtain first target data and second target data; train the first candidate model and the second candidate model according to the first target data and the second target data to obtain a first prediction model and a second prediction model, wherein the first prediction model is obtained by training based on the first target data and the second target data, and the second prediction model is obtained by training based on the second target data.

7. An insulin resistance prediction device, characterized in that: The insulin resistance prediction device comprises: a memory, a processor, and an insulin resistance prediction program stored in the memory and executable on the processor, wherein the insulin resistance prediction program is configured to implement the insulin resistance prediction method according to any one of claims 1 to 5.

8. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores an insulin resistance prediction program, and when the insulin resistance prediction program is executed by a processor, the insulin resistance prediction method according to any one of claims 1 to 5 is implemented.

9. A computer program product, characterized in that The computer program product comprises an insulin resistance prediction program, which implements the steps of the insulin resistance prediction method according to any one of claims 1 to 5 when executed by a processor.

Citation Information

Patent Citations

  • Construction system, method and application of insulin resistance prediction model

    CN115188493A

  • Machine learning-based early warning system for occurrence of acute kidney injury of critically ill patient

    CN118299054A

  • Diabetes risk prediction model based on multi-mode large model

    CN119274810A