Health risk detection system, method and computer program product

CN122552101APending Publication Date: 2026-08-11WEST CHINA HOSPITAL SICHUAN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-18
Publication Date
2026-08-11

AI Technical Summary

Technical Problem

但是,基于现有方法,在检测某些健康风险时,检测过程相对繁琐、复杂,检测成本相对较高,还很容易出现误差

Benefits of technology

[0015]Based on the health risk detection system, method, and computer program products provided in this manual, before specific implementation, a corresponding first health risk prediction model can be trained according to the preset joint training rules, based on the generalized linear mixed effects model and the Cox risk proportion model; at the same time, a corresponding second health risk prediction model can be trained based on the mixed model of the latent class mixed effects model and the logistic regression model. In practice, the following steps can be taken: First, blood biomarker parameters and feature parameters of various potential related factors at multiple time points can be collected from the target object. Then, composite biomarker parameters based on blood biomarkers can be obtained. Next, based on the blood biomarker parameters at multiple time points and the composite biomarker parameters, the change trajectory of the target object's key physical examination characteristics can be constructed. Then, based on the blood biomarker parameters at multiple time points, the blood biomarker parameters at the target object's target time point can be extracted. Finally, based on the feature parameters of various potential related factors and the change trajectory of the key physical examination characteristics, joint data of the target object can be constructed. A first health risk prediction model is used to process the blood biomarker parameters at the target object's target time point to obtain a first prediction result. A second health risk prediction model is used to process the joint data of the target object to obtain a second prediction result. Based on the first and second prediction results, it can be determined whether the target object has a health risk. This approach fully utilizes the target object's existing physical examination data, eliminating the need for complex and cumbersome testing, and enabling efficient and accurate prediction of the target object's health risk at a lower cost.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122552101A_ABST
    Figure CN122552101A_ABST
Patent Text Reader

Abstract

This specification provides a health risk detection system, method, and computer program product. Based on this method, before implementation, a first health risk prediction model can be trained according to preset joint training rules, using a generalized linear mixed-effects model combined with a Cox proportional hazards model; simultaneously, a second health risk prediction model can be trained using a hybrid model based on a latent class mixed-effects model and a logistic regression model. In implementation, blood biomarker parameters and characteristic parameters of multiple potential related factors at multiple time points can be collected from the target object; composite biomarker parameters based on blood biomarkers can be obtained; then, by utilizing the relatively easily obtainable data, the first and second health risk prediction models can be used in conjunction for prediction, thereby efficiently and accurately determining whether the target object has a health risk.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This manual belongs to the field of health big data technology, and in particular relates to health risk detection systems, methods and computer program products. Background Technology

[0002] As living standards improve, people are paying more and more attention to their health and hoping to understand their health risks in advance. However, based on existing methods, the detection process for certain health risks is relatively cumbersome and complex, the detection cost is relatively high, and errors are easily made.

[0003] There is currently no effective solution to the above problems. Summary of the Invention

[0004] This specification provides a health risk detection system, method, and computer program product that can predict the health risk status of target subjects efficiently and accurately at a low cost.

[0005] This manual provides a health risk detection system, which includes at least: a physical examination information collection module, a related information collection module, and a data processing module; wherein, The physical examination information collection module is used to collect blood biomarker parameters of the target object at multiple time points; and to obtain composite biomarker parameters based on blood biomarkers; The association information acquisition module is used to collect feature parameters of multiple potential association factors of the target object; The data processing module is used to construct the change trajectory of key physical examination characteristics of the target object based on blood biomarker parameters at multiple time points and composite biomarker parameters based on blood biomarkers; and to extract the blood biomarker parameters of the target object at a target time point based on the blood biomarker parameters of the target object at multiple time points; to construct joint data of the target object based on the feature parameters of multiple potential related factors of the target object and the change trajectory of the key physical examination characteristics; to obtain a first prediction result by processing the blood biomarker parameters of the target object at the target time point using a first health risk prediction model; and to obtain a second prediction result by processing the joint data of the target object using a second health risk prediction model; and to determine whether the target object has a health risk based on the first prediction result and the second prediction result.

[0006] In one embodiment, the first health risk prediction model is a model pre-trained based on a generalized linear mixed-effects model and a Cox risk proportion model according to a preset joint training rule. The second health risk prediction model is a model trained in advance according to a preset joint training rule, based on a hybrid model of latent class mixed effects model and logistic regression model.

[0007] In one embodiment, the multiple potential association factors include: basic attribute association factors, lifestyle association factors, metabolic state association factors, disease infection association factors, and disease genetic association factors; The basic attribute correlation factors include at least: age, gender, and body mass index; the lifestyle correlation factors include at least: smoking habits and drinking habits; the metabolic status correlation factors include at least: hypertension prevalence and diabetes prevalence; the infection, disease, and test correlation factors include at least: HBV infection status, fatty liver prevalence, and liver function indicators; and the family history correlation factors include at least: the current disease status of related relatives and the historical disease records of related relatives.

[0008] In one embodiment, the composite biomarker parameters include at least: basophil to lymphocyte ratio, eosinophil to lymphocyte ratio, monocyte to lymphocyte ratio, neutrophil to lymphocyte ratio, and platelet to lymphocyte ratio.

[0009] In one embodiment, determining whether the target object has a health risk based on the first prediction result and the second prediction result includes: Based on the preset model weight ratio coefficient mapping relationship, the target weight ratio coefficients for the first prediction result and the second prediction result are determined; Using a preset dynamic adjustment model, based on the change trajectory of the key physical examination characteristics of the target object, the blood biomarker parameters at the target time point, and the characteristic parameters of multiple potential related factors, the first degree of fit for the first health risk prediction model and the second degree of fit for the second health risk prediction model are determined. Using a preset dynamic adjustment model, the target weight ratio coefficient is adjusted based on the first and second fit degrees to obtain the adjusted target weight ratio coefficient. Based on the first prediction result, the second prediction result, and the adjusted target weight ratio coefficient, the corresponding target prediction result is determined. Based on the target prediction results, determine whether the target object has any health risks.

[0010] In one embodiment, the health risk detection system further includes a training module; wherein the training module is used for: Obtain the physical examination records, medical records, and investigation records of various potential related factors of the sample subjects; Based on the physical examination records, the sample subjects with multiple physical examinations and the sample subjects with a single physical examination were identified as the first type of sample subjects and the second type of sample subjects, respectively. A first sample dataset is constructed based on the physical examination records and medical records of the first type of sample subjects, as well as the investigation records of multiple potential related factors; a second sample dataset is constructed based on the physical examination records and medical records of the second type of sample subjects, as well as the investigation records of multiple potential related factors. Construct a first initial prediction model based on a generalized linear mixed effects model, and a second initial prediction model based on a mixture of latent class mixed effects model and logistic regression model; According to the preset joint training rules, the first sample dataset and the second sample dataset are used together, combined with the Cox risk ratio model, to train the first initial prediction model and obtain a first health risk prediction model that meets the requirements. According to the preset joint training rules, the second initial prediction model is trained using the first sample dataset to obtain a second health risk prediction model that meets the requirements. Based on the second sample dataset, a random test dataset is constructed; and using the random test dataset, the first health risk prediction model and the second health risk prediction model are randomly jointly tested to correct the first health risk prediction model and the second health risk prediction model, and a preset model weight ratio coefficient mapping relationship based on the first health risk prediction model and the second health risk prediction model is determined.

[0011] In one embodiment, the step of training the first initial prediction model by jointly utilizing the first sample dataset and the second sample dataset, combined with the Cox risk ratio model, according to a preset joint training rule, to obtain a first health risk prediction model that meets the requirements includes: By combining the first and second sample datasets, a third sample dataset is constructed. Based on the third sample dataset, a first training set and multiple first sub-sample datasets are constructed; wherein, the first training set does not contain survey records of multiple potential correlation factors; the first sub-sample datasets do not contain sample physical examination records, and each of the first sub-sample datasets corresponds to one type of potential correlation factor; The first initial prediction model is trained using the first training set to obtain the first intermediate prediction model that meets the requirements; the third initial prediction model based on the Cox risk ratio model is trained using the first training set to obtain the corresponding auxiliary model. According to the preset training order, the first intermediate prediction model is iteratively trained in multiple rounds using multiple first subsample datasets to obtain the corresponding first target prediction model. The first target prediction model is validated and corrected using an auxiliary model to obtain a first health risk prediction model that meets the requirements.

[0012] In one embodiment, training the second initial prediction model using the first sample dataset according to a preset joint training rule to obtain a second health risk prediction model that meets the requirements includes: Based on the second sample dataset, a second training set and multiple second sub-sample datasets are constructed; wherein, the second training set does not contain survey records of multiple potential correlation factors; the second sub-sample datasets do not contain sample physical examination records, and each of the second sub-sample datasets corresponds to one type of potential correlation factor; The second initial prediction model is trained using the second training set to obtain the second intermediate prediction model that meets the requirements. Following a preset training order, the second intermediate prediction model is iteratively trained multiple times using multiple second subsample datasets to obtain a second health risk prediction model that meets the requirements.

[0013] This manual also provides a method for health risk detection using a health risk detection system, including: Collect blood biomarker parameters of the target object at multiple time points, as well as feature parameters of multiple potential related factors of the target object; and obtain composite biomarker parameters based on blood biomarkers; Based on the blood biomarker parameters of the target object at multiple time points, and the composite biomarker parameters based on the blood biomarkers, the change trajectory of the key physical examination characteristics of the target object is constructed; and based on the blood biomarker parameters of the target object at multiple time points, the blood biomarker parameters of the target object at the target time point are extracted. Based on the characteristic parameters of multiple potential related factors of the target object and the change trajectory of the key physical examination characteristics, joint data of the target object is constructed. A first prediction result is obtained by processing the blood biomarker parameters of the target object at a target time point using a first health risk prediction model; and a second prediction result is obtained by processing the joint data of the target object using a second health risk prediction model. Based on the first prediction result and the second prediction result, it is determined whether the target object has a health risk.

[0014] This specification also provides a computer program product comprising a computer program that, when executed by a processor, implements the steps of the health risk detection method.

[0015] Based on the health risk detection system, method, and computer program products provided in this manual, before specific implementation, a corresponding first health risk prediction model can be trained according to the preset joint training rules, based on the generalized linear mixed effects model and the Cox risk proportion model; at the same time, a corresponding second health risk prediction model can be trained based on the mixed model of the latent class mixed effects model and the logistic regression model. In practice, the following steps can be taken: First, blood biomarker parameters and feature parameters of various potential related factors at multiple time points can be collected from the target object. Then, composite biomarker parameters based on blood biomarkers can be obtained. Next, based on the blood biomarker parameters at multiple time points and the composite biomarker parameters, the change trajectory of the target object's key physical examination characteristics can be constructed. Then, based on the blood biomarker parameters at multiple time points, the blood biomarker parameters at the target object's target time point can be extracted. Finally, based on the feature parameters of various potential related factors and the change trajectory of the key physical examination characteristics, joint data of the target object can be constructed. A first health risk prediction model is used to process the blood biomarker parameters at the target object's target time point to obtain a first prediction result. A second health risk prediction model is used to process the joint data of the target object to obtain a second prediction result. Based on the first and second prediction results, it can be determined whether the target object has a health risk. This approach fully utilizes the target object's existing physical examination data, eliminating the need for complex and cumbersome testing, and enabling efficient and accurate prediction of the target object's health risk at a lower cost. Attached Figure Description

[0016] To more clearly illustrate the embodiments of this specification, the accompanying drawings used in the embodiments will be briefly introduced below. The drawings described below are only some embodiments recorded in this specification. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0017] Figure 1 This is a schematic diagram of the structural composition of a health risk detection system provided in one embodiment of this specification; Figure 2 This is a schematic diagram illustrating one embodiment of the health risk detection system provided in this specification, applied in a scenario example. Figure 3 This is a schematic diagram illustrating one embodiment of the health risk detection system provided in this specification, applied in a scenario example. Figure 4 This is a schematic diagram illustrating one embodiment of the health risk detection system provided in this specification, applied in a scenario example. Figure 5This is a schematic diagram illustrating one embodiment of the health risk detection system provided in this specification, applied in a scenario example. Figure 6 This is a schematic flowchart of a health risk detection method provided in one embodiment of this specification; Figure 7 This is a schematic diagram of the structural composition of an electronic device provided in one embodiment of this specification; Figure 8 This is a schematic diagram illustrating one embodiment of the health risk detection system provided in the embodiments of this specification, applied in a scenario example. Detailed Implementation

[0018] To enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this specification, and not all embodiments. Based on the embodiments in this specification, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of this specification.

[0019] It should be noted that the information and data related to users involved in the embodiments of this specification are all information and data authorized by the user or fully authorized by the relevant parties. Furthermore, the collection, storage, use, processing, transmission, provision, disclosure, and application of the relevant data all comply with relevant laws, regulations, and standards, and necessary confidentiality measures have been taken. They do not violate public order and good morals, and corresponding operation entry points are provided for users or relevant parties to choose to authorize or refuse.

[0020] It should also be noted that in the embodiments of this specification, certain software, components, models and other existing solutions in the industry may be mentioned. These should be regarded as exemplary and are only intended to illustrate the feasibility of implementing the technical solution of this application. However, it does not mean that the applicant has used or necessarily used the solution.

[0021] See Figure 1 As shown in the embodiments of this specification, a health risk detection system is provided. This health risk detection system may include at least: a physical examination information collection module, a related information collection module, and a data processing module; wherein, The physical examination information collection module can specifically be used to collect blood biomarker parameters of the target object at multiple time points; and obtain composite biomarker parameters based on blood biomarkers; The associated information acquisition module can be specifically used to collect feature parameters of multiple potential associated factors of the target object; The data processing module can specifically be used to construct the change trajectory of key physical examination characteristics of the target object based on blood biomarker parameters at multiple time points and composite biomarker parameters based on blood biomarkers; and extract the blood biomarker parameters of the target object at a target time point based on the blood biomarker parameters of the target object at multiple time points; construct joint data of the target object based on the feature parameters of multiple potential related factors of the target object and the change trajectory of the key physical examination characteristics; obtain a first prediction result by processing the blood biomarker parameters of the target object at the target time point using a first health risk prediction model; obtain a second prediction result by processing the joint data of the target object using a second health risk prediction model; and determine whether the target object has a health risk based on the first prediction result and the second prediction result.

[0022] The aforementioned target objects can be understood as biological objects whose health risks are to be predicted. Specifically, these target objects can be human users or animals, such as monkeys and chimpanzees.

[0023] The aforementioned health risks may specifically include risks of abnormal states, such as the risk of sub-health, or risks of diseases, such as the risk of cancer.

[0024] The aforementioned health risk screening system can be deployed in a health checkup center or connected to the center's database. Accordingly, in practice, the target individuals can periodically visit the center for screenings every year or six months. During the screening, data such as blood test results can be collected and stored in the health checkup database.

[0025] When it is necessary to detect the current or future health risks of a target object, the application health risk detection system can be triggered to perform the corresponding health risk detection on the target object.

[0026] The aforementioned blood biomarker parameters can reflect the inflammatory and immune status of the target subject based on the dimensions of routine blood indicators.

[0027] The aforementioned blood biomarker parameters may include at least: key characteristic parameters of red blood cells, key characteristic parameters of white blood cells, etc.

[0028] The aforementioned composite biomarker parameters can be understood as characteristic parameters obtained by further processing blood biomarker parameters, which can reflect the inflammatory and immune status of the target body through more indicator dimensions beyond routine blood indicators.

[0029] The aforementioned composite biomarker parameters may include at least the following: basophil to lymphocyte ratio (BLR), eosinophil to lymphocyte ratio (ELR), monocyte to lymphocyte ratio (MLR), neutrophil to lymphocyte ratio (NLR), and platelet to lymphocyte ratio (PLR).

[0030] In practice, the aforementioned physical examination information collection module can obtain and extract the required blood biomarker parameters for multiple time points of the target object by acquiring and collecting blood routine test data collected during physical examinations at multiple time points. Furthermore, based on preset preprocessing rules, the module can calculate the composite biomarker parameters for multiple time points corresponding to the target object.

[0031] The aforementioned potential correlation factors can be specifically understood as factors that, through big data research and analysis combined with relevant experiments, are found to have a high probability of affecting the health status of the target object.

[0032] The aforementioned potential associated factors may include at least: basic attribute associated factors, lifestyle associated factors, metabolic status associated factors, disease infection associated factors, and disease genetic associated factors.

[0033] Furthermore, the basic attribute correlation factors include at least: age, gender, and body mass index; the lifestyle correlation factors include at least: smoking habits and drinking habits; the metabolic status correlation factors include at least: hypertension prevalence characteristics and diabetes prevalence characteristics; the infection, disease, and test correlation factors include at least: HBV infection status, fatty liver prevalence characteristics, and liver function indicators; and the family history correlation factors include at least: the current disease status of related relatives and the historical disease records of related relatives.

[0034] HBV specifically refers to the hepatitis B virus. The aforementioned liver function parameters include at least ALT (alanine aminotransferase) and AST (aspartate aminotransferase). The aforementioned related relatives specifically refer to the target individual's immediate family members.

[0035] Of course, it should be noted that the potential related factors listed above are only illustrative. In actual implementation, depending on the specific circumstances and processing needs, the above potential related factors may include other suitable factors. This specification does not limit this.

[0036] Specifically, the aforementioned health risk detection system can also be connected to a user database. This user database stores at least basic information about the target individuals, such as their age, gender, and medical history.

[0037] Accordingly, in specific implementation, the aforementioned related information collection module can obtain the first type of user data of the target object by querying the user database; extract the second type of user data of the target object by acquiring and based on other physical examination data collected from the target object at multiple time points during physical examinations; identify and extract the third type of user data of the target object by conducting a questionnaire survey on the target object and based on the collected questionnaires; and then integrate and use the first, second, and third type of user data, and through heterogeneous data integration and cross-validation, extract high-quality and comprehensive feature parameters of multiple potential related factors.

[0038] The change trajectory of the key physical examination characteristics of the target object can specifically include the change trajectory curves of the target object's blood biomarker parameters and composite biomarker parameters.

[0039] The aforementioned target time points include at least the time point most recent to the current time point. Furthermore, the aforementioned target time points may also include time points representative of health risks, such as the time point when the target subject undergoes a physical examination after reaching age 35; or, for example, the time point when the target subject undergoes a targeted physical examination at a hospital due to illness.

[0040] The aforementioned first health risk prediction model can be understood as a pre-trained algorithm model that can automatically analyze and predict the health risk of a user based on the relationship between the user's instantaneous key features (e.g., blood markers, composite markers), potential related factors, and health risks (e.g., cancer risk) at discrete time points.

[0041] The aforementioned second health risk prediction model can be understood as a pre-trained algorithm model that can automatically analyze and predict the health risk of a user based on the continuous changing trends of key characteristics (e.g., blood markers, composite markers) and potential related factors over a long period of time and the correlation with health risks (e.g., cancer risk).

[0042] In specific implementation, the aforementioned data processing module can first construct the change trajectory of the target object's key physical examination characteristics based on the blood biomarker parameters at multiple time points of the target object, as well as the composite biomarker parameters based on the blood biomarkers; and extract the blood biomarker parameters of the target object at the target time point based on the blood biomarker parameters at multiple time points of the target object; construct joint data of the target object based on the feature parameters of multiple potential related factors of the target object and the change trajectory of the key physical examination characteristics; use a first health risk prediction model to process the blood biomarker parameters of the target object at the target time point to obtain a first prediction result; and use a second health risk prediction model to process the joint data of the target object to obtain a second prediction result; finally, determine whether the target object has a health risk based on the first prediction result and the second prediction result.

[0043] In practice, the change trajectory of the key physical examination characteristics of the target object can be constructed by data fitting based on the blood biomarker parameters of the target object at multiple time points and the composite biomarker parameters based on the blood biomarkers. Then, the intercept and / or the slope of some key points of the change trajectory of the key physical examination characteristics can be adjusted according to the feature parameters of multiple potential related factors of the target object to construct the joint data of the target object.

[0044] When specifically detecting health risks, one can first determine the matching target weight ratio coefficient; then, based on the target weight ratio coefficient, use the first prediction result and the second prediction result together, with the second prediction result as the main factor and the first prediction result as the supplement, to more accurately determine whether the target subject has health risks, such as whether there is a risk of developing cancer.

[0045] Based on the above embodiments, the health risk detection system introduces and combines the first health risk prediction model and the second health risk prediction model. By effectively utilizing multiple dimensions of data information, such as blood biomarker parameters, composite biomarker parameters, and characteristic parameters of multiple potential related factors, and based on both instantaneous characteristics and long-term trends, it can comprehensively and accurately detect and judge the health risks of the target object.

[0046] In some embodiments, the first health risk prediction model can be a model trained in advance according to a preset joint training rule, based on a generalized linear mixed effects model and combined with a Cox risk proportion model. The second health risk prediction model can be a model trained in advance according to a preset joint training rule, based on a hybrid model of latent class mixed effects model and logistic regression model.

[0047] The Generalized Linear Mixed Model (GLMM) mentioned above can specifically be a type of hierarchical generalized linear model, which is an algorithmic model based on statistical distribution. By introducing and using the GLMM, more attention can be paid to binary classification (classification of two categories) involving complex multivariates, and data patterns for binary classification in complex data can be more effectively mined and utilized.

[0048] The aforementioned Cox Proportional Hazards Model can be specifically described as an algorithmic model that quantifies the risk of different factors affecting the occurrence of a endpoint event (e.g., death from disease) in research subjects simply by focusing on the relationship between independent variables and risk rates.

[0049] The aforementioned Latent Class Mixed Model (LCMM) specifically refers to an algorithmic model that simultaneously considers fixed effects (systematic influencing factors) and random effects (individual random variations) to classify potential categories. By introducing and using the Latent Class Mixed Model, it is more suitable for processing trajectory-type data and can also effectively take into account the data patterns of both binary and multi-class classification.

[0050] The aforementioned logistic regression model can be understood as an algorithmic model used to solve binary classification problems (which can also be extended to multi-class classification). It achieves probability prediction and classification judgment of sample categories by fitting the probability relationship between independent and dependent variables into an S-shaped logistic function.

[0051] Based on the above embodiments, by comprehensively utilizing the advantages and characteristics of different models, a first health risk prediction model with better performance and suitable for handling single-point instantaneous features, and a second health risk prediction model suitable for handling long-term changing trends can be constructed respectively.

[0052] In some embodiments, the multiple potential association factors may specifically include: basic attribute association factors, lifestyle association factors, metabolic state association factors, disease infection association factors, disease genetic association factors, etc. The basic attribute correlation factors include at least: age, gender, and body mass index; the lifestyle correlation factors include at least: smoking habits and drinking habits; the metabolic status correlation factors include at least: hypertension prevalence and diabetes prevalence; the infection, disease, and test correlation factors include at least: HBV infection status, fatty liver prevalence, and liver function indicators; and the family history correlation factors include at least: the current disease status of related relatives and the historical disease records of related relatives.

[0053] It should be noted that the aforementioned potential correlation factors were chosen because they can directly or indirectly affect or reflect the health status of the subjects. Therefore, by introducing and using the characteristic parameters of these potential correlation factors, we can indirectly assist in detecting the health risks of the subjects, and also assist in cross-validation, thereby making the prediction of the subjects' health risks more accurate.

[0054] In some embodiments, the blood biomarker parameters may include at least: key characteristic parameters of red blood cells, key characteristic parameters of white blood cells, etc.

[0055] The key characteristic parameters of red blood cells mentioned above may specifically include at least one of the following: red blood cell count (RBC), red blood cell distribution width SD (RDW-SD), red blood cell distribution width CV (RDW-CV), hematocrit (HCT), mean corpuscular hemoglobin concentration (MCHC), mean corpuscular hemoglobin content (MCH), mean corpuscular volume (MCV), hemoglobin concentration (Hb), etc.

[0056] The aforementioned red blood cell distribution width (SD) specifically refers to the standard deviation of red blood cell distribution width, a parameter reflecting the heterogeneity of red blood cell volume and used to reflect the uniformity of red blood cell size and shape. The aforementioned red blood cell distribution width (CV) specifically refers to the coefficient of variation of red blood cell distribution width, used to reflect the dispersion of red blood cell size.

[0057] Furthermore, the key characteristic parameters of red blood cells mentioned above may also include: erythrocyte sedimentation rate (ESR), erythrocyte sedimentation rate equation K value (ESR-K), etc.

[0058] The key characteristic parameters of the aforementioned white blood cells may include at least one of the following: key characteristics of granulocytes (e.g., neutrophils, eosinophils, basophils, etc.), key characteristics of lymphocytes (e.g., B lymphocytes, T lymphocytes, natural killer cells, etc.), key characteristics of monocytes (e.g., macrophages, dendritic cells, etc.), etc.

[0059] The key characteristic parameters of white blood cells mentioned above may specifically include at least one of the following: white blood cell count (WBC), neutrophil count (NeuC), neutrophil percentage (Neu%), lymphocyte count (LymC), lymphocyte percentage (Lym%), eosinophil count (EosC), eosinophil percentage (Eos%), basophil count (BasC), basophil percentage (Bas%), monocyte count (MonC), monocyte percentage (Mon%), and T lymphocyte CD3 count (CD3 count), T lymphocyte CD3 percentage (CD3%), T lymphocyte CD4 count (CD4 count), T lymphocyte CD4 percentage (CD4%), T lymphocyte CD8 count (CD8 count), T lymphocyte CD8 percentage (CD8%), and T lymphocyte CD4 / CD8 ratio.

[0060] The T lymphocytes mentioned above specifically refer to mature lymphocytes in the thymus, which are divided into helper T cells (CD4+) and cytotoxic T cells (CD8+). CD3 specifically refers to the CD3 protein marker commonly carried on the surface of mature T lymphocytes; its value reflects the body's cellular immune function. CD4 specifically refers to the CD4 protein marker carried on the surface of helper T cells. CD8 specifically refers to the CD8 protein marker carried on the surface of cytotoxic T cells.

[0061] Furthermore, the key characteristic parameters of the blood cells may also include: key characteristics of platelets and / or composite characteristics among multiple types of blood cells. Among these, the key characteristics of platelets may include at least: platelet count (PLT).

[0062] It should be noted that the key characteristic parameters of red blood cells mentioned above were chosen as the blood biomarkers to be acquired and used because of the correlation between systemic inflammation and changes in the subject's health. Therefore, by introducing and using the key characteristic parameters of red blood cells, changes in the subject's health can be analyzed and predicted based on the dimension of systemic inflammation correlation.

[0063] The key characteristic parameters of leukocytes mentioned above were chosen as the blood biomarkers to be acquired and used because of the correlation between immune status and changes in the subject's health, as well as the correlation between neuroendocrine interactions and changes in the subject's health. Therefore, by introducing and using the key characteristic parameters of leukocytes, changes in the subject's health can be analyzed and predicted based on the dimensions of correlation with immune status and correlation with neuroendocrine interactions.

[0064] In some embodiments, the composite biomarker parameters may include at least: basophil to lymphocyte ratio (BLR), eosinophil to lymphocyte ratio (ELR), monocyte to lymphocyte ratio (MLR), neutrophil to lymphocyte ratio (NLR), and platelet to lymphocyte ratio (PLR).

[0065] Specifically, the aforementioned composite biomarker parameters can be calculated by combining multiple different types of blood biomarker parameters, which can reflect the inflammatory and immune status of the subject's body from another dimension.

[0066] It should be noted that the aforementioned composite biomarker parameters were chosen because of the correlation between the body's inflammation and immune status and changes in the subject's health status. For example, relevant studies have found that health risks such as cancer are also associated with the subject's inflammation and immune status. Therefore, by introducing and using the aforementioned composite biomarker parameters, changes in the subject's health status can be analyzed and predicted based on the dimensions of the subject's inflammation and immune status.

[0067] In practice, the first and second health risk prediction models mentioned above can be used. By using the blood biomarker parameters and composite biomarker parameters as exposure variables (or main variables) and the characteristic parameters of multiple potential related factors as covariates (or auxiliary variables), the data feature information of the object from multiple different dimensions can be integrated. At the same time, based on instantaneous features and long-term continuous change trends, a comprehensive and detailed analysis can be conducted from multiple angles, with the change trajectory as the main focus and single-point features as a supplement. In this way, the health risk status of the object can be accurately predicted.

[0068] In some embodiments, see Figure 2 As shown, the above-mentioned determination of whether the target object has a health risk based on the first prediction result and the second prediction result may include the following in specific implementation: S2-1: Determine the target weight ratio coefficients for the first prediction result and the second prediction result based on the preset model weight ratio coefficient mapping relationship; S2-2: Using a preset dynamic adjustment model, based on the change trajectory of the key physical examination characteristics of the target object, the blood biomarker parameters at the target time point, and the characteristic parameters of multiple potential related factors, determine the first degree of fit for the first health risk prediction model and the second degree of fit for the second health risk prediction model. S2-3: Using a preset dynamic adjustment model, based on the first fit and the second fit, adjust the target weight ratio coefficient to obtain the adjusted target weight ratio coefficient. S2-4: Based on the first prediction result, the second prediction result, and the adjusted target weight ratio coefficient, determine the corresponding target prediction result; S2-5: Based on the target prediction results, determine whether the target object has any health risks.

[0069] The aforementioned preset model weight ratio coefficient mapping relationship is determined by conducting random test experiments on the first health risk prediction model and the second health risk prediction model using a large amount of sample data, and based on the results of the random test experiments, and considering the overall application effect of the model.

[0070] Specifically, for example, the aforementioned preset dynamic adjustment model is a pre-trained algorithm model based on a large language model. It is capable of targeted and refined adjustments based on the individual adaptation differences of the target object. By analyzing and determining the adaptation degree between the target object and the first health risk prediction model and the second health risk prediction model, and based on the weight ratio coefficient determined by the preset model weight ratio coefficient mapping relationship, it can make adjustments.

[0071] In practice, the target trajectory range to which the change trajectory of the key physical examination characteristics of the target object belongs can be determined first; at the same time, the target parameter range of the blood biomarker parameters at the target time point can be determined; then, according to the preset model weight ratio coefficient mapping relationship, the preset weight ratio coefficient corresponding to the combination of the above target trajectory range and target parameter range can be determined as the target weight ratio coefficient. The preset model weight ratio coefficient mapping relationship can include multiple preset weight ratio coefficients, and each preset weight ratio coefficient corresponds to at least one combination of trajectory range and parameter range.

[0072] In practice, a pre-set dynamic adjustment model can be used to analyze the data characteristics of the target object's individual data based on the change trajectory of key physical examination features, blood biomarker parameters at the target time point, and characteristic parameters of multiple potential related factors. Then, based on these individual data characteristics and the previously trained and mastered adaptation and coupling rules between individual data and the model, the first degree of fit of the target object's individual data to the first health risk prediction model and the second degree of fit to the second health risk prediction model can be determined. Furthermore, based on the first and second degree of fit, and combined with the specific characteristics of the target object's individual data, targeted dynamic adjustments can be made to the target weight ratio coefficient to ensure that the weight ratio coefficient matches the individual situation of the target object. This results in adjusted target weight ratio data that is more targeted and has a better fit for the target object.

[0073] In practice, the first weight coefficient for the first prediction result and the second weight coefficient for the second prediction result can be determined based on the adjusted target weight ratio coefficient. Then, the corresponding target prediction result can be obtained by weighting the first weight coefficient, the second weight coefficient, the first prediction result, and the second prediction result. Finally, based on the target prediction result, it can be determined whether the target object has a health risk.

[0074] Specifically, based on the target prediction results, a health risk probability value can be determined; then, it can be checked whether this health risk probability value is greater than a preset risk probability threshold. If it is greater than the preset risk probability threshold, it is determined that the target object has a health risk; conversely, if it is less than or equal to the preset risk probability threshold, it is determined that the target object does not have a health risk.

[0075] Based on the above embodiments, by introducing and utilizing the preset model weight ratio coefficient mapping relationship and the preset dynamic adjustment model, it is possible to effectively and reasonably combine the first prediction result obtained based on the first health risk prediction model and the second prediction result obtained based on the second health risk prediction model to accurately determine whether the target object has a health risk.

[0076] In some embodiments, the health risk detection system may further include a training module; wherein, for specific real-time details, see [link to documentation]. Figure 3 As shown, the training module can specifically train a first health risk prediction model and a second health risk prediction model that meet the requirements according to preset joint training rules in the following manner: S3-1: Obtain the physical examination records, medical records, and investigation records of various potential related factors of the sample subjects; S3-2: Based on the physical examination records, the sample subjects with multiple physical examinations and the sample subjects with a single physical examination are respectively identified as the first type of sample subjects and the second type of sample subjects. S3-3: Construct the first sample dataset based on the physical examination records and medical records of the first type of sample objects, as well as the investigation records of multiple potential related factors; construct the second sample dataset based on the physical examination records and medical records of the second type of sample objects, as well as the investigation records of multiple potential related factors. S3-4: Construct a first initial prediction model based on a generalized linear mixed effects model, and a second initial prediction model based on a mixed model of latent class mixed effects model and logistic regression model; S3-5: According to the preset joint training rules, the first sample dataset and the second sample dataset are used together, combined with the Cox risk ratio model, to train the first initial prediction model and obtain a first health risk prediction model that meets the requirements. S3-6: According to the preset joint training rules, the second initial prediction model is trained using the first sample dataset to obtain a second health risk prediction model that meets the requirements. S3-7: Based on the second sample dataset, construct a random test dataset; and using the random test dataset, perform random joint testing on the first health risk prediction model and the second health risk prediction model to correct the first health risk prediction model and the second health risk prediction model, and determine the preset model weight ratio coefficient mapping relationship based on the first health risk prediction model and the second health risk prediction model.

[0077] In practice, after obtaining the physical examination records, medical records, and investigation records of various potential related factors of the sample subjects, the sample subjects who already had health risks before the first physical examination record can be screened out based on the medical records and designated as invalid sample subjects. The physical examination records, medical records, and investigation records of various potential related factors of the invalid sample subjects are then removed.

[0078] In addition, after obtaining the physical examination records, medical records, and investigation records of various potential related factors of the sample objects, the physical examination records of each sample object can be time-aligned according to the preset timeliness rules to avoid time errors when using the training model in the future.

[0079] In practice, based on the physical examination records of the sample subjects, blood biomarker parameters and composite biomarker parameters at multiple time points can be obtained as exposure variables; based on the sample cases of the sample subjects, the target disease outcome of the sample subjects can be obtained as outcome variables; based on the investigation records of multiple potential related factors of the sample subjects, the characteristic parameters of multiple potential related factors of the sample subjects can be obtained as covariates.

[0080] Specifically, the aforementioned target disease outcome may include disease indicator markers for multiple representative diseases.

[0081] Specifically, taking cancer risk as an example of health risks, the aforementioned representative diseases can include: lung cancer, liver cancer, stomach cancer, colorectal cancer, esophageal cancer, pancreatic cancer, gallbladder cancer, brain tumors and central nervous system tumors, leukemia, lymphoma, bladder cancer, kidney cancer, cervical cancer, endometrial cancer, ovarian cancer, breast cancer, testicular cancer, prostate cancer, melanoma of the skin, cancer of the lips, mouth, and pharynx, laryngeal cancer, nasopharyngeal cancer, thyroid cancer, etc. Among these, multiple types of cancer can comprehensively cover cancer risk detection, ensuring the generalization ability of the trained model.

[0082] Before implementation, if permitted, a large number of sample medical examination records and case studies can be collected. Based on the case studies, candidate subjects related to the targeted health risks are selected from the sample subjects. Based on the sample medical examination records of the candidate subjects, the change trajectories of key medical examination characteristics of the candidate subjects are constructed. Based on the change trajectories of key medical examination characteristics of the candidate subjects, multiple trajectory type data groups are determined through clustering learning and data statistics. Each trajectory type data group corresponds to at least one trajectory type and contains multiple candidate subjects corresponding to that trajectory type. For each trajectory type group, trajectory feature distribution statistics are performed based on the change trajectories of key medical examination characteristics of multiple candidate subjects in that trajectory type group. Based on the results of the trajectory feature distribution statistics, multiple candidate subjects that meet the distribution consistency requirements in that trajectory type group are selected as potential subjects. The trajectory feature differences between the change trajectories of key medical examination characteristics of potential subjects in the same trajectory type group are then calculated and statistically analyzed. Potential subjects with trajectory feature differences less than a preset difference threshold are removed from the trajectory type groups, resulting in the final reference subjects for each trajectory type group. By acquiring and analyzing sample cases from reference subjects, the disease outcomes of these subjects are determined. Based on these outcomes, representative diseases representing health risks are identified. Subsequently, based on these representative diseases, targeted physical examination records, medical records, and investigation records of various potential related factors can be collected from relevant sample subjects to train a first and a second health risk prediction model that are well-suited for health risk detection and have good generalization capabilities.

[0083] In practice, based on the physical examination records of the sample subjects, sample subjects who have only participated in one physical examination and have a physical examination record at a single time point can be selected as the first type of sample subjects; at the same time, sample subjects who have participated in at least two physical examinations and have physical examination records at at least two time points can be selected as the second type of sample subjects.

[0084] In practical implementation, a second initial prediction model can be constructed, which is a hybrid model based on a latent class mixed effects model and a logistic regression model. This second initial prediction model includes at least a trajectory processing structure and a regression classification structure, where the trajectory processing structure is a modular structure based on the latent class mixed effects model, and the regression classification structure is a modular structure based on the logistic regression model. Furthermore, a first initial prediction model based on a generalized linear mixed effects model and a third initial prediction model based on the Cox proportional hazards model can also be constructed.

[0085] In practice, the second initial prediction model can be trained separately according to the preset joint training rules to obtain a second health risk prediction model that meets the requirements; at the same time, the first initial prediction model can be trained separately to obtain a first health risk prediction model that meets the requirements; and then the first health risk prediction model and the second health risk prediction model can be randomly jointly tested to obtain the random joint test results.

[0086] Furthermore, based on the results of random joint testing, the two models can be used as a reference for each other to interactively refine the two models, resulting in a first health risk prediction model and a second health risk prediction model that offer relatively better performance and higher accuracy. Simultaneously, based on the results of random joint testing, a pre-defined model weight ratio coefficient mapping relationship can be determined between the first and second health risk prediction models.

[0087] Based on the above embodiments, the advantages and characteristics of different types of model structures can be fully utilized to efficiently train a first health risk prediction model and a second health risk prediction model that are adapted to health risk detection scenarios and have good performance and high accuracy.

[0088] In some embodiments, see Figure 4 As shown, the above-mentioned method, based on preset joint training rules, jointly utilizes the first sample dataset and the second sample dataset, combined with the Cox risk ratio model, to train the first initial prediction model, thereby obtaining a first health risk prediction model that meets the requirements. In specific implementation, this may include the following: S4-1: Combine the first and second sample datasets to construct the third sample dataset; S4-2: Based on the third sample dataset, a first training set and multiple first sub-sample datasets are constructed; wherein, the first training set does not contain survey records of multiple potential correlation factors; the first sub-sample datasets do not contain sample physical examination records, and each of the first sub-sample datasets corresponds to one type of potential correlation factor; S4-3: Use the first training set to train the first initial prediction model to obtain the first intermediate prediction model that meets the requirements; use the first training set to train the third initial prediction model based on the Cox risk ratio model to obtain the corresponding auxiliary model. S4-4: According to the preset training order, the first intermediate prediction model is iteratively trained multiple times using multiple first subsample datasets to obtain the corresponding first target prediction model. S4-5: Use an auxiliary model to verify and correct the first target prediction model to obtain a first health risk prediction model that meets the requirements.

[0089] Before implementation, sample medical records and investigation records of potential related factors can be extracted from the sample subjects' physical examination records, medical records, and investigation records of various potential related factors as auxiliary sample data. Then, based on the medical records, the auxiliary sample data is divided into a first type of auxiliary sample data and a second type of auxiliary sample data. The first type of auxiliary sample data identifies health risks based on the medical records, while the second type identifies no health risks based on the medical records. Based on the first type of auxiliary sample data, a first type of correlation matrix heatmap is constructed by calculating and analyzing the correlation between various potential related factors and existing health risks. Simultaneously, based on the second type of auxiliary sample data, a second type of correlation matrix heatmap is constructed by calculating and analyzing the correlation between various potential related factors and no health risks. By combining the first and second type of correlation matrix heatmaps, the degree of correlation between various potential related factors in predicting health risks can be determined. Based on this degree of correlation, a preset training sequence for multi-round iterative training can be determined.

[0090] In practice, taking the calculation of the correlation analysis results between various potential related factors and existing health risks based on the first type of auxiliary sample data as an example, we can first calculate the correlation coefficient between each potential related factor and existing health risks based on multiple preset correlation analysis rules and the first type of auxiliary sample data; then, based on the correlation coefficient between each potential related factor and existing health risks, we can statistically obtain the correlation analysis results between various potential related factors and existing health risks.

[0091] Specifically, the aforementioned pre-defined correlation analysis rules may include: correlation analysis rules based on Pearson correlation coefficient, correlation analysis rules based on Spearman correlation coefficient, correlation analysis rules based on Kendall correlation coefficient, etc. It should be noted that the pre-defined correlation analysis rules listed above are merely illustrative. In actual implementation, other types of correlation analysis rules may be included depending on the specific circumstances and processing requirements. This specification does not limit this.

[0092] Specifically, the above-preset training sequence is as follows: first, training based on basic attribute-related factors; then, iterative training based on lifestyle-related factors; next, iterative training based on metabolic state-related factors; then, iterative training based on disease infection-related factors; and finally, iterative training based on disease genetic-related factors.

[0093] In practice, after constructing the third sample dataset, the blood biomarker parameters and composite biomarker parameters of the sample objects at a single time point are extracted from the third sample dataset and combined with the disease results to construct the first training set. Based on the third sample dataset, the feature parameters of different types of potential correlation factors of the sample objects are extracted and combined with the disease results to construct multiple first subsample datasets corresponding to multiple types of potential correlation factors.

[0094] In specific implementation, the following steps can be taken: First, a first intermediate prediction model can be trained using the first subsample dataset corresponding to the factors associated with basic attributes. This will be the first round of iterative training to learn and master the indirect influence of the factors associated with basic attributes on health risk prediction, resulting in a first improved model. Next, the first improved model can be trained using the first subsample dataset corresponding to the factors associated with lifestyle. This will be the second round of iterative training to learn and master the indirect influence of lifestyle factors on health risk prediction, resulting in a second improved model. Then, the second improved model can be trained using the first subsample dataset corresponding to the factors associated with metabolic state. This will be the third round of iterative training to learn and master the indirect influence of metabolic state factors on health risk prediction, resulting in a third improved model. Then, the third improved model can be trained using the first subsample dataset corresponding to the factors associated with disease infection. This will be the fourth round of iterative training to learn and master the indirect influence of disease infection factors on health risk prediction, resulting in a fourth improved model. Finally, the fourth improved model can be trained using the first subsample dataset corresponding to the factors associated with disease genetics. This will be the fifth improved model, serving as the corresponding first target prediction model.

[0095] In practice, under the condition of the same characteristic parameters of potential related factors, an auxiliary model can be used as a reference. By conducting sensitivity analysis and robustness verification on the model of the first target prediction model, the performance of the first target prediction model can be verified and the corresponding verification results can be obtained. Based on the verification results, the relevant model parameters in the first target prediction model can be fine-tuned and corrected to finally obtain a first health risk prediction model that meets the requirements.

[0096] Based on the above embodiments, the first sample dataset and the second sample dataset can be fully utilized, and an auxiliary model can be combined to train a first health risk prediction model with good performance based on single-point instantaneous feature angle.

[0097] In some embodiments, see Figure 5As shown, the second initial prediction model is trained using the first sample dataset according to the preset joint training rules to obtain a second health risk prediction model that meets the requirements. In specific implementation, it may include the following: S5-1: Based on the second sample dataset, construct a second training set and multiple second sub-sample datasets; wherein, the second training set does not contain survey records of multiple potential correlation factors; the second sub-sample datasets do not contain sample physical examination records, and each of the second sub-sample datasets corresponds to one type of potential correlation factor; S5-2: Use the second training set to train the second initial prediction model to obtain the second intermediate prediction model that meets the requirements; S5-3: Following the preset training order, the second intermediate prediction model is iteratively trained multiple times using multiple second subsample datasets to obtain a second health risk prediction model that meets the requirements.

[0098] In practice, based on the second sample dataset, blood biomarker parameters and composite biomarker parameters of the sample objects at multiple time points can be extracted to construct the change trajectory of the corresponding key physical examination features. Then, the change trajectory of the key physical examination features can be combined with the disease results to construct the second training set. Based on the second sample dataset, feature parameters of different types of potential related factors of the sample objects can be extracted and combined with the disease results to construct multiple second subsample datasets corresponding to multiple types of potential related factors.

[0099] In practice, the regression classification structure in the second initial prediction model can be locked and kept unchanged. The trajectory processing structure in the second initial prediction model can be trained and adjusted using the second training set until the trajectory processing structure meets the preset requirements. Then the locking of the regression classification structure can be released. The trajectory processing structure and the regression classification structure can be trained and adjusted simultaneously using the second training set to obtain a second intermediate prediction model that meets the requirements.

[0100] In practice, the intercept and / or slope of the key physical examination feature change trajectory can be adjusted based on the feature parameters of the potential related factors in the second subsample dataset corresponding to the basic attribute related factors, to obtain the first adjusted second subsample dataset corresponding to the basic attribute related factors. Then, the second intermediate prediction model can be trained using the first adjusted second subsample dataset corresponding to the basic attribute related factors, and the first round of iterative training can be carried out to learn and master the lateral influence of the basic attribute related factors on health risk prediction, thus obtaining the model after one round of iterative improvement.

[0101] Then, based on the feature parameters of potential related factors in the second subsample dataset corresponding to lifestyle-related factors, the intercept and / or slope of the change trajectory of key physical examination features are adjusted to obtain the second adjusted second subsample dataset corresponding to lifestyle-related factors. Then, the improved model is trained in one round using the second adjusted second subsample dataset corresponding to lifestyle-related factors, and a second round of iterative training is conducted to learn and master the lateral influence of lifestyle-related factors on health risk prediction, resulting in the improved model after two rounds of iterative training.

[0102] Next, based on the feature parameters of potential related factors in the second subsample dataset corresponding to metabolic state-related factors, the intercept and / or slope of the change trajectory of key physical examination features are adjusted to obtain the third adjusted second subsample dataset corresponding to metabolic state-related factors. Then, the model improved by two rounds of iterations is trained using the third adjusted second subsample dataset corresponding to metabolic state-related factors, and a third round of iterations is conducted to learn and master the lateral influence of metabolic state-related factors on health risk prediction, resulting in the model improved by three rounds of iterations.

[0103] Then, based on the feature parameters of potential associated factors in the second subsample dataset corresponding to disease infection-related factors, the intercept and / or slope of the change trajectory of key physical examination features are adjusted to obtain the fourth adjusted second subsample dataset corresponding to disease infection-related factors. The model is then trained using the fourth adjusted second subsample dataset corresponding to disease infection-related factors for three rounds of iterative improvement, and a fourth round of iterative training is conducted to learn and master the lateral influence of disease infection-related factors on health risk prediction, resulting in the model after four rounds of iterative improvement.

[0104] Finally, based on the feature parameters of potential associated factors in the second subsample dataset corresponding to disease genetic factors, the intercept and / or slope of the change trajectory of key physical examination features are adjusted to obtain the fifth adjusted second subsample dataset corresponding to disease genetic factors. Then, the model is trained using the fifth adjusted second subsample dataset corresponding to disease genetic factors for four rounds of iterative improvement, and a fifth round of iterative training is conducted to learn and master the lateral influence of disease genetic factors on health risk prediction, resulting in a model with five rounds of iterative improvement, which serves as the second health risk prediction model that meets the requirements.

[0105] Based on the above embodiments, the second sample dataset can be fully utilized to train a second health risk prediction model with better performance in terms of long-term trajectory change trend.

[0106] In some embodiments, when implementing a specific task, multiple sample data can be randomly selected from the second sample dataset as the initial test dataset; then, based on the initial test dataset, the corresponding random test dataset can be obtained by sample expansion.

[0107] In practice, the first health risk prediction model and the second health risk prediction model can be used to process the test sample data randomly selected from the random test dataset at the same time, and the corresponding first processing result and second processing result can be obtained. Then, by calculating the deviation values ​​between the first processing result, the second processing result and the result variable in the test sample data, the deviation values ​​of the first health risk prediction model and the second health risk prediction model based on multiple test sample data can be obtained, which are used as the random joint test results.

[0108] In practice, based on the results of random joint testing and various deviation values, the first and second health risk prediction models can be used together as references to interactively correct the model parameters of both models. This multi-round iterative correction yields the corrected first and second health risk prediction models, which are then used as the final first and second health risk prediction models. This approach fully utilizes the advantages of each model, allowing for further optimization and adjustment to obtain a first and second health risk prediction model with relatively higher accuracy and better performance.

[0109] In specific implementation, based on the results of random joint testing, the change trajectory of corresponding key features, blood biomarker parameters at the target time point, and corresponding deviation values ​​can be extracted and combined to obtain multiple test result data pairs. Based on the test result data pairs, multiple clusters are obtained by performing clustering processing based on deviation value types. Each cluster corresponds to a deviation value type. Based on each cluster, the combination of trajectory range and parameter range corresponding to each deviation value type is determined. At the same time, based on the deviation value type, the corresponding preset model weight ratio coefficient is determined. Then, the preset model weight ratio coefficient corresponding to the same deviation value type is associated with the combination of trajectory range and parameter range to establish a preset model weight ratio coefficient mapping relationship.

[0110] In some embodiments, the aforementioned preset dynamic adjustment model can be constructed as follows: an initial adjustment model based on a large language model is constructed; a random test dataset and corresponding random test results are obtained and used as an initial sample training set; based on the survey records of multiple potential correlation factors of the second type of sample objects, feature parameters of corresponding multiple potential correlation factors are added to the initial sample training set to obtain a target sample training set; using the target sample training set, the model parameters of the initial adjustment model are continuously trained and adjusted to obtain a preset dynamic adjustment model that meets the requirements.

[0111] As can be seen from the above, based on the health risk detection system provided in the embodiments of this specification, before specific implementation, according to the preset joint training rules, a corresponding first health risk prediction model is trained based on the generalized linear mixed effects model and the Cox risk proportion model; at the same time, a corresponding second health risk prediction model is trained based on the mixed model of the latent class mixed effects model and the logistic regression model. In practice, the following steps can be taken: First, blood biomarker parameters and feature parameters of various potential related factors at multiple time points can be collected from the target object. Then, composite biomarker parameters based on blood biomarkers can be obtained. Next, based on the blood biomarker parameters at multiple time points and the composite biomarker parameters, the change trajectory of the target object's key physical examination characteristics can be constructed. Then, based on the blood biomarker parameters at multiple time points, the blood biomarker parameters at the target object's target time point can be extracted. Based on the feature parameters of various potential related factors and the change trajectory of the key physical examination characteristics, joint data of the target object can be constructed. A first health risk prediction model is used to process the blood biomarker parameters at the target object's target time point to obtain a first prediction result. A second health risk prediction model is used to process the joint data of the target object to obtain a second prediction result. Based on the first and second prediction results, it can be determined whether the target object has a health risk. This approach fully utilizes existing physical examination data and enables efficient and accurate prediction of the target object's health risk at a lower cost.

[0112] See Figure 6 As shown in the embodiments of this specification, a health risk detection method based on the above-described health risk detection system is also provided. Specifically, this method may include the following: S601: Collect blood biomarker parameters of the target object at multiple time points, as well as feature parameters of multiple potential related factors of the target object; and obtain composite biomarker parameters based on blood biomarkers; S602: Based on the blood biomarker parameters of the target object at multiple time points, and the composite biomarker parameters based on the blood biomarkers, construct the change trajectory of the key physical examination characteristics of the target object; and extract the blood biomarker parameters of the target object at the target time point based on the blood biomarker parameters of the target object at multiple time points; S603: Construct joint data of the target object based on the characteristic parameters of multiple potential related factors of the target object and the change trajectory of the key physical examination characteristics; S604: Using a first health risk prediction model, a first prediction result is obtained by processing the blood biomarker parameters of the target object at a target time point; and using a second health risk prediction model, a second prediction result is obtained by processing the joint data of the target object. S605: Based on the first prediction result and the second prediction result, determine whether the target object has a health risk.

[0113] In practice, step S601 can be achieved using the physical examination information collection module and the associated information collection module in the health risk detection system, and steps S602 to S605 can be achieved using the data processing module.

[0114] Based on the above embodiments, a health risk detection system can be used to predict the health risk status of target individuals at a lower cost and with higher efficiency and accuracy.

[0115] This specification provides an electronic device through its embodiments. (See attached document.) Figure 7 As shown. The electronic device includes a network communication port 701, a processor 702, and a memory 703. These structures are connected by internal cables so that they can perform specific data interaction.

[0116] Specifically, the network communication port 701 can be used to acquire blood biomarker parameters of the target object at multiple time points, as well as feature parameters of multiple potential related factors of the target object; and to acquire composite biomarker parameters based on blood biomarkers.

[0117] The processor 702 is specifically configured to: construct a trajectory of change in key physical examination characteristics of the target object based on blood biomarker parameters at multiple time points and composite biomarker parameters based on blood biomarkers; extract blood biomarker parameters at a target time point based on the blood biomarker parameters at multiple time points of the target object; construct joint data of the target object based on feature parameters of multiple potential related factors of the target object and the trajectory of change in the key physical examination characteristics; obtain a first prediction result by processing the blood biomarker parameters at the target time point of the target object using a first health risk prediction model; obtain a second prediction result by processing the joint data of the target object using a second health risk prediction model; and determine whether the target object has a health risk based on the first prediction result and the second prediction result.

[0118] The memory 703 can be used to store the corresponding instruction program and related intermediate data.

[0119] Based on the above methods, the relevant structural performance of electronic devices can be effectively utilized to improve the data processing speed of electronic devices and efficiently realize the data processing for health risk detection.

[0120] In this embodiment, the network communication port 701 can be a virtual port bound to different communication protocols, thereby enabling the sending or receiving of different data. For example, the network communication port can be a port responsible for web data communication, a port responsible for FTP data communication, or a port responsible for email data communication. Furthermore, the network communication port can also be a physical communication interface or communication chip. For example, it can be a wireless mobile network communication chip, such as GSM or CDMA; it can also be a Wi-Fi chip; or it can be a Bluetooth chip.

[0121] In this embodiment, the processor 702 can be implemented in any suitable manner. For example, the processor can take the form of a microprocessor or processor and a computer-readable medium storing computer-readable program code (e.g., software or firmware) executable by the (micro)processor, logic gates, switches, application-specific integrated circuits (ASICs), programmable logic controllers, and embedded microcontrollers, etc. This specification is not limiting.

[0122] In this embodiment, the memory 703 may include multiple layers. In a digital system, anything that can store binary data can be a memory. In an integrated circuit, a circuit with storage function but no physical form is also called a memory, such as RAM, FIFO, etc. In a system, a storage device with a physical form is also called a memory, such as a memory stick, TF card, etc.

[0123] This specification also provides a computer-readable storage medium based on the above-described health risk detection method. The computer-readable storage medium stores computer program instructions that, when executed, implement the following: collecting blood biomarker parameters of a target object at multiple time points, and feature parameters of multiple potential related factors of the target object; obtaining composite biomarker parameters based on blood biomarkers; constructing a change trajectory of key physical examination characteristics of the target object based on the blood biomarker parameters of the target object at multiple time points and the composite biomarker parameters based on blood biomarkers; extracting blood biomarker parameters of the target object at a target time point based on the blood biomarker parameters of the target object at multiple time points; constructing joint data of the target object based on the feature parameters of multiple potential related factors of the target object and the change trajectory of the key physical examination characteristics; obtaining a first prediction result by processing the blood biomarker parameters of the target object at the target time point using a first health risk prediction model; obtaining a second prediction result by processing the joint data of the target object using a second health risk prediction model; and determining whether the target object has a health risk based on the first prediction result and the second prediction result.

[0124] In this embodiment, the storage medium includes, but is not limited to, Random Access Memory (RAM), Read-Only Memory (ROM), Cache, Hard Disk Drive (HDD), or Memory Card. The memory can be used to store computer program instructions. The network communication unit can be an interface configured according to standards specified in the communication protocol for network connection communication.

[0125] In this embodiment, the specific functions and effects implemented by the program instructions stored in the computer-readable storage medium can be explained in comparison with other embodiments, and will not be repeated here.

[0126] This specification also provides a computer program product, comprising at least a computer program, which, when executed by a processor, implements the following method steps: collecting blood biomarker parameters of a target object at multiple time points, and feature parameters of multiple potential related factors of the target object; obtaining composite biomarker parameters based on blood biomarkers; constructing a change trajectory of key physical examination characteristics of the target object based on the blood biomarker parameters of the target object at multiple time points, and the composite biomarker parameters based on blood biomarkers; extracting blood biomarker parameters of the target object at a target time point based on the blood biomarker parameters of the target object at multiple time points; constructing joint data of the target object based on the feature parameters of multiple potential related factors of the target object and the change trajectory of the key physical examination characteristics; obtaining a first prediction result by processing the blood biomarker parameters of the target object at the target time point using a first health risk prediction model; obtaining a second prediction result by processing the joint data of the target object using a second health risk prediction model; and determining whether the target object has a health risk based on the first prediction result and the second prediction result.

[0127] This specification also provides a health risk detection device, which may specifically include the following structural modules: The data acquisition module can be used to collect blood biomarker parameters of the target object at multiple time points, as well as feature parameters of multiple potential related factors of the target object; and to obtain composite biomarker parameters based on blood biomarkers. The first processing module can be specifically used to construct the change trajectory of the key physical examination characteristics of the target object based on the blood biomarker parameters of the target object at multiple time points and the composite biomarker parameters based on the blood biomarkers; and to extract the blood biomarker parameters of the target object at the target time point based on the blood biomarker parameters of the target object at multiple time points. The second processing module can be used to construct joint data of the target object based on the feature parameters of multiple potential related factors of the target object and the change trajectory of the key physical examination features; The prediction module can be specifically used to obtain a first prediction result by processing the blood biomarker parameters of the target object at a target time point using a first health risk prediction model; and to obtain a second prediction result by processing the joint data of the target object using a second health risk prediction model. The determination module can be used to determine whether the target object has a health risk based on the first prediction result and the second prediction result.

[0128] It should be noted that the units, devices, or modules described in the above embodiments can be implemented by computer chips or physical entities, or by products with certain functions. For ease of description, the above devices are described by dividing them into various modules according to their functions. Of course, in implementing this specification, the functions of each module can be implemented in one or more software and / or hardware, or the module that implements the same function can be implemented by a combination of multiple sub-modules or sub-units, etc. The device embodiments described above are merely illustrative. For example, the division of units is only a logical functional division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection between the devices or units shown or discussed can be through some interfaces, and the indirect coupling or communication connection between devices or units can be electrical, mechanical, or other forms.

[0129] As can be seen from the above, the health risk detection device provided in the embodiments of this specification can make full use of existing physical examination data and can predict the health risk status of the target object efficiently and accurately at a low cost.

[0130] In a specific scenario example, the health risk detection system provided in this manual can be used to predict the risk of cancer incidence (a health risk) based on blood cell biomarkers (e.g., blood marker parameters). For detailed implementation procedures, please refer to [the manual's documentation]. Figure 8 As shown, it includes the following content.

[0131] In this scenario example, existing cancer risk identification and prediction technologies mainly rely on genomics, proteomics, multi-omics integrated modeling, circulating tumor DNA detection, and imaging analysis. Although these methods have made significant progress in specific cancer types, they generally suffer from high detection costs, complex operations, and difficulties in sample acquisition, making large-scale application in the general population difficult. Furthermore, some studies attempt to assess cancer risk based on single hematological test results or individual immune inflammatory markers; however, due to the influence of temporary physiological states, environmental factors, or short-term illnesses, the prediction results fluctuate greatly and fail to accurately reflect long-term immune and inflammatory trends, resulting in limited stability and generalization ability of the prediction models.

[0132] In practical implementation, existing technologies still have the following shortcomings: First, most methods are based solely on static data, failing to capture time-series changes in blood cell parameters and revealing the dynamic evolution of immune and inflammatory states. Second, research subjects are mostly focused on single cancer types, lacking a unified analytical framework that can be generalized to multiple cancer types or even pan-cancer studies. Third, existing models largely rely on complex omics detection or image analysis, resulting in high computational demands and costs, making them unsuitable for widespread deployment in health checkups or population screening scenarios. These issues limit the practicality of existing technologies in early cancer risk identification, population-based prevention and control, and personalized monitoring.

[0133] Furthermore, considering that in recent years, blood cell parameters have gradually become a new direction in cancer risk research due to their advantages such as simple detection, low cost, and stable results, among which, white blood cells, red blood cells, platelets, and their derived ratios in the blood can reflect the body's systemic immune and inflammatory status, and multiple epidemiological studies have shown that they are significantly correlated with the risk of various cancers. However, most existing studies are based on single-test data and fail to fully consider the impact of physiological fluctuations and time-series changes on prediction results, resulting in insufficient model stability and generalization ability; at the same time, existing methods mostly focus on specific cancer types and lack a unified pan-cancer risk prediction system.

[0134] Based on the above considerations, this scenario example proposes a cancer incidence risk prediction method and system based on blood cell biomarkers. By analyzing the trajectory of blood cell parameters at multiple time points, a characteristic model reflecting the long-term dynamic changes in immunity and inflammation in the body is constructed, thereby achieving risk prediction for pan-cancer and various common cancers. Based on this method and system, automated analysis can be performed using routine physical examination blood data. It has advantages such as simple detection, low cost, high repeatability, and wide applicability, effectively overcoming the shortcomings of existing technologies in dynamic monitoring and pan-cancer prediction. Specifically, it may include the following steps.

[0135] 1. Data Collection

[0136] 1.1 Exposure Variables

[0137] Blood cell biomarker data and cancer diagnosis results were collected from a large-scale health checkup population as the basis for modeling analysis. The collected blood cell indicators not only cover routine red blood cell, white blood cell, and T lymphocyte parameters, but also include composite ratios reflecting the body's inflammation and immune status, which can support the early identification and dynamic prediction of cancer risk. Specifically, this includes the following 34 blood-related biomarkers (e.g., blood biomarker parameters): Red blood cell related parameters (8 items): Red blood cell count (RBC), red blood cell distribution width SD (RDW-SD), red blood cell distribution width CV (RDW-CV), hematocrit (HCT), mean corpuscular hemoglobin concentration (MCHC), mean corpuscular hemoglobin content (MCH), mean corpuscular volume (MCV), and hemoglobin concentration (Hb). White blood cell related indicators (11 items): white blood cell count (WBC), neutrophil count (NeuC), neutrophil percentage (Neu%), lymphocyte count (LymC), lymphocyte percentage (Lym%), eosinophil count (EosC), eosinophil percentage (Eos%), basophil count (BasC), basophil percentage (Bas%), monocyte count (MonC), monocyte percentage (Mon%). T lymphocyte-related indicators (7 items): CD3 count, CD3 percentage (CD3%), CD4 count, CD4 percentage (CD4%), CD8 count, CD8 percentage (CD8%), CD4 / CD8 ratio; Other parameters (3 items): platelet count (PLT), erythrocyte sedimentation rate (ESR), and erythrocyte sedimentation rate equation K value (ESR-K).

[0138] In addition to basic blood cell indicators, five ratio-based biomarkers related to inflammation response (e.g., composite biomarker parameters) are further introduced: basophil to lymphocyte ratio (BLR), eosinophil to lymphocyte ratio (ELR), monocyte to lymphocyte ratio (MLR), neutrophil to lymphocyte ratio (NLR), and platelet to lymphocyte ratio (PLR) to comprehensively assess the body's inflammation and immune status.

[0139] The data collection work has continued for many years, and some participants (e.g., sample subjects) have undergone multiple health examinations, resulting in repeated measurement data across time points. This data has good longitudinal tracking capabilities and can support modeling and analysis of the dynamic trajectory of blood cells.

[0140] For data acquisition, the XE-2100 and XE-5000 systems (Sysmex, Kobe, Japan) were used to determine hematological analytes, including RBC, WBC, PLT, and Hb. Flow cytometry (six-color flow cytometry, BD Company, USA) was used to detect T lymphocyte (CD3, CD4, and CD8) levels; ESR and ESR-K were detected using Alifax Test 1 (ALIFAX Company, Italy).

[0141] 1.2 Outcome Variables

[0142] Here, clinically confirmed diagnoses from hospital electronic medical records are used to determine outcomes (e.g., disease outcomes). These data are derived from the WHALE (West-China Hospital Alliance Longitudinal Epidemiology Wellness) study, a large-scale longitudinal health checkup cohort study involving approximately 700,000 individuals with a follow-up period of over 10 years.

[0143] In all regression-based analyses, including generalized linear mixed-effects models and standard Cox proportional hazards models, the outcome variable was binary, indicating whether cancer occurred. For trajectory analyses, participants were first categorized into different potential groups based on longitudinal patterns of biomarker levels. Subsequent analyses examined the association between trajectory group membership and binary cancer status. To achieve cancer-specific risk prediction, the method further subdivided the cancer outcome variable into 23 major cancer types based on data from diagnosed cancer cases in the WHALE study. The cancer types covered include: (1) Lung cancer; (2) Liver cancer; (3) Stomach cancer; (4) Colorectal cancer; (5) Esophageal cancer; (6) Pancreatic cancer; (7) Gallbladder cancer; (8) Brain tumors and central nervous system tumors; (9) Leukemia; (10) Lymphoma; (11) Bladder cancer; (12) Kidney cancer; (13) Cervical cancer; (14) Endometrial cancer; (15) Ovarian cancer; (16) Breast cancer; (17) Testicular cancer; (18) Prostate cancer; (19) Skin melanoma; (20) Lip, oral cavity and pharyngeal cancer; (21) Laryngeal cancer; (22) Nasopharyngeal cancer; (23) Thyroid cancer.

[0144] The classification results of the aforementioned cancer types can be used as binary variables for overall risk prediction, or as multi-class labels for trajectory modeling analysis of specific cancer types. By analyzing the trajectory correlations of samples from different cancer types, dynamic trajectory models for specific cancer types can be identified, supporting personalized early identification of cancer.

[0145] 1.3 Covariates

[0146] The development and progression of cancer are influenced by a variety of demographic characteristics, lifestyle factors, metabolic status, infectious factors, and disease heredity (e.g., multiple potential associated factors). These factors are not only closely related to cancer risk itself, but can also have a lasting impact on the baseline levels of hematological indicators and their long-term trends.

[0147] Specifically, age, sex, body mass index (BMI), smoking status, alcohol consumption, hypertension, diabetes, family history of cancer, hepatitis B virus (HBV) infection status, fatty liver grade, and liver function indicators aspartate aminotransferase (AST) and alanine aminotransferase (ALT) were selected as covariates. Age, sex, and BMI were used to characterize individual baseline physiological differences; smoking status and alcohol consumption reflected lifestyle-related exposures; hypertension and diabetes characterized chronic metabolic states; and family history of cancer reflected potential genetic background. Given the close association between HBV infection, fatty liver, and abnormal liver function with the occurrence of liver cancer, HBV infection status and fatty liver grade were selected to characterize specific infection and metabolic-related liver states, while AST and ALT were selected to reflect liver function levels and their abnormalities.

[0148] By incorporating the covariates and repeated measures data of hematological indicators into the longitudinal modeling process, the true trajectory of changes in hematological indicators can be more accurately depicted while controlling for interference from non-target factors, providing a reliable data foundation for subsequent cancer risk assessment based on trajectory features.

[0149] 2. Data Preprocessing

[0150] The collected data undergoes quality control to exclude data that does not meet requirements. For example, individuals diagnosed with cancer during their initial health check are excluded to focus on research into newly diagnosed cases. Furthermore, the data is standardized, such as using z-scores to standardize biomarker data, to facilitate subsequent analysis.

[0151] 3. Establish a prediction model

[0152] 3.1 Establishment Steps

[0153] In this scenario example, we first analyze the association between hematological indicators and cancer incidence using a generalized linear mixed-effects model based on longitudinal repeated measures data. Covariates are then gradually introduced to control for potential confounding factors such as demographic characteristics, lifestyle, and baseline health status. Subsequently, under the same covariate correction conditions, a Cox proportional hazards model is used to conduct a time-event analysis. Sensitivity analysis and robustness verification are performed on the above association results to test the consistency of conclusions under different statistical modeling assumptions.

[0154] Building upon this, for subjects with ≥2 physical examination records (e.g., the first type of sample), a latent class mixed-effects model was used to model the long-term trends of hematological indicators. The covariates were simultaneously incorporated into the modeling process, and individual differences were corrected by setting individual-level random intercepts and random time slopes. This allowed for the extraction of trajectory features reflecting long-term immune and inflammatory status changes after controlling for covariates. Finally, using these trajectory features as core input variables, a logistic regression model was employed to evaluate the association between different hematological trajectory features and cancer occurrence.

[0155] The health checkup data package contained multiple test results of participants from 2010 to 2023. First, a baseline descriptive statistical analysis was performed based on whether the subjects were screened for cancer. Continuous variables were expressed as median and interquartile range (IQR, 25th–75th percentile), and categorical variables were expressed as number (n) and percentage (%).

[0156] In practical implementation, to account for the impact of repeated measures on individuals in the time series, a generalized linear mixed-effects model (GLMM) is introduced. The binary outcome variable is set as whether cancer occurred at each physical examination, and a random intercept based on individual ID is introduced. This model can be expressed as:

[0157] in, This represents the probability value predicted for sample object i with number j, indicating a health risk. Represents the random intercept term. This indicates the factors affecting blood biomarker parameters. This represents the blood biomarker parameters for test subject i in sample j. The feature parameters (or covariates) represent the multiple potential association factors when considering sample object i with number j. This represents the influence of multiple potential association factors when sample object i is numbered j. This represents the random response of sample object i with number j.

[0158] In this model, hematological biomarkers are included in the analysis as time-varying variables. Each physical examination constitutes a repeated observation point to characterize the association between hematological indicators and cancer occurrence at different observation time points.

[0159] In practice, to verify the robustness of the results, a sensitivity analysis was conducted using a Cox proportional hazards regression model. This model uses follow-up time as the time scale and cancer incidence as the outcome variable, incorporating the same covariates to model the association between blood cell biomarkers and cancer incidence risk over time. The reliability and robustness of the main analysis results were further validated by comparing the consistency of the hazard ratio (HR) and its 95% confidence interval under different models. The model is specifically expressed in the following form:

[0160] in, This represents the instantaneous risk function of disease onset for sample object i at time t. This is the baseline risk function, reflecting the risk level when all covariates are zero; Indicates the blood biomarker levels of the sample subjects; Let k be the kth covariate (e.g., age, gender, BMI, smoking status, etc.). represents the corresponding regression coefficient.

[0161] This model primarily utilizes individual follow-up time information to assess the overall temporal association between hematological marker levels and cancer incidence risk, and is used to verify the robustness of results obtained based on repeated measures analysis.

[0162] In practice, to identify the long-term change trajectory of blood biomarkers, individuals with ≥2 health records are included. Latent Class Mixture Model (LCMM) is used to perform trajectory analysis on key biomarkers, fitting linear, quadratic, or cubic polynomial curves, and the number of trajectory categories is determined based on the following criteria: minimum Yess Information Criterion (BIC); posterior classification probability ≥0.70; and each category accounts for ≥2%.

[0163] In trajectory modeling, a latent class mixed-effects model is used to fit the changes of hematological biomarkers over time to characterize the overall changing trends of different latent trajectory categories. The model also incorporates an individual-level random effects structure. By setting a random intercept to represent individual differences at baseline levels of hematological indicators, and by setting a random time slope to characterize individual differences in the rate of change of indicators over time, the model systematically corrects for inter-individual heterogeneity during trajectory modeling, making the resulting trajectories more stably reflect the long-term changing patterns of hematological indicators.

[0164] Based on this, the longitudinal measurement data of the subjects are used to classify their trajectories and extract the corresponding trajectory feature parameters. These trajectory feature parameters serve as characteristic variables representing changes in long-term immune and inflammatory states and are incorporated into the risk assessment model after covariate adjustment. Specifically, a logistic regression model is used to model the relationship between the trajectory features and cancer occurrence. Under consistent covariate control conditions, the statistical association between different trajectory features and the risk of cancer is evaluated, thereby achieving cancer risk assessment based on long-term hematological trajectory features.

[0165] In all regression-based analyses, four stepwise adjusted models were constructed: Model 0 was the unadjusted model; Model 1 adjusted for age, sex, and BMI; Model 2 further adjusted for smoking and alcohol consumption factors based on Model 1; and Model 3 further incorporated hypertension, diabetes, and family history of cancer. For liver cancer, Model 3 additionally incorporated HBV infection, fatty liver grade, and AST and ALT to control for liver status factors closely related to the occurrence of liver cancer.

[0166] Furthermore, considering previous studies have shown that there may be gender and age differences in the correlation between blood biomarkers and cancer occurrence, this invention further conducts stratified analyses by gender and age to explore potential heterogeneity. Simultaneously, for the top 10 most common cancer types with high incidence rates, subgroup analyses of cancer types are conducted within the same trajectory modeling and risk assessment framework to verify the applicability and consistency of the proposed method across different cancer types.

[0167] In practice, all analyses were performed using R software (version 4.2.0, primarily relying on the "lme4", "TableOne", and "lcmm" packages), and a p-value < 0.05 was considered significant. The Benjamini-Hochberg procedure was used to adjust all p-values ​​for false discovery rate (FDR).

[0168] 3.2 Analytical methods involved: In practice, to model changes in cancer symptoms across multiple follow-ups and to consider the correlation of repeated measures within an individual, GLMM was used for analysis. This model uses a participant's unique identifier (such as ID) as a random intercept term, introducing intra-individual variation and improving its adaptability to hierarchical dependency structures.

[0169] Code example: glmer(status ~ biomarker + (1 | ID), data = ..., family = binomial).

[0170] In practice, to assess the temporal correlation between biomarker levels and cancer risk, a Cox proportional hazards regression model was used. This model uses individual follow-up time as the time scale and cancer incidence events (1 = cancer occurrence, 0 = no occurrence) as the outcome variable, characterizing the association between biomarkers and cancer risk over time. The model can estimate the relative impact of different biomarker levels on cancer risk while controlling for various confounding factors.

[0171] R code example: # Model Fitting fit <- coxph(Surv(time_to_event, status) ~ biomarker + age + sex +BMI + smoking + drinking + hypertension + diabetes, data = ...).

[0172] To characterize the dynamic changes of blood biomarkers in longitudinal data and explore their potential association with cancer risk, LCMM modeling was used.

[0173] This method can identify subgroups of individuals with different time trends and estimate the trajectory function corresponding to each group. The model form allows for a multinomial time structure (linear, quadratic, etc.) and allows for setting random intercepts and random time slopes to support individual-level variability.

[0174] Code example: m3.1 <- hlme( fixed = value ~ time + I(time^2), mixture = ~ time + I(time^2), random = ~ 1 + time, ng = 3, # Set the number of potential categories subject = "dah", data = rundata, B = initial_model ) .

[0175] The number of trajectory categories was determined based on a combination of the Bayesian Information Criterion (BIC), posterior classification probability (>0.7), and minimum category size (≥2% of the sample size). Finally, a logistic regression model was constructed to analyze the relationship between each trajectory group and cancer risk, and to explore the interaction of moderating factors such as gender.

[0176] 3.3 Code Usability

[0177] The analysis was performed in R using publicly available software packages, including lcmm for trajectory modeling, lme4 for generalized linear mixed-effects models, and survival for Cox risk proportion models.

[0178] 4. System Implementation

[0179] Based on the above ideas, a cancer incidence risk prediction system based on the aforementioned methods was also developed. This system can receive users' blood cell biomarker data, and through established prediction models and trajectory analysis methods, output the user's cancer risk prediction results, providing a basis for early intervention.

[0180] The above scenario examples validate the health risk detection system provided in this specification. First, it not only focuses on biomarker levels at a single time point but also considers the dynamic changes of biomarkers through trajectory analysis, enabling a more comprehensive reflection of the changing trends in cancer risk and improving prediction accuracy. Compared with existing technologies, this method is better able to capture the dynamic changes of biomarkers during the development of cancer, thus more accurately identifying high-risk individuals. Second, it can provide personalized cancer incidence risk prediction services for large-scale health check-up populations, facilitating early detection and intervention of cancer, and has significant social and economic benefits. Early identification of high-risk individuals allows for timely preventative measures, reducing the occurrence and progression of cancer, which is of great significance for improving public health. Furthermore, by employing various advanced statistical methods in data analysis, such as generalized linear mixed-effects models, Cox proportional hazards models, and latent class mixed-effects models, it can more effectively process repeated measures data and fully utilize individual data changes over time, further improving the accuracy and reliability of prediction. The comprehensive application of these methods makes this application more innovative and advantageous technically.

[0181] While this specification provides the steps of operation for the methods described in the embodiments or flowcharts, more or fewer steps may be included based on conventional or non-inventive means. The order of steps listed in the embodiments is merely one possible order of execution among many steps and does not represent the only possible order. In actual device or client product execution, the methods shown in the embodiments or drawings may be executed sequentially or in parallel (e.g., in a parallel processor or multi-threaded processing environment, or even a distributed data processing environment). The terms "comprising," "including," or any other variations thereof are intended to cover a non-exclusive inclusion, such that a process, method, product, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, product, or apparatus. Without further limitations, the presence of other identical or equivalent elements in a process, method, product, or apparatus that includes said elements is not excluded. The terms "first," "second," etc., are used to denote names and do not indicate any particular order.

[0182] Those skilled in the art will also know that, besides implementing the controller using purely computer-readable program code, the same functions can be achieved by logically programming the method steps, making the controller function as logic gates, switches, application-specific integrated circuits (ASICs), programmable logic controllers (PLCs), and embedded microcontrollers. Therefore, such a controller can be considered a hardware component, and the devices within it used to implement various functions can also be considered structures within that hardware component. Alternatively, the devices used to implement various functions can be considered as both software modules implementing the method and structures within a hardware component.

[0183] This specification can be described in the general context of computer-executable instructions that are executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, classes, etc., that perform a specific task or implement a specific abstract data type. This specification can also be practiced in distributed computing environments, where tasks are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer-readable storage media, including storage devices.

[0184] As can be seen from the above description of the embodiments, those skilled in the art can clearly understand that this specification can be implemented by means of software plus necessary general-purpose hardware platforms. Based on this understanding, the technical solutions of this specification can essentially be embodied in the form of a software product. This computer software product can be stored in a storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, mobile terminal, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments of this specification.

[0185] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to interchangeably. Each embodiment focuses on its differences from other embodiments. This specification can be used in numerous general-purpose or special-purpose computer system environments or configurations. Examples include: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable electronic devices, network PCs, minicomputers, mainframe computers, and distributed computing environments including any of the above systems or devices, etc.

[0186] Although this specification has been described by way of examples, those skilled in the art will recognize that many variations and modifications are possible without departing from the spirit of this specification, and it is intended that the appended claims cover such variations and modifications without departing from the spirit of this specification.

Claims

1. A health risk detection system, characterized in that, At least including: The system includes a physical examination information collection module, a related information collection module, and a data processing module; among them, The physical examination information collection module is used to collect blood biomarker parameters of the target object at multiple time points; and to obtain composite biomarker parameters based on blood biomarkers; The association information acquisition module is used to collect feature parameters of multiple potential association factors of the target object; The data processing module is used to construct the change trajectory of key physical examination characteristics of the target object based on blood biomarker parameters at multiple time points and composite biomarker parameters based on blood biomarkers; and to extract the blood biomarker parameters of the target object at a target time point based on the blood biomarker parameters of the target object at multiple time points; to construct joint data of the target object based on the feature parameters of multiple potential related factors of the target object and the change trajectory of the key physical examination characteristics; to obtain a first prediction result by processing the blood biomarker parameters of the target object at the target time point using a first health risk prediction model; and to obtain a second prediction result by processing the joint data of the target object using a second health risk prediction model; and to determine whether the target object has a health risk based on the first prediction result and the second prediction result.

2. The health risk detection system according to claim 1, characterized in that, The first health risk prediction model is a model that is trained in advance according to a preset joint training rule, based on a generalized linear mixed effects model and combined with a Cox risk proportion model. The second health risk prediction model is a model trained in advance according to a preset joint training rule, based on a hybrid model of latent class mixed effects model and logistic regression model.

3. The health risk detection system according to claim 1, characterized in that, The various potential association factors include: basic attribute association factors, lifestyle association factors, metabolic status association factors, infection, disease and test results association factors, and family history association factors; The basic attribute correlation factors include at least: age, gender, and body mass index; the lifestyle correlation factors include at least: smoking habits and drinking habits; the metabolic status correlation factors include at least: hypertension prevalence and diabetes prevalence; the infection, disease, and test correlation factors include at least: HBV infection status, fatty liver prevalence, and liver function indicators; and the family history correlation factors include at least: the current disease status of related relatives and the historical disease records of related relatives.

4. The health risk detection system according to claim 1, characterized in that, The composite biomarker parameters include at least the following: basophil to lymphocyte ratio, eosinophil to lymphocyte ratio, monocyte to lymphocyte ratio, neutrophil to lymphocyte ratio, and platelet to lymphocyte ratio.

5. The health risk detection system according to claim 1, characterized in that, The step of determining whether the target object has a health risk based on the first prediction result and the second prediction result includes: Based on the preset model weight ratio coefficient mapping relationship, the target weight ratio coefficients for the first prediction result and the second prediction result are determined. Using a preset dynamic adjustment model, based on the change trajectory of the key physical examination characteristics of the target object, the blood biomarker parameters at the target time point, and the characteristic parameters of multiple potential related factors, the first degree of fit for the first health risk prediction model and the second degree of fit for the second health risk prediction model are determined. Using a preset dynamic adjustment model, the target weight ratio coefficient is adjusted based on the first and second fit degrees to obtain the adjusted target weight ratio coefficient. Based on the first prediction result, the second prediction result, and the adjusted target weight ratio coefficient, the corresponding target prediction result is determined. Based on the target prediction results, determine whether the target object has any health risks.

6. The health risk detection system according to claim 2, characterized in that, The health risk detection system further includes a training module; wherein the training module is used for: Obtain the physical examination records, medical records, and investigation records of various potential related factors of the sample subjects; Based on the physical examination records, the sample subjects with multiple physical examinations and the sample subjects with a single physical examination were identified as the first type of sample subjects and the second type of sample subjects, respectively. A first sample dataset is constructed based on the physical examination records and medical records of the first type of sample subjects, as well as the investigation records of multiple potential related factors; a second sample dataset is constructed based on the physical examination records and medical records of the second type of sample subjects, as well as the investigation records of multiple potential related factors. Construct a first initial prediction model based on a generalized linear mixed effects model, and a second initial prediction model based on a mixture of latent class mixed effects model and logistic regression model; According to the preset joint training rules, the first sample dataset and the second sample dataset are used together, combined with the Cox risk ratio model, to train the first initial prediction model and obtain a first health risk prediction model that meets the requirements. According to the preset joint training rules, the second initial prediction model is trained using the first sample dataset to obtain a second health risk prediction model that meets the requirements. Based on the second sample dataset, a random test dataset is constructed; and using the random test dataset, the first health risk prediction model and the second health risk prediction model are randomly jointly tested to correct the first health risk prediction model and the second health risk prediction model, and a preset model weight ratio coefficient mapping relationship based on the first health risk prediction model and the second health risk prediction model is determined.

7. The health risk detection system according to claim 6, characterized in that, The step of training the first initial prediction model according to preset joint training rules, using the first sample dataset and the second sample dataset in conjunction with the Cox risk ratio model, to obtain a first health risk prediction model that meets the requirements includes: By combining the first and second sample datasets, a third sample dataset is constructed. Based on the third sample dataset, a first training set and multiple first sub-sample datasets are constructed; wherein, the first training set does not contain survey records of multiple potential correlation factors; the first sub-sample datasets do not contain sample physical examination records, and each of the first sub-sample datasets corresponds to one type of potential correlation factor; The first initial prediction model is trained using the first training set to obtain the first intermediate prediction model that meets the requirements; the third initial prediction model based on the Cox risk ratio model is trained using the first training set to obtain the corresponding auxiliary model. According to the preset training order, the first intermediate prediction model is iteratively trained in multiple rounds using multiple first subsample datasets to obtain the corresponding first target prediction model. The first target prediction model is validated and corrected using an auxiliary model to obtain a first health risk prediction model that meets the requirements.

8. The health risk detection system according to claim 6, characterized in that, The step of training the second initial prediction model using the first sample dataset according to preset joint training rules to obtain a second health risk prediction model that meets the requirements includes: Based on the second sample dataset, a second training set and multiple second sub-sample datasets are constructed; wherein, the second training set does not contain survey records of multiple potential correlation factors; the second sub-sample datasets do not contain sample physical examination records, and each of the second sub-sample datasets corresponds to one type of potential correlation factor; The second initial prediction model is trained using the second training set to obtain the second intermediate prediction model that meets the requirements. Following a preset training order, the second intermediate prediction model is iteratively trained multiple times using multiple second subsample datasets to obtain a second health risk prediction model that meets the requirements.

9. A health risk detection method using the health risk detection system according to any one of claims 1 to 8, characterized in that, include: Collect blood biomarker parameters of the target object at multiple time points, as well as feature parameters of multiple potential related factors of the target object; And obtain the parameters of composite biomarkers based on blood biomarkers; Based on the blood biomarker parameters of the target object at multiple time points, and the composite biomarker parameters based on the blood biomarkers, the change trajectory of the key physical examination characteristics of the target object is constructed; Based on the blood biomarker parameters of the target object at multiple time points, the blood biomarker parameters of the target object at the target time point are extracted; Based on the characteristic parameters of multiple potential related factors of the target object and the change trajectory of the key physical examination characteristics, joint data of the target object is constructed. The first prediction result is obtained by processing the blood biomarker parameters of the target object at the target time point using the first health risk prediction model; The second health risk prediction model is then used to process the joint data of the target objects to obtain a second prediction result. Based on the first prediction result and the second prediction result, it is determined whether the target object has a health risk.

10. A computer program product, characterized in that, It includes a computer program that, when executed by a processor, implements the steps of the method of claim 9.