Method for predicting new hypertension

Through deep learning architecture and feature interaction algorithms, multi-category feature data is processed, and the problem of low accuracy in predicting new hypertension in the prior art is solved, achieving more efficient and economical prediction results.

CN120072298APending Publication Date: 2025-05-30THE SEVENTH MEDICAL CENTER OF PLA GENERAL HOSPITAL
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510135044.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-07
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

When predicting new hypertension, the prior art has low economic efficiency, low work efficiency and low accuracy, and cannot meet work needs.

Method used

Using the deep learning architecture, the historical data of the sample population is preprocessed, the first sample feature data is decomposed to obtain the second sample feature data with a larger number of category features, and the feature interaction algorithm is integrated for deep learning training to generate a predicted new hypertension model.

Benefits of technology

Improve the accuracy of predicting new hypertension results, reduce dependence on characteristic engineering, improve work efficiency and reduce economic costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120072298A_ABST
    Figure CN120072298A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of medical diagnosis, and discloses a method for predicting new hypertension, which is applied to a deep learning architecture, and comprises the following steps: preprocessing historical data of a sample population to obtain first sample feature data; decomposing the first sample feature data to obtain second sample feature data, the number of category features of the second sample feature data being greater than the number of category features of the first sample feature data, the first sample feature data including a plurality of different category feature data, and the category feature data including category features; fusing a feature interaction algorithm, and performing deep learning training by using the second sample feature data to obtain a new hypertension prediction model; and on the basis of the new hypertension prediction model, inputting feature data of a person with new hypertension to be predicted, and generating a prediction result of the new hypertension. The method for predicting the new hypertension provided by the invention has better generalization, and the accuracy of predicting the new hypertension is greatly improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of medical diagnosis, and particularly relates to a method for predicting new-onset hypertension. Background Art

[0002] In the prior art, traditional machine learning models are mainly used to predict new-onset hypertension, such as logistic regression, SVM (Support Vector Machine), XGBoost, and LightGBM tree ensemble models. These machine learning models require a large number of manual feature selection operations and have a strong dependence on feature engineering. They are suitable for learning and training using relatively small-scale feature data. Since there are many factors affecting new-onset hypertension, it is necessary to further predict whether the interviewee has the risk of newly developing hypertension under the influence of these factors. In this case, the traditional machine learning model for predicting new-onset hypertension has low economy, low work efficiency, and low accuracy, and cannot meet the work requirements.

[0003] With the development of deep learning, some deep learning models have been applied to process tabular data in the medical field. However, when using more feature data for learning and training, compared with the above traditional machine learning models, the accuracy of the generated results is not significant, and the persistence of the accuracy of the generated results is poor, resulting in low accuracy of the prediction results. Summary of the Invention

[0004] Based on the above technical requirements, the present invention provides a method for predicting new-onset hypertension, which at least has the technical effect of being able to greatly improve the accuracy of the prediction result of new-onset hypertension.

[0005] A method for predicting new-onset hypertension provided by an embodiment of the present invention is applied to a deep learning architecture. The method includes:

[0006] S1: Preprocess the historical data of the sample population to obtain first sample feature data;

[0007] S2: Decompose the first sample feature data to obtain second sample feature data. The number of category features of the second sample feature data is greater than the number of category features of the first sample feature data, where the first sample feature data includes multiple different category feature data, and the category feature data includes the category features;

[0008] S3: Incorporate a feature interaction algorithm, and use the second sample feature data for deep learning training to obtain a model for predicting new-onset hypertension;

[0009] S4: Based on the predicted new-onset hypertension model, input the characteristic data of the person to be predicted for new-onset hypertension, and generate the prediction result of their new-onset hypertension, where the characteristic data of the person to be predicted for new-onset hypertension corresponds to the category characteristic data of the second sample characteristic data.

[0010] Further, according to the method for predicting new-onset hypertension provided by the embodiments of the present invention, the first sample characteristic data in S1 includes: age category characteristic data, hip circumference category characteristic data, waist circumference category characteristic data, waist-hip ratio category characteristic data, and triceps skinfold thickness category characteristic data; wherein, the age category characteristic data includes a first age numerical interval, the hip circumference category characteristic data includes a first hip circumference numerical interval, the waist circumference category characteristic data includes a first waist circumference numerical interval, the waist-hip ratio category characteristic data includes a first waist-hip ratio numerical interval, and the triceps skinfold thickness category characteristic data includes a first triceps skinfold thickness numerical interval.

[0011] Further, according to the method for predicting new-onset hypertension provided by the embodiments of the present invention, decomposing the first sample characteristic data in S2 to obtain the second sample characteristic data includes the following steps:

[0012] Decompose the first age numerical interval according to a first preset rule to obtain a plurality of different second age numerical intervals, and each second age numerical interval in the plurality of different second age numerical intervals is a category characteristic data in the second sample characteristic data;

[0013] Decompose the first hip circumference numerical interval according to a second preset rule to obtain a plurality of different second hip circumference numerical intervals, and each second hip circumference numerical interval in the plurality of different second hip circumference numerical intervals is a category characteristic data in the second sample characteristic data;

[0014] Decompose the first waist circumference numerical interval according to a third preset rule to obtain a plurality of different second waist circumference numerical intervals, and each second waist circumference numerical interval in the plurality of different second waist circumference numerical intervals is a category characteristic data in the second sample characteristic data;

[0015] Decompose the first waist-hip ratio numerical interval according to a fourth preset rule to obtain a plurality of different second waist-hip ratio numerical intervals, and each second waist-hip ratio numerical interval in the plurality of different second waist-hip ratio numerical intervals is a category characteristic data in the second sample characteristic data;

[0016] Decompose the first triceps skinfold thickness numerical interval according to the fifth preset rule to obtain multiple different second triceps skinfold thickness numerical intervals, and each second waist-hip ratio numerical interval in the multiple different second triceps skinfold thickness numerical intervals is a category feature data in the second sample feature data.

[0017] Further, according to the method for predicting new-onset hypertension provided by the embodiments of the present invention, the second sample feature data further includes the following category feature data: province, urban-rural location, gender, education level, employment status, personal total income, marital status, medical insurance status, diabetes diagnosis, smoking status, drinking status, BMI situation, tea-drinking habit, daily bedridden time, height, weight, BMI, and per capita household income; wherein, the BMI represents the square of the calculation result of dividing the weight by the height.

[0018] Further, according to the method for predicting new-onset hypertension provided by the embodiments of the present invention, the feature interaction algorithm in S3 includes a parallel additive attention mechanism and multiplicative attention mechanism, as well as a prompt token, and the feature interaction algorithm is incorporated into the encoder layer of the deep learning architecture to obtain the target deep learning model.

[0019] Further, according to the method for predicting new-onset hypertension provided by the embodiments of the present invention, the deep learning architecture includes a Transformer architecture.

[0020] Further, according to the method for predicting new-onset hypertension provided by the embodiments of the present invention, segment the second sample feature data to obtain a test data set and a validation data set, and use the test data set and the validation data set to train and validate the target deep learning model to obtain a new-onset hypertension prediction model.

[0021] Further, according to the method for predicting new-onset hypertension provided by the embodiments of the present invention, the new-onset hypertension prediction model calculates the marginal contribution value of each category feature data in the second sample feature data through a SHAP block, and the marginal contribution value represents the influence degree of each category feature data in the second sample feature data on the result of predicting new-onset hypertension.

[0022] The present invention has at least the following beneficial effects:

[0023] Use the second sample feature data with more category feature data to further train the impact of different category feature data on predicting new-onset hypertension. Moreover, by integrating a feature interaction algorithm into the deep learning architecture and using a parallel attention mechanism and prompt tokens, the interaction between different category feature data is enhanced, the underfitting or overfitting phenomenon when processing category feature data is alleviated, the generalization ability of the model is improved, and the accuracy of the prediction result is greatly improved. At the same time, end-to-end deep learning is achieved, the dependence on feature engineering when processing the tabular data of predicting new-onset hypertension is reduced, the work efficiency is improved, and the economic cost is reduced. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] Figure 1 is a schematic flowchart of the method for predicting new-onset hypertension provided by an embodiment of the present invention;

[0025] Figure 2 is an ROC curve diagram provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0026] The technical solutions in the embodiments of the present invention will be clearly described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention belong to the scope of protection of the present invention.

[0027] The terms "first", "second", etc. in the specification and claims of the present invention are used to distinguish similar objects, rather than to describe a specific order or sequence. It should be understood that such terms can be interchanged under appropriate circumstances so that the embodiments of the present invention can be implemented in an order different from those illustrated or described herein, and the objects distinguished by "first" and "second" are usually of the same category, and the number of objects is not limited. For example, the first object can be one or more. In addition, "and / or" in the specification and claims means at least one of the connected objects, and the character " / " generally means that the associated objects before and after are in an "or" relationship.

[0028] The method for predicting new-onset hypertension provided by the embodiments of the present invention will be described in detail below with reference to the accompanying drawings through some embodiments.

[0029] See Figure 1-2 , a method for predicting new-onset hypertension provided by an embodiment of the present invention, is applied to a deep learning architecture, and the method includes:

[0030] S1: Preprocess the historical data of the sample population to obtain the first sample feature data;

[0031] It is understandable that the database adopted in the embodiments of the present invention is based on the sample population database already collected in the medical field.

[0032] It should be noted that the above-mentioned first sample feature data includes: age category feature data, hip circumference category feature data, waist circumference category feature data, waist-hip ratio category feature data, and triceps skinfold thickness category feature data; wherein, the age category feature data includes a first age numerical range, the hip circumference category feature data includes a first hip circumference numerical range, the waist circumference category feature data includes a first waist circumference numerical range, the waist-hip ratio category feature data includes a first waist-hip ratio numerical range, and the triceps skinfold thickness category feature data includes a first triceps skinfold thickness numerical range.

[0033] That is to say, age, triceps skinfold thickness, waist circumference, hip circumference, and waist-hip ratio each correspond to a numerical range.

[0034] It should be noted that the above-mentioned existing sample population database further includes multiple category feature data, which are respectively province, urban-rural location, medical insurance status, employment status, diabetes diagnosis, gender, smoking status, drinking status, education level, marital status, BMI category, personal total income category, and tea-drinking habit, daily bed rest time, age, triceps skinfold thickness TSF (triceps skinfold thickness), hip circumference, waist circumference, WHR (waist-hip ratio), height, weight; wherein, BMI is calculated as the square of the result of dividing the weight value by the height value, the unit of weight is kilogram, and the unit of height is meter.

[0035] It is understandable that there are some abnormal characteristic values in the historical data of the sample population, which need to be preprocessed. In a possible implementation manner, the abnormal characteristic values are replaced by the values calculated by KNN (k-nearest neighbor algorithm).

[0036] S2: Decompose the first sample feature data to obtain second sample feature data, where the number of category features of the second sample feature data is greater than the number of category features of the first sample feature data, and the first sample feature data includes multiple different category feature data, and the category feature data includes the category features;

[0037] In a possible embodiment, the decomposing the first sample feature data in S2 to obtain second sample feature data includes the following steps:

[0038] Decompose the first age value range according to a first preset rule to obtain multiple different second age value ranges, and each second age value range in the multiple different second age value ranges is a category feature data in the second sample feature data; in a possible embodiment, the value range corresponding to the age category feature is decomposed and divided into five different value ranges, namely less than 25 years old, 25 to 40 years old, 40 to 55 years old, 55 to 65 years old, and greater than 65 years old;

[0039] It can be understood that in the first sample feature data, the first age value range is greater than or equal to 18 years old.

[0040] Decompose the first hip circumference value range according to a second preset rule to obtain multiple different second hip circumference value ranges, and each second hip circumference value range in the multiple different second hip circumference value ranges is a category feature data in the second sample feature data; in a possible embodiment, the hip circumference category feature is divided into four value ranges, namely less than 80 cm, 80 to 90 cm, 90 to 100 cm, and greater than 100 cm;

[0041] It can be understood that in the first sample feature data, the first hip circumference value range is 40 to 180 cm.

[0042] Decompose the first waist circumference value range according to a third preset rule to obtain multiple different second waist circumference value ranges, and each second waist circumference value range in the multiple different second waist circumference value ranges is a category feature data in the second sample feature data; in a possible embodiment, the waist circumference category feature is divided into four value ranges, namely less than 70 cm, 70 to 80 cm, 80 to 90 cm, and greater than 90 cm;

[0043] It can be understood that in the first sample feature data, the first waist circumference value range is 40 to 180 cm.

[0044] Decompose the first waist-hip ratio value range according to a fourth preset rule to obtain multiple different second waist-hip ratio value ranges, and each second waist-hip ratio value range in the multiple different second waist-hip ratio value ranges is a category feature data in the second sample feature data; in a possible embodiment, the WHR category feature is divided into five value ranges, namely less than 0.8, 0.8 to 0.85, 0.85 to 0.9, 0.9 to 1, and greater than 1;

[0045] It can be understood that in the first sample feature data, the first waist-hip ratio value range is 0.2 to 5.

[0046] Decompose the first triceps skinfold thickness numerical interval according to the fifth preset rule to obtain multiple different second triceps skinfold thickness numerical intervals. Each second waist-hip ratio numerical interval among the multiple different second triceps skinfold thickness numerical intervals is a category feature data in the second sample feature data. In a possible embodiment, the TSF category features are divided into four numerical intervals, namely less than 10 mm, 10 to 16 mm, 16 to 22 mm, and greater than 22 mm.

[0047] It can be understood that in the first sample feature data, the first triceps skinfold thickness numerical interval is 2 to 40 mm.

[0048] Based on historical sample data, more specific category feature data are obtained through decomposition, and new sample feature data, that is, the second sample feature data, are obtained, which improves work efficiency, reduces economic costs, and is also convenient for further analyzing the results of predicting the occurrence of new-onset hypertension under different category feature data, thereby improving the accuracy of predicting new-onset hypertension.

[0049] S3: Incorporate a feature interaction algorithm, and use the second sample feature data to perform deep learning training to obtain a model for predicting new-onset hypertension.

[0050] Furthermore, according to the method for predicting new-onset hypertension provided by the embodiments of the present invention, in a possible embodiment, the feature interaction algorithm in S3 includes a parallel additive attention mechanism and multiplicative attention mechanism, as well as a prompt token. The feature interaction algorithm is incorporated into the encoder layer of the deep learning architecture to obtain a target deep learning model.

[0051] It should be noted that in the process of deep learning, there are often problems of overfitting or underfitting. The embodiments of the present invention extend the feature interaction algorithm to the deep learning architecture to perform feature data operation processing, reduce the impact of overfitting or underfitting on the model performance, and thus improve the accuracy of the model in predicting the results of new-onset hypertension.

[0052] It should be noted that the feature interaction algorithm includes a parallel attention mechanism and a prompt token. The parallel attention mechanism and the prompt token are used to enhance the interaction between the data features corresponding to the category features. This parallel mechanism processes the embeddings of the input feature data in two independent attention paths in parallel, one responsible for calculating additive interactions and the other for simulating multiplicative interactions, and then combines these two types of interaction candidates and integrates them through a fully connected layer.

[0053] The prompt token is a learnable prompt token, and the attention mechanism is a multi-head attention mechanism.

[0054] That is to say, the feature interaction algorithm includes two data processing mechanism paths. One is to perform additive interaction on feature data to obtain an additive interaction result, and the other is to perform multiplicative interaction on feature data to obtain a multiplicative interaction result. Then, the additive interaction result and the multiplicative interaction result are combined and integrated and output through a fully connected layer. That is to say, the feature interaction algorithm completes the integration of two arithmetic operations on the input feature data.

[0055] It can be understood that based on the parallel additive attention mechanism and multiplicative attention mechanism, feature interaction is performed on each category feature in the second sample feature data, and a new feature representation is generated and transmitted to the next layer in the deep learning model; wherein, the new feature representation corresponds to the information of the category feature data in the second sample feature data.

[0056] It should be noted that the feature interaction algorithm is incorporated into the encoder layer of the deep learning architecture. The encoder layer is responsible for converting the external input sequence into a high-dimensional representation. In one possible embodiment, the encoder layer includes a feature interaction algorithm sub-layer and a feed-forward neural network sub-layer. Residual connection and layer normalization are adopted after each sub-layer to stabilize training and accelerate convergence; the feature interaction algorithm sub-layer is connected to the feed-forward neural network sub-layer.

[0057] It can be understood that underfitting means that insufficient features are learned from the training data, resulting in poor performance of the model on both the training set and the test set; overfitting means that normal data as well as noise and random fluctuations in the training set are captured on the training data, rather than the true relationship between the data, resulting in poor performance of the model on the unseen test set.

[0058] It should be noted that the risk of underfitting usually stems from the lack of necessary feature data. In a specific embodiment, the feature interaction algorithm is to identify and extract key additive and multiplicative interaction candidate features, and then perform logical operations to solve the underfitting problem. At the same time, the prompt tokens in the feature interaction algorithm are used as queries for addition and multiplication to solve the potential overfitting problem caused by redundant feature data.

[0059] Furthermore, in one possible embodiment, in order to enable each layer of the neural network to learn the same interaction pattern, the Top-k algorithm is adopted to ignore the correlation of unimportant data information, which can prevent overfitting and enhance the resistance to data noise. The Top-k algorithm is located in the layer below the above-mentioned prompt token query.

[0060] It can be understood that the Top-k algorithm means only considering a number of candidate features with the highest affinity.

[0061] It should be noted that the feature interaction algorithm can learn the interactions between candidate features, that is, the weights of the correlations between candidate features are continuously adjusted to appropriate values during the training process. In some specific embodiments, for the input feature embedding matrix X, embeddings of queries, keys, and values are generated through trainable parameters, and their affinities are calculated to obtain the candidate features for additive interaction or multiplicative interaction. In the additive path, weights are assigned to the candidate features for additive arithmetic integration. In the multiplicative path, the candidate feature data is processed for multiplicative interaction on a logarithmic scale, and then the multiplicative interaction result is obtained.

[0062] To solve the problem of interaction sparsity, a hard attention mechanism is adopted, and only the top n affinity entries are retained, where the value of n is set according to the actual training model situation.

[0063] It can be understood that the candidate features correspond to the categorical feature data in the second sample feature data.

[0064] It can be understood that the prompt token is the translation of Prompt token, the key is the translation of Key, and the value is the translation of Value, which are well-known concepts in deep learning models and will not be elaborated here.

[0065] Furthermore, according to the method for predicting new-onset hypertension provided by the embodiments of the present invention, the deep learning architecture includes a Transformer architecture.

[0066] It should be noted that in a possible embodiment, a deep learning model is obtained by creating a model based on the Transformer architecture. The deep learning model further includes a decoder layer, and the decoder layer includes a multi-head attention mechanism sub-layer, an encoder-decoder attention mechanism layer, and a feed-forward neural network sub-layer. Residual connections and layer normalization are adopted after each sub-layer.

[0067] It can be understood that the input of the decoder layer includes the output of the encoder layer to generate a final output sequence.

[0068] Furthermore, according to the method for predicting new-onset hypertension provided by the embodiments of the present invention, the second sample feature data is segmented to obtain a test data set and a validation data set, and the target deep learning model is trained and validated using the test data set and the validation set to obtain a new-onset hypertension prediction model.

[0069] S4: Based on the new-onset hypertension prediction model, input the feature data of the person to be predicted for new-onset hypertension to generate a prediction result for their new-onset hypertension, where the feature data of the person to be predicted for new-onset hypertension corresponds to the categorical feature data of the second sample feature data.

[0070] It can be understood that based on the prediction model for newly developed hypertension provided by the embodiments of the present invention, by inputting the province, urban-rural location, medical insurance status, employment status, diabetes diagnosis, gender, smoking status, drinking status, education level, marital status, BMI category, personal total income category, tea-drinking habit, daily bed rest time, age, triceps skinfold thickness TSF (triceps skinfold thickness), hip circumference, waist circumference, WHR (waist-hip ratio), height, weight and corresponding values of the population to be predicted for newly developed hypertension, the risk result of predicting newly developed hypertension can be obtained.

[0071] It can be understood that based on the characteristic data of hypertension to be predicted input into the prediction model for newly developed hypertension, the risk value of predicting whether a person has hypertension can be obtained.

[0072] In a possible embodiment, the first sample population is 5,000 people. The dimension of the embedding vector in the hyperparameter setting is 336. The model depth is 3 encoder layers and 3 decoder layers. The number of features of each neural network layer in the 3 encoder layers and 3 decoder layers is set to 200, 100, and 50 respectively. The number of heads in the multi-head attention sublayer is 12. The dropout rates of the multi-head attention sublayer and the feed-forward neural network sublayer are 0.5 and 0.4 respectively. The sigmoid function is used in the output layer.

[0073] As Figure 2 shown, the curve in the figure is the Receiver Operating Characteristic curve, hereinafter referred to as the ROC curve, which represents the relationship between the True Positive Rate (hereinafter referred to as TPR) and the False Positive Rate (hereinafter referred to as FPR). The area under the ROC curve (AUC) value of the prediction model for newly developed hypertension provided by the embodiments of the present invention is 0.85. Based on the same sample characteristic data, when the conventional machine learning models Logistic Regression, SVM (Support Vector Machine), XGBoost, and LightGBM are at their optimal settings, the range of their AUC values is between 0.65 and 0.75.

[0074] Since the value range of AUC is greater than or equal to 0 and less than or equal to 1, the larger the AUC value, the more accurate the prediction result of the prediction model for newly developed hypertension.

[0075] TPR represents the ratio of the number of samples misjudged as positive classes to the number of all negative class samples; FPR represents the ratio of the number of samples correctly judged as positive classes to the number of all positive class samples. TPR and FPR are common knowledge in the art and will not be elaborated here.

[0076] The method for predicting new-onset hypertension provided by the embodiments of the present invention has significantly improved the accuracy of predicting the results of new-onset hypertension.

[0077] Furthermore, according to the method for predicting new-onset hypertension provided by the embodiments of the present invention, the prediction model for new-onset hypertension calculates the marginal contribution value of each categorical feature data in the second sample feature data through the SHAP block, and the marginal contribution value represents the influence degree of each categorical feature data in the second sample feature data on the prediction result of new-onset hypertension.

[0078] It can be understood that SHAP (SHapley Additive exPlanations) is a model interpretation method that can quantify the contribution of each categorical feature data to the prediction result of new-onset hypertension, and further obtain the degree of influence of each categorical feature data on new-onset hypertension.

[0079] The working principle of the method for predicting new-onset hypertension in this embodiment is as follows: based on historical sample feature data, it is decomposed to obtain second sample feature data with more categorical feature data; based on the Transformer architecture, a feature interaction algorithm is added, and this model is trained based on the second sample feature data to obtain a prediction model for new-onset hypertension.

[0080] It should be noted that in this article, the term "including", "comprising" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or also includes elements inherent to this process, method, article or device. Without further limitations, the element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, article or device including the element. In addition, it should be pointed out that the scope of the methods and devices in the embodiments of the present invention is not limited to performing functions in the order shown or discussed, and may also include performing functions in a substantially simultaneous manner or in the reverse order according to the functions involved. For example, the described method may be performed in an order different from that described, and various steps may be added, omitted, or combined. Additionally, the features described with reference to certain examples may be combined in other examples.

[0081] The embodiments of the present invention have been described above in conjunction with the accompanying drawings, but the present invention is not limited to the above specific embodiments. The above specific embodiments are merely illustrative and not restrictive. Under the inspiration of the present invention, those of ordinary skill in the art can also make many forms without departing from the purpose of the present invention and the scope protected by the claims, and all of them belong to the protection scope of the present invention.

Claims

1. A method for predicting new-onset hypertension, applied to a deep learning architecture, characterized in that: The method comprises: S1: preprocess the historical data of the sample population to obtain the first sample characteristic data; S2: decomposing the first sample feature data to obtain second sample feature data, wherein the number of category features of the second sample feature data is greater than the number of category features of the first sample feature data, wherein the first sample feature data includes a plurality of different category feature data, and the category feature data includes the category feature; S3: Incorporating a feature interaction algorithm, using the second sample feature data, performing deep learning training, and obtaining a model for predicting new-onset hypertension; S4: Based on the model for predicting new-onset hypertension, characteristic data of the person with new-onset hypertension to be predicted is input to generate a prediction result of the new-onset hypertension, wherein the characteristic data of the person with new-onset hypertension to be predicted corresponds to the category characteristic data of the second sample characteristic data.

2. The method for predicting new-onset hypertension according to claim 1, characterized in that: The first sample feature data in S1 includes: age category feature data, hip circumference category feature data, waist circumference category feature data, waist-to-hip ratio category feature data and triceps skinfold thickness category feature data; wherein, the age category feature data includes a first age value interval, the hip circumference category feature data includes a first hip circumference value interval, the waist circumference category feature data includes a first waist circumference value interval, the waist-to-hip ratio category feature data includes a first waist-to-hip ratio value interval, and the triceps skinfold thickness category feature data includes a first triceps skinfold thickness value interval.

3. The method for predicting new-onset hypertension according to claim 2, characterized in that: Decomposing the first sample feature data to obtain second sample feature data in S2 includes the following steps: Decomposing the first age value interval according to a first preset rule to obtain a plurality of different second age value intervals, each of the plurality of different second age value intervals being a category feature data in the second sample feature data; Decomposing the first hip circumference value interval according to a second preset rule to obtain a plurality of different second hip circumference value intervals, each of the plurality of different second hip circumference value intervals being a category feature data in the second sample feature data; Decomposing the first waist circumference value interval according to a third preset rule to obtain a plurality of different second waist circumference value intervals, each of the plurality of different second waist circumference value intervals being a category feature data in the second sample feature data; Decomposing the first waist-to-hip ratio value interval according to a fourth preset rule to obtain a plurality of different second waist-to-hip ratio value intervals, each of the plurality of different second waist-to-hip ratio value intervals being a category feature data in the second sample feature data; The first triceps skin fold thickness numerical interval is decomposed by the fifth preset rule to obtain a plurality of different second triceps skin fold thickness numerical intervals, and each second waist-to-hip ratio numerical interval in the plurality of different second triceps skin fold thickness numerical intervals is a category feature data in the second sample feature data.

4. The method for predicting new-onset hypertension according to claim 3, characterized in that: The second sample characteristic data also includes the following category characteristic data: province, urban and rural location, gender, education level, employment status, total personal income, marital status, medical insurance status, diabetes diagnosis, smoking status, drinking status, BMI status, tea drinking habits, daily bed time, height, weight, BMI and per capita household income; wherein the BMI represents the square of the calculation result of weight divided by height.

5. The method for predicting new-onset hypertension according to claim 1, characterized in that: The feature interaction algorithm of S3 includes a parallel additive attention mechanism and a multiplicative attention mechanism, as well as prompt tags. The feature interaction algorithm is integrated into the encoder layer of the deep learning architecture to obtain the target deep learning model.

6. The method for predicting new-onset hypertension according to claim 5, characterized in that: The deep learning architecture includes a Transformer architecture.

7. The method for predicting new-onset hypertension according to claim 6, characterized in that: The second sample feature data is segmented to obtain a test data set and a validation data set, and the test data set and the validation set are used to train and validate the target deep learning model to obtain a model for predicting new-onset hypertension.

8. The method for predicting new-onset hypertension according to claim 7, characterized in that: The model for predicting new-onset hypertension calculates the marginal contribution value of each of the category feature data in the second sample feature data through the SHAP block, and the marginal contribution value represents the degree of influence of each category feature data in the second sample feature data on the result of predicting new-onset hypertension.