Electronic devices and methods for screening features for predicting physiological states

Through the processor and storage medium in the electronic device, multiple models are used to screen out features that significantly affect the physiological state, select common features, and train the prediction model. This solves the problem in the existing technology that body mass index cannot take metabolic status into account, and achieves more accurate physiological state judgment.

CN114694838BActive Publication Date: 2025-09-23NATIONAL HEALTH RESEARCH INSTITUTE +3
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202110695287.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-12-29
Filing Date
2021-06-03
Publication Date
2025-09-23
Estimated Expiration
2041-08-26

AI Technical Summary

Technical Problem

When judging the physiological status of a subject, existing technologies rely solely on body mass index and fail to consider metabolic status, resulting in inaccurate judgments.

Method used

Through the processor and storage medium in the electronic device, multiple models are used to screen out features that significantly affect the physiological state, and common features are selected based on the relationship indicators between the features, and the prediction model is trained to improve the accuracy of judgment.

Benefits of technology

It can more accurately screen out features that are highly correlated with physiological status, assist doctors in monitoring the physiological status of subjects, and reduce the analytical burden of irrelevant metabolites.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114694838B_ABST
    Figure CN114694838B_ABST
Patent Text Reader

Abstract

The present invention provides an electronic device and method for screening features for predicting physiological states. The method includes: obtaining a plurality of physiological data corresponding to a plurality of features; generating a plurality of first subsets of the plurality of features based on the plurality of physiological data based on a first model, wherein the plurality of first subsets respectively correspond to the plurality of physiological data; selecting a first feature from a plurality of features based on the plurality of first subsets; calculating a first relationship index corresponding to a second feature of the plurality of features and the first feature, and selecting the second feature as a co-feature of the first feature based on the first relationship index; and outputting the first feature and the co-feature.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an electronic device and method for screening features for predicting physiological states. Background Art

[0002] The method of obtaining multiple metabolic indicators from a patient through a blood draw, allowing doctors to assess the patient's physiological status based on these indicators, is a growing and important medical technology. For example, to determine a patient's obesity level, existing technology uses the body mass index (BMI). However, the BMI calculation only considers the subject's weight and height, and fails to account for the subject's metabolic status.

[0003] The human body produces a wide variety of metabolites, each of which has different correlations with different physiological states. Therefore, developing a method to screen for metabolites that are highly correlated with specific physiological states would help doctors more accurately monitor the physiological status of the subject. Summary of the Invention

[0004] The present invention provides an electronic device and method for screening features for predicting physiological states, which can screen out features significantly related to a specific physiological state from multiple features.

[0005] The present invention provides an electronic device for screening features for predicting physiological states, comprising a processor, a storage medium, and a transceiver. The storage medium stores a plurality of modules. The processor couples the storage medium and the transceiver, and accesses and executes the plurality of modules, wherein the plurality of modules include a data collection module, a training module, a calculation module, and an output module. The data collection module obtains a plurality of physiological data corresponding to a plurality of features through the transceiver. The training module generates a plurality of first subsets of a plurality of features based on a plurality of physiological data based on a first model, wherein the plurality of first subsets respectively correspond to the plurality of physiological data. The calculation module selects a first feature from a plurality of features based on the plurality of first subsets, calculates a first relationship index corresponding to a second feature of the plurality of features and the first feature, and selects the second feature as a co-feature of the first feature based on the first relationship index. The output module outputs the first feature and the co-feature through the transceiver.

[0006] In one embodiment of the present invention, the above-mentioned training module generates multiple second subsets of multiple features based on multiple physiological data based on the second model, wherein the multiple second subsets correspond to the multiple physiological data respectively; and the operation module selects the first feature from the multiple features based on the multiple first subsets and the multiple second subsets.

[0007] In one embodiment of the present invention, the computing module calculates a first number of first features in the plurality of first subsets and calculates a second number of first features in the plurality of second subsets; and the computing module selects the first feature from the plurality of features according to the first number and the second number.

[0008] In one embodiment of the present invention, the above-mentioned operation module calculates a first score of the first feature based on the first quantity, the first weight corresponding to the first model, the second quantity, and the second weight corresponding to the second model; and the operation module selects the first feature from the multiple features in response to the first score being greater than a first threshold.

[0009] In one embodiment of the present invention, the above-mentioned operation module calculates a first score of the first feature based on the first quantity, the first weight corresponding to the first model, the second quantity and the second weight corresponding to the second model; the operation module calculates a third quantity of the third features in multiple first subsets, and calculates a fourth quantity of the third features in multiple second subsets; the operation module calculates a second score of the third feature based on the third quantity, the first weight, the fourth quantity and the second weight; and the operation module selects the first feature from the first feature and the third feature in response to the first score being greater than the second score.

[0010] In one embodiment of the present invention, the above-mentioned operation module obtains a first number of first features in each of the multiple first subsets to generate a first vector; the operation module obtains a second number of second features in each of the multiple first subsets to generate a second vector; and the operation module calculates a first relationship index based on the first vector and the second vector.

[0011] In one embodiment of the present invention, the operation module selects the second feature as a common feature of the first feature in response to the first relationship index being greater than a second threshold.

[0012] In one embodiment of the present invention, the computing module calculates a second relationship index corresponding to a third feature and a fourth feature among the plurality of features; and the computing module selects the second feature as a companion feature of the first feature in response to the first relationship index being greater than the second relationship index.

[0013] In one embodiment of the present invention, the above-mentioned training module trains at least one first prediction model of the physiological state based on multiple physiological data, the first feature and the shared feature, and calculates at least one first performance indicator corresponding to the at least one first prediction model; the training module randomly selects a third feature and a fourth feature from multiple features, wherein any one of the third feature and the fourth feature is different from any one of the first feature and the second feature; the training module trains at least one second prediction model of the physiological state based on multiple physiological data, the third feature and the fourth feature, and calculates at least one second performance indicator of the at least one second prediction model; the operation module determines that the first feature and the shared feature are available in response to the at least one first performance indicator being greater than the at least one second performance indicator; and the output module outputs the first feature and the shared feature in response to the first feature and the shared feature being available.

[0014] In one embodiment of the present invention, the aforementioned multiple features correspond to multiple metabolites of the human body.

[0015] In one embodiment of the present invention, the above-mentioned data collection module receives a physiological data set through a transceiver, and divides the physiological data set into a plurality of training data and a plurality of test data corresponding to the plurality of physiological data according to a self-help re-extraction method; the training module generates a plurality of first subsets based on the plurality of training data; the training module generates at least one first prediction model based on the plurality of training data; and the training module calculates at least one first performance indicator based on the plurality of test data.

[0016] In one embodiment of the present invention, the first model or the second model is associated with one of the following: a random forest algorithm, a logistic regression, and a support vector machine.

[0017] In one embodiment of the present invention, the first model generates the plurality of first subsets based on one of the following: stepwise selection and feature importance.

[0018] A method of screening features for predicting physiological states according to the present invention includes: obtaining multiple physiological data corresponding to multiple features; generating multiple first subsets of multiple features based on the multiple physiological data based on a first model, wherein the multiple first subsets respectively correspond to the multiple physiological data; selecting a first feature from multiple features based on the multiple first subsets, calculating a first relationship index corresponding to a second feature in the multiple features and the first feature, and selecting the second feature as a co-feature of the first feature based on the first relationship index; and outputting the first feature and the co-feature.

[0019] Based on the above, the present invention can select features that significantly influence the predicted physiological state of the subject and select co-features corresponding to these features. The present invention can output these features and co-features for the user's reference. For example, if the user is a doctor, they can simply refer to the metabolites and co-metabolites output by the present invention to determine the subject's obesity level, without having to waste time analyzing metabolites that are irrelevant to obesity level. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] Figure 1 A schematic diagram illustrating an electronic device for screening features for predicting physiological states according to an embodiment of the present invention is shown;

[0021] Figure 2 A flowchart showing a method for screening features for predicting physiological states according to an embodiment of the present invention;

[0022] Figure 3 A flowchart showing a method for determining whether selected features and co-features are available according to an embodiment of the present invention;

[0023] Figure 4 According to another embodiment of the present invention, a flow chart of a method for screening features for predicting a physiological state is shown.

[0024] Description of Reference Numerals

[0025] 100: Electronic devices

[0026] 110: Processor

[0027] 120: Storage media

[0028] 121: Data collection module

[0029] 122: Training Module

[0030] 123: Operation module

[0031] 124: Output module

[0032] 130: Transceiver

[0033] S201, S202, S203, S204, S205, S206, S207, S208, S209, S210, S301, S302, S303, S304, S305, S306, S307, S308, S401, S402, S403, S404: Steps DETAILED DESCRIPTION

[0034] Reference will now be made in detail to exemplary embodiments of the present invention, examples of which are illustrated in the accompanying drawings. Whenever possible, the same reference numerals are used in the drawings and the description to refer to the same or like parts.

[0035] Figure 1 A schematic diagram of an electronic device 100 for screening features for predicting physiological states according to an embodiment of the present invention is shown. The electronic device 100 may include a processor 110 , a storage medium 120 , and a transceiver 130 .

[0036] The processor 110 may be, for example, a central processing unit (CPU), or other programmable general-purpose or special-purpose microcontroller unit (MCU), microprocessor, digital signal processor (DSP), programmable controller, application-specific integrated circuit (ASIC), graphics processing unit (GPU), image signal processor (ISP), image processing unit (IPU), arithmetic logic unit (ALU), complex programmable logic device (CPLD), field programmable gate array (FPGA), or other similar components or combinations thereof. The processor 110 may be coupled to the storage medium 120 and the transceiver 130 and access and execute multiple modules and various applications stored in the storage medium 120.

[0037] The storage medium 120 may be, for example, any type of fixed or removable random access memory (RAM), read-only memory (ROM), flash memory, hard disk drive (HDD), solid state drive (SSD), or similar device, or a combination thereof, and is used to store multiple modules or various application programs executable by the processor 110. In this embodiment, the storage medium 120 may store multiple modules including a data collection module 121, a training module 122, a calculation module 123, and an output module 124, the functions of which will be described later.

[0038] The transceiver 130 transmits and receives signals wirelessly or by wire. The transceiver 130 may also perform operations such as low noise amplification, impedance matching, frequency mixing, up or down frequency conversion, filtering, amplification, and the like.

[0039] Figure 2 According to one embodiment of the present invention, a flow chart of a method for screening features for predicting physiological states is shown, wherein the method can be performed as follows: Figure 1 The electronic device 100 is shown as an implementation.

[0040] In step S201, the data collection module 121 receives a physiological data set of the subject through the transceiver 130. The physiological data set may include K feature values ​​corresponding to K features, wherein the K features (features f1, f2, ..., f K ) can correspond to K metabolites in the human body. K can be any positive integer.

[0041] In step S202, the data collection module 121 may divide the physiological data set into N pieces of physiological data corresponding to K features, where N can be any positive integer. Specifically, the data collection module 121 may divide the physiological data set into N pieces of physiological data based on a bootstrap method, where each of the N pieces of physiological data may include training data and test data. In other words, the data collection module 121 may divide the physiological data set into N pieces of training data and N pieces of test data corresponding to the N pieces of physiological data, respectively.

[0042] In one embodiment, the data collection module 121 may separate the physiological data set into individual physiological data corresponding to K features, where the physiological data may include training data and test data. Specifically, the data collection module 121 may separate the physiological data set into training data and test data based on K-fold cross-validation to generate the physiological data.

[0043] In step S203, the training module 122 can generate N subsets SB1 of K features based on the N training data based on the first model, wherein the N subsets SB1 can correspond to the N training data respectively. The first model can be configured to select one or more features that significantly affect a specific physiological state (e.g., obesity level) from the K features based on a training data, wherein the one or more features are the subset SB1 of the K features. Accordingly, the N training data can generate N subsets SB1 of K features, which are the subsets

[0044] Similarly, the training module 122 can generate N subsets SB2 of K features based on N training data based on the second model, wherein the N subsets SB2 can correspond to the N training data respectively. The second model can be configured to select one or more features that significantly affect a specific physiological state (e.g., obesity level) from the K features based on one training data, wherein the one or more features are the subset SB1 of the K features. Accordingly, the N training data can generate N subsets SB2 of K features, which are the subsets

[0045] Similarly, the training module 122 can generate N subsets SB3 of K features based on N training data based on the third model, wherein the N subsets SB3 can correspond to the N training data respectively. The third model can be configured to select one or more features that significantly affect a specific physiological state (e.g., obesity level) from the K features based on one training data, wherein the one or more features are the subset SB3 of the K features. Accordingly, the N training data can generate N subsets SB3 of K features, namely, subsets

[0046] The number M of models used in step S203 can be defined by the user based on their needs. Although M is equal to 3 in this embodiment (i.e., three models, namely, the first model, the second model, and the third model), the present invention is not limited thereto. For example, M can be any positive integer greater than 1.

[0047] In one embodiment, the first model (or the second model, the third model) may correspond to a random forest algorithm (RF), logistic regression or support vector machine (SVM), but the present invention is not limited thereto. The first model (or the second model, the third model) may, for example, use a stepwise selection method or feature importance to select one or more features that significantly affect a specific physiological state from K features. For example, the first model may use a stepwise selection method based on P-value or Akaike information criterion (AIC) to select one or more features from K features. Different models may use the same or different algorithms. For example, the first model, the second model and the third model may use the same or different algorithms.

[0048] In step S204, the operation module 123 may obtain N subsets SB1 corresponding to the first model (respectively, the subsets ), corresponding to the N subsets SB2 of the second model (respectively, subsets ) and N subsets SB3 corresponding to the third model (respectively subsets ).

[0049] In step S205, the operation module 123 may calculate the number of N subsets SB according to the number of m Calculate the score for each of the K features, where m is the index of the model. For example, m=1 corresponds to the first model, m=2 corresponds to the second model, and m=3 corresponds to the third model.

[0050] Assume that the operation module 123 wants to calculate the feature f among the K features j The score Z j , the operation module 123 can calculate N subsets SB according to formula (1) m The feature f in j Number in For N subsets SB m The i-th subset SB in m The feature f in j The number of ( can be 0 or 1).

[0051]

[0052] For example, assuming that the total number of physiological data is 3 (N=3) and the total number of features is 5 (K=5), the operation module 123 can generate Table 1, Table 2 and Table 3 as shown below according to formula (1), where Table 1 corresponds to the first model, Table 2 corresponds to the second model and Table 3 corresponds to the third model. Taking feature f1 in Table 1 as an example, the number corresponding to feature f1 is

[0053] Table 1

[0054]

[0055] Table 2

[0056]

[0057] Table 3

[0058]

[0059] After obtaining N subsets SB m The feature f in j Number After that, the operation module 123 can calculate the corresponding quantity according to formula (2). Ratio in For N subsets SB m And the feature f j The corresponding quantity.

[0060]

[0061] For example, assuming N=3, the operation module 123 can generate Table 4 shown below according to Table 1, Table 2, and Table 3 based on formula (2), where the number and ratio Corresponding to the first model, the number and ratio corresponds to the second model, and the number and ratio Corresponding to the third model.

[0062] Table 4

[0063]

[0064] In obtaining the ratio Then, the operation module 123 can calculate the corresponding feature f according to formula (3). j The score Z j , where w m is the weight corresponding to the mth model, and For the characteristic fj The ratio corresponding to the mth model. For example, the weights corresponding to the first model, the second model, and the third model may be w1 = 0, w2 = 0, and w3 = 1, respectively. For another example, the weights corresponding to the first model, the second model, and the third model may be w1 = 1 / 3, w2 = 1 / 3, and w3 = 1 / 3, respectively.

[0065]

[0066] For example, assuming that weight w1=1 / 3, weight w2=1 / 3, and weight w3=1 / 3, the operation module 123 can generate Table 5 shown below according to Table 4 based on formula (3).

[0067] Table 5

[0068]

[0069] In obtaining the feature f among the K features j The corresponding score Z j Then, in step S206, the operation module 123 can calculate the value of the threshold m1 and the score Z j Determine feature f j Whether it is a selected feature.

[0070] In one embodiment, the threshold m1 may be associated with the feature f j Score ranking among the K features. For example, threshold m1 may indicate that the features with the highest scores among the K features are selected as the features. Taking Table 5 as an example, operation module 123 may select feature f1 with the highest score from features f1 to f5 based on threshold m1 as the selected feature. In other words, operation module 123 may select feature f1 corresponding to score Z1 from features f1 to f5 as the selected feature in response to score Z1 being greater than scores Z2, Z3, Z4, and Z5.

[0071] In one embodiment, the operation module 123 may respond to the score Z j Exceeding the threshold m1 and feature f j Taking Table 5 as an example, assuming that the threshold m1 is equal to 5 / 9, the operation module 123 may select the feature f1 as the selected feature in response to the score Z1 of the feature f1 being greater than 5 / 9.

[0072] In step S207, the calculation module 123 may calculate a relation index between each of the K features and the other features. Specifically, the calculation module 123 may obtain K vectors corresponding to the K features, and select two vectors from the K vectors to calculate the relation index between the two vectors.

[0073] If the operation module 123 wants to calculate the feature f among the K features A and feature f B The relationship index between the N subsets SB m The feature f in A The number of produces a vector as shown in formula (4) And according to N subsets SB m The feature f in B The number of produces a vector as shown in formula (5) in For N subsets SB m The i-th subset SB in m The feature f in A the number of, and For N subsets SB m The i-th subset SB in m The feature f in B Then, the operation module 123 can calculate the number of vectors. and vectors Calculate feature f A and feature f B The relationship indicator between the m-th model and the m-th model is a Pearson correlation coefficient (PCC), but the present invention is not limited thereto.

[0074]

[0075]

[0076] For example, the operation module 123 can generate Table 6 shown below based on Tables 1, 2, and 3 based on Formulas (4) and (5). Table 6 includes M*K vectors corresponding to K features (K=5) and M models (M=3). Each vector can include N elements (N=3) corresponding to N pieces of training data.

[0077] Table 6

[0078]

[0079] The operation module 123 can calculate the relationship index corresponding to the at least two features. For example, if the at least two features only include two features, the operation module 123 can calculate the two features (for example, feature f A and f B ) between the relationship index RI(f A,f B ), where C(V x ,V y ) is the vector V x and vector V y The correlation coefficient of W m For another example, if the at least two features are more than two features, the operation module 123 can calculate the P value of the at least two features based on the analysis of variance (ANOVA) test as the relationship indicator.

[0080]

[0081] In one embodiment, the weights corresponding to the first model, the second model, and the third model may be weight W1 = 0, weight W2 = 0, and weight W3 = 1, respectively. In one embodiment, the weights corresponding to the first model, the second model, and the third model may be weight W1 = 1 / 3, weight W2 = 1 / 3, and weight W3 = 1 / 3, respectively. For example, if weight W1 = 1 / 3, weight W2 = 1 / 3, and weight W3 = 1 / 3, then the operation module 123 may generate Table 7 based on the vectors in Table 6 based on formula (6).

[0082] Table 7

[0083]

[0084] In step S208, the operation module 123 may calculate the relationship index RI (f A ,f B ) determines the feature f A and feature f B Whether it is a selected feature pair.

[0085] In one embodiment, the threshold m2 may be associated with a feature pair in For example, the threshold m2 may indicate that The feature pairs with the top few highest relationship indices among the feature pairs are selected as the selected feature pairs. Taking Table 7 as an example, the operation module 123 can select the feature pair (f1, f2) with the highest relationship index from the 10 feature pairs according to the threshold m2 as the selected feature pair. In other words, the operation module 123 can select the feature pair (f1, f2) from the 10 feature pairs in Table 7 as the selected feature pair in response to the relationship index RI(f1, f2) being greater than the relationship indices RI(f1, f3), RI(f1, f4), RI(f1, f5), RI(f2, f3), RI(f2, f4), RI(f2, f5), RI(f3, f4), RI(f3, f5), and RI(f4, f5).

[0086] In one embodiment, the operation module 123 may respond to the relationship index RI (f A ,f B ) exceeds the threshold m2 and the feature pair (f A ,f B ) is selected as the selected feature pair. Taking Table 7 as an example, assuming that the threshold m2 is equal to 1 / 4, the operation module 123 can select the feature pair (f1, f2) as the selected feature pair in response to the relationship index RI (f1, f2) being greater than 1 / 4, but the selected feature pair is not limited to one pair.

[0087] After executing step S206 and step S208 to obtain the selected feature and the selected feature pair respectively, in step S209, the operation module 123 may obtain the feature corresponding to the selected feature from the selected feature pair as an accompanying feature. A is the selected feature and the feature pair (f A ,f B ) is the selected feature pair, the operation module 123 can be A ,f B ) and feature f A The corresponding feature f B As feature f A The accompanying characteristics.

[0088] In one embodiment, the common feature may be selected by professionals based on experience from K features to select features corresponding to the selected features as the common feature.

[0089] In step S210, the output module 124 may output the selected features and the shared features via the transceiver 130. In one embodiment, the operation module 123 may determine whether the selected features and the shared features are available. If the selected features and the shared features are available, the output module 124 may output the selected features and the shared features. If the selected features and the shared features are unavailable, the output module 124 may not output the selected features and the shared features. Figure 3 A flowchart of a method for determining whether selected features and co-features are available is shown according to an embodiment of the present invention.

[0090] In step S301 , the operation module 123 obtains a selected feature and a co-feature corresponding to the selected feature.

[0091] In step S302, the training module 122 may extract portions corresponding to the selected features and co-features from the N training data sets to train at least one first prediction model for predicting physiological states. The at least one first prediction model may correspond to a random forest algorithm, logistic regression, or a support vector machine, but the present invention is not limited thereto.

[0092] In step S303, the training module 122 may obtain portions corresponding to the selected features and the co-features from the N test data sets to calculate at least one first performance metric corresponding to at least one first prediction model. The at least one first performance metric may correspond to a parameter in a confusion matrix, such as accuracy (ACC), precision, recall rate, false positives (FP), or F1 score.

[0093] In step S304, the training module 122 may select two random features from the K features, wherein any one of the two random features is different from any one of the selected features and the shared features. Next, the training module 122 may obtain a portion corresponding to the two random features from the N training data to train at least one second prediction model for predicting the physiological state. The at least one second prediction model may correspond to a random forest algorithm, logistic regression, or a support vector machine, but the present invention is not limited thereto. In one embodiment, the training module 122 may select a plurality of random features corresponding to the number of selected features and shared features from the K features to train the at least one second prediction model. For example, if the total number of selected features and shared features obtained by the operation module 123 in step S301 is 4, the training module 122 may select 4 random features from the K features to train the at least one second prediction model.

[0094] In step S305, the training module 122 may extract portions corresponding to the two random features (or a number of random features corresponding to the number of selected features and co-occurring features) from the N test data sets to calculate at least one second performance metric corresponding to at least one second prediction model. The at least one second performance metric may correspond to a parameter in a confusion matrix, such as accuracy, precision, recall, false positives, or an F1 score.

[0095] In step S306, the computing module 123 determines whether the at least one first performance indicator is greater than the at least one second performance indicator. If the at least one first performance indicator is greater than the at least one second performance indicator, the process proceeds to step S307. If the at least one first performance indicator is less than or equal to the at least one second performance indicator, the process proceeds to step S308.

[0096] In step S307, the operation module 123 may determine that the selected feature and the shared feature are available. In step S308, the operation module 123 may determine that the selected feature and the shared feature are unavailable.

[0097] Figure 4According to another embodiment of the present invention, a flow chart of a method for screening features for predicting physiological states is shown, wherein the method can be performed as follows: Figure 1 The electronic device 100 shown is implemented. In step S401, multiple physiological data corresponding to multiple features are obtained. In step S402, multiple first subsets of the multiple features are generated based on the multiple physiological data based on a first model, wherein the multiple first subsets respectively correspond to the multiple physiological data. In step S403, a first feature is selected from the multiple features based on the multiple first subsets, a first relationship index corresponding to a second feature in the multiple features and the first feature is calculated, and the second feature is selected as a co-feature of the first feature based on the first relationship index. In step S404, the first feature and the co-feature are output.

[0098] In summary, the present invention can utilize a variety of different models to select features that significantly influence the prediction results of physiological characteristics from multiple features, and can select co-features corresponding to the features based on relationship indicators between the features and other features. Co-features selected in this way can also significantly influence the prediction results of physiological characteristics. After obtaining at least one feature and at least one corresponding co-feature, the present invention can train a prediction model based on the at least one feature and at least one co-feature, and calculate an efficiency index of the prediction model. If the efficiency index shows that the at least one feature and at least one co-feature can significantly influence the prediction model's prediction results for the physiological characteristics, the present invention can output the at least one feature and at least one co-feature for user reference.

[0099] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. An electronic device for screening features for predicting physiological states, characterized in that: include: transceiver; Storage medium, storing multiple modules; as well as A processor is coupled to the storage medium and the transceiver, and accesses and executes the plurality of modules, wherein the plurality of modules include: a data collection module, for acquiring a plurality of physiological data corresponding to a plurality of characteristics through the transceiver; a training module configured to generate a plurality of first subsets of the plurality of features from the plurality of physiological data based on a first model, wherein the plurality of first subsets respectively correspond to the plurality of physiological data, and the training module to generate a plurality of second subsets of the plurality of features from the plurality of physiological data based on a second model, wherein the plurality of second subsets respectively correspond to the plurality of physiological data, wherein the second model is different from the first model; a calculation module that calculates the number of each feature in the plurality of first subsets to obtain a first number of first features, calculates the number of each feature in the plurality of second subsets to obtain a second number of the first features, selects the first feature from the plurality of features based on the first number and the second number, calculates a relationship index between each feature other than the first feature in the plurality of features and the first feature to obtain a first relationship index corresponding to a second feature in the plurality of features and the first feature, and selects the second feature as a companion feature of the first feature based on the first relationship index, wherein the relationship index includes a correlation coefficient; and An output module outputs the first feature and the common feature through the transceiver.

2. The electronic device according to claim 1, wherein The operation module calculates a plurality of scores corresponding to the plurality of features, including: calculating a first score for the first feature based on the first quantity, a first weight corresponding to the first model, the second quantity, and a second weight corresponding to the second model; as well as The operation module selects the first feature from the plurality of features in response to the first score being greater than a first threshold.

3. The electronic device according to claim 1, wherein The operation module calculates a plurality of scores corresponding to the plurality of features, including: calculating a first score for the first feature based on the first quantity, a first weight corresponding to the first model, the second quantity, and a second weight corresponding to the second model; The operation module calculates a third number of third features in the plurality of first subsets, and calculates a fourth number of the third features in the plurality of second subsets; The calculation module calculates a second score of the third feature according to the third number, the first weight, the fourth number, and the second weight; and The operation module selects the first feature from the first feature and the third feature in response to the first score being greater than the second score.

4. The electronic device according to claim 1, wherein The operation module obtains a third quantity of the first feature in each of the plurality of first subsets to obtain a plurality of third quantities respectively corresponding to the plurality of first subsets, and generates a first vector according to the plurality of third quantities; The operation module obtains a fourth quantity of the second feature in each of the plurality of first subsets to obtain a plurality of fourth quantities respectively corresponding to the plurality of first subsets, and generates a second vector according to the plurality of fourth quantities; and The operation module calculates the first relationship index according to the first vector and the second vector.

5. The electronic device according to claim 1, wherein The operation module selects the second feature as the companion feature of the first feature in response to the first relationship index being greater than a second threshold. The electronic device according to claim 1 , wherein The operation module calculates a second relationship index corresponding to a third feature and a fourth feature among the plurality of features; and The operation module selects the second feature as the companion feature of the first feature in response to the first relationship index being greater than the second relationship index.

7. The electronic device according to claim 1, wherein The training module trains at least one first prediction model of the physiological state according to the plurality of physiological data, the first feature, and the shared feature, and calculates at least one first performance indicator corresponding to the at least one first prediction model; The training module randomly selects a third feature and a fourth feature from the plurality of features, wherein any one of the third feature and the fourth feature is different from any one of the first feature and the second feature; The training module trains at least one second prediction model of the physiological state according to the plurality of physiological data, the third feature, and the fourth feature, and calculates at least one second performance indicator of the at least one second prediction model; The computing module determines that the first feature and the companion feature are available in response to the at least one first performance indicator being greater than the at least one second performance indicator; and The output module outputs the first feature and the companion feature in response to the first feature and the companion feature being available. The electronic device according to claim 1 , wherein the plurality of features respectively correspond to a plurality of metabolites of a human body.

9. The electronic device according to claim 7, wherein The data collection module receives a physiological data set through the transceiver, and divides the physiological data set into a plurality of training data and a plurality of test data respectively corresponding to the plurality of physiological data according to a self-help re-extraction method; The training module generates the plurality of first subsets according to the plurality of training data; The training module generates the at least one first prediction model according to the plurality of training data; as well as The training module calculates the at least one first performance indicator according to the plurality of test data.

10. The electronic device of claim 1, wherein the first model or the second model is associated with one of the following: Random forest algorithms, logistic regression, and support vector machines.

11. The electronic device of claim 10, wherein the first model generates the plurality of first subsets based on one of the following: Stepwise selection and feature importance.

12. A method for screening features for predicting physiological states, characterized in that: include: obtaining a plurality of physiological data corresponding to a plurality of characteristics; generating a plurality of first subsets of the plurality of features according to the plurality of physiological data based on a first model, wherein the plurality of first subsets respectively correspond to the plurality of physiological data; generating a plurality of second subsets of the plurality of features from the plurality of physiological data based on a second model, wherein the plurality of second subsets respectively correspond to the plurality of physiological data, wherein the second model is different from the first model; Calculating the number of each feature in the plurality of first subsets to obtain a first number of first features, calculating the number of each feature in the plurality of second subsets to obtain a second number of the first feature, selecting the first feature from the plurality of features based on the first number and the second number, calculating a relationship index between each feature other than the first feature in the plurality of features and the first feature to obtain a first relationship index corresponding to a second feature in the plurality of features and the first feature, and selecting the second feature as a companion feature of the first feature based on the first relationship index, wherein the relationship index comprises a correlation coefficient; and The first feature and the co-feature are output.

Citation Information

Patent Citations

  • Disease-associated microbiome characterization process

    WO2019036176A1