Risk behavior identification method, system and device based on decision tree and medium

Through the decision tree-based risk behavior identification method, using multi-source data for feature analysis and risk identification, the problems of less utilization of mental health assessment data and low accuracy of risk identification in the existing technology are solved, and a more comprehensive and accurate mental health risk assessment is achieved.

CN120048511APending Publication Date: 2025-05-27SOUTH CHINA NORMAL UNIV
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510009241.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-03
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

Existing assessment of mental health problems relies on clinician experience and patient self-report, with less data utilization and less comprehensiveness and accuracy of risk identification.

Method used

The risk behavior identification method based on the decision tree is adopted, and the characteristic data set is determined by obtaining multi-source data for preprocessing and analysis, and classified and identified according to the built risk identification model to obtain risk levels and/or risk categories.

Benefits of technology

The comprehensiveness and accuracy of risk identification are improved, and more accurate mental health risk assessment and monitoring are achieved through the integration of multi-source data and the extraction of features.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120048511A_ABST
    Figure CN120048511A_ABST
Patent Text Reader

Abstract

The invention discloses a risk behavior identification method and system based on a decision tree, electronic equipment and a storage medium, and the method comprises the steps: collecting multi-source data, carrying out the preprocessing, and integrating the data into a first data set; analyzing the first data set, determining a feature data set, and classifying and identifying the feature data set through a constructed risk identification model to obtain a risk type or a risk level corresponding to the multi-source data; by collecting multi-source data and integrating data of different data sources for risk identification, the data utilization rate is improved, and the comprehensiveness of an evaluation result is further improved; on the basis of the characteristics of the multi-source data, the multi-source data is classified and recognized by constructing the risk recognition model, psychological health changes reflected by the multi-source data can be accurately recognized and monitored, and then the assessment accuracy is improved; the embodiment of the invention can be widely applied to the technical field of mental health.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of testing technologies, and in particular, to a risk behavior recognition method, system, electronic device, and storage medium based on a decision tree. Background Art

[0002] With the development of society and the intensification of competition, mental health problems have become increasingly important. Conducting mental health assessments, promptly detecting mental health problems, accurately quantifying the risk levels of mental health problems, and then performing corresponding interventions or psychological counseling can greatly reduce the occurrence of personal safety problems. Existing mental health problem assessments mainly rely on the experience of clinicians and the self-report of patients, with less data utilization, low comprehensiveness of risk identification, and low accuracy. Summary of the Invention

[0003] The main objective of the embodiments of the present invention is to provide a risk behavior recognition method, system, electronic device, and storage medium based on a decision tree, which can improve the comprehensiveness of risk identification and the accuracy of risk identification.

[0004] To achieve the above objective, on the one hand, an embodiment of the present invention provides a risk behavior recognition method based on a decision tree, and the method includes:

[0005] Obtain multi-source data, preprocess the multi-source data to obtain a first data set;

[0006] Analyze the first data set to determine a feature data set;

[0007] Perform classification recognition according to the constructed risk recognition model and the feature data set to obtain a recognition result; wherein, the recognition result includes a risk level and / or a risk category; wherein, the risk recognition model includes a plurality of leaf nodes and non-leaf nodes.

[0008] In some embodiments, the performing classification according to the constructed risk recognition model and the first data set to obtain a recognition result specifically includes:

[0009] Analyze the first data set to determine a feature data set; calculate a first parameter value according to the constructed risk recognition model and the feature data set;

[0010] Divide the feature data set according to the first parameter value and a first preset threshold to obtain a plurality of feature data subsets, and determine the node types corresponding to the plurality of feature data subsets;

[0011] If the node type is a leaf node, determine the recognition result according to the node information corresponding to the feature data subset;

[0012] If the node type is a non - leaf node, use the subset of feature data as the feature data set, and return to perform calculations using the constructed risk identification model and the feature data set to obtain a first parameter value. Continue this process until the node type is a leaf node, and then determine the identification result based on the node information corresponding to the subset of feature data.

[0013] In some embodiments, the risk identification model is determined in the following manner:

[0014] Obtain a sample data set, analyze the sample data set to obtain a sample feature set; where the sample data set includes any combination of electronic health records, psychological assessment results, social media behavior data, or physiological monitoring data;

[0015] Perform calculations on the sample feature set to determine a third parameter value, and determine the target sample feature based on the third parameter value;

[0016] Segment the sample feature set according to the target sample feature to determine a number of sample feature subsets and construct a tree - like structure;

[0017] Perform calculations on a number of the sample feature subsets respectively to determine a number of fourth parameter values, and determine the risk identification model based on the number of fourth parameter values and a third preset threshold; where the fourth parameter value includes node purity or decision depth.

[0018] In some embodiments, the determining the risk identification model based on the number of fourth parameter values and the third preset threshold specifically includes:

[0019] If the number of fourth parameter values is greater than or equal to the third preset threshold, use the tree - like structure as the risk identification model;

[0020] If there is a fourth parameter value less than the third preset threshold, use the sample feature subset corresponding to the fourth parameter value as the sample feature set, and return to perform calculations according to a first preset formula and the sample feature set to determine a third parameter value. Continue this process until the number of fourth parameter values is greater than or equal to the third preset threshold, and then use the tree - like structure as the risk identification model.

[0021] In some embodiments, the performing calculations on the sample feature set to determine a third parameter value specifically includes:

[0022] Calculate the first information entropy of the sample feature set, and segment the sample feature set according to the sample features of the sample feature set to determine a number of sample subsets;

[0023] Calculate the second information entropy and the first ratio corresponding to a plurality of the sample subsets respectively, and calculate the sum of the products of the second information entropy and the first ratio to determine a fifth parameter; wherein, the first ratio represents the proportion of the sample subset in the sample feature set;

[0024] Calculate the difference between the first information entropy and the fifth parameter as the value of the third parameter.

[0025] In some embodiments, the method further includes:

[0026] Determine a risk level according to the recognition result, and compare the risk level with a preset warning threshold;

[0027] If the risk level is greater than or equal to the preset warning threshold, generate a warning message and generate an intervention plan according to the recognition result.

[0028] In some embodiments, the method further includes:

[0029] Integrate historical patient data and the first data set according to a preset period to determine a second data set;

[0030] Construct a classification model according to the second data set, and perform cross-validation on the classification model to determine a first model performance value;

[0031] Determine a second model performance value according to the risk recognition model, and compare the first model performance value with the second model performance value;

[0032] If the first model performance value is greater than the second model performance value, update the risk recognition model according to the classification model.

[0033] To achieve the above object, another aspect of the embodiments of the present invention provides a risk behavior recognition system based on a decision tree, including:

[0034] A first module, configured to obtain multi-source data, preprocess the multi-source data to obtain a first data set;

[0035] A second module, configured to analyze the first data set to determine a feature data set;

[0036] A third module, configured to perform classification recognition according to the constructed risk recognition model and the feature data set to obtain a recognition result; wherein, the recognition result includes a risk level and / or a risk category; wherein, the risk recognition model includes a plurality of leaf nodes and non-leaf nodes.

[0037] To achieve the above object, on the other hand, an embodiment of the present application provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the foregoing method is implemented.

[0038] To achieve the above object, on the other hand, an embodiment of the present application provides a computer-readable storage medium, which stores a computer program, and when the computer program is executed by a processor, the foregoing method is implemented.

[0039] Implementing the embodiments of the present invention includes the following beneficial effects: The present embodiment provides a risk behavior recognition method, system, electronic device, and storage medium based on a decision tree. This solution collects multi-source data and performs preprocessing to integrate it into a first data set; analyzes the first data set to determine a feature data set, and classifies and recognizes the feature data set through a constructed risk recognition model to obtain the risk type or risk level corresponding to the multi-source data; by collecting multi-source data and integrating data from different data sources for risk recognition, the data utilization rate is improved, and thus the comprehensiveness of the risk recognition result is improved; based on the characteristics of multi-source data, by constructing a risk recognition model to classify and recognize multi-source data, the mental health changes reflected by the multi-source data can be accurately recognized and monitored, and thus the accuracy of risk recognition is improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] Figure 1 is a schematic flowchart of the steps of a risk behavior recognition method based on a decision tree provided by an embodiment of the present invention;

[0041] Figure 2 is a schematic flowchart of the steps of generating an intervention plan in a risk behavior recognition method based on a decision tree provided by an embodiment of the present invention;

[0042] Figure 3 is a schematic flowchart of the steps of regularly updating a risk recognition model in a risk behavior recognition method based on a decision tree provided by an embodiment of the present invention;

[0043] Figure 4 is a schematic flowchart of the steps of a risk recognition model determining a recognition result in a risk behavior recognition method based on a decision tree provided by an embodiment of the present invention;

[0044] Figure 5 is a schematic flowchart of the steps of constructing a risk recognition model in a risk behavior recognition method based on a decision tree provided by an embodiment of the present invention;

[0045] Figure 6 is a schematic flowchart of the steps of determining whether the construction of a risk recognition model is completed in a risk behavior recognition method based on a decision tree provided by an embodiment of the present invention;

[0046] Figure 7 It is a schematic flowchart of the steps for calculating the third parameter value in a risk behavior recognition method based on a decision tree provided by an embodiment of the present invention;

[0047] Figure 8 It is a schematic flowchart of the steps of a specific embodiment provided by an embodiment of the present invention;

[0048] Figure 9 It is a structural block diagram of a risk behavior recognition model system based on a decision tree provided by an embodiment of the present invention;

[0049] Figure 10 It is a schematic diagram of the hardware structure of an electronic device provided by an embodiment of the present invention. Detailed implementation manners

[0050] The following further elaborates the present invention in detail with reference to the accompanying drawings and specific embodiments. For the step numbers in the following embodiments, they are only set for the convenience of description and explanation, and no limitation is imposed on the order between the steps. The execution order of each step in the embodiments can be adaptively adjusted according to the understanding of those skilled in the art.

[0051] In the following description, "some embodiments" are involved, which describe a subset of all possible embodiments. However, it can be understood that "some embodiments" can be the same subset or different subsets of all possible embodiments, and can be combined with each other without conflict.

[0052] In the following description, the terms "first / second / third" only distinguish similar objects and do not represent a specific order for the objects. It can be understood that "first / second / third" can be interchanged with a specific order or sequence when allowed, so that the embodiments of the present invention described here can be implemented in an order other than that illustrated or described here.

[0053] Unless otherwise defined, all technical and scientific terms used in the embodiments of the present invention have the same meaning as commonly understood by those skilled in the technical field to which the present invention belongs. The terms used in the embodiments of the present invention are only for the purpose of describing the embodiments of the present invention and are not intended to limit the present invention.

[0054] Figure 1 It is an optional flowchart of a risk behavior recognition method based on a decision tree provided by an embodiment of the present application, Figure 1 The method in may include but is not limited to steps S101 to S103.

[0055] Step S101, obtain multi-source data, preprocess the multi-source data to obtain a first data set;

[0056] Step S102: Analyze the first data set to determine the feature data set.

[0057] Step S103: Perform classification and identification based on the constructed risk identification model and the feature data set to obtain an identification result. The identification result includes a risk level and / or a risk category.

[0058] In steps S101 to S103 illustrated in the embodiments of the present invention, the system obtains relevant data of a patient from multiple different data sources through an interface or a data collection tool. Exemplarily, the health examination records of the patient, psychological evaluation results, social media behavior data, etc. The system preprocesses the collected multi-source data, including data cleaning, handling missing values, etc. At the same time, variable transformation is performed on different data, continuous data is standardized or normalized to unify different data to the same scale; categorical variables are converted into numerical forms. For example, gender is represented by numerical values 0 and 1. In this embodiment, the system uses one-hot encoding or label encoding to perform variable transformation on the obtained multi-source data. The system integrates the preprocessed multi-source data into a data set, analyzes the data set, extracts variables in the data set as features. Exemplarily, the psychological state variable in the multi-source data is used as feature 1, and the life background is used as feature 2, etc. The system integrates the extracted feature data into a set as the feature data set. The system inputs the feature data set into the constructed risk identification model for identification and classification, and outputs the risk type and / or risk level of the current patient. In this embodiment, a decision tree algorithm is used to construct the risk identification model. The risk identification model divides the feature data set according to the feature information in the input feature data set to obtain at least two subsets, and gradually and deeply mines the obtained subsets. By performing segmentation to a certain depth through the risk identification model, feature labels or subsets with highly overlapping information are obtained, thereby obtaining the identification result.

[0059] In step S101 of some embodiments, a decision tree algorithm can be used to construct a risk identification model and process the feature data based on the risk identification model. A risk identification model can also be constructed through other types of machine learning models, such as support vector machines or neural networks, which are not limited thereto.

[0060] Please refer to Figure 2 , in some embodiments, a risk behavior identification method based on a decision tree provided by the embodiments of the present invention may include, but is not limited to, steps S201 to S202:

[0061] Step S201: Determine the risk level according to the identification result and compare the risk level with a preset warning threshold.

[0062] In step S202, if the risk level is greater than or equal to the preset warning threshold, a warning message is generated, and an intervention plan is generated according to the recognition result.

[0063] In step S201 of some embodiments, after the system outputs the recognition result through the risk recognition model, it can also compare the recognition result with the preset evaluation relationship to provide further evaluation services for the patient; in this embodiment, the system determines the risk level from the recognition result output by the risk recognition model, compares the risk level with the risk warning threshold set in the system, and determines whether it is necessary to intervene in the patient.

[0064] In step S202 of some embodiments, if the risk level is greater than the risk warning threshold set in the system, the system determines that the mental health problem of the current patient is relatively serious, the system generates a warning message and sends it to the relevant medical service providers for relevant processing; at the same time, the system generates a corresponding intervention plan according to the information such as the risk type in the recognition result by matching the corresponding database, and provides it to the medical service providers or users.

[0065] Please refer to Figure 3 , in some embodiments, a risk behavior recognition method based on a decision tree provided by an embodiment of the present invention may include but is not limited to steps S301 to S304:

[0066] Step S301, integrating historical patient data and a first data set according to a preset period to determine a second data set;

[0067] Step S302, constructing a classification model according to the second data set, and performing cross-validation on the classification model to determine a first model performance value;

[0068] Step S303, determining a second model performance value according to the risk recognition model, and comparing the first model performance value with the second model performance value;

[0069] Step S304, if the first model performance value is greater than the second model performance value, updating the risk recognition model according to the classification model.

[0070] In step S301 of some embodiments, the system regularly collects and integrates the input patient data, preprocesses the collected patient data to obtain data with high consistency; integrates and merges the preprocessed patient data with the historical data stored in the system to obtain a new data set; the system regularly updates the parameters of the risk recognition model according to the integrated and merged data set to maintain the accuracy of risk recognition and system security, and ensure long-term effective operation.

[0071] In step S302 of some embodiments, the system constructs a new decision tree model based on the newly integrated data set, classifies the newly integrated data set, and trains the constructed new model. After the training is completed, the trained classification model is verified through cross-validation to evaluate the performance of the model. The system determines whether the performance of the current risk identification model has decreased or "model drift" has occurred based on the performance evaluation of the constructed new model, so as to determine whether it is necessary to update the current risk identification model.

[0072] In step S303 of some embodiments, the performance of the system's current risk identification model is evaluated to determine the current performance value of the risk identification model. In this embodiment, the system calculates performance indicators such as the accuracy rate, precision rate, and recall rate of the risk identification model, and comprehensively determines the current performance value of the risk identification model to reflect the current performance status of the risk identification model.

[0073] In step S304 of some embodiments, the system compares the currently calculated performance value of the risk identification model with the performance value of the constructed new model. If the performance value of the risk identification model is still greater than the performance value of the new model, it indicates that the current risk identification model still has a high recognition accuracy rate; otherwise, it indicates that the performance of the current risk identification model has decreased, making it difficult to ensure the effective operation of the system, and incremental learning is required. The current risk identification model is updated according to the constructed new model.

[0074] Please refer to Figure 4 , in some embodiments, step S103 may include but is not limited to steps S401 to S404:

[0075] Step S401, calculate based on the constructed risk identification model and the feature data set to obtain a first parameter value;

[0076] Step S402, divide the feature data set according to the first parameter value and the first preset threshold to obtain several feature data subsets, and determine the node types corresponding to the several feature data subsets;

[0077] Step S403, if the node type is a leaf node, determine the recognition result according to the node information corresponding to the feature data subset;

[0078] Step S404, if the node type is a non-leaf node, use the feature data subset as the feature data set, and return to execute the calculation based on the constructed risk identification model and the feature data set to obtain a first parameter value until the node type is a leaf node, and determine the recognition result according to the node information corresponding to the feature data subset.

[0079] In step S401 of some embodiments, the risk identification model calculates the splitting parameter values of each variable or feature in the input feature dataset according to the judgment condition of the current node. In this embodiment, the system uses the information gain value or the Gini impurity as the splitting parameter value.

[0080] In step S402 of some embodiments, the risk identification model compares the splitting parameter values corresponding to each feature with the set splitting threshold, and divides the current feature into the left subtree or the right subtree according to the comparison result; the risk identification model sequentially divides each feature in the input feature dataset, so as to realize the segmentation of the feature dataset and obtain two feature subsets; in this embodiment, a stepped splitting threshold can be set to divide the feature dataset into more than two feature subsets; after the risk identification model obtains at least two feature subsets through segmentation, it queries the node type of the current node of each feature subset to determine whether to continue to segment the feature subset.

[0081] In step S403 of some embodiments, if it is queried that the current node of the feature subset is a leaf node, it means that the risk identification model has separated the current feature subset along a certain path to a certain depth. At this time, the features in the feature subset have a high purity, and the risk identification model has completed the identification and classification of the current feature subset, and outputs the label of the current leaf node as the identification result.

[0082] In step S404 of some embodiments, if it is queried that the current node of the feature subset is a non-leaf node, the risk identification model divides the current feature subset based on the judgment condition of the node until the divided feature subset is at a leaf node. The risk identification model has completed the identification and classification of the current feature subset, and outputs the label of the current leaf node as the identification result.

[0083] Please refer to Figure 5 , in some embodiments, the construction of the risk identification model in a risk behavior identification method based on a decision tree provided by the embodiments of the present invention includes but is not limited to steps S501 to S504:

[0084] Step S501, obtain a sample dataset, analyze the sample dataset, and obtain a sample feature set; wherein, the sample dataset includes any combination of electronic health records, psychological evaluation results, social media behavior data, or physiological monitoring data;

[0085] Step S502, calculate the sample feature set to determine the third parameter value, and determine the target sample feature according to the third parameter value;

[0086] Step S503, segment the sample feature set according to the target sample feature, determine a plurality of sample feature subsets, and construct a tree structure;

[0087] In step S504, calculate for each of a number of sample feature subsets to determine a number of fourth parameter values, and determine a risk identification model based on the number of fourth parameter values and a third preset threshold; wherein, the fourth parameter value includes node purity or decision depth.

[0088] In step S501 of some embodiments, collect data from multiple different data sources of different patients, including electronic health records, psychological assessment results, social media behavior data, and real-time physiological monitoring data, etc., integrate data from multiple different perspectives to accurately reflect the changes in the psychological state of patients; the system analyzes the integrated multi-source data, extracts the feature variables therein to obtain a sample feature set, and the system constructs a risk identification model based on the obtained sample feature set.

[0089] In step S502 of some embodiments, the system calculates the split parameter value of each feature variable in the sample feature set, and determines the judgment condition of the current node according to the split parameter values of different feature variables; in this embodiment, the system selects the feature variable corresponding to the extreme value in the split parameter values as the judgment condition of the current node, such as maximizing the information gain or minimizing the Gini impurity.

[0090] In step S503 of some embodiments, after determining the judgment condition, the system sets a judgment threshold and a splitting operation for the judgment condition, and constructs the root node of the decision tree model; after the system determines the judgment condition of the root node, it splits the sample feature set according to the judgment condition to obtain at least two feature subsets; repeat the above operations for the obtained feature subsets, calculate the classification parameter values, determine the judgment condition of the current node according to the extreme values, and gradually construct the tree structure of the decision tree.

[0091] In step S504 of some embodiments, when determining the node judgment condition according to the feature subset and constructing the structure of the decision tree model, after the system splits each feature subset to obtain a new node, before calculating the split parameter value of the feature variable in the new feature subset after splitting, the system judges the node purity or the decision tree depth of the current node, so as to judge whether to stop splitting the feature subset and create a new decision tree node, thereby obtaining the constructed risk identification model.

[0092] Please refer to Figure 6 , in some embodiments, step S504 may include but is not limited to steps S601 to S602:

[0093] In step S601, if a number of fourth parameter values are all greater than or equal to the third preset threshold, use the tree structure as the risk identification model;

[0094] Step S602: If there is a fourth parameter value less than the third preset threshold, use the sample feature subset corresponding to the fourth parameter value as the sample feature set, and return to perform calculations according to the first preset formula and the sample feature set to determine the third parameter value until several fourth parameter values are all greater than or equal to the third preset threshold, and use the tree structure as the risk identification model.

[0095] In step S601 of some embodiments, the system constructs the root node and tree nodes of the decision tree model by splitting the original feature dataset; to improve the efficiency of the decision tree model, the system sets a threshold to limit the depth of the constructed decision tree model or the node purity of the leaf nodes; after splitting the feature subset, the system determines whether the labels of the feature variables in the split feature subset are consistent as the node purity; or determines the depth of the current leaf node, and compares the determined node purity or leaf node depth with the set threshold. If the node purity or leaf node depth is greater than or equal to the set threshold, stop splitting the current leaf node; if the node depth or node purity of all leaf nodes is greater than or equal to the set threshold, the system stops constructing the decision tree model and uses the current tree structure as the risk identification model.

[0096] In step S602 of some embodiments, if there are still branches in the constructed tree model or the node depth or node purity of the leaf nodes of all branches is less than the set threshold, the system continues to construct child nodes for the branches with node depth or node purity less than the set threshold until the node depth or node purity of the leaf nodes of the branch is greater than or equal to the set threshold, thereby obtaining the risk identification model.

[0097] Please refer to Figure 7 , in some embodiments, step S502 may include but is not limited to steps S701 to S703:

[0098] Step S701: Calculate the first information entropy of the sample feature set, and split the sample feature set according to the sample features of the sample feature set to determine several sample subsets;

[0099] Step S702: Calculate the second information entropy and the first ratio corresponding to several sample subsets respectively, and calculate the sum of the products of the second information entropy and the first ratio to determine the fifth parameter; wherein, the first ratio represents the proportion of the sample subset in the sample feature set;

[0100] Step S703: Calculate the difference between the first information entropy and the fifth parameter as the third parameter value.

[0101] In step S701 of some embodiments, the system uses information gain as the splitting parameter for each feature variable in the feature dataset or each feature subset. The system divides the feature dataset through the information gain of each feature variable. First, the system calculates the information entropy of the feature dataset. Then, the feature variables in the feature dataset are used to divide the feature dataset to obtain at least two feature subsets.

[0102] In step S702 of some embodiments, for the feature subsets obtained by the division, the system calculates the corresponding information entropy respectively and calculates the proportion of each feature subset in the feature dataset before the division. The system calculates the sum of the products of the information entropy of each feature subset and the corresponding proportion as the conditional information entropy of the feature subsets after the feature dataset is divided according to the selected feature variable.

[0103] In step S703 of some embodiments, the difference between the information entropy of the feature dataset before the division and the conditional information entropy of the feature subsets after the division calculated above is used as the information gain of the feature variable of the current division condition. The system repeats the above information gain calculation for all feature variables in the feature dataset before the division to obtain the corresponding information gain values. The feature variable corresponding to the maximum information gain value is selected as the judgment condition for dividing the feature dataset at the current node.

[0104] Next, in combination with specific application examples, the solutions of the embodiments of the present invention will be introduced and described in detail:

[0105] Please refer to Figure 8 , a risk behavior recognition method based on decision tree provided by the present invention is implemented through a school mental health assessment system that integrates mental health risk assessment and comprehensive mental health monitoring. This system connects to the student information system data to obtain relevant student data; through the API interface, it obtains regular mental health questionnaire data and optional social behavior data. The system integrates the collected multi-dimensional student data, performs preprocessing such as data cleaning and missing value processing on the data of different dimensions, and then inputs it into the risk recognition model for processing. The risk recognition model set in the system is constructed by the decision tree algorithm CART; the nodes in the risk recognition model calculate the information gain value or Gini impurity value for the features of the input data according to the judgment conditions of the current node. The specific calculation formulas are as follows,

[0106]

[0107]

[0108] Among them, G(D,f) is the information gain value of the dataset D divided according to the feature f, H(D) is the information entropy of the dataset D, D j is the j-th subset obtained by dividing according to the feature f, H(Dj ) is the information entropy of the j-th subset obtained by splitting according to the feature f, m is the number of subsets obtained by splitting, G(D) is the Gini impurity value of the data D, k is the number of features in the data set D, and P i is the proportion of the i-th type of feature in the data set D; the risk identification model divides the data according to the threshold of the current node and the calculated information gain value or Gini impurity value, and outputs the divided data to the corresponding next node until the data is divided into a leaf node. Take a specific category or risk level corresponding to the leaf node as one of the identification results, and integrate the specific categories or risk levels corresponding to the leaf nodes assigned with the input data to obtain the risk identification result. For example: anxiety disorder, low risk. When outputting the risk identification result, the risk identification model also evaluates the importance of each feature in the input data according to the following formula,

[0109]

[0110] where I(f) is the importance value of the feature f, Tf is all the tree nodes using the feature f, IG(t,f) is the information gain using the feature f at the node t, and N(t) is the sample proportion at the node t; the importance of the feature reflects the depth and frequency of use of the feature in the risk identification model during risk identification, as well as the degree of influence on the risk identification result; medical service providers or psychological teachers determine the main factors affecting the mental health status of students according to the importance of each feature; at the same time, the system monitors in real time according to the risk identification result output by the risk identification model, compares the risk identification result with the set multi-level early warning system, and when a high-risk signal is detected through the comparison, the system automatically notifies the psychological teacher or the head teacher, and automatically generates personalized intervention suggestions according to the result output by the risk identification model, including participation in counseling groups, one-on-one consultations or online mental health courses, to provide timely help for students.

[0111] Implementing the embodiments of the present invention includes the following beneficial effects: The present embodiment provides a risk behavior identification method, system, electronic device and storage medium based on a decision tree. This solution collects multi-source data and performs preprocessing to integrate it into a first data set; analyzes the first data set to determine a feature data set, and classifies and identifies the feature data set through the constructed risk identification model to obtain the risk type or risk level corresponding to the multi-source data; by collecting multi-source data and integrating data from different data sources for risk identification, the data utilization rate is improved, and thus the comprehensiveness of the risk identification result is improved; based on the characteristics of multi-source data, by constructing a risk identification model to classify and identify multi-source data, the mental health changes reflected by multi-source data can be accurately identified and monitored, thereby improving the accuracy of risk identification; at the same time, by using the risk level or risk type obtained through classification and identification as the identification result, the objectivity and understandability of risk assessment are improved.

[0112] As shown Figure 9 in the figure, an embodiment of the present invention further provides a risk behavior recognition system based on a decision tree, including:

[0113] A first module, configured to obtain multi-source data, preprocess the multi-source data, and obtain a first data set;

[0114] A second module, configured to analyze the first data set to determine a feature data set;

[0115] A third module, configured to perform classification and recognition according to the constructed risk recognition model and the feature data set to obtain a recognition result; wherein, the recognition result includes a risk level and / or a risk category.

[0116] It can be seen that the content in the above method embodiments is applicable to the system embodiments of the present invention. The functions specifically implemented by the system embodiments of the present invention are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those of the above method embodiments.

[0117] An embodiment of the present application further provides an electronic device. The electronic device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it implements the above-mentioned risk behavior recognition method based on a decision tree. The electronic device can be any intelligent terminal including a tablet computer, an in-vehicle computer, etc.

[0118] It can be understood that the content in the above method embodiments is applicable to the device embodiments of the present invention. The functions specifically implemented by the device embodiments of the present invention are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those of the above method embodiments.

[0119] Please refer to Figure 10 , Figure 10 which shows the hardware structure of an electronic device according to another embodiment. The electronic device includes:

[0120] A processor 1001, which can be implemented in a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, etc., and is configured to execute relevant programs to implement the technical solutions provided by the embodiments of the present application;

[0121] The memory 1002 can be implemented in the form of a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM), etc. The memory 1002 can store an operating system and other application programs. When implementing the technical solutions provided in the embodiments of this specification through software or firmware, the relevant program codes are stored in the memory 1002 and are called by the processor 1001 to execute a risk behavior recognition method based on a decision tree according to an embodiment of the present application;

[0122] The input / output interface 1003 is used to implement information input and output;

[0123] The communication interface 1004 is used to implement communication interaction between this device and other devices. It can achieve communication through a wired manner (such as USB, network cable, etc.) or through a wireless manner (such as mobile network, WIFI, Bluetooth, etc.);

[0124] The bus 1005 transmits information between various components of the device (such as the processor 1001, the memory 1002, the input / output interface 1003, and the communication interface 1004);

[0125] Among them, the processor 1001, the memory 1002, the input / output interface 1003, and the communication interface 1004 achieve communication connections with each other inside the device through the bus 1005.

[0126] Among them, as a non-transitory computer-readable storage medium, the memory can be used to store non-transitory software programs and non-transitory computer-executable programs. The memory can include high-speed random access memory, and can also include non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory optionally includes a remote memory remotely set relative to the processor, and these remote memories can be connected to the processor through a network. Examples of the above networks include, but are not limited to, the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.

[0127] In addition, an embodiment of the present application also discloses a computer program product or a computer program. The computer program product or the computer program is stored in a computer-readable storage medium. The processor of the computer device can read the computer program from the computer-readable storage medium, and the processor executes the computer program, so that the computer device executes the above method. Similarly, the content in the above method embodiments is applicable to this storage medium embodiment. The functions specifically implemented by this storage medium embodiment are the same as those in the above method embodiments, and the beneficial effects achieved are also the same as those in the above method embodiments.

[0128] An embodiment of the present application also provides a computer-readable storage medium storing a computer program, which when executed by a processor, implements the above-mentioned risk behavior recognition method based on a decision tree.

[0129] It can be understood that the content in the above method embodiments is applicable to this storage medium embodiment. The functions specifically implemented by this storage medium embodiment are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those of the above method embodiments.

[0130] It can be understood that all or some of the steps and systems in the methods disclosed above can be implemented as software, firmware, hardware, and their appropriate combinations. Some physical components or all physical components can be implemented as software executed by a processor, such as a central processing unit, a digital signal processor, or a microprocessor, or implemented as hardware, or implemented as an integrated circuit, such as an application-specific integrated circuit. Such software can be distributed on a computer-readable medium, which can include a computer storage medium (or non-transitory medium) and a communication medium (or transitory medium). As is well known to those of ordinary skill in the art, the term computer storage medium includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information, such as computer-readable instructions, data structures, program modules, or other data. Computer storage media include, but are not limited to, RAM, ROM, EEPROM, flash memory, or other memory technologies, CD-ROM, digital versatile disk (DVD), or other optical disk storage, magnetic cartridges, tapes, magnetic disk storage, or other magnetic storage devices, or any other medium that can be used to store the desired information and can be accessed by a computer. In addition, as is well known to those of ordinary skill in the art, a communication medium typically includes computer-readable instructions, data structures, program modules, or other data in a modulated data signal such as a carrier wave or other transmission mechanism, and can include any information delivery medium.

[0131] The above is a specific description of the preferred embodiments of the present invention. However, the present invention is not limited to the described embodiments. Those skilled in the art can make various equivalent deformations or substitutions without departing from the spirit of the present invention, and these equivalent deformations or substitutions are all included within the scope defined by the claims of this application.

Claims

1. A risk behavior identification method based on decision tree, characterized in that: The method comprises: Acquire multi-source data, and pre-process the multi-source data to obtain a first data set; Analyze the first data set to determine a feature data set; Classification and identification are performed according to the constructed risk identification model and the feature data set to obtain an identification result; wherein the identification result includes a risk level and / or a risk category; wherein the risk identification model includes a plurality of leaf nodes and non-leaf nodes.

2. The method according to claim 1, characterized in that The root is classified and identified according to the constructed risk identification model and the feature data set to obtain an identification result, which specifically includes: Calculating according to the constructed risk identification model and the characteristic data set to obtain a first parameter value; Segmenting the feature data set according to the first parameter value and a first preset threshold value to obtain a plurality of feature data subsets, and determining the node types corresponding to the plurality of feature data subsets; If the node type is a leaf node, determining the recognition result according to the node information corresponding to the feature data subset; If the node type is a non-leaf node, the feature data subset is used as the feature data set, and the calculation based on the constructed risk identification model and the feature data set is returned to obtain a first parameter value until the node type is a leaf node, and the identification result is determined according to the node information corresponding to the feature data subset.

3. The method according to claim 1, characterized in that The risk identification model is determined in the following way: Acquire a sample data set, and analyze the sample data set to obtain a sample feature set; wherein the sample data set includes any combination of electronic health records, psychological assessment results, social media behavior data, or physiological monitoring data; Calculating the sample feature set to determine a third parameter value, and determining a target sample feature according to the third parameter value; Segmenting the sample feature set according to the target sample feature, determining a plurality of sample feature subsets, and constructing a tree structure; Calculate several sample feature subsets respectively to determine several fourth parameter values, and determine the risk identification model according to the several fourth parameter values ​​and a third preset threshold; wherein the fourth parameter value includes node purity or decision depth.

4. The method according to claim 3, characterized in that Determining the risk identification model according to the plurality of fourth parameter values ​​and the third preset threshold value specifically includes: If a plurality of the fourth parameter values ​​are greater than or equal to the third preset threshold, using the tree structure as the risk identification model; If there is a fourth parameter value that is less than the third preset threshold, the sample feature subset corresponding to the fourth parameter value is used as the sample feature set, and the process returns to execute the calculation based on the first preset formula and the sample feature set to determine the third parameter value, until several fourth parameter values ​​are greater than or equal to the third preset threshold, and the tree structure is used as the risk identification model.

5. The method according to claim 3, characterized in that: The calculating the sample feature set to determine the third parameter value specifically includes: Calculating a first information entropy of the sample feature set, segmenting the sample feature set according to sample features of the sample feature set, and determining a plurality of sample subsets; Calculate the second information entropy and the first ratio corresponding to a number of the sample subsets respectively, and calculate the sum of the products of the second information entropy and the first ratio to determine a fifth parameter; wherein the first ratio represents the proportion of the sample subset to the sample feature set; The difference between the first information entropy and the fifth parameter is calculated as the third parameter value.

6. The method according to claim 1, characterized in that The method further comprises: Determine a risk level based on the identification result, and compare the risk level with a preset warning threshold; If the risk level is greater than or equal to the preset warning threshold, a warning message is generated, and an intervention plan is generated according to the identification result.

7. The method according to claim 1, characterized in that The method further comprises: Integrate the historical patient data and the first data set according to a preset period to determine a second data set; Building a classification model according to the second data set, and cross-validating the classification model to determine a first model performance value; determining a second model performance value according to the risk identification model, and comparing the first model performance value with the second model performance value; If the first model performance value is greater than the second model performance value, the risk identification model is updated according to the classification model.

8. A risk behavior identification system based on decision tree, characterized in that: include: The first module is used to acquire multi-source data and pre-process the multi-source data to obtain a first data set; A second module is used to analyze the first data set to determine a feature data set; The third module is used to perform classification and identification based on the constructed risk identification model and the feature data set to obtain an identification result; wherein the identification result includes a risk level and / or a risk category; wherein the risk identification model includes a number of leaf nodes and non-leaf nodes.

9. An electronic device, characterized in that: include: at least one processor; at least one memory for storing at least one program; When the at least one program is executed by the at least one processor, the at least one processor implements the method according to any one of claims 1 to 7.

10. A computer-readable storage medium storing a program executable by a processor, characterized in that: The processor-executable program is used to perform the method according to any one of claims 1 to 7 when executed by the processor.

Citation Information

Patent Citations

  • Large health intelligent analysis and evaluation system based on multi-source data fusion

    CN118737464A

  • Psychological health condition general screening and evaluation method, system, equipment and medium

    CN118866364A

  • Mental health monitoring and early warning method and system based on big data

    CN119007998A

  • Intelligent classification early warning method and system for psychological health risk of college students

    CN119028574A