Influence factor evaluation method and device, equipment, storage medium and program product

By classifying patient characteristics and conducting correlation analysis, influencing factors were identified, the problem of subjective bias in home assessment was resolved, and more accurate adjustments to rehabilitation plans were achieved.

CN121964167APending Publication Date: 2026-05-01CHINA MOBILE COMM LTD RES INST +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CHINA MOBILE COMM LTD RES INST
Filing Date
2024-10-30
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Existing technologies rely on subjective recording in patient home assessments, which are prone to recall bias and omissions, and lack objective and continuous comprehensive assessment methods. In particular, they are difficult to capture intermittent and sudden symptoms in chronic disease rehabilitation monitoring.

Method used

By classifying business-related characteristics, calculating correlations and clustering analysis, determining the influencing factors of static and dynamic attribute characteristics, and combining importance calculations, objective and continuous assessment of patients in the home environment can be achieved.

Benefits of technology

It improves the accuracy of cluster analysis of population characteristics, provides more objective and continuous home assessment support, and reduces the burden on medical staff and patients.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121964167A_ABST
    Figure CN121964167A_ABST
Patent Text Reader

Abstract

The invention discloses an influence factor evaluation method and device, equipment, a storage medium and a program product, and the method comprises the steps: classifying business related features, and obtaining a plurality of class features; the plurality of category features comprise static attribute features and dynamic attribute features; performing correlation calculation and clustering analysis on feature parameters in the plurality of category features to obtain a plurality of first evaluation results; performing importance calculation on the plurality of category features to obtain a plurality of second evaluation results; and based on the plurality of first evaluation results and the plurality of second evaluation results, determining a target influence factor related to the business.
Need to check novelty before this filing date? Find Prior Art

Description

An impact factor assessment method, apparatus, equipment, storage medium, and program product. Technical Field

[0001] This application relates to the fields of artificial intelligence and big data, and in particular to an impact factor assessment method, apparatus, device, storage medium, and program product. Background Technology

[0002] Currently, evaluating the efficacy of medications and adjusting rehabilitation plans for patients requires collecting cyclical patterns of gait characteristics over multiple days. Furthermore, due to the time constraints of outpatient clinics and the "observer effect" in some diseases, home-based assessment and testing could significantly reduce the burden on both doctors and patients. However, current home-based disease assessment technologies primarily rely on subjective records such as patient diaries, which suffer from drawbacks such as recall bias and omissions. Therefore, there is a strong demand for more objective and continuous comprehensive assessment methods suitable for the daily home environment. Summary of the Invention

[0003] To address the aforementioned technical problems, embodiments of the present invention provide an impact factor assessment method, apparatus, device, storage medium, and program product.

[0004] The impact factor evaluation method provided in this application includes:

[0005] Business-related features are classified to obtain multiple category features; these multiple category features include static attribute features and dynamic attribute features.

[0006] Correlation calculation and cluster analysis are performed on the feature parameters of the multiple category features to obtain multiple first evaluation results; importance calculation is performed on the multiple category features to obtain multiple second evaluation results;

[0007] Based on the multiple first evaluation results and the multiple second evaluation results, business-related target impact factors are determined.

[0008] The impact factor evaluation device provided in this application includes:

[0009] A classification unit is used to classify business-related features to obtain multiple category features; the multiple category features include static attribute features and dynamic attribute features;

[0010] The evaluation unit is used to perform correlation calculation and cluster analysis on the feature parameters of the multiple category features to obtain multiple first evaluation results; and to perform importance calculation on the multiple category features to obtain multiple second evaluation results.

[0011] The determining unit is used to determine business-related target impact factors based on the plurality of first evaluation results and the plurality of second evaluation results.

[0012] The processing device provided in this application includes a processor and a memory. The memory is used to store computer programs, and the processor is used to call and run the computer programs stored in the memory to execute any of the above-described impact factor evaluation methods.

[0013] The computer-readable storage medium provided in this application embodiment is used to store a computer program that causes a computer to execute any of the above-described impact factor evaluation methods.

[0014] The computer program product provided in this application includes computer program instructions that cause a computer to execute any of the above-described impact factor evaluation methods.

[0015] In the technical solution of this application embodiment, business-related features are classified to obtain multiple category features. Correlation calculations and cluster analysis are performed on the feature parameters within these multiple category features to obtain multiple first evaluation results. Importance calculations are also performed on the multiple category features to obtain multiple second evaluation results. Based on these multiple first and second evaluation results, business-related target influencing factors are determined. The multiple category features include static attribute features and dynamic attribute features. Thus, by determining the degree of influence between related parameters through correlation and cluster analysis of feature parameters within category features, and by determining the importance of different categories through importance analysis of different category features, and by combining the degree of influence between related parameters and the importance of different categories, the weights of different influencing factors can be analyzed under a detailed quantification based on the mixed influence of static and dynamic attribute features. This improves the accuracy of population feature clustering analysis and provides strong support for more objective and continuous comprehensive assessment of patients in a home environment. Attached Figure Description

[0016] Figure 1 is a schematic diagram of the architecture of the body area network provided in an embodiment of this application;

[0017] Figure 2 is a schematic diagram of the system components in which the hospital-oriented clustering capability provided in the embodiments of this application is located;

[0018] Figure 3 is a flowchart illustrating the impact factor evaluation method provided in the embodiments of this application;

[0019] Figure 4 is a flowchart illustrating the method for evaluating the impact factors of twin portraits of human physical characteristics based on mixed effects, as provided in the embodiments of this application.

[0020] Figure 5 is a schematic diagram of the data distribution after performing cluster analysis on the coarse-grained dataset according to the clustering model provided in the embodiments of this application;

[0021] Figure 6 is a schematic diagram of the clustering results analysis of the original dataset and the coarse-grained processed dataset provided in the embodiments of this application;

[0022] Figure 7 is a schematic diagram of clustering result analysis for the three distance calculation methods provided in the embodiments of this application;

[0023] Figure 8 is a schematic diagram of clustering results analysis when the weighting coefficients in the custom weighted distance function provided in the embodiments of this application are different values;

[0024] Figure 9 is a schematic diagram of the clustering results analysis when the weighting coefficient λ3 of the gender difference function provided in the embodiments of this application has different values;

[0025] Figure 10 is a schematic diagram of the impact factor evaluation device provided in an embodiment of this application;

[0026] Figure 11 is a schematic diagram of the processing device provided in an embodiment of this application. Detailed Implementation

[0027] The technical solutions of the embodiments of this application will now be described with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are within the scope of protection of this application.

[0028] In the description of the embodiments of this application, the term "correspondence" may indicate that there is a direct or indirect correspondence between two things, or that there is an association between two things, or that there is a relationship of instruction and being instructed, configuration and being configured, etc.

[0029] To facilitate understanding of the technical solutions of the embodiments of this application, the relevant technologies of the embodiments of this application are described below. The following relevant technologies are optional solutions and can be combined with the technical solutions of the embodiments of this application in any way, and they all fall within the protection scope of the embodiments of this application.

[0030] Body area networks (BANs) are short-range, low-power, high-speed wireless communication networks primarily used to connect body-worn sensors and medical devices. With the development of BANs, body-worn sensors have brought more real-time and comprehensive data to human digital twins. In addition to analyzing traditional static profiles of people (relatively stable age, height, and other behavioral habits over a period of time), they can also more accurately analyze real-time dynamic profiles of people (subtle physical signs under different movement states). Figure 1 shows a schematic diagram of the body area network (BAN) architecture. The BAN is an application network composed of a body soft gateway, terminals connected to the body soft gateway, and core nodes. The body soft gateway is a crucial component of the BAN, primarily responsible for processing and forwarding data from connected terminal devices. These terminals are typically various body-worn sensors and medical devices, such as heart rate monitors, blood pressure monitors, and blood glucose meters. These terminal devices connect to the soft gateway via wireless communication to achieve data transmission and exchange. Core nodes are nodes with core functions within the BAN, typically including the body soft gateway and other critical equipment. Core nodes are responsible for processing data flows, implementing data routing, protocol conversion, and data encryption to ensure the normal operation of the BAN. In certain application scenarios, core nodes can also interconnect with external networks (such as hospital information systems and home networks) to achieve data exchange and information sharing.

[0031] Figure 2 shows a schematic diagram of the system components in which the clustering capability for hospitals is located. The adaptive open clustering middleware based on the daily home vital signs profile of the population can realize the health monitoring function of users at home. Doctors can remotely analyze the user's daily home vital signs data to understand the user's health status more quickly and provide the user with personalized medical advice.

[0032] The dynamic vital signs parameters of wearable systems are affected by multiple factors, including: (1) demographic factors such as gender, age, height and body mass index (BMI); (2) dynamic movement status, such as stride length, which varies in different exercise scenarios such as running, walking, and going up and down stairs; (3) drug effects; and (4) exercise rehabilitation treatment plans.

[0033] Chronic diseases require rehabilitation assessment and management. Based on symptom evaluation, disease progression is assessed, and adjustments are made to rehabilitation programs, including medication type, dosage, and exercise rehabilitation. During rehabilitation, it's necessary to monitor the cyclical changes in gait characteristics at different times each day after medication administration, as well as over multiple days, to assess drug efficacy and adjust rehabilitation programs. Furthermore, some symptoms, such as frozen gait, have intermittent and sudden onset, making them difficult to detect within the limited time available in an outpatient setting. For neurological disorders, the "observer effect" exists; walking under the supervision of medical staff may not reveal the full gait characteristics of a naturally relaxed state. Therefore, compared to inpatient observation, home-based monitoring would significantly reduce the burden on both patients and healthcare providers.

[0034] However, current home-based disease assessment technologies primarily rely on subjective records such as patient diaries, which suffer from drawbacks such as recall bias and omissions. Therefore, there is a strong demand for more objective and continuous comprehensive assessment methods suitable for the daily home environment. Given that dynamic user profiles of vital signs are influenced by a combination of factors and the complexity of the body area network, analyzing the weights of different influencing factors by breaking down the components of these factors would be of great significance for the development of home-based rehabilitation programs.

[0035] To address the aforementioned technical problems and facilitate understanding of the technical solutions in this application, the technical solutions of this application are described in detail below through specific embodiments. The above-mentioned related technologies are optional solutions and can be arbitrarily combined with the technical solutions in the embodiments of this application, all of which fall within the protection scope of the embodiments of this application. The embodiments of this application include at least some of the following contents.

[0036] This application proposes an impact factor evaluation method. Figure 3 is a flowchart illustrating the impact factor evaluation method provided in this application. As shown in Figure 3, the method includes the following steps:

[0037] Step 301: Classify the business-related features to obtain multiple categories of features.

[0038] Among them, multiple categorical features include static attribute features and dynamic attribute features.

[0039] In this embodiment, feature parameters related to business requirements can be classified according to different attribute characteristics. Feature parameters with the same attribute characteristics are grouped into one category, resulting in multiple category features. Specifically, feature parameters related to business requirements can be classified according to static and dynamic attribute classification methods, resulting in two category features: static attribute features and dynamic attribute features. Static attribute features and dynamic attribute features each include multiple feature parameters with the same or similar attributes. For example, static attribute features may include relatively fixed feature parameters such as age, height, and weight that are not easily changed over time, while dynamic feature parameters may include feature parameters that describe the behavior or state of an object and change over time, such as walking, running, and jumping.

[0040] Here, dynamic attribute features also include gait attribute features, which are used to describe the characteristics of an object during walking. These include gait feature parameters such as stride length, stride frequency, stride speed, stride width, gait cycle, and stride length, which change with the object's walking behavior.

[0041] Step 302: Perform correlation calculation and cluster analysis on the feature parameters of multiple categories to obtain multiple first evaluation results.

[0042] In this embodiment of the application, after classifying business-related features to obtain multiple category features, correlation calculation and cluster analysis can be performed on the feature parameters within each category feature, and the obtained correlation calculation results and clustering results can be fused to obtain multiple first evaluation results.

[0043] In some implementations, static attribute features include an age parameter, dynamic attribute features include gait attribute features, and gait attribute features include multiple gait feature parameters; step 302 can be implemented through steps S3021-S3024, specifically including:

[0044] Step S3021: Perform correlation calculations on the age parameter and each gait feature parameter among the multiple gait feature parameters to obtain multiple correlation coefficients between the age parameter and the multiple gait feature parameters.

[0045] In this embodiment, an age parameter can be extracted from static attribute features, and multiple gait feature parameters can be extracted from gait attribute features included in dynamic attribute features. Correlation calculations are then performed on the age parameter and each gait feature parameter to obtain multiple correlation coefficients between the age parameter and the multiple gait feature parameters. Specifically, the formula for calculating the correlation coefficient between the age parameter and the multiple gait feature parameters is as follows:

[0046]

[0047] Among them, X kY is the k-th data point corresponding to the age parameter. j,k For the k-th data corresponding to the j-th gait feature parameter, This represents the average value of the data corresponding to the age parameter. Let r be the average value of the data corresponding to the j-th gait feature parameter, N be the number of data points, and r be the average value of the data points. j is the Pearson correlation coefficient between the age parameter and the j-th gait feature parameter.

[0048] Step S3022: Determine multiple contour coefficients based on multiple gait feature parameters.

[0049] In the embodiments of this application, each of the multiple gait feature parameters can be used as a clustering criterion to calculate the clustering result of multiple gait feature parameters under each gait feature parameter, and multiple contour coefficients can be determined based on multiple different clustering results.

[0050] In some implementations, step S1022 may specifically include:

[0051] Using each gait feature parameter as input, a clustering model is used to cluster multiple gait feature parameters, resulting in multiple clusters for each gait feature parameter.

[0052] Multiple contour coefficients are determined based on multiple clusters of each gait feature parameter.

[0053] Here, each gait feature parameter is used as a clustering criterion. Each gait feature parameter is input into the clustering model, and multiple gait feature parameters are clustered to obtain a clustering result with each gait feature parameter as the clustering criterion. This clustering result consists of multiple clusters, each containing at least one gait feature parameter. The clustering model is a pre-trained model, including K-Means clustering and Density-Based Spatial Clustering of Applications with Noise (DBSCAN), etc., and is not limited to any particular model here.

[0054] Specifically, for multiple clusters based on each gait feature parameter, determining multiple contour coefficients can include:

[0055] For each gait feature parameter, determine the similarity coefficient within each cluster and the dissimilarity coefficient between each cluster in the multiple clusters of that gait feature parameter;

[0056] Based on the similarity coefficient within each cluster and the dissimilarity coefficient between each cluster, the mean similarity coefficient and the mean dissimilarity coefficient of multiple clusters for this gait feature parameter are determined.

[0057] Based on the mean similarity coefficient and mean dissimilarity coefficient of multiple clusters of the gait feature parameter, the silhouette coefficient of the gait feature parameter is determined.

[0058] Here, under the m-th gait feature parameter, D is obtained from training the clustering model. m For each of the three clusters, obtain the intra-cluster data for each cluster, and iterate through D. m The cluster centers and intra-cluster data of each cluster are analyzed. The similarity coefficient within each cluster is calculated, and the average of the similarity coefficients across all clusters is calculated using the following formula:

[0059]

[0060] Among them, C m,i Y represents the number of data points in the i-th cluster under the m-th gait feature parameter. m,i,k Y represents the k-th data point in the i-th cluster under the m-th gait feature parameter. m,i,j For the j-th data point in the i-th cluster under the m-th gait feature parameter, a m It is the average similarity coefficient within all clusters under the m-th gait feature parameter.

[0061] Then, iterate through D. m Within each cluster, the dissimilarity coefficient between each cluster is calculated, and the average dissimilarity coefficient between all clusters is calculated using the following formula:

[0062]

[0063] Among them, C nearest Y represents the number of data points in the cluster closest to the cluster of the i-th data point (gait feature parameter). m,i Y represents the i-th data point under the m-th gait feature parameter. m,j This represents the j-th data point in the cluster closest to the i-th data point under the m-th gait feature parameter, where M represents the number of gait feature parameters, and b m This represents the average dissimilarity coefficient among all clusters under the m-th gait feature parameter.

[0064] Finally, based on the average similarity coefficients within all clusters and the average dissimilarity coefficients between all clusters for the m-th gait feature parameter, the silhouette coefficient for the m-th gait feature parameter is calculated using the following formula:

[0065]

[0066] Where max(·) represents finding the maximum value of the input data, s m It represents the silhouette coefficient under the m-th gait feature parameter.

[0067] Step S3023: Determine multiple fusion weight coefficients based on multiple correlation coefficients and multiple profile coefficients.

[0068] In this embodiment of the application, after calculating the correlation coefficient between the age parameter and each gait feature parameter and the contour coefficient of each gait feature parameter to obtain multiple correlation coefficients and multiple contour coefficients, multiple first-level weight coefficients and multiple second-level weight coefficients can be obtained based on the multiple correlation coefficients and multiple contour coefficients, and the multiple first-level weight coefficients and multiple second-level weight coefficients can be fused to obtain multiple fused weight coefficients.

[0069] In some implementations, step S1023 may specifically include:

[0070] The absolute values ​​of multiple correlation coefficients are sorted in ascending order to obtain multiple first-level weight coefficients;

[0071] Multiple contour coefficients are sorted in ascending order to obtain multiple secondary weight coefficients;

[0072] Multiple primary weight coefficients and multiple secondary weight coefficients are merged to obtain multiple merged weight coefficients.

[0073] Specifically, the absolute values ​​of the correlation coefficients between the age parameter and each gait feature parameter are taken and sorted in ascending order to obtain multiple first-level weight coefficients. The calculation formula is as follows:

[0074] [r (1) ,r (2) ,…,r (M) =Sort([|r1|,|r2|,…,|r...]) M |]) (5)

[0075]

[0076] Where, |r j| indicates that the absolute value of the correlation coefficient between the age parameter and the j-th gait feature parameter is taken. Since the Pearson correlation coefficient ranges from [-1, 1], a correlation coefficient less than zero indicates a negative correlation between the two data points. Here, the sign of the correlation is disregarded, only the degree of correlation is considered; therefore, the absolute value of the correlation coefficient is used. Sort(·) indicates that the input numerical sequence is sorted in ascending order, and the sorted result is [r...]. (1) ,r (2) ,…,r (M) ], where r (1) r is the minimum value in the original sequence. (M) The maximum value is r. (1) ≤r (2) ≤…≤r (M) M represents the number of gait feature parameters; Let be the weight coefficient corresponding to the gait feature parameter at position (i) in the sorted relevant sequence, with the value being .

[0077] Then, the contour coefficients of each gait feature parameter are sorted in ascending order to obtain multiple secondary weight coefficients, the calculation formula of which is as follows:

[0078] [s (1) ,s (2) ,…,s (M) =Sort([r1,r2,…,r)) M (7)

[0079]

[0080] Among them, [s (1) ,s (2) ,…,s (M) [ ] represents the sequence of silhouette coefficients of different time-state feature parameters sorted from smallest to largest, where s (1) s is the minimum value in the original sequence. (M) The maximum value is s. (1) ≤s (2) ≤…≤s (M) ; The weight coefficient corresponding to the gait feature parameter at the (i)th position in the sorted contour coefficient sequence is:

[0081] Finally, a fusion weighting method is used to fuse multiple primary weight coefficients and multiple secondary weight coefficients to obtain multiple fused weight coefficients, which measure the degree of influence of age on gait feature parameters. The calculation formula is as follows:

[0082]

[0083] in, and These represent the first-level weighting coefficient and the second-level weighting coefficient under the m-th gait feature parameter, respectively. This represents the fusion weight coefficient for the m-th gait feature parameter, with a value range of...

[0084] Step S3024: Determine multiple first evaluation results based on multiple fusion weight coefficients.

[0085] In the embodiments of this application, since different fusion weight coefficients have different sizes, the different sizes of fusion weight coefficients represent the correlation between age and different time-phase characteristic parameters, that is, they can be used to measure the degree of influence of age on different time-phase characteristic parameters. The correlation between age and different time-phase characteristic parameters represents different evaluation results. Therefore, multiple first evaluation results can be determined based on multiple fusion weight coefficients.

[0086] In some implementations, step S3024 may specifically include:

[0087] Multiple fusion weight coefficients are reverse-converted to obtain multiple reverse fusion weight coefficients;

[0088] Multiple reverse fusion weight coefficients are sorted in descending order to obtain multiple first evaluation results.

[0089] Here, since the fusion weight coefficients represent the positive correlation between age parameters and gait feature parameters—that is, the influence of age on gait feature parameters is directly proportional—it is necessary to convert this direct correlation into an inverse correlation and then sort them in reverse order to obtain gait feature parameters with low correlation to age. Specifically, multiple fusion weight coefficients are converted using a negative logarithm method to obtain multiple inverse fusion weight coefficients, and the calculation formula is as follows:

[0090]

[0091] in, The fusion weight coefficient (reverse fusion weight coefficient) after logarithmic transformation under the m-th gait feature parameter has a value range of [0, 2logM]. In this case, the smaller the value of the reverse fusion weight coefficient, the greater the influence of age factors, and the larger the value, the smaller the influence of age factors.

[0092] Then, the multiple reverse fusion weight coefficients are sorted in descending order to obtain multiple first evaluation results. Among them, the reverse fusion weight coefficient with the largest value indicates that its corresponding gait feature parameters are least affected by age.

[0093] Step 303: Calculate the importance of multiple category features to obtain multiple second evaluation results.

[0094] In this embodiment, after classifying business-related features to obtain multiple category features, the importance of each category feature can be calculated to obtain an importance score for each category feature, thereby obtaining multiple second evaluation results. Specifically, the importance calculation method of decision tree impurity can be used to calculate the importance of each category feature, that is, the importance score of each category feature is obtained by calculating the average contribution of each category feature on all decision trees.

[0095] In some implementations, step 303 can be achieved through steps S3031-S3032, specifically including:

[0096] Step S3031: For each category feature, determine the average impurity reduction of that category feature through multiple decision trees.

[0097] Multiple decision trees are used to evaluate the importance of each category feature.

[0098] In this embodiment, the impurity change (impurity reduction) caused by splitting nodes in each decision tree for each category feature can be calculated first, and the total impurity change caused by splitting nodes in all decision trees for each category feature can be accumulated. Then, the total impurity change for each category feature is divided by the number of decision trees to obtain the average impurity reduction for each category feature. This average impurity reduction can reflect the importance of the category feature.

[0099] In some implementations, step S3031 may specifically include:

[0100] Determine the impurity reduction of each decision tree in a set of multiple decision trees for that category feature;

[0101] Based on the reduction in impurity of each decision tree for that category of feature, the average reduction in impurity of that category of feature is determined.

[0102] Specifically, determining the reduction of impurity for each of multiple decision trees for that category feature can include:

[0103] For each decision tree, when the category feature is used to split a node in the decision tree, the difference in impurity before and after the split is calculated to obtain the reduction in impurity caused by the split.

[0104] Based on the reduction in impurity caused by each split in the decision tree, the reduction in impurity of the decision tree for this category feature is determined.

[0105] Here, when calculating the change in impurity caused by splitting each node in each decision tree for each category feature, for each node in each decision tree, when a category feature splits that node in the decision tree, the difference in impurity before and after the split is calculated. This difference is the change in impurity caused by the split. Then, the changes in impurity caused by splitting the category feature into other nodes in the decision tree are accumulated in turn, so as to obtain the total reduction in impurity of the category feature in the decision tree.

[0106] It should be noted that the change in impurity can be calculated based on impurity functions such as information entropy or the Gini index. These functions quantify the impurity of data. By comparing the impurity before and after splitting the decision tree, the quality of the split can be evaluated. The greater the change in impurity, the better the split, and the higher the purity of the resulting data subset. The formula for calculating Gini impurity based on the Gini index is as follows:

[0107]

[0108] Where, p i This represents the proportion of samples in the i-th category feature in the dataset out of the total number of samples.

[0109] For a specific split, the formula for calculating the change in impurity caused by that split is:

[0110]

[0111] Where ΔGini represents the change in impurity caused by splitting, Gini parent It's the impurity of the parent node, Gini left and Gini right It represents the impurity of the left and right child nodes, N. left N right , and N parent These are the sample numbers for the left node, right node, and parent node, respectively.

[0112] For each category feature, the reduction in impurity caused by that category feature across all decision trees is summed, then divided by the number of decision trees to obtain the average reduction in impurity for that category feature. This value is considered the importance score of that category feature, and its calculation formula is as follows:

[0113]

[0114] in, This indicates a decrease in the average impurity (importance score) of the i-th category feature, N. trees It is the number of decision trees.

[0115] Step S3032: Reduce the average impurity of the category feature and sort it in descending order to obtain the second evaluation result of the category feature.

[0116] In the embodiments of this application, after obtaining the average impurity reduction of each category feature, the average impurity reduction of each category feature can be used as the importance assessment result of that category feature. The average impurity reduction of each category feature is sorted from largest to smallest, thereby obtaining multiple second assessment results.

[0117] Step 304: Based on multiple first assessment results and multiple second assessment results, determine the target impact factors relevant to the business.

[0118] In this embodiment, since the first evaluation result is the inverse fusion weight coefficient between age and gait feature parameters, and the second evaluation result is the reduction in the average impurity of the category features, which represents the importance score of the category features, different weights can be assigned to static and dynamic attribute features based on multiple first and second evaluation results. Then, the data of business-related features are clustered based on the static and dynamic attribute features with different weights, thereby determining the target impact factor related to the business based on different clustering results. The target impact factor optimizes the clustering effect of the dataset.

[0119] In some implementations, step 304 can be achieved through steps S3041-S3042, specifically including:

[0120] Step S3041: Based on multiple first evaluation results and multiple second evaluation results, assign different weights to static attribute features and dynamic attribute features respectively to obtain multiple weighted static attribute features and multiple weighted dynamic attribute features.

[0121] In this embodiment of the application, it is necessary to first normalize the multiple first evaluation results and multiple second evaluation results to ensure that the numerical range of each evaluation result is normalized to the [0,1] interval. Then, based on the normalized multiple first evaluation results and multiple second evaluation results, different weights are assigned to the static attribute features and the dynamic attribute features to obtain multiple weighted static attribute features and multiple weighted dynamic attribute features with different weighting coefficients.

[0122] Specifically, firstly, the values ​​corresponding to the normalized first evaluation results are arranged in order to form a first weighted coefficient sequence. Since the normalized second evaluation results represent the importance scores of multiple category features, the importance score corresponding to each category feature (static attribute feature and dynamic attribute feature) is used as a second weighted coefficient and assigned to the corresponding category feature. Then, according to the first weighted coefficient sequence, the values ​​in the first weighted coefficient sequence are added to the second weighted coefficient of each category feature in turn, and the average value is taken as the final weighted coefficient of each category feature, thus obtaining multiple weighted static attribute features and multiple dynamic attribute features with different weighted coefficients.

[0123] Assume that the static attribute features are represented as X = [Gender, Age, Height, BMI], and the dynamic attribute features are represented as G = [g 1 ,g 2 …,g 30 The first weighted coefficient sequence is represented as [N1, N2, N3, ..., N]. n Let M1 and M2 be the second weighting coefficients corresponding to the static and dynamic attribute features, respectively. Then, the final weighting coefficients are obtained by combining the values ​​in the first weighting coefficient sequence with the second weighting coefficients of the static attribute features. Represented as:

[0124]

[0125] Similarly, the final weighted coefficients are obtained by combining the values ​​in the first weighted coefficient sequence with the second weighted coefficients of the dynamic attribute features. Represented as:

[0126]

[0127] Step S3042: Determine the target influencing factor based on multiple weighted static attribute features and multiple weighted dynamic attribute features.

[0128] In the embodiments of this application, after obtaining multiple weighted static attribute features and multiple dynamic attribute features with different weighting coefficients, the multiple weighted static attribute features and multiple weighted dynamic attribute features can be combined into different weighted distance functions according to different combination methods. Then, cluster analysis is performed on the data of business-related features according to each weighted distance function, so as to determine the target influence factor based on different clustering results.

[0129] In some implementations, step S3042 may specifically include:

[0130] Based on multiple weighted static attribute features and multiple weighted dynamic attribute features, multiple weighted distance functions are determined;

[0131] Cluster analysis of business-related features is performed based on multiple weighted distance functions, resulting in multiple clustering results;

[0132] Based on multiple clustering results, the target influencing factor was determined.

[0133] Here, based on the static feature data and corresponding gait feature data of different samples, multiple weighted distance functions are determined using multiple weighted static attribute features and multiple weighted dynamic attribute features. Then, based on each weighted distance function, a clustering model is used to perform clustering analysis on the business-related feature data, resulting in multiple different clustering results. Finally, the multiple clustering results are evaluated based on clustering evaluation indicators, and the weighted distance function with the best clustering effect is selected as the target weighted distance function. The weighting coefficients in the target weighted distance function are the target influence factors. The clustering model is a pre-trained model, including K-Means models, DBSCAN models, etc., and is not limited here.

[0134] Suppose the static feature data of the sample is represented as X i =(Gender) i Age i Height i BMI i The corresponding dynamic feature data is represented as follows: Therefore, based on multiple weighted static attribute features and multiple weighted dynamic attribute features, the distance function between static feature data X1 and X2 and their corresponding dynamic feature data G1 and G2 can be defined as:

[0135]

[0136] Among them, D 12 This represents the distance between static feature data X1 and X2 and their corresponding dynamic feature data G1 and G2. These are the weighting coefficients for static attribute features. The weighting coefficients are the dynamic attribute features. D1(X1,X2) represents the static feature distance, and D2(G1,G2) represents the dynamic feature distance.

[0137] The function D1 represents the difference in static feature information between the static feature data X1 and X2 corresponding to two samples, and its definition formula is:

[0138]

[0139] Where, ΔXi This represents the difference between the corresponding values ​​of the feature parameters in the two static feature data sets, X1 and X2. This represents the Euclidean distance between X1 and X2, specifically the age, height, and body mass index components. λ3 is the weighting coefficient of the sex difference function f(Gen1, Gen2), defined by the following formula:

[0140]

[0141] The function D2 represents the difference in gait feature information between the dynamic feature data G1 and G2 corresponding to two samples, and its definition formula is:

[0142]

[0143] Wherein, Δg i This represents the difference in the i-th gait feature between two samples.

[0144] After determining several different weighted distance functions, clustering analysis is performed on the business-related feature data using a clustering model based on each weighted distance function, resulting in multiple different clustering results. These results are then evaluated using clustering evaluation metrics. To quantify the performance of the clustering algorithm, the Davis-Bouldin index (DBI) is used as an evaluation metric for the clustering results. DBI is an internal evaluation metric for measuring cluster quality; its value is calculated by comparing the compactness within clusters and the separation between clusters. Specifically, the DBI value is obtained by calculating the average relative distance between each cluster and its nearest cluster. The formula is as follows:

[0145]

[0146] Wherein d(c i ,c j ) represents the cluster centers c of different clusters. i and c j The separation degree between them, σ i Let be the average distance of data within the i-th cluster, and let represent the compactness within the i-th cluster. Similarly, σ j σ represents the average distance between data points within the j-th cluster, indicating the compactness within the j-th cluster. i The calculation formula is:

[0147]

[0148] Among them, Ci It is the set of points within the i-th cluster. d(x,c) i ) represents the distance from point x within the i-th cluster to the cluster center c within the i-th cluster. i The distance d(c) i ,c j d(c) is the separation degree used to calculate the distance between the cluster centers of different clusters. i ,c j The formula for calculating ) is:

[0149] d(c i ,c j )=‖c i -c j ‖ (twenty two)

[0150] Among them, c i c represents the cluster center within the i-th cluster. j Let c represent the cluster center within the j-th cluster. i -c j ‖ represents the cluster center c i and c j The distance between them.

[0151] A lower DBI value indicates a more compact data structure within clusters and a higher degree of separation between different clusters, thus reflecting better clustering performance. Based on this, by evaluating multiple clustering results using the DBI metric, multiple DBI values ​​can be obtained. The smallest DBI value is then selected. Since the clustering result corresponding to the smallest DBI value is optimal, the weighted distance function corresponding to the optimal clustering result can be used as the target weighted distance function. The weighting coefficients in this target weighted distance function are the target impact factors. These target impact factors reflect the importance of static and dynamic attribute features with different weighting coefficients to the data clustering results.

[0152] It's important to note that before performing clustering analysis on business-related feature data based on a weighted distance function, coarse-grained data processing can be performed on this data to avoid the impact of individual differences and outliers on the clustering results, thus improving the robustness of the clustering algorithm. Specifically, the data is first divided into two parts based on gender, and then further divided based on age, height, and weight. This precise segmentation of the data across four dimensions allows for the calculation of the average gait feature data belonging to each segment, serving as the representative value for that range. This achieves the effect of coarse-grained data processing.

[0153] In the technical solution of this application embodiment, business-related features are classified to obtain multiple category features. Correlation calculations and cluster analysis are performed on the feature parameters within these multiple category features to obtain multiple first evaluation results. Importance calculations are also performed on the multiple category features to obtain multiple second evaluation results. Based on these multiple first and second evaluation results, business-related target influencing factors are determined. The multiple category features include static attribute features and dynamic attribute features. Thus, by determining the degree of influence between related parameters through correlation and cluster analysis of feature parameters within category features, and by determining the importance of different categories through importance analysis of different category features, and by combining the degree of influence between related parameters and the importance of different categories, the weights of different influencing factors can be analyzed under a detailed quantification based on the mixed influence of static and dynamic attribute features. This improves the accuracy of population feature clustering analysis and provides strong support for more objective and continuous comprehensive assessment of patients in a home environment.

[0154] This application also proposes a method for evaluating the impact factor of twin profiles of human physical characteristics based on mixed effects. Figure 4 is a schematic diagram of the hierarchical structure of the method for evaluating the impact factor of twin profiles of human physical characteristics based on mixed effects provided in this application. As shown in Figure 4, the hierarchical structure of this method is divided into four levels, specifically including:

[0155] First level (overall scheduling iterative unit)

[0156] The characteristic parameters related to the population are classified, and the different categories of characteristic parameters after classification are distributed in a multi-level network element architecture or in the functional modules within the network element according to the data server / unit where the data is located.

[0157] Specifically, population-related feature parameters are classified according to static and dynamic attribute classification methods, resulting in static attribute features and dynamic attribute features. Both static and dynamic attribute features include multiple feature parameters. For example, static attribute features can be represented as X = [Gender, Age, Height, BMI], and dynamic attribute features can be represented as G = [g...]. 1 ,g 2 …,g 30 ].

[0158] Second level (fusion of multiple parameters within different categories)

[0159] The degree of influence between multiple parameters within different categories is determined by inverse correlation, and feature category fusion is performed. At the same time, feature parameters within different categories are distributed in a multi-level network element architecture or a functional module within a network element according to the data server / unit where the data is located.

[0160] Specifically, firstly, age parameters are extracted from static attribute features, and multiple gait feature parameters are extracted from gait attribute features included in dynamic attribute features. Correlation calculations are performed on the age parameter and each gait feature parameter to obtain the Pearson correlation coefficient between age and each gait feature parameter. The calculation formula can be found in the above formula (1).

[0161] Secondly, each gait feature parameter is used as a clustering criterion. Each gait feature parameter is sequentially input into the clustering model, allowing the model to cluster multiple gait feature parameters, resulting in multiple clustering results. Each clustering result includes multiple clusters, and each cluster contains at least one gait feature parameter. For the clustering result (D) corresponding to the m-th gait feature parameter... m (1 cluster), obtain the intra-cluster data corresponding to each cluster, calculate the similarity coefficient within each cluster, and take the average of the similarity coefficients of all clusters to obtain the average of the intra-cluster similarity coefficients under the m-th gait feature parameter. The calculation formula can be found in the above formula (2).

[0162] Then, iterate through the D of the m-th gait feature parameters. m For each cluster, calculate the dissimilarity coefficient between each cluster. Take the average of the dissimilarity coefficients between all clusters to obtain the average dissimilarity coefficients between all clusters under the m-th gait feature parameter. The calculation formula can be found in the formula (3) above.

[0163] Finally, based on the average similarity coefficient within all clusters and the average dissimilarity coefficient between all clusters under the m-th gait feature parameter, the silhouette coefficient of the m-th gait feature parameter is calculated. The calculation formula can be found in the formula (4) above.

[0164] Thirdly, after calculating multiple Pearson correlation coefficients between the age parameter and multiple gait feature parameters, the absolute values ​​of the multiple Pearson correlation coefficients can be sorted in ascending order to obtain multiple first-level weight coefficients. The calculation formulas can refer to the above formulas (5) and (6). After calculating multiple contour coefficients of multiple gait feature parameters, the multiple contour coefficients can be sorted in ascending order to obtain multiple second-level weight coefficients. The calculation formulas can refer to the above formulas (7) and (8).

[0165] Then, multiple primary weight coefficients and multiple secondary weight coefficients are merged using the fusion weight method to obtain multiple fusion weight coefficients. The calculation formula can be found in the above formula (9).

[0166] Fourthly, after obtaining multiple fusion weight coefficients, since the fusion weight coefficients represent the positive correlation between age parameters and gait feature parameters, that is, the influence of age on gait feature parameters is directly proportional, the larger the value of a fusion weight coefficient, the greater the influence of age on the gait feature parameter corresponding to the fusion weight coefficient. In order to obtain gait feature parameters with low correlation to age, multiple fusion weight coefficients can be converted by using negative logarithms to convert the direct proportional relationship into an inverse proportional relationship, thereby obtaining multiple reverse fusion weight coefficients. These coefficients are then sorted in reverse order (from maximum to minimum) to obtain multiple first evaluation results. The specific conversion formula can be found in the formula (10) above.

[0167] The third level (integration between different categories)

[0168] The impurity importance calculation method of decision tree is used to fuse the feature categories. At the same time, but not limited to, the different fusion results between the categories are distributed in the network elements of the multi-level architecture or the functional modules within the network elements according to the data server / unit where the data is located.

[0169] First, for each category feature, we can calculate the change in impurity (reduction in impurity) caused by splitting nodes in each decision tree, and accumulate the total change in impurity caused by splitting nodes in all decision trees. Then, we divide the total change in impurity of the category feature by the number of decision trees to obtain the average reduction in impurity of the category feature. This average reduction in impurity can reflect the importance of the category feature.

[0170] Specifically, when calculating the reduction in impurity caused by splitting a node in each decision tree for a given category feature, for each node in each decision tree, when the category feature splits that node, the difference in impurity before and after the split is calculated. This difference is the reduction in impurity caused by the split. Then, the reductions in impurity caused by splitting the category feature into other nodes in the decision tree are accumulated to obtain the total reduction in impurity of the category feature in that decision tree. The total reduction in impurity of the category feature in all decision trees is accumulated and divided by the number of decision trees to obtain the average reduction in impurity of the category feature. This process is repeated to obtain the average reduction in impurity for each category feature. The formula for calculating the reduction in impurity caused by each split can be found in the formula (12) above.

[0171] Then, the average impurity reduction of each category feature is sorted in descending order to obtain multiple second evaluation results.

[0172] Fourth level (intra-category and inter-category fusion iteration)

[0173] A comprehensive evaluation of multiple first and second assessment results is conducted to determine the fusion parameters (target impact factor), which are then fed back to the user module and continuously iterated.

[0174] Specifically, in the first aspect, multiple first evaluation results and multiple second evaluation results are normalized to ensure that the numerical range of each evaluation result is standardized to the [0,1] interval. The numerical values ​​corresponding to the normalized multiple first evaluation results are used to construct a first weighted coefficient sequence in sequence, and the normalized multiple second evaluation results are used as second weighted coefficients to be assigned to the corresponding category features. Then, the numerical values ​​in the first weighted coefficient sequence are added to the second weighted coefficient of each category feature in turn, and the average value is taken as the final weighted coefficient of each category feature, thereby obtaining multiple weighted static attribute features and multiple weighted dynamic attribute features with different weighted coefficients. The calculation formula for the final weighted coefficient of each category feature can be referred to the above formulas (14) and (15).

[0175] Secondly, the dataset is first divided into two parts based on gender, and then further subdivided based on age, height, and weight, resulting in a dataset segmented along four dimensions. The average gait characteristic data belonging to each segment is then calculated as the representative value for that segment. All data in the dataset are replaced with these representative values, thus performing coarse-grained data processing. This approach avoids the influence of individual differences and outliers on the clustering results, improving the robustness of the clustering algorithm. For example, after coarse-grained data processing, the original dataset yielded 1169 representative values.

[0176] Thirdly, based on the static feature data of different samples in the population feature dataset and their corresponding gait feature data, multiple different weighted distance functions are determined based on multiple weighted static attribute features and multiple weighted dynamic attribute features. The calculation formula can be found in the above formula (16).

[0177] Then, based on each weighted distance function, cluster analysis is performed on the coarse-grained dataset using a clustering model to obtain multiple clustering results. These results are then evaluated using the Dispersion Index (DBI). The DBI value for each clustering result is calculated by comparing the compactness within each cluster with the separation between clusters; the calculation formula can be found in formula (20) above. A lower DBI value indicates more compact data within each cluster and higher separation between different clusters, reflecting a better clustering effect.

[0178] Finally, the smallest DBI value is selected from the multiple DBI values ​​corresponding to the multiple clustering results. The clustering result corresponding to this DBI value is the optimal one. The weighted distance function corresponding to the optimal clustering result is used as the target weighted distance function. The weighting coefficients in the target weighted distance function are the target influence factors, which can reflect the importance of static and dynamic attribute features with different weighting coefficients to the data clustering results. Figure 5 shows the data distribution after clustering analysis of the coarse-grained dataset according to the clustering model. From the figure, it can be seen that in each category with similar gait, the distribution of basic information such as gender, height, and weight is relatively concentrated. This proves that individuals with similar static attribute features are likely to exhibit similar gait features.

[0179] Fourthly, the target impact factors are stored, and an update trigger mechanism is set up. When the number of newly added data in the population characteristic dataset reaches a certain threshold, the update mechanism is triggered. Starting from the first level, the updated dataset is reclassified, correlation calculated, clustered, importance calculated, and fusion analyzed, generating new target impact factors. By continuously iterating the above process, the impact factors related to population characteristics can be refined, and the error of population characteristic clustering results can be continuously reduced, thereby improving the accuracy of population characteristic analysis under the mixed influence of static and dynamic attribute characteristics.

[0180] It should be noted that, to demonstrate the effectiveness of coarse-grained processing, cluster analysis was performed on both the original dataset and the coarse-grained processed dataset through experiments. Figure 6 shows a schematic diagram of the clustering results analysis for the original dataset and the coarse-grained processed dataset. The number of cluster centers in the experiment ranged from 5 to 30. The figure contains six curves, corresponding to the DBI values ​​of gait feature data in the original dataset, the DBI values ​​of the original static feature data, the DBI values ​​of gait feature data in the coarse-grained processed dataset, the DBI values ​​of the static feature data in the coarse-grained processed dataset, the average DBI value of the original dataset, and the average DBI value of the coarse-grained processed dataset. The DBI was calculated by performing cluster training on the dataset to obtain the category labels, then calculating the DBI values ​​of the gait feature data and the static feature data separately, and finally averaging the two DBI values ​​to obtain the average DBI value of the dataset. The experimental results show that, across all cluster center counts, the DBI values ​​of the coarse-grained dataset are lower than those of the original dataset. Furthermore, while the DBI values ​​of the gait feature data in the coarse-grained dataset are relatively close to those in the original dataset, the DBI values ​​of the static feature data are significantly reduced. This is the main reason for the better average final result, indicating that coarse-grained processing of the data can effectively improve the performance of clustering results.

[0181] Furthermore, to demonstrate the effectiveness of the custom weighted distance function clustering, two other distance calculation methods are presented: calculating only the distance between static feature data and calculating only the distance between gait feature data (hereinafter referred to as the first and second distance calculation methods). Experiments were conducted for each method, and Figure 7 shows a schematic diagram of the clustering results analysis for the three distance calculation methods. The figure contains nine curves, corresponding to the DBI values ​​of gait feature data, static feature data, and their average DBI value under the first distance calculation method; the DBI values ​​of gait feature data, static feature data, and their average DBI value under the second distance calculation method; and the DBI values ​​of gait feature data, static feature data, and their average DBI value under the custom weighted distance function calculation method. The experimental results show that the custom weighted distance function clustering method outperforms the other two methods. Furthermore, the difference in DBI values ​​between gait feature data and static feature data under the custom weighted distance function calculation method is smaller than that under the other two methods. This demonstrates that the custom weighted distance function clustering method can effectively balance the two types of features, resulting in better clustering results.

[0182] Figure 8 shows a schematic diagram of clustering results analysis for different values ​​of the weighting coefficient in the custom weighted distance function. The figure shows that the DBI value is smallest when λ1 = 0.5 and λ2 = 0.5, corresponding to the best clustering effect. Furthermore, the DBI value on the main diagonal is significantly smaller than the DBI values ​​above and below it; that is, the closer the values ​​of λ1 and λ2 are, the smaller the DBI value. This demonstrates the effectiveness of the custom weighted distance function calculation method in coordinating gait attribute features and static attribute features. In addition to setting the values ​​of λ1 and λ2, the weighting coefficient λ3 of the gender difference function included in function D1 of the custom weighted distance function can also be set. Figure 9 shows a schematic diagram of clustering results analysis for different values ​​of the weighting coefficient of the gender difference function. The figure shows that the value range of λ3 is 0 to 10, with a step size of 0.1. As λ3 is gradually increased, the DBI value reaches its minimum when λ3 = 6.7.

[0183] In the technical solution of this application embodiment, an impact factor evaluation method for population phenotypic twin profiles under mixed effects is proposed. Population characteristics are divided into multiple categories, and multiple parameters are defined within each category. The influence degree between the multiple parameters included in each category is determined by inverse correlation. The importance of different categories is determined by decision tree impurity reduction. The results of the influence degree between related parameters and the importance results of different categories are fused together, thereby determining a method for determining different impact factors and their influence degree under the influence of mixed effects on population phenotypic twin profiles. It also combines the analysis results of the weight of different impact factors under the subdivision quantification based on the mixed influence of static and dynamic demographic characteristics, which can improve the accuracy of population characteristic clustering analysis and provide strong support for home-based disease diagnosis.

[0184] This application also proposes an impact factor assessment device. Figure 10 is a schematic diagram of the structure of the impact factor assessment device provided in this application embodiment. As shown in Figure 10, the device includes:

[0185] Classification unit 1001 is used to classify business-related features to obtain multiple category features; the multiple category features include static attribute features and dynamic attribute features.

[0186] Evaluation unit 1002 is used to perform correlation calculation and cluster analysis on feature parameters in multiple categories to obtain multiple first evaluation results; and to perform importance calculation on multiple categories to obtain multiple second evaluation results.

[0187] Unit 1003 is used to determine business-related target impact factors based on multiple first evaluation results and multiple second evaluation results.

[0188] In some embodiments, the evaluation unit 1002 may include a first evaluation unit and a second evaluation unit; wherein,

[0189] The first evaluation unit is used to calculate the correlation between the age parameter and each of the multiple gait feature parameters to obtain multiple correlation coefficients between the age parameter and the multiple gait feature parameters; determine multiple contour coefficients based on the multiple gait feature parameters; determine multiple fusion weight coefficients based on the multiple correlation coefficients and multiple contour coefficients; and determine multiple first evaluation results based on the multiple fusion weight coefficients; wherein, the static attribute features include the age parameter, the dynamic attribute features include the gait attribute features, and the gait attribute features include multiple gait feature parameters.

[0190] The second evaluation unit is used to determine the average impurity reduction of each category feature through multiple decision trees; wherein, multiple decision trees are used to evaluate the importance of each category feature; the average impurity reduction of the category feature is sorted in descending order to obtain the second evaluation result of the category feature.

[0191] In some implementations, the first evaluation unit is specifically used to cluster multiple gait feature parameters using a clustering model, with each gait feature parameter as an input feature, to obtain multiple clusters for each gait feature parameter; and to determine multiple contour coefficients based on the multiple clusters for each gait feature parameter.

[0192] In some implementations, the first evaluation unit is further specifically configured to, for each gait feature parameter, determine the similarity coefficient within each cluster and the dissimilarity coefficient between each cluster in a plurality of clusters of the gait feature parameter; based on the similarity coefficient within each cluster and the dissimilarity coefficient between each cluster, determine the mean similarity coefficient and the mean dissimilarity coefficient of the plurality of clusters of the gait feature parameter; and based on the mean similarity coefficient and the mean dissimilarity coefficient of the plurality of clusters of the gait feature parameter, determine the contour coefficient of the gait feature parameter.

[0193] In some implementations, the first evaluation unit is further specifically used to sort the absolute values ​​of multiple correlation coefficients in ascending order to obtain multiple primary weight coefficients; sort multiple profile coefficients in ascending order to obtain multiple secondary weight coefficients; and fuse the multiple primary weight coefficients and multiple secondary weight coefficients to obtain multiple fused weight coefficients.

[0194] In some implementations, the first evaluation unit is also specifically used to reverse transform multiple fusion weight coefficients to obtain multiple reverse fusion weight coefficients; and to sort the multiple reverse fusion weight coefficients in descending order to obtain multiple first evaluation results.

[0195] In some implementations, the second evaluation unit is specifically used to determine the reduction in impurity of each of the multiple decision trees for the category feature; and based on the reduction in impurity of each decision tree for the category feature, to determine the average reduction in impurity of the category feature.

[0196] In some implementations, the second evaluation unit is further specifically configured to, for each decision tree, calculate the difference in impurity before and after the split when the category feature is used to split a node in the decision tree, and obtain the reduction in impurity caused by the split; based on the reduction in impurity caused by each split in the decision tree, determine the reduction in impurity of the decision tree for the category feature.

[0197] In some implementations, the determining unit 1003 is specifically used to assign different weights to static attribute features and dynamic attribute features based on multiple first evaluation results and multiple second evaluation results, respectively, to obtain multiple weighted static attribute features and multiple weighted dynamic attribute features; and to determine the target influence factor based on the multiple weighted static attribute features and multiple weighted dynamic attribute features.

[0198] In some implementations, the determining unit 1003 is further specifically used to determine multiple weighted distance functions based on multiple weighted static attribute features and multiple weighted dynamic attribute features; to perform cluster analysis on business-related features based on the multiple weighted distance functions to obtain multiple clustering results; and to determine the target influence factor based on the multiple clustering results.

[0199] In the technical solution of this application embodiment, business-related features are classified to obtain multiple category features. Correlation calculations and cluster analysis are performed on the feature parameters within these multiple category features to obtain multiple first evaluation results. Importance calculations are also performed on the multiple category features to obtain multiple second evaluation results. Based on these multiple first and second evaluation results, business-related target influencing factors are determined. The multiple category features include static attribute features and dynamic attribute features. Thus, by determining the degree of influence between related parameters through correlation and cluster analysis of feature parameters within category features, and by determining the importance of different categories through importance analysis of different category features, and by combining the degree of influence between related parameters and the importance of different categories, the weights of different influencing factors can be analyzed under a detailed quantification based on the mixed influence of static and dynamic attribute features. This improves the accuracy of population feature clustering analysis and provides strong support for more objective and continuous comprehensive assessment of patients in a home environment.

[0200] Those skilled in the art should understand that the functions of each unit in the impact factor evaluation device shown in Figure 10 can be understood with reference to the relevant description of the aforementioned method. The functions of each unit in the impact factor evaluation device shown in Figure 10 can be implemented by a program running on a processor or by specific logic circuits.

[0201] Figure 11 is a schematic diagram of the processing device provided in an embodiment of this application. The processing device may be a terminal device or a network device. The processing device shown in Figure 11 includes a processor 1101, which can call and run computer programs from memory to implement the methods in the embodiments of this application.

[0202] Optionally, as shown in FIG11, the processing device may further include a memory 1102. The processor 1101 may retrieve and run computer programs from the memory 1102 to implement the methods in the embodiments of this application.

[0203] The memory 1102 can be a separate device independent of the processor 1101, or it can be integrated into the processor 1101.

[0204] Optionally, as shown in FIG11, the processing device may further include a transceiver 1103, which the processor 1101 can control to communicate with other devices. Specifically, it can send information or data to other devices or receive information or data sent by other devices.

[0205] The transceiver 1103 may include a transmitter and a receiver. The transceiver 1103 may further include an antenna, and the number of antennas may be one or more.

[0206] The processing device may specifically be the impact factor evaluation device in the embodiments of this application, and the processing device can implement the corresponding processes implemented by the impact factor evaluation device in the various methods of the embodiments of this application. For the sake of brevity, it will not be described in detail here.

[0207] It should be understood that the processor in the embodiments of this application may be an integrated circuit chip with signal processing capabilities. In implementation, the steps of the above method embodiments can be completed by integrated logic circuits in the processor's hardware or by instructions in software form. The processor described above can be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this application. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of this application can be directly embodied in the execution of a hardware decoding processor, or executed by a combination of hardware and software modules in the decoding processor. The software modules can be located in random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, or other mature storage media in the art. The storage medium is located in memory, and the processor reads information from the memory and, in conjunction with its hardware, completes the steps of the above method.

[0208] It is understood that the memory in the embodiments of this application can be volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. The non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. The volatile memory can be random access memory (RAM), which is used as an external cache. By way of example, but not limitation, many forms of RAM are available, such as Static Random Access Memory (SRAM), Dynamic Random Access Memory (DRAM), Synchronous DRAM (SDRAM), Double Data Rate SDRAM (DDR SDRAM), Enhanced Synchronous DRAM (ESDRAM), Synchlink DRAM (SLDRAM), and Direct Rambus RAM (DR RAM). It should be noted that the memory used in the systems and methods described herein is intended to include, but is not limited to, these and any other suitable types of memory.

[0209] It should be understood that the above-described memory is exemplary and not a limiting description. For example, the memory in the embodiments of this application may also be static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct memory bus RAM (DR RAM), etc. That is to say, the memory in the embodiments of this application is intended to include, but is not limited to, these and any other suitable types of memory.

[0210] This application also provides a computer-readable storage medium for storing a computer program. This computer-readable storage medium can be applied to the processing device in the embodiments of this application, and the computer program causes the computer to execute the corresponding processes implemented by the impact factor evaluation device in the various methods of the embodiments of this application; for the sake of brevity, these will not be elaborated further here.

[0211] This application also provides a computer program product, including computer program instructions. This computer program product can be applied to the processing device in the embodiments of this application, and the computer program instructions cause the computer to execute the corresponding processes implemented by the impact factor evaluation device in the various methods of the embodiments of this application; for the sake of brevity, these will not be elaborated further here.

[0212] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0213] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0214] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.

[0215] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0216] In addition, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.

[0217] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0218] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A method for evaluating impact factors, characterized in that, The method includes: classifying business-related features to obtain multiple category features; the multiple category features include static attribute features and dynamic attribute features; performing correlation calculation and cluster analysis on the feature parameters in the multiple category features to obtain multiple first evaluation results; performing importance calculation on the multiple category features to obtain multiple second evaluation results; and determining the business-related target influence factor based on the multiple first evaluation results and the multiple second evaluation results.

2. The method according to claim 1, characterized in that, The static attribute features include an age parameter, and the dynamic attribute features include gait attribute features, which in turn include multiple gait feature parameters. The process of performing correlation calculations and cluster analysis on the feature parameters among the multiple categorical features to obtain multiple first evaluation results includes: calculating the correlation between the age parameter and each of the multiple gait feature parameters to obtain multiple correlation coefficients between the age parameter and the multiple gait feature parameters; determining multiple contour coefficients based on the multiple gait feature parameters; determining multiple fusion weight coefficients based on the multiple correlation coefficients and the multiple contour coefficients; and determining multiple first evaluation results based on the multiple fusion weight coefficients.

3. The method according to claim 2, characterized in that, The step of determining multiple contour coefficients based on the multiple gait feature parameters includes: using each gait feature parameter as an input feature, clustering the multiple gait feature parameters through a clustering model to obtain multiple clusters for each gait feature parameter; and determining the multiple contour coefficients based on the multiple clusters for each gait feature parameter.

4. The method according to claim 3, characterized in that, The step of determining the multiple contour coefficients based on multiple clusters of each gait feature parameter includes: for each gait feature parameter, determining the similarity coefficient within each cluster and the dissimilarity coefficient between each cluster; based on the similarity coefficient within each cluster and the dissimilarity coefficient between each cluster, determining the mean similarity coefficient and the mean dissimilarity coefficient of the multiple clusters of the gait feature parameter; and based on the mean similarity coefficient and the mean dissimilarity coefficient of the multiple clusters of the gait feature parameter, determining the contour coefficient of the gait feature parameter.

5. The method according to claim 2, characterized in that, The step of determining multiple fusion weight coefficients based on the multiple correlation coefficients and the multiple contour coefficients includes: sorting the absolute values ​​of the multiple correlation coefficients in ascending order to obtain multiple primary weight coefficients; sorting the multiple contour coefficients in ascending order to obtain multiple secondary weight coefficients; and fusing the multiple primary weight coefficients and the multiple secondary weight coefficients to obtain the multiple fusion weight coefficients.

6. The method according to claim 2, characterized in that, The step of determining multiple first evaluation results based on the multiple fusion weight coefficients includes: performing a reverse transformation on the multiple fusion weight coefficients to obtain multiple reverse fusion weight coefficients; and sorting the multiple reverse fusion weight coefficients in descending order to obtain the multiple first evaluation results.

7. The method according to claim 1, characterized in that, The step of calculating the importance of the multiple category features to obtain multiple second evaluation results includes: for each category feature, determining the reduction in the average impurity of the category feature through multiple decision trees; wherein, the multiple decision trees are used to evaluate the importance of each category feature; and sorting the reduction in the average impurity of the category feature in descending order to obtain the second evaluation result of the category feature.

8. The method according to claim 7, characterized in that, The step of determining the average impurity reduction of each category feature through multiple decision trees includes: determining the impurity reduction of each decision tree for the category feature; and determining the average impurity reduction of the category feature based on the impurity reduction of each decision tree for the category feature.

9. The method according to claim 8, characterized in that, Determining the reduction in impurity of each of the plurality of decision trees for the category feature includes: for each decision tree, when the category feature is used to split a node in the decision tree, calculating the difference in impurity before and after the split to obtain the reduction in impurity caused by the split; and determining the reduction in impurity of the decision tree for the category feature based on the reduction in impurity caused by each split in the decision tree.

10. The method according to any one of claims 1 to 9, characterized in that, The step of determining the target impact factor related to the business based on the multiple first evaluation results and the multiple second evaluation results includes: assigning different weights to the static attribute features and the dynamic attribute features based on the multiple first evaluation results and the multiple second evaluation results, respectively, to obtain multiple weighted static attribute features and multiple weighted dynamic attribute features; and determining the target impact factor based on the multiple weighted static attribute features and the multiple weighted dynamic attribute features.

11. The method according to claim 10, characterized in that, The step of determining the target influence factor based on the multiple weighted static attribute features and the multiple weighted dynamic attribute features includes: determining multiple weighted distance functions based on the multiple weighted static attribute features and the multiple weighted dynamic attribute features; performing cluster analysis on the business-related features based on the multiple weighted distance functions to obtain multiple clustering results; and determining the target influence factor based on the multiple clustering results.

12. An impact factor assessment device, characterized in that, The apparatus includes: a classification unit for classifying business-related features to obtain multiple category features; the multiple category features include static attribute features and dynamic attribute features; an evaluation unit for performing correlation calculation and cluster analysis on the feature parameters of the multiple category features to obtain multiple first evaluation results; and performing importance calculation on the multiple category features to obtain multiple second evaluation results; and a determination unit for determining a business-related target influence factor based on the multiple first evaluation results and the multiple second evaluation results.

13. A processing apparatus, characterized in that, include: A processor and a memory for storing a computer program, the processor for calling and running the computer program stored in the memory to perform the method as described in any one of claims 1 to 11.

14. A computer-readable storage medium, characterized in that, Used to store a computer program that causes a computer to perform the method as described in any one of claims 1 to 11.

15. A computer program product, characterized in that, It includes computer program instructions that cause a computer to perform the method as described in any one of claims 1 to 11.