Intraoperative stress injury intelligent decision-making method and system based on machine learning
Through sample selection analysis, multicollinearity assessment and information loss assessment, the problem of low screening accuracy of high-risk factors in intelligent decision-making for intraoperative stress injury was solved, and the accuracy of high-risk factor screening and the stability of multi-factor regression model was achieved.
Patent Information
- Application Number
- CN202510347008.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-24
- Publication Date
- 2025-07-11
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
In the prior art, the highly correlated interfering analysis results of independent variables in multi-factor Logistic regression analysis have caused low accuracy in screening high-risk factors in intelligent decision-making for intraoperative stress injury.
Through sample selection analysis, multicollinearity assessment and information loss assessment, sample selection adjustment, principal component dimensionality reduction and gradual introduction of potential factors, obtain the risk weight of high-risk factors, and intelligent decision-making for intraoperative stress injury is combined with high-risk factors.
It improves the screening accuracy of high-risk factors in intelligent decision-making of intraoperative stress injury, reduces multicollinearity, and enhances the stability and accuracy of the multi-factor regression model.
Smart Images

Figure CN120299702A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of electrical digital data processing, and particularly to an intelligent decision-making method and system for intraoperative pressure injury based on machine learning. Background Art
[0002] Posterior spinal scoliosis orthopedic surgery is one of the main treatment methods. However, due to the long operation time and special body position, children have a relatively high risk of intraoperative pressure injury. Its occurrence not only affects the postoperative recovery of patients, but also increases medical costs and nursing burdens. Traditional intraoperative pressure injury risk assessment mainly relies on clinical experience and manual assessment tools, such as the Braden scale. However, these methods have problems such as strong subjectivity and insufficient prediction accuracy. Currently, in the intelligent decision-making system for intraoperative pressure injury based on machine learning, univariate Logistic regression is often used to preliminarily screen potential risk factors related to pressure injury. The multivariate Logistic regression model can integrate multiple risk factors to more accurately predict the occurrence probability of pressure injury.
[0003] In the prior art, by analyzing the electronic medical record data and clinical information data of patients before, during, and after surgery, risk prediction, intelligent decision-making, and intervention suggestions for intraoperative pressure injury are realized.
[0004] For example, the prognostic prediction model for patients with hilar cholangiocarcinoma disclosed in the patent application with the publication number: CN107305596A includes: a carrier for postoperative prognosis of patients with hilar cholangiocarcinoma, and the carrier is used to calculate the scores of risk factors and the 3-year survival rate Y3 and / or 5-year survival rate Y5 of the patients; among them, the risk factors at least include the patient's age X, and the scores of the patient's age, 3-year survival rate, and 5-year survival rate satisfy the relationships in the text.
[0005] For example, the method for establishing a disease risk adjustment model disclosed in the invention patent announcement with the announcement number: CN104992058B includes: historical data of all inpatients in a certain hospital or all hospitals in a certain region. The comorbidities / complications accompanied by the patients when admitted to the hospital, the demographic characteristics of the patients themselves, and the source of the admission status, etc. are integrated into the influencing variable factors of disease treatment. According to the disease diagnosis-related group (DRG) category and the final treatment information of these patients, statistical models are established respectively to numerically predict and analyze the mortality rate, length of hospital stay, and inpatient medical costs of hospital patients.
[0006] However, in the process of implementing the technical solutions of the embodiments of the present application, it is found that the above technologies have at least the following technical problems:
[0007] In the prior art, in multiple-factor Logistic regression analysis, the high correlation among independent variables will interfere with the accuracy of the analysis results. Although the principal component analysis method can reduce the dimension to lower the high correlation among independent variables, it will also cause information loss, resulting in a decrease in the accuracy of the analysis results, and there is a problem of low screening accuracy of high-risk factors in the intelligent decision-making of intraoperative pressure injuries. Summary of the Invention
[0008] By providing an intelligent decision-making method and system for intraoperative pressure injuries based on machine learning, the embodiments of the present application solve the problem of low screening accuracy of high-risk factors in the intelligent decision-making of intraoperative pressure injuries in the prior art, and achieve an improvement in the screening accuracy of high-risk factors in the intelligent decision-making of intraoperative pressure injuries.
[0009] The embodiments of the present application provide an intelligent decision-making method for intraoperative pressure injuries based on machine learning, including the following steps: S1, after performing univariate analysis on the obtained intraoperative pressure injury decision data, screening to obtain potential factors, and analyzing the sample selection analysis data of the potential factors to determine whether to adjust the sample selection; S2, if the sample selection is adjusted, screening the potential factors after the sample selection adjustment to obtain high-risk factors, and evaluating the multicollinearity evaluation data of the obtained high-risk factors to determine whether to perform principal component dimension reduction; S3, if the principal component dimension reduction is not performed, obtaining the risk weights of the high-risk factors, otherwise analyzing the dimension-reduced high-risk factors obtained after the principal component dimension reduction to obtain information loss evaluation data, evaluating the information loss evaluation data to determine whether to gradually introduce potential factors, and determining whether to obtain the risk weights of the high-risk factors based on the multicollinearity evaluation value obtained after the introduction of the potential factors; S4, performing intelligent decision-making on intraoperative pressure injuries according to the risk weights of the high-risk factors and the high-risk factors.
[0010] Further, the sample selection analysis data includes a data category balance value, a minimum observed frequency of factor combinations, a maximum observed frequency under a factor, and a table conversion time; the data category balance value represents quantitative data of the influence degree of the maximum data volume and the minimum data volume under the factor category on the data category balance situation; the multicollinearity evaluation data includes a variance inflation factor, a maximum correlation coefficient, a maximum eigenvalue, and a minimum eigenvalue; the information loss evaluation data includes the image similarity before and after dimension reduction, a main factor retention coefficient, the information entropy before dimension reduction, the information entropy after dimension reduction, and the dimension reduction time.
[0011] Further, the specific method for analyzing the sample selection analysis data of potential factors is as follows: obtaining a category balance coefficient by performing numerical range processing on the data category balance value, where the category balance coefficient represents the quantitative data of the influence degree of the data category balance value on the sample selection effectiveness; obtaining a factor observation coefficient according to the relative deviation relationship between the minimum observed frequency of the factor combination, the maximum observed frequency under the factor, the maximum observed frequency under the preset factor obtained from the preset database, and the minimum observed frequency of the preset factor combination, where the factor observation coefficient represents the quantitative data of the joint influence degree of the minimum observed frequency of the factor combination and the maximum observed frequency under the factor on the sample selection effectiveness; if the table conversion time is within the preset table conversion time, obtaining a conversion time coefficient by performing numerical range restriction on the table conversion time, otherwise setting the conversion time coefficient to 1, where the conversion time coefficient represents the quantitative data of the influence degree of the table conversion time on the sample selection effectiveness; introducing the sample selection weight obtained from the preset database to correct the category balance coefficient, the factor observation coefficient, and the conversion time coefficient to obtain a sample selection evaluation value.
[0012] Further, the specific method for determining whether to perform sample selection adjustment is as follows: A1, determining whether the sample selection evaluation value is less than the second preset sample selection threshold obtained from the preset database. If the sample selection evaluation value is not less than the second preset sample selection threshold obtained from the preset database, perform source file adjustment, otherwise execute A2; A2, determining whether the sample selection evaluation value is less than the first preset sample selection threshold obtained from the preset database. If the sample selection evaluation value is not less than the first preset sample selection threshold obtained from the preset database, perform sample format adjustment, otherwise, after data cleaning, screen the potential factors after data cleaning to obtain high-risk factors; the source file adjustment includes converting the source file format, adjusting the encoding format, and re-obtaining data; the sample format adjustment includes merging missing categories and mean imputation; the data cleaning includes outlier deletion and mean imputation.
[0013] Further, the specific method for evaluating the multicollinearity evaluation data of the obtained high-risk factors is as follows: obtaining a variance inflation coefficient by performing variance inflation approach degree quantification processing according to the variance inflation factor and the preset variance inflation factor obtained from the preset database, where the variance inflation approach degree quantification processing is used to quantify the deviation between the variance inflation factor and the preset variance inflation factor; obtaining a characteristic coefficient by performing eigenvalue approach degree quantification processing according to the minimum eigenvalue and the maximum eigenvalue, where the eigenvalue approach degree quantification processing is used to quantify the deviation between the minimum eigenvalue and the maximum eigenvalue; introducing the multicollinearity weight obtained from the preset database to correct the maximum correlation coefficient, the variance inflation coefficient, and the characteristic coefficient to obtain a multicollinearity evaluation value.
[0014] Further, the specific method for determining whether to perform principal component dimensionality reduction is as follows: Determine whether the multicollinearity evaluation value is less than the preset multicollinearity threshold obtained from the preset database. If the multicollinearity evaluation value is not less than the preset multicollinearity threshold obtained from the preset database, perform principal component dimensionality reduction on the high-risk factors; otherwise, obtain the risk weights of the high-risk factors. The specific method for obtaining the risk weights of the high-risk factors is as follows: Analyze the odds ratio based on the regression coefficients corresponding to the high-risk factors in the multiple factor regression model, and set the proportion of the odds ratio of each high-risk factor in the sum of all odds ratios as the risk weight of the corresponding high-risk factor.
[0015] Further, the specific method for evaluating the information loss evaluation data is as follows: Obtain the information entropy coefficients before and after dimensionality reduction through normalization of the information entropy before dimensionality reduction and the information entropy after dimensionality reduction. The information entropy coefficients before and after dimensionality reduction represent the quantitative data of the influence degree of the information entropy before dimensionality reduction and the information entropy after dimensionality reduction on the information loss degree together; Quantify the proximity degree of the dimensionality reduction time to obtain the dimensionality reduction time coefficient according to the dimensionality reduction time and the preset dimensionality reduction time obtained from the preset database. The quantification process of the proximity degree of the dimensionality reduction time is used to quantify the deviation between the dimensionality reduction time and the preset dimensionality reduction time; Introduce the information loss weight obtained from the preset database to correct the image similarity before and after dimensionality reduction, the main factor retention coefficient, the information entropy coefficients before and after dimensionality reduction, and the dimensionality reduction time coefficient to obtain the information loss evaluation value.
[0016] Further, the specific process for determining whether to gradually introduce potential factors is as follows: C1, Determine whether the information loss evaluation value is less than the preset information loss threshold obtained from the preset database. If the information loss evaluation value is less than the preset information loss threshold obtained from the preset database, obtain the risk weights of the high-risk factors; otherwise, execute C2; C2, Determine whether the significance level of the potential factor is less than the initial significance level threshold. If the significance level of the potential factor is less than the initial significance level threshold, introduce the corresponding potential factor into the multiple factor regression model; otherwise, do not introduce the potential factor and execute C3; C3, Determine whether a potential factor has been introduced. If no potential factor has been introduced, adjust the initial significance level threshold to obtain the significance level threshold; otherwise, determine whether to obtain the risk weights of the high-risk factors based on the multicollinearity evaluation value obtained after the introduction of the potential factor. The significance level threshold is obtained by correcting the correlation coefficient of the introduced variable, the sample size coefficient, the information loss deviation coefficient, and the initial significance level threshold with the significant weight obtained from the preset database. The sample size coefficient represents the quantitative data of the influence degree of the sample size after introducing the variable and the preset sample size on the significance level threshold together. The information loss deviation coefficient represents the quantitative data of the influence degree of the information loss evaluation value and the preset information loss threshold on the significance level threshold together.
[0017] Further, the specific method for determining whether to obtain the risk weight of high-risk factors based on the multicollinearity evaluation value obtained after introducing potential factors is as follows: B1, evaluate the multicollinearity evaluation value based on the multicollinearity evaluation data obtained after introducing potential factors; B2, determine whether the multicollinearity evaluation value obtained after introducing potential factors is less than the preset multicollinearity threshold obtained from the preset database. If the multicollinearity evaluation value obtained after introducing potential factors is not less than the preset multicollinearity threshold obtained from the preset database, perform principal component dimensionality reduction on high-risk factors and gradually adjust the dimensionality reduction parameters of the principal component dimensionality reduction. Otherwise, execute B4; B3, if the information loss evaluation value obtained after adjusting the dimensionality reduction parameters is not less than the preset information loss threshold obtained from the preset database, gradually adjust the preset information loss threshold. Otherwise, execute B4; B4, determine the optimal number of high-risk factors based on the number of high-risk factors with non-zero coefficients under the optimal regularization parameter obtained by the cross-validation method, and set the high-risk factors with non-zero coefficients under the optimal regularization parameter as the final high-risk factors, and obtain the risk weight accordingly.
[0018] The embodiment of the present application provides an intraoperative pressure injury intelligent decision-making system based on machine learning, including: a sample selection and analysis module, a multicollinearity evaluation module, an information loss evaluation module, and a threshold adjustment module; wherein, the sample selection and analysis module is used to perform single-factor analysis on the obtained intraoperative pressure injury decision data to screen out potential factors, and analyze the sample selection and analysis data of the potential factors to determine whether to adjust the sample selection; the multicollinearity evaluation module is used to, if the sample selection is adjusted, screen the potential factors after the sample selection adjustment to obtain high-risk factors, and evaluate the multicollinearity evaluation data of the obtained high-risk factors to determine whether to perform principal component dimensionality reduction; the information loss evaluation module is used to, if the principal component dimensionality reduction is not performed, obtain the risk weight of the high-risk factors, otherwise analyze the dimensionality-reduced high-risk factors obtained after the principal component dimensionality reduction to obtain information loss evaluation data, and evaluate the information loss evaluation data to determine whether to gradually introduce potential factors, and determine whether to obtain the risk weight of the high-risk factors based on the multicollinearity evaluation value obtained after introducing the potential factors; the intelligent decision-making module is used to perform intraoperative pressure injury intelligent decision-making based on the risk weight of the high-risk factors and the high-risk factors.
[0019] One or more technical solutions provided in the embodiments of the present application have at least the following technical effects or advantages:
[0020] 1. Adjust the sample selection through sample selection analysis, then perform principal component dimensionality reduction based on the evaluation of multicollinearity, and finally gradually introduce potential factors based on the evaluation of information loss. Obtain the risk weights of high-risk factors based on the multicollinearity evaluation value after the introduction of potential factors, and combine with high-risk factors to make intelligent decisions on intraoperative pressure injuries. This reduces the multicollinearity between high-risk factors, thereby improving the screening accuracy of high-risk factors in the intelligent decision-making of intraoperative pressure injuries, and effectively solving the problem of low screening accuracy of high-risk factors in the intelligent decision-making of intraoperative pressure injuries in the prior art.
[0021] 2. Obtain the category balance coefficient by performing numerical range processing on the data category balance value, then obtain the factor observation coefficient according to the minimum observed frequency of the factor combination and the maximum observed frequency under the factor. Next, perform numerical range limitation on the table conversion time to obtain the conversion time coefficient. Finally, comprehensively process the category balance coefficient, factor observation coefficient, conversion time coefficient, and the sample selection weight obtained from the preset database to obtain the sample selection evaluation value, thereby effectively identifying sample anomalies and improving the effectiveness of sample selection.
[0022] 3. Obtain the variance inflation coefficient through the relative influence relationship between the variance inflation factor and the preset variance inflation factor obtained from the preset database, then analyze the data distribution situation according to the minimum eigenvalue and the maximum eigenvalue to obtain the eigenvalue coefficient. Finally, process the maximum correlation coefficient, variance inflation coefficient, eigenvalue coefficient, and the multicollinearity weight obtained from the preset database to obtain the multicollinearity evaluation value, thereby quantifying the collinearity influence between high-risk factors and improving the stability of multi-factor analysis. Description of the Drawings
[0023] Figure 1 It is a flowchart of the intelligent decision-making method for intraoperative pressure injury based on machine learning provided by the embodiment of the present application;
[0024] Figure 2 It is a schematic structural diagram of the intelligent decision-making system for intraoperative pressure injury based on machine learning provided by the embodiment of the present application. Detailed Embodiments
[0025] The embodiments of the present application provide an intelligent decision-making method and system for intraoperative pressure injuries based on machine learning, which solve the problem of low screening accuracy of high-risk factors in intelligent decision-making of intraoperative pressure injuries in the prior art. By analyzing the sample selection analysis data of potential factors to determine whether to adjust the sample selection, then evaluating the multicollinearity of high-risk factors to determine whether to perform principal component dimensionality reduction, then evaluating the information loss to determine whether to gradually introduce potential factors, and determining whether to obtain the risk weights of high-risk factors based on the multicollinearity evaluation value obtained after the introduction of potential factors, and finally making an intelligent decision on intraoperative pressure injuries based on the risk weights and high-risk factors of high-risk factors, the screening accuracy of high-risk factors in intelligent decision-making of intraoperative pressure injuries is improved.
[0026] The technical solution in the embodiments of the present application aims to solve the problem of low screening accuracy of high-risk factors in the intelligent decision-making of intraoperative pressure injuries. The general idea is as follows:
[0027] Adjust the sample selection through sample selection analysis, then perform principal component dimensionality reduction based on multicollinearity evaluation, and finally gradually introduce potential factors based on information loss evaluation, obtain the risk weights of high-risk factors based on the multicollinearity evaluation value after the introduction of potential factors, and make an intelligent decision on intraoperative pressure injuries in combination with high-risk factors, achieving the effect of improving the screening accuracy of high-risk factors in intelligent decision-making of intraoperative pressure injuries.
[0028] To better understand the above technical solution, the above technical solution will be described in detail below in combination with the accompanying drawings of the specification and specific implementation manners.
[0029] As Figure 1 shown, it is a flowchart of an intelligent decision-making method for intraoperative pressure injuries based on machine learning provided by the embodiments of the present application. The method includes the following steps: S1, sample selection analysis: perform univariate analysis on the obtained intraoperative pressure injury decision data to screen out potential factors, and analyze the sample selection analysis data of the potential factors to determine whether to adjust the sample selection; S2, multicollinearity evaluation: if the sample selection is adjusted, screen the potential factors after sample selection adjustment to obtain high-risk factors, and evaluate the multicollinearity evaluation data of the obtained high-risk factors to determine whether to perform principal component dimensionality reduction; S3, information loss evaluation: if the principal component dimensionality reduction is not performed, obtain the risk weights of high-risk factors, otherwise analyze the reduced high-risk factors obtained after principal component dimensionality reduction to obtain information loss evaluation data, evaluate the information loss evaluation data to determine whether to gradually introduce potential factors, and determine whether to obtain the risk weights of high-risk factors based on the multicollinearity evaluation value obtained after the introduction of potential factors; S4, intelligent decision-making: make an intelligent decision on intraoperative pressure injuries based on the risk weights and high-risk factors of high-risk factors.
[0030] It should be added that the intraoperative pressure injury decision-making data represents the preoperative, intraoperative, and postoperative data of children collected from multiple sources such as the hospital information system, electronic medical records, surgical anesthesia system, and monitoring equipment. Specifically, it includes basic information such as age, gender, weight, Body Mass Index (BMI), preoperative examination results (blood routine, biochemical indicators), surgical information (surgical time, anesthesia method, body position), intraoperative monitoring data (body temperature, blood pressure, heart rate, oxygen saturation), and nursing records (skin condition, decompression measures); potential factors and risk factors represent the intraoperative pressure injury decision-making data related to intraoperative pressure injury, such as surgical time, preoperative hypoproteinemia, BMI, intraoperative blood loss, etc.
[0031] It should be added that the sample selection analysis data includes the data category balance value, the minimum observed frequency of factor combinations, the maximum observed frequency under a factor, and the table conversion time; the data category balance value represents the quantitative data of the influence degree of the maximum and minimum data volumes under the factor category on the data category balance situation; the data category balance value represents the balance degree of the data of a certain factor, which can be considered by taking the ratio of the maximum and minimum frequencies of different data occurrences; the multicollinearity evaluation data includes the variance inflation factor, the maximum correlation coefficient, the maximum eigenvalue, and the minimum eigenvalue; the information loss evaluation data includes the image similarity before and after dimensionality reduction, the main factor retention coefficient, the information entropy before dimensionality reduction, the information entropy after dimensionality reduction, and the dimensionality reduction time.
[0032] Specifically, the minimum observed frequency of factor combinations is obtained by counting the observed frequencies under any two factor category combinations in the table. For example, the minimum frequency of the data in the cell of the combination of "gender" and "age"; the maximum observed frequency under a factor is obtained by counting the observed frequencies under any two factor category combinations in the table. For example, in the category of "18 years old", the frequency of "female" is 20 and the frequency of "male" is 10, then the maximum observed frequency under the factor is 20; the table conversion time is obtained by calculating the time recorded by the timer at the start and end of the table conversion operation (end time of the table conversion operation - start time of the table conversion operation).
[0033] Specifically, the variance inflation factor is calculated by using the VIF function in the Statistical Product and Service Solutions (SPSS) software; the maximum correlation coefficient is represented by the maximum value of the Pearson correlation coefficients obtained by statistically calculating the absolute values of the Pearson correlation coefficients between any two factors; the maximum eigenvalue and the minimum eigenvalue are obtained by performing factor analysis in the SPSS software after encoding the non-numerical data for features.
[0034] Feature encoding means converting non-numerical data into numerical data; for categorical variables (such as gender), one-hot encoding can be used to convert them into numerical values; for time variables, time features (such as date, time, day of the week, etc.) can be extracted and converted into numerical values; for text data, text vectorization techniques (such as TF-IDF, word embeddings, etc.) can be used to convert it into numerical features.
[0035] Specifically, the image similarity before and after dimensionality reduction is obtained by calculating the cosine similarity of the kernel density estimation maps before and after dimensionality reduction; the main factor retention coefficient is obtained by taking the ratio of the sum of the peak area and the tail area of the kernel density estimation map before dimensionality reduction to the sum of the peak area and the tail area after dimensionality reduction. The sum of the peak area and the tail area before dimensionality reduction is obtained by summing the peak area and the tail area; the peak area is represented by the product of the peak half-width and the peak; the peak half-width refers to the distance between the abscissas of two points corresponding to the height that is half of the peak height found by interpolation on both sides of the determined peak position (maximum value position); the tail area is obtained by integrating the probability density greater than the preset tail threshold; the preset tail threshold is obtained from the preset database and can be set to the 90th percentile; the kernel density estimation map is drawn using the Seaborn library in Python.
[0036] The information entropy is obtained by summing the single-factor information entropies of all factors. The single-factor information entropy is obtained by substituting the probability distribution of all possible values of any factor into the information entropy calculation formula for summation; the calculation formula for the single-factor information entropy is: where i represents the number of possible values of any factor, i = 1, 2, …, n, n represents the total number of possible values of any factor, H(X) represents the single-factor information entropy, and P(x i ) represents the probability of the possible value x i of any factor X.
[0037] The dimensionality reduction time is obtained by calculating the dimensionality reduction start time and the dimensionality reduction end time recorded by the timer (the dimensionality reduction end time minus the dimensionality reduction start time).
[0038] In this embodiment, the intraoperative pressure injury decision-making data stored in the data platform is converted into a table format through SPSS software, that is, the above table is obtained; the data platform stores structured and unstructured data by combining a relational database (such as MySQL) and a non-relational database (such as MongoDB), providing support for subsequent analysis.
[0039] Through single-factor analysis and multi-factor Logistic regression analysis, potential factors and high-risk factors related to intraoperative pressure injury can be screened, thereby improving the accuracy of intelligent decision-making for intraoperative pressure injury; through multicollinearity assessment and principal component dimensionality reduction, the correlation between high-risk factors can be reduced, and the factor screening ability of the multi-factor regression model can be improved; by gradually introducing potential factors and adjusting the significance level threshold, the information loss caused by principal component dimensionality reduction can be reduced; furthermore, the screening accuracy of high-risk factors in intelligent decision-making for intraoperative pressure injury is improved.
[0040] Furthermore, the specific method for analyzing the sample selection analysis data of potential factors is as follows: The category balance coefficient is obtained by performing numerical range processing on the data category balance value, that is The category balance coefficient represents the quantitative data of the influence degree of the data category balance value on the effectiveness of sample selection; the factor observation coefficient is obtained according to the relative deviation relationship between the minimum observed frequency of the factor combination, the maximum observed frequency under the factor, the maximum observed frequency under the preset factor obtained from the preset database, and the minimum observed frequency of the preset factor combination, that is The factor observation coefficient represents the quantitative data of the influence degree of the minimum observed frequency of the factor combination and the maximum observed frequency under the factor on the effectiveness of sample selection; if the table conversion time is within the preset table conversion time, the conversion time coefficient, that is SUD, is obtained by performing numerical range limitation on the table conversion time, otherwise the conversion time coefficient is recorded as 1, and the conversion time coefficient represents the quantitative data of the influence degree of the table conversion time on the effectiveness of sample selection; the sample selection weight obtained from the preset database is introduced to correct the category balance coefficient, the factor observation coefficient, and the conversion time coefficient to obtain the sample selection evaluation value.
[0041] Among them, the specific limiting expression of the sample selection evaluation value:
[0042]
[0043] In the formula, YB represents the sample selection evaluation value, PHE represents the data category balance value, SUD represents the table conversion time coefficient, α1 represents the category balance weight, α2 represents the factor observation weight, α3 represents the table conversion time weight, YGC max represents the maximum observed frequency under the factor, YGC0 represents the maximum observed frequency under the preset factor, YZG min represents the minimum observed frequency of the factor combination, YZG0 represents the minimum observed frequency of the preset factor combination, SUDU represents the table conversion time, SUDU min represents the preset minimum table conversion time, SUDU max represents the preset maximum table conversion time.
[0044] In this embodiment, whether there is a complete separation phenomenon is observed through the maximum observed frequency under a factor. Therefore, the maximum observed frequency under the preset factor is represented by 90% of the total data volume in this category; whether there is an empty cell is observed through the minimum observed frequency of the factor combination. Therefore, the preset minimum observed frequency of the factor combination is set to a value close to 0, generally taking 0.01; the preset minimum table conversion time is represented by the minimum value of the table conversion times when the tables are successfully converted within the historical time period; the preset maximum table conversion time is represented by the maximum value of the table conversion times when the tables are successfully converted within the historical time period.
[0045] This algorithm is obtained by processing multiple independent variables (data category balance value, minimum observed frequency of factor combination, maximum observed frequency under a factor, table conversion time). There are mutual influence relationships among these independent variables; when the distance between the data category balance value and 1 is greater, it means that the number of samples in some categories is smaller, which may lead to a lower minimum observed frequency of the factor combination; when the minimum observed frequency of the factor combination is lower, it means that the number of samples in some factor combinations in the data set is smaller, and when the maximum observed frequency under a factor is higher, it means that some categories or characteristic values dominate the data, which may cause the model to overfit the factor categories with higher observed frequencies and ignore the factor categories with lower observed frequencies; the smaller the minimum observed frequency of the factor combination, special processing methods (such as filling, imputation, or recombination) may be required during data conversion, thus increasing the table conversion time; when the distance between the data category balance value and 1 is farther, more adjustments may be needed during the data preprocessing stage (such as oversampling the minority class or undersampling the majority class), which may increase the table conversion time; when the maximum observed frequency under a factor is larger, it may lead to a very large data table generated during data conversion, which may increase the table conversion time; the larger the sample selection evaluation value, the higher the effectiveness of sample selection.
[0046] The category balance weight, factor observation weight, and table conversion time weight are obtained from the preset database, and the sum of the three is 1. For example, the data category balance value and the preset category balance weight form a mapping set, and the real-time data category balance value is input into the mapping set to obtain the corresponding category balance weight; the minimum observed frequency of the factor combination, the maximum observed frequency under a factor, and the preset factor observation weight form a mapping set, and the real-time minimum observed frequency of the factor combination and the maximum observed frequency under a factor are input into the mapping set to obtain the corresponding factor observation weight; the table conversion time and the preset table conversion time weight form a mapping set, and the real-time table conversion time is input into the mapping set to obtain the corresponding table conversion time weight; the mapping relationships therein can be one-to-one or many-to-one relationships.
[0047] By quantifying the data category balance value and the factor observation frequency, situations of data category imbalance and abnormal factor observation frequency can be identified to ensure the representativeness of sample selection; by converting the table time, abnormal table conversion situations can be identified. When the table conversion time is less than the preset minimum table conversion time, it may indicate that the source file format does not match the SPSS software or there are abnormalities in the source file encoding. When the table conversion time is greater than the preset maximum table conversion time, it may indicate that there are more abnormal situations in the source file and more data processing time is required, thereby improving the effectiveness of sample selection and avoiding low accuracy of multi-factor analysis caused by subsequent sample data problems.
[0048] Further, the specific method for determining whether to adjust sample selection is as follows: A1, determine whether the sample selection evaluation value is less than the second preset sample selection threshold obtained from the preset database. If the sample selection evaluation value is less than the second preset sample selection threshold obtained from the preset database, perform source file adjustment; otherwise, execute A2; A2, determine whether the sample selection evaluation value is less than the first preset sample selection threshold obtained from the preset database. If the sample selection evaluation value is less than the first preset sample selection threshold obtained from the preset database, perform sample format adjustment; otherwise, perform data cleaning, and screen the potential factors after data cleaning to obtain high-risk factors; source file adjustment includes converting the source file format, adjusting the encoding format, and re-obtaining data; sample format adjustment includes merging missing categories and mean imputation; data cleaning includes outlier deletion and mean imputation.
[0049] In this embodiment, the second preset sample selection threshold is represented by the sum value between the average value of the qualified sample selection evaluation values in the historical database and twice the variance; the first preset sample selection threshold is represented by the difference value between the average value of the qualified sample selection evaluation values in the historical database and twice the variance.
[0050] Converting the source file format converts the database into a file format that can be matched by the SPSS software (such as CSV format) through the phpMyAdmin tool.
[0051] Before adjusting the encoding format, back up the entire database to prevent data loss or damage during the adjustment process; adjust the encoding format of the database by modifying the NLS parameter in the "Advanced" tab of "Oracle SQL Developer" in the Oracle software.
[0052] Merging missing categories merges the categories with a data missing amount greater than the preset missing threshold under the same factor into the "other" category through the "Recode into Different Variables" function of the SPSS software; the preset missing threshold is set according to specific application requirements. For example, it can be set to 40% of the sample data volume of this factor.
[0053] Mean imputation performs mean imputation on numerical variables through the "Replace Missing Values" function in SPSS software, that is, replacing the missing values with the average of the non-missing values of the variable.
[0054] Outlier deletion identifies outliers through methods such as the Z-score in SPSS software and allows for deletion. For example, in the "Analyze" menu of SPSS software, select "Descriptive Statistics", then select the "Descriptives" option. In the pop-up dialog box, select the variables to be analyzed and check the option "Save standardized values as variables". SPSS software can add a new column containing Z-scores to the original data table. After obtaining the Z-scores, outliers can be identified through sorting or filtering functions. For example, using a conditional statement (such as "IF ABS(Z_score)>3") to filter out data points with an absolute Z-score greater than 3, and these data points are the possible outliers.
[0055] In the multiple-factor regression model, the potential factors after data cleaning are screened to obtain high-risk factors. The t-statistic and degrees of freedom corresponding to the potential factors are obtained through SPSS software, and the t-statistic and degrees of freedom corresponding to the potential factors are input into SPSS software to obtain the P-value. It is judged whether the P-value of the corresponding potential factor is less than the initial significance level threshold. If the P-value of the corresponding potential factor is less than the initial significance level threshold, the corresponding potential factor is set as a high-risk factor; otherwise, it is not set as a high-risk factor.
[0056] Through source file adjustment, sample format adjustment, and data cleaning, the data quality can be significantly improved; through measures such as merging missing categories and mean imputation, the analysis bias caused by missing values or outliers can be reduced, thereby improving the accuracy of subsequent multiple-factor analysis.
[0057] Furthermore, the specific method for evaluating the multicollinearity evaluation data of the obtained high-risk factors is as follows: The variance inflation coefficient is obtained by quantifying the variance inflation approximation degree according to the variance inflation factor and the preset variance inflation factor obtained from the preset database, that is The variance inflation approximation degree quantification is used to quantify the deviation between the variance inflation factor and the preset variance inflation factor; the eigenvalue coefficient is obtained by quantifying the eigenvalue approximation degree according to the minimum eigenvalue and the maximum eigenvalue, that is The eigenvalue approximation degree quantification is used to quantify the deviation between the minimum eigenvalue and the maximum eigenvalue; the multicollinearity evaluation value is obtained by introducing the multicollinearity weight obtained from the preset database to correct the maximum correlation coefficient, variance inflation coefficient, and eigenvalue coefficient.
[0058] Among them, the specific limiting expression of the multicollinearity evaluation value is:
[0059]
[0060] In the formula, DUO represents the multicollinearity evaluation value, VIF represents the variance inflation factor, XGX represents the maximum correlation coefficient, TZZ min represents the minimum eigenvalue, TZZ max represents the maximum eigenvalue, VIF0 represents the preset variance inflation factor, β1 represents the variance inflation weight, β2 represents the correlation coefficient weight, and β3 represents the eigenvalue weight.
[0061] In this embodiment, the preset variance inflation factor is set according to specific application requirements. For example, the preset variance inflation factor can be set to 10.
[0062] This algorithm involves processing multiple independent variables (maximum correlation coefficient, variance inflation coefficient, characteristic coefficient), and there are mutual influence relationships among these independent variables; the larger the maximum correlation coefficient, the stronger the possible linear relationship between variables, which may lead to an increase in the variance inflation coefficient, that is, there may be a multicollinearity problem; the larger the variance inflation coefficient, the more multicollinearity exists among the independent variables, which may lead to a larger characteristic coefficient; if the maximum correlation coefficient increases, the importance of the corresponding two factors in the multiple factor regression model may increase, which may lead to an increase in the maximum eigenvalue, and then increase the characteristic coefficient; the larger the multicollinearity evaluation value, the greater the degree of multicollinearity.
[0063] The variance inflation weight, correlation coefficient weight, and eigenvalue weight are obtained from the preset database, and the sum of the three is 1. For example, the variance inflation coefficient and the preset variance inflation weight form a mapping set, and the real-time variance inflation coefficient is input into the mapping set to obtain the corresponding variance inflation weight; the maximum correlation coefficient and the preset correlation coefficient weight form a mapping set, and the real-time maximum correlation coefficient is input into the mapping set to obtain the corresponding correlation coefficient weight; the minimum eigenvalue and the maximum eigenvalue and the preset eigenvalue weight form a mapping set, and the real-time minimum eigenvalue and maximum eigenvalue are input into the mapping set to obtain the corresponding eigenvalue weight; the mapping relationship therein can be one-to-one or many-to-one.
[0064] Through the variance inflation coefficient, characteristic coefficient, and maximum correlation coefficient, the collinearity influence among high-risk factors is quantified, and then data support is provided for dimensionality reduction of high-risk factors, improving the stability of multi-factor analysis.
[0065] Further, the specific method for determining whether to perform principal component dimensionality reduction is as follows: Determine whether the multicollinearity evaluation value is less than the preset multicollinearity threshold obtained from the preset database. If the multicollinearity evaluation value is not less than the preset multicollinearity threshold obtained from the preset database, perform principal component dimensionality reduction on the high-risk factors; otherwise, obtain the risk weights of the high-risk factors. The specific method for obtaining the risk weights of the high-risk factors is as follows: Analyze the odds ratio based on the regression coefficient corresponding to the high-risk factor in the multiple-factor regression model, and set the proportion of the odds ratio of each high-risk factor in the sum of all odds ratios as the risk weight of the corresponding high-risk factor.
[0066] In this embodiment, the preset multicollinearity threshold is represented by the average value of the qualified multicollinearity evaluation values within the historical time period.
[0067] Principal component dimensionality reduction is achieved through principal component analysis; principal component analysis is a commonly used data dimensionality reduction technique that projects the original data into a new low-dimensional space to retain as much data variability as possible while removing redundant information.
[0068] The calculation formula for the odds ratio is: OR = exp(β), where β represents the regression coefficient and OR represents the odds ratio.
[0069] Dimensionality reduction through principal component analysis ensures that the multicollinearity problem among high-risk factors is effectively solved, improving the interpretability and factor screening ability of the multiple-factor regression model; through the calculation of risk weights, the influence degree of each high-risk factor on postoperative pressure injury is quantified, thereby optimizing the feature selection in the multiple-factor regression model.
[0070] Further, the specific method for evaluating the information loss evaluation data is as follows: Normalize the information entropy before dimensionality reduction and the information entropy after dimensionality reduction to obtain the information entropy coefficient before and after dimensionality reduction, that is The information entropy coefficient before and after dimensionality reduction represents the quantitative data of the influence degree of the information entropy before dimensionality reduction and the information entropy after dimensionality reduction on the information loss degree; Quantify the proximity degree of the dimensionality reduction time based on the dimensionality reduction time and the preset dimensionality reduction time obtained from the preset database to obtain the dimensionality reduction time coefficient, that is The quantification process of the proximity degree of the dimensionality reduction time is used to quantify the deviation between the dimensionality reduction time and the preset dimensionality reduction time; Introduce the information loss weight obtained from the preset database to correct the image similarity before and after dimensionality reduction, the main factor retention coefficient, the information entropy coefficient before and after dimensionality reduction, and the dimensionality reduction time coefficient to obtain the information loss evaluation value.
[0071] Among them, the specific limit expression of the information loss evaluation value is:
[0072]
[0073] In the formula, XS represents the information loss evaluation value, TXS represents the image similarity before and after dimensionality reduction, YBL represents the main factor retention coefficient, XXS q represents the information entropy before dimensionality reduction, XXS h represents the information entropy after dimensionality reduction, JWS represents the dimensionality reduction time, JWS0 represents the preset dimensionality reduction time, γ1 represents the image similarity weight, γ2 represents the factor retention weight, γ3 represents the information entropy weight, and γ4 represents the dimensionality reduction time weight.
[0074] In this embodiment, the preset dimensionality reduction time is represented by the average value of the dimensionality reduction times in the qualified information loss evaluation values in the historical database.
[0075] This algorithm is obtained by processing multiple independent variables (information entropy coefficients before and after dimensionality reduction, dimensionality reduction time, image similarity before and after dimensionality reduction, main factor retention coefficient). There are mutual influence relationships among these independent variables; the longer the dimensionality reduction time, the more information loss may occur, resulting in the information entropy coefficients before and after dimensionality reduction and the image similarity before and after dimensionality reduction being less close to 1; the larger the main factor retention coefficient, the higher the retention degree of key information in the dimensionality reduction process, which may lead to the information entropy coefficients before and after dimensionality reduction and the image similarity before and after dimensionality reduction being closer to 1; the closer the information entropy coefficients before and after dimensionality reduction are to 1, the higher the information retention degree in the dimensionality reduction process, which may lead to the image similarity before and after dimensionality reduction and the main factor retention coefficient being closer to 1; the smaller the information loss evaluation value, the smaller the information loss degree between the high-risk factors after principal component dimensionality reduction and the high-risk factors before principal component dimensionality reduction.
[0076] The image similarity weight, factor retention weight, information entropy weight, and dimensionality reduction time weight are obtained from a preset database, and the sum of the four is 1. For example, the image similarity before and after dimensionality reduction forms a mapping set with the preset image similarity weight, and the real-time image similarity before and after dimensionality reduction is input into the mapping set to obtain the corresponding image similarity weight; the main factor retention coefficient forms a mapping set with the preset factor retention weight, and the real-time main factor retention coefficient is input into the mapping set to obtain the corresponding factor retention weight; the information entropy coefficients before and after dimensionality reduction form a mapping set with the preset information entropy weight, and the real-time information entropy coefficients before and after dimensionality reduction are input into the mapping set to obtain the corresponding information entropy weight; the dimensionality reduction time forms a mapping set with the preset dimensionality reduction time weight, and the real-time dimensionality reduction time is input into the mapping set to obtain the corresponding dimensionality reduction time weight; the mapping relationships therein can be one-to-one or many-to-one relationships.
[0077] By comprehensively considering the information entropy before and after dimensionality reduction, dimensionality reduction time, image similarity, and main factor retention, the information loss degree in the dimensionality reduction process can be comprehensively evaluated; through data analysis means such as normalization processing, the evaluation process is simplified, the evaluation efficiency is improved, and the efficiency and accuracy of data processing are improved.
[0078] Further, the specific process for determining whether to gradually introduce potential factors is as follows: C1. Determine whether the information loss evaluation value is less than the preset information loss threshold obtained from the preset database. If the information loss evaluation value is less than the preset information loss threshold obtained from the preset database, obtain the risk weight of the high-risk factor; otherwise, execute C2. C2. Determine whether the significance level of the potential factor is less than the initial significance level threshold. If the significance level of the potential factor is less than the initial significance level threshold, introduce the corresponding potential factor into the multi-factor regression model; otherwise, do not introduce the potential factor and execute C3. C3. Determine whether a potential factor has been introduced. If no potential factor has been introduced, adjust the initial significance level threshold to obtain the significance level threshold; otherwise, determine whether to obtain the risk weight of the high-risk factor based on the multicollinearity evaluation value obtained after the introduction of the potential factor. The significance level threshold is obtained by correcting the correlation coefficient of the introduced variable, the sample size coefficient, the information loss deviation coefficient, and the initial significance level threshold with the significance weight obtained from the preset database. The sample size coefficient (i.e., ) is obtained from the relative deviation relationship between the sample size after the introduction of the variable and the preset sample size obtained from the preset database. The sample size coefficient represents the quantitative data of the degree of influence of the sample size after the introduction of the variable and the preset sample size on the significance level threshold. The information loss deviation coefficient (i.e., ) is obtained from the relative deviation relationship between the information loss evaluation value and the preset information loss threshold obtained from the preset database. The information loss deviation coefficient represents the quantitative data of the degree of influence of the information loss evaluation value and the preset information loss threshold on the significance level threshold.
[0079] Among them, the specific method for obtaining the significance level threshold is:
[0080]
[0081] In the formula, XZ0 ' represents the significance level threshold, XZ0 represents the initial significance level threshold, BXG represents the correlation coefficient of the introduced variable, YAL represents the sample size after the introduction of the variable, XS represents the information loss evaluation value, XS0 represents the preset information loss threshold, YAL0 represents the preset sample size, δ1 represents the weight of the correlation coefficient of the introduced variable, δ2 represents the sample size weight, and δ3 represents the information loss weight.
[0082] In this embodiment, if no potential factor is introduced after the adjustment of the significance level threshold, feedback is performed. Here, the feedback means notifying the preset personnel that no potential factor is introduced and it is impossible to optimize the information loss caused by the principal component dimensionality reduction.
[0083] The preset information loss threshold is represented by the average value of the qualified information loss evaluation values within the historical time period; the preset sample size is represented by the amount of sample data collected; the initial significance level threshold is set by the preset personnel. For example, the common significance level threshold is 0.05; the maximum value of the common significance level threshold is 0.1. Therefore, the maximum value multiplied by the initial significance level threshold is set to 2, so that the significance level threshold does not exceed 0.1 at most.
[0084] The introduced variable correlation coefficient represents the average value of the Pearson correlation coefficients between the introduced potential factors and each high-risk factor after dimensionality reduction; the sample size is obtained through the "Descriptive Statistics" menu of SPSS software.
[0085] This algorithm is obtained by processing multiple independent variables (introduced variable correlation coefficient, sample size after introducing variables, information loss evaluation value). There is an interaction relationship among these independent variables; as the sample size increases, more information can be provided for the multi-factor regression model, which may lead to a decrease in the introduced variable correlation coefficient; the higher the correlation coefficient between the new variable and the existing variables, the more likely the introduction of this variable will cause an increase in the information loss evaluation value; the larger the sample size after introducing variables, the smaller the information loss evaluation value, because a larger sample size can provide more comprehensive information, enabling the multi-factor regression model to better fit the data, thereby reducing information loss.
[0086] The introduced variable correlation coefficient weight, sample size weight, and information loss weight are obtained from the preset database, and the sum of the three is 1. For example, the introduced variable correlation coefficient forms a mapping set with the preset introduced variable correlation coefficient weight, and the real-time introduced variable correlation coefficient is input into the mapping set to obtain the corresponding introduced variable correlation coefficient weight; the sample size after introducing variables forms a mapping set with the preset sample size weight, and the real-time sample size after introducing variables is input into the mapping set to obtain the corresponding sample size weight; the information loss evaluation value forms a mapping set with the preset information loss weight, and the real-time information loss evaluation value is input into the mapping set to obtain the corresponding information loss weight; the mapping relationship therein can be one-to-one or many-to-one.
[0087] The value range of the sample size coefficient is adjusted to between 0 and 1 through the exponential function for the calculation adjustment of the significance level threshold; by gradually introducing potential factors and judging their significance, it is possible to more precisely identify which factors have a significant impact on the target variable, thereby avoiding the introduction of too many irrelevant or redundant factors, which helps to improve the generalization ability of the multi-factor regression model; by comprehensively considering the information loss evaluation value, introduced variable correlation coefficient, and sample size after introducing variables, it is possible to more comprehensively evaluate the impact of potential factors, and thus achieve more accurate multi-factor analysis.
[0088] Further, the specific method for determining whether to obtain the risk weights of high-risk factors based on the multicollinearity evaluation value obtained after introducing potential factors is as follows: B1, evaluate the multicollinearity evaluation value based on the multicollinearity evaluation data obtained after introducing potential factors; B2, determine whether the multicollinearity evaluation value obtained after introducing potential factors is less than the preset multicollinearity threshold obtained from the preset database. If the multicollinearity evaluation value is not less than the preset multicollinearity threshold obtained from the preset database, perform principal component dimensionality reduction on the high-risk factors and gradually adjust the dimensionality reduction parameters of the principal component dimensionality reduction. Otherwise, execute B4; B3, if the information loss evaluation value obtained after adjusting the dimensionality reduction parameters is not less than the preset information loss threshold obtained from the preset database, gradually adjust the preset information loss threshold. Otherwise, execute B4; B4, determine the optimal number of high-risk factors according to the number of high-risk factors with non-zero coefficients under the optimal regularization parameter obtained by the cross-validation method, and set the high-risk factors with non-zero coefficients under the optimal regularization parameter as the final high-risk factors, and obtain the risk weights accordingly.
[0089] In this embodiment, the dimensionality reduction parameters of the principal component dimensionality reduction are adjusted by gradually reducing the number of principal components until it is reduced to the preset number of principal components; the preset number of principal components is represented by the minimum value of the qualified multicollinearity evaluation values within the historical time period.
[0090] The gradual adjustment of the preset information loss threshold is represented by gradually increasing the preset ratio of the preset information loss threshold until it is increased to the preset information loss adjustment threshold; the preset information loss adjustment threshold is represented by the maximum value of the preset information loss adjustment thresholds within the historical time period; the preset ratio is set by the preset personnel and can be set to 10%, for example.
[0091] By introducing potential factors and adjusting the dimensionality reduction parameters, information loss can be effectively avoided, ensuring the stability and reliability of the multi-factor regression model; by adjusting the dimensionality reduction parameters and the information loss threshold, useful feature information can be retained to the greatest extent, avoiding overfitting of the multi-factor regression model caused by the introduction of potential factors, thereby improving the accuracy of the multi-factor regression model; through cross-validation and regularization techniques, high-risk factors related to intraoperative pressure injury can be accurately selected, providing a basis for factor screening.
[0092] Such as Figure 2As shown in the figure, it is a schematic structural diagram of an intelligent decision-making system for intraoperative pressure injury based on machine learning provided by an embodiment of the present application. The intelligent decision-making system for intraoperative pressure injury based on machine learning provided by the embodiment of the present application includes: a sample selection and analysis module, a multicollinearity evaluation module, an information loss evaluation module, and an intelligent decision-making module; among them, the sample selection and analysis module is used to perform univariate analysis on the obtained intraoperative pressure injury decision data to screen out potential factors, and determine whether to adjust the sample selection by analyzing the sample selection and analysis data of the potential factors; the multicollinearity evaluation module is used to, if the sample selection is adjusted, screen the potential factors after the sample selection adjustment to obtain high-risk factors, and determine whether to perform principal component dimensionality reduction by evaluating the multicollinearity evaluation data of the obtained high-risk factors; the information loss evaluation module is used to, if the principal component dimensionality reduction is not performed, obtain the risk weights of the high-risk factors, otherwise analyze the reduced high-risk factors obtained after the principal component dimensionality reduction to obtain information loss evaluation data, and determine whether to gradually introduce potential factors by evaluating the information loss evaluation data, and determine whether to obtain the risk weights of the high-risk factors based on the multicollinearity evaluation value obtained after the introduction of the potential factors; the intelligent decision-making module is used to perform intelligent decision-making on intraoperative pressure injury according to the risk weights and high-risk factors of the high-risk factors.
[0093] In this embodiment, the above-mentioned multiple factor regression model refers to the multiple factor logistic regression analysis algorithm; based on the high-risk factors, the random forest method is selected to construct a risk prediction model. The random forest algorithm has the ability to process high-dimensional data and non-linear relationships, and is suitable for the complex intraoperative pressure injury decision data in this embodiment.
[0094] Steps for constructing the risk prediction model: First step, divide the data set into a training set and a test set. Usually, the data with a preset training ratio is used for training, and the remaining data is used to test and verify the effect of the model. The preset training ratio is set by a preset person, for example, it can be set to 70%; Second step, use the training set data to train the random forest model. Each decision tree will be trained on a random subset of the data; Third step, determine how many trees to generate (usually 100 or more), and each tree has an independent learning process.
[0095] Steps for training the risk prediction model: According to the screened high-risk factors, determine the input variables as features; train the random forest model by using the training set data. Each decision tree is trained according to a randomly selected subset of features and a subset of data, and finally the prediction result is output through a voting mechanism; in the feature design stage, if the values of the features are adjusted based on the risk weights (for example, by combining the values of the high-risk factors with the corresponding weights), then the features with higher risk weights may be selected more during the decision tree splitting process, thus affecting the final prediction result to a greater extent.
[0096] The risk prediction model is the core technical support of the intelligent decision-making information platform. Through the real-time evaluation and early warning functions of the risk prediction model, it can quickly identify high-risk children and provide personalized prevention suggestions for medical staff. The accuracy and stability of the risk prediction model directly determine the practicality and reliability of the platform and are the key to realizing the early warning of intraoperative pressure injury.
[0097] Based on the risk prediction model, the intraoperative pressure injury risk of children is evaluated in real time; when the risk value exceeds the preset threshold, an early warning signal is automatically sent out, and prevention suggestions are provided. The warning information is displayed in the form of a pop-up window, the risk levels are distinguished by different colors, and the function of generating prevention suggestions with one key is provided; combined with the individual characteristics of the children and the warning information, personalized nursing decision support is provided for medical staff, including suggestions for position adjustment, decompression measures, skin care, etc.
[0098] Analyze the sample selection analysis data of potential factors to judge whether to adjust the sample selection. This process also reflects the consideration of data sensitivity, which is a common method in model tuning in machine learning; multicollinearity is a common problem in machine learning, which may lead to the instability of the model and the reduction of interpretability. By evaluating the multicollinearity of high-risk factors, it can be determined whether principal component dimensionality reduction is needed; principal component dimensionality reduction is a commonly used method in machine learning for processing high-dimensional data and improving the efficiency of the model; evaluating information loss to determine whether to gradually introduce potential factors is an important link in machine learning model tuning.
[0099] Through the sample selection analysis module, the effectiveness of sample selection is improved; through the multicollinearity evaluation module, the possibility of multicollinearity in the multiple factor logistic regression analysis is evaluated, and the high-risk factors selected are reduced in dimension by the principal component analysis method; through the information loss evaluation module, the information loss generated after the principal component dimensionality reduction is compensated; through the threshold adjustment module, the overfitting of the multiple factor regression model that may exist after the introduction of potential factors is adjusted, thereby realizing the accuracy of the intelligent decision-making analysis of intraoperative pressure injury.
[0100] In summary, in the embodiment of the present application, the sample selection is adjusted through sample selection analysis, then the principal component dimensionality reduction is performed based on the multicollinearity evaluation, and finally the potential factors are gradually introduced based on the information loss evaluation, and the risk weights of high-risk factors are obtained based on the multicollinearity evaluation value after the introduction of potential factors, and intelligent decision-making for intraoperative pressure injury is carried out in combination with high-risk factors, thereby reducing the multicollinearity between high-risk factors, and further realizing the improvement of the screening accuracy of high-risk factors in the intelligent decision-making of intraoperative pressure injury, effectively solving the problem of low screening accuracy of high-risk factors in the intelligent decision-making of intraoperative pressure injury in the prior art.
[0101] Those skilled in the art will appreciate that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.
[0102] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and combinations of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing device produce means for implementing the functions specified in Figure 1 one or more of the flows Figure 1 or blocks or combinations of blocks.
[0103] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a particular manner, such that the instructions stored in the computer-readable memory produce a manufacture including instruction means that implement the functions specified in Figure 1 one or more of the flows Figure 1 or blocks or combinations of blocks.
[0104] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are performed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in Figure 1 one or more of the flows Figure 1 or blocks or combinations of blocks.
[0105] Although the preferred embodiments of the present invention have been described, those skilled in the art can make additional changes and modifications once they learn the basic creative concepts. Therefore, the appended claims are intended to be construed to include the preferred embodiments and all changes and modifications that fall within the scope of the present invention.
[0106] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalent technologies, the present invention is also intended to include these modifications and variations.
Claims
1. An intelligent decision-making method for intraoperative pressure injury based on machine learning, characterized in that, It includes the following steps: S1. After performing univariate analysis on the obtained intraoperative pressure injury decision data, potential factors are screened out. By analyzing the sample selection analysis data of the potential factors, it is determined whether to adjust the sample selection; S2. If the sample selection is adjusted, the potential factors after the sample selection adjustment are screened to obtain high-risk factors. By evaluating the multicollinearity evaluation data of the obtained high-risk factors, it is determined whether to perform principal component dimensionality reduction; S3. If the principal component dimensionality reduction is not performed, the risk weights of the high-risk factors are obtained. Otherwise, based on the reduced high-risk factors obtained after the principal component dimensionality reduction, information loss evaluation data is analyzed. By evaluating the information loss evaluation data, it is determined whether to gradually introduce potential factors, and based on the multicollinearity evaluation value obtained after the introduction of the potential factors, it is determined whether to obtain the risk weights of the high-risk factors; S4. Perform intelligent decision-making on intraoperative pressure injury according to the risk weights of the high-risk factors and the high-risk factors.
2. The intelligent decision-making method for intraoperative pressure injury based on machine learning according to claim 1, wherein: The sample selection analysis data includes data category balance value, minimum observed frequency of factor combination, maximum observed frequency under factor, and table conversion time; The data category balance value represents the quantitative data of the influence degree of the maximum data volume and the minimum data volume under the factor category on the data category balance situation; The multicollinearity evaluation data includes variance inflation factor, maximum correlation coefficient, maximum eigenvalue, and minimum eigenvalue; The information loss evaluation data includes image similarity before and after dimensionality reduction, main factor retention coefficient, information entropy before dimensionality reduction, information entropy after dimensionality reduction, and dimensionality reduction time.
3. The intelligent decision-making method for intraoperative pressure injury based on machine learning according to claim 2, characterized in that: The specific method for analyzing the sample selection analysis data of the potential factors is as follows: The category balance coefficient is obtained by performing numerical range processing on the data category balance value. The category balance coefficient represents the quantitative data of the influence degree of the data category balance value on the effectiveness of sample selection; According to the relative deviation relationship between the minimum observed frequency of factor combination, the maximum observed frequency under factor, the maximum observed frequency under the preset factor, and the minimum observed frequency of preset factor combination obtained from the preset database, the factor observation coefficient is obtained. The factor observation coefficient represents the quantitative data of the influence degree of the minimum observed frequency of factor combination and the maximum observed frequency under factor on the effectiveness of sample selection; If the table conversion time is within the preset table conversion time, the conversion time coefficient is obtained by performing numerical range limitation on the table conversion time. Otherwise, the conversion time coefficient is recorded as 1. The conversion time coefficient represents the quantitative data of the influence degree of the table conversion time on the effectiveness of sample selection; The sample selection weight obtained from the preset database is introduced to correct the category balance coefficient, factor observation coefficient, and conversion time coefficient to obtain the sample selection evaluation value.
4. The intelligent decision-making method for intraoperative pressure injury based on machine learning according to claim 3, characterized in that: The specific method for determining whether to adjust the sample selection is as follows: A1. Determine whether the sample selection evaluation value is less than the second preset sample selection threshold obtained from the preset database. If the sample selection evaluation value is not less than the second preset sample selection threshold obtained from the preset database, source file adjustment is performed. Otherwise, execute A2; A2. Determine whether the sample selection evaluation value is less than the first preset sample selection threshold obtained from the preset database. If the sample selection evaluation value is not less than the first preset sample selection threshold obtained from the preset database, perform sample format adjustment. Otherwise, after data cleaning, screen the potential factors after data cleaning to obtain high-risk factors; The source file adjustment includes converting the source file format, adjusting the encoding format, and re-obtaining data; The sample format adjustment includes merging missing categories and mean imputation; The data cleaning includes outlier deletion and mean imputation.
5. The intelligent decision-making method for intraoperative pressure injury based on machine learning according to claim 2, wherein: The specific method for evaluating the multicollinearity evaluation data of the obtained high-risk factors is as follows: Quantify the variance inflation coefficient according to the variance inflation factor and the preset variance inflation factor obtained from the preset database. The variance inflation approximation quantification process is used to quantify the deviation between the variance inflation factor and the preset variance inflation factor; Quantify the eigenvalue coefficient according to the minimum eigenvalue and the maximum eigenvalue. The eigenvalue approximation quantification process is used to quantify the deviation between the minimum eigenvalue and the maximum eigenvalue; Introduce the multicollinearity weight obtained from the preset database to correct the maximum correlation coefficient, variance inflation coefficient, and eigenvalue coefficient to obtain the multicollinearity evaluation value.
6. The intelligent decision-making method for intraoperative pressure injury based on machine learning according to claim 5, wherein: The specific method for determining whether to perform principal component dimensionality reduction is as follows: Determine whether the multicollinearity evaluation value is less than the preset multicollinearity threshold obtained from the preset database. If the multicollinearity evaluation value is not less than the preset multicollinearity threshold obtained from the preset database, perform principal component dimensionality reduction on the high-risk factors. Otherwise, obtain the risk weight of the high-risk factors; The specific method for obtaining the risk weight of the high-risk factors is: Analyze the odds ratio according to the regression coefficient corresponding to the high-risk factor in the multiple factor regression model, and set the proportion of the odds ratio of each high-risk factor in the sum of all odds ratios as the risk weight of the corresponding high-risk factor.
7. The intelligent decision-making method for intraoperative pressure injury based on machine learning according to claim 2, characterized in that: The specific method for evaluating the information loss evaluation data is as follows: Obtain the information entropy coefficient before and after dimensionality reduction through normalization of the information entropy before dimensionality reduction and the information entropy after dimensionality reduction. The information entropy coefficient before and after dimensionality reduction represents the quantification data of the influence degree of the information entropy before dimensionality reduction and the information entropy after dimensionality reduction on the information loss degree; Quantify the dimensionality reduction time coefficient according to the dimensionality reduction time and the preset dimensionality reduction time obtained from the preset database. The dimensionality reduction time approximation quantification process is used to quantify the deviation between the dimensionality reduction time and the preset dimensionality reduction time; Introduce the information loss weight obtained from the preset database to correct the image similarity before and after dimensionality reduction, the main factor retention coefficient, the information entropy coefficient before and after dimensionality reduction, and the dimensionality reduction time coefficient to obtain the information loss evaluation value.
8. The intelligent decision-making method for intraoperative pressure injury based on machine learning according to claim 7, characterized in that: The specific process for determining whether to gradually introduce potential factors is as follows: C1. Determine whether the information loss evaluation value is less than the preset information loss threshold obtained from the preset database. If the information loss evaluation value is less than the preset information loss threshold obtained from the preset database, obtain the risk weight of the high-risk factors. Otherwise, execute C2; C2. Determine whether the significance level of the potential factor is less than the initial significance level threshold. If the significance level of the potential factor is less than the initial significance level threshold, introduce the corresponding potential factor into the multi-factor regression model; otherwise, do not introduce the potential factor and execute C3. C3. Determine whether any potential factor has been introduced. If no potential factor has been introduced, adjust the initial significance level threshold to obtain the significance level threshold; otherwise, judge whether to obtain the risk weight of the high-risk factor based on the multicollinearity evaluation value obtained after the introduction of the potential factor. The significance level threshold is obtained by correcting the correlation coefficient of the introduced variable, the sample size coefficient, the information loss deviation coefficient, and the initial significance level threshold with the significant weight obtained from the preset database. The sample size coefficient represents the quantitative data of the influence degree of the sample size after the introduction of the variable and the preset sample size on the significance level threshold. The information loss deviation coefficient represents the quantitative data of the influence degree of the information loss evaluation value and the preset information loss threshold on the significance level threshold.
9. The intelligent decision-making method for intraoperative pressure injury based on machine learning according to claim 8, wherein: The specific method for judging whether to obtain the risk weight of the high-risk factor based on the multicollinearity evaluation value obtained after the introduction of the potential factor is as follows: B1. Evaluate based on the multicollinearity evaluation data obtained after the introduction of the potential factor to obtain the multicollinearity evaluation value. B2. Judge whether the multicollinearity evaluation value obtained after the introduction of the potential factor is less than the preset multicollinearity threshold obtained from the preset database. If the multicollinearity evaluation value obtained after the introduction of the potential factor is not less than the preset multicollinearity threshold obtained from the preset database, perform principal component dimensionality reduction on the high-risk factor and gradually adjust the dimensionality reduction parameter of the principal component dimensionality reduction; otherwise, execute B4. B3. If the information loss evaluation value obtained after the adjustment of the dimensionality reduction parameter is not less than the preset information loss threshold obtained from the preset database, gradually adjust the preset information loss threshold; otherwise, execute B4. B4. Determine the optimal number of high-risk factors according to the number of high-risk factors with non-zero coefficients under the optimal regularization parameter obtained by the cross-validation method, and set the high-risk factors with non-zero coefficients under the optimal regularization parameter as the final high-risk factors, and thus obtain the risk weight.
10. An intelligent decision-making system for intraoperative pressure injury based on machine learning, characterized in that, Including: Sample selection analysis module, multicollinearity evaluation module, information loss evaluation module, and intelligent decision-making module; Among them, the sample selection analysis module is used to perform univariate analysis on the obtained intraoperative pressure injury decision data to screen out potential factors, and analyze the sample selection analysis data of the potential factors to judge whether to adjust the sample selection. The multicollinearity evaluation module is used to, if the sample selection is adjusted, screen the potential factors after the sample selection adjustment to obtain high-risk factors, and evaluate the multicollinearity evaluation data of the obtained high-risk factors to judge whether to perform principal component dimensionality reduction. The information loss assessment module is used to obtain the risk weights of high-risk factors if principal component dimensionality reduction is not performed. Otherwise, it analyzes the dimensionality-reduced high-risk factors obtained after principal component dimensionality reduction to obtain information loss assessment data, evaluates the information loss assessment data to determine whether to gradually introduce potential factors, and determines whether to obtain the risk weights of high-risk factors based on the multicollinearity assessment value obtained after the introduction of potential factors. The intelligent decision-making module is used to make an intelligent decision on intraoperative pressure injury according to the risk weights of high-risk factors and high-risk factors.
Citation Information
Patent Citations
Method for establishing disease risk adjustment model
CN104992058B
Hilar cholangiocarcinoma patient prognosis prediction model
CN107305596A
Cited By
Damage assessment method and device for composite material structure and electronic equipment
CN121011278A