Artificial intelligence prediction model system and establishment method

By combining large language models and SHAP analysis, the format limitation problem of AI medical prediction model is solved, flexible processing and detailed interpretation of data in different formats is achieved, and the credibility and application scope of the model is improved.

CN120299724APending Publication Date: 2025-07-11QUANTA COMPUTER INC
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202410100752.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-01-10
Filing Date
2024-01-24
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

The existing AI medical prediction models have format limitations in terms of input and output, cannot process natural language data, and lack clear interpretation mechanisms, resulting in limited application scope and low credibility.

Method used

Combining large language models, we use multiple feature values to build a predictive model using machine learning algorithms, and use SHAP analysis to generate detailed interpreted content, accept natural language input and provide detailed output.

Benefits of technology

It realizes flexible processing of data in different formats, provides detailed explanations and transparent prediction results, and improves the credibility and application scope of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120299724A_ABST
    Figure CN120299724A_ABST
Patent Text Reader

Abstract

The invention provides a system and an establishment method of an artificial intelligence prediction model. The system comprises a user interface, a large language model module, a machine learning module, an SHAP analysis module and a judgment module. The user interface is used for receiving a plurality of input characteristic values. And the large language model module is in signal connection with the user interface and the judgment module, and is used for receiving the input characteristic value and providing the output content of the judgment module to the user interface for display. And the machine learning module is in signal connection with the large language model module and is used for receiving the characteristic value and analyzing by using a machine learning algorithm to obtain a prediction result. And the SHAP analysis module is in signal connection with the machine learning module and is used for carrying out SHAP analysis on the characteristic values. The judgment module is in signal connection with the SHAP analysis module and is used for generating explanation content of the prediction result of the individual according to the result of the SHAP analysis. The invention also provides an establishment method of the artificial intelligence prediction model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to artificial intelligence or machine learning, and particularly to a system for an artificial intelligence prediction model and a method for establishing the same. Background Art

[0002] Known problems faced by artificial intelligence (AI) medical prediction models are not only reflected in the input limitations but also in the output. In terms of input, these AI medical prediction models can only accept data in specific formats, such as JSON (JavaScript Object Notation, a lightweight data interchange format) or other specific format files, and cannot process any format such as natural language. This limits the convenience in applications. Users must convert the data into a format that meets the requirements of the AI medical prediction model for input, which not only increases the complexity of operations but also may cause errors or losses during data transmission.

[0003] In terms of output, the results of AI medical prediction models are also restricted. The output of AI medical prediction models is usually limited to specific formats, such as JSON or other specific format files, and only provides the probability value of the prediction result. However, such output lacks important information. First, the lack of judgment basis information means that users cannot know the specific basis for the AI medical prediction model to make a prediction, which makes it difficult for users to determine the reliability of the prediction result. Second, the AI medical prediction model cannot provide the reason for explaining the prediction result, that is, it cannot clearly present the decision-making process of the AI medical prediction model. This situation makes it difficult for users to understand the operation logic of the AI medical prediction model and also difficult to trust the prediction results of the AI medical prediction model.

[0004] Generally speaking, AI medical prediction models face format limitations when processing input and output data, which restricts the application scope and credibility of AI medical prediction models. Future research and development should focus on solving these problems so that AI medical prediction models can more flexibly process different formats of data and provide clear and credible prediction results, so as to truly exert the potential of AI medical prediction models in the medical field and provide better support and assistance for medical health. Summary of the Invention

[0005] Based on the above, the present disclosure provides a method for establishing an artificial intelligence prediction model, which combines with a large language model and includes the following steps. According to the analysis target, determine multiple features for which data is to be collected. Collect the feature values of these features for multiple samples. Divide the feature values of these samples into a training group and a test group. Use multiple machine learning algorithms to analyze the feature values of the training group to establish multiple prediction models for the analysis target. Use the feature values of the test group to test the prediction accuracy of the multiple prediction results of these prediction models. Select the one with the highest prediction accuracy as the target model from these prediction models. Use the target model to calculate multiple SHAP values of the feature values of the training group. Use the SHAP values of the feature values of the training group to create a swarm plot and multiple partial dependence plots. For the SHAP values of the feature values of an individual among these samples in the training group, create a force plot, and let the target model generate an explanation content for the prediction result of this individual based on the force plot. Use a large language model to generate input text and output text in natural language. Apply the feature values of these samples to the input text as prompts for the target model. Apply the prediction result of this individual obtained from the force plot to the output text as the output content of the target model for this individual.

[0006] According to an embodiment of the present disclosure, the analysis target includes the risk of suffering from heart disease or the treatment efficacy of sudden deafness.

[0007] According to an embodiment of the present disclosure, when the analysis target is the risk of suffering from heart disease, these features include at least one of the factors of age, diabetes, smoking, blood pressure, blood lipid, and obesity.

[0008] According to an embodiment of the present disclosure, when the analysis target is the treatment efficacy of sudden deafness, these features include at least one of the factors of age, vestibular system symptoms / signs, steroid dose, and number of steroid injections.

[0009] According to an embodiment of the present disclosure, these machine learning algorithms include Logistic regression (LR), Decision Tree, Random Forest, Adaptive Boosting (AdaBoost), or eXtreme Gradient Boosting (XGBoost).

[0010] According to an embodiment of the present disclosure, the sorting basis of the swarm plot is the average absolute value of the SHAP values of the feature values of each of these features. The larger the average absolute value of the SHAP value, the greater the impact on the prediction result.

[0011] According to an embodiment of the present disclosure, a method for determining a threshold value of an interesting feature using a partial dependence graph of the interesting feature with respect to these features includes performing curve fitting on these feature values distributed in the partial dependence graph to obtain a fitted curve of the interesting feature and finding the intersection point of the fitted curve and the horizontal line with a SHAP value of zero in the partial dependence graph as the threshold value of the interesting feature, which is used as the basis for interpreting the analysis target of the individual.

[0012] Based on the above, the present disclosure further provides an artificial intelligence prediction model system, which is combined with a large language model and includes a user interface, a large language model module, a machine learning module, a SHAP analysis module, and a judgment module. The user interface is configured to receive an input content composed of multiple feature values of these features of an individual input by a user according to multiple features of the analysis target. The large language model module is connected to the user interface in a signal manner to receive the input content and provide an input text to analyze the input content to obtain these feature values. The machine learning module is connected to the large language model module in a signal manner to receive these feature values and use machine learning algorithms to analyze these feature values to obtain a prediction result of the individual. The SHAP analysis module is connected to the machine learning module in a signal manner to perform SHAP analysis on these feature values. The judgment module is connected to the SHAP analysis module and the large language model module in a signal manner to generate an explanatory content of the prediction result of the individual according to the result of the SHAP analysis, and let the explanatory content be incorporated into an output text provided by the large language model module to generate an output content for display on the user interface.

[0013] According to an embodiment of the present disclosure, the artificial intelligence prediction model system of the present disclosure further includes a database, which is connected to the large language model module and the machine learning module in a signal manner to receive and store these feature values of the individual from the large language model module for the machine learning module to access these feature values of the individual.

[0014] According to an embodiment of the present disclosure, the artificial intelligence prediction model system of the present disclosure further includes a verification module, which is connected to the machine learning module and the SHAP analysis module in a signal manner. When the machine learning module uses multiple machine learning algorithms to analyze these feature values, the verification module is configured to verify the accuracy rate of these machine learning algorithms, and use the prediction model established by the machine learning algorithm with the highest accuracy rate as the target model to provide the prediction result of the target model to the SHAP analysis module for SHAP analysis to obtain the explanatory content of the prediction result of the individual.

[0015] According to an embodiment of the present disclosure, these machine learning algorithms include logistic regression, decision tree, random forest, adaptive boosting, or extreme gradient boosting.

[0016] According to an embodiment of the present disclosure, the SHAP analysis module generates swarm plots of multiple samples, multiple partial dependence plots of these eigenvalue, and multiple force plots of these features of the individual.

[0017] According to an embodiment of the present disclosure, the swarm plots are sorted based on the average absolute value of the SHAP values of the eigenvalue of each of these features. The larger the average absolute value of the SHAP value, the greater the impact on the prediction result.

[0018] According to an embodiment of the present disclosure, the method for determining a critical value of an interesting feature using the partial dependence plot of the interesting feature of these features includes: performing curve fitting on the eigenvalue distributed in the partial dependence plot to obtain a fitting curve of the interesting feature and finding the intersection point of the fitting curve and the horizontal line with a SHAP value of zero in the partial dependence plot as the critical value of the interesting feature, as the basis for interpreting the analysis target of the individual.

[0019] According to an embodiment of the present disclosure, the analysis target includes the risk of suffering from heart disease or the treatment effect of sudden deafness.

[0020] According to an embodiment of the present disclosure, when the analysis target is the risk of suffering from heart disease, these features include at least one of the factors of age, diabetes, smoking, blood pressure, blood lipid, and obesity.

[0021] According to an embodiment of the present disclosure, when the analysis target is the treatment effect of sudden deafness, these features include at least one of the factors of age, vestibular system symptoms / signs, steroid dose, and number of steroid injections.

[0022] According to the above artificial intelligence prediction model system and its establishment method, it can be seen that it has flexibility in input data format, detailed explanation of prediction results, and improvement of user trust in prediction results. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] Figure 1 It is a schematic flowchart of a screening method for a prediction model in an artificial intelligence prediction model establishment method according to an embodiment of the present disclosure.

[0024] Figure 2A It is a schematic flowchart of a method for finding important features in an artificial intelligence prediction model establishment method according to an embodiment of the present disclosure.

[0025] Figure 2B It is a schematic flowchart of a method for finding the critical value of each feature in an artificial intelligence prediction model establishment method according to an embodiment of the present disclosure.

[0026] Figure 2CA method flowchart for finding the explanation content of an individual prediction result in a method for establishing an artificial intelligence prediction model according to an embodiment of the present disclosure.

[0027] Figure 3 A method flowchart for using a large language model to adjust input and output content in a method for establishing an artificial intelligence prediction model according to an embodiment of the present disclosure.

[0028] Figure 4 An architecture schematic diagram of an artificial intelligence prediction model system according to an embodiment of the present disclosure.

[0029] Figure 5A An example of displaying a structured data table.

[0030] Figure 5B An example of a swarm plot showing SHAP values.

[0031] Figure 5C An example of a partial dependence plot showing SHAP values.

[0032] Figure 5D An example of a force plot showing SHAP values.

[0033] Figure 6A An example of a SHAP summary plot showing the sum of the feature importances of patients case_0, case_1, and case_2.

[0034] Figure 6B-6D Separate SHAP summary plots of patients case_0, case_1, and case_2 are shown.

[0035] Figure 7A-7C Separate SHAP partial dependence plots of the feature Interval_HLto1stITSI of patients case_0, case_1, and case_2 are shown.

[0036] Figure 8A-8C Separate SHAP partial dependence plots of the feature Number_of_injections of patients case_0, case_1, and case_2 are shown.

[0037] Figure 9 A force plot of the features of case_0 is shown.

[0038]

Symbol Explanation

[0039] 110-170, 210a-230a, 210b-240b, 210c-230c, 310-360: Step 400: Artificial Intelligence Prediction Model System

[0040] 410: User Interface

[0041] 420: Large Language Model Module

[0042] 430: Database 2

[0043] 440: Machine Learning Module

[0044] 450: Verification Module

[0045] 460: SHAP Analysis Module

[0046] 470: Judgment Module Detailed Implementation Manner

[0047] To solve the problem of format limitations faced by known artificial intelligence prediction models when processing input and output data, the present disclosure proposes a method for establishing an artificial intelligence prediction model, which combines with a large language model (LLM) in artificial intelligence. The method for establishing an artificial intelligence prediction model includes establishing a prediction model, obtaining a prediction result using the prediction model, using SHAP analysis to explain the prediction result, and using a large language model to break through the problem of format limitations faced when dealing with known input and output data. Each of these will be described in detail below.

[0048] Method for Establishing an Artificial Intelligence Prediction Model

[0049] First, Figure 1 is a schematic flowchart of a method for screening a prediction model in a method for establishing an artificial intelligence prediction model according to an embodiment of the present disclosure. In Figure 1 step 110, according to the analysis target, determine multiple features of the data to be collected. Then, collect the feature values of these features for multiple samples.

[0050] In an optionally implementable step 120, these feature values can be converted into multiple structured data, for example, organize these feature values in the form of a Figure 5A table for data storage. Figure 5AThe table lists several characteristics that may affect the analysis target of "risk of heart disease", such as age, sex, chest pain type, resting blood pressure, cholesterol, fasting blood sugar, resting ECG, maximum heart rate, exercise angina, ST segment depression value in exercise electrocardiogram (Oldpeak), and ST slope in electrocardiogram (ST_Slope), etc. The results of whether each individual has heart disease are also listed. The above ST segment depression value refers to the manifestation of myocardial ischemia during exercise relative to rest (ST Depression).

[0051] In step 130, these samples are divided into a training group and a test group. The sample ratio of the training group to the test group can be, for example, 70:30 to 90:10, such as 70:30, 75:25, 80:20, 85:15, or 90:10.

[0052] Next, multiple machine learning algorithms are used to analyze these feature values of the training group to establish multiple prediction models for the analysis target. The prediction accuracy of the multiple prediction results of these prediction models is tested using the feature values of the test group.

[0053] Therefore, in step 140, the feature values of the training group are input into a machine learning algorithm for machine learning to establish a prediction model for the analysis target. Then, in step 150, the prediction accuracy of the prediction model is verified using the feature values of the test group. Then, in step 160, it is checked whether there are still candidate machine learning algorithms that have not been trained. If so, step 140 is executed again until all candidate machine learning algorithms have been trained at least once. The above machine learning algorithms include Logistic regression (LR), Decision Tree, RandomForest, Adaptive Boosting (AdaBoost), or eXtreme Gradient Boosting (XGBoost).

[0054] Then, in step 170, the prediction model with the highest prediction accuracy is selected from these prediction models as the target model. After that, all analysis predictions use the target model.

[0055] Next, using the target model, multiple SHAP values of these feature values of the training group are calculated. Please refer to Figure 2A-2C。 Figure 2A Schematic flow chart of a method for finding important features of an artificial intelligence prediction model establishment method according to an embodiment of the present disclosure.

[0056] In Figure 2A step 210a, calculate the average absolute value of the SHAP values of the respective feature values of the test group.

[0057] In step 220a, use the average absolute value of the SHAP values of the respective features of the training group to create a beeswarm plot as Figure 5B shown, and it can be seen how the most important features affect Figure 1 the obtained target model. The x-axis of the beeswarm plot is the SHAP value, and the respective features are arranged in sequence on the y-axis. The beeswarm plot can also use colors to indicate the magnitudes of the feature values of the respective features. Generally, red is used to indicate larger feature values, and blue is used to indicate smaller feature values.

[0058] In step 230a, based on the beeswarm plot, find several important features that have a positive impact on the analysis target. In the beeswarm plot, the wider the distribution of the SHAP value of a feature and the larger the SHAP value, the greater the impact of this feature on the analysis target. When the SHAP value is greater than zero, it has a positive impact on the analysis target. Conversely, when the SHAP value is less than zero, it has a negative impact on the analysis target. For example, in Figure 5B it, the feature of max_heart_rate_achieved has an important positive impact on the analysis target of "risk of suffering from heart disease".

[0059] Figure 2B Schematic flow chart of a method for finding the critical values of respective features of an artificial intelligence prediction model establishment method according to an embodiment of the present disclosure. In Figure 2B step 210b, create a partial dependence plot (PDP) of the SHAP values of the respective features of the training group, as Figure 5C shown. The PDP is a scatter plot that shows the impact of a single feature of interest on the analysis target and ignores the impact of other features on the analysis target. The x-axis of the PDP is the feature value of the feature of interest, and the y-axis is the SHAP value. In Figure 5C it, the x-axis is the value of the maximum heart rate, and the y-axis is the SHAP value of the maximum heart rate.

[0060] In step 220b, perform curve fitting on multiple feature values in the PDP to obtain the fitting curves of these feature values. Figure 5CDisplay the fitting curve of the maximum heart rate and the SHAP value. This method uses polynomial regression to fit the scatter data. After fitting, it also includes finding the intersection point of the fitting curve and a specific horizontal line (e.g., SHAP value = 0).

[0061] In step 230b, find the eigenvalue at the intersection of the fitting curve and the horizontal line "SHAP value = 0", which is the critical value of the feature of most interest. For example, in Figure 5C , the critical value of the maximum heart rate found is 150.

[0062] In step 240b, find the eigenvalue interval of "SHAP value > 0" in the fitting curve, that is, the interval where the feature has a positive impact on the prediction result as the basis for interpreting the individual prediction result. For example, in Figure 5C , it is shown that when the maximum heart rate is greater than 150, it has a positive impact on the "risk of having heart disease", that is, it will increase the "risk of having heart disease".

[0063] Figure 2C It is a schematic flowchart of a method for finding the interpretation content of an individual prediction result of an artificial intelligence prediction model establishment method according to an embodiment of the present disclosure. In Figure 2C step 210c, arbitrarily select an individual to be predicted.

[0064] In step 220c, use the Figure 1 target model to analyze the eigenvalues of the individual and obtain the prediction result.

[0065] In step 230c, calculate the SHAP value of each eigenvalue of the individual and create a Force plot to show how the prediction result of the individual is contributed by each eigenvalue, as Figure 5D shown. In Figure 5D , it can be seen that on the horizontal axis, the SHAP values of each feature are summed (f(x)), and a right arrow is used to show the positive SHAP value, and a left arrow is used to show the negative SHAP value. Therefore, it can be seen from the Force plot which features have the greatest positive influence on the analysis target, and it can be clearly known how much each feature specifically contributes to the analysis target. Therefore, it can be used as an auxiliary basis for the interpretation content of the individual prediction result.

[0066] Next, a large language model in artificial intelligence will be used so that the artificial intelligence prediction model can accept input content in natural language and can generate output content in natural language. Figure 3 It is a schematic flowchart of a method for using a large language model to adjust input and output content of an artificial intelligence prediction model establishment method according to an embodiment of the present disclosure.

[0067] As described above, since the data of each sample is usually stored in the form of structured data, when the artificial intelligence prediction model system to be trained receives input content in natural language, the structured data needs to be converted into natural language before use.

[0068] Therefore, in Figure 3 Step 310, a large language model is first used to generate an input text in natural language, leaving blanks for each eigenvalue to facilitate the filling of eigenvalues of different individuals. For example, based on Figure 5A the content of the table, the following input text can be generated: "A [xx]-year-old [xx] gender, suffering from [xx], with a blood pressure at rest of [xx], a serum cholesterol level of [xx], a fasting blood glucose of [xx], [xx] exceeding 120 mg / dl. The electrocardiogram shows [xx], and the maximum heart rate reached is [xx]. When this person exercises [xx] angina attacks, the ST-segment deviation value caused by exercise is [xx], the ST-segment slope is in the [xx] state, and [xx] has heart disease." Among them, [xx] represents the blanks to be filled in the input text.

[0069] In Step 320, each eigenvalue in the structured data of one of these samples is inserted into the input text to form an input content. For example, based on Figure 5A the data in the second row of the table, fill in the blanks of the above input text to get the following content: "A [49-year-old] [female], suffering from [non-anginal pain (NAP)], with a blood pressure at rest of

[160] , a serum cholesterol level of

[180] , a fasting blood glucose of [normal], [not] exceeding 120 mg / dl. The electrocardiogram shows [normal], and the maximum heart rate reached is

[156] . When this person exercises [no] angina attacks, the ST-segment deviation value caused by exercise is [1], the ST-segment slope is in the [flat] state, and [has] heart disease." Among them, the content enclosed in square brackets is Figure 5A each eigenvalue in the second row of the table.

[0070] In Step 330, let Figure 1 the obtained target model receive the input content obtained in Step 320, perform analysis and prediction, and obtain a prediction result.

[0071] In Step 340, according to the results of each SHAP analysis, the explanatory content of the prediction result of this individual is obtained.

[0072] In Step 350, the large language model generates an output text in natural language based on the prediction result obtained in Step 330 and the explanatory content obtained in Step 340. Since the generation method of the output text is similar to that of the input text, it will not be elaborated here.

[0073] In step 360, the prediction result and the explanation content are inserted into the output text to form an output content. For example, please refer to the output content in Table 1. At this point, the artificial intelligence prediction model combined with the large language model is established, and the work of analyzing and predicting new individuals can begin.

[0074] Table 1: Example of a patient's heart disease risk assessment report

[0075]

[0076] *Risk value comes from the SHAP value of the feature

[0077] Artificial Intelligence Prediction Model System

[0078] Next, we will introduce the artificial intelligence (AI) prediction model system. Figure 4 FIG. 1 is a schematic diagram of the architecture of an artificial intelligence prediction model system according to an embodiment of the present disclosure. Figure 4 In the figure, the artificial intelligence prediction model system 400 includes a user interface 410, a large language model module 420, a database 430, a machine learning module 440, a verification module 450, a SHAP analysis module 460 and a judgment module 470.

[0079] The user interface 410 is used to receive an input content composed of a plurality of characteristic values ​​of the characteristics of an entity input by a user according to the plurality of characteristics of the analysis target.

[0080] The large language model module 420 is connected to the user interface 410 for receiving the input content. The large language model module 420 is also responsible for providing an input text to analyze the input content to obtain the feature values.

[0081] The machine learning module 440 is connected to the large language model module 420 for receiving the feature values ​​and using a machine learning algorithm to analyze the feature values ​​to obtain a prediction result of the individual.

[0082] The SHAP analysis module 460 is connected to the machine learning module 440 for performing SHAP analysis on the feature values. The SHAP analysis module 460 generates a bee colony diagram of multiple samples and multiple partial dependency diagrams of the feature values, as well as multiple force diagrams of the features of the individual. The bee colony diagram, partial dependency diagram, and force diagram have been explained above, so they will not be repeated here.

[0083] The judgment module 470 is signal-connected to the SHAP analysis module 460 and the large language model module 420, so as to generate an explanation of the prediction result of the individual according to the result of the SHAP analysis, and let the explanation be incorporated into an output text provided by the large language model module 420 to generate an output content for display on the user interface 410.

[0084] In addition, a database 430 can be selectively configured, which is signal-connected to the large language model module 420 and the machine learning module 440. The database 430 is used to receive and store these characteristic values of the individual from the large language model module 420 for the machine learning module 440 to access these characteristic values of the individual.

[0085] A verification module 450 can also be selectively configured, which is signal-connected to the machine learning module 440 and the SHAP analysis module 460. When the machine learning module 440 uses multiple machine learning algorithms to analyze these characteristic values, the verification module 450 is used to verify the accuracy rate of these machine learning algorithms, and use the prediction model established by the machine learning algorithm with the highest accuracy rate as the target model to provide the prediction result of the target model to the SHAP analysis module 460 for SHAP analysis to obtain an explanation of the prediction result of the individual. The above machine learning algorithms include logistic regression, decision tree, random forest, adaptive boosting or extreme gradient boosting.

[0086] Experimental example: Prediction of "Therapeutic effect of sudden deafness"

[0087] Next, take "Therapeutic effect of sudden deafness" as an example to illustrate the establishment process of the above artificial intelligence prediction model.

[0088] First, select the characteristics required to predict the "Therapeutic effect of sudden deafness", as shown in Table 2 below.

[0089] Table 2: 21 characteristics selected for the "Therapeutic effect of sudden deafness".

[0090]

[0091]

[0092] In this experimental example, the prediction accuracies of five machine learning algorithms, namely logistic regression, decision tree, random forest, adaptive boosting or extreme gradient boosting, were evaluated and listed in Table 3 below. It can be seen from Table 3 that the "random forest" has the highest prediction accuracy. Therefore, the "random forest" algorithm is used subsequently for individual prediction and explanation of the "Therapeutic effect of sudden deafness".

[0093] Table 3: Prediction accuracies of machine learning algorithms

[0094]

[0095]

[0096] Then, SHAP value analysis is performed on each eigenvalue of each sample.

[0097] Figure 6A Show the SHAP summary plot of the sum of the importance of each feature that can display patients with complete recovery case_0 (Complete Recovery, CR), partial recovery case_1 (Partial Recovery, PR), and no recovery case_2 (No Recovery, NR). Each feature is arranged from top to bottom according to importance. From Figure 6A it can be known that for the prediction of "the treatment effect of sudden deafness", the most important 5 features (that is, the 5 features at the Figure 6A top) are Interval_HLto1stITSI, Initial_dB_difference_2k, Initial_dB_difference_4k, SD_Mean, and Age.

[0098] Figure 6A The respective SHAP summary plots of patients case_0, case_1, and case_2 in Figure 6B-6D are shown in Figure 6B-6D In Figure 6B-6D , each feature is sorted according to the magnitude of its influence, and the influence degree of different features on the model prediction result can be understood at the same time, including whether it is a positive influence or a negative influence.

[0099] Figure 7A-7C Show the respective SHAP partial dependence plots of patients case_0, case_1, and case_2. Figure 7A-7C Show the relationship between a specific feature (Interval_HLto1stITSI) and the prediction result, whether this relationship is linear, non-linear, or more complex.

[0100] In Figure 7A of the complete recovery case_0, find Figure 7A the fitting curve (not shown) and the intersection point where the SHAP value is zero, and its eigenvalue is 10. Regarding the influence of the feature Interval_HLto1stITSI on predicting the complete recovery (CR, that is, class_0) of patients, combined with the interval where the SHAP value > 0, from Figure 7AIt can be seen from the SHAP partial dependence graph that when the value of the feature Interval_HLto1stITSI is less than 10, the model predicts a relatively high probability that the patient will achieve complete recovery. This means that if the first middle ear steroid injection treatment is carried out as soon as possible within a relatively short time after the onset of the disease, the possibility of complete recovery of the patient can be increased.

[0101] In the case of no recovery case_2 Figure 7C find Figure 7C The fitting curve (not shown) and the intersection point where the SHAP value is zero, and its feature value is 14. Regarding the influence of the feature Interval_HLto1stITSI on predicting that the patient has no hearing recovery (NR, i.e., class_2), combined with the interval where the SHAP value > 0, according to Figure 7C the analysis of the SHAP partial dependence graph, when the value of Interval_HLto1stITSI is greater than 14, the model tends to predict that the patient will belong to the NR (no recovery) situation. This indicates that if the patient receives the first middle ear steroid injection treatment more than 14 days after the onset of deafness, it may significantly reduce the probability of recovery and increase the risk of no hearing recovery.

[0102] Figure 8A-8C Show the respective SHAP partial dependence graphs of patients case_0, case_1 and case_2. Figure 8A-8C Show the relationship between a specific feature (Number_of_injections) and the prediction result, whether this relationship is linear, non-linear or more complex.

[0103] In the case of complete recovery case_0 Figure 8A find Figure 8A The fitting curve (not shown) and the intersection point where the SHAP value is zero, and its feature value is 3. Regarding the influence of the feature Number_of_injections on predicting the patient's complete recovery (CR, i.e., class_0), combined with the interval where the SHAP value > 0, from the SHAP partial dependence graph, it can be observed that when the number of injections Number_of_injections is less than 3 times, the model predicts a relatively high probability that the patient will achieve complete recovery. This may mean that a relatively small number of treatment injections, that is, 3 times or less, can achieve effective treatment results.

[0104] In the case of no recovery case_2 Figure 8C find Figure 8CThe intersection point of the fitting curve (not shown) and the SHAP value of zero, with a characteristic value of 3. When evaluating the impact of the feature Number_of_injections (number of injections) on predicting that the patient's hearing has no recovery (NR, i.e., class_2), combined with the interval where the SHAP value > 0, it can be seen that when the number of injections of the patient exceeds 3 times, the model tends to classify these patients as NR (no recovery). This indicates that as the number of injections increases, the possibility of the patient fully recovering their hearing decreases. This may be because patients with poor recovery keep coming back for further treatment, so the number of injections is also higher.

[0105] Figure 9 Show the force diagram of each feature of patient case_0. From Figure 9 It can be seen that the three features with the greatest contribution are Interval_HLto1stITSI, Number_of_injections, and ITSI_Protocol, and their contribution degrees (i.e., SHAP values) to the analysis target "treatment efficacy of sudden deafness" are +0.12, +0.11, and +0.09 in sequence.

[0106] From the above, it can be seen that the method system for establishing the artificial intelligence prediction model provided by this disclosure has at least the following advantages.

[0107] Improved flexibility of data format: This method can handle data entries in different formats, including natural language format, improving the flexibility of data entry.

[0108] Provide detailed explanations and basis: Through methods such as SHAP values, swarm plots, partial dependence plots, and force diagrams, detailed explanations of the prediction results are provided, increasing the credibility of the prediction results.

[0109] Expansion of application fields: Overcoming the format limitations of traditional AI medical prediction models, it can be widely applied in the medical field, including risk prediction and treatment effect analysis of different diseases. It can also be extended to other application fields with similar requirements.

[0110] Improve the trust of users: The detailed explanations and transparent decision-making basis improve the trust of users in the prediction results of the artificial intelligence prediction model.

Claims

1. A method for establishing an artificial intelligence prediction model, which is combined with a large language model. The artificial intelligence prediction model includes: Determine multiple features of the data to be collected according to the analysis target; Collect the feature values of the multiple samples for the said features; Divide the feature values of the said samples into a training group and a test group; Use multiple machine learning algorithms to analyze the feature values of the training group to establish multiple prediction models for the analysis target; Use the feature values of the test group to test the prediction accuracy of the multiple prediction results of the prediction model; Select the one with the highest prediction accuracy from the said prediction models as the target model; Use the target model to calculate multiple SHAP values of the feature values of the training group; Use the SHAP values of the feature values of the training group to create a swarm plot and multiple partial dependence plots; For the SHAP value of the feature value of an individual of the said samples in the training group, create a force plot, and let the target model generate the explanatory content of the prediction result of the individual according to the force plot; Use a large language model to generate input text and output text in natural language; Apply the feature values of the said samples to the input text as the prompt of the target model; And Apply the prediction result of the individual obtained from the force plot to the output text as the output content of the target model for the individual.

2. The method for establishing an artificial intelligence prediction model according to claim 1, wherein the analysis target includes the risk of suffering from heart disease or the treatment effect of sudden deafness.

3. The method for establishing an artificial intelligence prediction model according to claim 1, wherein the machine learning algorithms include logistic regression, decision tree, random forest, adaptive boosting or extreme gradient boosting.

4. The method for establishing an artificial intelligence prediction model according to claim 1, wherein the method for determining the critical value of the feature of interest using the partial dependence plot of the feature of interest includes: Perform curve fitting on the feature values distributed in the partial dependence plot to obtain the fitting curve of the feature of interest; and Find the intersection point of the fitting curve and the horizontal line with a SHAP value of zero in the partial dependence plot as the critical value of the feature of interest, as the basis for interpreting the analysis target of the individual.

5. An artificial intelligence prediction model system, which is combined with a large language model. The artificial intelligence prediction model system includes: A user interface, according to multiple features of the analysis target, for receiving input content composed of multiple feature values of the features of an individual input by the user; A large language model module, signal-connected to the user interface, for receiving the input content and providing input text to analyze the input content to obtain the feature values; A machine learning module, signal-connected to the large language model module, for receiving the feature values and using machine learning algorithms to analyze the feature values to obtain the prediction result of the individual; A SHAP analysis module, signal-connected to the machine learning module, for performing SHAP analysis on the feature values; And A judgment module, which is connected to the SHAP analysis module and the large language model module by signals, is used to generate an explanation of the prediction result of the individual based on the result of the SHAP analysis, and let the explanation be incorporated into the output text provided by the large language model module to generate output content for display on the user interface.

6. The artificial intelligence prediction model system according to claim 5 further includes a verification module, which is connected to the machine learning module and the SHAP analysis module by signals. When the machine learning module uses multiple machine learning algorithms to analyze the feature values, the verification module is used to verify the accuracy rate of the machine learning algorithms, and use the prediction model established by the machine learning algorithm with the highest accuracy rate as the target model to provide the prediction result of the target model to the SHAP analysis module for SHAP analysis to obtain the explanation of the prediction result of the individual.

7. The artificial intelligence prediction model system according to claim 5, wherein the machine learning algorithms include logistic regression, decision tree, random forest, adaptive boosting or extreme gradient boosting.

8. The artificial intelligence prediction model system according to claim 5, wherein the SHAP analysis module generates swarm plots of multiple samples and partial dependence plots of the feature values, and multiple force plots of the features of the individual.

9. The method for determining the critical value of the feature of interest using the partial dependence plot of the feature of interest according to claim 8 includes: Performing curve fitting on the feature values distributed in the partial dependence plot to obtain a fitting curve of the feature of interest; and Finding the intersection point of the fitting curve and the horizontal line with a SHAP value of zero in the partial dependence plot as the critical value of the feature of interest, as the basis for interpreting the analysis target of the individual.

10. The artificial intelligence prediction model system according to claim 5, wherein the analysis target includes the risk of suffering from heart disease or the treatment effect of sudden deafness.

Citation Information

Cited By

  • Raw tea sensory classification and flavor critical threshold extraction method and system

    CN122132928A