Lung function diagnosis and prediction system of improved random forest model based on weighted regularization information gain
Through the improved random forest model based on weighted regularized information gain, the problems of low efficiency and low accuracy of traditional pulmonary function diagnosis are solved, and automated, standardized and high-precision pulmonary function diagnosis is achieved, reducing the burden on doctors.
Patent Information
- Application Number
- CN202510792985.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-13
- Publication Date
- 2025-09-12
Smart Images

Figure CN120636774A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of diagnostic system construction, and specifically relates to a lung function diagnosis and prediction system based on an improved random forest model with weighted regularized information gain. Background Art
[0002] Traditional pulmonary function diagnosis relies primarily on manual interpretation of multiple indicators in pulmonary function test reports, such as vital capacity (VC), forced expiratory volume in one second (FEV1), and total lung capacity (TLC). Doctors make comprehensive judgments on the patient's ventilation and diffusion functions based on their experience and guidelines. While this approach relies on the doctor's professional expertise, it also faces the following challenges:
[0003] (1) Low efficiency and high cost: It requires a lot of manual reading and analysis, which increases the burden on doctors and reduces work efficiency;
[0004] (2) Highly subjective and prone to error: The diagnostic results are easily affected by differences in physician experience, and there is a risk of misdiagnosis or omission;
[0005] (3) Difficulty in processing large-scale data: Traditional methods are difficult to implement batch analysis and automated processing, which limits the application of big data in lung function diagnosis.
[0006] To address these issues, artificial intelligence and machine learning methods have been gradually introduced into the medical field in recent years. Ensemble learning methods, such as Random Forest, have particularly demonstrated their simplicity, robustness, and ease of interpretation, and have been used in a variety of medical prediction tasks. However, traditional Random Forests still have the following shortcomings when it comes to pulmonary function diagnosis:
[0007] (1) The model has poor recognition ability for minority class samples and tends to be biased towards the majority class;
[0008] (2) Failure to constrain feature complexity may lead to overfitting;
[0009] (3) The feature selection mechanism is not optimized enough, which affects the model generalization ability and diagnostic accuracy.
[0010] Therefore, there is an urgent need for an improved random forest model that combines sample balancing strategy and feature regularization mechanism to build a lung function diagnosis and prediction system to achieve automated, standardized and high-precision diagnosis. Summary of the Invention
[0011] The purpose of the present invention is to provide a lung function diagnosis and prediction system based on an improved random forest model with weighted regularized information gain. By introducing a regularization term in the information gain calculation to measure the influence of features on model complexity, overfitting can be effectively prevented, medical diagnosis efficiency can be significantly improved, and the system has good application prospects.
[0012] To achieve the above objectives, the technical solution of the present invention is: a lung function diagnosis and prediction system based on an improved random forest model with weighted regularized information gain, comprising a data extraction module, a data preprocessing module, a feature selection module, a random forest training module, a diagnosis module, and a result output module; wherein,
[0013] The data extraction module is used to extract data from pulmonary function test report forms from test report files in PDF, TXT, or image formats. The data extraction module includes a text recognition function that can identify and extract data from PDF and TXT files, as well as a rectangular contour recognition algorithm that can identify the areas where various pulmonary function indicators are located in image files and extract the indicator data from the case report form.
[0014] The data preprocessing module is used to clean and standardize the indicators extracted by the data extraction module. This module includes data cleaning functions to remove invalid or erroneous data and data standardization functions.
[0015] A feature selection module is used to determine some key indicators for pulmonary function diagnosis from the indicators. This module can analyze the correlation between the indicators in the data set and pulmonary function diagnosis and select key indicators for diagnosis.
[0016] Random forest training module, used to train random forest prediction models using data sets with correct diagnostic conclusions;
[0017] The diagnosis module is used to diagnose lung function based on the trained random forest prediction model. This module applies the trained model to the diagnosis of the lung test report currently input by the user and converts the prediction result into a diagnosis conclusion.
[0018] The result output module is used to display the diagnosis results; this module is used to generate a detailed diagnosis report to be displayed on the system interface, and output a completed pulmonary test report.
[0019] Compared to the prior art, the present invention has the following beneficial effects: The system of the present invention includes a data extraction module, a data preprocessing module, a feature selection module, a model training module, a diagnosis module, and a result output module. The present invention inputs pulmonary examination reports in three formats: text files, PDF files, and image files, reads the data values of each indicator in the report, trains and loads the pre-trained prediction model, selects the indicators required for predicting ventilation function and diffusion function, inputs them into the corresponding models for prediction, and finally saves the diagnosis conclusion in a new file as a complete diagnosis report. To address the overfitting phenomenon that occurs when the random forest algorithm is applied to pulmonary function diagnosis due to the uneven distribution of sample types in the data set, the present invention proposes to use the SMOTE algorithm to upsample minority class samples, thereby improving diagnostic accuracy and the generalization ability of the model. Furthermore, by introducing a regularization term in the information gain calculation to measure the degree of influence of features on model complexity, overfitting can be effectively prevented. The present invention applies the weighted regularized information gain random forest algorithm to pulmonary function diagnosis, significantly improving medical diagnosis efficiency, having good application prospects and effectively reducing the workload of doctors. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] Figure 1 This is a flow chart of the execution of each module of the system of the present invention. DETAILED DESCRIPTION
[0021] The technical solution of the present invention will be described in detail below with reference to the accompanying drawings.
[0022] like Figure 1 As shown, the present invention provides a lung function diagnosis and prediction system based on an improved random forest model with weighted regularized information gain, comprising a data extraction module, a data preprocessing module, a feature selection module, a random forest training module, a diagnosis module, and a result output module; wherein,
[0023] The data extraction module is used to extract data from pulmonary function test report forms from test report files in PDF, TXT, or image formats. The data extraction module includes a text recognition function that can identify and extract data from PDF and TXT files, as well as a rectangular contour recognition algorithm that can identify the areas where various pulmonary function indicators are located in image files and extract the indicator data from the case report form.
[0024] The data preprocessing module is used to clean and standardize the indicators extracted by the data extraction module. This module includes data cleaning functions to remove invalid or erroneous data and data standardization functions.
[0025] A feature selection module is used to determine some key indicators for pulmonary function diagnosis from the indicators. This module can analyze the correlation between the indicators in the data set and pulmonary function diagnosis and select key indicators for diagnosis.
[0026] Random forest training module, used to train random forest prediction models using data sets with correct diagnostic conclusions;
[0027] The diagnosis module loads the trained random forest prediction model through the random forest model loading module and is used to diagnose lung function based on the trained random forest prediction model. This module applies the trained model to the diagnosis of the pulmonary test report currently input by the user and converts the prediction result into a diagnosis conclusion.
[0028] The result output module is used to display the diagnosis results; this module is used to generate a detailed diagnosis report to be displayed on the system interface, and output a completed pulmonary test report.
[0029] In one embodiment of the present invention, the data extraction module has the following specific functions:
[0030] (1) If the report to be processed is a text file, read the data information through the following steps:
[0031] Step a1: Read the file name suffix of the report, which is label F s , F s If it is 1, it means that the file has a detection conclusion. s If it is 0, it means that there is no detection conclusion for the file;
[0032] Step a2: Read the full text of the report into the string str f ;
[0033] Step a3: From str f Extract the character string str1 representing personal information and the character string str2 representing the indicator detection value;
[0034] Step a4: Read further indicators from str2, including unit, actual value, predicted value, and actual / predicted value information, and record them in array A str ;
[0035] (2) If the report to be processed is an image file, the rectangular area contour recognition algorithm is used to extract the indicator data from the image. The steps are as follows:
[0036] Step b1: Read the file name suffix of the report, which is label F P , F P If it is 1, it means that the file has a detection conclusion. P If it is 0, it means that there is no detection conclusion for the file;
[0037] Step b1: converting the report to be processed into a grayscale image G;
[0038] Step b2: Input algorithm parameters Δx and Δy, which represent the length and width of the area where the indicator value is located respectively;
[0039] Step b3: Perform edge detection on G to obtain image G containing all line segments c ;
[0040] Step b4: Identify G c Calculate the angle between each line segment and the horizontal line, delete the line segments with an angle greater than 5°, and keep the remaining line segments in the set n l is the number of line segments;
[0041] Step b5: Pair the line segments in L in pairs. a ,l b ) Get the starting and ending coordinates; where the straight line l a ={(x a1 ,y a1 ),(x a2 ,y a2 )},(x a1 ,y a1 ) and (x a2 ,y a2 ) is l a The starting and ending coordinates of the straight line l b ={(x b1 ,y b1 ),(x b2 ,y b2 )}, where (x b1 ,y b1 ) and (x b2 ,y b2 ) is l b The starting and ending coordinates of l a 、l b , calculate the maximum value of the horizontal coordinate x max =max{x a1 ,x a2 ,x b1 ,x a2} and the minimum value of the horizontal axis x min =min{x a1 ,x a2 ,x b1 ,x a2}, calculate the maximum value of the vertical coordinate y max =max{y a1 ,y a2 ,y b1 ,y a2} and the minimum value of the vertical coordinate y min =min{y a1 ,ya2 ,y b1 ,y a2}; Select (x min ,y min ),(x min ,y max ),(x max ,y min ),(x max ,y max ) Four coordinates in image G c The area in which the image is located forms an image segment s l , all line segment pairs (l a ,l b ) constitutes an image segment collection n s is the number of line segment pairs;
[0042] Step b9: For any image segment s in S i , calculate s i The horizontal axis difference and the vertical coordinate difference The first one to satisfy and s i The image segment s identified as containing indicator data o , δ is the allowable error of the horizontal axis, ε is the allowable error of the vertical axis;
[0043] Step b10: Input algorithm parameter d, where d represents the pixel width of an indicator value on the report image, and s o From left to right are the indicator name, indicator unit, indicator actual value, indicator predicted value and the ratio of the indicator actual value to the indicator predicted value (hereinafter referred to as the indicator actual / predicted value). Scan s from right to left according to the width d. o , extract the three sub-image segments s representing the ratio of the actual value of the indicator to the predicted value of the indicator, the predicted value of the indicator and the actual value of the indicator ap ,s p ,s a ;
[0044] Step b11: Use OCR tools to classify s ap ,s p ,s a Convert to 3 strings and record them in string array A pic middle.
[0045] In one embodiment of the present invention, the data preprocessing module uses the following steps to clean and standardize the extracted data, including removing invalid or erroneous data and converting the data into a unified format:
[0046] Step c1: Apic The incorrectly identified indicator values are replaced with the pre-given indicator default values;
[0047] Step c2: When using OCR tools to recognize data, the decimal point on the test report only occupies a small number of pixels, so existing OCR tools are prone to problems with recognition. pic For the indicator value that exceeds γ, the value is divided by 10 until it is less than γ, where γ is the preset upper limit of the indicator value;
[0048] Step c3: According to F P and F s Determine whether the corresponding test report file has a diagnosis conclusion. If A str or A pic The corresponding test report file has a diagnosis conclusion. str or A pic Generates a collection of indicator values O p ={C1,C2,...,C nc}, C1, C2 corresponds to A str or A pic The index value in n c Indicates the number of indicator values; the diagnostic conclusion is converted into L b ; O p and L b Combination is (O p ,L b ) output to the set S t In the random selection S t 70% of the dataset is used as the training set, and the remaining 30% is used as the validation set; if A str or A pic The corresponding test report file does not have a diagnosis conclusion. str or A pic Input the report set S to be predicted p .
[0049] In one embodiment of the present invention, the feature selection module determines the most influential key diagnostic indicators based on the correlation between the indicators and the pulmonary function diagnosis. The feature selection module inputs the key indicator set S k ={VC (vital capacity), MVV (maximal ventilation), FEV1 (forced expiratory volume in the first second), FEV1 / FVC (one-second rate), FV curve value, TLC (total lung capacity), RV (residual volume), RV / TLC (residual to total lung capacity ratio)}, output key diagnostic indicators.
[0050] In one embodiment of the present invention, the random forest training module is trained on a set F of test report files with correct diagnostic labels, wherein the key diagnostic indicators obtained by the feature selection module are used as input feature data, and the correct diagnostic results are used as labels to train the random forest prediction model for diagnosis;
[0051] Most existing random forest algorithms use traditional metrics such as the Gini index or information gain to measure the quality of node splitting. However, these metrics still have the following problems:
[0052] (1) Traditional indicators are very sensitive to noisy data, which will cause the decision trees in the random forest to split on the noise features, resulting in poor generalization ability of the model.
[0053] (2) Traditional indicators usually do not consider the interactions between features, which results in the model being unable to capture the complex relationships between features and may affect the prediction accuracy.
[0054] (3) Traditional indicators tend to select features with more values for splitting, but such features are not always the best splitting features, and a large number of child nodes may be split based on these features, making the model complicated.
[0055] To address the above problems, the present invention designs a new weighted regularized information gain in the random forest training module. First, a weight ω(f) is assigned to feature f, and a regularization penalty term is defined. The penalty term considers two factors: the number of values of feature f and the frequency of use of feature f in the decision tree. In this way, if the possible value of feature f exceeds a predetermined value, or the number of times feature f is used in the decision tree exceeds a predetermined value, the value of the regularization term C(f) will increase, and the model will be penalized. At the same time, SMOTE upsampling is performed on the data set X, which is the data set obtained after processing the test report file set F with correct diagnostic labels by the data extraction module and the data preprocessing module. The weighted regularized information gain is used to measure the quality of node splitting in the random forest. The weighted regularized information gain calculation formula is:
[0056]
[0057] in, is the weighted regularized information gain of feature f on dataset X, ω(f) is the weight of feature f, λ is a parameter, G(X,f) is the information gain of feature f on dataset X, which is used to measure the quality of splitting using feature f. The calculation formula is:
[0058]
[0059] Among them, E(S) is the entropy of the parent node, E(S i ) is the entropy of the i-th child node, |Si | and |S| represent the number of samples of parent nodes and child nodes respectively, v represents the total number of child nodes; E(S i ) and C(f) are calculated as follows:
[0060]
[0061] C(f)=α·N(f)+β·F(f)
[0062] Where K is the total number of all possible categories of samples in the dataset, p(k) is the proportion of samples of the kth category in the dataset, α and β are parameters, N(f) represents the number of all possible values of feature f in the dataset, and F(f) is the frequency of feature f in the decision tree, which is used to measure the influence of feature f on the complexity of the model. The calculation formula is:
[0063]
[0064] Among them, C used (f, t) is the number of times feature f is used in the tth tree, and T is the total number of decision trees. Clearly, the greater the number of possible values for feature f, the larger the value of N(f), and the more complex the decision tree may become when feature f is used. Furthermore, the more times feature f is used in a decision tree, the larger the value of F(f), and the more complex the tth decision tree will be. Using the new weighted regularized information gain to evaluate the quality of node splitting can effectively prevent overfitting caused by overly complex models.
[0065] In one embodiment of the present invention, the random forest training module uses the improved random forest algorithm based on weighted information gain to train the random forest prediction model in the following specific steps:
[0066] Step d1: using the data extraction module to read the data set D from the report file set F;
[0067] Step d2: Use the data preprocessing module to preprocess the data set D to obtain the data set D';
[0068] Step d3: Perform SMOTE upsampling on the minority class whose sample number in the dataset D' is less than the predetermined value to obtain the preprocessed dataset D";
[0069] Step d31: Determine the upsampling ratio to be 1:1, so that the number of minority class samples is increased to be equal to the number of majority class samples;
[0070] Step d32: Generate synthetic minority class samples according to the following interpolation formula and add them to the dataset;
[0071] SynS i =M i+γ·(N i -M i )
[0072] Among them, SynS i is a synthetic sample, M i is the selected minority class sample, N i is the nearest neighbor of the selected sample, γ is a random number from 0 to 1;
[0073] Step d4: Use Boostrap sampling, that is, randomly select samples from the dataset D' with replacement until a subsample set N of the same size as the original dataset is obtained;
[0074] Step d5: Use N to select features based on recursive feature elimination and train a decision tree;
[0075] Step d6: Use step d5 to train k decision trees in parallel to obtain random forest prediction models M1 and M2 for predicting ventilation function and diffusion function, respectively.
[0076] In one embodiment of the present invention, step d5 specifically includes the following steps:
[0077] Step d51: Set the number of features to be reduced by 10% in each iteration, set the target number of features m, and use all the features in the dataset to train the decision tree;
[0078] Step d52: Calculate the weighted regularized information gain for node splitting using each feature f
[0079] Step d53: Remove weighted regularization information gain The lowest 10% of features;
[0080] Step d54: Repeat steps d52 to d53 until the number of features is reduced to a preset number m.
[0081] In one embodiment of the present invention, the diagnosis module uses a newly designed diagnostic rule library to convert the test results into test index evaluations, and simultaneously calls the random forest prediction module to generate a diagnosis conclusion based on the test index values in the report to be predicted. The specific implementation steps are as follows:
[0082] Step e1: Design a diagnostic rule base consisting of the following rules to convert key diagnostic indicator values into detection indicator evaluations; the diagnostic rule base can add, modify, and delete key diagnostic indicators in real time;
[0083] Rule 1: For vital capacity (VC), the ratio of the measured value to the predicted value is:
[0084]
[0085] Rule 2: For MVV actual / predicted, the ratio of the measured value to the predicted value of MVV is:
[0086]
[0087] Rule 3: For the FEV1 in the first second, the actual / predicted value is the ratio of the measured value to the predicted value in the first second:
[0088]
[0089] Rule 4: For the measured value of FEV1 / FVC in one second, that is, the ratio of the measured value to the predicted value of the one-second rate:
[0090]
[0091] Rule 5: For TLC actual / predicted, the ratio of the measured value to the predicted value of the total lung capacity is:
[0092]
[0093] Rule 6: For residual gas value RV actual / predicted, that is, the ratio of the measured value to the predicted value of the residual gas value:
[0094]
[0095] Rule 7: For the residual-to-total ratio RV / TLC actual / predicted, that is, the ratio of the measured value to the predicted value of the residual-to-total ratio:
[0096]
[0097] Step e2: Use the rule base mapping to obtain the description of each indicator and store it in the string S c middle;
[0098] Step e3: If the values of the key diagnostic indicators in the report sheet, including total lung capacity (TLC) actual / predicted and residual volume (RV) actual / predicted, are null, then the corresponding report sheet only needs to predict ventilation function; otherwise, both ventilation function and diffusion function need to be predicted;
[0099] (1) The steps to predict ventilation function are as follows:
[0100] Step f1: Load the pre-trained random forest prediction model M1 for predicting ventilation function;
[0101] Step f2: Input the data set D0 to be diagnosed into the model M1 and obtain the output value Y v ;
[0102] Step f3: Change Y v Mapped to a string S describing the ventilation function v ;
[0103] (2) The steps for predicting diffusion function are as follows:
[0104] Step g1: Load the pre-trained random forest prediction model M2 for predicting diffusion function;
[0105] Step g2: Input the data set D0 to be diagnosed into the model M1 to obtain the output value Y d ;
[0106] Step g3: Y d Mapped to a string S describing the diffusion function d .
[0107] In one embodiment of the present invention, the result output module reads the report file to be filled in with the test conclusion, combines the test index evaluation and the diagnosis conclusion, fills them into the diagnosis conclusion box of the test report, and outputs the prediction result through the following steps:
[0108] Step h1: Use the data extraction module to read the report form PDF file for the test conclusion to be filled in;
[0109] Step h2: The string S output by the diagnostic module c 、S v 、S d Merge into the diagnostic result string S r ;
[0110] Step h3: Use the contour recognition algorithm to identify the position of the black rectangular box in the PDF file. This black rectangular box is formally represented as a point set P = {P1, P2, P3, P4} containing 4 nodes, where P i (i=1,2,3,4) represents the coordinates of the four vertices of the black rectangle in clockwise order. The black rectangle is used to output the text information of the diagnosis result;
[0111] Step h4: Translate the diagnostic result string S r Write to the black rectangular area P in the report PDF file, generate a report file with the diagnosis conclusion and save it.
[0112] The above are preferred embodiments of the present invention. Any changes made according to the technical solution of the present invention, as long as the resulting functions and effects do not exceed the scope of the technical solution of the present invention, shall fall within the scope of protection of the present invention.
Claims
1. A lung function diagnosis and prediction system based on an improved random forest model with weighted regularized information gain, characterized in that: It includes data extraction module, data preprocessing module, feature selection module, random forest training module, diagnosis module and result output module; among them, Data extraction module, used to extract data from pulmonary function test report files in PDF, TXT or image formats; Data preprocessing module, used to clean and standardize the indicators extracted by the data extraction module; A feature selection module is used to determine some key indicators for lung function diagnosis from the indicators; Random forest training module, used to train random forest prediction models using data sets with correct diagnostic conclusions; The diagnosis module is used to diagnose lung function based on the trained random forest prediction model; The result output module is used to display the diagnosis results.
2. The lung function diagnosis and prediction system based on the improved random forest model with weighted regularized information gain according to claim 1, characterized in that: The data extraction module includes a text recognition function that can identify and extract data from PDF files and TXT files, and designs a rectangular contour recognition algorithm that can identify the areas where various indicators of lung function are located from image files and extract indicator data from case reports; the diagnosis module applies the trained model to the diagnosis of the pulmonary test report currently input by the user, and converts the predicted results into diagnostic conclusions; the result output module is used to generate a detailed diagnostic report displayed on the system interface, and output the diagnosed pulmonary test report.
3. A lung function diagnosis and prediction system based on an improved random forest model with weighted regularized information gain according to claim 1 or 2, characterized in that: Data extraction module, the specific functions are as follows: (1) If the report to be processed is a text file, read the data information through the following steps: Step a1: Read the file name suffix of the report, which is label F s , F s If it is 1, it means that the file has a detection conclusion. s If it is 0, it means that there is no detection conclusion for the file; Step a2: Read the full text of the report into the string str f ; Step a3: From str f Extract the character string str1 representing personal information and the character string str2 representing the indicator detection value; Step a4: Read further indicators from str2, including unit, actual value, predicted value, and actual / predicted value information, and record them in array A str ; (2) If the report to be processed is an image file, the rectangular area contour recognition algorithm is used to extract the indicator data from the image. The steps are as follows: Step b1: Read the file name suffix of the report, which is label F P , F P If it is 1, it means that the file has a detection conclusion. P If it is 0, it means that there is no detection conclusion for the file; Step b1: converting the report to be processed into a grayscale image G; Step b2: Input algorithm parameters Δx and Δy, which represent the length and width of the area where the indicator value is located respectively; Step b3: Perform edge detection on G to obtain image G containing all line segments c ; Step b4: Identify G c Calculate the angle between each line segment and the horizontal line, delete the line segments with an angle greater than 5°, and keep the remaining line segments in the set n l is the number of line segments; Step b5: Pair the line segments in L in pairs. a ,l b ) Get the starting and ending coordinates; where the straight line l a ={(x a1 ,y a1 ),(x a2 ,y a2 )},(x a1 ,y a1 ) and (x a2 ,y a2 ) is l a The starting and ending coordinates of the straight line l b ={(x b1 ,y b1 ),(x b2 ,y b2 )}, where (x b1 ,y b1 ) and (x b2 ,y b2 ) is l b The starting and ending coordinates of l a 、l b , calculate the maximum value of the horizontal coordinate x max =max{x a1 ,x a2 ,x b1 ,x a2 } and the minimum value of the horizontal axis x min =min{x a1 ,x a2 ,x b1 ,x a2 }, calculate the maximum value of the vertical coordinate y max =max{y a1 ,y a2 ,y b1 ,y a2 } and the minimum value of the vertical coordinate y min =min{y a1 ,y a2 ,y b1 ,y a2 }; Select (x min ,y min ),(x min ,y max ),(x max ,y min ),(x max ,y max ) Four coordinates in image G c The area in which the image is located forms an image segment s l , all line segment pairs (l a ,l b ) constitutes an image segment collection n s is the number of line segment pairs; Step b9: For any image segment s in S i , calculate s i The horizontal axis difference and the vertical coordinate difference The first one to satisfy and s i The image segment s identified as containing indicator data o , δ is the allowable error of the horizontal axis, ε is the allowable error of the vertical axis; Step b10: Input algorithm parameter d, where d represents the pixel width of an indicator value on the report image, and s o From left to right are the indicator name, indicator unit, indicator actual value, indicator predicted value and the ratio of indicator actual value to indicator predicted value. Scan s from right to left according to width d. o , extract the three sub-image segments s representing the ratio of the actual value of the indicator to the predicted value of the indicator, the predicted value of the indicator and the actual value of the indicator ap ,s p ,s a ; Step b11: Use OCR tools to classify s ap ,s p ,s a Convert to 3 strings and record them in string array A pic middle.
4. The lung function diagnosis and prediction system based on the improved random forest model with weighted regularized information gain according to claim 3, characterized in that: The data preprocessing module uses the following steps to clean and standardize the extracted data, including removing invalid or erroneous data and converting the data into a unified format: Step c1: A pic The incorrectly identified indicator values are replaced with the pre-given indicator default values; Step c2: A pic For the indicator value that exceeds γ, the value is divided by 10 until it is less than γ, where γ is the preset upper limit of the indicator value; Step c3: According to F P and F s Determine whether the corresponding test report file has a diagnosis conclusion. If A str or A pic The corresponding test report file has a diagnosis conclusion. str or A pic Generates a collection of indicator values O p ={C1,C2,...,C nc }, C1, C2 corresponds to A str or A pic The index value in n c Indicates the number of indicator values; The diagnosis is converted into L b ; O p and L b Combination is (O p ,L b ) output to the set S t In the random selection S t 70% of the dataset is used as the training set, and the remaining 30% is used as the validation set; if A str or A pic The corresponding test report file does not have a diagnosis conclusion. str or A pic Input the report set S to be predicted p .
5. A lung function diagnosis and prediction system based on an improved random forest model with weighted regularized information gain according to claim 1 or 2, characterized in that: The feature selection module determines the most influential key diagnostic indicators based on the correlation between the indicators and lung function diagnosis. The feature selection module inputs the key indicator set S k ={VC (vital capacity), MVV (maximal ventilation), FEV1 (forced expiratory volume in the first second), FEV1 / FVC (one-second rate), FV curve value, TLC (total lung capacity), RV (residual volume), RV / TLC (residual to total lung capacity ratio)}, output key diagnostic indicators.
6. A lung function diagnosis and prediction system based on an improved random forest model with weighted regularized information gain according to claim 1 or 2, characterized in that: The random forest training module is trained on the detection report file set F with correct diagnosis labels. The key diagnostic indicators obtained by the feature selection module are used as input feature data, and the correct diagnostic results are used as labels to train the random forest prediction model for diagnosis; The random forest training module designs a new weighted regularized information gain. First, the feature f is assigned a weight ω(f) and a regularization penalty term is defined. The penalty term considers two factors: the number of values of feature f and the frequency of use of feature f in the decision tree. In this way, if the possible value of feature f exceeds the predetermined value, or the number of times feature f is used in the decision tree exceeds the predetermined value, the value of the regularization term C(f) will increase, and the model will be penalized. At the same time, SMOTE upsampling is performed on the dataset X. The dataset X is the dataset obtained by processing the test report file set F with correct diagnostic labels through the data extraction module and the data preprocessing module. The weighted regularized information gain is used to measure the quality of node splitting in the random forest. The weighted regularized information gain calculation formula is: in, is the weighted regularized information gain of feature f on dataset X, ω(f) is the weight of feature f, λ is a parameter, G(X,f) is the information gain of feature f on dataset X, which is used to measure the quality of splitting using feature f. The calculation formula is: Among them, E(S) is the entropy of the parent node, E(S i ) is the entropy of the i-th child node, |S i | and |S| represent the number of samples of parent nodes and child nodes respectively, v represents the total number of child nodes; E(S i ) and C(f) are calculated as follows: C(f)=α·N(f)+β·F(f) Where K is the total number of possible categories of samples in the dataset, p(k) is the proportion of samples in the kth category in the dataset, α and β are parameters, N(f) represents the number of possible values of feature f in the dataset, and F(f) is the frequency of feature f in the decision tree, which is used to measure the influence of feature f on the complexity of the model. The calculation formula is: Among them, C used (f,t) is the number of times feature f is used in the tth tree, and T is the total number of decision trees.
7. The lung function diagnosis and prediction system based on the improved random forest model with weighted regularized information gain according to claim 6, characterized in that: The specific steps of the random forest training module using the improved random forest algorithm based on weighted information gain to train the random forest prediction model are as follows: Step d1: using the data extraction module to read the data set D from the report file set F; Step d2: Use the data preprocessing module to preprocess the data set D to obtain the data set D'; Step d3: Perform SMOTE upsampling on the minority class whose sample number in the dataset D' is less than the predetermined value to obtain the preprocessed dataset D"; Step d31: Determine the upsampling ratio to be 1:1, so that the number of minority class samples is increased to be equal to the number of majority class samples; Step d32: Generate synthetic minority class samples according to the following interpolation formula and add them to the dataset; SynS i =M i +γ·(N i -M i ) Among them, SynS i is a synthetic sample, M i is the selected minority class sample, N i is the nearest neighbor of the selected sample, γ is a random number from 0 to 1; Step d4: Use Boostrap sampling, that is, randomly select samples from the dataset D' with replacement until a subsample set N of the same size as the original dataset is obtained; Step d5: Use N to select features based on recursive feature elimination and train a decision tree; Step d6: Use step d5 to train k decision trees in parallel to obtain random forest prediction models M1 and M2 for predicting ventilation function and diffusion function, respectively.
8. The lung function diagnosis and prediction system based on the improved random forest model with weighted regularized information gain according to claim 7, characterized in that: Step d5 specifically includes the following steps: Step d51: Set the number of features to be reduced by 10% in each iteration, set the target number of features m, and use all the features in the dataset to train the decision tree; Step d52: Calculate the weighted regularized information gain for node splitting using each feature f Step d53: Remove weighted regularization information gain The lowest 10% of features; Step d54: Repeat steps d52 to d53 until the number of features is reduced to a preset number m.
9. A lung function diagnosis and prediction system based on an improved random forest model with weighted regularized information gain according to claim 1 or 2, characterized in that: The diagnosis module uses a newly designed diagnostic rule base to convert the test results into test index evaluations. At the same time, it calls the random forest prediction module to generate diagnostic conclusions based on the test index values in the report to be predicted. The specific implementation steps are as follows: Step e1: Design a diagnostic rule base consisting of the following rules to convert key diagnostic indicator values into detection indicator evaluations; the diagnostic rule base can add, modify, and delete key diagnostic indicators in real time; Rule 1: For vital capacity (VC), the ratio of the measured value to the predicted value is: Rule 2: For MVV actual / predicted, the ratio of the measured value to the predicted value of MVV is: Rule 3: For the FEV1 in the first second, the actual / predicted value is the ratio of the measured value to the predicted value in the first second: Rule 4: For the measured value of FEV1 / FVC in one second, that is, the ratio of the measured value to the predicted value of the one-second rate: Rule 5: For TLC actual / predicted, the ratio of the measured value to the predicted value of the total lung capacity is: Rule 6: For residual gas value RV actual / predicted, that is, the ratio of the measured value to the predicted value of the residual gas value: Rule 7: For the residual-to-total ratio RV / TLC actual / predicted, that is, the ratio of the measured value to the predicted value of the residual-to-total ratio: Step e2: Use the rule base mapping to obtain the description of each indicator and store it in the string S c middle; Step e3: If the values of the key diagnostic indicators in the report sheet, including total lung capacity (TLC) actual / predicted and residual volume (RV) actual / predicted, are null, then the corresponding report sheet only needs to predict ventilation function; otherwise, both ventilation function and diffusion function need to be predicted; (1) The steps for predicting ventilation function are as follows: Step f1: Load the pre-trained random forest prediction model M1 for predicting ventilation function; Step f2: Input the data set D0 to be diagnosed into the model M1 and obtain the output value Y v ; Step f3: Change Y v Mapped to a string S describing the ventilation function v ; (2) The steps for predicting diffusion function are as follows: Step g1: Load the pre-trained random forest prediction model M2 for predicting diffusion function; Step g2: Input the data set D0 to be diagnosed into the model M1 to obtain the output value Y d ; Step g3: Y d Mapped to a string S describing the diffusion function d .
10. The lung function diagnosis and prediction system based on the improved random forest model with weighted regularized information gain according to claim 9, characterized in that: The result output module reads the report file to be filled in with the test conclusion, combines the test index evaluation and the diagnosis conclusion, fills it into the diagnosis conclusion box of the test report, and outputs the prediction result through the following steps: Step h1: Use the data extraction module to read the report form PDF file for the test conclusion to be filled in; Step h2: The string S output by the diagnostic module c 、S v 、S d Merge into the diagnostic result string S r ; Step h3: Use the contour recognition algorithm to identify the position of the black rectangular box in the PDF file. This black rectangular box is formally represented as a point set P = {P1, P2, P3, P4} containing 4 nodes, where P i , i = 1, 2, 3, 4, are the coordinates of the four vertices of the black rectangular box represented in clockwise order, and the black rectangular box is used to output the text information of the diagnosis result; Step h4: Translate the diagnostic result string S r Write to the black rectangular area P in the report PDF file, generate a report file with the diagnosis conclusion and save it.