Image recognition method and device and computer equipment

By screening and optimizing multiple initial image recognition models and combining with integrated strategies, the accuracy of benign and malignant identification of bile duct dilation in MRCP images is improved, the problem of low accuracy in the prior art is solved, and more efficient diagnostic effects are achieved.

CN120495809APending Publication Date: 2025-08-15SOUTHWEST MEDICAL UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510483898.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-17
Publication Date
2025-08-15

AI Technical Summary

Technical Problem

In the prior art, the accuracy of judging benign and malignant bile duct dilation through MRCP images is low.

Method used

Image recognition is performed through multiple initial image recognition models, based on preset filtering strategies and parameter optimization strategies, the target image recognition model with performance optimization is selected, the MRCP image data is identified, and the identification accuracy is improved in combination with the integrated strategy.

Benefits of technology

It improves the accuracy of benign and malignant identification of bile duct dilation in MRCP images and the robustness of the model, and enhances the accuracy and reliability of the diagnosis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120495809A_ABST
    Figure CN120495809A_ABST
Patent Text Reader

Abstract

The invention relates to an image recognition method and device and computer equipment. The method comprises the following steps: respectively carrying out image recognition on sample image data through a plurality of initial image recognition models to obtain initial recognition results respectively output by each initial image recognition model; in the plurality of initial image recognition models, based on a preset screening strategy and each initial recognition result, carrying out screening to obtain a plurality of first image recognition models; optimizing at least one hyper-parameter in each first image recognition model through a parameter optimization strategy to obtain a second image recognition model corresponding to each first image recognition model; in the plurality of second image recognition models, screening based on a preset performance condition to obtain a target image recognition model; and identifying the to-be-detected image data based on the target image identification model to obtain an image identification result. By adopting the method, the accuracy of judging benign and malignant biliary duct expansion can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of medical image analysis, and in particular to an image recognition method, apparatus, and computer equipment. Background Art

[0002] The bile duct is an important channel for transporting bile secreted by the liver to the small intestine, and is mainly composed of bile duct epithelial cells. Due to bile duct stenosis, obstruction or other pathological changes caused by various reasons, both intrahepatic and extrahepatic bile ducts may be abnormally dilated. The diameter of the common bile duct in healthy adults is less than 8mm, the normal left hepatic duct diameter is about 3.5mm, and the right hepatic duct diameter is about 3.3mm. A diameter exceeding this value is called bile duct dilatation (BDD). There are many causes for intrahepatic and extrahepatic BDD, which can be roughly divided into benign and malignant dilatations. Benign causes mainly include congenital biliary dilatation, cholelithiasis and inflammatory stenosis, while malignant causes mainly include bile duct cancer, pancreatic cancer, ampullary cancer and metastatic tumors.

[0003] Current diagnostic methods for biliary diseases include non-invasive techniques such as ultrasound (US), computed tomography (CT), and magnetic resonance imaging (MRI), as well as invasive methods such as endoscopic retrograde cholangiopancreatography (ERCP), endoscopic ultrasound (EUS), and intraductal ultrasound (IDUS). US offers a low-cost, real-time, portable, and radiation-free imaging option, but its image quality is affected by multiple factors, including patient fitness, gastrointestinal gas, and patient cooperation. CT offers high spatial resolution but low contrast and involves ionizing radiation. ERCP is considered the gold standard for detecting biliary strictures, but this invasive diagnostic method can lead to complications such as infection, pancreatitis, and bleeding. MRI, especially Magnetic Resonance Cholangiopancreatography (MRCP), is widely used in the diagnosis of biliary diseases because of its non-invasive, relatively accurate, and ability to clearly depict the pancreatic duct and the entire biliary tree.

[0004] However, in the related art, conventional image recognition algorithms are usually used to determine the benign or malignant nature of BDD through MRCP images. However, conventional image recognition algorithms have low accuracy in determining the benign or malignant nature of BDD through MRCP images. Summary of the Invention

[0005] Based on this, it is necessary to provide an image recognition method, device and computer equipment that can improve the accuracy of judging the benign or malignant nature of bile duct dilatation in order to address the above technical problems.

[0006] In a first aspect, the present application provides an image recognition method, comprising:

[0007] Performing image recognition on the sample image data through multiple initial image recognition models to obtain initial recognition results output by each initial image recognition model;

[0008] Screening the multiple initial image recognition models based on a preset screening strategy and the initial recognition results to obtain multiple first image recognition models;

[0009] optimizing at least one hyperparameter in each of the first image recognition models using a parameter optimization strategy to obtain a second image recognition model corresponding to each of the first image recognition models; and selecting the target image recognition model from the plurality of second image recognition models based on preset performance conditions;

[0010] Based on the target image recognition model, the image data to be detected is recognized to obtain an image recognition result, wherein the image recognition result represents a good or bad probability value corresponding to the image data to be detected.

[0011] In one embodiment, the method further comprises:

[0012] Preprocessing a plurality of to-be-detected image data of the target object to obtain a plurality of preprocessed image data;

[0013] The plurality of pre-processed image data are respectively inputted into the target image recognition model to obtain a plurality of image recognition results of the target object.

[0014] In one embodiment, the method further comprises:

[0015] Based on the target integration strategy, prediction processing is performed on each of the image recognition results to obtain the actual prediction result of the target object.

[0016] In one embodiment, the method further comprises:

[0017] Identify multiple sample image data of the same target object based on the target image recognition model to obtain a predicted recognition result corresponding to each sample image data;

[0018] A plurality of integrated strategies are used to perform prediction processing on multiple prediction and recognition results of the same target object respectively, and the prediction results corresponding to the target object of each preset integrated strategy are obtained; based on the prediction results, the integrated strategy that meets the preset prediction conditions is determined as the target integrated strategy.

[0019] In one embodiment, determining the integration strategy that meets the preset prediction condition based on the prediction result as the target integration strategy includes:

[0020] determining a first true positive rate and a first false positive rate of the sample image data according to the prediction result and a first label threshold corresponding to the prediction result; and calculating a first performance indicator of the integration strategy based on the first true positive rate and the first false positive rate;

[0021] The integration strategy corresponding to the highest first performance indicator is determined as the target integration strategy.

[0022] In one embodiment, the hyperparameter includes multiple values; the hyperparameter includes at least one or more of a parameter optimization algorithm, a learning rate, a learning rate adjustment algorithm, a resolution normalization parameter, and an image cropping size; and optimizing at least one hyperparameter in the first image recognition model through the parameter optimization strategy to obtain the second image recognition model includes:

[0023] For each hyperparameter, the first image recognition model is adjusted using multiple values corresponding to the hyperparameter to obtain an adjusted first image recognition model, and performance test results corresponding to the adjusted first image recognition model are obtained, and the values corresponding to the performance test results that meet the preset performance screening conditions are determined as the optimized values corresponding to the hyperparameter;

[0024] Based on the optimized values corresponding to each of the hyperparameters, the hyperparameters of the first image recognition model are optimized to obtain an optimized first image recognition model, and the optimized first image recognition model is determined as the second image recognition model.

[0025] In one embodiment, the screening is performed based on a preset screening strategy and each of the initial recognition results in a plurality of initial image recognition models to obtain a plurality of first image recognition models, including:

[0026] For each of the initial image recognition models, determining a second true positive rate and a second false positive rate of the sample image data based on the initial recognition result and a second label threshold corresponding to the initial recognition result; and calculating a second performance indicator of the initial image recognition model based on the second true positive rate and the second false positive rate;

[0027] A target number of target performance indicators are screened from the plurality of second performance indicators based on a preset screening strategy, and an initial image recognition model corresponding to each target performance indicator is determined as the first image recognition model.

[0028] In one embodiment, screening the plurality of second image recognition models based on preset performance conditions to obtain a target image recognition model includes:

[0029] Inputting the sample image data into each of the second image recognition models to obtain a plurality of first recognition results;

[0030] determining a third true positive rate and a third false positive rate of the sample image data based on the plurality of first recognition results and a third label threshold corresponding to each of the first recognition results; and calculating a third performance indicator of the second image recognition model based on the third true positive rate and the third false positive rate;

[0031] The second image recognition model corresponding to the highest third performance indicator among the plurality of third performance indicators is determined as the target image recognition model.

[0032] In a second aspect, the present application further provides an image recognition device, comprising:

[0033] An image recognition module is used to perform image recognition on sample image data using multiple initial image recognition models to obtain initial recognition results output by each initial image recognition model;

[0034] A first screening module is configured to screen a plurality of initial image recognition models based on a preset screening strategy and each of the initial recognition results to obtain a plurality of first image recognition models;

[0035] a second screening module configured to optimize at least one hyperparameter in each of the first image recognition models using a parameter optimization strategy to obtain a second image recognition model corresponding to each of the first image recognition models; and to screen the plurality of second image recognition models based on preset performance conditions to obtain a target image recognition model;

[0036] The recognition module is used to recognize the image data to be detected based on the target image recognition model to obtain an image recognition result, wherein the image recognition result represents the good or bad probability value corresponding to the image data to be detected.

[0037] In a third aspect, the present application further provides a computer device comprising a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the following steps are implemented:

[0038] Performing image recognition on the sample image data through multiple initial image recognition models to obtain initial recognition results output by each initial image recognition model;

[0039] Screening the multiple initial image recognition models based on a preset screening strategy and the initial recognition results to obtain multiple first image recognition models;

[0040] optimizing at least one hyperparameter in each of the first image recognition models using a parameter optimization strategy to obtain a second image recognition model corresponding to each of the first image recognition models; and selecting the target image recognition model from the plurality of second image recognition models based on preset performance conditions;

[0041] Based on the target image recognition model, the image data to be detected is recognized to obtain an image recognition result, wherein the image recognition result represents a good or bad probability value corresponding to the image data to be detected.

[0042] In a fourth aspect, the present application further provides a computer-readable storage medium having a computer program stored thereon, wherein when the computer program is executed by a processor, the following steps are implemented:

[0043] Performing image recognition on the sample image data through multiple initial image recognition models to obtain initial recognition results output by each initial image recognition model;

[0044] Screening the multiple initial image recognition models based on a preset screening strategy and the initial recognition results to obtain multiple first image recognition models;

[0045] optimizing at least one hyperparameter in each of the first image recognition models using a parameter optimization strategy to obtain a second image recognition model corresponding to each of the first image recognition models; and selecting the target image recognition model from the plurality of second image recognition models based on preset performance conditions;

[0046] Based on the target image recognition model, the image data to be detected is recognized to obtain an image recognition result, wherein the image recognition result represents a good or bad probability value corresponding to the image data to be detected.

[0047] In a fifth aspect, the present application further provides a computer program product, comprising a computer program, which, when executed by a processor, implements the following steps:

[0048] Performing image recognition on the sample image data through multiple initial image recognition models to obtain initial recognition results output by each initial image recognition model;

[0049] Screening the multiple initial image recognition models based on a preset screening strategy and the initial recognition results to obtain multiple first image recognition models;

[0050] optimizing at least one hyperparameter in each of the first image recognition models using a parameter optimization strategy to obtain a second image recognition model corresponding to each of the first image recognition models; and selecting the target image recognition model from the plurality of second image recognition models based on preset performance conditions;

[0051] Based on the target image recognition model, the image data to be detected is recognized to obtain an image recognition result, wherein the image recognition result represents a good or bad probability value corresponding to the image data to be detected.

[0052] The above-mentioned image recognition method, apparatus and computer equipment screen multiple first image recognition models from multiple initial image recognition models based on the initial recognition results output by the sample image data in multiple initial image recognition models and a preset screening strategy, and optimize the hyperparameters in each first image recognition model through a parameter optimization strategy to obtain a first image recognition model after hyperparameter optimization, determine the first image recognition model after hyperparameter optimization as the second image recognition model, and screen a target image recognition model that meets the preset performance conditions in each second image recognition model, and process the image data to be detected based on the target image recognition model to obtain an image recognition result, which can improve the accuracy of the target image recognition model in recognizing the image, improve the robustness of the model, and improve the accuracy of malignant and benign bile duct dilatation in MRCP images. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following briefly introduces the drawings required for use in the embodiments of the present application or related technical descriptions. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without paying any creative work.

[0054] Figure 1 1 is a flow chart of an image recognition method according to an embodiment;

[0055] Figure 2 Schematic diagram of the structure of a target image recognition model in one embodiment;

[0056] Figure 3A A schematic diagram of performance test results of a first image recognition model corresponding to different parameter optimization algorithms in one embodiment;

[0057] Figure 3B is a schematic diagram of performance test results of a first image recognition model corresponding to different learning rates in one embodiment;

[0058] Figure 3CA schematic diagram of performance test results of a first image recognition model corresponding to different learning rate adjustment algorithms in one embodiment;

[0059] Figure 3D A schematic diagram of performance test results of a first image recognition model corresponding to different resolution normalization parameters in one embodiment;

[0060] Figure 3E A schematic diagram of performance test results of a first image recognition model for different image cropping sizes in one embodiment;

[0061] Figure 4 is a schematic diagram of performance indicators of an initial image recognition model in one embodiment;

[0062] Figure 5A Schematic diagram of the loss function of the target image recognition model corresponding to the training set and the validation set during the training process in one embodiment;

[0063] Figure 5B A schematic diagram of the accuracy curves of the target image recognition model corresponding to the training set and the validation set during the training process in one embodiment;

[0064] Figure 6A A schematic diagram of the performance indicators of each layer of the training set in the target image recognition model in one embodiment;

[0065] Figure 6B A schematic diagram of the performance indicators of each layer of the target image recognition model on the validation set in one embodiment;

[0066] Figure 6C A schematic diagram of the performance indicators of each layer of the target image recognition model for a test set in one embodiment;

[0067] Figure 7 is a confusion matrix of five different integration strategies and results in sample image data in one embodiment;

[0068] Figure 8 1 is a flow chart of an image recognition method in one embodiment;

[0069] Figure 9 is a structural block diagram of an image recognition device in one embodiment;

[0070] Figure 10 FIG. 1 is a diagram showing the internal structure of a computer device in one embodiment. DETAILED DESCRIPTION

[0071] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.

[0072] In an exemplary embodiment, Figure 1 As shown, an image recognition method is provided. This embodiment uses the method applied to a terminal as an example for illustration. It is understandable that the method can also be applied to a server, or to a system including a terminal and a server, and implemented through interaction between the terminal and the server. In this embodiment, the method includes the following steps:

[0073] Step 101: Perform image recognition on sample image data using multiple initial image recognition models to obtain initial recognition results output by each initial image recognition model.

[0074] The initial image recognition model may be an initial neural network model, and the types of the multiple initial image recognition models may be different. The sample image data may be sample image data with a known label threshold, and the label threshold may be used to characterize the type of the sample image data. The sample image data may be a magnetic resonance cholangiopancreatography (MRCP) image. The sample image data may include multiple sample image data for multiple different target objects and the same target object. The initial recognition result may reflect the probability value of benign or malignant bile duct dilatation on the MRCP image.

[0075] Specifically, the terminal can obtain sample image data corresponding to different target objects, as well as multiple sample image data of the same target object. For each initial image recognition model, the terminal performs image recognition on the sample image data using the initial image recognition model to obtain an initial recognition result for each sample image data.

[0076] Furthermore, the sample image data are all preprocessed sample image data. Specifically, the preprocessing process may include: the terminal may obtain initial sample images and label thresholds corresponding to each initial sample image. The terminal may process the window width, window position, crop, and pixels of each initial sample image to obtain sample image data. For example, the window width and window position of the initial sample image may be adjusted to [1500, 800], and the pixels may be normalized to [0, 255].

[0077] Step 102 : Filter the multiple initial image recognition models based on a preset screening strategy and each initial recognition result to obtain multiple first image recognition models.

[0078] Specifically, the terminal can determine the performance indicators of each initial image recognition model based on the initial recognition results in each initial image recognition model and the label threshold corresponding to the initial recognition results, and filter out multiple first image recognition models from multiple initial image recognition models based on the performance indicators and preset screening strategies.

[0079] Step 103: Optimize at least one hyperparameter in each first image recognition model through a parameter optimization strategy to obtain a second image recognition model corresponding to each first image recognition model; and screen the multiple second image recognition models based on preset performance conditions to obtain a target image recognition model.

[0080] Among them, the parameter optimization strategy may refer to a strategy for adjusting and optimizing the hyperparameters in the first image recognition model. The second image recognition model is an image recognition model that has undergone hyperparameter optimization. The preset performance condition may refer to a pre-set image recognition model performance evaluation standard, and the performance indicators for evaluating the performance may include one or more of the area under the curve (AUC) value, accuracy (Acc), sensitivity (Sen), specificity (Spe), positive predictive value (PPV), and negative predictive value (NPV).

[0081] Specifically, for each first image recognition model, the terminal can optimize at least one hyperparameter in the first image recognition model through a parameter optimization strategy to obtain an optimized first image recognition model, and determine the optimized first image recognition model as the corresponding second image recognition model. The terminal can input sample image data into each second image recognition model respectively to obtain image recognition results corresponding to each second image recognition model. The terminal can test the performance of each second image recognition model based on the image recognition results corresponding to the second image recognition model to obtain performance indicators of each second image recognition model. The terminal can select performance indicators that meet preset performance conditions from the performance indicators of each second image recognition model, and determine the second image recognition model corresponding to the performance indicators that meet the preset performance conditions as the target image recognition model.

[0082] In addition, the terminal can also resample the sample image data through a preset resampling algorithm to obtain resampled sample image data. The terminal can use the resampled sample image data to retrain the second image recognition model and obtain the performance indicators corresponding to the current second image recognition model. And perform the process of resampling, training and evaluation for a preset number of times to obtain the distribution of each performance indicator, and based on the distribution of the performance indicator, calculate the confidence interval of the performance indicator, and determine whether the model is stable based on whether the confidence interval is within the preset range. The preset resampling algorithm can be a bootstrap method. For the pairwise second image recognition models, the terminal uses the DeLong test to compare the performance indicators of the two second image recognition models to obtain a test value. If it is stable and meets the preset performance conditions, the second image recognition model is determined as the target image recognition model.

[0083] For the second image recognition model determined above, the terminal can use the Hosmer and Lemeshow goodness-of-fit test to evaluate the goodness of fit of the two second image recognition models, draw a calibration curve showing the relationship between the predicted probability and the actual probability, and further evaluate the clinical utility of the model by drawing a clinical decision curve and a clinical impact curve. If the discrimination, calibration, and clinical applicability all meet the preset performance conditions, the second image recognition model is determined as the target image recognition model.

[0084] The terminal can also obtain sample image data from different sources, which can be divided into internal test sets and external test sets based on the source. The terminal can input the sample image data from different sources into each second image recognition model, obtain the recognition results of each sample image data in each image recognition model, and determine the performance indicators of each second image recognition model based on the performance of each second image recognition model in recognizing each sample image data. Based on the performance indicators of the second image recognition model for each sample image data, the second image recognition model corresponding to the performance indicator with the best performance for each sample image data is determined as the target image recognition model.

[0085] Step 104: Based on the target image recognition model, the image data to be detected is recognized to obtain an image recognition result.

[0086] The image data to be detected is image data for which the recognition result is unknown, and the image data to be detected may be an MRCP image. The image recognition result represents a good or bad probability value corresponding to the image data to be detected. The target image recognition module may include an entry flow module, a middle flow module, and an exit flow module.

[0087] Specifically, the terminal can input the image data to be detected into the inlet flow module of the target image recognition model for preliminary feature extraction and pooling processing to obtain a preliminary image feature map, and input the preliminary image feature map into the intermediate flow module for a preset number of convolution processing, activation processing, jump connection processing and pooling operations in sequence to obtain a processed feature map, and input the processed feature map into the outlet flow module for global average pooling and full connection layer processing to generate an image recognition result. Among them, the preset number of times in the intermediate flow module can be determined according to the actual application scenario, for example, it can be 8 times. The target image recognition model can include 14 modules and 36 convolution layers. Figure 2 As shown, Figure 2 It is a structural diagram of the target image recognition model.

[0088] The above-mentioned image recognition method screens multiple first image recognition models from multiple initial image recognition models based on the initial recognition results output by the sample image data in multiple initial image recognition models and a preset screening strategy, and optimizes the hyperparameters in each first image recognition model through a parameter optimization strategy to obtain a first image recognition model after hyperparameter optimization, determines the first image recognition model after hyperparameter optimization as the second image recognition model, and screens a target image recognition model that meets the preset performance conditions in each second image recognition model, and processes the image data to be detected based on the target image recognition model to obtain an image recognition result, which can improve the accuracy of the target image recognition model in recognizing the image, improve the robustness of the model, and improve the accuracy of malignant and benign bile duct dilatation in MRCP images.

[0089] In an exemplary embodiment, the image recognition method further includes:

[0090] Preprocessing a plurality of to-be-detected image data of a target object to obtain a plurality of preprocessed image data; inputting the plurality of preprocessed image data into a target image recognition model respectively to obtain a plurality of image recognition results of the target object.

[0091] Specifically, the terminal may preprocess the plurality of to-be-detected image data of the target object to obtain a plurality of preprocessed image data. The preprocessing process may specifically include one or more of the following: uniformly adjusting the window width and window level of the plurality of to-be-detected image data according to preset data, pixel normalization adjustment, and uniformly adjusting the image crop size.

[0092] The terminal can input the preprocessed image data into the entry flow module of the target image recognition model for preliminary feature extraction and pooling processing to obtain a preliminary image feature map, input the preliminary image feature map into the intermediate flow module for a preset number of convolution processing, activation processing, jump connection processing and pooling operations in sequence to obtain a processed feature map, input the processed feature map into the exit flow module for global average pooling and full connection layer processing to generate an image recognition result.

[0093] In this embodiment, multiple image data to be detected of the target object are preprocessed and the preprocessed data are input into the target image recognition model to generate multiple image recognition results. This not only improves the quality of the image data and the accuracy of recognition, but also supports a comprehensive analysis of the target object and significantly improves the accuracy and reliability of diagnosis.

[0094] In an exemplary embodiment, the image recognition method further includes:

[0095] Based on the target integration strategy, each image recognition result is predicted and processed to obtain the actual prediction result of the target object.

[0096] Among them, the target integration strategy can be a method for comprehensively predicting multiple image recognition results. The target integration strategy can be screened from multiple different integration strategies based on the first performance indicator predicted by each integration strategy. For example, the target integration strategy can include one of a direct averaging method, a voting method, an experience weighted average (WAE) method, a model weighted average (WAM) method, and a logistic regression method. The actual prediction result can represent the benign or malignant probability value of the target object.

[0097] Specifically, the terminal may perform prediction processing on each image recognition result of the same target object based on the target integration strategy to obtain an actual prediction result of the same target object.

[0098] When the target integration strategy is the direct averaging method, the terminal may calculate the average value of the image recognition results corresponding to the target object, and determine the average value as the actual prediction result of the target object.

[0099] When the target integration strategy is the voting method, the terminal can determine the image recognition result of the target object, calculate the probability value of the recognition result type corresponding to the image recognition result of the target object, obtain the average probability value of the recognition result type, and determine the actual prediction result of the target object based on the average probability value.

[0100] When the target integration strategy is the empirical weighted average method, the terminal can assign weights based on the importance of the images, for example, giving higher weights to images near the middle plane that more comprehensively displays the biliary tree structure, and decreasing weights toward the periphery. The weighted average of the individual recognition results for each target object is then calculated and used as the actual prediction result for the target object. For example, the weights corresponding to Layers 1 to 9 are 1, 2, 3, 4, 5, 4, 3, 2, and 1, respectively.

[0101] When the target integration strategy is the model weighted averaging method, the terminal can use the area under the curve corresponding to each image recognition result in the first image recognition model as a weight, and perform weighted averaging on the AUC corresponding to each image recognition result of the target object as a weight to obtain a weighted average result, and determine the weighted average result as the actual prediction result.

[0102] When the target integration strategy is the Logistic regression method, the benign or malignant prediction probability (image recognition result) is used as the dependent variable, and the image data to be detected is used as the independent variable. A Logistic regression equation is constructed for secondary modeling and integration to obtain a Logistic regression model. The actual prediction result of the target object is calculated based on the constructed Logistic regression model formula. The specific formula of the Logistic regression model can be:

[0103] Log(p / (1-p))=β0+β1×Layer1+β2×Layer2+β3×Layer3+……βn×Layern

[0104] Where p is the probability value of the actual prediction result being malignant; β0, ..., βn are the coefficients of the logistic regression model; Layer1, ..., Layern are the corresponding layer numbers in the internal dataset.

[0105] In this embodiment, by optimizing multiple image recognition models and combining them with a target integration strategy, the model with the best performance is selected and the image recognition results are predicted. This not only improves the accuracy of the image recognition results, but also enhances the generalization ability and robustness of the model prediction.

[0106] In an exemplary embodiment, the image recognition method further includes:

[0107] Based on the target image recognition model, multiple sample image data of the same target object are identified to obtain the predicted recognition results corresponding to each sample image data; multiple integration strategies are used to perform prediction processing on the multiple predicted recognition results of the same target object respectively to obtain the prediction results corresponding to the target object of each preset integration strategy; based on the prediction results, the integration strategy that meets the preset prediction conditions is determined as the target integration strategy.

[0108] The predicted recognition result may represent the goodness or badness probability value of the target object. The multiple integration strategies may include at least one of a direct averaging method, a voting method, a weighted average of individual estimates (WAE) method, a weighted average maturity (WAM) method, and a logistic regression method.

[0109] Specifically, the terminal can identify multiple sample image data of the same target object based on the target image recognition model to obtain a predicted recognition result corresponding to each sample image data. The specific recognition process is consistent with the process of identifying the image recognition result of the image to be detected in the above embodiment, and will not be repeated here.

[0110] The terminal can use multiple integration strategies to predict multiple prediction and recognition results for the same target object, respectively, to obtain prediction results corresponding to each prediction integration strategy. The specific process of using multiple integration strategies to predict the prediction and recognition results is consistent with the process of using the target integration strategy in the above embodiment, and will not be repeated here.

[0111] The terminal can determine the first performance indicator corresponding to each integration strategy based on the prediction results corresponding to each integration strategy, select the highest first performance indicator from each first performance indicator, and determine the integration strategy corresponding to the highest first performance indicator as the target integration strategy.

[0112] In this embodiment, by selecting a target integration strategy through multiple integration strategies, the recognition of the target image recognition model can be further predicted, the accuracy of image recognition prediction can be improved, and the generalization ability and robustness of the model can be enhanced.

[0113] In an exemplary embodiment, the specific implementation process of the step of “determining, based on the prediction results, an integration strategy that meets the preset prediction conditions as a target integration strategy” may include:

[0114] According to the prediction result and the first label threshold corresponding to the prediction result, a first true positive rate and a first false positive rate of the sample image data are determined; and a first performance indicator of the integration strategy is calculated based on the first true positive rate and the first false positive rate; and the integration strategy corresponding to the highest first performance indicator is determined as the target integration strategy.

[0115] Among them, the first label threshold can represent the actual probability value of the preset type of the sample image data corresponding to the prediction result. The preset type can be benign or malignant. The first true positive rate is the proportion of samples that are actually positive that are determined to be positive by the integrated strategy, and the first false positive rate is the proportion of samples that are actually negative that are determined to be positive by the integrated strategy. The first performance indicator can be used to evaluate the quantitative indicators of the integrated strategy, for example, one or more of AUC, ACC, Sen, Spe, PPV and NPV.

[0116] Specifically, the terminal can obtain a first label threshold for each sample image data. The terminal can also determine the confusion matrix of each integrated strategy based on the first label threshold and the prediction result. The confusion matrix can be a 2×2 matrix. The confusion matrix can include the number of sample images correctly predicted as benign by each integrated strategy (true positive) (True Positive, TP), the number of samples incorrectly predicted as benign (false positive) (False Positive, FP), the number of samples incorrectly predicted as negative (false negative) (False Negative, FN), and the number of samples correctly predicted as negative (true negative) (True Negative, TN). The terminal can determine the ratio of the number of true positives to the actual number of benign samples (the sum of the number of true positives and the number of false negatives) as a first true positive rate, and determine the ratio of the number of false positives to the actual number of malignant samples (the sum of the number of false positives and the number of true negatives) as a first false positive rate.

[0117] The terminal can use the false positive rate as the horizontal coordinate and the true positive rate as the vertical coordinate to generate a receiver operating characteristic curve (ROC) curve of the first true positive rate and the first false positive rate, and calculate the area value under the ROC curve, determine the area value as the AUC value, and determine the AUC value as the first performance indicator. The terminal can filter out the highest first performance indicator from the first performance indicators corresponding to each integration strategy, and determine the integration strategy corresponding to the highest first performance indicator as the target integration strategy. The specific method of filtering the highest first performance indicator can be determined by a sorting method, and the specific method of filtering is not limited here. The specific calculation of the area value under the ROC curve can be determined by accumulating the trapezoidal areas between each adjacent point, which is not specifically limited here.

[0118] In addition, the terminal may calculate a first sum of TP and TN, and a second sum of TP, TN, FP, and FN, and determine the ratio of the first sum to the ratio of the second sum as the ACC value of the integrated strategy. The terminal may calculate the ratio of TP to the sum of TP and FN and determine it as the Sen value of the integrated strategy. The terminal may determine the Spe value of the integrated strategy by the ratio of TN to the sum of TN and FP. The terminal may determine the PPV value of the integrated strategy by the ratio of TP to the sum of TP and FP. The terminal may determine the NPV value of the integrated strategy by the ratio of TN to the sum of TN and FN.

[0119] When comprehensively evaluating each integration strategy through AUC, ACC, Sen, Spe, PPV, and NPV, the terminal can determine the key performance indicators based on actual application requirements and determine the integration strategy corresponding to the highest performance indicator as the target integration strategy.

[0120] In this embodiment, by selecting the best integration strategy from multiple integration strategies as the target integration strategy, the image recognition results in the target image recognition model are further predicted, thereby improving the accuracy of the prediction of the image recognition results and improving the robustness and generalization of the model.

[0121] In an exemplary embodiment, the specific implementation process of "optimizing at least one hyperparameter in the first image recognition model using a parameter optimization strategy to obtain a second image recognition model" in step 103 may include:

[0122] For each hyperparameter, the first image recognition model is adjusted using multiple values corresponding to the hyperparameter to obtain an adjusted first image recognition model, and performance test results corresponding to the adjusted first image recognition model are obtained, and the values corresponding to the performance test results that meet the preset performance screening conditions are determined as the optimized values corresponding to the hyperparameter;

[0123] Based on the optimized values corresponding to each hyperparameter, the hyperparameters of the first image recognition model are optimized to obtain an optimized first image recognition model, and the optimized first image recognition model is determined as the second image recognition model.

[0124] Among them, hyperparameters include multiple values. Hyperparameters include at least one or more of parameter optimization algorithm, learning rate, learning rate adjustment algorithm, resolution normalization parameter, and image cropping size. For example, parameter optimization algorithms may include Stochastic Gradient Descent (SGD), Adaptive Delta (Adadelta), Adaptive Gradient (AdaGrad), and Adaptive Moment Estimation (Adam), etc. The value of the learning rate can be set according to the type of model or application scenario, such as learning rates can include 0.01, 0.001, 0.0001, etc., and learning rate adjustment algorithms can include step learning rate (StepLR), exponential learning rate (Exponential LR), and reduced learning rate on the platform (ReduceLRnPlateau), etc. The resolution normalization parameter can be a resolution normalization standard, and the value of the resolution normalization standard can include 0.5×0.5×0.5, 1×1×1, 2×2×2, 3×3×3, 4×4×4, 5×5×5, etc. The image cropping size may refer to a region cut out from the original image, for example, the image cropping size may be 32, 64, 128, 256, 512, 1024, etc. The performance test results may include one or more of AUC, ACC, Sen, Spe, PPV, and NPV.

[0125] Specifically, the terminal can determine the hyperparameters that need to be optimized in each first image recognition model, and determine the data of each hyperparameter. For each value of each hyperparameter, the hyperparameters of the first image recognition model are adjusted respectively to obtain the first image recognition model after adjusting the parameters. The terminal can input the sample image data into the first image recognition model after adjusting the parameters to obtain the recognition result corresponding to the sample image data, and the terminal can determine the performance test result of the first image recognition model after adjusting the parameters based on the recognition result. The calculation method of the performance test result is consistent with the process of calculating the performance parameters of each integration strategy in the above embodiment, and will not be repeated here.

[0126] For each hyperparameter, the terminal can select the highest performance test result from the performance test results of the first image recognition model after adjusting the parameters corresponding to different values, and use the first image recognition model after adjusting the parameters corresponding to the highest performance test result as the optimized value corresponding to the hyperparameter. Figure 3A-Figure 3E As shown, Figure 3AIt is the performance indicator of the first image recognition model of different parameter optimization algorithms. The horizontal axis is the type of the first image recognition model, including ResNet101, VggNet, DenseNet121, Xception, and ResNet152 in sequence; the vertical axis is the performance test result (AUC); the parameter optimization algorithm corresponding to the black icon is adaptive gradient (AdaGrad), the parameter optimization algorithm corresponding to the slash icon is adaptive moment estimation (Adam), the parameter optimization algorithm corresponding to the grid icon is stochastic gradient descent (SGD), and the parameter optimization algorithm corresponding to the sparse point icon is adaptive Delta (Adadelta). Figure 3B This is a schematic diagram of the performance test results of the first image recognition model corresponding to different learning rates; the horizontal axis is the type of the first image recognition model, including ResNet101, VggNet, DenseNet121, Xception, and ResNet152 in order; the vertical axis is the performance test result (AUC); the slash icon corresponds to a learning rate LR=0.01, the black icon corresponds to a learning rate LR=0.001, and the grid icon corresponds to a learning rate LR=0.0001; Figure 3C These are the performance test results of the first image recognition model with different learning rate adjustment algorithms. The horizontal axis represents the type of the first image recognition model, which includes ResNet101, VggNet, DenseNet121, Xception, and ResNet152 in sequence; the vertical axis represents the performance test result (AUC); the learning rate adjustment algorithm corresponding to the slash icon is ReduceLRnPlateau, the learning rate adjustment algorithm corresponding to the black icon is ExponentialLR, and the learning rate adjustment algorithm corresponding to the grid icon is StepLR. Figure 3D is the performance test result of the first image recognition model with different resolution normalization parameters, where the horizontal axis is the type of the first image recognition model, including ResNet101, VggNet, DenseNet121, Xception, and ResNet152 in order; the vertical axis is the performance test result (AUC); the resolution normalization parameter corresponding to the diagonal icon is 0.5×0.5×0.5; the resolution normalization parameter corresponding to the black icon is 1×1×1; the resolution normalization parameter corresponding to the grid icon is 2×2×2; the resolution normalization parameter corresponding to the sparse dot icon with a white background is 3×3×3; the resolution normalization parameter corresponding to the sparse dot icon with a gray background is 4×4×4; the resolution normalization parameter corresponding to the dense dot icon with a white background is 5×5×5; Figure 3EThese are the performance test results of the first image recognition model with different image cropping sizes, where the horizontal axis is the type of the first image recognition model, including ResNet101, VggNet, DenseNet121, Xception, and ResNet152 in sequence; the vertical axis is the performance test result (AUC); the image cropping size corresponding to the slash icon is 32; the image cropping size corresponding to the black icon is 64; the image cropping size corresponding to the grid icon is 128; the image cropping size corresponding to the sparse dot icon with a white background is 256; the image cropping size corresponding to the sparse dot icon with a gray background is 512; and the image cropping size corresponding to the dense dot icon with a white background is 1024.

[0127] The terminal can optimize the hyperparameters of the first image recognition model based on the optimized values corresponding to each hyperparameter to obtain an optimized first image recognition model, and determine the optimized first image recognition model as the second image recognition model.

[0128] For example, the terminal can test different parameter optimization algorithms, including SGD, Adadelta, AdaGrad, and Adam, based on the selected first image recognition model, and test different learning rates and learning rate adjustment algorithms, including StepLR, Exponential LR, and ReduceLRnPlateau. In order to ensure the compatibility between the input data and the model while improving computational efficiency, different resolution normalization methods are tested. In the image preprocessing stage, the original image is cropped and different crop sizes are selected to enhance the model's focus on key image areas. Finally, by observing the damage curve and various performance curves of each model, hyperparameters are adjusted for each model.

[0129] In this embodiment, by optimizing hyperparameters, the recognition accuracy and generalization ability of the model are significantly improved. By experimentally testing multiple values, the optimal hyperparameter configuration is automatically selected. It is also determined that the optimal hyperparameter value can meet actual application needs, thereby improving the practicality of the model.

[0130] In an exemplary embodiment, the specific implementation process of step 102 of "screening multiple initial image recognition models based on a preset screening strategy and each initial recognition result to obtain multiple first image recognition models" may include:

[0131] For each initial image recognition model, determining a second true positive rate and a second false positive rate of the sample image data based on the initial recognition result and a second label threshold corresponding to the initial recognition result; and calculating a second performance indicator of the initial image recognition model based on the second true positive rate and the second false positive rate;

[0132] A target number of target performance indicators are screened from the plurality of second performance indicators based on a preset screening strategy, and an initial image recognition model corresponding to each target performance indicator is determined as the first image recognition model.

[0133] Among them, the second label threshold can represent the actual probability value of the preset type of the sample image data corresponding to the initial recognition result. The preset type can be benign or malignant. The second true positive rate is the proportion of samples that are actually positive that are determined to be positive by the initial image recognition model, and the second false positive rate is the proportion of samples that are actually negative that are determined to be positive by the initial image recognition model. The second performance indicator can be used to evaluate the quantitative indicators of the initial image recognition model, for example, one or more of AUC, ACC, Sen, Spe, PPV and NPV.

[0134] Specifically, the terminal can obtain a second label threshold for each sample image data. The terminal can also determine the confusion matrix of each initial image recognition model based on the second label threshold and the initial recognition result. The confusion matrix can be a 2×2 matrix. The confusion matrix can include the number of sample images correctly predicted as benign by each initial image recognition model (true positive) (True Positive, TP), the number of samples incorrectly predicted as benign (false positive) (False Positive, FP), the number of samples incorrectly predicted as negative (false negative) (False Negative, FN), and the number of samples correctly predicted as negative (true negative) (True Negative, TN). The terminal can determine the ratio of the number of true positives to the actual number of benign samples (the sum of the number of true positives and the number of false negatives) as the second true positive rate, and determine the ratio of the number of false positives to the actual number of malignant samples (the sum of the number of false positives and the number of true negatives) as the second false positive rate.

[0135] The terminal can use the false positive rate as the horizontal coordinate and the true positive rate as the vertical coordinate to generate a receiver operating characteristic curve (ROC) curve of the second true positive rate and the second false positive rate, and calculate the area value under the ROC curve, determine the area value as the AUC value, and determine the AUC value as the second performance indicator. The terminal can filter out the highest second performance indicator from the second performance indicators corresponding to each initial image recognition model, and determine the initial image recognition model corresponding to the highest second performance indicator as the first image recognition model. The specific method of filtering the highest second performance indicator can be determined by a sorting method, and the specific method of filtering is not limited here. The specific calculation of the area value under the ROC curve can be determined by accumulating the trapezoidal areas between each adjacent point, which is not specifically limited here.

[0136] In addition, the terminal may calculate a first sum of TP and TN, and a second sum of TP, TN, FP, and FN, and determine the ratio of the first sum to the ratio of the second sum as the ACC value of the initial image recognition model. The terminal may calculate the ratio of TP to the sum of TP and FN as the Sen value of the initial image recognition model. The terminal may determine the ratio of TN to the sum of TN and FP as the Spe value of the initial image recognition model. The terminal may determine the ratio of TP to the sum of TP and FP as the PPV value of the initial image recognition model. The terminal may determine the ratio of TN to the sum of TN and FN as the NPV value of the initial image recognition model.

[0137] When comprehensively evaluating each initial image recognition model through AUC, ACC, Sen, Spe, PPV and NPV, the terminal can determine the key performance indicators based on actual application requirements, and determine the initial image recognition model corresponding to the highest performance indicator as the first image recognition model.

[0138] For example, the terminal can determine 14 initial image recognition models with default parameters, including an 18-layer residual network (ResNet18), a 34-layer residual network (ResNet34), a 50-layer residual network (ResNet50), a 101-layer residual network (ResNet101), a 152-layer residual network (ResNet152), AlexNet, Visual Geometry Group VGG network (VggNet), SqueezeNet, Inception, Depthwise Separable Convolutional Network (Xception), a 121-layer densely connected network (DenseNet121), a 169-layer densely connected network (DenseNet169), a 50-layer extended version of the residual network (ResNext50), and a 101-layer extended version of the residual network (ResNext101). The terminal can obtain the initial recognition results of the sample image data recognized by each model. The terminal can perform performance tests on each model to obtain performance indicators of each model, select the initial image recognition model corresponding to the top 5 performance indicators, and determine the initial image recognition model corresponding to the top 5 performance indicators as the first image recognition model. Figure 4 As shown, Figure 4It is the performance index corresponding to each initial image recognition model. The horizontal axis is the different initial image recognition models, ResNet18, ResNet34, ResNet50, ResNet101, ResNet152, AlexNet, VggNet, SqueezeNet, Inception, Xception, DenseNet121, DenseNet169, ResNext50 and ResNext101, and the vertical axis is the performance index value (AUC).

[0139] In an embodiment, several image recognition models with relatively good performance indicators are determined from multiple initial image recognition models to facilitate subsequent screening of target image recognition models. Through layer-by-layer screening, the robustness of the model and the accuracy of image recognition are improved.

[0140] In an exemplary embodiment, the specific implementation process of step 103 of "screening the plurality of second image recognition models based on preset performance conditions to obtain a target image recognition model" may include:

[0141] The sample image data is input into each second image recognition model to obtain multiple first recognition results; the third true positive rate and the third false positive rate of the sample image data are determined based on the multiple first recognition results and the third label threshold corresponding to each first recognition result; and the third performance indicator of the second image recognition model is calculated based on the third true positive rate and the third false positive rate; the second image recognition model corresponding to the highest third performance indicator among the multiple third performance indicators is determined as the target image recognition model.

[0142] Among them, the third label threshold can represent the actual probability value of the preset type of the sample image data corresponding to the prediction result. The preset type can be benign or malignant. The third true positive rate is the proportion of samples that are actually positive that are determined to be positive by the second image recognition model, and the third false positive rate is the proportion of samples that are actually negative that are determined to be positive by the second image recognition model. The third performance indicator can be used to evaluate the quantitative indicators of the second image recognition model, for example, one or more of AUC, ACC, Sen, Spe, PPV and NPV.

[0143] Specifically, the terminal can obtain a third label threshold for each sample image data. The terminal can also determine the confusion matrix of each second image recognition model based on the third label threshold and the prediction result. The confusion matrix can be a 2×2 matrix. The confusion matrix can include the number of sample images correctly predicted as benign by each second image recognition model (true positive) (TruePositive, TP), the number of samples incorrectly predicted as benign (false positive) (False Positive, FP), the number of samples incorrectly predicted as negative (false negative) (False Negative, FN), and the number of samples correctly predicted as negative (true negative) (True Negative, TN). The terminal can determine the ratio of the number of true positives to the actual number of benign samples (the sum of the number of true positives and the number of false negatives) as the third true positive rate, and determine the ratio of the number of false positives to the actual number of malignant samples (the sum of the number of false positives and the number of true negatives) as the third false positive rate.

[0144] The terminal can use the false positive rate as the horizontal coordinate and the true positive rate as the vertical coordinate to generate a receiver operating characteristic curve (ROC) curve of the third true positive rate and the third false positive rate, and calculate the area value under the ROC curve, determine the area value as the AUC value, and determine the AUC value as the third performance indicator. The terminal can filter out the highest third performance indicator from the third performance indicators corresponding to each second image recognition model, and determine the second image recognition model corresponding to the highest third performance indicator as the target image recognition model. The specific method of filtering the highest third performance indicator can be determined by a sorting method, and the specific method of filtering is not limited here. The specific calculation of the area value under the ROC curve can be determined by accumulating the trapezoidal areas between each adjacent point, which is not specifically limited here.

[0145] In addition, the terminal may calculate a first sum of TP and TN, and a second sum of TP, TN, FP, and FN, and determine the ratio of the first sum to the ratio of the second sum as the ACC value of the second image recognition model. The terminal may calculate the ratio of TP to the sum of TP and FN as the Sen value of the second image recognition model. The terminal may determine the ratio of TN to the sum of TN and FP as the Spe value of the second image recognition model. The terminal may determine the ratio of TP to the sum of TP and FP as the PPV value of the second image recognition model. The terminal may determine the ratio of TN to the sum of TN and FN as the NPV value of the second image recognition model.

[0146] When comprehensively evaluating each second image recognition model through AUC, ACC, Sen, Spe, PPV and NPV, the terminal can determine the performance indicators to focus on based on actual application requirements, and determine the second image recognition model corresponding to the highest performance indicator among the performance indicators as the target image recognition model.

[0147] Additionally, to further improve the performance and generalization of the image recognition model while reducing overfitting, we used a variety of data augmentation methods, including scaling, resampling, contrast adjustment, and Gaussian noise, as well as transfer learning with pre-training on ImageNet, a large visual database used for visual object recognition software research.

[0148] In this embodiment, the present application improves the accuracy of image recognition by determining a target image recognition model among multiple second image recognition models.

[0149] In one embodiment, an image recognition model (deep learning model) involves selecting the optimal architecture from a set of baseline neural networks based on their performance on an internal test set. The AUCs of the selected architectures ResNet101, VggNet, DenseNet121, Xception, and ResNet152 all exceed 0.8. All models exhibited optimal prediction performance on the internal test set with a learning rate of 0.001. However, different neural networks tended to prefer different methods when selecting parameter optimization and learning rate adjustment algorithms. In general, Adagrad and ReduceLROnPlateau performed best. Subsequent adjustments included resolution normalization parameters and crop size. For ResNet101, VggNet, DenseNet121, and ResNet152, the optimal resolution was 1×1×1, while for Xception, the optimal resolution was 0.5×0.5×0.5. Different crop sizes are suitable for different models. The performance of the internal test set at each step is shown below. Figure 3A-Figure 3E shown.

[0150] After the terminal was able to employ data augmentation, ImageNet transfer learning, and hyperparameter tuning, ResNet101 demonstrated excellent performance on the internal test set, with an AUC of 0.846 (95% CI, 0.817-0.874). However, its performance on the external test set was significantly lower, with an AUC of 0.689 (95% CI, 0.656-0.723). While ResNet152, VggNet, and DenseNet121 also demonstrated strong predictive performance on the internal test set, their performance on the external test set declined. Notably, the Xception model achieved excellent predictive performance on both the internal and external test sets, with AUCs of 0.816 (95% CI, 0.788-0.844) and 0.807 (95% CI, 0.779-0.835), respectively (Table 1). The terminal selected the Xception neural network as the final deep learning model. The model was trained with a learning rate of 0.001 and adjustments were managed by ReduceLROnPlatform (factor = 0.1, patience = 10). The model used cross entropy loss to minimize the difference between the predicted and actual probability distributions. Adagrad was used as the optimizer for fast parameter adjustment. The model was trained for 50 epochs with a batch size of 32 images and early stopping was used to prevent overfitting. Figure 5A and Figure 5B As shown in the figure, the horizontal axis is the number of training steps (Epochs) of the target image recognition model, the vertical axis is the loss function value (LOSS), the black solid line is the loss function value of the training set in the target image recognition model, and the gray solid line is the loss function value of the validation set in the target image recognition model; Figure 5B It is the accuracy curve of the training set and the validation set, where the horizontal axis is the number of training steps (Epochs) of the target image recognition model, the vertical axis is the ACC value, the black solid line is the loss function value of the training set in the target image recognition model, and the gray solid line is the loss function value of the validation set in the target image recognition model.

[0151] Table 1 Performance indicators of target image recognition models on various datasets

[0152]

[0153] The terminal integrates the predictions of deep learning models using five different integration strategies, including direct averaging, voting, WAE, WAM, and logistic regression models. The AUCs of Layer 1 to Layer 9 in the internal test set are 0.809, 0.785, 0.784, 0.837, 0.866, 0.855, 0.809, 0.814, and 0.774, respectively. The specific performance of each layer in the training set, validation set, and internal test set is as follows: Figures 6A-6C shown. Figure 6A The performance of the target image recognition model in different layers on the training set in terms of various evaluation indicators. The horizontal coordinates represent different layers, from Layer 1 to Layer 9, and the vertical coordinates represent performance indicator values. The dot combination line represents the AUC value, the triangle combination line represents the ACC value, the square combination line represents the Sen value, the cross combination line represents the Spe value, the circled multiplication sign combination line represents the F1-score, the cross combination line represents the PPV value, and the gray solid line represents the NPV value. Figure 6B The horizontal coordinates represent the performance of the target image recognition model in different layers on the validation set in terms of various evaluation indicators. The horizontal coordinates represent different layers, from Layer 1 to Layer 9, and the vertical coordinates represent the performance indicator values. The dot combination line represents the AUC value, the triangle combination line represents the ACC value, the square combination line represents the Sen value, the cross combination line represents the Spe value, the circled multiplication sign combination line represents the F1-score, the cross combination line represents the PPV value, and the gray solid line represents the NPV value. Figure 6C The horizontal coordinates represent the performance of the target image recognition model in different layers in terms of various evaluation indicators on the test set. The horizontal coordinates represent different layers, from Layer 1 to Layer 9, and the vertical coordinates represent the performance indicator values. The dot combination line represents the AUC value, the triangle combination line represents the ACC value, the square combination line represents the Sen value, the cross combination line represents the Spe value, the circled multiplication sign combination line represents the F1-score value, the cross combination line represents the PPV value, and the gray solid line represents the NPV value.

[0154] The specific formula of the logistic regression model can be (p is the probability of malignant BDD):

[0155] log(p1-p)=-3.396+1.294×Layer1+0.616×Layer2-1.755×Layer3+0.336×Layer4+3.42×Layer5+2.172×Layer6-1.071×Layer7+1.171×Layer8+0.71×Layer9

[0156] The AUC values of the five integrated strategies for patient-level benign and malignant prediction in each dataset exceeded the AUC values of the deep learning model for single MRCP image prediction, indicating that the integrated strategy improved the accuracy of benign and malignant prediction at the patient level. The internal test set showed that the prediction accuracy differences among the five integrated strategies were minimal, and the pairwise DeLong test p values were all greater than 0.05. In addition, the confusion matrix describes the actual and predicted results of each integrated strategy for benign and malignant BDD patients on the dataset. Figure 7 As shown, Figure 7This is the confusion matrix of five different ensemble strategies and their results in the sample image data. Each row is the confusion matrix corresponding to the true labels and preset labels of the training set, validation set, internal test set, and external test set for each ensemble strategy. The first row is the confusion matrix corresponding to direct averaging, the second row is the confusion matrix corresponding to voting, the third row is the confusion matrix corresponding to WAE, the fourth row is the confusion matrix corresponding to WAM, and the fifth row is the confusion matrix corresponding to the logistic regression model. The number in the upper left corner of each confusion matrix is the number of true positive samples, the number in the upper right corner is the number of false negative samples, the number in the lower left corner is the number of false positive samples, and the number in the lower right corner is the number of true negative samples. Table 2 lists in detail the specific prediction performance of the four ensemble strategies on the four datasets, showing that the ensemble strategy outperforms the deep learning model.

[0157] The ensemble strategy with the highest AUC in both the internal and external test sets was identified as logistic regression. The final ensemble model, Xception and Logistic, was called Xce-LR, and Xce-LR was calibrated and evaluated for clinical utility. The Brier score on the internal test set was 0.14, and the Hosmer-Lemeshow goodness-of-fit test p-value was 0.109, indicating good accuracy. In this example, clinical decision and clinical impact curves were used to evaluate the utility of the Xce-LR model in distinguishing benign from malignant BDD based on predicted risk thresholds. These curves were constructed to assess the potential benefit of using the model for initial risk stratification, rather than to represent specific clinical treatment decisions. The dichotomization of benign and malignant BDD can serve as a preliminary screening tool to identify cases that may require further diagnosis and specific staging procedures.

[0158] Table 2 Prediction performance of five ensemble strategies on each dataset

[0159]

[0160] The diagnostic performance of the Xce-LR model was compared with that of three radiologists of varying experience levels using a prospective cohort dataset. Radiologist 1, with 2 years of experience, correctly classified 21 of 30 benign cases and 25 of 30 malignant cases, with an accuracy of 76.7%, a sensitivity of 83.3%, and a specificity of 70.0%. Radiologist 2 (5 years) performed slightly better, correctly classifying 22 cases as benign and 26 as malignant, with an overall accuracy of 80.0%, a sensitivity of 86.7%, and a specificity of 73.3%. Radiologist 3, with the most experience (10 years), achieved the highest diagnostic accuracy, correctly classifying 24 cases as benign and 28 cases as malignant, with an accuracy of 86.7%, a sensitivity of 93.3%, and a specificity of 80.0%. The Xce-LR model, optimized using a threshold of 0.636 determined by the maximum Youden index, outperformed all three radiologists in the cohort. The model correctly identified 27 of 30 benign cases and 27 of 30 malignant cases, for an overall accuracy of 90.0% and a sensitivity and specificity of 90.0%.

[0161] The comparative analysis highlighted a key finding: In six cases (four benign and two malignant), all three radiologists made incorrect assessments, but the Xce-LR model correctly identified these cases. This discrepancy suggests that the Xce-LR model may provide additional value in cases where the imaging findings are subtle or ambiguous, making traditional interpretation challenging. Challenging cases misclassified by three radiologists but correctly identified by the model; cases were analyzed by MRCP, axial fat-suppressed T2-weighted imaging (T2WI), diffusion-weighted imaging (DWI), and T1-weighted imaging (T1WI). The first case was a 49-year-old female with ampullary carcinoma, and no malignant signs were found on T2WI, DWI, and T1WI; cases 2-4 were a 63-year-old female, a 58-year-old female, and a 71-year-old male, all diagnosed with inflammatory strictures; case 2 had circumferential thickening and stenosis of the distal common bile duct with mild DWI hyperintensity; case 3 had soft tissue signal at the portal vein on T2WI and no DWI hyperintensity; case 4 had mixed signal on T2WI and hyperintensity on DWI in the left hepatic lobe.

[0162] Notably, there were two cases in which all three radiologists correctly diagnosed the condition, but the Xce-LR model provided an incorrect classification. In one instance, a young patient with benign BDD was incorrectly classified as malignant by the model, likely due to the model's reliance solely on MRCP images without considering the patient's age and obvious signs of malignancy on other MRI sequences. In another case involving a patient with gallbladder cancer, the radiologists were able to accurately diagnose malignancy based on the presence of a visible mass on non-MRCP sequences, whereas the model, focusing solely on MRCP features, incorrectly classified the case as benign due to the absence of bile duct damage.

[0163] In one embodiment, Figure 8 As shown, Figure 8 This is a flowchart of an image recognition method, including: sample image collection and preprocessing, baseline neural network model selection, algorithm and parameter adjustment in the neural network model, model evaluation, and selection of an integration strategy.

[0164] It should be understood that, although the various steps in the flowcharts involved in the various embodiments described above are displayed in sequence according to the instructions of the arrows, these steps are not necessarily executed in sequence in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and these steps can be executed in other orders. Moreover, at least a portion of the steps in the flowcharts involved in the various embodiments described above can include multiple steps or multiple stages, and these steps or stages are not necessarily executed and completed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily to be carried out in sequence, but can be executed in turn or alternately with other steps or at least a portion of steps or stages in other steps.

[0165] Based on the same inventive concept, embodiments of the present application also provide an image recognition device for implementing the aforementioned image recognition method. The solution provided by this device is similar to the solution described in the aforementioned method. Therefore, the specific limitations in one or more of the following image recognition device embodiments can be found in the above-described limitations on the image recognition method and will not be further elaborated here.

[0166] In an exemplary embodiment, Figure 9 As shown, an image recognition device 90 is provided, comprising: an image recognition module 91, a first screening module 92, a second screening module 93 and a recognition module 94, wherein:

[0167] The image recognition module 91 is used to perform image recognition on the sample image data using multiple initial image recognition models to obtain initial recognition results output by each initial image recognition model;

[0168] A first screening module 92 is configured to screen multiple initial image recognition models based on a preset screening strategy and each initial recognition result to obtain multiple first image recognition models;

[0169] A second screening module 93 is configured to optimize at least one hyperparameter in each first image recognition model using a parameter optimization strategy to obtain a second image recognition model corresponding to each first image recognition model; and to screen the plurality of second image recognition models based on preset performance conditions to obtain a target image recognition model;

[0170] The recognition module 94 is used to recognize the image data to be detected based on the target image recognition model to obtain an image recognition result, which represents the good or bad probability value corresponding to the image data to be detected.

[0171] In one embodiment, the recognition module 94 is further configured to pre-process the plurality of to-be-detected image data of the target object to obtain a plurality of pre-processed image data;

[0172] The plurality of pre-processed image data are respectively input into the target image recognition model to obtain a plurality of image recognition results of the target object.

[0173] In one embodiment, the recognition module 94 is further configured to perform prediction processing on each image recognition result based on a target integration strategy to obtain an actual prediction result of the target object.

[0174] In one embodiment, the recognition module 94 is further configured to recognize multiple sample image data of the same target object based on the target image recognition model, and obtain a predicted recognition result corresponding to each sample image data;

[0175] A variety of integrated strategies are used to perform prediction processing on multiple prediction and recognition results of the same target object, and the prediction results corresponding to the target object of each preset integrated strategy are obtained; based on the prediction results, the integrated strategy that meets the preset prediction conditions is determined as the target integrated strategy.

[0176] In one embodiment, the second screening module 93 is configured to determine a first true positive rate and a first false positive rate of the sample image data based on the prediction result and a first label threshold corresponding to the prediction result; and calculate a first performance indicator of the integration strategy based on the first true positive rate and the first false positive rate;

[0177] The integration strategy corresponding to the highest first performance indicator is determined as the target integration strategy.

[0178] In one embodiment, the second screening module 93 is configured to adjust the first image recognition model for each hyperparameter using multiple values corresponding to the hyperparameter to obtain an adjusted first image recognition model, and obtain performance test results corresponding to the adjusted first image recognition model, and determine the value corresponding to the performance test result that meets the preset performance screening condition as the optimized value corresponding to the hyperparameter;

[0179] Based on the optimized values corresponding to each hyperparameter, the hyperparameters of the first image recognition model are optimized to obtain an optimized first image recognition model, and the optimized first image recognition model is determined as the second image recognition model.

[0180] In one embodiment, the first screening module 92 is configured to determine, for each initial image recognition model, a second true positive rate and a second false positive rate of the sample image data based on the initial recognition result and a second label threshold corresponding to the initial recognition result; and calculate a second performance indicator of the initial image recognition model based on the second true positive rate and the second false positive rate;

[0181] A target number of target performance indicators are screened from the plurality of second performance indicators based on a preset screening strategy, and an initial image recognition model corresponding to each target performance indicator is determined as the first image recognition model.

[0182] In one embodiment, the second screening module 93 is configured to input the sample image data into each second image recognition model to obtain a plurality of first recognition results;

[0183] determining a third true positive rate and a third false positive rate of the sample image data based on the plurality of first recognition results and a third label threshold corresponding to each first recognition result; and calculating a third performance indicator of the second image recognition model based on the third true positive rate and the third false positive rate;

[0184] The second image recognition model corresponding to the highest third performance indicator among the plurality of third performance indicators is determined as the target image recognition model.

[0185] Each module in the above-mentioned image recognition device can be implemented in whole or in part through software, hardware, or a combination thereof. Each module can be embedded in or independent of a processor in a computer device in the form of hardware, or can be stored in a memory in the computer device in the form of software, so that the processor can call and execute the corresponding operations of each module.

[0186] In an exemplary embodiment, a computer device is provided. The computer device may be a terminal, and its internal structure diagram may be as shown in FIG. Figure 10As shown. The computer device includes a processor, a memory, an input / output interface, a communication interface, a display unit and an input device. The processor, the memory and the input / output interface are connected via a system bus, and the communication interface, the display unit and the input device are connected to the system bus via the input / output interface. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The input / output interface of the computer device is used to exchange information between the processor and an external device. The communication interface of the computer device is used to communicate with an external terminal in a wired or wireless manner, and the wireless manner can be implemented through WIFI, a mobile cellular network, near field communication (NFC) or other technologies. When the computer program is executed by the processor, an image recognition method is implemented. The display unit of the computer device is used to form a visually visible picture, which can be a display screen, a projection device or a virtual reality imaging device. The display screen can be a liquid crystal display screen or an electronic ink display screen, and the input device of the computer device can be a touch layer covering the display screen, or a button, trackball or touchpad set on the computer device casing, or an external keyboard, touchpad or mouse.

[0187] Those skilled in the art will understand that Figure 10 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.

[0188] In one embodiment, a computer device is provided, including a memory and a processor. The memory stores a computer program, and the processor implements the steps in the above method embodiments when executing the computer program.

[0189] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps in the above-mentioned method embodiments are implemented.

[0190] In one embodiment, a computer program product is provided, including a computer program, which implements the steps in the above method embodiments when executed by a processor.

[0191] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant regulations.

[0192] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiment methods can be implemented by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, database or other media used in the embodiments provided in this application may include at least one of non-volatile memory and volatile memory. Non-volatile memory may include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory may include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The database involved in the various embodiments provided herein may include at least one of a relational database and a non-relational database. Non-relational databases may include, but are not limited to, distributed databases based on blockchains. The processor involved in the various embodiments provided herein may be, but are not limited to, a general-purpose processor, a central processing unit (CPU), a graphics processing unit (GPU), a digital signal processor (DSP), a programmable logic unit (PLC), a data processing logic unit based on quantum computing, an artificial intelligence (AI) processor, and the like.

[0193] The technical features of the above embodiments can be combined arbitrarily. In order to make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this application.

[0194] The above-described embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present application. It should be noted that a person of ordinary skill in the art may make various modifications and improvements without departing from the spirit of the present application, and these modifications and improvements fall within the scope of protection of the present application. Therefore, the scope of protection of the present application shall be determined by the appended claims.

Claims

1. An image recognition method, characterized in that: The method comprises: Performing image recognition on the sample image data through multiple initial image recognition models to obtain initial recognition results output by each initial image recognition model; Screening the multiple initial image recognition models based on a preset screening strategy and the initial recognition results to obtain multiple first image recognition models; optimizing at least one hyperparameter in each of the first image recognition models using a parameter optimization strategy to obtain a second image recognition model corresponding to each of the first image recognition models; and selecting the target image recognition model from the plurality of second image recognition models based on preset performance conditions; Based on the target image recognition model, the image data to be detected is recognized to obtain an image recognition result, wherein the image recognition result represents a good or bad probability value corresponding to the image data to be detected.

2. The method according to claim 1, characterized in that The method further comprises: Preprocessing a plurality of to-be-detected image data of the target object to obtain a plurality of preprocessed image data; The plurality of pre-processed image data are respectively inputted into the target image recognition model to obtain a plurality of image recognition results of the target object.

3. The method according to claim 2, characterized in that The method further comprises: Based on the target integration strategy, prediction processing is performed on each of the image recognition results to obtain the actual prediction result of the target object.

4. The method according to claim 3, characterized in that The method further comprises: Identify multiple sample image data of the same target object based on the target image recognition model to obtain a predicted recognition result corresponding to each sample image data; A plurality of integrated strategies are used to perform prediction processing on multiple prediction and recognition results of the same target object respectively, and the prediction results corresponding to the target object of each preset integrated strategy are obtained; based on the prediction results, the integrated strategy that meets the preset prediction conditions is determined as the target integrated strategy.

5. The method according to claim 4, characterized in that The step of determining an integration strategy that satisfies a preset prediction condition based on the prediction result as a target integration strategy includes: determining a first true positive rate and a first false positive rate of the sample image data according to the prediction result and a first label threshold corresponding to the prediction result; and calculating a first performance indicator of the integration strategy based on the first true positive rate and the first false positive rate; The integration strategy corresponding to the highest first performance indicator is determined as the target integration strategy.

6. The method according to claim 1, characterized in that The hyperparameters include multiple values; the hyperparameters include at least one or more of parameter optimization algorithm, learning rate, learning rate adjustment algorithm, resolution normalization parameter, and image cropping size; The step of optimizing at least one hyperparameter in the first image recognition model by using a parameter optimization strategy to obtain a second image recognition model includes: For each hyperparameter, the first image recognition model is adjusted using multiple values corresponding to the hyperparameter to obtain an adjusted first image recognition model, and performance test results corresponding to the adjusted first image recognition model are obtained, and the values corresponding to the performance test results that meet the preset performance screening conditions are determined as the optimized values corresponding to the hyperparameter; Based on the optimized values corresponding to each of the hyperparameters, the hyperparameters of the first image recognition model are optimized to obtain an optimized first image recognition model, and the optimized first image recognition model is determined as the second image recognition model.

7. The method according to claim 1, characterized in that The method of screening multiple initial image recognition models based on a preset screening strategy and each of the initial recognition results to obtain multiple first image recognition models includes: For each of the initial image recognition models, determining a second true positive rate and a second false positive rate of the sample image data based on the initial recognition result and a second label threshold corresponding to the initial recognition result; and calculating a second performance indicator of the initial image recognition model based on the second true positive rate and the second false positive rate; A target number of target performance indicators are screened from the plurality of second performance indicators based on a preset screening strategy, and an initial image recognition model corresponding to each target performance indicator is determined as the first image recognition model.

8. The method according to claim 1, characterized in that The step of screening the plurality of second image recognition models based on preset performance conditions to obtain a target image recognition model includes: Inputting the sample image data into each of the second image recognition models to obtain a plurality of first recognition results; determining a third true positive rate and a third false positive rate of the sample image data based on the plurality of first recognition results and a third label threshold corresponding to each of the first recognition results; and calculating a third performance indicator of the second image recognition model based on the third true positive rate and the third false positive rate; The second image recognition model corresponding to the highest third performance indicator among the plurality of third performance indicators is determined as the target image recognition model.

9. An image recognition device, characterized in that: The device comprises: An image recognition module is used to perform image recognition on sample image data using multiple initial image recognition models to obtain initial recognition results output by each initial image recognition model; A first screening module is configured to screen a plurality of initial image recognition models based on a preset screening strategy and each of the initial recognition results to obtain a plurality of first image recognition models; a second screening module configured to optimize at least one hyperparameter in each of the first image recognition models using a parameter optimization strategy to obtain a second image recognition model corresponding to each of the first image recognition models; and to screen the plurality of second image recognition models based on preset performance conditions to obtain a target image recognition model; The recognition module is used to recognize the image data to be detected based on the target image recognition model to obtain an image recognition result, wherein the image recognition result represents the good or bad probability value corresponding to the image data to be detected.

10. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 8 are implemented.