An angiopoietin detection method, device, electronic device and storage medium
By obtaining the pathological images of hematoxylin-eosin staining of hepatocellular carcinoma patients, segmenting the cancer tissue area, extracting features and constructing a pathologic model, the problem that detection accuracy in the prior art is affected by the operator and antibodies is solved, and the accurate detection of angiopoietin expression level is achieved.
Patent Information
- Application Number
- CN202411497551.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-25
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2044-10-25
AI Technical Summary
In the prior art, when detecting the expression level of angiopoietin 2, there is a problem that detection accuracy is affected by the antibody of the test operator and the subject, especially in the detection based on fresh tissue specimens and paraffin tissue specimens.
By obtaining the pathological images of hematoxylin-eosin staining of hepatocellular carcinoma patients, segmenting the cancer tissue area, extracting multiple features, screening the optimal subset of features, constructing a pathologic model, and using the model to calculate the angiopoietin expression probability to detect its expression level.
It improves the accuracy of angiopoietin detection, reduces the error impact caused by the operator and the subject's antibodies, and realizes objective and real detection based on pathological images.
Smart Images

Figure CN119379647B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image processing technology, and in particular, to a method, device, electronic device, and storage medium for detecting angiopoietin. Background Art
[0002] Currently, there are the following several schemes for detecting the expression level of angiopoietin 2 (ANGPT2): real-time detection of peripheral blood cytokines, but it cannot reflect the situation of the tumor parenchyma; detection based on fresh tissue specimens, but there are difficulties in collecting tissue specimens, and the detection accuracy is affected by the antibodies of the detection operator and the subject; detection based on paraffin tissue specimens, and the detection accuracy is also affected by the antibodies of the detection operator and the subject. Summary of the Invention
[0003] The main purpose of the embodiments of this application is to propose a method, device, electronic device, and storage medium for detecting angiopoietin to improve the detection accuracy of angiopoietin.
[0004] To achieve the above object, on the one hand, an angiopoietin detection method is proposed in an embodiment of this application, and the method includes the following steps:
[0005] Obtain hematoxylin-eosin staining pathological images of tissues of different hepatocellular carcinoma patients;
[0006] Segment the cancer tissue regions in each of the pathological images and randomly select them as sub-images;
[0007] Extract multiple features of each of the sub-images, and screen out the optimal feature subset from each of the features;
[0008] Build a pathological omics model based on the optimal feature subset;
[0009] Use the pathological omics model to calculate the angiopoietin expression probability of the corresponding hepatocellular carcinoma patient tissue according to the features, and detect the expression level of angiopoietin according to the angiopoietin expression probability.
[0010] In some embodiments, the obtaining of hematoxylin-eosin staining pathological images of tissues of different hepatocellular carcinoma patients includes the following steps:
[0011] Obtain hematoxylin-eosin staining pathological images of tissues of different hepatocellular carcinoma patients that have been fixed with formalin and embedded in paraffin.
[0012] In some embodiments, the segmenting of the cancer tissue regions in each of the pathological images and randomly selecting them as sub-images includes the following steps:
[0013] Using an image binarization segmentation threshold algorithm, multiple cancer tissue regions are segmented from each of the pathological images according to a preset threshold;
[0014] Each of the cancer tissue regions is segmented into multiple sub-images of a preset size;
[0015] For each of the pathological images, multiple representative sub-images are randomly selected.
[0016] In some embodiments, the extracting multiple features of each of the sub-images includes the following steps:
[0017] The original features and high-order features of each of the sub-images are extracted as the features; wherein, the original features include first-order statistics, gray-level co-occurrence matrix, gray-level run length matrix, gray-level size zone, adjacent gray-level difference matrix, and gray-level dependence relationship matrix; the high-order features are obtained by wavelet transformation of the original features.
[0018] In some embodiments, the screening the optimal feature subset from each of the features includes the following steps:
[0019] The average information gain between each of the features is calculated as the feature category correlation coefficient;
[0020] The sum of the mutual information between each of the features is calculated, and then the sum of the mutual information is divided by the square of the total number of features to obtain the feature redundancy coefficient;
[0021] According to the feature category correlation coefficient and the feature redundancy coefficient, a preset number of features are screened from each of the features as a candidate feature subset;
[0022] The importance scores of each of the features in the candidate feature subset are calculated;
[0023] Several features with relatively low importance score rankings are removed, and the remaining features in the candidate feature subset are used as the optimal feature subset.
[0024] In some embodiments, the constructing a pathomics model based on the optimal feature subset includes the following steps:
[0025] A training set is divided from the images of hepatocellular carcinoma patients obtained from a professional database, and the optimal feature subset is extracted from the training set;
[0026] Based on the optimal feature subset extracted from the training set, the pathomics model is constructed;
[0027] Using the pathomics model, the angiopoietin expression probability corresponding to each of the features in the training set is calculated as the prediction value;
[0028] Calculate the residuals and gradients between the predicted values and the actual values;
[0029] Construct a weak learner based on the residuals, fit the weak learner to the residuals, and train the newly added weak learner along the negative direction of the gradient;
[0030] Add the weak learner to the pathomics model as the new pathomics model;
[0031] Return the angiopoietin expression probabilities corresponding to each of the features in the training set calculated using the pathomics model as predicted values until a preset condition is reached.
[0032] In some embodiments, the method further includes the following steps:
[0033] Use the survminer package in R language to perform binary classification on the angiopoietin expression probabilities output by the pathomics model.
[0034] To achieve the above object, on the other hand, an embodiment of the present application provides an angiopoietin detection device, the device includes:
[0035] An image acquisition unit, configured to acquire hematoxylin-eosin staining pathological images of tissues of different hepatocellular carcinoma patients;
[0036] An image segmentation unit, configured to segment the cancer tissue regions in each of the pathological images and randomly select them as sub-images;
[0037] A feature processing unit, configured to extract multiple features of each of the sub-images and screen out an optimal feature subset from each of the features;
[0038] A model construction unit, configured to construct a pathomics model based on the optimal feature subset;
[0039] An expression level detection unit, configured to calculate the angiopoietin expression probability of the corresponding hepatocellular carcinoma patient tissue according to the features using the pathomics model, and detect the expression level of angiopoietin according to the angiopoietin expression probability.
[0040] To achieve the above object, on the other hand, an embodiment of the present application provides an electronic device, the electronic device includes a memory and a processor, the memory stores a computer program, and when the processor executes the computer program, the above method is implemented.
[0041] To achieve the above object, on the other hand, an embodiment of the present application provides a computer-readable storage medium, the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the above method is implemented.
[0042] The embodiments of the present application at least include the following beneficial effects:
[0043] The present application can obtain hematoxylin-eosin staining pathological images of tissues of different hepatocellular carcinoma patients; segment the cancer tissue regions in each pathological image, and randomly select them as sub-images; extract multiple features of each sub-image, and screen out the optimal feature subset from each feature; construct a pathological omics model based on the optimal feature subset; use the pathological omics model to calculate the angiopoietin expression probability of the corresponding hepatocellular carcinoma patient tissue according to the features, and detect the expression level of angiopoietin according to the angiopoietin expression probability. By obtaining rich features of pathological images and then constructing a pathological omics model to predict the corresponding angiopoietin expression probability for each feature, it is possible to objectively and truly detect the expression level of angiopoietin based on pathological images, reducing the error influence brought by operators and subjects' antibodies compared with the prior art and improving the detection accuracy of angiopoietin. Description of the Drawings
[0044] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0045] Figure 1 It is a schematic flowchart of a method for detecting angiopoietin provided by an embodiment of the present application;
[0046] Figure 2 It is an example flowchart of a method for detecting angiopoietin provided by an embodiment of the present application;
[0047] Figure 3 It is an example diagram of the optimal feature subset provided by an embodiment of the present application;
[0048] Figure 4 It is the ROC curve, calibration curve and decision curve diagram of the training set, test set and external test set provided by an embodiment of the present application;
[0049] Figure 5 It is a schematic structural diagram of a device for detecting angiopoietin provided by an embodiment of the present application;
[0050] Figure 6 It is a schematic hardware structure diagram of an electronic device provided by an embodiment of the present application. Detailed Embodiments
[0051] In order to make the objectives, technical solutions and advantages of the present application more clearly understood, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application. When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the embodiments of the present application. They are only examples of devices and methods consistent with some aspects of the embodiments of the present application as detailed in the appended claims.
[0052] It can be understood that the terms "first", "second", etc. used in the present application may be used herein to describe various concepts, but unless otherwise specified, these concepts are not limited by these terms. These terms are only used to distinguish one concept from another. For example, without departing from the scope of the embodiments of the present application, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the words "if", "when" as used herein may be interpreted as "when...", "while...", or "in response to determining".
[0053] The terms "at least one", "a plurality of", "each", "any one", etc. used in the present application, at least one includes one, two or more than two, a plurality includes two or more than two, each refers to each of the corresponding plurality, and any one refers to any one of the plurality.
[0054] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which the present application belongs. The terms used herein are only for the purpose of describing the embodiments of the present application and are not intended to limit the present application.
[0055] The embodiments of the present application provide an angiopoietin detection method, which relates to the technical field of image processing. The angiopoietin detection method provided by the embodiments of the present application can be applied to a terminal, or to a server, or can also be software running on a terminal or a server. In some embodiments, the terminal may be a smart phone, a tablet computer, a notebook computer, a desktop computer, a smart speaker, a smart watch, a vehicle-mounted terminal, etc., but is not limited thereto; the server side may be configured as an independent physical server, or may be configured as a server cluster or a distributed system composed of multiple physical servers, or may also be configured as a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. The server may also be a node server in a blockchain network; the software may be an application implementing the angiopoietin detection method, etc., but is not limited to the above forms.
[0056] This application can be used in numerous general or special computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet devices, multi-processor systems, microprocessor-based systems, set-top boxes, programmable consumer electronic devices, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and so on. This application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. This application can also be practiced in a distributed computing environment where tasks are executed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.
[0057] Referring to Figure 1 , an embodiment of this application provides a method for detecting angiopoietin, and this method may include but is not limited to including S100 to S140, specifically as follows:
[0058] S100: Obtain hematoxylin-eosin staining pathological images of tissues from different hepatocellular carcinoma patients.
[0059] Specifically, for tissues from different hepatocellular carcinoma patients, this embodiment can respectively obtain multiple HE (Hematoxylin-Eosin) staining sections, that is, pathological images.
[0060] To enrich the pathological images and obtain different pathological images, this embodiment can obtain multiple pathological images from a professional database and obtain the pathological images obtained from the real-time examination of patients.
[0061] As a further implementation manner, S100 may include:
[0062] Obtain hematoxylin-eosin staining pathological images of tissues from different hepatocellular carcinoma patients that have been fixed with formalin and embedded in paraffin.
[0063] Specifically, the tissues of hepatocellular carcinoma patients are fixed with formalin and embedded in paraffin, and then the hematoxylin-eosin staining pathological images of the tissues of hepatocellular carcinoma patients are obtained.
[0064] S110: Segment the cancer tissue regions in each of the pathological images and randomly select them as sub-images.
[0065] It can be understood that the regions containing tissue sections in the pathological images are used as cancer tissue regions, and the remaining regions are used as background regions. The cancer tissue regions are segmented from the pathological images as sub-images.
[0066] Further, S110 may include S111 to S113:
[0067] S111: Use the image binarization segmentation threshold algorithm to segment multiple cancer tissue regions from each of the pathological images according to a preset threshold.
[0068] Exemplarily, in this embodiment, the OTSU algorithm may be used as the image binarization segmentation threshold algorithm to segment the pathological images.
[0069] S112: Segment each of the cancer tissue regions into multiple sub-images of a preset size.
[0070] For unified processing of image standards, in this embodiment, the cancer tissue regions may be segmented into multiple sub-images of a preset size.
[0071] S113: Randomly select multiple representative sub-images for each of the pathological images.
[0072] Exemplarily, manually review each sub-image and exclude sub-images with poor image quality (contaminated, blurred, or blank area exceeding 50%). More specifically, randomly select 10 sub-images from each pathological image for subsequent analysis.
[0073] S120: Extract multiple features of each of the sub-images and screen out the optimal feature subset from each of the features.
[0074] Specifically, the pathological images of tissues contain rich tumor genetic information and can provide a large number of pathological features, such as tumor tissue structure, cell morphology, tumor microenvironment, immune cell infiltration, tumor differentiation degree, microvascular invasion degree, etc. In this embodiment, the above features may be extracted from the pathological images, and the extracted features are screened, and then each of the screened features is used as the optimal feature subset.
[0075] Further, extracting multiple features of each of the sub-images includes the following steps:
[0076] Extract the original features and high-order features of each of the sub-images as the features; wherein, the original features include first-order statistics, gray-level co-occurrence matrix, gray-level run length matrix, gray-level size zone, adjacent gray-level difference matrix, gray-level dependence relationship matrix; the high-order features are obtained by wavelet transformation of the original features.
[0077] Specifically, first-order statistics: Describe the distribution of voxel intensities within an image region defined by a mask through common and basic metrics (18, the standard deviation is the same as the variance, take the variance).
[0078] Gray Level Co-occurrence Matrix (24): Describes the second-order joint probability function of the image area constrained by the mask.
[0079] Gray Level Run Length Matrix (16): Quantifies the gray level runs, which are defined as the number of pixel lengths of consecutive pixels with the same gray level value.
[0080] Gray Level Size Zone Matrix (16): Quantifies the gray level zones in the image.
[0081] Neigbouring Gray Tone Difference Matrix (5): Quantifies the difference between the gray level value within a distance δ and the average gray level value of its adjacent area.
[0082] Gray Level Dependence Matrix (14): Quantifies the gray level dependence relationships in the image. The gray level dependence relationship is defined as the number of connected voxels that depend on the central voxel within a distance δ.
[0083] In addition, this embodiment can also extract high-order features of the pathological image. The high-order features are obtained from the original features through wavelet transformation. LL, LH, HL, and HH are four transformation methods, namely Wavelet(LL, LH, HL, HH). Finally, 93 (Original features) + 93 * 4 (High-order features) = 465 features are obtained.
[0084] Furthermore, an optimal feature subset is screened from each of the said features, including S121 to S125:
[0085] S121: Calculate the mean information gain between each of the said features as the feature class correlation coefficient;
[0086] S122: Calculate the sum of mutual information between each of the said features, and then divide the sum of mutual information by the square of the total number of features to obtain the feature redundancy coefficient;
[0087] S123: Screen a preset number of features from each of the said features as a candidate feature subset according to the feature class correlation coefficient and the feature redundancy coefficient;
[0088] S124: Calculate the importance scores of each of the said features in the candidate feature subset;
[0089] S125: Remove several of the features with relatively low importance score rankings, and the remaining features in the candidate feature subset are used as the optimal feature subset.
[0090] It can be understood that in this embodiment, the correlation and redundancy between each feature are first calculated, and then the candidate feature subset is preliminarily screened. Next, the importance scores of each feature in the candidate feature subset are calculated. The features with higher importance scores contribute more to the prediction of angiopoietin, so they are retained, while the features with lower importance scores can be removed.
[0091] S130: Build a pathomics model based on the optimal feature subset.
[0092] Specifically, the optimal feature subset may include features that are highly correlated with angiopoietin, and a pathomics model capable of predicting the expression probability of angiopoietin is built based on the optimal feature subset.
[0093] Furthermore, S130 may include S131 - S137:
[0094] S131: Divide a training set from the images of hepatocellular carcinoma patients obtained from a professional database, and extract the optimal feature subset from the training set;
[0095] S132: Build the pathomics model based on the optimal feature subset extracted from the training set;
[0096] S133: Use the pathomics model to calculate the expression probability of angiopoietin corresponding to each feature in the training set as the predicted value;
[0097] S134: Calculate the residual and gradient between the predicted value and the actual value;
[0098] S135: Build a weak learner according to the residual, fit the weak learner with the residual, and train the newly added weak learner along the negative direction of the gradient;
[0099] S136: Add the weak learner to the pathomics model as the new pathomics model;
[0100] S137: Return the expression probability of angiopoietin corresponding to each feature in the training set calculated by the pathomics model as the predicted value until a preset condition is reached.
[0101] Specifically, the preset conditions for stopping training the pathomics model may include reaching a preset number of times or the performance parameters of the pathomics model no longer improving significantly, that is, the improvement amplitude of the performance parameters is small or relatively stable.
[0102] S140: Calculate the angiopoietin expression probability of the tissue of the corresponding hepatocellular carcinoma patient according to the feature by using the pathological omics model, and detect the expression level of angiopoietin according to the angiopoietin expression probability.
[0103] Specifically, the higher the angiopoietin expression probability, the higher the probability of the presence of angiopoietin in the tissue of the hepatocellular carcinoma patient. Therefore, the expression level of angiopoietin can be detected based on the angiopoietin expression probability.
[0104] After detecting angiopoietin, the embodiment of the present application may further include the step of classifying the detection result, specifically including:
[0105] S150: Use the survminer package of R language to perform binary classification on each of the angiopoietin expression probabilities output by the pathological omics model.
[0106] Specifically, each angiopoietin expression probability obtained according to the pathological omics model of this embodiment is used as a predicted value, and each predicted value is subjected to binary classification based on the minimum P value calculated by using the survminer package of R language.
[0107] Next, the solution of the embodiment of the present application will be introduced and described in detail in combination with specific application examples.
[0108] Figure 2 It is an example flowchart of a method for detecting angiopoietin provided for this embodiment. Refer to Figure 2 for description. This embodiment may include Step 1 to Step 5:
[0109] Step 1: Obtain hematoxylin-eosin (HE) stained sections of the tumor tissue of hepatocellular carcinoma patients, providing a basis for subsequent digital processing and analysis.
[0110] Specifically, download the pathological images of 267 hepatocellular carcinoma patients from a professional database, and obtain the tissue section images of 91 hepatocellular carcinoma patients from the actual measured cases of the patients: pathological tissue sections fixed with formalin and embedded in paraffin, in svs format, and HE stained tissue pathological images with a maximum magnification of 20× or 40× (20× or 40× magnification).
[0111] Randomly divide the images downloaded from the professional database into a training set and a test set according to 7:3. Among them, 70% of the data in the training set is used to train the pathological omics model; the remaining 30% of the data is used as the test set. The baseline conditions of the patients in the training set and the test set are similar, and there is comparability between groups. The pathological images obtained from the actual measured cases of the patients are used as an external test set.
[0112] Step 2: Process and segment the obtained pathological images to obtain sub-images of each pathological image for subsequent feature extraction.
[0113] Specifically, the OTSU algorithm is used to obtain the cancer tissue region of the pathological image, that is, the tissue region. The OTSU algorithm, also known as the maximum inter-class variance method, is an image binary segmentation threshold algorithm that uses a threshold to divide the image into a background region and the tissue region required for research.
[0114] Segment the pathological tissue image into sub-images: Segment the 40× pathological image into multiple 1024×1024 pixel sub-images; segment the 20× pathological image into multiple 512×512 pixel sub-images and upsample them to 1024×1024 pixels.
[0115] Screen the sub-images: Optionally, manually review and exclude sub-images with poor image quality (contaminated, blurred, blank area exceeding 50%). Randomly select 10 sub-images from each pathological image for subsequent analysis.
[0116] Step 3: Standardize the sub-images and extract features, and use an algorithm to screen out the optimal feature subset.
[0117] Specifically, in this embodiment, pathological image feature extraction can be performed. Use the PyRadiomics open-source package to standardize the sub-images and extract 93 original features (including first-order features and second-order features), and extract high-order features (Wavelet(LL,LH,HL,HH)), for a total of 465 features. After extracting features from 10 sub-images of each patient's pathological image respectively, the average value is taken correspondingly as the pathological omics features of each sample for subsequent data analysis.
[0118] Among them, the types of Original features include:
[0119] First Order Statistics: Describe the distribution of voxel intensities within the image region defined by the mask through common and basic metrics (18, the standard deviation is the same as the variance, take the variance).
[0120] Gray Level Co-occurrence Matrix (24): Describe the second-order joint probability function of the image region constrained by the mask.
[0121] Gray Level Run Length Matrix (16): Quantifies gray level runs, which are defined as the number of pixels in a run of consecutive pixels with the same gray level value.
[0122] Gray Level Size Zone Matrix (16): Quantifies gray level zones in the image.
[0123] Neigbouring Gray Tone Difference Matrix (5): Quantifies the difference between the gray level value within a distance δ and the average gray level value of its neighbouring region.
[0124] Gray Level Dependence Matrix (14): Quantifies the gray level dependence relationships in the image. The gray level dependence relationship is defined as the number of connected voxels that depend on the central voxel within a distance δ.
[0125] The high-order features are obtained by wavelet transformation of the Original features. LL, LH, HL, and HH are four transformation methods, i.e., Wavelet(LL, LH, HL, HH). Finally, 93 (Original features) + 93 * 4 (high-order features) = 465 features are obtained.
[0126] The steps to select the optimal feature subset include: using an algorithm to select the optimal feature subset to help the pathomics model reduce overfitting to the training data. First, the mRMR (Maximum Relevance, Minimum Redundancy) algorithm is used to select the top 30 important features. The Maximum Relevance Minimum Redundancy algorithm mRMR (Maximum relevance relevance, minimum redundancy redundancy) selects features, taking into account not only the correlation between the features and the variable to be predicted, but also the correlation between the features. The metric used is Mutual Information. For the mRMR method, the correlation between the feature subset and the class is calculated by the average information gain of each feature with the class, and the redundancy between the features is calculated by summing the mutual information between the features and dividing by the square of the total number of features in the optimal feature subset.
[0127] Subsequently, the RFE (Recursive Feature Elimination) algorithm is used to filter the obtained 30 features. RFE (Recursive Feature Elimination) feature screening sorts the features before modeling, and the unimportant features are eliminated one by one. Its goal is to find the optimal feature subset that can be used to generate an accurate model. The pathological omics model is trained multiple times. Each time after training, n features with low importance scores are deleted, and then the new features are trained again to obtain the importance scores of the features again. Then, n features with low importance scores are deleted again until the optimal feature subset is obtained. Exemplarily, Figure 3 This is an example diagram of the optimal feature subset provided in this embodiment.
[0128] Step 4: Based on the optimal feature subset, construct a pathological omics model, and use machine learning algorithms to deeply analyze pathological features to improve the generalization ability and prediction accuracy of the pathological omics model.
[0129] Specifically, the features screened by the mRMR–RFE algorithm are used to construct a pathological omics model in the training set through the Gradient Boosting Machine algorithm.
[0130] The specific process of the Gradient Boosting Machine algorithm is as follows:
[0131] Step 1: Initialization - Set the initial prediction value as a constant (for example, in a regression problem, all prediction values are the target mean of the training data).
[0132] Step 2: Loop iteration - Calculate the residual or gradient between the prediction result of the current pathological omics model and the actual target value.
[0133] According to the residual, construct a new weak learner (such as a decision tree) to make it fit the residual, that is, update the parameters of the pathological omics model along the negative gradient direction.
[0134] Step 3: Update the pathological omics model - Add the new weak learner to the current pathological omics model to obtain a new pathological omics model.
[0135] Step 4: Adjust the learning rate (shrinkage) - To control the learning speed of each step and reduce overfitting.
[0136] Repeat the above steps until the predetermined number of iterations is reached or the performance parameters of the pathological omics model no longer improve significantly.
[0137] Step 5: Output the probability Pathomics score (PS) of predicting the expression level of angiopoietin 2 (ANGPT2) (mRNA level or protein level detected by immunohistochemistry), and group the pathological images to predict the ANGPT2 expression level.
[0138] Output of the predicted value: Using the finally trained pathomics model, generate a probability (0 - 100%) of high molecular expression according to the eigenvalue of different pathological images, that is, the pathomics score (PS). PS is the probability of the pathomics model predicting the ANGPT2 expression level, and its value ranges from 0 to 1. The closer the PS value is to 0, the lower the ANGPT2 expression level predicted by the pathomics model; on the contrary, the closer the PS value is to 1, the higher the ANGPT2 expression level predicted by the pathomics model.
[0139] Division of the predicted value: Calculate the pathomics score through the pathomics model, combine the pathomics score with the clinical data, and calculate the cutoff value of the Pathomics score, that is, the classification threshold, through the R package (mainly ggplot2 [version 3.3.6]) "survminer". Because the predicted level is different, the pathological images downloaded from the professional database are the mRNA level of ANGPT2, and the cutoff value of the predicted value Pathomics score of the GBM (Gradient Boosting Machines) model based on the minimum P value is 0.512; for the external test set, based on the immunohistochemical protein level of ANGPT2, the cutoff value of the predicted value Pathomics score of the GBM model is taken as 0.500 using the median. Divide the pathological images into a binary classification of high / low expression levels of angiopoietin according to the cutoff value.
[0140] Demonstration of the prediction performance of the pathomics model:
[0141] Establish and validate the pathomics model based on the pathological images of the professional database and actual measurement cases. The pathomics model shows high prediction accuracy in the training set, test set, and external test set.
[0142] Specific evaluation indicators for the performance of the pathological omics model include: (1) ROC curve: showing the true positive rate (sensitivity) and false positive rate (1 - specificity) of the pathological omics model at a set threshold. (2) Area under the ROC curve: measuring the ability of the pathological omics model to distinguish different efficacy prediction groups. The larger the ROC - AUC, the larger the area under the curve, the more convex the curve is towards the upper left corner, and the better the effect of the pathological omics model. The AUC value of a perfect classifier is 1, while the AUC value of a model without discrimination ability is close to 0.5. (3) Drawing a calibration curve to evaluate the calibration of the pathological omics model; quantifying the comprehensive performance of the imaging prediction model through the Brier score, and the smaller the value, the better the consistency of the prediction of the pathological omics model; (4) Decision curve (DCA), showing the clinical benefit of the pathological omics prediction model. Figure 4 The ROC curves, calibration curves, and decision curves for the training set, test set, and external test set. Figure 4 A - C in it: The prognostic performance of the model shown by the ROC in the training, test, and external test sets. The AUC value of this model is 0.811, the AUC of the test set is 0.726, and the AUC of the external test set is 0.710. Figure 4 D - F in it: The calibration curve shows that the predicted probability and the actual state of the model are consistent in the training, test, and external test sets. Figure 4 G - I in it: The DCA shows the clinical benefits of the prediction model in the training, test, and external test sets.
[0143] The above results not only prove that the pathological omics model has good prediction effects, but also has good consistency between the predicted probability and the true value of whether the gene is highly expressed, and has high clinical practicability and universality.
[0144] Beneficial effects of this embodiment:
[0145] (1) Precise prediction: It can comprehensively extract the tumor feature information contained in pathological images. The analysis results are not affected by individual difference factors such as age and gender, overcoming the limitation that existing methods fail to fully consider tumor heterogeneity, and at the same time improving the objectivity and accuracy of prediction. (2) Efficient image segmentation: In the process of pathological omics analysis, the OTSU algorithm is used to accurately segment the tumor tissue in pathological images, realizing the automatic delineation of the tumor area, laying a solid foundation for in - depth study of the morphological characteristics of tumors and predicting the survival period of patients. (3) Good generalization: The features screened by the mRMR–RFE algorithm are used to construct a pathological omics model in the training set through the GBM algorithm, demonstrating the versatility of the pathological omics model and its powerful feature extraction ability. Through external validation, the trained model can adapt to specific pathological image analysis tasks and show good performance on new data sets.
[0146] This embodiment can comprehensively extract the morphological features of tumors, taking into account tumor heterogeneity, while existing technologies may not fully consider this complexity of tumors. For the computational challenges brought by the high pixel count of pathological tissue images, this embodiment uses an algorithm to efficiently distinguish tumor and background regions, not only simplifying the analysis process but also significantly reducing the computational burden of high-resolution histopathological analysis. The pathomics model was trained on pathological images in a professional database and passed internal and external validations, demonstrating its broad application potential and high prediction accuracy in various clinical settings.
[0147] In addition to the pathological image processing solution of the embodiments of the present application, the following alternative solutions can also be considered to achieve the detection of angiopoietin: (1) Genomics analysis: It involves gene sequencing of tumor tissues of liver cancer patients and analyzing gene expression patterns to directly reflect the molecular expression level. This solution can directly reveal the molecular characteristics of tumors. (2) Radiomics technology: Using imaging technologies such as computed tomography or magnetic resonance imaging, combined with radiomics analysis, to extract features from imaging data and predict molecular expression. This solution can provide morphological information of tumors. (3) Multimodal data fusion: Combining pathomics data with radiomics and genomics data to improve the accuracy of the prediction model through ensemble learning or data fusion techniques. This solution may improve the prediction precision.
[0148] Although the above alternative solutions have their respective advantages, this embodiment provides a more direct and cost-effective solution through pathomics analysis, which is particularly suitable for detection scenarios lacking high-end imaging equipment or gene sequencing technologies. In addition, this embodiment shows higher prediction accuracy than existing technologies in internal and external tests and is easy to implement.
[0149] Referring to Figure 5 , the embodiments of the present application also provide an angiopoietin detection device that can implement the above angiopoietin detection method. The device includes:
[0150] An image acquisition unit for acquiring hematoxylin-eosin staining pathological images of tissues of different hepatocellular carcinoma patients;
[0151] An image segmentation unit for segmenting the cancer tissue regions in each of the pathological images and randomly selecting them as sub-images;
[0152] A feature processing unit for extracting multiple features of each of the sub-images and screening an optimal feature subset from each of the features;
[0153] A model construction unit for constructing a pathomics model based on the optimal feature subset;
[0154] An expression level detection unit is configured to calculate the angiopoietin expression probability of the tissue of a hepatocellular carcinoma patient corresponding to the features according to the pathological omics model, and detect the expression level of angiopoietin according to the angiopoietin expression probability.
[0155] It can be understood that the content in the above method embodiments is applicable to the device embodiments. The functions specifically implemented by the device embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those of the above method embodiments.
[0156] An embodiment of the present application further provides an electronic device. The electronic device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the above angiopoietin detection method is implemented. The electronic device can be any intelligent terminal including a tablet computer, an in-vehicle computer, etc.
[0157] It can be understood that the content in the above method embodiments is applicable to the device embodiments. The functions specifically implemented by the device embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those of the above method embodiments.
[0158] Please refer to Figure 6 , Figure 6 which schematically shows the hardware structure of an electronic device according to another embodiment. The electronic device includes:
[0159] A processor 601, which can be implemented in a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, etc., and is configured to execute relevant programs to implement the technical solutions provided by the embodiments of the present application;
[0160] A memory 602, which can be implemented in the form of a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM), etc. The memory 602 can store an operating system and other application programs. When implementing the technical solutions provided by the embodiments of the present specification through software or firmware, the relevant program codes are stored in the memory 602, and the processor 601 is called to execute the angiopoietin detection method of the embodiments of the present application;
[0161] An input / output interface 603, which is configured to implement information input and output;
[0162] A communication interface 604 for implementing communication interaction between this device and other devices, which can achieve communication through wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.);
[0163] A bus 605 for transmitting information between various components of the device (such as a processor 601, a memory 602, an input / output interface 603, and a communication interface 604);
[0164] Among them, the processor 601, the memory 602, the input / output interface 603, and the communication interface 604 are communicatively connected to each other inside the device through the bus 605.
[0165] The embodiment of the present application also provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the above-mentioned angiopoietin detection method is implemented.
[0166] It can be understood that the content in the above method embodiment is applicable to the present storage medium embodiment. The functions specifically implemented by the present storage medium embodiment are the same as those of the above method embodiment, and the beneficial effects achieved are also the same as those of the above method embodiment.
[0167] As a non-transitory computer-readable storage medium, the memory can be used to store non-transitory software programs and non-transitory computer-executable programs. In addition, the memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory may optionally include a memory remotely disposed relative to the processor, and these remote memories can be connected to the processor through a network. Examples of the above network include, but are not limited to, the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.
[0168] The embodiments described in the embodiments of the present application are for more clearly illustrating the technical solutions of the embodiments of the present application, and do not constitute a limitation on the technical solutions provided by the embodiments of the present application. Those skilled in the art know that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of the present application are equally applicable to similar technical problems.
[0169] Those skilled in the art can understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of the present application, and may include more or fewer steps than those shown, or combine certain steps, or different steps.
[0170] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0171] Those of ordinary skill in the art can understand that all or some of the steps in the methods disclosed above, and the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, and their appropriate combinations.
[0172] The terms "first", "second", "third", "fourth", etc. (if any) in the specification of this application and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of this application described here can be implemented in an order other than those illustrated or described here. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.
[0173] It should be understood that in this application, "at least one (item)" means one or more, and "multiple" means two or more. "And / or" is used to describe the association relationship of associated objects and indicates that three relationships can exist. For example, "A and / or B" can mean: only A exists, only B exists, and both A and B exist at the same time. Among them, A and B can be singular or plural. The character " / " generally means that the associated objects before and after are in an "or" relationship. "At least one (one) of the following" or a similar expression means any combination of these items, including any combination of single item (one) or plural items (ones). For example, at least one (one) of a, b, or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.
[0174] In several embodiments provided by the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the above division of units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of devices or units can be in electrical, mechanical or other forms.
[0175] The units described above as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0176] In addition, in each embodiment of the present application, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.
[0177] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in each embodiment of the present application. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store programs.
[0178] The preferred embodiments of the embodiments of the present application have been described above with reference to the accompanying drawings. However, this does not limit the scope of the rights of the embodiments of the present application. Any modifications, equivalent replacements, and improvements made by those skilled in the art without departing from the scope and essence of the embodiments of the present application shall fall within the scope of the rights of the embodiments of the present application.
Claims
1. A method for detecting angiopoietin, characterized in that, The method includes the following steps: Obtain hematoxylin-eosin staining pathological images of tissues from different hepatocellular carcinoma patients; Segment the cancer tissue regions in each of the pathological images, and randomly select them as sub-images; Extract multiple features of each of the sub-images, and screen out the optimal feature subset from each of the features; Construct a pathological omics model based on the optimal feature subset; Use the pathological omics model to calculate the angiopoietin expression probability of the corresponding hepatocellular carcinoma patient tissue according to the features, and detect the expression level of angiopoietin according to the angiopoietin expression probability; The extracting multiple features of each of the sub-images includes the following steps: Extract the original features and high-order features of each of the sub-images as the features; wherein, the original features include first-order statistics, gray-level co-occurrence matrix, gray-level run length matrix, gray-level size zone, adjacent gray-level difference matrix, gray-level dependence relationship matrix; the high-order features are obtained by wavelet transformation of the original features; The screening out the optimal feature subset from each of the features includes the following steps: Calculate the mean value of the information gain between each of the features as the feature category correlation coefficient; Calculate the sum of the mutual information between each of the features, and then divide the sum of the mutual information by the square of the total number of the features to obtain the feature redundancy coefficient; Screen out a preset number of features from each of the features as a candidate feature subset according to the feature category correlation coefficient and the feature redundancy coefficient; Calculate the importance scores of each of the features in the candidate feature subset; Remove several of the features with relatively low importance score rankings, and the remaining features in the candidate feature subset are used as the optimal feature subset.
2. The angiopoietin detection method according to claim 1, wherein The obtaining hematoxylin-eosin staining pathological images of tissues from different hepatocellular carcinoma patients includes the following steps: Obtain hematoxylin-eosin staining pathological images of tissues from different hepatocellular carcinoma patients fixed with formalin and embedded in paraffin.
3. The angiopoietin detection method according to claim 1, characterized in that, The segmenting the cancer tissue regions in each of the pathological images and randomly selecting them as sub-images includes the following steps: Use the image binary segmentation threshold algorithm to segment multiple cancer tissue regions from each of the pathological images according to a preset threshold; Divide each of the cancer tissue regions into multiple sub-images of a preset size; Randomly select multiple representative sub-images for each of the pathological images.
4. The angiopoietin detection method according to claim 1, characterized in that, The constructing a pathological omics model based on the optimal feature subset includes the following steps: Divide a training set from the hepatocellular carcinoma patient images obtained from a professional database, and extract the optimal feature subset from the training set; Construct the pathological omics model based on the optimal feature subset extracted from the training set; Use the pathological omics model to calculate the angiopoietin expression probability corresponding to each of the features in the training set as the predicted value; Calculate the residual and gradient between the predicted value and the actual value; Construct a weak learner according to the residual, fit the weak learner with the residual, and train the newly added weak learner along the negative direction of the gradient; Add the weak learner to the pathological omics model as the new pathological omics model; Return the angiopoietin expression probability corresponding to each of the features in the training set calculated using the pathological omics model as a prediction value until a preset condition is reached.
5. A method for detecting angiopoietin according to any one of claims 1 to 4, characterized in that, The method further includes the following steps: Use the survminer package in R language to perform binary classification on the angiopoietin expression probabilities output by the pathological omics model.
6. An angiopoietin detection device, characterized in that, The device includes: An image acquisition unit for acquiring hematoxylin-eosin staining pathological images of tissues of different hepatocellular carcinoma patients; An image segmentation unit for segmenting the cancer tissue regions in each of the pathological images and randomly selecting them as sub-images; A feature processing unit for extracting multiple features of each of the sub-images and screening out an optimal feature subset from each of the features; A model construction unit for constructing a pathological omics model based on the optimal feature subset; An expression level detection unit for calculating the angiopoietin expression probability of the hepatocellular carcinoma patient tissue corresponding to the feature using the pathological omics model and detecting the expression level of angiopoietin according to the angiopoietin expression probability; The extraction of multiple features of each of the sub-images includes the following steps: Extract the original features and high-order features of each of the sub-images as the features; wherein, the original features include first-order statistics, gray-level co-occurrence matrix, gray-level run length matrix, gray-level size zone, adjacent gray-level difference matrix, gray-level dependence relationship matrix; the high-order features are obtained by wavelet transformation of the original features; The screening of the optimal feature subset from each of the features includes the following steps: Calculate the mean value of the information gain between each of the features as the feature category correlation coefficient; Calculate the sum of the mutual information between each of the features, and then divide the sum of the mutual information by the square of the total number of features to obtain the feature redundancy coefficient; Screen out a preset number of features from each of the features as a candidate feature subset according to the feature category correlation coefficient and the feature redundancy coefficient; Calculate the importance scores of each of the features in the candidate feature subset; Remove several features with relatively low importance score rankings, and the remaining features in the candidate feature subset are used as the optimal feature subset.
7. An electronic device, characterized in that, The electronic device includes a memory and a processor, the memory stores a computer program, and when the processor executes the computer program, the method described in any one of claims 1 to 5 is implemented.
8. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, the method described in any one of claims 1 to 5 is implemented.