An image-based feature extraction and prognosis model establishment method and device

By using image-based feature extraction and prognostic models, the problem of insufficient quantitative analysis of microvessels in existing technologies has been solved, enabling the development of individualized treatment plans for glioma patients and improving the accuracy and efficiency of prognostic assessment.

CN115457069BActive Publication Date: 2025-12-30THE FIRST AFFILIATED HOSPITAL OF ARMY MEDICAL UNIV +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211128481.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-16
Publication Date
2025-12-30
Estimated Expiration
2042-09-16

AI Technical Summary

Technical Problem

Existing technologies for quantitative analysis of microvessels in gliomas neglect the structural characteristics of tumor microvessels, resulting in an inability to effectively assist in clinical assessment of patient treatment responsiveness and prognosis. Furthermore, existing methods are time-consuming, labor-intensive, and have low efficiency.

Method used

By collecting pathological tissue slice images, extracting the location and omics characteristics of microvessels, and combining deep learning models and penalty Cox proportional hazards models, a predictive model for patient prognostic risk is constructed to achieve a comprehensive description of microvascular characteristics and prognostic assessment.

Benefits of technology

It provides an objective, feasible, and highly reproducible method that can automatically process and analyze pathological images, comprehensively describe the spatial distribution, morphological phenotype, and nuclear characteristics of microvessels, and help develop individualized treatment plans.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115457069B_ABST
    Figure CN115457069B_ABST
Patent Text Reader

Abstract

The application provides an image-based feature extraction and prognosis model establishment method and device for judging the prognosis of glioma. Microvessels on a pathological section H&E staining digital image are segmented by a deep learning algorithm, the nuclei inside the microvessels are segmented by a watershed algorithm, and the features of the microvessels are calculated by a pathology method. The features related to the prognosis of patients are selected by a machine learning method, and a relationship model of the features and the actual survival of tumor patients is constructed. The application provides an automatic scheme for selecting and extracting key image regions of patients, and a feature set beneficial to the prognosis evaluation of patients is selected and combined by a machine learning method.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of electronic information, and in particular relates to a method and apparatus for feature extraction and prognostic model establishment based on images. Background Technology

[0002] Gliomas are tumors originating from glial cells in the brain and are the most common primary intracranial tumors. The World Health Organization (WHO) classification of central nervous system tumors divides gliomas into grades I-IV, with grades I and II being low-grade gliomas and grades III and IV being high-grade gliomas. In my country, the annual incidence of gliomas is 5-8 per 100,000, and the 5-year mortality rate is second only to pancreatic and lung cancer among all cancers. Diagnosis of gliomas requires obtaining specimens through tumor resection or biopsy for histological and molecular pathological examination to determine the pathological grade and molecular subtype. Treatment for gliomas primarily involves surgical resection combined with radiotherapy, chemotherapy, and other comprehensive treatment methods to alleviate clinical symptoms and prolong survival. Traditional pathological diagnosis mainly relies on the characteristics of the tumor parenchyma (i.e., tumor cells) to determine the tumor type and grade. Currently, individualized treatment and clinical prognostic assessment of gliomas are mainly based on tumor histological classification and molecular subtyping. Key molecular pathological markers include isocitrate dehydrogenase (IDH) mutations, chromosomal 1p / 19q co-deletion, and telomerase reverse transcriptase (TERT) promoter mutations. However, tumors of the same histological type and molecular subtyping exhibit varying responses to treatment, resulting in heterogeneous prognoses, and a lack of indicators to guide individualized treatment plans.

[0003] Abundant angiogenesis is an important morphological feature of malignant tumors and a basis for tumor evolution. Microvessels in the tumor microenvironment are regulated by multiple molecular signaling pathways, exhibiting differences in vascular architecture phenotypes across different regions, times, and spaces within the tumor—a phenomenon known as vascular architecture heterogeneity. The significance of tumor microvessels in tumor grading, prognosis, and treatment strategy selection is widely recognized. Due to the unique structure of the blood-brain barrier, cerebral vessels significantly affect the penetration of chemotherapeutic drugs, thus influencing treatment efficacy. Multiple teams both domestically and internationally have been exploring parameters and standardized implementation methods for quantitatively reflecting tumor microvascular characteristics. For example, the Mayo Clinic in the United States quantitatively evaluated tumor microvessels in 24 patients with recurrent high-grade gliomas. Using CD31 immunohistochemical staining to label blood vessels, they obtained microvessel area (MVA) and microvessel density (MVD) values, finding that MVA was correlated with patient prognosis, while MVD was not. The Humanitas Clinical Research Institute applied fractal geometry to quantify brain tumor microvessels, finding that fractal dimension (FD) was superior to the traditional indicator MVD in reflecting vascular characteristics. The aforementioned research methods provide a methodological basis for quantitative analysis of tumor microvessels. Current quantitative analysis of microvessels simplifies the state of vessels to single parameters such as counting or fractal dimension, losing a significant amount of morphological and distributional characteristics. Furthermore, counting tiny targets and manually selecting vascular hotspots are time-consuming and labor-intensive, increasing costs and workload. Some studies have used deep learning techniques to segment tumor microvessels, but the segmentation efficiency is low and cannot reach a practical level. Moreover, these methods neglect the architectural characteristics of tumor microvessels, limiting the description of microvessels to the counting level, which is insufficient to assist clinical assessment of patient responsiveness to treatment and prognosis. Summary of the Invention

[0004] To address the aforementioned technical problems, this invention provides a method and apparatus for image-based feature extraction and prognostic model establishment, the technical solution of which is as follows:

[0005] A method for image-based feature extraction and prognostic model building includes the following steps:

[0006] Step 1: Collect diagnostic pathological tissue sections H&E stained images, clinical information, prognostic information, and molecular subtyping information from tumor patients; use the diagnostic pathological tissue sections H&E stained images as images to be segmented;

[0007] Step 2: Obtain the location information of microvessels of multiple morphological types on the image to be segmented;

[0008] Step 3: Obtain the location information of cell nuclei within microvessels in the image to be segmented;

[0009] Step 4: Extract the segmented microvascular-related omics feature data, that is, analyze the microvessels and the cell nuclei within the microvessels separately, including spatial distribution, morphology, color and texture features;

[0010] Step 5: Combine microvascular omics feature data from multiple diagnostic H&E stained tissue slice images of tumor patients as patient-dimensional microvascular feature data to construct a patient-dimensional feature set;

[0011] Step 6: Integrate microvascular omics feature data with patient clinical and molecular subtyping information, input the penalty Cox proportional hazards model, and calculate the feature subsets that effectively predict the prognostic risk combination of patients and the risk coefficient corresponding to each feature through cross-validation and gridded search methods, and construct a predictive model for patient prognostic risk.

[0012] Furthermore, in step 1, the image to be segmented is a panoramic image, a preprocessed pathological image, or a part of the original pathological image.

[0013] Furthermore, in step 2, a trained deep learning model is applied to extract microvessels of four different morphological types from the image to be segmented, thereby obtaining a mask of microvessels corresponding to the slice size.

[0014] Furthermore, in step 3, the image of the tissue region where the microvessels are located is decomposed into H channel, E channel and residual channel; the watershed algorithm is used to calculate the H channel image to extract the outline of the cell nucleus.

[0015] Furthermore, in step 4, the omics feature data specifically includes: microvascular spatial distribution feature set, microvascular morphological feature set, microvascular local distribution pattern feature, and cell nuclear feature set.

[0016] Furthermore, in step 5, when a patient has multiple diagnostic images, a representative image of the patient is selected for prognostic analysis based on the extracted omics feature data. The selection methods include selecting the image with the most statistically significant regions, the image with the highest proportion of microvessels in a certain subtype, and the image with the highest cell nucleus aggregation index. The methods for generating patient dimensional feature data include histogram algorithms and descriptive parameter calculation methods.

[0017] Furthermore, in step 6, the penalty Cox proportional hazards model is as follows:

[0018]

[0019] Where argmax represents the parameter β that maximizes the function, γ is the weight ratio of L1 and L2, PL(β) refers to the partial likelihood function of the Cox model, and β1, β2, ..., β... pThese are the coefficients of p features, and α≥0 is a hyperparameter for adjusting the shrinkage rate;

[0020] By integrating patients' clinical information, molecular subtyping information, and microvascular-related features at the patient level, an input feature dataset is formed. This feature dataset is then input into a penalized Cox proportional hazards model. Each input feature is standardized to allow direct comparison of the coefficients of each feature. A gridded search method is used to determine the range and value of α, and the maximum number of iterations for fitting each α to the penalized Cox proportional hazards model is set to 100. Furthermore, using the average consistency index as an evaluation metric, five-fold cross-validation is performed on each α to determine the α value with the best generalization and its corresponding feature subset. The risk coefficient corresponding to each feature is calculated, and the feature combination with non-zero coefficients constitutes the predictive model for patient prognostic risk.

[0021] The present invention also provides an apparatus for implementing an image-based feature extraction and prognostic model building method, comprising:

[0022] The first acquisition unit is used to acquire the location information of a first region, i.e., an object of interest, in an H&E-stained image of a diagnostic pathological tissue section; the object of interest includes microvessels.

[0023] The second acquisition unit is used to acquire the location information of the cell nuclei inside the second region of the diagnostic pathological tissue section H&E staining image, i.e., the object of interest.

[0024] The calculation unit is used to calculate parameters based on the location information of the first region, the location information of the second region, and the information of the corresponding H&E staining image; the parameters include: microvessel spatial distribution feature set, microvessel morphological feature set, microvessel local distribution pattern feature, and cell nucleus feature set;

[0025] A determination unit is used to determine a selected subset of features and corresponding risk parameters for a tumor prognosis model based on feature values ​​and a preset model; the feature values ​​include at least the image-based microvascular omics features and clinical information and molecular subtyping information associated with the patient; the preset model is a penalty Cox proportional hazards model.

[0026] The present invention also provides an electronic device, including a processor and a memory for storing executable instructions of the processor; the processor is configured to execute the executable instructions to implement the image-based feature extraction and prognostic model building method.

[0027] The present invention also provides a storage medium comprising a stored program, wherein, when the program is executed, it controls the device where the storage medium is located to execute the image-based feature extraction and prognostic model establishment method.

[0028] Beneficial effects:

[0029] The purpose of this invention is to provide an objective, feasible, and highly reproducible method and apparatus for feature extraction and prognostic assessment based on tumor microvessels. This invention performs global-to-local data processing and analysis on tissue slice images, comprehensively describing the spatial distribution patterns, morphological phenotypes, and cell nucleus texture and color features of the tissue of interest, significantly expanding the methods for quantitative research on regions of interest (ROIs) in pathological images. This invention provides an automated scheme for selecting key image regions for patients, using machine learning algorithms to select and combine feature sets beneficial for patient prognostic assessment. Risk stratification of tumor patients based on the assessment system can help develop different management strategies and treatment plans. Attached Figure Description

[0030] Figure 1 A schematic diagram illustrating an image-based feature extraction and prognostic model establishment method provided in an embodiment of this application;

[0031] Figure 2 This is a schematic diagram of the structure of an apparatus for implementing an image-based feature extraction and prognostic model establishment method, as provided in an embodiment of this application. Detailed Implementation

[0032] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention. Furthermore, the technical features involved in the various embodiments of this invention described below can be combined with each other as long as they do not conflict with each other.

[0033] like Figure 1 As shown, according to one aspect of the present invention, the present invention provides a method for image-based feature extraction and prognostic model establishment, specifically including the following steps:

[0034] Step 1: Collect H&E stained images of pathological tissue sections, clinical information, prognostic information, and molecular subtyping information of tumor patients. The images to be segmented are all diagnostic H&E stained images of tissue sections from the patient. These images can be brain pathological images or pathological images of other organs or tissues; this embodiment of the invention does not specifically limit this. This embodiment of the invention also does not limit the specific form of the images to be segmented; they can be whole-slide images, preprocessed pathological images, or a portion of the original pathological image.

[0035] Step 2: Obtain the location information of microvessels of multiple morphological types on the image to be segmented;

[0036] By applying a trained deep learning model, microvessels of four different morphological types are extracted from the image to be segmented, and a mask of microvessels corresponding to the size of the slice image is obtained.

[0037] Step 3: Obtain the location information of cell nuclei within microvessels in the image to be segmented, including: decomposing the tissue region image containing microvessels into H channel, E channel and residual channel; using the watershed algorithm to calculate the H channel image and extract the contour of the cell nuclei.

[0038] The watershed algorithm is an automatic image region segmentation method, a mathematical morphology segmentation method based on topological theory. Its basic idea is to view an image as a topographical landform in geodesy, where the gray value of each pixel represents its altitude, each local minimum and its affected area are called a catchment basin, and the boundary of the catchment basin forms a watershed. During segmentation, it uses the similarity between neighboring pixels as a reference, connecting pixels that are spatially close and have similar gray values ​​to form a closed contour. This nucleus extraction algorithm can also use thresholding or the Canny contour detection algorithm, etc.

[0039] Step 4: Extract the segmented microvessel-related omics feature data, that is, analyze the microvessels and the cell nuclei (including endothelial cells, pericytes, and tumor cells) within the microvessels separately, including features such as spatial distribution, morphology, color, and texture; the omics feature data specifically includes:

[0040] (1) Microvascular spatial distribution feature set, including the spatial distribution heterogeneity of microvessels and the location and size of hotspots, coldspots, diamond spots, and doughnut spots, as well as quantitative features based on each specific spatial region; including:

[0041] Quantifying the heterogeneity of microvessel distribution: Based on the coordinates of the tissue H&E stained image, the image is divided into square slices with a length of 330.25 μm. Slices with an area less than 50% of the square are removed, and the microvessel area within each slice is calculated. This invention uses the global Moran'I parameter (spatial autocorrelation) to quantify the spatial distribution pattern of microvessels in tissue slices, i.e., clustered, random, or diffuse distribution, which can reflect the heterogeneity of microvessel distribution. Spatial autocorrelation is characterized by the correlation of signals between nearby locations in space; the signal is the microvessel area within each slice. Spatial autocorrelation is multidimensional and multidirectional, and can be expressed as:

[0042]

[0043] Among them, Zi W is the deviation of the value of signal i from its average value. i,j Let represent the spatial weights between signals i and j, n represent the total number of features, and S0 be the sum of all spatial weights. A Moran'I parameter of 0 indicates that the signal is randomly distributed; the closer the parameter is to 1, the higher the signal's clustering; and the closer the parameter is to -1, the more the signal tends to have a diffuse distribution.

[0044] Identifying Special Regions of Vascular Distribution: Local Moran's I parameters were used to analyze vascular area, identifying statistically significant vascular hotspots, cold spots, diamond zones, and sandwich zones on tissue sections. This quantifies the signal strength of these regions relative to their surrounding areas. The signal values ​​were categorized into high (H) and low (L) values ​​based on the median signal. Therefore, each section, combined with the signals from its neighboring sections, can be classified as HH, LL, HL, or LH. The signal values ​​of hotspots, cold spots, diamond zones, and sandwich zones were significant in the overall image signal distribution (p<0.01), and the combined signals from neighboring sections belonged to HH, LL, HL, and LH, respectively. These special regions indicate areas of extensive microvascular proliferation, areas of widespread microvascular necrosis, areas of localized microvascular proliferation near necrotic areas, and areas of localized necrosis, respectively. This helps to objectively and accurately describe the local state of the tissue. Quantitative parameters such as the size, number, adjacency, and microvascular classification ratio of these special regions constitute a microvascular spatial distribution feature set.

[0045] (2) Microvascular morphological feature set, including:

[0046] Quantitative vascular parameters: These are quantitative parameters used to evaluate the degree of tumor microvascular proliferation. The main quantitative parameters include: microvessel density (MVD), microvessel area (MVA), and microvessel fractal dimension (mvFD), which represent the number, area, and fractal dimension of microvessels in tumor tissue, respectively.

[0047] The fractal dimension mvFD is calculated using a box-counting algorithm, as shown in the following formula: Where mvFD is the fractal dimension, ε represents the box side length, and N(ε) represents the minimum number of consecutive, non-overlapping boxes required to cover a blood vessel. To achieve the goal of calculating mvFD on H&E-stained images, this invention provides a blood vessel lumen extraction algorithm. The given image is converted from RGB to HSV color space, and the color range of red blood cells is defined as (0, 70, 50) to (12, 255, 255) and (170, 100, 50) to (180, 255, 255). The region containing red blood cells is extracted and set to white. The image is further converted to grayscale, and the region containing the blood vessel lumen is extracted using the Otsu binarization algorithm. False positives are removed through morphological post-processing. The contour of the microvascular lumen mask is further extracted to perform mvFD calculation based on the H&E-stained image.

[0048] Based on each specific spatial region, the quantitative parameters of each region are calculated: local microvessel density, local microvessel area, and local microvessel fractal dimension (local-MVD, local-MVA, local-mvFD).

[0049] Perimeter Ratio (PA Ratio) 2 The area ratio (ARR) is defined as the ratio of the square of an object's perimeter to its area. It is often used to quantify the compactness of two-dimensional objects. In this project, it is used to characterize the irregularity of blood vessel shapes.

[0050] Morphological features of blood vessel cavities: By applying the blood vessel cavity extraction method described above, the red blood cell region and blank region of the microvascular image can be extracted, and the morphological features of the blood vessel cavity can be analyzed. The feature set includes the number of cavities in a single microvessel, the ratio of cavity area to microvessel area, and the ratio of cavity perimeter to area, which reflect properties such as the degree of blood vessel aggregation and the regularity of blood vessel cavity shape, and constitute the morphological feature set of blood vessel cavities.

[0051] (3) Local distribution pattern feature set of microvessels: The tumor microenvironment surrounds microvessels in various organizational modes, including perivascular niche, hypoxic niche, and invasive niche, which are manifested by significant differences in the morphological classification and distribution of microvessels in different local areas. This invention develops a tile-based mask clustering algorithm based on tissue image slices. The tissue image is sliced, and each slice is statistically classified into five patterns: low microvessel density, extensive proliferation of multilayered microvessels, abundant thin-walled microvessels, coexistence of multiple types of microvessels, and moderate density of thin-walled microvessels. The statistical results of slices with different patterns on the image are used as the characteristics of the local tumor microenvironment.

[0052] (4) Cell nuclear feature set: mainly includes: Fourier Shape Descriptors, Global Cell Graph Features, Intensity Features, and Haralick Texture Features, which describe the shape, texture, color, and aggregation degree of the cell nucleus, reflecting information on cell proliferation, basophilia, and nuclear atypia; the cell nuclear aggregation index reflects the multilayering and proliferation degree of cell nuclei related to microvessels; the shortest distance from the cell nucleus to the blood vessel boundary reflects the basement membrane thickness of the blood vessel, i.e., whether there is a phenomenon of basement membrane thickening.

[0053] The Clustering Index is a measurement algorithm based on Ripley's K function to measure the spatial distribution characteristics of points. In this invention, it can describe the multi-layered features of microvessels, as shown in the following formula:

[0054]

[0055] Where CI(τ) represents the nuclear aggregation index within the search radius of τ, K represents the number of cell nuclei within the blood vessel, and d a,b Ω represents the Euclidean distance between cell nuclei a and b. a Indicates the distance d from the cell nucleus a. a,b The algorithm can effectively calculate the number of cell nuclei within a specific distance, such as 50 μm, from each vascular cell nucleus by finding the number of cell nuclei b that are less than or equal to the search radius τ.

[0056] Distance to microvessel boundary: This is used to calculate the shortest distance between each cell nucleus within a blood vessel and the vessel boundary, and can represent the basement membrane thickness characteristics of microvessels.

[0057] Step 5: Combine slide-level omics feature data in one manner as patient-level feature data; including:

[0058] When a patient has multiple diagnostic images, representative images are selected for prognostic analysis based on extracted omics feature data. Selection methods include images with the most statistically significant regions, images with the highest proportion of microvessels in a particular subtype, and images with the highest nuclear aggregation index. This invention does not limit the method for selecting representative images.

[0059] The methods for generating patient dimension feature data include: (1) using the histogram method for the feature data corresponding to the selected image. In order to construct the histogram, the first step is to segment the range of values, that is, to divide the entire range of feature values ​​into a series of intervals, and then calculate how many values ​​are in each interval and calculate the frequency within the interval; (2) calculating the distribution curve and descriptive parameters of each feature, including skewness, kurtosis, median, mean, standard deviation, 25th percentile, and 75th percentile.

[0060] Step 6: Input the penalized Cox Proportional Hazard Model, and use cross-validation and grid search methods to calculate the feature subset of the prognostic risk combination that effectively predicts the patient's prognostic risk and the risk coefficient corresponding to each feature, and construct a predictive model for the patient's prognostic risk.

[0061] Cox proportional hazards models are commonly used for predicting prognostic risk. Their parameters can be understood as hazard ratios (HR), helping clinicians assess a patient's risk factors. However, when there are many input features and strong correlations between them, the non-singularity of the feature matrix prevents the Cox proportional hazards model from calculating. Lasso penalty can effectively filter feature subsets, but it has two drawbacks: 1. It cannot handle high-dimensional data and cannot select feature subsets exceeding the number of training samples; 2. For highly correlated feature groups, Lasso will randomly select features from one group and discard features from other groups. In statistics, especially when fitting linear or logistic regression models, ElasticNet is a regularized regression method that linearly combines the L1 and L2 norms of Lasso and Ridge methods, forming a single model with two penalty factors: one proportional to the L1 norm and the other proportional to the L2 norm. The model obtained using this method is sparse, similar to pure lasso regression, but possesses the same regularization capability as ridge regression. The function to be solved using this method is:

[0062]

[0063] Where argmax represents the parameter β that maximizes the function, γ is the weight ratio of L1 and L2, PL(β) refers to the partial likelihood function of the Cox model, and β1, β2, ..., β... p These are the coefficients of the p features, and α≥0 is a hyperparameter for adjusting the shrinkage rate.

[0064] By inputting patients' clinical indicators, molecular subtypes, and microvascular-related features into a penalized Cox proportional hazards model, each feature is standardized to allow direct comparison of feature coefficients. A gridded search method is used to determine the range and value of α, and the maximum number of iterations for fitting each α to the model is set to 100. Furthermore, using the average concordance index as an evaluation metric, five-fold cross-validation is performed on each α to determine the α value with the best generalization and its corresponding feature subset. The risk coefficient corresponding to each feature is calculated, and the combination of features with non-zero coefficients constitutes the predictive model for patient prognostic risk.

[0065] like Figure 2 As shown, the apparatus for implementing the image-based feature extraction and prognostic model establishment method of the present invention includes:

[0066] The first acquisition unit 100 is used to acquire the location information of a first region in the image, i.e., an object of interest; the object of interest includes microvessels.

[0067] The second acquisition unit 200 is used to acquire the location information of the cell nucleus inside the second region of the image, i.e., the object of interest.

[0068] The calculation unit 300 is used to calculate parameters based on the location information of the first region, the location information of the second region, and the information of the image. The parameters include: the spatial distribution pattern of the object, quantization parameters of hot and cold regions, morphological features of the object, and characteristics of the shape, texture, color, and aggregation degree of cell nuclei.

[0069] The determination unit 400 is used to determine the selected feature combination and corresponding risk parameters of the tumor prognosis model based on feature values ​​and a preset model; the feature values ​​include at least the image-based microvascular omics features and the clinical information and molecular subtyping information associated with the patient; the preset model is a penalty Cox proportional hazards model.

[0070] The present invention also includes an electronic device comprising a processor and a memory for storing processor-executable instructions; the processor is configured to execute the executable instructions to implement the feature extraction and prognostic model building method of the present invention.

[0071] The present invention also includes a storage medium comprising a stored program, wherein, when the program is executed, the device on which the storage medium is located executes the feature extraction and prognostic model building method of the present invention.

[0072] Those skilled in the art will readily understand that the above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. An image-based feature extraction and prognostic model building method, characterized by, The method comprises the following steps: Step 1, collecting diagnostic pathological tissue section H&E staining images of tumor patients, clinical information, prognosis information and molecular typing information; the diagnostic pathological tissue section H&E staining image is used as a to-be-segmented image; Step 2, obtaining position information of a plurality of morphological typing microvessels on the to-be-segmented image; Step 3, obtaining position information of cell nuclei in the microvessels on the to-be-segmented image; Step 4, extracting segmented microvessel-related omics feature data, that is, analyzing the microvessels and the cell nuclei in the microvessels respectively, including spatial distribution, morphology, color and texture features; Step 5, combining the microvessel omics feature data on a plurality of diagnostic H&E stained tissue section images of the tumor patient as patient-dimension microvessel feature data, and constructing a patient-dimension feature set; Step 6, integrating the microvessel omics feature data and the patient clinical and molecular typing information, inputting into a penalty Cox proportional risk model, calculating an effective prediction feature subset of the patient's prognosis risk and a risk coefficient corresponding to each feature through cross-validation and gridding search method, and constructing a prediction model of the patient's prognosis risk; The penalty Cox proportional risk model is: ; where argmax denotes the argument of the function that makes the function get the maximum value, γ is the weight ratio of L1 norm and L2 norm, PL(β) is the partial likelihood function of the penalty Cox proportional risk model, β1, β2…β p are the coefficients of p features, and α≥0 is a hyperparameter that adjusts the shrinkage rate.

2. The image-based feature extraction and prognostic model building method of claim 1, wherein, In the step 1, the to-be-segmented image is a panoramic image, a preprocessed pathological image or a part of an original pathological image.

3. The method of claim 1, wherein, In the step 2, a trained deep learning model is applied to extract microvessels of four different morphological typing on the to-be-segmented image, and a mask corresponding to the size of the section is obtained.

4. The image-based feature extraction and prognostic model building method of claim 2, wherein, In the step 3, the tissue region image where the microvessel is located is decomposed into an H channel, an E channel and a residual channel; the H channel image is calculated using a watershed algorithm to extract the outline of the cell nucleus.

5. The image-based feature extraction and prognostic model building method of claim 3, wherein, In the step 4, the omics feature data specifically includes: a microvessel spatial distribution feature set, a microvessel morphology feature set, a microvessel local distribution pattern feature, and a cell nucleus feature set.

6. The image-based feature extraction and prognostic model building method of claim 4, wherein, In the step 5, in the case that the patient has a plurality of diagnostic images, a representative image of the patient is selected for prognosis analysis according to the extracted omics feature data, and the selection method includes selecting an image with the most statistically significant regions, an image with the highest proportion of a certain typing microvessel, and an image with the highest cell nucleus aggregation index; the generation method of the patient-dimension feature data includes a histogram algorithm and a descriptive parameter calculation method.

7. The image-based feature extraction and prognostic model building method of claim 5, wherein, In the step 6, the input feature data set is formed by integrating the clinical information, molecular typing information and patient-dimension microvessel-related features of the patient; the feature data set is input into the penalty Cox proportional risk model, each input feature is standardized to allow the coefficients of each feature to be directly compared; the gridding search method is used to determine the range and value of α, and the maximum number of iterations of the penalty Cox proportional risk model fitting each α is set to 100 times; Taking the average consistency index as the evaluation index, five-fold cross-validation is performed on each α to determine the α value with the best generalization and the corresponding feature subset, and the risk coefficient corresponding to each feature is calculated, and the combination of the features with non-zero coefficients is the prediction model of the patient's prognosis risk.

8. An apparatus for implementing the image-based feature extraction and prognosis model development method of any one of claims 1-7, characterized in that, It comprises: A first acquisition unit is configured to acquire position information of a first region, i.e., a position of an object of interest, in a diagnostic H&E staining image of a pathological tissue section; The object of interest includes microvessels; A second acquisition unit is configured to acquire position information of a second region, i.e., a position of a cell nucleus inside the object of interest, in the diagnostic H&E staining image of the pathological tissue section; A calculation unit is configured to calculate parameters according to the position information of the first region, the position information of the second region, and information of the corresponding H&E staining image; the parameters include: a microvessel spatial distribution feature set, a microvessel morphological feature set, a microvessel local distribution pattern feature, and a cell nucleus feature set; A determination unit is configured to determine a selected feature subset of a tumor prognosis model and a corresponding risk parameter based on feature values and a preset model; the feature values at least include the image-based microvessel omics features and clinical information and molecular typing information associated with the patient; and the preset model is a penalty Cox proportional hazards model; The penalty Cox proportional hazards model is: ; where argmax denotes the argument of the function that gives the maximum value of the function, γ is the weight ratio of L1 and L2, PL(β) is the partial likelihood function of the Cox model, β1, β2, …, β p are coefficients of p features, and α ≥ 0 is a hyperparameter that adjusts the shrinkage rate.

9. An electronic device, comprising: A device includes a processor and a memory for storing executable instructions of the processor; the processor is configured to execute the executable instructions to implement the image-based feature extraction and prognosis model establishment method of any one of claims 1-7.

10. A storage medium, characterized by The storage medium includes a stored program, wherein the program controls a device in which the storage medium is located to execute the image-based feature extraction and prognosis model establishment method of any one of claims 1-7 when the program is running.

Citation Information

Patent Citations

  • Method for establishing bleeding risk predicting model of acute coronary syndrome after interventional therapy

    CN110364261A

  • Biomarker prediction system, method and equipment based on Unet model

    CN114121226A