A method, device, equipment and medium for prognostic assessment of hepatocellular carcinoma patients
By performing image block cropping and cell nucleus segmentation on digital pathological slides of hepatocellular carcinoma patients, extracting multi-dimensional features, and combining cross-validation and Akaike Information Criterion screening, a prognostic model is trained. This solves the problems of time-consuming and low-accuracy prognostic assessment of hepatocellular carcinoma patients in existing technologies, and achieves efficient and accurate prognostic assessment.
Patent Information
- Application Number
- CN202411875096.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-19
- Publication Date
- 2025-10-31
- Estimated Expiration
- 2044-12-19
AI Technical Summary
In existing technologies, prognostic assessment methods for hepatocellular carcinoma patients are time-consuming and have low accuracy. Human assessment is easily affected by subjectivity and is difficult to effectively capture the sub-visual characteristics of tissue cells.
By acquiring digital pathological slides from hepatocellular carcinoma patients, image block cropping and cell nucleus segmentation were performed. Cell nucleus shape, texture, and interaction features were extracted. Features were screened using cross-validation and the Akaike Information Criterion, and a prognostic model was trained for evaluation.
It enables efficient and accurate prognostic assessment of hepatocellular carcinoma patients, improves the automation and accuracy of prognostic assessment, and reduces subjective errors in human assessment.
Smart Images

Figure CN119742066B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of prognostic analysis technology, and in particular, to a method, apparatus, device, and medium for prognostic assessment of patients with hepatocellular carcinoma. Background Technology
[0002] At the cellular and tissue level, histopathological sections contain rich characteristic information about tumors and their microenvironment. In clinical practice, pathologists observe histopathological sections under high-powered microscopes to qualitatively assess histopathological patterns and predict cancer behavior to some extent. However, manual assessment is time-consuming and easily affected by human subjectivity, resulting in low accuracy in assessing the prognostic performance of tissue cells. Summary of the Invention
[0003] This application provides a method, device, equipment, and medium for prognostic assessment of hepatocellular carcinoma patients. It can accurately assess the prognostic performance of hepatocellular carcinoma patients through a trained prognostic model, thereby improving the accuracy of prognostic assessment.
[0004] Other features and advantages of this application will become apparent from the following detailed description, or may be learned in part from practice of this application.
[0005] According to one aspect of the embodiments of this application, a method for prognostic assessment of hepatocellular carcinoma patients is provided, the method comprising:
[0006] Obtain the first target feature vector of at least one hepatocellular carcinoma patient;
[0007] Based on cross-validation and the Akaike Information Criterion, the first target feature vectors of each of the hepatocellular carcinoma patients are filtered to obtain target features;
[0008] The target features are used as training samples for a pre-defined prognostic model to train the prognostic model.
[0009] Obtain histopathological section data of the target hepatocellular carcinoma patient and input the histopathological section data into the trained prognostic model to obtain the prognostic assessment score of the target hepatocellular carcinoma patient;
[0010] The first target feature vector of a single hepatocellular carcinoma patient can be obtained through the following steps:
[0011] Obtain digital pathological sections from the hepatocellular carcinoma patient, wherein the digital pathological sections are tissue pathological sections of the hepatocellular carcinoma patient stained with hematoxylin and eosin;
[0012] The digital pathological slides are cropped into image blocks to obtain a first preset number of image blocks;
[0013] Each image block is segmented into cell nuclei to obtain a cell nucleus mask that corresponds one-to-one with each image block. The cell nucleus mask is image data that has the features of all cell nuclei in the image block.
[0014] Based on each of the cell nucleus masks, feature extraction is performed to obtain a second target feature vector, each of which has cell nucleus shape features, cell nucleus texture features, and cell nucleus interaction features. Each of the image blocks has a second target feature vector.
[0015] A second target feature vector of the target image patch is obtained to generate a feature matrix corresponding to the target image patch and the second target feature vector, and the feature matrix is concatenated to obtain the first target feature vector.
[0016] In one embodiment of this application, based on the foregoing scheme, a single second target feature vector can be obtained through the following steps:
[0017] Based on the cell nucleus mask, the shape features of each cell nucleus in the image block are extracted to obtain a first feature vector with a first dimension of information. The first feature vector is used to characterize the cell nucleus shape features of the image block.
[0018] Based on the cell nucleus mask, the texture features of each cell nucleus in the image block are extracted to obtain a second feature vector with second-dimensional information. The second feature vector is used to characterize the cell nucleus texture features of the image block.
[0019] The interaction features of each cell nucleus in the image patch are extracted based on the cell nucleus mask to obtain a third feature vector with third-dimensional information. The third feature vector is used to characterize the cell nucleus interaction features of the image patch.
[0020] The first feature vector, the second feature vector, and the third feature vector are concatenated laterally to obtain the second target feature vector, which has the sum of the first dimension information, the second dimension information, and the third dimension information.
[0021] In one embodiment of this application, based on the foregoing scheme, the first feature vector having first dimension information can be obtained through the following steps:
[0022] Based on the cell nucleus mask, the first target parameter feature of each cell nucleus is extracted, wherein the first target parameter feature is a feature vector with a first initial dimension information;
[0023] Perform a first preset statistical operation on each of the first target parameter features to obtain the second target parameter features of each of the cell nuclei. The second target parameter features are feature vectors with second initial dimension information.
[0024] The first target parameter feature and the second target parameter feature are concatenated laterally to obtain a first feature vector with the first dimension information.
[0025] Wherein, the first dimension information is the product of the first initial dimension information and the second initial dimension information.
[0026] In one embodiment of this application, based on the foregoing scheme, the second feature vector having second-dimensional information can be obtained through the following steps:
[0027] The third target parameter features of each cell nucleus are extracted based on the cell nucleus mask. The third target parameter features are feature vectors with third initial dimension information.
[0028] A second preset statistical operation is performed on each of the third target parameter features to obtain the fourth target parameter features of each of the cell nuclei. The fourth target parameter features are feature vectors with fourth initial dimension information.
[0029] The third target parameter feature and the fourth target parameter feature are concatenated laterally to obtain a second feature vector with the second dimension information;
[0030] The second dimension information is the product of the third initial dimension information and the fourth initial dimension information.
[0031] In one embodiment of this application, based on the foregoing scheme, the third feature vector with third-dimensional information can be obtained through the following steps:
[0032] The fifth target parameter features of each cell nucleus are extracted based on the cell nucleus mask, and the fifth target parameter features are feature vectors with fifth initial dimension information;
[0033] A third preset statistical operation is performed on each of the fifth target parameter features to obtain the sixth target parameter features of each of the cell nuclei. The sixth target parameter features are feature vectors with sixth initial dimension information.
[0034] The fifth target parameter feature and the sixth target parameter feature are concatenated laterally to obtain a third feature vector containing the third dimension information;
[0035] The third dimension information is obtained based on the fifth initial dimension information and the sixth initial dimension information.
[0036] In one embodiment of this application, based on the foregoing scheme, the second target feature vector is a feature vector having the sum of the first dimension information, the second dimension information, and the third dimension information. The step of concatenating the feature matrix to obtain the first target feature vector includes:
[0037] A third preset statistical operation is performed on each of the second target feature vectors in the feature matrix to obtain a third target feature vector with fourth dimension information;
[0038] The second target feature vector and the third target feature vector are concatenated to obtain the first target feature vector.
[0039] In one embodiment of this application, based on the foregoing scheme, the feature filtering of the first target feature vector of each of the hepatocellular carcinoma patients based on cross-validation and the Akaike Information Criterion to obtain target features includes:
[0040] The first target feature vectors of each hepatocellular carcinoma patient were subjected to 10-fold cross-validation in a preset round, and the target screening features obtained in each round were determined by the bidirectional selection stepwise regression method based on the Akaike information criterion. Each round corresponds to an Akaike information criterion value.
[0041] In each round, the target round with the smallest Akaike Information Criterion Value is selected, and the target screening features selected by the target round are obtained as the target features.
[0042] According to one aspect of the embodiments of this application, a prognostic assessment device for patients with hepatocellular carcinoma is provided, the device comprising:
[0043] The first acquisition unit is used to acquire the first target feature vector of at least one hepatocellular carcinoma patient;
[0044] The screening unit is used to perform feature screening on the first target feature vector of each of the hepatocellular carcinoma patients based on cross-validation and the Akaike Information Criterion to obtain target features;
[0045] The training unit is used to use the target features as training samples for a preset prognostic model to train the prognostic model.
[0046] The second acquisition unit is used to acquire histopathological section data of the target hepatocellular carcinoma patient and input the histopathological section data into the trained prognostic model to obtain the prognostic assessment score of the target hepatocellular carcinoma patient.
[0047] The first target feature vector of a single hepatocellular carcinoma patient can be obtained through the following steps:
[0048] Obtain digital pathological sections from the hepatocellular carcinoma patient, wherein the digital pathological sections are tissue pathological sections of the hepatocellular carcinoma patient stained with hematoxylin and eosin;
[0049] The digital pathological slides are cropped into image blocks to obtain a first preset number of image blocks;
[0050] Each image block is segmented into cell nuclei to obtain a cell nucleus mask that corresponds one-to-one with each image block. The cell nucleus mask is image data that has the features of all cell nuclei in the image block.
[0051] Based on each of the cell nucleus masks, feature extraction is performed to obtain a second target feature vector, each of which has cell nucleus shape features, cell nucleus texture features, and cell nucleus interaction features. Each of the image blocks has a second target feature vector.
[0052] A second target feature vector of the target image patch is obtained to generate a feature matrix corresponding to the target image patch and the second target feature vector, and the feature matrix is concatenated to obtain the first target feature vector.
[0053] According to one aspect of the embodiments of this application, a computer-readable storage medium is provided that stores a computer program thereon, the computer program including executable instructions that, when executed by a processor, implement the method described in the above embodiments.
[0054] According to one aspect of the embodiments of this application, an electronic device is provided, including: one or more processors; and a memory for storing executable instructions of the processors, which, when executed by the one or more processors, cause the one or more processors to perform the method as described in the above embodiments.
[0055] This application obtains a first target feature vector from at least one hepatocellular carcinoma patient. The process involves cropping image blocks and segmenting cell nuclei from digitized pathological slices of the hepatocellular carcinoma patient to obtain a cell nucleus mask. Analyzing the cell nucleus mask provides a clearer understanding of cell nucleus shape features, texture features, and interaction features. Each image block corresponds to a second target feature vector, which is a multi-dimensional feature vector capable of representing various cell nuclei features. Therefore, the first target feature vector obtained by concatenating the feature matrix generated from the second target feature vector can be used to represent the set of cell nuclei features of a single patient.
[0056] Furthermore, by using the first target feature vector to train the pre-set prognostic model with samples, the prognostic assessment of the prognostic model can be made more accurate. Only the histopathological slide data of the target hepatocellular carcinoma patient needs to be input into the trained prognostic model, and an accurate prognostic assessment score will be automatically output, achieving efficient and automated prognostic assessment. This solves the problems of inefficiency and low accuracy caused by manual assessment in the existing technology.
[0057] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and do not limit this application. Attached Figure Description
[0058] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application. It is obvious that the drawings described below are merely some embodiments of this application, and those skilled in the art can obtain other drawings based on these drawings without any inventive effort. In the drawings:
[0059] Figure 1 This is a flowchart illustrating a prognostic assessment method for patients with hepatocellular carcinoma according to an embodiment of this application;
[0060] Figure 2 This is a full logic diagram of the prognostic model shown according to an embodiment of this application;
[0061] Figure 3 This is a block diagram of a prognostic assessment device for hepatocellular carcinoma patients according to an embodiment of this application;
[0062] Figure 4 This is a schematic diagram of the system structure of an electronic device according to an embodiment of this application. Detailed Implementation
[0063] Exemplary embodiments will now be described more fully with reference to the accompanying drawings. However, these exemplary embodiments can be implemented in many forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided to make this application more comprehensive and complete, and to fully convey the concept of the exemplary embodiments to those skilled in the art.
[0064] Furthermore, the described features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. Numerous specific details are provided in the following description to give a thorough understanding of embodiments of this application. However, those skilled in the art will recognize that the technical solutions of this application can be practiced without one or more of the specific details, or other methods, components, apparatuses, steps, etc., can be employed. In other instances, well-known methods, apparatuses, implementations, or operations are not shown or described in detail to avoid obscuring various aspects of this application.
[0065] The block diagrams shown in the accompanying drawings are merely functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities can be implemented in software, in one or more hardware modules or integrated circuits, or in different network and / or processor devices and / or microcontroller node devices.
[0066] The flowcharts shown in the accompanying drawings are merely illustrative and do not necessarily include all content and operations / steps, nor do they necessarily have to be performed in the described order. For example, some operations / steps can be broken down, while others can be combined or partially combined; therefore, the actual execution order may change depending on the specific circumstances.
[0067] It should be noted that "multiple" in this article refers to two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, or B alone. The character " / " generally indicates that the preceding and following related objects have an "or" relationship.
[0068] The following is a detailed description of the background technology corresponding to the embodiments of this application:
[0069] At the cellular and tissue levels, histopathological sections contain rich characteristic information about tumors and their microenvironment. In clinical practice, pathologists observe H&E-stained sections under high-powered microscopes to qualitatively assess histopathological patterns and predict cancer behavior to some extent. However, manual assessment is time-consuming, subjective, and fails to capture sub-visual features. To address these issues, many prognostic models have been developed to quantify the prognostic status of patients with hepatocellular carcinoma after surgery. These include clinicopathological models based on clinical and pathological information, deep learning models based on deep neural networks, and manual feature models based on domain knowledge and mathematical statistics. Each of these prognostic models has its advantages and disadvantages in practical applications, as described below:
[0070] Clinicopathological models: Clinicopathological models are prognostic models built based on a patient's clinical and pathological information, such as age and gender. They primarily quantify the macroscopic and histological characteristics of hepatocellular carcinoma (HCC). Combining more clinical and pathological information can achieve a higher prognostic performance in postoperative HCC patients. However, more effective clinical and pathological information often means higher costs for detection technologies (such as molecular and genetic testing) and materials, which limits their widespread application in clinical practice. Furthermore, due to the heterogeneity of tumors and their microenvironment, patients with the same clinicopathological information can have significantly different postoperative prognoses, further limiting the development of clinicicopathological models in clinical applications.
[0071] Deep learning models: Deep learning models refer to prognostic models built based on deep neural networks. With the advancement of big data analytics and the increasing number of digitized pathological slides, deep learning has become the mainstream method for automated tissue image analysis. Deep convolutional neural networks, deep residual networks, and other deep learning models have indeed improved prognostic performance in diagnosis. However, due to the "black box" nature of deep learning, its decision-making process is difficult to interpret and understand, which makes it difficult for deep learning models to be widely accepted in clinical practice.
[0072] Handcrafted feature models: These are prognostic models developed based on domain knowledge and mathematical statistics. Indicators such as the shape, texture, cell density, and spatial distribution of the region of interest (ROI) provide clear outputs with a given input, exhibiting high interpretability and ease of understanding by physicians. In recent years, numerous studies have emerged constructing prognostic models based on deep features of tumors and their microenvironment extracted using handcrafted feature models. Examples include prognostic models for postoperative lung adenocarcinoma patients developed based on the texture features of tumor regions, and prognostic models for hepatocellular carcinoma patients developed based on the shape and texture features of tumors and their microenvironment. However, the prognostic efficacy of comprehensive nuclear features in postoperative hepatocellular carcinoma patients remains unknown.
[0073] Therefore, based on the problems existing in the above background technology, this application proposes a pre-defined prognostic model (multi-dimensional fusion cell nuclear feature model) for the prognosis of patients with hepatocellular carcinoma after surgery. It quantifies the prognostic status of patients from three perspectives: cell nuclear shape, texture and interaction, using quantitative features to improve the accuracy of prognostic assessment.
[0074] The implementation details of the technical solutions in the embodiments of this application are described in detail below:
[0075] According to one aspect of this application, a method for prognostic assessment of patients with hepatocellular carcinoma is provided. Figure 1The flowchart below illustrates a prognostic assessment method for hepatocellular carcinoma patients according to an embodiment of this application. This prognostic assessment method for hepatocellular carcinoma patients includes at least steps 110 to 140, detailed below:
[0076] In step 110, a first target feature vector of at least one hepatocellular carcinoma patient is obtained.
[0077] Specifically, the first target feature vector is a feature vector used to characterize the set of cell nuclear features in hepatocellular carcinoma patients. This first target feature vector is a feature vector with multiple dimensions of information, such as the area ratio of the cell nucleus to the cell nucleus mask, Haralick texture features, cell nucleus interaction features, and so on. To prevent overfitting caused by excessive dimensionality in the first target feature vector, cross-validation and the Akaike Information Criterion are used to filter the first target feature vectors of each hepatocellular carcinoma patient. This process yields target features that reduce overfitting and are more representative.
[0078] In one embodiment of this application, the first target feature vector of a single hepatocellular carcinoma patient can be obtained through the following steps:
[0079] Obtain digital pathological sections from the hepatocellular carcinoma patient, wherein the digital pathological sections are tissue pathological sections of the hepatocellular carcinoma patient stained with hematoxylin and eosin;
[0080] The digital pathological slides are cropped into image blocks to obtain a first preset number of image blocks;
[0081] Each image block is segmented into cell nuclei to obtain a cell nucleus mask that corresponds one-to-one with each image block. The cell nucleus mask is image data that has the features of all cell nuclei in the image block.
[0082] Based on each of the cell nucleus masks, feature extraction is performed to obtain a second target feature vector, each of which has cell nucleus shape features, cell nucleus texture features, and cell nucleus interaction features. Each of the image blocks has a second target feature vector.
[0083] A second target feature vector of the target image patch is obtained to generate a feature matrix corresponding to the target image patch and the second target feature vector, and the feature matrix is concatenated to obtain the first target feature vector.
[0084] Specifically, the digital pathological sections of hepatocellular carcinoma patients are histopathological sections of hepatocellular carcinoma patients stained with hematoxylin and eosin, which are image data that can be recognized by computers and other devices.
[0085] For reference Figure 2 As shown, Figure 2 As a schematic diagram of the complete processing flow, the digital pathological slide is cropped into image blocks to obtain a first preset number of image blocks. Specifically, this can be achieved by: determining the edges of the tissue region and segmenting the tissue region using the Otsu method (OTSU), and then generating a tissue region mask. Based on the tissue region mask, the tissue region in the digital pathological slide is cropped into 1000×1000 pixel non-overlapping PNG format image blocks. To obtain relatively dense tissue regions, image blocks with a tissue region ratio of less than 70% are filtered out. Subsequently, a pre-trained segmentation model, HoverNet, is used to segment these image blocks into cell nuclei, generating a cell nucleus mask corresponding to each image block. It should be noted that an image block can contain multiple cell nuclei, so the first preset number mentioned in this embodiment can be any value. In this application, the digital pathological slide can be cropped into 1000 image blocks with 1000×1000 pixels. In other embodiments, the first preset number can also be set to other values, and the pixel setting of the image block can also be 500×500. Here, there is no limitation on the number of image blocks or the number of image blocks.
[0086] Furthermore, each of the image blocks is segmented into cell nuclei to obtain cell nuclei masks that correspond one-to-one with each of the image blocks. Each cell nuclei mask is image data that has the features of all cell nuclei in the image block. In other words, each cell nuclei mask is actually image data with only black and white pixels, where black is the background and white is the cell nucleus.
[0087] Furthermore, feature extraction is performed based on each of the cell nucleus masks to obtain a second target feature vector, each of which has cell nucleus shape features, cell nucleus texture features, and cell nucleus interaction features. Each image patch has a second target feature vector. That is, we extracted three different quantitative features from the image patch, corresponding to the cell nucleus shape features, texture features, and interaction features, respectively, as described in detail below:
[0088] In one embodiment of this application, a single second target feature vector can be obtained through the following steps:
[0089] Based on the cell nucleus mask, the shape features of each cell nucleus in the image block are extracted to obtain a first feature vector with a first dimension of information. The first feature vector is used to characterize the cell nucleus shape features of the image block.
[0090] Based on the cell nucleus mask, the texture features of each cell nucleus in the image block are extracted to obtain a second feature vector with second-dimensional information. The second feature vector is used to characterize the cell nucleus texture features of the image block.
[0091] The interaction features of each cell nucleus in the image patch are extracted based on the cell nucleus mask to obtain a third feature vector with third-dimensional information. The third feature vector is used to characterize the cell nucleus interaction features of the image patch.
[0092] The first feature vector, the second feature vector, and the third feature vector are concatenated laterally to obtain the second target feature vector, which has the sum of the first dimension information, the second dimension information, and the third dimension information.
[0093] Specifically, the first feature vector of the first dimension information can be a 100-dimensional cell nucleus shape feature. The cell nucleus shape feature is used to quantify the morphological changes of a single cell nucleus. Specific descriptive parameters include area ratio, distance ratio, distance standard deviation, perimeter-area ratio, Fourier descriptor of boundary points, etc. Finally, 25-dimensional features are extracted from each cell nucleus, and each image patch will obtain an N-row 25-column feature matrix, where N represents the number of cell nuclei in the image patch. The mean, standard deviation, median, minimum-to-maximum ratio, and other four statistical measures are calculated for each column of this feature matrix, resulting in four feature vectors. These four feature vectors are horizontally concatenated to obtain a 100-dimensional feature vector to represent an image patch.
[0094] In one embodiment of this application, the first feature vector having first dimension information can be obtained through the following steps:
[0095] Based on the cell nucleus mask, the first target parameter feature of each cell nucleus is extracted, wherein the first target parameter feature is a feature vector with a first initial dimension information;
[0096] Perform a first preset statistical operation on each of the first target parameter features to obtain the second target parameter features of each of the cell nuclei. The second target parameter features are feature vectors with second initial dimension information.
[0097] The first target parameter feature and the second target parameter feature are concatenated laterally to obtain a first feature vector with the first dimension information.
[0098] Wherein, the first dimension information is the product of the first initial dimension information and the second initial dimension information.
[0099] Specifically, the first target parameter features are area ratio, distance ratio, distance standard deviation, perimeter-area ratio, Fourier descriptor of boundary points, etc., and the 25-dimensional (i.e., first initial dimension information) feature vector corresponding to each cell nucleus. The first preset statistical operation is to calculate the mean, standard deviation, median, and ratio of minimum to maximum value of each of the 25-dimensional feature vectors corresponding to each cell nucleus. These four statistics are the second target parameter features of four dimensions (i.e., second initial dimension information), thus obtaining four feature vectors. Then, each of the 25-dimensional feature vectors is concatenated with the four statistics to obtain a 25×4=100-dimensional first feature vector.
[0100] In one embodiment of this application, the second feature vector having second-dimensional information can be obtained through the following steps:
[0101] The third target parameter features of each cell nucleus are extracted based on the cell nucleus mask. The third target parameter features are feature vectors with third initial dimension information.
[0102] A second preset statistical operation is performed on each of the third target parameter features to obtain the fourth target parameter features of each of the cell nuclei. The fourth target parameter features are feature vectors with fourth initial dimension information.
[0103] The third target parameter feature and the fourth target parameter feature are concatenated laterally to obtain a second feature vector with the second dimension information;
[0104] The second dimension information is the product of the third initial dimension information and the fourth initial dimension information.
[0105] Specifically, the second feature vector is used to characterize cell nucleus texture features: cell nucleus texture features quantify the distribution and co-occurrence of pixels in a single cell nucleus. The specific descriptive parameters include 5-dimensional Haralick texture features, which include information entropy, energy, information metric 1, and information metric 2. Based on the gray-level co-occurrence matrix of each cell nucleus, a contrast matrix and an average intensity matrix are obtained. The mean, variance, energy, information entropy, and inverse moment of the contrast matrix are calculated to obtain 5-dimensional features. The mean, variance, and information entropy of the average intensity matrix are calculated to obtain 3-dimensional features. Finally, a 5+5+3, or 13-dimensional (third initial dimension information) feature vector is extracted from each cell nucleus, which is the third target parameter feature.
[0106] Each image patch will yield an N-row, 13-column feature matrix, where N represents the number of cell nuclei in the image patch. The mean, standard deviation, median, range, skewness, and kurtosis of each column of this feature matrix are calculated, resulting in six statistical measures (the fourth initial dimension information), which in turn yield six feature vectors, also known as the fourth target parameter features. These six feature vectors are then concatenated horizontally, resulting in a 6×13 matrix, which ultimately yields a 78-dimensional second feature vector to represent an image patch.
[0107] In one embodiment of this application, the third feature vector having third-dimensional information can be obtained through the following steps:
[0108] The fifth target parameter features of each cell nucleus are extracted based on the cell nucleus mask, and the fifth target parameter features are feature vectors with fifth initial dimension information;
[0109] A third preset statistical operation is performed on each of the fifth target parameter features to obtain the sixth target parameter features of each of the cell nuclei. The sixth target parameter features are feature vectors with sixth initial dimension information.
[0110] The fifth target parameter feature and the sixth target parameter feature are concatenated laterally to obtain a third feature vector containing the third dimension information;
[0111] The third dimension information is obtained based on the fifth initial dimension information and the sixth initial dimension information.
[0112] Specifically, the third feature vector is used to characterize the nucleus interaction features: the nucleus interaction features are used to quantify the structural characteristics of feature-driven local cell maps (FLocKs) and the interactions between feature-driven local cell maps. First, based on the centroid and area of the nuclei in each image patch, the nuclei are clustered using the mean-shift clustering algorithm within a bandwidth of 150, resulting in different categories of FLOCKs. Based on the clustered FLOCKs, specific descriptive parameters are calculated, including the absolute area of the intersection of two FLOCKs, the proportion of the number of intersecting FLOCKs to the total number of FLOCKs, the area ratio between the intersecting region and each FLOCK, Voronoi global graph related parameters constructed based on the FLOCK centroids, Delaunay triangulation related parameters, minimum spanning tree related parameters, etc. (fifth initial dimension information), etc., as the fifth objective parameter features. The mean, standard deviation, and range of these fifth objective parameter features (sixth initial dimension information) are then calculated as the sixth objective parameter features. Finally, a 222-dimensional third feature vector is obtained to characterize an image patch.
[0113] In one embodiment of this application, the second target feature vector is a feature vector having the sum of the first dimension information, the second dimension information, and the third dimension information. The step of concatenating the feature matrix to obtain the first target feature vector includes:
[0114] A third preset statistical operation is performed on each of the second target feature vectors in the feature matrix to obtain a third target feature vector with fourth dimension information;
[0115] The second target feature vector and the third target feature vector are concatenated to obtain the first target feature vector.
[0116] Specifically, by horizontally concatenating the first feature vector, the second feature vector, and the third feature vector, a second target feature vector (400 dimensions) is obtained, which has the sum of the first dimension information (100 dimensions), the second dimension information (78 dimensions), and the third dimension information (dimensional).
[0117] Furthermore, among the cropped image patches, the top 400 patches with the highest number of cell nuclei are selected to represent a digital pathological slide / patient, forming a 400×400 matrix. The first 400 represents the number of target image patches, and the second 400 represents the dimension of the second target feature vector, i.e., the first dimension information. For each column of this feature matrix, five statistical measures are calculated: mean, median, standard deviation, kurtosis, and skewness, resulting in five feature vectors. These five second feature vectors, possessing the second dimension information (5 dimensions), are then horizontally concatenated, resulting in a 2000-dimensional feature vector representing a digital pathological slide / patient.
[0118] In step 120, the first target feature vectors of each hepatocellular carcinoma patient are screened based on cross-validation and the Akaike Information Criterion to obtain target features.
[0119] Specifically, the first target feature vector obtained by feature concatenation has 2000-dimensional feature vectors, which has the problem of overfitting. Therefore, in this step 120, the first target feature vectors of each hepatocellular carcinoma patient are screened using cross-validation and the Akaike Information Criterion to obtain target features. The most representative target features are selected as training samples for the prognostic model, thereby making the training of the prognostic model more efficient and accurate.
[0120] In one embodiment of this application, the step of performing feature filtering on the first target feature vector of each of the hepatocellular carcinoma patients based on cross-validation and the Akaike Information Criterion to obtain target features includes:
[0121] The first target feature vectors of each hepatocellular carcinoma patient were subjected to 10-fold cross-validation in a preset round, and the target screening features obtained in each round were determined by the bidirectional selection stepwise regression method based on the Akaike information criterion. Each round corresponds to an Akaike information criterion value.
[0122] In each round, the target round with the smallest Akaike Information Criterion Value is selected, and the target screening features selected by the target round are obtained as the target features.
[0123] Specifically, the first target feature vector of each hepatocellular carcinoma patient is used as a discovery queue, which contains multiple first target feature vectors. Then, feature screening operations based on cross-validation and the Akaike Information Criterion are performed on the discovery queue.
[0124] The purpose of this application embodiment is to select the target feature with the highest prognostic efficacy. The target feature is also known as the Top feature, and there can be one or more target features.
[0125] First, Z-scores are performed on each feature value of the first target feature vector. Second, a LASSO Cox regression model is built on the discovery queue and 500 rounds (preset rounds) of 10-fold cross-validation are run. Top features can be added or subtracted during cross-validation. After each round of 10-fold cross-validation, further feature selection is performed through bidirectional selection stepwise regression based on the Akaike Information Criterion (AIC). Finally, the top feature with the smallest AIC value (Akaike Information Criterion value) in 500 rounds (preset rounds) is selected to construct a multidimensional fusion cell nucleus feature prognostic model, which is the trained prognostic model described in the embodiments of this application.
[0126] In step 130, the target features are used as training samples for a preset prognostic model to train the prognostic model.
[0127] Specifically, by using the target features as training samples for the pre-defined prognostic model, the trained prognostic model can more accurately and quickly assess the prognostic performance of patients with the target hepatocellular carcinoma.
[0128] In step 140, histopathological section data of the target hepatocellular carcinoma patient are obtained and input into the trained prognostic model to obtain the prognostic assessment score of the target hepatocellular carcinoma patient.
[0129] Specifically, the weighted linear combination of the Top features and their corresponding regression coefficients is used as the patient's risk score (S), calculated as follows:
[0130]
[0131] in xj The feature value representing the j-th top feature (i.e., one of the target features) j represents the regression coefficient corresponding to the j-th Top feature, and N represents the number of Top features (target features). In the discovery cohort, we determined the median risk score as the optimal cutoff threshold for classifying patients as high-risk or low-risk. Patients with risk scores less than or equal to this threshold were classified as low-risk, and those with scores greater than or equal to this threshold were classified as high-risk. This optimal cutoff threshold determined from the discovery cohort was then applied to the external validation set to classify patients as high- or low-risk.
[0132] Therefore, this application can obtain histopathological section data of patients with target hepatocellular carcinoma, quickly find the feature value corresponding to the target feature, input it into the above formula, obtain the corresponding prognostic assessment score, and evaluate the risk coefficient corresponding to the score based on the above optimal cutoff threshold.
[0133] To verify whether combining the prognostic model proposed in this application with clinicopathological variables can significantly enhance prognostic stratification, we performed the following operations: First, we included a total of five clinicopathological variables, including BCLC (Barcelona) stage, tumor grade (differentiation), MVI, age, and gender; second, we screened out predictors with p < 0.05 among the five clinicopathological variables through univariate analysis; finally, we performed Cox multivariate analysis on the model developed in this invention together with the predictors with p < 0.05, and used bidirectional stepwise regression based on AIC to determine the prognostic factor of the model with the smallest final AIC value, thus determining the complete model formed by combining the prognostic model proposed in this application with clinicopathological variables.
[0134] Figure 3 This is a block diagram of a prognostic assessment device 300 for hepatocellular carcinoma patients according to an embodiment of this application. The prognostic assessment device 300 for hepatocellular carcinoma patients according to an embodiment of this application includes: a first acquisition unit 301, a screening unit 302, a training unit 303, and a second acquisition unit 304.
[0135] The first acquisition unit 301 is used to acquire the first target feature vector of at least one hepatocellular carcinoma patient;
[0136] The screening unit 302 is used to perform feature screening on the first target feature vector of each of the hepatocellular carcinoma patients based on cross-validation and the Akaike Information Criterion to obtain target features;
[0137] Training unit 303 is used to use the target features as training samples for a preset prognostic model to train the prognostic model;
[0138] The second acquisition unit 304 is used to acquire histopathological section data of the target hepatocellular carcinoma patient and input the histopathological section data into the trained prognostic model to obtain the prognostic assessment score of the target hepatocellular carcinoma patient.
[0139] The first target feature vector of a single hepatocellular carcinoma patient can be obtained through the following steps:
[0140] Obtain digital pathological sections from the hepatocellular carcinoma patient, wherein the digital pathological sections are tissue pathological sections of the hepatocellular carcinoma patient stained with hematoxylin and eosin;
[0141] The digital pathological slides are cropped into image blocks to obtain a first preset number of image blocks;
[0142] Each image block is segmented into cell nuclei to obtain a cell nucleus mask that corresponds one-to-one with each image block. The cell nucleus mask is image data that has the features of all cell nuclei in the image block.
[0143] Based on each of the cell nucleus masks, feature extraction is performed to obtain a second target feature vector, each of which has cell nucleus shape features, cell nucleus texture features, and cell nucleus interaction features. Each of the image blocks has a second target feature vector.
[0144] A second target feature vector of the target image patch is obtained to generate a feature matrix corresponding to the target image patch and the second target feature vector, and the feature matrix is concatenated to obtain the first target feature vector.
[0145] In another aspect, this application also provides a computer-readable storage medium storing a program product capable of implementing the methods provided above in this specification. In some possible implementations, various aspects of this application may also be implemented as a program product comprising program code that, when run on a terminal device, causes the terminal device to perform the steps described in the "Embodiment Methods" section of this specification according to various exemplary embodiments of this application.
[0146] The program product for implementing the above-described method according to the embodiments of this application may employ a portable compact disc read-only memory (CD-ROM) and include program code, and may run on a terminal device, such as a personal computer. However, the program product of this application is not limited thereto. In this document, the readable storage medium may be any tangible medium containing or storing a program that may be used by or in conjunction with an instruction execution system, apparatus, or device.
[0147] The program product may employ any combination of one or more readable media. A readable medium may be a readable signal medium or a readable storage medium. A readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of readable storage media (a non-exhaustive list) include: an electrical connection having one or more wires, a portable disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.
[0148] Computer-readable signal media may include data signals propagated in baseband or as part of a carrier wave, carrying readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A readable signal medium may also be any readable medium other than a readable storage medium, capable of sending, propagating, or transmitting programs for use by or in conjunction with an instruction execution system, apparatus, or device.
[0149] The program code contained on the readable medium may be transmitted using any suitable medium, including but not limited to wireless, wired, optical fiber, RF, etc., or any suitable combination thereof.
[0150] Program code for performing the operations of this application can be written in any combination of one or more programming languages, including object-oriented programming languages such as Java and C++, and conventional procedural programming languages such as C or similar languages. The program code can execute entirely on the user's computing device, partially on the user's computing device, as a standalone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server. In cases involving remote computing devices, the remote computing device can be connected to the user's computing device via any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computing device (e.g., via the Internet using an Internet service provider).
[0151] In another respect, this application also provides an electronic device capable of implementing the above-described method.
[0152] Those skilled in the art will understand that various aspects of this application can be implemented as a system, method, or program product. Therefore, various aspects of this application can be specifically implemented in the following forms: a completely hardware implementation, a completely software implementation (including firmware, microcode, etc.), or a combination of hardware and software implementations, collectively referred to herein as a "circuit," "module," or "system."
[0153] The following reference Figure 4 To describe an electronic device 400 according to this embodiment of the present application. Figure 4 The electronic device 400 shown is merely an example and should not impose any limitations on the functionality and scope of use of the embodiments of this application.
[0154] like Figure 4 As shown, the electronic device 400 is manifested in the form of a general-purpose computing device. The components of the electronic device 400 may include, but are not limited to: at least one processing unit 410, at least one storage unit 420, and a bus 430 connecting different system components (including storage unit 420 and processing unit 410).
[0155] The storage unit stores program code that can be executed by the processing unit 410, causing the processing unit 410 to perform the steps described in the "Embodiment Methods" section above according to various exemplary embodiments of this application.
[0156] Storage unit 420 may include readable media in the form of volatile storage units, such as random access memory (RAM) 421 and / or cache memory 422, and may further include read-only memory (ROM) 423.
[0157] Storage unit 420 may also include a program / utility 424 having a set (at least one) of program modules 425, including but not limited to: an operating system, one or more application programs, other program modules, and program data, each or some combination of these examples may include an implementation of a network environment.
[0158] Bus 430 can represent one or more of several types of bus structures, including a memory cell bus or memory cell control node, a peripheral bus, a graphics acceleration port, a processing unit, or a local bus using any of the various bus structures.
[0159] Electronic device 400 can also communicate with one or more external devices 1200 (e.g., keyboard, pointing device, Bluetooth device, etc.), and with one or more devices that enable a user to interact with electronic device 400, and / or with any device that enables electronic device 400 to communicate with one or more other computing devices (e.g., router, modem, etc.). This communication can be performed via input / output (I / O) interface 450. Furthermore, electronic device 400 can also communicate with one or more networks (e.g., local area network (LAN), wide area network (WAN), and / or public networks, such as the Internet) via network adapter 460. As shown, network adapter 460 communicates with other modules of electronic device 400 via bus 430. It should be understood that, although not shown in the figures, other hardware and / or software modules can be used in conjunction with electronic device 400, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems.
[0160] Through the above description of the embodiments, those skilled in the art will readily understand that the exemplary embodiments described herein can be implemented by software or by combining software with necessary hardware. Therefore, the technical solutions according to the embodiments of this application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, USB flash drive, external hard drive, etc.) or on a network, including several instructions to cause a computing device (such as a personal computer, server, terminal device, or network device, etc.) to execute the methods according to the embodiments of this application.
[0161] Furthermore, the above figures are merely illustrative of the processes included in the method according to exemplary embodiments of this application, and are not intended to be limiting. It is readily understood that the processes shown in the above figures do not indicate or limit the temporal order of these processes. Additionally, it is readily understood that these processes may be executed synchronously or asynchronously, for example, in multiple modules.
[0162] It should be understood that this application is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this application is limited only by the appended claims.
Claims
1. A method for prognostic assessment of hepatocellular carcinoma patients, characterized in that, The method includes: Obtain the first target feature vector of at least one hepatocellular carcinoma patient; Based on cross-validation and the Akaike Information Criterion, the first target feature vectors of each of the hepatocellular carcinoma patients are filtered to obtain target features; The target features are used as training samples for a pre-defined prognostic model to train the prognostic model. Obtain histopathological section data of the target hepatocellular carcinoma patient and input the histopathological section data into the trained prognostic model to obtain the prognostic assessment score of the target hepatocellular carcinoma patient; The first target feature vector of a single hepatocellular carcinoma patient is obtained through the following steps: Obtain digital pathological sections from the hepatocellular carcinoma patient, wherein the digital pathological sections are tissue pathological sections of the hepatocellular carcinoma patient stained with hematoxylin and eosin; The digital pathological slides are cropped into image blocks to obtain a first preset number of image blocks; Each image block is segmented into cell nuclei to obtain a cell nucleus mask that corresponds one-to-one with each image block. The cell nucleus mask is image data that has the features of all cell nuclei in the image block. Feature extraction is performed based on each of the cell nucleus masks to obtain a second target feature vector, each of which has cell nucleus shape features, cell nucleus texture features and cell nucleus interaction features. Each of the image blocks has a second target feature vector. A second target feature vector of the target image patch is obtained to generate a feature matrix corresponding to the target image patch and the second target feature vector, and the feature matrix is concatenated to obtain the first target feature vector; A single second target feature vector is obtained through the following steps: Based on the cell nucleus mask, the shape features of each cell nucleus in the image block are extracted to obtain a first feature vector with a first dimension of information. The first feature vector is used to characterize the cell nucleus shape features of the image block. Based on the cell nucleus mask, the texture features of each cell nucleus in the image block are extracted to obtain a second feature vector with second-dimensional information. The second feature vector is used to characterize the cell nucleus texture features of the image block. The interaction features of each cell nucleus in the image patch are extracted based on the cell nucleus mask to obtain a third feature vector with third-dimensional information. The third feature vector is used to characterize the cell nucleus interaction features of the image patch. The first feature vector, the second feature vector, and the third feature vector are horizontally concatenated to obtain the second target feature vector, which has the sum of the first dimension information, the second dimension information, and the third dimension information. The first target feature vector of each hepatocellular carcinoma patient is filtered based on cross-validation and the Akaike Information Criterion to obtain target features, including: The first target feature vectors of each hepatocellular carcinoma patient were subjected to 10-fold cross-validation in a preset round, and the target screening features obtained in each round were determined by the bidirectional selection stepwise regression method based on the Akaike information criterion. Each round corresponds to an Akaike information criterion value. In each round, the target round with the smallest Akaike Information Criterion Value is selected, and the target screening features selected by the target round are obtained as the target features.
2. The prognostic assessment method for hepatocellular carcinoma patients according to claim 1, characterized in that, The first feature vector with the first dimension information is obtained through the following steps: Based on the cell nucleus mask, the first target parameter feature of each cell nucleus is extracted, wherein the first target parameter feature is a feature vector with a first initial dimension information; Perform a first preset statistical operation on each of the first target parameter features to obtain the second target parameter features of each of the cell nuclei. The second target parameter features are feature vectors with second initial dimension information. The first target parameter feature and the second target parameter feature are concatenated laterally to obtain a first feature vector with the first dimension information. Wherein, the first dimension information is the product of the first initial dimension information and the second initial dimension information.
3. The prognostic assessment method for hepatocellular carcinoma patients according to claim 2, characterized in that, The second feature vector with second-dimensional information is obtained through the following steps: The third target parameter features of each cell nucleus are extracted based on the cell nucleus mask. The third target parameter features are feature vectors with third initial dimension information. A second preset statistical operation is performed on each of the third target parameter features to obtain the fourth target parameter features of each of the cell nuclei. The fourth target parameter features are feature vectors with fourth initial dimension information. The third target parameter feature and the fourth target parameter feature are concatenated laterally to obtain a second feature vector with the second dimension information; The second dimension information is the product of the third initial dimension information and the fourth initial dimension information.
4. The prognostic assessment method for hepatocellular carcinoma patients according to claim 3, characterized in that, The third feature vector with third-dimensional information is obtained through the following steps: The fifth target parameter features of each cell nucleus are extracted based on the cell nucleus mask, and the fifth target parameter features are feature vectors with fifth initial dimension information; A third preset statistical operation is performed on each of the fifth target parameter features to obtain the sixth target parameter features of each of the cell nuclei. The sixth target parameter features are feature vectors with sixth initial dimension information. The fifth target parameter feature and the sixth target parameter feature are concatenated laterally to obtain a third feature vector containing the third dimension information; The third dimension information is obtained based on the fifth initial dimension information and the sixth initial dimension information.
5. The prognostic assessment method for hepatocellular carcinoma patients according to claim 4, characterized in that, The second target feature vector is a feature vector containing the sum of the first dimension information, the second dimension information, and the third dimension information. The step of concatenating the feature matrix to obtain the first target feature vector includes: A third preset statistical operation is performed on each of the second target feature vectors in the feature matrix to obtain a third target feature vector with fourth dimension information; The second target feature vector and the third target feature vector are concatenated to obtain the first target feature vector.
6. A prognostic assessment device for patients with hepatocellular carcinoma, characterized in that, The device includes: The first acquisition unit is used to acquire the first target feature vector of at least one hepatocellular carcinoma patient; The screening unit is used to perform feature screening on the first target feature vector of each of the hepatocellular carcinoma patients based on cross-validation and the Akaike Information Criterion to obtain target features; The training unit is used to use the target features as training samples for a preset prognostic model to train the prognostic model. The second acquisition unit is used to acquire histopathological section data of the target hepatocellular carcinoma patient and input the histopathological section data into the trained prognostic model to obtain the prognostic assessment score of the target hepatocellular carcinoma patient. The first target feature vector of a single hepatocellular carcinoma patient is obtained through the following steps: Obtain digital pathological sections from the hepatocellular carcinoma patient, wherein the digital pathological sections are tissue pathological sections of the hepatocellular carcinoma patient stained with hematoxylin and eosin; The digital pathological slides are cropped into image blocks to obtain a first preset number of image blocks; Each image block is segmented into cell nuclei to obtain a cell nucleus mask that corresponds one-to-one with each image block. The cell nucleus mask is image data that has the features of all cell nuclei in the image block. Feature extraction is performed based on each of the cell nucleus masks to obtain a second target feature vector, each of which has cell nucleus shape features, cell nucleus texture features and cell nucleus interaction features. Each of the image blocks has a second target feature vector. A second target feature vector of the target image patch is obtained to generate a feature matrix corresponding to the target image patch and the second target feature vector, and the feature matrix is concatenated to obtain the first target feature vector; A single second target feature vector is obtained through the following steps: Based on the cell nucleus mask, the shape features of each cell nucleus in the image block are extracted to obtain a first feature vector with a first dimension of information. The first feature vector is used to characterize the cell nucleus shape features of the image block. Based on the cell nucleus mask, the texture features of each cell nucleus in the image block are extracted to obtain a second feature vector with second-dimensional information. The second feature vector is used to characterize the cell nucleus texture features of the image block. The interaction features of each cell nucleus in the image patch are extracted based on the cell nucleus mask to obtain a third feature vector with third-dimensional information. The third feature vector is used to characterize the cell nucleus interaction features of the image patch. The first feature vector, the second feature vector, and the third feature vector are horizontally concatenated to obtain the second target feature vector, which has the sum of the first dimension information, the second dimension information, and the third dimension information. The first target feature vector of each hepatocellular carcinoma patient is filtered based on cross-validation and the Akaike Information Criterion to obtain target features, including: The first target feature vectors of each hepatocellular carcinoma patient were subjected to 10-fold cross-validation in a preset round, and the target screening features obtained in each round were determined by the bidirectional selection stepwise regression method based on the Akaike information criterion. Each round corresponds to an Akaike information criterion value. In each round, the target round with the smallest Akaike Information Criterion Value is selected, and the target screening features selected by the target round are obtained as the target features.
7. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores at least one piece of program code, which is loaded and executed by a processor to perform the operations performed by the method as described in any one of claims 1 to 5.
8. An electronic device, characterized in that, The electronic device includes one or more processors and one or more memories, wherein at least one piece of program code is stored in the one or more memories, and the at least one piece of program code is loaded and executed by the one or more processors to perform the operation performed by the method as described in any one of claims 1 to 5.
Citation Information
Patent Citations
Cell nucleus segmentation method, system and device and cancer auxiliary analysis system and device based on pathological image
CN113222944A
Gene space expression prediction method based on tumor pathological image
CN116844631A