A method for extracting unlabeled pathological image features based on spatial location information
By using a self-supervised learning framework to generate a training dataset using the spatial location information between pathological images, the problem of relying on manual annotation for pathological image feature extraction is solved, and efficient, accurate feature extraction and consistency models are achieved.
Patent Information
- Application Number
- CN202111454090.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-01
- Publication Date
- 2025-11-14
- Estimated Expiration
- 2041-12-01
AI Technical Summary
Existing methods for extracting features from pathological images rely on manual annotation, which is inefficient and inconsistent. They also ignore the positional and spatial relationships between image blocks, resulting in incomplete feature extraction.
A self-supervised pre-task module is constructed to generate a training dataset using spatial location information between images. The model is trained through a self-supervised learning framework without manual annotation, and features are extracted using similar/difference data pairs.
It achieves pathological image feature extraction without manual annotation, reduces the difficulty and cost of dataset acquisition, improves the accuracy and consistency of feature extraction, and has good model compatibility.
Smart Images

Figure CN115439843B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to a method for extracting unlabeled features from pathological images based on spatial location information, belonging to the field of digital image processing. Background Technology
[0002] Pathological slides accurately reflect the state of human tissue. Doctors can diagnose lesions by assessing the tissue structure, distribution density, and other characteristics of cells, glands, etc., in pathological samples. However, manually summarizing the features of pathological samples is too subjective, and the diagnostic results of different doctors vary significantly, resulting in low consistency. With the development of fully digital pathology technology and artificial intelligence technology, it has become possible to perform feature analysis on digital pathological slides using computers. For example, using graphics-related algorithms to analyze the texture and spatial features of pathological images, and using deep convolutional neural networks to analyze the multidimensional and frequency domain features of pathological images, can provide objective and accurate feature extraction results, compensating for the instability, inconsistency, and lack of objectivity of manual analysis.
[0003] However, current computer-based feature extraction from pathological images typically requires massive amounts of data to train the corresponding models. Generating this training data necessitates experienced pathologists manually outlining lesion areas on the images, overly relying on the doctor's prior knowledge of pathological diagnosis. This annotation process is inefficient and inconsistent, placing an excessive burden on doctors. Furthermore, common deep learning models are trained using pathological images and corresponding annotations, analyzing each image patch separately. This ignores the positional relationships between image patches within a pathological image and fails to consider the spatial relationships at different magnifications, resulting in existing feature extraction methods missing some key features of pathological images.
[0004] Chinese patent CN108229576A proposes a method for learning cross-magnification pathological image features. It mentions using high-magnification, precisely labeled data for pre-training, followed by reconstruction training using low-magnification data to obtain cross-magnification features. Although it doesn't consider the spatial relationship between high and low magnification, it demonstrates the feasibility of using overlapping high and low magnification data in supervised learning. Chinese patent CN107480702A proposes a feature selection and fusion method for HCC pathological image recognition. It integrates multiple features such as deep learning, encoding, and texture for subsequent tasks. Although it doesn't consider the spatial relationship between images, it also achieves good model performance. Summary of the Invention
[0005] The purpose of this invention is to provide a method for extracting pathological image features that does not rely on manual lesion annotation but can make full use of the relative spatial location information between images.
[0006] To achieve the above objectives, the technical solution of the present invention provides a method for extracting unlabeled pathological image features based on spatial location information, characterized by comprising the following steps:
[0007] Step S101: Construct a self-supervised pre-task module to generate a training dataset based on spatial location relationships, including the following steps:
[0008] Step S101-1: Obtain the original set of pathological images;
[0009] Step S101-2: Overlap and slice the original pathological image set at different resolutions and starting coordinates to form an image patch set, denoted as x, then x = {x1, x2, ..., x} N}, x N This represents the Nth image patch in the set of image patches;
[0010] Step S101-3: Record the spatial location information of each image patch to form an image patch information vector. All image patch information vectors form an image patch information vector set, denoted as i, then i = {i1, i2, ..., i...} N}, i N This represents the image block information vector of the Nth image block.
[0011] Step S101-4: Analyze the spatial relationships of the images, define the rules for generating similar data pairs and different data pairs, and construct the dataset. Wherein: if two map tiles are spatially adjacent, intersecting, or contained, then these two map tiles are defined as a similar data pair; otherwise, the two map tiles are a different data pair.
[0012] Step S102: Construct a self-supervised learning framework and train the model:
[0013] The initial input to the self-supervised learning framework is the data pair obtained in step S101, i.e., paired images, denoted as a and b respectively. If a and b are similar data pairs, the label is denoted as Y=1; if a and b are different data pairs, the label is denoted as Y=0. The self-supervised learning framework processes images a and b as follows:
[0014] Perform random data augmentation on images a and b, and denote the augmented images as a′ and b′;
[0015] Input image a into deep neural network g θ Feature extraction is performed to obtain feature representation F0, which can be used for various subsequent tasks.
[0016] Shared deep neural network g θ Weights, gradient removal, and the resulting deep neural network g′ θ ;
[0017] Input images b, a′, and b′ into deep neural network g′ θ The corresponding feature representations are obtained respectively, denoted as F1, F2 and F3;
[0018] Input feature F0 into feature projection network p θ We perform feature dimensionality reduction projection to obtain the two-dimensional projection vector E0 of feature F0;
[0019] Shared feature projection network p θ Weights, gradient removal, and the resulting feature projection network p′ θ ;
[0020] Features F1, F2, and F3 are input into the feature projection network p′ θ The corresponding two-dimensional projection vectors E1, E2 and E3 are obtained respectively;
[0021] The similarity between two-dimensional projection vectors E0 and E1, E2 and E3 is calculated separately. The similarity between E0 and E1 is used as the loss for similar images. The overall image similarity loss function is defined as follows:
[0022]
[0023] In equation (2), N is the sum of the original images and the augmented images used in the current training;
[0024] For the others, pairwise image similarity loss functions are established with the label Y. The pairwise image similarity loss function is defined as follows:
[0025]
[0026] In equation (3), i = 2, 3; It is E0 and E i Similarity measurement algorithm
[0027] Finally, only deep neural networks g are discussed. θ and feature projection network p θ Perform optimization and gradient backpropagation, and iterate training until the model achieves the expected results;
[0028] Step S103: Preprocess the input pathological image to obtain pathological patches;
[0029] Step S104: Input the pathological map obtained in step S103 into the deep neural network g obtained in step S102. θ The prediction is performed to obtain the feature representation F0, and this feature representation is used for various subsequent tasks.
[0030] Preferably, in steps S101-3, each image patch includes starting coordinates, image size, slice resolution, and the original pathological image to which it belongs.
[0031] Preferably, in steps S101-4, the rule functions for generating similar data pairs and generating difference data pairs are defined as follows:
[0032]
[0033] In equation (1), i i ∈i、i j ∈i, where i is the image block information vector of the i-th image block and j is the image block information vector of the j-th image block; function f l (i i i j The function f represents determining whether the i-th image patch and the j-th image patch are adjacent or intersecting under the same original pathological image and the same scaling factor; s (i i i j The function f represents determining whether the i-th image patch and the j-th image patch contain each other under different scaling ratios for the same original pathological image; o (i i i j This indicates whether the i-th image patch and the j-th image patch come from the same source.
[0034] Preferably, in step S102, the similarity measurement algorithm Defined as two-dimensional projection vector E0 and two-dimensional projection vector E i In a two-dimensional projected vector space, the cosine distance indicates similarity. The value should ideally be 1, and the difference should be... The value should be as low as possible (-1); when Y = 1, minimize it as much as possible. To achieve the maximum similarity; when Y=0, minimize the similarity as much as possible. To achieve the greatest degree of difference.
[0035] Preferably, step S103 specifically includes the following steps:
[0036] Step S103-1: Obtain the original pathological images;
[0037] Step S103-2: Extract the low-magnification thumbnail and calculate the tissue region outline on it;
[0038] Step S103-3: Count the number of pixels in the organization region on the thumbnail and calculate the coordinates of the lower image block corresponding to the maximum magnification for each pixel;
[0039] Step S103-4: Cut the map at the maximum magnification according to the coordinates to generate the overall dataset of map tiles to be predicted;
[0040] Step S103-5: Standardize the color of each tile.
[0041] This invention eliminates the need for manual labeling of lesions. By employing a self-supervised algorithm, it constructs a dataset and labels using information such as the spatial location and scaling factor of different pathological image blocks, and trains relevant models to complete the extraction of pathological image features and subsequent tasks.
[0042] Compared to existing supervised and weakly supervised deep learning algorithms, this invention is easier to implement. On one hand, it utilizes self-supervised learning to automatically construct similar / difference data pairs using spatial location information, eliminating the need for manual annotation and significantly reducing the difficulty of acquiring datasets, thus saving substantial manpower and computing power costs. On the other hand, this invention constructs a novel training framework that fully leverages the similarity / difference levels of different data to automatically train the feature extraction model. Furthermore, it places no specific structural requirements on the feature extraction model itself, making this invention highly compatible, easy to implement, and allowing for the replacement of specific networks according to different tasks. Attached Figure Description
[0043] Figure 1 This is a flowchart illustrating the overall process of the label-free pathological image feature extraction method based on spatial location information according to the present invention.
[0044] Figure 2 The flowchart illustrates the method for constructing a self-supervised pre-task module to generate a training dataset based on spatial location relationships in this invention.
[0045] Figure 3 This is a schematic diagram illustrating the rules for generating similar / difference data pairs as defined in this invention.
[0046] Figure 4 A schematic diagram of the self-supervised learning framework constructed in this invention;
[0047] Figure 5 This is a flowchart of the preprocessing method for input pathological images according to the present invention. Detailed Implementation
[0048] The present invention will be further illustrated below with reference to specific embodiments. It should be understood that these embodiments are for illustrative purposes only and are not intended to limit the scope of the invention. Furthermore, it should be understood that after reading the teachings of this invention, those skilled in the art can make various alterations or modifications to the invention, and these equivalent forms also fall within the scope defined by the appended claims.
[0049] like Figure 1 As shown in the figure, this embodiment discloses a method for extracting unlabeled pathological image features based on spatial location information, which specifically includes the following steps:
[0050] Step S101: Construct a self-supervised pre-task module to generate a training dataset based on spatial location relationships.
[0051] like Figure 2 As shown, step S101 includes the following steps:
[0052] Step S101-1: Obtain the raw pathological images (WSI) set;
[0053] Step S101-2: Overlap and slice the original pathological image set at different resolutions and starting coordinates to form an image patch set, denoted as x, then x = {x1, x2, ..., x} N}, x N This represents the Nth image patch in the set of image patches;
[0054] Step S101-3: Record the spatial location information of each image patch to form an image patch information vector. All image patch information vectors form an image patch information vector set, denoted as i, then i = {i1, i2, ..., i...} N}, i N This represents the image patch information vector of the Nth image patch, where:
[0055] Each image tile contains, but is not limited to, vector information such as: starting coordinates, image size, slice resolution, and the original pathological image to which it belongs;
[0056] Step S101-4: Analyze the spatial relationships of images, define the generation rules for similar data pairs and different data pairs, and construct the dataset;
[0057] like Figure 3 As shown, this invention designs a rule function to measure the spatial similarity s of two images. Specifically, if two image patches are adjacent, intersecting, or contain each other in spatial position, then the two image patches are defined as a similar data pair; otherwise, the two image patches are a difference data pair.
[0058] The rule functions for generating rules for similar data pairs and different data pairs are defined as follows:
[0059]
[0060] In equation (1), i i ∈i、i j ∈i, where i is the image block information vector of the i-th image block and j is the image block information vector of the j-th image block; function f l (i i i j The function f represents determining whether the i-th image patch and the j-th image patch are adjacent or intersecting under the same original pathological image and the same scaling factor; s (ii i j The function f represents determining whether the i-th image patch and the j-th image patch contain each other under different scaling ratios for the same original pathological image; o (i i i j This indicates whether the i-th image patch and the j-th image patch come from the same source.
[0061] Step S102: Construct a self-supervised learning framework and train the model.
[0062] like Figure 4 As shown, the original input of the self-supervised learning framework is the data pair obtained in step S101, i.e., the paired images, denoted as a and b respectively. If a and b are similar data pairs, the label is denoted as Y=1; if a and b are different data pairs, the label is denoted as Y=0. The self-supervised learning framework processes images a and b as follows:
[0063] Perform random data augmentation on images a and b, such as translation, rotation, cropping, adding noise, and color transformation. The augmented images are denoted as a′ and b′.
[0064] Input image a into deep neural network g θ Feature extraction is performed to obtain feature representation F0, which can be used for various subsequent tasks.
[0065] Shared deep neural network g θ Weights, gradient removal, and the resulting deep neural network g′ θ ;
[0066] Input images b, a′, and b′ into deep neural network g′ θ The corresponding feature representations are obtained respectively, denoted as F1, F2 and F3;
[0067] Input feature F0 into feature projection network p θ We perform feature dimensionality reduction projection to obtain the two-dimensional projection vector E0 of feature F0;
[0068] Shared feature projection network p θ Weights, gradient removal, and the resulting feature projection network p′ θ ;
[0069] Features F1, F2, and F3 are input into the feature projection network p′ θ The corresponding two-dimensional projection vectors E1, E2 and E3 are obtained respectively;
[0070] The similarity between two-dimensional projection vectors E0 and E1, E2 and E3 is calculated respectively. Since image a′ is an enhanced image of image a, the two-dimensional projection vectors E0 and E1 are used as similar images to calculate the loss. The others are used to establish pairwise image similarity loss functions with label Y.
[0071] Finally, only deep neural networks g are discussed. θ and feature projection network p θ Optimization and gradient backpropagation are performed, and the model is iteratively trained until it achieves the expected results. Here, the deep neural networks and feature projection networks used are not limited to a specific model.
[0072] In the above steps, the pairwise image similarity loss function is defined as follows:
[0073]
[0074] In equation (2), i = 2, 3; It is E0 and E i The similarity measurement algorithm is defined here as the similarity between the two-dimensional projection vector E0 and the two-dimensional projection vector E i In a two-dimensional projected vector space, the cosine distance indicates similarity. The value should ideally be 1, and the difference should be... The value should be as low as possible (-1); when Y = 1, minimize it as much as possible. To achieve the maximum similarity; when Y=0, minimize the similarity as much as possible. To achieve the greatest degree of difference.
[0075] The similarity measurement algorithm used here is not limited to any one specific algorithm.
[0076] In the above steps, the overall image similarity loss function used to calculate the loss for similar images is defined as follows:
[0077]
[0078] In equation (3), N is the sum of the original images and the augmented images used in the current training.
[0079] Step S103: Preprocess the input pathological images.
[0080] like Figure 5 As shown, step S103 specifically includes the following steps:
[0081] Step S103-1: Obtain raw pathological images (WSI);
[0082] Step S103-2: Extract the low-magnification thumbnail and calculate the tissue region outline on it;
[0083] Step S103-3: Count the number of pixels in the organization region on the thumbnail and calculate the coordinates of the lower image block corresponding to the maximum magnification for each pixel;
[0084] Step S103-4: Cut the map at the maximum magnification according to the coordinates to generate the overall dataset of map tiles to be predicted;
[0085] Step S103-5: Standardize the color of each tile.
[0086] Step S104: Calculate the feature representation of the real-time input pathological image using the model, and use it for subsequent tasks:
[0087] The pathological image obtained in step S103 is fed into the deep neural network g obtained in step S102. θ The prediction is performed to obtain the feature representation F0, and this feature representation is used for various subsequent tasks.
Claims
1. A method for extracting unlabeled features from pathological images based on spatial location information, characterized in that, Includes the following steps: Step S101: Construct a self-supervised pre-task module to generate a training dataset based on spatial location relationships, including the following steps: Step S101-1: Obtain the original set of pathological images; Step S101-2: Overlap and slice the original pathological image set at different resolutions and starting coordinates to form an image patch set, denoted as x, then x = {x1, x2, ..., x} N }, x N This represents the Nth image patch in the set of image patches; Step S101-3: Record the spatial location information of each image patch to form an image patch information vector. All image patch information vectors form an image patch information vector set, denoted as i, then i = {i1, i2, ..., i...} N }, i N This represents the image block information vector of the Nth image block. Step S101-4: Analyze the spatial relationships of the images, define the rules for generating similar data pairs and different data pairs, and construct the dataset. Wherein: if two map tiles are spatially adjacent, intersecting, or contained, then these two map tiles are defined as a similar data pair; otherwise, the two map tiles are a different data pair. The rule functions for generating rules for similar data pairs and different data pairs are defined as follows: In equation (1), i i ∈i、i j ∈i, where i is the image block information vector of the i-th image block and j is the image block information vector of the j-th image block; function f l (i i i j The function f represents determining whether the i-th image patch and the j-th image patch are adjacent or intersecting under the same original pathological image and the same scaling factor; s (i i i j The function f represents determining whether the i-th image patch and the j-th image patch contain each other under different scaling ratios for the same original pathological image; o (i i i j This indicates whether the i-th image patch and the j-th image patch come from the same source; Step S102: Construct a self-supervised learning framework and train the model: The initial input to the self-supervised learning framework is the data pair obtained in step S101, i.e., paired images, denoted as a and b respectively. If a and b are similar data pairs, the label is denoted as Y=1; if a and b are different data pairs, the label is denoted as Y=0. The self-supervised learning framework processes images a and b as follows: Perform random data augmentation on images a and b. The augmented image is denoted as a. i and b i ; Input image a into deep neural network g θ Feature extraction is performed to obtain feature representation F0, which can be used for various subsequent tasks. Shared deep neural network g θ Weights, gradient removal, and the resulting deep neural network g. ′ θ ; Image b and image a ′ and image b ′ Input deep neural network g ′ θ The corresponding feature representations are obtained respectively, denoted as F1, F2 and F3; Input feature F0 into feature projection network p θ We perform feature dimensionality reduction projection to obtain the two-dimensional projection vector E0 of feature F0; Shared feature projection network p θ Weights, gradient removal, and the resulting feature projection network p ′ θ ; Features F1, F2, and F3 are input into the feature projection network p. ′ θ The corresponding two-dimensional projection vectors E1, E2 and E3 are obtained respectively; The similarity between two-dimensional projection vectors E0 and E1, E2 and E3 is calculated separately. The similarity between E0 and E1 is used as the loss for similar images. The overall image similarity loss function is defined as follows: In equation (2), N is the sum of the original images and the augmented images used in the current training; For the others, pairwise image similarity loss functions are established with the label Y. The pairwise image similarity loss function is defined as follows: In equation (3), i = 2, 3; It is E0 and E i Similarity measurement algorithms; Similarity measurement algorithm Defined as two-dimensional projection vector E0 and two-dimensional projection vector E i In a two-dimensional projected vector space, the cosine distance indicates similarity. The value approaches 1, and the difference is... The value approaches -1; when Y = 1, minimize it as much as possible. To achieve the maximum similarity; when Y=0, minimize the similarity as much as possible. To achieve the maximum degree of difference; Finally, only deep neural networks g are discussed. θ and feature projection network p θ Perform optimization and gradient backpropagation, and iterate training until the model achieves the expected results; Step S103: Preprocess the input pathological image to obtain pathological patches; Step S104: Input the pathological map obtained in step S103 into the deep neural network g obtained in step S102. θ The prediction is performed to obtain the feature representation F0, and this feature representation is used for various subsequent tasks.
2. The method for extracting unlabeled pathological image features based on spatial location information as described in claim 1, characterized in that, In steps S101-3, each image patch includes starting coordinates, image size, slice resolution, and the original pathological image to which it belongs.
3. The method for extracting unlabeled pathological image features based on spatial location information as described in claim 1, characterized in that, Step S103 specifically includes the following steps: Step S103-1: Obtain the original pathological images; Step S103-2: Extract the low-magnification thumbnail and calculate the tissue region outline on it; Step S103-3: Count the number of pixels in the organization region on the thumbnail and calculate the coordinates of the lower image block corresponding to the maximum magnification for each pixel; Step S103-4: Cut the map at the maximum magnification according to the coordinates to generate the overall dataset of map tiles to be predicted; Step S103-5: Standardize the color of each tile.
Citation Information
Patent Citations
Characteristic selection and characteristic fusion method for HCC pathological image identification
CN107480702A
Cross multiplying power pathological image feature learning method
CN108229576A
Liver tumor recognition method based on self-supervised dense convolutional neural network
CN113362295A