A method for automatically analyzing pathological sections of digestive tract biopsy

By dividing the pathological slice images into sub-regions, combining feature extraction, distance transformation and graph convolution networks, the problem of low classification accuracy in the prior art is solved, and more accurate pathological slice diagnosis and lesion area positioning are achieved.

CN112927215BActive Publication Date: 2025-05-06MOTIC XIAMEN MEDICAL DIAGNOSTICS SYST
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110281259.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-03-16
Publication Date
2025-05-06
Estimated Expiration
2041-03-16

AI Technical Summary

Technical Problem

The existing full-section classification strategy is difficult to accurately diagnose the problem of reduced classification accuracy due to the ambiguous number of decisive image blocks or the loss of position information when processing pathological sections of digestive tract biopsy.

Method used

The pathological slice images are divided into multiple sub-regions, and the pre-constructed feature extraction model and image classification model are combined with the sliding window method and the distance transformation algorithm to extract and quantify the features, generate an organizational structure chart, and finally use the graph convolution network model for classification prediction.

Benefits of technology

It improves the classification accuracy of pathological sections of digestive tract biopsy, can help doctors complete diagnosis more accurately, and clarify the lesion area through heat maps.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN112927215B_ABST
    Figure CN112927215B_ABST
Patent Text Reader

Abstract

The present invention discloses an automatic analysis method for digestive tract biopsy pathological sections, comprising: dividing the pathological section image into multiple sub-regions; based on a feature extraction model, extracting features from each sub-region one by one to obtain a feature matrix of each sub-region; classifying each feature vector in each feature matrix based on an image classification model to obtain a corresponding lesion probability matrix; based on a distance transformation algorithm, quantizing the distance of each feature vector in each feature matrix to obtain a distance quantization feature matrix; cascading the feature matrix of each sub-region and the distance quantization feature matrix to obtain a fused feature matrix; based on the fused feature matrix and the lesion probability matrix of each sub-region, generating an organizational structure diagram of each sub-region; and classifying and predicting each organizational structure diagram based on a graph convolutional network model to generate a diagnosis result of each sub-region. The present invention can screen biopsy pathology, automatically output diagnostic conclusions, and assist doctors to complete their work accurately and efficiently.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of medical image processing, and more particularly to an automatic analysis method for digestive tract biopsy pathological sections. Background Art

[0002] Gastrointestinal endoscopic tissue biopsy is an important means of screening for gastric cancer, intestinal cancer and other digestive system cancers. Pathologists complete screening by examining the microscopic sections of biopsy tissue one by one. However, there are many cases of gastrointestinal endoscopic examinations, and pathologists have a heavy workload for diagnosis. Long-term and high-intensity work can easily lead to misdiagnosis. With the rapid development of computer and microscopic imaging technology, digital pathology images are easy to obtain and fast. Computer automatic analysis algorithms suitable for digital pathology full-slice images have become a research hotspot in this direction in recent years.

[0003] Since the resolution of digital pathology images is much higher than that of natural scene images, it is difficult for computer vision algorithms to directly process the entire pathology image. In order to classify the entire slice, most existing algorithms use the full slice image block method, first classifying the local area, and then formulating a specific strategy to complete the classification of the entire slice based on the local prediction results. Common strategies include the following three:

[0004] 1) Full-slice classification strategy based on majority voting

[0005] This strategy regards the classification results of the image blocks contained in the slice as a vote for the category of the entire slice, and directly outputs the category with the most votes as the category of the slice. This method is simple and intuitive, and can obtain correct results when the image blocks that play a decisive role in diagnosis in the slice occupy a majority in number. However, in pathological diagnosis, sometimes the image blocks that determine the diagnosis result of the entire slice do not occupy a majority in number, and even the number accounts for less than 1% of the total number of image blocks. In such cases, the majority voting strategy is difficult to give a correct classification result for the entire slice.

[0006] 2) Full-slice classification strategy based on convolutional neural network

[0007] This strategy arranges the image block classification results or image block features into a three-dimensional tensor according to their positional relationship in the full slice, and then uses the three-dimensional tensor as a sample and the slice category as a label to train a convolutional neural network (CNN) network to achieve full slice classification, such as Figure 1 This method can effectively alleviate the problem of unequal number of image blocks in strategy 1), but this strategy is limited by the CNN model and has poor adaptability to the pixel resolution and aspect ratio of the whole slice, making it difficult to meet the needs of practical applications.

[0008] 3) Classification strategy based on key image block sampling

[0009] After completing the classification of image blocks, this strategy samples the image blocks according to certain rules (for example, selecting image blocks with classification confidence higher than the threshold T) to reduce the number of image blocks that have little effect on decision-making or have side effects; then, relying on the sampled image block set, a set classification model such as multi-instance learning is established to achieve the classification of the entire slice. Compared with strategy 2), this method increases the adaptability to the slice size, but the absolute position information of the image block in the entire slice and the relative position information between the image blocks are discarded during the sampling process, resulting in a decrease in classification accuracy.

[0010] In order to solve the problems existing in the existing whole-slice classification strategy, how to provide a method that can effectively screen gastrointestinal biopsy pathology and improve classification accuracy, and automatically analyze gastrointestinal biopsy pathology slices based on the computer automatic classification method of digital pathology whole-slice images has become an urgent problem that technical personnel in this field need to solve. Summary of the invention

[0011] In view of this, the present invention provides an automatic analysis method for digestive tract biopsy pathology sections, which can be used for digestive tract biopsy pathology screening and automatically output diagnosis conclusions to assist doctors in completing their work accurately and efficiently.

[0012] In order to achieve the above object, the present invention adopts the following technical solution:

[0013] A method for automatically analyzing digestive tract biopsy pathological sections, comprising:

[0014] Dividing the pathological slice image into multiple sub-regions, and diagnosing each of the sub-regions one by one;

[0015] Determine the current diagnosis sub-region a, extract features of the current diagnosis sub-region a based on the pre-built feature extraction model and the sliding window method, and obtain the feature matrix F (a) ;

[0016] Based on the pre-built image classification model, the feature matrix F (a) Classify each feature vector in and get the lesion probability matrix P (a) ;

[0017] Based on the distance transformation algorithm, the feature matrix F (a) The distance quantization is performed on each feature vector in to obtain the distance quantization feature matrix H (a) ; The feature matrix F (a) and the distance quantization feature matrix H (a) Cascade to obtain the fusion feature matrix

[0018] Based on the fusion feature matrix and the lesion probability matrix P (a) , generate the organizational structure diagram G of the current diagnosis sub-area a (a) ;

[0019] The organizational structure diagram G of the current diagnosis sub-area a is obtained based on the pre-trained graph convolutional network model. (a) Make classification predictions;

[0020] Based on the classification prediction results, generate the diagnosis result c of the current diagnosis sub-region a (a) .

[0021] Preferably, in the above-mentioned method for automatic analysis of digestive tract biopsy pathological sections, dividing the pathological section image into a plurality of sub-regions and diagnosing each of the sub-regions one by one includes:

[0022] Convert the pathological slice image from RGB three-channel image to grayscale image;

[0023] Processing the grayscale image using a threshold segmentation method to obtain a tissue region binary template M;

[0024] Performing a closing operation on the tissue region binary template M to obtain a closing operation result;

[0025] The closed operation result is subjected to connected area detection, and a rectangular area is intercepted from the tissue area binary template M with the circumscribed rectangle of each connected area as the boundary to obtain a sub-area binary template M (a) ;

[0026] Each of the sub-region binary templates M is (a) Make a diagnosis.

[0027] Preferably, in the above-mentioned method for automatic analysis of digestive tract biopsy pathological sections, the feature matrix F (a) The representation is:

[0028]

[0029] In the above formula, Indicates that the length of the window corresponding to the i-th row and j-th column is d f The feature vector of [*] indicates rounding down; when the window of row i and column j does not contain the current diagnosis sub-region, no feature extraction is performed on the window, and F is directly assigned ij =0.

[0030] Preferably, in the above-mentioned method for automatic analysis of digestive tract biopsy pathological sections, the lesion probability matrix P (a) The representation is:

[0031]

[0032] In the above formula, C represents the number of lesion types involved in the automatic classification task; Represents the predicted probability of the image block in the window of row i and column j for C lesion types.

[0033] Preferably, in the above-mentioned method for automatic analysis of digestive tract biopsy pathological sections, the distance transformation algorithm is used to perform distance quantization on each feature vector in the feature matrix to obtain a distance quantization feature matrix; and the feature matrix and the distance quantization feature matrix are cascaded to obtain a fusion feature matrix, including:

[0034] Based on the distance transform method, the feature vector F of any image block in the current diagnosis sub-region is obtained. ij In the feature matrix F (a) The shortest coordinate distance d from the zero vector or the boundary ij ;

[0035] Use the following formula to calculate d ij Perform a distance transform:

[0036]

[0037] In the above formula, Indicates the degree to which the current image block is close to the boundary of the tissue region, τ represents the temperature coefficient, which is set according to the actual application effect, and τ=16;

[0038] Use the following formula Perform distance quantization encoding;

[0039]

[0040]

[0041] In the above formula, d h Indicates the length of the quantization code; H ij represents the distance quantized feature vector; h ijk Indicates H ij The value of the kth element in ;

[0042] The feature vectors F of each image block in the current diagnosis sub-region are respectively ij and the distance quantized feature vector H ij Cascade to obtain the fusion feature vector of each image block Its length is d = d f +d h ;

[0043] Fusion feature vector based on each image block Constructing the fusion feature matrix

[0044] Preferably, in the above-mentioned method for automatic analysis of digestive tract biopsy pathological sections, the fusion feature matrix and the lesion probability matrix P (a) , generate the organizational structure diagram G of the current diagnosis sub-area a (a) ,include:

[0045] The fusion feature matrix of the current diagnosis sub-region is calculated using the following formula: Conduct network sampling;

[0046]

[0047] Among them, χ grid represents the network sampling set; s is the sampling step, which is based on the number of image blocks N contained in the current diagnosis sub-region and the upper limit of the expected number of image blocks after grid sampling N max Obtained by calculation; |*| indicates the length of the set;

[0048] Get the lesion probability matrix P of the current diagnosis sub-region (a) The top N with the highest probability of lesions conf image blocks and construct a confidence sampling set; the expression of the confidence sampling set is as follows:

[0049]

[0050] Among them, χ conf represents the confidence sampling set; α represents the confidence sampling threshold; Represents a pre-built image classification model;

[0051] The network sampling set χ grid and the confidence sampling set χ conf Perform the union to obtain the set χ, which can be expressed as follows:

[0052]

[0053] Among them, N g =|χ|, which indicates the length of the set χ;

[0054] Construct the adjacency matrix A using the following formula;

[0055]

[0056]

[0057]

[0058] a pq represents any element in the adjacency matrix A; Represents the fused feature vector and In the original feature matrix Euclidean distance in coordinate space; (i p ,j p ) and (i q ,j q )express and In the original feature matrix The coordinates in ;

[0059] Construct the organizational chart G of the current diagnosis sub-area a (a) , G (a) =(A,X).

[0060] Preferably, in the above-mentioned method for automatic analysis of digestive tract biopsy pathological sections, the calculation formula of the sampling step length S is as follows:

[0061]

[0062]

[0063] In the above formula, N represents the number of image blocks contained in the current diagnosis sub-region; N max Indicates the upper limit of the expected number of image blocks after grid sampling;

[0064] The calculation process of the confidence sampling threshold α is:

[0065] The prediction probability P of the sliding window image block in the i-th row and j-th column on C lesion types is calculated using the following formula: ij and the disease probability r ij ;

[0066] P ij =[p ij1 ,p ij2 ,...,p ijC ];

[0067] r ij =1-p ij1 ;

[0068] The lesion probability of all image blocks obtained by sliding the window {r ij} Sort by large to small to get the set {r' 1 ,r' 2 ,...};

[0069] According to the lesion probability set {r' 1 ,r'2 ,...}Get the confidence sampling threshold α,

[0070] Preferably, in the above-mentioned method for automatic analysis of digestive tract biopsy pathological sections, the diagnosis result c of the current diagnosis sub-region a is obtained based on the classification prediction result. (a) Afterwards, the method further includes: generating the diagnosis results of other sub-regions one by one, and arranging and outputting the diagnosis results of each of the sub-regions according to the preset rules; the diagnosis result c of the current diagnosis sub-region a (a) The diagnostic process is:

[0071] Build a graph convolutional network model based on DiffPool;

[0072] Obtain a training sample set, and preprocess the training sample set using a sliding window method; the preprocessed training sample set The representation method is as follows:

[0073]

[0074] Among them, G k represents the tissue region map of the kth pathological section in the training sample set; l k Represents G k The corresponding slice label;

[0075] Using the preprocessed training sample set to train the graph convolutional network model;

[0076] The organizational structure diagram of the current diagnosis sub-area a is classified based on the trained graph convolutional network model, and its expression is as follows:

[0077]

[0078] Among them, z (a) ∈(0,1) C , represents the probability that the organizational chart G is divided into C categories; Represents the trained graph convolutional network model;

[0079] The diagnostic result c of the current diagnostic sub-area a is generated using the following formula: (a) ;

[0080] c (a) = argmax(z (a) ).

[0081] 10. Preferably, in the above-mentioned method for automatic analysis of digestive tract biopsy pathological sections, the pre-trained graph convolutional network model is used to classify and predict the tissue structure diagram of the current diagnosis sub-region a, and obtain the diagnosis result c of the current diagnosis sub-region a. (a) , also includes:

[0082] Based on the lesion probability matrix P (a) The probability value in is used to generate a heat map of each sub-region, and the heat map is superimposed on the original image surface of the corresponding sub-region to obtain a diagnostic area map; the process of generating the heat map is:

[0083] If c (a) >1, extract the lesion probability matrix P (a) The c (a) channels as heat maps, if c (a) =1, indicating that there is no lesion in the corresponding sub-region and no heat map is output.

[0084] Preferably, in the above-mentioned method for automatic analysis of digestive tract biopsy pathological sections, it also includes:

[0085] Arrange and output the diagnosis results and / or diagnosis area map of each sub-region according to preset rules; Use the following formula to express the diagnosis results of each sub-region;

[0086]

[0087] In the above formula, n a Indicates the number of sub-regions contained in the pathological slice image;

[0088] The diagnostic result of the pathological slice image is expressed by the following formula:

[0089]

[0090]

[0091] Z c represents the c-th row of Z; Represents the probability that the pathological section image belongs to each category.

[0092] It can be seen from the above technical solutions that, compared with the prior art, the present invention discloses a method for automatically analyzing digestive tract biopsy pathological sections, which has the following beneficial effects:

[0093] The present invention can classify digestive tract biopsy pathologies and can also present the lesion area in the form of a heat map in the full slice image to assist doctors in diagnosing the disease.

[0094] The present invention utilizes a combination of grid sampling and confidence sampling to balance the amount of information in tissue regions of different sizes, and the method has a wider range of applicability.

[0095] The present invention introduces the quantitative distance of tissue boundaries in the classification of tissue regions to represent the absolute position of features in the tissue, uses the tissue structure diagram to describe the spatial relative position of the features, and uses the graph convolutional network model to integrate the above information to finally complete the classification of tissue regions. The introduction of absolute position and relative position information enables the method of the present invention to significantly improve the classification accuracy of certain lesion types that rely on different tissue morphological distributions for diagnosis. BRIEF DESCRIPTION OF THE DRAWINGS

[0096] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying creative work.

[0097] Figure 1 The accompanying drawing is a flowchart of the method for automatically analyzing digestive tract biopsy pathological sections provided by the present invention;

[0098] Figure 2 The accompanying drawing is an application flow chart of the method for automatically analyzing digestive tract biopsy pathological sections provided by the present invention;

[0099] Figure 3 The accompanying drawing is a flow chart of establishing a full-slice tissue structure diagram provided by the present invention;

[0100] Figure 4 The accompanying drawings are a schematic diagram of the pathological section area, doctor's annotation results and lesion area model of a certain pathological section in the training sample set provided by the present invention. DETAILED DESCRIPTION

[0101] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0102] like Figure 1-3 As shown, the embodiment of the present invention discloses a method for automatically analyzing digestive tract biopsy pathological sections, comprising the following steps:

[0103] S1, divide the pathological slice image into multiple sub-regions, such as Figure 2As shown in b, each sub-area is diagnosed one by one;

[0104] S2, determine the current diagnosis sub-region a, based on the pre-built feature extraction model, and use the sliding window method to extract the current diagnosis sub-region a (such as Figure 2 c) to extract features and obtain the feature matrix F (a) ,like Figure 2 As shown in d;

[0105] S3, based on the pre-built image classification model, the feature matrix F (a) Classify each feature vector in and get the lesion probability matrix P (a) ,like Figure 2 As shown in e;

[0106] S4, based on the distance transformation algorithm to the feature matrix F (a) The distance quantization is performed on each feature vector in to obtain the distance quantization feature matrix H (a) ;,like Figure 2 f; the feature matrix F (a) And the distance quantization feature matrix H (a) Cascade to obtain the fusion feature matrix

[0107] S5. Based on fusion feature matrix and the lesion probability matrix P (a) , generate the organizational structure diagram G of the current diagnosis sub-area a (a) ,like Figure 2 As shown in h;

[0108] S6. Organizational structure diagram G of the current diagnosis sub-area a based on the pre-trained graph convolutional network model (a) Make classification predictions;

[0109] S7. Based on the classification prediction result, the diagnosis result c of the current diagnosis sub-region a is obtained. (a) ,like Figure 2 As shown in i.

[0110] S8. Repeat S2-S6 to generate diagnosis results for other sub-regions one by one.

[0111] S9. Arrange and output the diagnosis results of each sub-region according to preset rules.

[0112] More advantageously, in order to assist doctors in quickly locating the lesion area, it also includes:

[0113] S10, based on the probability value of the lesion probability matrix of each sub-region, generate a heat map of each sub-region one by one, and superimpose the heat map on the original image surface of the corresponding sub-region to obtain a diagnostic area map of each sub-region, such as Figure 2 j; the process of generating the heat map is:

[0114] If c (a) >1, extract the lesion probability matrix P (a) The c (a) channels as heat maps, if c (a) =1, indicating that there is no lesion in the current diagnosis sub-region and no heat map is output.

[0115] The diagnostic results and diagnostic area map of the entire pathological section image are sorted and output according to preset rules.

[0116] The following is a detailed description of each of the above steps:

[0117] S1. Divide the pathological slice image into multiple sub-regions and diagnose each sub-region one by one, including:

[0118] When preparing biopsy pathology slides, it is common to place multiple layers of tissue side by side on the same glass slide (e.g. Figure 2 a), the embodiment of the present invention needs to predict all sub-regions one by one. Specifically including:

[0119] S111, converting the pathological slice image I from an RGB three-channel image to a grayscale image;

[0120] S112, using the threshold segmentation method (Otsu threshold segmentation algorithm) to process the grayscale image to obtain the tissue region binary template M, such as Figure 2 As shown in b;

[0121] S113, performing a closing operation on the tissue region binary template M to obtain a closing operation result;

[0122] S114, perform connected area detection on the closed operation result, and cut out a rectangular area from the tissue area binary template M with the circumscribed rectangle of each connected area as the boundary to obtain a sub-area binary template, such as Figure 2 c or 3b; define the binary template of the ath sub-region as M (a) , and define the corresponding area in the original image I as I (a) ,like Figure 2 As shown in the rectangular box in a;

[0123] S115: Binary template M of each sub-region is mapped one by one (a) Make a diagnosis.

[0124] The construction process of the feature extraction model in S2 and the image classification model in S3 is as follows:

[0125] 1. Slice annotation and data collection

[0126] like Figure 4 As shown, the method of the present invention relies on the annotation of pathologists, requiring pathologists to outline the typical lesion area in the original slice image, and then convert the area outlined by the doctor into a lesion area template through a closed curve filling algorithm. Represents a RGB three-channel pathological full-slice image with a pixel resolution of w×h at a high microscopic magnification (such as a resolution of 0.46um / pixel). After annotating K slices, a training dataset is established Among them I k represents the kth slice image in the training set, E k Represents the generated lesion area template, whose size is the same as that of image I k The same is used to record the lesion category number corresponding to each pixel in the image, l k Category number indicating the overall diagnosis of the annotated slice.

[0127] At the same time, the training set needs to be Specifically, the slices are divided into image blocks in the form of sliding windows, and the window size is defined as t×t and the sliding window step is t / 2. It represents the nth image block obtained by sliding the window. The submatrix of this image block in the same area of ​​the lesion area template is Y n , the image block dataset established by sliding window is expressed as where y n The value of Y n The values ​​are determined by majority vote.

[0128] 2. Model building

[0129] Training set of image patches Any image block T in n The extracted image features can be expressed as the formula:

[0130]

[0131] in represents the feature extraction model, Indicates the length is d f The present invention has no special restrictions on the feature extraction model, and digital image feature extraction methods can be selected as needed, such as traditional image features such as texture features, color histogram features, shape features, frequency domain transformation features, or feature extraction methods based on machine learning such as autoencoder networks and convolutional neural networks.

[0132] In the above image feature vector f n Based on this, an image classification model is established to classify the lesion category of the image block. The expression of the image classification model is as follows:

[0133]

[0134] in represents the image classification model, p n ∈(0,1) C represents the predicted probability, C represents the number of lesion types involved in the automatic classification task, and p nc represents the probability that the nth image block belongs to the cth class and satisfies For the convenience of subsequent description, c=1 is specified to represent the normal area category, and c>1 is specified to represent the lesion category. There are no special requirements. You can choose support vector machine, random forest, or neural network-based classification model as needed.

[0135] In order to achieve higher automatic classification accuracy, the present invention uses convolutional neural network (CNN) to establish the above and CNN is an end-to-end machine learning model that requires the use of the above Complete the CNN training. After the training is completed, the last fully connected-Softmax structure in the CNN structure is used as the classification model The entire network except the classification layer is used as a feature extraction model

[0136] The feature matrix F of the current diagnosis sub-region a in S2 (a) The construction process is:

[0137] Using the above feature extraction model The image features of the current diagnostic sub-region are extracted with the sliding window method, where the window size is t×t, which is the same as the window size in the training set preprocessing process, and the sliding window step is t. The feature matrix of the whole slice obtained by feature extraction is expressed as Where [*] means round down. Matrix elements That is, the feature corresponding to the window of row i and column j. In order to avoid unnecessary calculations, when the window of row i and column j does not contain tissue areas (according to Figure 3 b) when judging the tissue area template, no feature extraction is performed on the window, and F is directly assigned ij =0.

[0138] In S3, the lesion probability matrix P (a) The specific construction process is:

[0139] The above image classification model Acts on F (a) Any F ij , obtain the lesion probability matrix of the current diagnosis sub-region a The matrix elements That is, the predicted probability of the image block in the window of the i-th row and j-th column on C types of lesions, such as Figure 3 D

[0140] The distance transformation process of the current diagnosis sub-region a in S4 is:

[0141] In order to describe the position of the local tissue area in the tissue block, the present invention adds the tissue boundary distance quantization feature in the image block feature extraction link, and uses the distance transformation to obtain any feature F ij In the feature matrix F (a) The shortest coordinate distance from the zero vector or the boundary, expressed as d ij The implementation effect is as follows Figure 3 As shown in e.

[0142] In order to make the distance coding more sensitive to the boundary of the tissue region (corresponding to the possible epithelial part in tissue pathology), the following formula is used to calculate d ij Implementing the transformation

[0143]

[0144] In the above formula, Indicates the degree to which the current image block is close to the boundary of the tissue region, τ indicates the temperature coefficient, which is set according to the actual application effect. For digestive tract biopsy pathological sections, τ=16 is preferably taken;

[0145] In order to make the subsequent machine learning model better use this feature, the following formula is used Distance quantization coding

[0146]

[0147]

[0148] Among them, d h Indicates the length of the quantization code, H ij The value of the kth element in H ij represents the distance quantized feature vector, h ijk Indicates H ij The value of the kth element in .

[0149] For ease of description, the distance quantization encoding process of the image block at the window of the i-th row and j-th column is recorded as:

[0150]

[0151] In order to avoid unnecessary calculations, when the window in the i-th row and j-th column does not contain a tissue region (judged according to the tissue region template shown in FIG3b ), the distance quantization is not performed on the window, and H is directly assigned. ij =0.

[0152] The feature vectors F of each image block in the current diagnosis sub-region are respectively ij and the distance quantized feature vector H ij Cascade to obtain the fusion feature vector of each image block Its length is d = d f +d h ;

[0153] Fusion feature vector based on each image block Constructing the fusion feature matrix

[0154] The specific process of generating an organizational chart in S5 is as follows:

[0155] Considering that the size of tissue regions in pathological images varies greatly, in extreme cases, the tissue region may contain hundreds of thousands of image blocks (windows). Directly using all image block features to construct the graph will result in the graph being too large, affecting the calculation of the subsequent graph convolutional network model. Therefore, the present invention balances the number of image blocks used for graph construction in different tissue slices by sampling.

[0156] 1) First, grid sampling is performed, and the fusion feature matrix of the current diagnosis sub-region is calculated using the following formula: Conduct network sampling;

[0157]

[0158] Among them, χ grid represents the network sampling set; s is the sampling step, which is based on the number of image blocks N contained in the current diagnosis sub-region and the upper limit of the expected number of image blocks after grid sampling N max Obtained by calculation; |*| indicates the length of the set.

[0159] The calculation formula of sampling step length S is as follows:

[0160]

[0161]

[0162] In the above formula, N represents the number of image blocks contained in the current diagnosis sub-region; N maxIndicates the upper limit of the expected number of image blocks after grid sampling.

[0163] 2) χ grid Only the number of image blocks in the spatial structure is considered to be reduced. In order to ensure that the image blocks that play a decisive role in classification in the slice can participate in the composition, the embodiment of the present invention additionally adopts confidence sampling to obtain the lesion probability matrix P of the current diagnosis sub-region (a) The top N with the highest probability of lesions conf image patches and build a confidence sampling set.

[0164] For ease of description, P ij Detailed expression is P ij =[p ij1 ,p ij2 ,...,p ijC ], the lesion probability of the image block in row i and column j obtained by sliding the window is defined as r ij =1-p ij1 , the lesion probability of all image blocks obtained by sliding the window {r ij} Sort from large to small to get the set {r' 1 ,r' 2 ,...}, and the confidence sampling threshold is defined accordingly According to the threshold α, the confidence sampling result is defined as:

[0165]

[0166] Among them, χ conf represents the confidence sampling set; α represents the confidence sampling threshold; Represents a pre-built image classification model.

[0167] The network sampling set χ grid and the confidence sampling set χ conf Perform the union to obtain the set χ, which can be expressed as follows:

[0168]

[0169] Among them, N g =|χ|, indicating the length of the set χ.

[0170] Define the adjacency matrix A, The element a pq The value selection rules are as follows:

[0171]

[0172]

[0173] apq represents any element in the adjacency matrix A; Represents the fused feature vector and In the original feature matrix Euclidean distance in coordinate space; (i p ,j p ) and (i q ,j q )express and In the original feature matrix The coordinates in .

[0174] In summary, the construction process of the organizational chart is completed, denoted as G (a) =(A,X), such as Figure 3 (g) as shown.

[0175] For ease of expression, the feature matrix The whole process of constructing the organizational regional structure diagram with the prediction probability matrix P is defined as the graph construction model The composition process is recorded as

[0176] Training and classification prediction of graph convolutional network model in S6.

[0177] Building a graph convolutional network model for graph classification prediction It is used to classify the pathological type of the tissue region map G. When the present invention is used for digestive tract biopsy pathological sections, the DiffPool model is preferably used. The diagnosis result c of the current diagnosis sub-region a (a) The diagnostic process is:

[0178] Build a graph convolutional network model based on DiffPool;

[0179] Obtain a training sample set and preprocess the training sample set using a sliding window method; the preprocessed training sample set The representation method is as follows:

[0180]

[0181] Among them, G k represents the tissue region map of the kth pathological section in the training sample set; l k Represents G k The corresponding slice label;

[0182] Use the preprocessed training sample set to train the graph convolutional network model;

[0183] Based on the trained graph convolutional network model, the organizational structure diagram of the current diagnosis sub-area a is classified, and its expression is as follows:

[0184]

[0185] Among them, z (a) ∈(0,1) C , represents the probability that the organizational chart G is divided into C categories; Represents the trained graph convolutional network model;

[0186] The diagnostic result c of the current diagnostic sub-area a is generated using the following formula: (a) ;

[0187] c (a) = argmax(z (a) ).

[0188] According to S2-S6, the diagnosis results of each sub-region are generated one by one.

[0189] S10. Output case diagnosis results

[0190] The prediction results z of all sub-regions (sometimes distributed in multiple full slices) of the screening pathology slice image are (a) Arrange as a matrix where n a represents the number of sub-regions contained in the case, then the screening result of the entire case is defined as in

[0191]

[0192] Where Z c represents the c-th row of Z. That is, the probability that a case belongs to each category, which can be output to doctor users according to certain rules based on application needs.

[0193] In this specification, each embodiment is described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the embodiments can be referred to each other. For the device disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the method part.

[0194] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but will conform to the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A method for automatically analyzing digestive tract biopsy pathological sections, characterized in that: include: Dividing the pathological slice image into multiple sub-regions, and diagnosing each of the sub-regions one by one; Determine the current diagnosis sub-region a, extract features of the current diagnosis sub-region a based on the pre-built feature extraction model and the sliding window method, and obtain the feature matrix F (a) ; Based on the pre-built image classification model, the feature matrix F (a) Classify each feature vector in and get the lesion probability matrix P (a) ; Based on the distance transformation algorithm, the feature matrix F (a) The distance quantization is performed on each feature vector in to obtain the distance quantization feature matrix H (a) ; The feature matrix F (a) and the distance quantization feature matrix H (a) Cascade to obtain the fusion feature matrix Based on the fusion feature matrix and the lesion probability matrix P (a) , generate the organizational structure diagram G of the current diagnosis sub-area a (a) ; The organizational structure diagram G of the current diagnosis sub-area a is obtained based on the pre-trained graph convolutional network model. (a) Make classification predictions; Based on the classification prediction results, generate the diagnosis result c of the current diagnosis sub-region a (a) ; The feature matrix F is transformed based on the distance transformation algorithm. (a) The distance quantization is performed on each feature vector in to obtain the distance quantization feature matrix H (a) ; The feature matrix F (a) and the distance quantization feature matrix H (a) Cascade to obtain the fusion feature matrix include: Based on the distance transform method, the feature vector F of any image block in the current diagnosis sub-region is obtained. ij In the feature matrix F (a) The shortest coordinate distance d from the zero vector or the boundary ij ; Indicates that the length of the window corresponding to the i-th row and j-th column is d f The eigenvector of Use the following formula to calculate d ij Perform a distance transform: In the above formula, Indicates the degree to which the current image block is close to the boundary of the tissue region, τ represents the temperature coefficient, which is set according to the actual application effect, and τ=16; Use the following formula Perform distance quantization encoding; In the above formula, d h Indicates the length of the quantization code; H ij represents the distance quantized feature vector; h ijk Indicates H ij The value of the kth element in ; The feature vectors F of each image block in the current diagnosis sub-region are respectively ij and the distance quantized feature vector H ij Cascade to obtain the fusion feature vector of each image block Its length is d = d f +d h ; Fusion feature vector based on each image block Constructing the fusion feature matrix 2. The method for automatically analyzing digestive tract biopsy pathological sections according to claim 1, characterized in that: The step of dividing the pathological slice image into a plurality of sub-regions and diagnosing each of the sub-regions one by one includes: Convert the pathological slice image from RGB three-channel image to grayscale image; Processing the grayscale image using a threshold segmentation method to obtain a tissue region binary template M; Performing a closing operation on the tissue region binary template M to obtain a closing operation result; The closed operation result is subjected to connected area detection, and a rectangular area is intercepted from the tissue area binary template M with the circumscribed rectangle of each connected area as the boundary to obtain a sub-area binary template M (a) ; Each of the sub-region binary templates M is (a) Make a diagnosis.

3. The method for automatically analyzing digestive tract biopsy pathological sections according to claim 1, characterized in that: The feature matrix F (a) The representation is: In the above formula, [*] indicates rounding down; w×h indicates pixel resolution; t×t indicates window size; when the window in the i-th row and j-th column does not contain the tissue area of ​​the current diagnosis sub-area, feature extraction is not performed on the window, and F is directly assigned ij =0.

4. The method for automatically analyzing digestive tract biopsy pathological sections according to claim 3, characterized in that: The lesion probability matrix P (a) The representation is: In the above formula, C represents the number of lesion types involved in the automatic classification task; Represents the predicted probability of the image block in the window of row i and column j for C lesion types.

5. The method for automatically analyzing digestive tract biopsy pathological sections according to claim 4, characterized in that: Based on the fusion feature matrix and the lesion probability matrix P (a) , generate the organizational structure diagram G of the current diagnosis sub-area a (a) ,include: The fusion feature matrix of the current diagnosis sub-region is calculated using the following formula: Conduct network sampling; in, represents the network sampling set; s is the sampling step, which is based on the number of image blocks N contained in the current diagnosis sub-region and the upper limit of the expected number of image blocks after grid sampling N max Obtained by calculation; |*| indicates the length of the set; Get the lesion probability matrix P of the current diagnosis sub-region (a) The top N with the highest probability of lesions conf image blocks and construct a confidence sampling set; the expression of the confidence sampling set is as follows: in, represents the confidence sampling set; α represents the confidence sampling threshold; Represents a pre-built image classification model; The network sampling set With confidence sampling set Perform the union to obtain set X, which can be expressed as follows: Among them, N g =|X|, which indicates the length of set X; Construct the adjacency matrix A using the following formula; a pq represents any element in the adjacency matrix A; Represents the fused feature vector and In the original feature matrix Euclidean distance in coordinate space; (i p ,j p ) and (i q ,j q )express and In the original feature matrix The coordinates in ; Construct the organizational chart G of the current diagnosis sub-area a (a) , G (a) =(A,X).

6. The method for automatically analyzing digestive tract biopsy pathological sections according to claim 5, characterized in that: The calculation formula of sampling step length S is as follows: In the above formula, N represents the number of image blocks contained in the current diagnosis sub-region; N max Indicates the upper limit of the expected number of image blocks after grid sampling; The calculation process of the confidence sampling threshold α is: The prediction probability P of the sliding window image block in the i-th row and j-th column on C lesion types is calculated using the following formula: ij and the disease probability r ij ; P ij =[p ij1 ,p ij2 ,...,p ijC ]; r ij =1-p ij1 ; The lesion probability of all image blocks obtained by sliding the window {r ij } Sort from largest to smallest to obtain the set {r'1,r'2,...}; According to the lesion probability set {r'1, r'2, ...}, the confidence sampling threshold α is obtained.

7. The method for automatically analyzing digestive tract biopsy pathological sections according to claim 5, characterized in that: The diagnosis result c of the current diagnosis sub-area a (a) The diagnostic process is: Build a graph convolutional network model based on DiffPool; Obtain a training sample set, and preprocess the training sample set using a sliding window method; the preprocessed training sample set The representation method is as follows: Among them, G k represents the tissue region map of the kth pathological section in the training sample set; l k Represents G k Corresponding slice labels; K means that the training sample set contains a total of K pathological slices; Using the preprocessed training sample set to train the graph convolutional network model; The organizational structure diagram of the current diagnosis sub-area a is classified based on the trained graph convolutional network model, and its expression is as follows: Among them, z (a) ∈(0,1) C , represents the probability that the organizational chart G is divided into C categories; Represents the trained graph convolutional network model; The diagnostic result c of the current diagnostic sub-area a is generated using the following formula: (a) ; c (a) =argmax(z (a) )。 8. The method for automatically analyzing digestive tract biopsy pathological sections according to claim 5, characterized in that: The pre-trained graph convolutional network model is used to classify and predict the tissue structure diagram of the current diagnosis sub-region a, and obtain the diagnosis result c of the current diagnosis sub-region a. (a) , also includes: Based on the probability values ​​in the lesion probability matrix of each sub-region, a heat map of each sub-region is generated one by one, and the heat map is superimposed on the original image surface of the corresponding sub-region to obtain a diagnostic area map of each sub-region; the process of generating the heat map is: If c (a) >1, extract the lesion probability matrix P (a) The c (a) channels as heat maps, if c (a) =1, indicating that there is no lesion in the corresponding sub-region and no heat map is output.

9. The method for automatically analyzing digestive tract biopsy pathological sections according to claim 8, characterized in that: Also includes: Arrange and output the diagnosis results and / or diagnosis area map of each sub-region according to preset rules; use the following formula to express the diagnosis results of each sub-region; In the above formula, n a Indicates the number of sub-regions contained in the pathological slice image; The diagnostic result of the pathological slice image is expressed by the following formula: Z c represents the c-th row of Z; Represents the probability that the pathological section image belongs to each category.

Citation Information

Patent Citations

  • Gastroscope pathological image classification method based on weak supervised learning

    CN111985536A

  • Classification apparatus for pathologic diagnosis of medical image, and pathologic diagnosis system using the same

    US20170236271A1