Method and system for multi-resolution image reconstruction of pathological sections based on sparse sampling
By combining sparse sampling and multi-resolution reconstruction methods with conditional vector coding, the problems of long scanning time and high storage cost in the digitization of pathological slides are solved, and rapid, high-quality digitization of pathological slides is achieved.
Patent Information
- Application Number
- CN202211375557.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-04
- Publication Date
- 2026-01-02
- Estimated Expiration
- 2042-11-04
AI Technical Summary
Existing digitization technologies for pathological slides suffer from problems such as long scanning times, high storage costs, and low reconstruction quality, failing to simultaneously meet the requirements for real-time performance and accuracy.
Low-resolution images of pathological sections are acquired using a low-magnification imaging system, and high-resolution image patches of a small number of representative grids are obtained through sparse sampling. By combining conditional vector encoding of information on the source and preparation method of the pathological sections, multi-resolution reconstruction methods are used to improve image quality.
This approach improves the speed of digitizing pathological slides while reducing storage requirements, and simultaneously enhances the accuracy of high-resolution images, meeting the real-time and accuracy requirements for pathological slide digitization.
Smart Images

Figure CN115908283B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of image processing, and in particular to a pathological section multi-resolution image reconstruction method and system based on sparse sampling. BACKGROUND
[0002] The digitization of pathological sections refers to scanning the sections with a high magnification optical scanner. The specific process is as follows: the pathological section is regarded as being composed of multiple grids, a high magnification optical lens is focused on and magnifies the pathological tissue in each grid, and after scanning, sampling and quantization, the image of each grid is obtained. The images of all the grids are spliced to obtain the digital pathological image of the pathological section. The main defects of the current digitization technology are as follows: first, due to the different thicknesses of pathological tissues and the small depth of field of high magnification lenses, focusing is required for each grid, and thus the scanning time of a single section is as long as 5 minutes. Since each patient has multiple pathological sections, the digitization of the pathological sections of a single patient takes tens of minutes. Second, the number of grids of pathological sections is large, and the number of pixels of each high magnification image block is large, so the number of pixels of a single pathological section image is extremely large, and the storage capacity of a single image can be as high as 5 GB. A hospital performs pathological diagnosis on several thousand to tens of thousands of people each year, and considering the requirement of the state that pathological information needs to be saved for 20 years, the storage capacity of the pathological images generated each year is large, and a large amount of storage cost needs to be spent. Therefore, high magnification imaging cannot guarantee the real-time performance of pathological section digitization and the data storage capacity is huge.
[0003] In view of the high time cost and storage cost of the existing pathological section digitization technology, which is difficult for hospitals to accept. A solution is to scan the pathological sections with a low magnification lens and save low resolution images; and to reconstruct the low resolution images into high resolution images by using a super-resolution image processing algorithm during diagnosis. Since the depth of field of the low magnification lens is large, the generated image is small, and thus the technical bottlenecks of slow scanning speed caused by multiple focusing and large image storage capacity can be solved. It is considered that the super-resolution image processing algorithm for quickly obtaining high magnification images with quality comparable to that of real scanning is a promising digitization technology.
[0004] However, the low-resolution image loses a large amount of high-resolution prior information, and the existing super-resolution image processing algorithm can only approximately reconstruct when the high-resolution prior information is lost, and it is difficult to completely reconstruct a high-resolution image with quality comparable to that of a real high-magnification scanning image. Common problems of the reconstructed high-resolution image include color deviation and deformation of the internal structure of cells. Secondly, the images of pathological sections obtained by different parts and different preparation methods are quite different, and the existing super-resolution image processing algorithm does not distinguish the source of the section, so it cannot obtain high-quality reconstructed images for all sources of images when the prior information of the source of the pathological section is missing. In summary, the existing pathological section digitization method of high-magnification optical imaging has the problems of low efficiency and large storage capacity, and cannot meet the real-time requirement of pathological section digitization. The pathological section digitization method of low-magnification imaging and super-resolution reconstruction improves the imaging speed and reduces the storage capacity, but has the technical bottlenecks of low reconstruction quality, image color deviation, cell deformation, and inability to achieve high-quality image reconstruction for all sources of images, and cannot meet the accuracy requirement of pathological section digitization. SUMMARY
[0005] The present application provides a pathological section multi-resolution image reconstruction method and system based on sparse sampling, which uses a low-magnification imaging system to obtain a low-resolution image of a pathological section, and uses a high-magnification imaging system to sparsely obtain high-resolution image blocks of a small number of representative grids of the pathological section. By reducing the number of high-magnification imaging and establishing a multi-resolution reconstruction method of high-resolution images, the technical problem that the current pathological section digitization technology cannot simultaneously consider real-time and accuracy is solved.
[0006] To solve the above technical problems, the technical scheme provided by the present application is as follows:
[0007] A pathological section multi-resolution image reconstruction method based on sparse sampling, comprising the following steps:
[0008] obtaining a low-resolution image of a pathological section by using a low-magnification imaging system, and dividing the low-resolution image into a grid to obtain a set of grid image blocks L of the low-resolution image, wherein L 00 , L 01, ...L MN}, and M and N are the number of horizontal and vertical grid divisions, respectively;
[0009] dividing the low-resolution image into n non-overlapping image regions, each image region including a plurality of grid image blocks, wherein n << (M+1)(N+1), and (M+1)(N+1) is the number of grids;
[0010] for any kth region, wherein k≤n, the following steps are performed:
[0011] calculating effective information score of all grid in the kth region, taking the grid with the maximum effective information score as the representative grid of the region, and acquiring high resolution image block H k ;
[0012] calculating effective information score of all grid in the kth region, taking the grid with the maximum effective information score as the representative grid of the region, and acquiring high resolution image block H ij ; ij ; k ; ij ;
[0013] encoding the pathological section source and manufacturing method information with the condition vector, and implementing multi-resolution reconstruction of the high resolution image block of each grid according to the condition vector and the hidden vector; and splicing the high resolution image blocks of all grids reconstructed from the low resolution image to obtain the high resolution image of the pathological section reconstructed from the low resolution image.
[0014] Preferably, the hidden vector of the grid image block L ij is constructed by an encoder, and the encoder comprises a local encoding module, a prior encoding module and an information fusion module.
[0015] The local encoding module is used to extract low resolution local information hidden vector C ij from the grid image block L ij .
[0016] The prior encoding module is used to extract high resolution prior information hidden vector k from the high resolution image block H
[0017] The information fusion module fuses the low resolution local hidden vector C ij and the high resolution prior hidden vector to obtain a hidden vector v ij that fuses the low resolution local information and the high resolution prior information, and the length L of the hidden vector v ij is much smaller than the number of pixels in the grid; and the multi-resolution image reconstruction reconstructs the high resolution image by using the low resolution local information and the high resolution prior information.
[0018] Preferably, the condition vector encodes the pathological section source and manufacturing method information, and provides pathological section prior information of category encoding of the manufacturing method and category encoding of the human body part from which the sample is taken.
[0019] Preferably, the conditional vector of the pathological slice is obtained, and a high-resolution image patch of each grid is reconstructed based on the conditional vector of the pathological slice and the latent vector of each grid of the pathological slice, which is achieved by a decoder; the decoder is a conditional generation network, including a conditional embedding module and a decoding module.
[0020] The condition embedding module is used to embed the condition vector to obtain the condition embedding vector w;
[0021] The decoding module is used to determine the conditional embedding vector w and the latent vector v. ij Reconstructing high-resolution image block B ij .
[0022] Preferably, the training of the encoder and decoder specifically involves:
[0023] The decoder and encoder are trained separately, with the decoder trained first. After the decoder is trained, its weights remain unchanged, and the encoder is trained end-to-end, allowing the encoder to progress from L... ij and H k Generate the latent vector v ij Together with the conditional embedding vector w, it is input into the decoding module to generate a high-resolution image patch B. ij And B ij The generated high-resolution image patch has the greatest similarity to the real high-resolution image patch; the greatest similarity means that the pixel, texture, perceptual cost function and discriminator cost function of the generated high-resolution image patch and the real high-resolution image patch are minimized.
[0024] Preferably, the training of the decoder includes the following steps:
[0025] S61: Scan pathological slides using a high-magnification imaging system to obtain high-resolution image patches, and divide the image patches into several sets of training datasets according to different conditional vectors:
[0026] S62: Establish a generative adversarial network, where the generator is a decoding module and the discriminator is used to distinguish the authenticity of the generated image; randomly initialize the generator and discriminator.
[0027] S63: From any training dataset The training data is extracted from the data to train the conditional embedding module, the decoding module, and the discriminator, where t i This represents the conditional vector of the i-th training dataset. This represents a true high-resolution image patch in the i-th training dataset;
[0028] S63-1: Set the condition vector t i The input is fed into the conditional embedding module; the latent vectors are randomly initialized as noise, and the two are then concatenated.
[0029] S63-2: Input the connection vector into the decoding module and output high-resolution image block B. ij ;
[0030] S63-3: Input the conditional embedding vector into the discriminator, and use the discriminator to calculate the true high-resolution image patches. and output high-resolution image patch B ij The cost, as the first identification cost;
[0031] S63-4: With the goal of minimizing the first discrimination cost, the weights of the conditional embedding module, the decoding module, and the discriminator are updated using the backward gradient diffusion method.
[0032] S64: Take out the other training data and repeat S63.
[0033] Preferably, the training of the encoder includes the following steps:
[0034] S71: Establish a training dataset d = {d ij} and the set of real high-resolution images {H ij}, where d ij ={t ij ,L ij H k H ij}, L ij The grid image patch obtained by the low-magnification imaging system has coordinates i,j; t ij H is the conditional vector of the pathological slice to which the grid belongs; ij Grid image patch L acquired for high magnification imaging system ij True high-resolution image patches, H k For L ij A high-resolution image patch of a representative grid in the k-th region;
[0035] S72: Construct the network structure of the encoder and decoder, wherein the encoder includes a local coding module, a priori coding module, and an information fusion module; the decoder includes a conditional embedding module and a decoding module;
[0036] S73: Initialize the decoder with the trained decoder weights and keep the decoder weights unchanged during training. Randomly initialize the encoder, including the local coding module, the prior coding module, and the information fusion module.
[0037] S74: Retrieve training data d ij ={t ij ,L ij H k H ij};
[0038] S74-1: inputting the grid image block L ij into a local encoding module to obtain a low-resolution local information latent vector C ij ; inputting the high-resolution image block H k into a prior encoding module to obtain a high-resolution prior information latent vector
[0039] S74-2: inputting the local information latent vector C ij and the prior information latent vector into an information fusion module to obtain a latent vector v ij fusing the low-resolution local information and the high-resolution prior information;
[0040] S74-3: inputting the condition vector t ij into a condition embedding module to obtain a condition embedding vector w;
[0041] S74-4: inputting the condition embedding vector w and the latent vector v ij into a decoding module to generate a high-resolution image block B ij ;
[0042] S74-5: calculating pixel, texture and perception costs of the high-resolution image block B ij and the real high-resolution image block H ij , and adding the three costs with certain weights to obtain a first cost; calculating a cost of the high-resolution image block B ij and a real high-resolution image set {H ij} by the discriminator to obtain a second discriminator cost;
[0043] S74-6: updating weights of the local encoding module, the prior encoding module, the information fusion module and the discriminator by using a back propagation method, with a sum of the first cost and the second discriminator cost being a training target;
[0044] S75: replacing the training data and repeating S74.
[0045] Preferably, the effective information score is defined as a weighted sum of a tissue coverage score, a cell number score, a cell coverage score and a color distribution score; the tissue coverage score is defined as a percentage of a tissue coverage area in the grid; the cell number score is defined as a calculated function value of a cell number in the grid; the cell coverage score is defined as a percentage of a cell coverage area in the grid; the color distribution score is defined as a sum of information entropy of all pixel channel information in the grid; the representative grid is a grid with a maximum effective information score in the region; and the high-resolution image block H k of the representative grid comprises the following steps:
[0046] taking out any one grid L in the kth image area ij , judging whether there is L ij in the cache module as a representative grid high-resolution image block H k ; if there is, taking out the representative grid high-resolution image block; if there is not, selecting an image block L k of the grid with the largest valid information score from the image area, using a high-magnification imaging system to accurately focus on the tissue in the grid corresponding to L k , obtaining a representative grid high-resolution image H k , and saving H k to the cache module; the cache module is a memory for saving a set of high-resolution image blocks {H1, H2,...H n} of the current pathological section or historical pathological sections of the same source; when digitizing a pathological section of the same source, searching for a high-resolution image block similar to L ij from the cache module as H k , so as to further reduce the imaging times of high magnification and accelerate the imaging speed.
[0047] Preferably, the high-resolution image block of each grid is reconstructed according to the condition vector of the pathological section and the latent vector of each grid, and the method comprises the following steps:
[0048] The latent vectors of all the grids are combined to form a latent vector matrix, and the condition vector of the pathological section is stored as a pathological data file together, and the data capacity of the file is far less than that of a file storing image pixels;
[0049] The latent vector matrix and the condition vector in the pathological data file are read; and the high-resolution image block is reconstructed according to the condition vector of the pathological section and the latent vector of each grid by using a decoder;
[0050] The high-resolution image blocks are spliced to obtain a high-resolution image reconstructed by the low-resolution image.
[0051] A computer system comprises a memory, a processor, and a computer program stored on the memory and capable of running on the processor, and the processor implements the steps of the above method when executing the computer program.
[0052] The present application has the following beneficial effects:
[0053] 1. The application discloses a pathological section multi-resolution image reconstruction method and system based on sparse sampling, which acquires a low-resolution image of a pathological section by using a low magnification imaging system, acquires a high-resolution image block of a small number of representative grids of the pathological section by using a high magnification imaging system, thereby reducing the number of high magnification imaging, improving the digitization speed and reducing the data storage amount.
[0054] 2. In the preferred scheme, the application further reduces the number of high magnification imaging by reading the high-resolution image block of the current pathological section or the historical pathological section of the same source from the cache module, thereby improving the pathological section imaging speed.
[0055] In addition to the purposes, features and advantages described above, the application has other purposes, features and advantages. The application will be further described in detail below with reference to the accompanying drawings. BRIEF DESCRIPTION OF DRAWINGS
[0056] The accompanying drawings, which form a part of the present application, are intended to provide further understanding of the present application, and the illustrative embodiments of the present application and their description serve the purpose of explaining the present application. The present application is not limited by the improper interpretation of the accompanying drawings. In the drawings:
[0057] Figure 1 The flowchart of the pathological section multi-resolution image reconstruction method and system based on sparse sampling provided by the embodiment of the application;
[0058] Figure 2 The low magnification imaging and image segmentation flowchart of the pathological section provided by the embodiment of the application;
[0059] Figure 3 The high magnification imaging flowchart of the sparse sampling representative grid provided by the embodiment of the application;
[0060] Figure 4 The conditional vector schematic diagram provided by the embodiment of the application;
[0061] Figure 5 The working flowchart of the encoder provided by the embodiment of the application;
[0062] Figure 6is a workflow diagram of the decoder provided by an embodiment of the present application;
[0063] Figure 7 is a training process diagram of the decoder provided by an embodiment of the present application;
[0064] Figure 8 is a training process diagram of the encoder provided by an embodiment of the present application. DETAILED DESCRIPTION
[0065] Embodiments of the present application are described in detail below with reference to the accompanying drawings, but the present application can be implemented in various different ways as defined and covered by the claims.
[0066] Embodiment one:
[0067] As shown in the Figure 1 , the present application discloses a sparse sampling pathological section multi-resolution image reconstruction method and system, the main steps include:
[0068] 1. Low magnification imaging and image segmentation of pathological sections
[0069] The workflow of low magnification imaging and image segmentation of pathological sections is as shown in Figure 2 , including:
[0070] S1: magnify the human tissue on the pathological section with a low magnification optical lens to obtain an optical magnified image. Obtain the electrical signal of the magnified image with a sensor, and after analog-to-digital conversion, sampling and quantization, obtain the low-resolution image of the pathological section.
[0071] S2: divide the low-resolution image with a grid, and set the number of horizontal and vertical grids as M+1 and N+1 respectively. Each grid is an image block, and the set of grid image blocks is L={L 00 ,L 01 ,...,L MN}, wherein the number of pixels of each grid image block is Z.
[0072] S3: use an image segmentation algorithm to segment the low-resolution image into n non-overlapping regions, each region including a certain number of grid image blocks, wherein n<< (M+1)(N+1).
[0073] Preferably, the low magnification imaging can use the preview camera of the pathological scanner to capture the image of the pathological section.
[0074] Preferably, image calibration technology is adopted to calibrate the image space of the low magnification imaging system and the high magnification imaging system, so that the image coordinate systems of the high / low resolution images obtained by the two systems are consistent.
[0075] Preferably, the magnification of the low-magnification imaging system is set to 5x to accommodate the minimum observation magnification required for pathological diagnosis of human cell structures.
[0076] Preferably, during image segmentation, the color pathological image is converted to a grayscale image. A fast segmentation algorithm, such as clustering or thresholding, is used to obtain n regions, and region merging is performed so that each region includes at least one grid image patch.
[0077] 2. High-magnification imaging of sparsely sampled representative grids
[0078] The high-magnification imaging workflow of sparsely sampled representative grids is as follows: Figure 3 As shown.
[0079] S1: Process n regions of the low-resolution image sequentially. For the k-th region, search the cache module for a high-resolution image patch H representing the grid of that region. k .
[0080] The representative grid for this region is obtained through the following steps:
[0081] Calculate the effective information score for all low-resolution image patches in the k-th region, and take the grid with the largest effective information score as the representative grid of the region. The effective information score is defined as the weighted sum of tissue coverage score, cell number score, cell coverage score, and color distribution score. The tissue coverage score is defined as the percentage of human tissue coverage area in the grid. The cell number score is defined as the calculated function value of the number of cells in the grid. The cell coverage score is defined as the percentage of cell coverage area in the grid. The color distribution score is defined as the sum of the red, green, and blue channel information entropies in the grid.
[0082] S2: If the high-resolution image patch H representing the grid does not exist in the cache module. k Then select a low-resolution image patch L representing the grid. k .
[0083] S3: Extract L k The coordinates are converted to the corresponding grid coordinates in the coordinate system of the high magnification imaging system.
[0084] S3: Image the human tissue within the grid coordinates using a high-magnification lens to obtain a high-resolution grid image patch H. k .
[0085] S4: Repeat S1 and S3 to obtain high-resolution grid image patches of representative grids for each region, save them to the cache module, and construct a set {H1, H2, ... H... n}
[0086] Preferably, the grid with the largest number of valid information in the kth region is selected as the representative grid of the region.
[0087] Preferably, for each L ij , if there is already a grid in the cache, then directly use it to reduce the number of high magnification imaging. ij , then directly use it to reduce the number of high magnification imaging. k , then directly use it to reduce the number of high magnification imaging.
[0088] Preferably, the high-resolution image blocks of the previously scanned pathological sections are established as a historical image block set {H1, H2,...H h} and saved in the cache module. When scanning new pathological sections from the same source, search for similar high-resolution image blocks from the cache module and L ij . If there is, use the historical grid image block as H k ; if not, use the high magnification imaging system to image the representative grid of the kth region where L ij is located to obtain H k , to further reduce the number of high magnification imaging.
[0089] Preferably, the similarity is the feature distance of the down-sampled image of H k and L ij , and less than a predetermined threshold is considered similar.
[0090] Preferably, the features of H k and L ij are extracted using a pre-trained VGG19, Inception V3, etc. image feature extractor, and the Euclidean distance is calculated as the image feature distance.
[0091] 3. Construct the condition vector
[0092] Different sources of pathological images, such as different slice preparation methods and different body parts, have significant differences in cell shape and color. Using prior information about the source of the pathological section and the preparation method helps to reduce the color bias and cell deformation of the reconstructed image, and improves the accuracy of image reconstruction. The condition vector represents the slice preparation method and the part taken from the human body, and its design is shown in Figure 4 .
[0093] S1: The condition vector is designed as a high-dimensional vector coded by 0 and 1, including two parts of information of slice preparation method and body part.
[0094] S2: The preparation method part covers the current existing slice preparation methods. When the vector element is coded as 1, it represents a preparation method, such as FFPE preparation method, frozen method, etc.
[0095] S3: The human body parts section covers human organ systems based on pathological classification. When a vector element is 1, it indicates that the slice is taken from a human body part, such as the brain, lymph nodes, breast, kidney, ovary, prostate, lung, liver, cervix, bone, thyroid, intestine, etc.
[0096] Preferably, the film preparation method and the human body parts are combined to form a high-dimensional conditional vector.
[0097] Preferably, the conditional vectors are designed with redundancy, and their total length is longer than the sum of the number of currently known slide preparation methods and human organ systems, in order to extend to future slide preparation methods and more detailed organ system classifications.
[0098] 4. Encoder Working Process
[0099] Use an encoder to combine low-resolution image blocks L ij The high-resolution image patch H of the k representative grids in the region. k Or the high-resolution image block H in the cache module k Information is encoded to obtain its latent vector, which can then be used by the decoder to reconstruct image patch L. ij High-resolution image patches. Workflow as follows: Figure 5 As shown:
[0100] S1: Extract any grid image block L from the k-th region of the low-resolution image. ij Its coordinates are i,j. It is converted into a low-resolution local information latent vector C using a local coding module. ij .
[0101] S2: Obtain L using a high-magnification imaging system or from the cache module. ij High-resolution image patch H of the representative grid of the k-th region k The prior coding module converts the information into a high-resolution prior information latent vector.
[0102] S3: C ij and The input is connected along the channel dimension and fed into the information fusion module, outputting a latent vector v of length L. ij v ij The decoder will be used to reconstruct L ij High-resolution image patches.
[0103] S4: Repeat S1-S3 for all grids of the low-resolution image to obtain its latent vector set and establish the latent vector matrix.
[0104] S5: Save the latent vector matrix and the conditional vectors of the pathological slides as a pathological data file.
[0105] Preferably, the local encoding module is implemented by a multi-layer convolutional neural network, including convolutional layers and pooling layers, gradually reducing the size of the grid image block and increasing the number of channels until outputting the local information hidden vector C ij .
[0106] Preferably, the prior encoding module is implemented by a multi-layer convolutional neural network, including convolutional layers and pooling layers, gradually reducing the size of the grid image block and increasing the number of channels until outputting the prior information hidden vector
[0107] Preferably, the information fusion module is implemented by a multi-layer fully connected network, first reducing the dimension, and then expanding the dimension until outputting v ij .
[0108] Preferably, the length L of v ij is much smaller than the number of pixels of the grid image block L ij , and much smaller than the number of pixels of H k , so as to reduce the data storage amount.
[0109] Preferably, the encoder modules can introduce an attention network module to improve performance.
[0110] Preferably, the training method of the encoder refers to step 7.
[0111] 5. Workflow of the decoder
[0112] The decoder is used to convert the hidden vector and the condition vector into a high-resolution image block, and its workflow is as shown in Figure 6 .
[0113] S1: According to the preparation method of the pathological section and the part of the human body taken, or from the pathological data file, obtain the condition vector.
[0114] S2: Input the condition vector into the condition embedding module to complete the embedding of the condition vector and obtain the condition embedding vector w.
[0115] S3: Extract the hidden vector v ij of L ij from the hidden vector matrix of the pathological data file, and connect it with w to obtain the condition hidden vector.
[0116] S4: Input the condition hidden vector into the decoding module to output a high-resolution image block B ij similar to the real high-resolution image block.
[0117] S5: Repeat S3-S4 to obtain the high-resolution image block of each grid and splice them into a high-resolution image of the pathological section.
[0118] Preferably, the conditional embedding module is a one-dimensional convolutional neural network, including convolutional layers and pooling layers.
[0119] Preferably, the decoding module adopts a generative adversarial network structure, such as StyleGAN. It is trained on a dataset of high-resolution image blocks, and the training method is referred to Step 6.
[0120] 6. Training process of the decoder
[0121] The training process of the decoder is shown in Figure 7 .
[0122] S1: Prepare the training dataset.
[0123] S1-1: Prepare pathological sections of various staining methods and human body parts, and set the condition vector as t.
[0124] S1-2: Scan the pathological sections with a high magnification imaging system to obtain high-resolution images and cut them into image blocks.
[0125] S1-3: Divide the image blocks into several sets of training datasets according to the different condition vectors: where t i represents the condition vector of the i-th training dataset, represents the real high-resolution image block in the i-th training dataset.
[0126] S2: Establish a generative adversarial network, where the generator is the decoding module and the discriminator is used to distinguish the authenticity of the generated image. Randomly initialize the condition embedding module, the generator and the discriminator.
[0127] S3: Take training data from any one of the training datasets , and train the condition embedding module, the decoding module and the discriminator.
[0128] S3-1: Input the condition vector t i to the condition embedding module to obtain the condition embedding vector w; randomly initialize the hidden vector as noise, and connect the two.
[0129] S3-2: Input the connected vector to the decoding module to output the high-resolution image block B ij .
[0130] S3-3: Input the condition embedding vector w to the discriminator; calculate the cost of the real high-resolution image block and the output high-resolution image block B ij using the discriminator as the first discrimination cost Loss D,1 .
[0131] The first discrimination cost Loss D,1 is defined as follows:
[0132]
[0133] wherein E is an expectation function; and D is a discriminator.
[0134] S3-4: update the weight values of the conditional embedding module, the decoding module and the discriminator by using the backpropagation through time method with the first discrimination cost as the training target.
[0135] S4: take out other training data, and repeat S3.
[0136] 7. Training process of the encoder
[0137] The training process of the encoder is shown in Figure 8 .
[0138] S1: prepare training data.
[0139] S1-1: scan the pathological section by using a low magnification imaging system to obtain a low resolution image, and divide the grid to obtain a low resolution image block L ij , which is at the coordinate i, j of the grid.
[0140] S1-2: scan the pathological section by using a high magnification imaging system to obtain a high resolution image, and divide the same number of horizontal and vertical times to obtain a high resolution image block H ij , which is at the coordinate i, j of the grid, corresponding to the image block L ij .
[0141] S1-3: divide the low resolution image into n non-overlapping image regions.
[0142] S1-4: calculate the effective information score for all grids in the kth image region where L ij is located, and take the grid with the maximum effective information score as the representative grid, and obtain the image block H k of the grid by using the high magnification imaging system.
[0143] The effective information score is defined as:
[0144] score = λ1s1 + λ2s2 + λ3s3 + λ4s4
[0145] wherein λ 1-4 is a weight.
[0146] s1 = pixel(h_saturation) / (pixel(l_saturation) + pixel(h_saturation))
[0147] s1 is the tissue coverage score, defined as the percentage of the area covered by human tissue in the grid. pixel(h_saturation) is the number of high saturation pixels (human tissue); pixel(l_saturation) is the number of low saturation pixels (background).
[0148] s2 = [1 / (1+exp(-a*cell_counts))-0.5]*2
[0149] s2 is the cell count score, the more the cell count, the closer s2 is to 1. cell_counts is the number of cells in the grid, and 0<α<1 is a trust parameter representing the degree of trust in s2; the smaller the value, the more cell counts are needed to obtain a larger cell count score.
[0150] s3 = pixel(cell) / pixel(grid)
[0151] s3 is the cell coverage score, defined as the percentage of the area covered by cells (pixels) in the grid. pixel(cell) is the number of cell pixels, and pixel(grid) is the number of grid pixels.
[0152]
[0153] s4 is the color distribution score, defined as the sum of the red, green and blue channel information entropy. Where p(k) is the probability value of color value k on the three channel normalized histogram, k = rgb is the color value; ln is the logarithm with base 2.
[0154] S1-5: Establish training data d ij = {t ij , L ij , H k , H ij}.
[0155] S1-6: Repeat scanning multiple pathological sections to take out multiple pathological section image blocks, and establish training data set d = {d ij} and real high-resolution image block set {H ij}.
[0156] S2: Construct the network structure of the encoder and the decoder, wherein the encoder includes a local encoding module, a prior encoding module and an information fusion module; the decoder includes a conditional embedding module and a decoding module. As shown in Figure 8 .
[0157] S3: Initialize the decoder with the decoder weights trained in step 6, and keep the decoder weights unchanged during the training process, while the discriminator weights can change. Randomly initialize the encoder, including the local encoding module, the prior encoding module, and the information fusion module.
[0158] S4: Take out the training data d ij = {t ij , L ij , H k , H ij}.
[0159] S4-1: Input L ij to the local encoding module to obtain C ij ; input H k to the prior encoding module to obtain
[0160] S4-2: Connect C ij and to the information fusion module to obtain v ij .
[0161] S4-3: Input t ij to the conditional embedding module to obtain the conditional embedding vector w.
[0162] S4-4: Connect the conditional embedding vector w and the hidden vector v ij to the decoding module to generate the high-resolution image block B ij .
[0163] S4-5: Calculate the pixel, texture, and perceptual cost of the high-resolution image block B ij and the real high-resolution image block H ij , and add them together with a certain weight as the first cost Loss1.
[0164] S4-6: Input the conditional embedding vector w to the discriminator, and the discriminator calculates the cost of the high-resolution image block B ij and the real high-resolution image set {H ij} as the second discriminant cost Loss D,2 .
[0165] The first cost Loss1 is defined as follows:
[0166] Loss1 = Diff1(B ij , H ij ) + Diff2(B ij , H ij ) + Diff3(B ij , H ij )
[0167] where Diff1(B ij ,H ij ) is a pixel cost; Diff2(B ij ,H ij ) is a texture cost; and Diff3(B ij ,H ij ) is a perception cost. Diff1 is defined as follows:
[0168]
[0169] where m is the number of pixels in the input image.
[0170] Diff2(B ij ,H ij ) is defined as follows:
[0171] Diff2(B ij ,H ij ) = θMean(B ij ,H ij ) + aCon(B ij ,H ij ) + βEntropy(B ij ,H ij )
[0172] where θ, a, and β are constant coefficients; Mean(B ij ,H ij ) is the average of the gray scale difference between the two input feature images B ij and H ij , Con(B ij ,H ij ) is the contrast of the gray scale difference between B ij and H ij , Con(B ij ,H ij ) = |∑ l l 2 p(l) - ∑ k k 2 p(k) |; and Entropy(B ij ,H ij ) is the entropy of the gray scale difference between B ij and H ij , where l and k are certain gray scale values of B ij and H ij , respectively; and p(l) and p(k) represent the probabilities of the gray scale values l and k in the image blocks B ij and H ij , respectively.
[0173] Diff3(Bij ,H ij ) is defined as follows:
[0174]
[0175] wherein is a feature extractor; MSE is a mean square error function.
[0176] Second discriminative cost Loss D,2 is defined as follows:
[0177] Loss D,2 = -E[log(D({H ij},w))]-E[log(1-D(B ij ,w))]
[0178] wherein, E is an expectation function; D is a discriminator; {H ij} is any image patch in the set of real high-resolution image patches.
[0179] S4-7: Update the weights of the local encoding module, the prior encoding module, the information fusion module, and the discriminator using the backpropagation-through-time method with the training objective of minimizing the sum of the first cost and the second discriminative cost.
[0180] S5: Replace the training data and repeat S4.
[0181] 8. Digitization and storage of sparse sampled pathology slides
[0182] After the training of the encoding module and the decoding module, the encoder is used to realize the digitization of the pathology slides, and the workflow is as follows:
[0183] S1: A low magnification imaging system acquires a low-resolution image of the pathology slide and divides the grid image block L ij , and the image is segmented into n regions.
[0184] S2: Obtain {H1, H2,...H n} from the cache module of high-resolution image patches of historical pathology slides or current pathology slides; or use a high magnification imaging system to obtain high-resolution image patches of representative grids of image regions.
[0185] S3: Save the low-resolution image and the set of representative grids {H1, H2,...H n}, record the information of the source of the pathology slide and the method of making, as intermediate data for digitization.
[0186] S4: Repeat S1-S3 to complete the rapid digitization of batches of pathology slides.
[0187] S5: Using the encoder to convert L ij and H k into latent vector v ij , and get the latent vector matrix for all grids.
[0188] S6: Get the condition vector according to the information of pathological section source and preparation method.
[0189] S7: Save the latent vector matrix and the condition vector, and get the pathological data file.
[0190] S8: Repeat S5-S7 to complete the batch of low-storage pathological data files.
[0191] 9. Image reconstruction process
[0192] Using the decoder to reconstruct the pathological data file into an image. The workflow is as follows:
[0193] S1: Read the pathological data file to get the latent vector matrix and the condition vector.
[0194] S2: Take out the latent vector v ij of a grid from the matrix, and input it into the decoder together with the condition vector to get the high-resolution image block B ij of the grid.
[0195] S3: Splice the image blocks to get the high-resolution image of the pathological section.
[0196] In summary, the present application modifies the time-consuming high-magnification sampling of all grids of the pathological section, and the low-magnification sampling and inaccurate super-resolution reconstruction, into a new scheme of high-magnification sparse sampling of representative grids and multi-resolution reconstruction using high and low resolution information, to improve the real-time performance and accuracy of pathological section digitization.
[0197] The present application introduces the condition vector and the high-resolution image of the representative grid, and improves the technical problem of reducing the image reconstruction quality due to different sources of pathological sections or lack of high-resolution prior information, and establishes a multi-resolution reconstruction algorithm to improve the image reconstruction accuracy.
[0198] The present application saves the image pixel data, and improves it to save the shorter latent vector of the grid image, to reduce the data storage amount.
[0199] The above only describes the preferred embodiments of the present application and is not intended to limit the present application. For those skilled in the art, the present application can have various modifications and changes. Any modification, equivalent replacement, improvement, etc. made within the spirit and principles of the present application shall be included in the protection scope of the present application.
Claims
1. A method for reconstructing multi-resolution images of pathological sections based on sparse sampling, characterized in that, Includes the following steps: Low-resolution images of pathological slides are acquired using a low-magnification imaging system, and the low-resolution images are then divided into grids to obtain a set of grid image blocks L = {L 00 L 01 , ...L MN }, where M and N are the number of horizontal and vertical grid divisions, respectively; The low-resolution image is divided into n non-overlapping image regions, each image region including multiple grid image blocks, n << (M+1)(N+1), where (M+1)(N+1) is the number of grids; For any k-th region, where k≤n, the following steps are performed: Calculate the effective information score for all low-resolution image patches in the k-th region, select the grid with the largest effective information score as the representative grid of the region, and acquire the high-resolution image patch H of the representative grid using a high-magnification imaging system. k For any grid image patch L within the k-th region ij With the grid image block L ij The image information is low-resolution local information, with the high-resolution image block H as the reference. k The image information is high-resolution prior information, and the grid image block L is constructed accordingly. ij The implicit vector; The source and preparation method information of the pathological slide are encoded using conditional vectors; high-resolution image blocks of each grid are reconstructed using multi-resolution methods based on the conditional vectors and latent vectors; and the high-resolution image blocks reconstructed from all grids of the low-resolution image are stitched together to obtain a high-resolution image of the pathological slide reconstructed from the low-resolution image.
2. The method for multi-resolution image reconstruction of pathological sections based on sparse sampling according to claim 1, characterized in that, The grid image block L is constructed using an encoder. ij The encoder comprises: a local coding module, a priori coding module, and an information fusion module; The local encoding module is used to extract from the grid image block L ij Extracting low-resolution local information latent vector C ij ; The prior coding module is used to obtain information from the high-resolution image block H. k Extracting high-resolution prior information latent vectors The information fusion module will use the low-resolution local latent vector C ij and high-resolution prior latent vectors Information fusion is performed to obtain a latent vector v that combines low-resolution local information and high-resolution prior information. ij v ij The length L is much smaller than the number of pixels in the grid; the multi-resolution image reconstruction utilizes low-resolution local information and high-resolution prior information to jointly reconstruct a high-resolution image.
3. The method for multi-resolution image reconstruction of pathological sections based on sparse sampling according to claim 2, characterized in that, The conditional vector encoding provides information on the source and preparation method of the pathological slides, including prior information on the pathological slides such as the category code of the preparation method and the category code of the human body part from which the specimen was taken.
4. The method for multi-resolution image reconstruction of pathological sections based on sparse sampling according to claim 3, characterized in that, Obtain the condition vector of the pathological section, and based on the condition vector and latent vector v of each pathological section... ij High-resolution image patches for each grid are reconstructed via a decoder; the decoder is a conditional generation network, including a conditional embedding module and a decoding module. The condition embedding module is used to embed the condition vector to obtain the condition embedding vector w; The decoding module is used to determine the conditional embedding vector w and the latent vector v. ij Reconstructing high-resolution image block B ij .
5. The method for multi-resolution image reconstruction of pathological sections based on sparse sampling according to claim 4, characterized in that, The training of the encoder and decoder specifically involves: The decoder and encoder are trained separately; the decoder is trained first. After the decoder is trained, its weights remain unchanged, and the encoder is trained end-to-end, allowing the encoder to progress from L... ij and H k Generate the latent vector v ij Together with the conditional embedding vector w, it is input into the decoding module to generate a high-resolution image patch B. ij And B ij The generated high-resolution image patch has the greatest similarity to the real high-resolution image patch; the greatest similarity means that the pixel, texture, perceptual cost function and discriminator cost function of the generated high-resolution image patch and the real high-resolution image patch are minimized.
6. The method for multi-resolution image reconstruction of pathological sections based on sparse sampling according to claim 5, characterized in that, The training of the decoder includes the following steps: S61: Scan pathological slides using a high-magnification imaging system to obtain high-resolution image patches, and divide the image patches into several sets of training datasets according to different conditional vectors: S62: Establish a generative adversarial network, where the generator is a decoding module and the discriminator is used to distinguish the authenticity of the generated image; randomly initialize the generator and discriminator. S63: From any training dataset The training data is extracted from the data to train the conditional embedding module, the decoding module, and the discriminator, where t i This represents the conditional vector of the i-th training dataset. This represents a true high-resolution image patch in the i-th training dataset; S63-1: Set the condition vector t i The input is fed into the conditional embedding module; the latent vectors are randomly initialized as noise, and the two are then concatenated. S63-2: Input the connection vector into the decoding module and output high-resolution image block B. ij ; S63-3: Input the conditional embedding vector into the discriminator, and use the discriminator to calculate the true high-resolution image patches. and output high-resolution image patch B ij The cost, as the first identification cost; S63-4: With the goal of minimizing the first discrimination cost, the weights of the conditional embedding module, the decoding module, and the discriminator are updated using the backward gradient diffusion method. S64: Take out the other training data and repeat S63.
7. The method for reconstructing multi-resolution images of pathological sections based on sparse sampling according to claim 6, characterized in that, The training of the encoder includes the following steps: S71: Establish a training dataset d = {d ij } and the set of real high-resolution images {H ij }, where d ij ={t ij L ij H k H ij }, L ij The grid image patch obtained by the low-magnification imaging system has coordinates i, j; t ij H is the conditional vector of the pathological slice to which the grid belongs; ij Grid image patch L acquired for high magnification imaging system ij True high-resolution image patches, H k For L ij A high-resolution image patch of a representative grid in the k-th region; S72: Construct the network structure of the encoder and decoder, wherein the encoder includes a local coding module, a priori coding module, and an information fusion module; the decoder includes a conditional embedding module and a decoding module; S73: Use the trained decoder as the initial decoder, and keep the decoder weights unchanged during training. Randomly initialize the encoder, including the local coding module, the prior coding module, and the information fusion module. S74: Retrieve training data d ij ={t ij L ij H k H ij }; S74-1: Move the grid image block L ij Input the local encoding module to obtain the low-resolution local information latent vector C. ij ; to transfer high-resolution image blocks H k Input the prior coding module to obtain high-resolution prior information latent vectors. S74-2: The local information latent vector C ij and prior information latent vector The input is connected to the information fusion module to obtain the latent vector v, which combines low-resolution local information and high-resolution prior information. ij ; S74-3: Set the condition vector t ij The input is fed into the conditional embedding module to obtain the conditional embedding vector w; S74-4: Integrating the conditional embedding vector w and the latent vector v ij After connection, the data is input to the decoding module to generate high-resolution image patch B. ij ; S74-5: Calculate high-resolution image patch B ij and real high-resolution image blocks H ij The pixel, texture, and perceptual costs are added together with certain weights to form the first cost; the discriminator calculates the high-resolution image patch B. ij and the set of real high-resolution images {H ij The cost of} serves as the second discrimination cost; S74-6: With the training objective of minimizing the sum of the first cost and the second discrimination cost, the weights of the local coding module, the prior coding module, the information fusion module, and the discriminator are updated using the backward gradient diffusion method. S75: Change the training data and repeat S74.
8. The method for multi-resolution image reconstruction of pathological sections based on sparse sampling according to any one of claims 1-7, characterized in that, The effective information score is defined as the weighted sum of tissue coverage score, cell number score, cell coverage score, and color distribution score; the tissue coverage score is defined as the percentage of human tissue coverage area in the grid; the cell number score is defined as the calculated function value of the number of cells in the grid; the cell coverage score is defined as the percentage of cell coverage area in the grid; the color distribution score is defined as the sum of the information entropy of all pixel channels in the grid; the representative grid is the grid with the highest effective information score within the region; and the high-resolution image patch H of the representative grid is obtained. k This includes the following steps: Take any grid L from the k-th image region ij Determine if L exists in the cache module. ij The high-resolution image patch H in the region is used as a representative grid. k If it exists, extract the high-resolution image patch of that representative grid; if it does not exist, select the image patch L of the grid with the largest effective information score from the image region. k Using a high-magnification imaging system to study L k Precisely focus on the tissue within the corresponding grid to acquire a high-resolution image H of a representative grid. k And H k Saved to the cache module; the cache module is a set of high-resolution image blocks {H1, H2, ... H} that stores the current pathological slide or historical pathological slides from the same source. n The memory of}; when digitizing pathological slides from the same source, search for and L from the cache module. ij Similar high-resolution image patches as H k This further reduces the number of high-magnification imaging attempts and speeds up the imaging process.
9. The method for multi-resolution image reconstruction of pathological sections based on sparse sampling according to claim 8, characterized in that, Reconstructing high-resolution image patches for each grid cell based on the conditional vectors of the pathological sections and the latent vectors of each grid cell in the pathological sections includes the following steps: The latent vectors of all grids are combined into a latent vector matrix, which is then stored together with the conditional vectors of the pathological slices as a pathological data file. Its data size is much smaller than that of a file storing image pixels. Read the latent vector matrix and conditional vector from the pathology data file; use a decoder to reconstruct high-resolution image patches for each grid based on the conditional vector of the pathology slice and the latent vector of each grid. By stitching together the high-resolution image blocks of each grid, a high-resolution image reconstructed from the low-resolution image is obtained.
10. A computer system comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 9.
Citation Information
Patent Citations
Methods and systems for image matting and foreground estimation based on hierarchical graphs
US20160078634A1
Pathological section analyzer with large field of view, high throughput and high resolution
WO2022121284A1