Training Method, Device, Equipment and Medium of Image Encoder

Through data augmentation and clustering technology, weight and group loss functions are generated for the image encoder, which solves the problem of inappropriate negative sample assumptions, improves the encoding effect of the image encoder, reduces the impact of false negative samples, and achieves better feature distinction.

CN115115855BActive Publication Date: 2025-07-25TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210531184.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-05-16
Publication Date
2025-07-25
Estimated Expiration
2042-05-16

AI Technical Summary

Technical Problem

In the prior art, when training an image encoder, the negative sample assumption is inappropriate, resulting in the positive sample being erroneously pulled away, affecting the encoding effect of the image encoder.

Method used

Through data augmentation and clustering technology, the image encoder is trained to generate weight loss function and group loss function, distinguish the 'negative degree' of negative samples, and accurately stretch the anchor image from the negative samples.

Benefits of technology

The encoding effect of the image encoder is improved, the influence of false negative samples is reduced, and the characteristics between anchor images and negative samples can be better distinguished.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115115855B_ABST
    Figure CN115115855B_ABST
Patent Text Reader

Abstract

The present application discloses a training method, device, equipment and medium of an image encoder, belonging to the field of artificial intelligence. The method includes: performing data augmentation on a first sample tissue image twice to obtain a first image and a second image; inputting the first image into a first image encoder to obtain a first feature vector; inputting the second image into a second image encoder to obtain a second feature vector; inputting multiple second sample tissue images into the first image encoder to obtain multiple feature vectors; clustering the multiple feature vectors to obtain multiple cluster centers; generating multiple weights based on the similarity values between the multiple cluster centers and the first feature vector; generating a first sub-function based on the first feature vector and the second feature vector; generating a second sub-function by combining the second feature vector and the multiple feature vectors with the multiple weights; generating a first weight loss function based on the first sub-function and the second sub-function; and training the first image encoder based on the first weight loss function.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence, and particularly relates to a training method, device, equipment and medium for an image encoder. Background Art

[0002] In the medical field, there is a scenario of searching for whole slide images (WSIs) similar to a given whole slide image. Each whole slide image (large image) includes a huge number of histopathological images (small images).

[0003] In the related art, the small image with the most representative ability in the large image is used to represent the entire large image, and then the target small image most similar to it is searched in the database according to the feature vector of the small image, and the large image corresponding to the target small image is used as the final search result. The above process requires an image encoder to extract the feature vector of the small image. In the related art, contrast learning is used to train the image encoder. Contrast learning aims to learn the common features of the anchor image and the positive sample, and distinguish the different features between the anchor image and the negative sample (often simply referred to as bringing the anchor image closer to the positive sample and pulling the anchor image away from the negative sample).

[0004] When the related art uses contrast learning to train the image encoder, for image X, the images X1 and X2 obtained by performing data augmentation on image X twice are used as a pair of positive samples, and image X and image Y are used as a pair of negative samples. However, the positive and negative sample assumptions in the related art are inappropriate in special scenarios. In one scenario, when the tissue regions to which the small images selected from one WSI belong are the same as those of the small images selected from another WSI, these two small images are considered a pair of negative samples; in another scenario, when two adjacent small images are selected from the same WSI, these two small images are also considered a pair of negative samples. Obviously, the two small images selected in the above two scenarios should form a pair of positive samples, and the related art will incorrectly pull the positive samples away during the training of the image encoder. Therefore, how to set the correct negative sample assumption in contrast learning has become an urgent technical problem to be solved. Summary of the Invention

[0005] The present application provides a training method, device, equipment and medium for an image encoder, which can improve the encoding effect of the image encoder. The technical solution is as follows:

[0006] According to one aspect of the present application, a training method for an image encoder is provided. The method includes:

[0007] Obtain a first sample tissue image and multiple second sample tissue images, where the second sample tissue images are negative samples in contrast learning;

[0008] Perform data augmentation on the first sample tissue image to obtain a first image; input the first image into a first image encoder to obtain a first feature vector; the first image is a positive sample in contrastive learning;

[0009] Perform data augmentation on the first sample tissue image to obtain a second image; input the second image into a second image encoder to obtain a second feature vector; the second image is an anchor image in contrastive learning;

[0010] Input multiple second sample tissue images into the first image encoder to obtain multiple feature vectors of the multiple second sample tissue images; cluster the multiple feature vectors to obtain multiple cluster centers; generate multiple weights based on the similarity values between the multiple cluster centers and the first feature vector;

[0011] Generate a first sub-function for characterizing the error between the anchor image and the positive sample based on the first feature vector and the second feature vector; generate a second sub-function for characterizing the error between the anchor image and the negative sample by combining the second feature vector and the multiple feature vectors and the multiple weights; generate a first weight loss function based on the first sub-function and the second sub-function;

[0012] Train the first image encoder and the second image encoder based on the first weight loss function; update the first image encoder based on the second image encoder.

[0013] According to another aspect of the present application, there is provided a method for training an image encoder, the method comprising:

[0014] Obtain a first sample tissue image;

[0015] Perform data augmentation on the first sample tissue image to obtain a second image; input the second image into a second image encoder to obtain a fourth feature vector;

[0016] Perform data augmentation on the first sample tissue image to obtain a third image; input the third image into a third image encoder to obtain a fifth feature vector;

[0017] Determine the fourth feature vector as a contrast vector for contrastive learning, and determine the fifth feature vector as an anchor vector for contrastive learning;

[0018] Cluster multiple fourth feature vectors of different first sample tissue images to obtain multiple first cluster centers; determine the feature vector with the largest similarity value between the multiple first cluster centers and the fifth feature vector as the positive sample vector among the multiple fourth feature vectors; determine the remaining first feature vectors as the negative sample vectors among the multiple fourth feature vectors, where the remaining first feature vectors refer to the feature vectors among the multiple fourth feature vectors other than the feature vector with the largest similarity value between the multiple first cluster centers and the fifth feature vector;

[0019] Generate a fifth sub - function based on the fifth eigenvector and the positive sample vectors among multiple fourth eigenvectors; generate a sixth sub - function based on the fifth eigenvector and the negative sample vectors among multiple fourth eigenvectors; generate a first group loss function based on the fifth sub - function and the sixth sub - function;

[0020] Train a second image encoder and a third image encoder based on the first group loss function; determine the third image encoder as the finally trained image encoder.

[0021] According to another aspect of the present application, there is provided a method for searching whole - slide pathology sections, the method comprising:

[0022] Obtain a whole - slide pathology section and crop the whole - slide pathology section into multiple tissue images;

[0023] Generate multiple image feature vectors of the multiple tissue images through an image encoder;

[0024] Determine multiple key images from the multiple tissue images by clustering the multiple image feature vectors;

[0025] Query multiple candidate image packs from a database based on the image feature vectors of the multiple key images, the multiple candidate image packs correspond to the multiple key images one by one, and any one candidate image pack contains at least one candidate tissue image;

[0026] Filter the multiple candidate image packs according to the attributes of the candidate image packs to obtain multiple target image packs;

[0027] Determine the whole - slide pathology sections to which the multiple target tissue images in the multiple target image packs belong as the final search result.

[0028] According to another aspect of the present application, there is provided a training device for an image encoder, the device comprising:

[0029] An acquisition module, configured to acquire a first sample tissue image and multiple second sample tissue images, and the second sample tissue images are negative samples in contrast learning;

[0030] A processing module, configured to perform data augmentation on the first sample tissue image to obtain a first image; input the first image into a first image encoder to obtain a first feature vector; the first image is a positive sample in contrast learning;

[0031] The processing module is further configured to perform data augmentation on the first sample tissue image to obtain a second image; input the second image into a second image encoder to obtain a second feature vector; the second image is an anchor image in contrast learning;

[0032] The processing module is further configured to input multiple second sample tissue images into the first image encoder to obtain multiple feature vectors of the multiple second sample tissue images; cluster the multiple feature vectors to obtain multiple cluster centers; and generate multiple weights based on the similarity values between the multiple cluster centers and the first feature vector.

[0033] The generation module is configured to generate a first sub-function for characterizing the error between the anchor image and the positive sample based on the first feature vector and the second feature vector; generate a second sub-function for characterizing the error between the anchor image and the negative sample by combining the second feature vector and the multiple feature vectors and the multiple weights; and generate a first weight loss function based on the first sub-function and the second sub-function.

[0034] The training module is configured to train the first image encoder and the second image encoder based on the first weight loss function; and update the first image encoder based on the second image encoder.

[0035] According to another aspect of the present application, there is provided a training device for an image encoder, the device including:

[0036] The acquisition module is configured to acquire a first sample tissue image.

[0037] The processing module is configured to perform data augmentation on the first sample tissue image to obtain a second image; and input the second image into the second image encoder to obtain a fourth feature vector.

[0038] The processing module is further configured to perform data augmentation on the first sample tissue image to obtain a third image; and input the third image into the third image encoder to obtain a fifth feature vector.

[0039] The determination module is configured to determine the fourth feature vector as a contrast vector for contrast learning, and determine the fifth feature vector as an anchor vector for contrast learning.

[0040] The clustering module is configured to cluster the fourth feature vectors of different first sample tissue images to obtain multiple first cluster centers; determine the feature vector with the largest similarity value between the multiple first cluster centers and the fifth feature vector as the positive sample vector among the multiple fourth feature vectors; and determine the remaining feature vectors of the multiple first cluster centers as the negative sample vectors among the multiple fourth feature vectors.

[0041] The generation module is configured to generate a fifth sub-function based on the fifth feature vector and the positive sample vector among the multiple fourth feature vectors; generate a sixth sub-function based on the fifth feature vector and the negative sample vectors among the multiple fourth feature vectors; and generate a first group loss function based on the fifth sub-function and the sixth sub-function.

[0042] A training module, configured to train a second image encoder and a third image encoder based on a first group loss function; and determine the third image encoder as the finally trained image encoder.

[0043] According to another aspect of the present application, there is provided a search device for whole-slide pathology sections, the device comprising:

[0044] An acquisition module, configured to acquire whole-slide pathology sections and crop the whole-slide pathology sections into multiple tissue images;

[0045] A generation module, configured to generate multiple image feature vectors of the multiple tissue images through an image encoder;

[0046] A clustering module, configured to determine multiple key images from the multiple tissue images by clustering the multiple image feature vectors;

[0047] A query module, configured to query multiple candidate image packages from a database based on the image feature vectors of the multiple key images, where the multiple candidate image packages correspond to the multiple key images one by one, and any one of the candidate image packages contains at least one candidate tissue image;

[0048] A screening module, configured to screen the multiple candidate image packages according to the attributes of the candidate image packages to obtain multiple target image packages;

[0049] A determination module, configured to determine the whole-slide pathology sections to which the multiple target tissue images in the multiple target image packages belong as the final search result.

[0050] According to one aspect of the present application, there is provided a computer device, the computer device comprising: a processor and a memory, the memory storing a computer program, and the computer program being loaded and executed by the processor to implement the above method for training an image encoder, or, the method for searching whole-slide pathology sections.

[0051] According to another aspect of the present application, there is provided a computer-readable storage medium, storing a computer program, and the computer program being loaded and executed by the processor to implement the above method for training an image encoder, or, the method for searching whole-slide pathology sections.

[0052] According to another aspect of the present application, there is provided a computer program product or a computer program, the computer program product or the computer program comprising computer instructions, and the computer instructions being stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the above method for training an image encoder, or, the method for searching whole-slide pathology sections.

[0053] The beneficial effects brought by the technical solution provided in the embodiments of the present application at least include:

[0054] By assigning weights to the negative samples identified in the related art and further distinguishing the "negative degree" of the negative samples among the negative samples, the loss function used in contrastive learning (also known as the contrastive learning paradigm) can more precisely pull apart the anchor image and the negative samples, reducing the influence of potential false negative samples. Furthermore, it can better train the image encoder, and the trained image encoder can better distinguish the different features between the anchor image and the negative samples. BRIEF DESCRIPTION OF THE DRAWINGS

[0055] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0056] Figure 1 It is a schematic diagram of the related introduction of contrastive learning provided by an embodiment of the present application;

[0057] Figure 2 It is a schematic diagram of a computer system provided by an embodiment of the present application;

[0058] Figure 3 It is a schematic diagram of the training architecture of an image encoder provided by an embodiment of the present application;

[0059] Figure 4 It is a flowchart of the training method of an image encoder provided by an embodiment of the present application;

[0060] Figure 5 It is a schematic diagram of the training architecture of an image encoder provided by another embodiment of the present application;

[0061] Figure 6 It is a schematic diagram of the training architecture of an image encoder provided by another embodiment of the present application;

[0062] Figure 7 It is a flowchart of the training method of an image encoder provided by another embodiment of the present application;

[0063] Figure 8 It is a schematic diagram of the training architecture of an image encoder provided by another embodiment of the present application;

[0064] Figure 9 It is a flowchart of the training method of an image encoder provided by another embodiment of the present application;

[0065] Figure 10It is a schematic diagram of the training architecture of an image encoder provided by another embodiment of the present application;

[0066] Figure 11 It is a schematic diagram of the training architecture of an image encoder provided by another embodiment of the present application;

[0067] Figure 12 It is a flowchart of a training method for an image encoder provided by another embodiment of the present application;

[0068] Figure 13 It is a schematic diagram of the training architecture of an image encoder provided by another embodiment of the present application;

[0069] Figure 14 It is a flowchart of a search method for a whole-field pathological section provided by an embodiment of the present application;

[0070] Figure 15 It is a schematic diagram of the construction architecture of a database provided by an embodiment of the present application;

[0071] Figure 16 It is a block diagram of the structure of a training device for an image encoder provided by an embodiment of the present application;

[0072] Figure 17 It is a block diagram of the structure of a training device for an image encoder provided by an embodiment of the present application;

[0073] Figure 18 It is a block diagram of the structure of a search device for a whole-field pathological section provided by an embodiment of the present application;

[0074] Figure 19 It is a block diagram of the structure of a computer device provided by an embodiment of the present application. Detailed implementation manners

[0075] To make the objectives, technical solutions, and advantages of the present application clearer, the following will further describe the embodiments of the present application in detail with reference to the accompanying drawings.

[0076] First, briefly introduce the nouns involved in the embodiments of the present application:

[0077] Whole Slide Image (WSI): WSI uses a digital scanner to scan traditional pathology slides, collects high-resolution images, and then uses a computer to seamlessly stitch the collected fragmented images to produce a visual digital image. WSI can be enlarged and reduced in any proportion, moved and browsed in any direction, etc. by using specific software. Usually the data size of a WSI is between hundreds of megabytes (MB) and even several gigabytes (GB). In this application, WSI is often referred to as a large picture. The relevant technology focuses on the selection and analysis of local tissue areas within the WSI when processing WSI. In this application, the local tissue areas within the WSI are often referred to as small pictures.

[0078] Contrastive learning (also known as contrastive learning): please refer to Figure 1 , deep learning can be divided into supervised learning and unsupervised learning according to whether the data is labeled. Supervised learning requires labeling of massive amounts of data, while unsupervised learning allows autonomous discovery of potential structures in the data. Unsupervised learning can be further divided into generative learning and contrastive learning. Generative learning is represented by methods such as autoencoders (such as GAN, VAE, etc.), which generate data from data to make it similar to the training data in terms of overall or high-level semantics. For example, multiple horse images in the training set are used to learn the characteristics of horses through a generative model, and then new horse images can be generated.

[0079] Contrastive learning focuses on learning the common features between samples of the same type and distinguishing the different features between samples of different types. In contrastive learning, the encoder is often trained through sample triplets (anchor image, negative sample, positive sample). Figure 1 As shown, circle A is an anchor image in contrastive learning, circle A1 is a positive sample in contrastive learning, and square B is a negative sample in contrastive learning. The contrastive learning aims to shorten the distance between circle A and circle A1 and to lengthen the distance between circle A and square B through the trained encoder. That is, the trained encoder supports similar encoding of similar data and makes the encoding results of different types of data as different as possible. In this application, a method for training an image encoder through contrastive learning will be introduced.

[0080] Next, the implementation environment of this application is introduced.

[0081] Figure 2 FIG. 1 is a schematic diagram of a computer system according to an exemplary embodiment. Figure 2 As shown, the image encoder training device 21 is used to train the image encoder, and then the image encoder training device 21 sends the image encoder to the image encoder using device 22, and the image encoder using device 22 uses the image encoder to search for full-field pathological slices.

[0082] During the training phase of the image encoder, as Figure 2 shown, the image encoder is trained by contrastive learning. The distance between the anchor image 210 and the positive samples is less than the distance between the anchor image 210 and the negative samples. In Figure 2 , the positive samples include the positive sample clusters 211 and 212 obtained through clustering, and the negative samples include the negative sample clusters 213 and 214 obtained through clustering. The distance between the clustering center of the positive sample cluster 211 and the anchor image 210 is L1, the distance between the clustering center of the positive sample cluster 212 and the anchor image 210 is L2, the distance between the clustering center of the negative sample cluster 213 and the anchor image 210 is L3, and the distance between the clustering center of the negative sample cluster 214 and the anchor image 210 is L4.

[0083] In this application, after clustering multiple positive samples, multiple positive sample clusters are obtained. The distance between the clustering center of the cluster most similar to the anchor image and the anchor image is set as L2, and the distances between the other positive samples in the multiple positive samples and the anchor image are set as L1 (note: Figure 2 the L2 shown is only the distance between the clustering center of the positive sample cluster 212 and the anchor image, and the distances between the other positive samples in the positive sample cluster 212 and the anchor image are L1). According to the redefined distances between the multiple positive samples and the anchor image, the anchor image and the multiple positive samples are pulled closer. In the related art, it is considered that the distances between all positive samples and the anchor image are the same.

[0084] In this application, after clustering multiple negative samples, multiple negative sample clusters are obtained. Weights are assigned to each cluster based on the similarity between the clustering center of each cluster and the anchor image, and the anchor image and the negative samples are pulled farther apart according to the cluster weights. Figure 2 The distances L3 and L4 shown are the distances after weighting. In the related art, it is considered that the distances between all negative samples and the anchor image are the same.

[0085] During the usage phase of the image encoder, as Figure 2 shown, in this application, the usage phase of the image encoder is the search process for whole-slide pathology sections.

[0086] First, a WSI is cropped to obtain multiple tissue images (small images); then, the multiple tissue images are clustered to obtain multiple key images, and the multiple key images are jointly used to represent a WSI. Next, for one of the key images (small image A), the small image A is input into the image encoder to obtain the image feature vector of the small image A; finally, according to the image feature vector of the small image A, the database is queried to obtain small images A1 to AN, and the WSIs corresponding to the small images A1 to AN are used as the search results. All the multiple key images are used as query images to determine the WSIs from the database.

[0087] Optionally, the training device 21 of the above image encoder and the using device 22 of the image encoder may be computer devices with machine learning capabilities. For example, the computer device may be a terminal or a server.

[0088] Optionally, the training device 21 of the above image encoder and the using device 22 of the image encoder may be the same computer device, or the training device 21 of the image encoder and the using device 22 of the image encoder may also be different computer devices. Moreover, when the training device 21 of the image encoder and the using device 22 of the image encoder are different devices, the training device 21 of the image encoder and the using device 22 of the image encoder may be devices of the same type. For example, the training device 21 of the image encoder and the using device 22 of the image encoder may both be servers; or the training device 21 of the image encoder and the using device 22 of the image encoder may also be devices of different types. The above server may be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or may also be a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms. The above terminal may be a smart phone, in-vehicle terminal, smart TV, wearable device, tablet computer, notebook computer, desktop computer, smart speaker, smart watch, etc., but is not limited thereto. The terminal and the server may be directly or indirectly connected through wired or wireless communication methods, and this application does not make any restrictions here.

[0089] The following will be introduced in detail in the following order.

[0090] Training stage of the image encoder - 1;

[0091] - Related content of pulling the anchor image away from the negative sample - 1-1;

[0092] - Related content of the first weight loss function - 1-1-1;

[0093] - Related content of the second weight loss function - 1-1-2;

[0094] - Related content of pulling the anchor image closer to the positive sample - 1-2;

[0095] - Related content of the first group loss function - 1-2-1;

[0096] - Related content of the second group loss function - 1-2-2;

[0097] -Related content of the complete loss function - 1-3;

[0098] Usage stage of the image encoder (search process for whole-field pathological sections) - 2;

[0099] Related content of the first weight loss function - 1-1-1:

[0100] Figure 3 Shows a training framework of an image encoder provided by an exemplary embodiment, taking the application of this framework to Figure 2 the illustrated image encoder training device 21 as an example for illustration.

[0101] Figure 3 Shows that: Multiple second sample tissue images 301 pass through the first image encoder 305 to generate multiple feature vectors 307; The first sample tissue image 302 undergoes data augmentation to obtain the first image 303, and the first image 303 passes through the first image encoder 305 to generate the first feature vector 308; The first sample tissue image 302 undergoes data augmentation to obtain the second image 304, and the second image 304 passes through the second image encoder 306 to generate the second feature vector 309; Based on the first feature vector 308 and the second feature vector 309, a first sub-function 310 is generated; Based on the multiple feature vectors 307 and the second feature vector 309, a second sub-function 311 is generated; Based on the first sub-function 310 and the second sub-function 311, a first weight loss function 312 is generated.

[0102] Among them, the first weight loss function 312 is used to increase the distance between the anchor image and the negative samples.

[0103] Figure 4 Shows a flowchart of a training method of an image encoder provided by an exemplary embodiment, taking the application of this method to Figure 3 the illustrated image encoder training framework as an example for illustration. This method includes:

[0104] Step 401, obtain a first sample tissue image and multiple second sample tissue images, where the second sample tissue images are negative samples in contrast learning;

[0105] The first sample tissue image refers to the image used to train the image encoder in this application; The second sample tissue image refers to the image used to train the image encoder in this application. Among them, the first sample tissue image and the second sample tissue image are different small images, that is, the first sample tissue image and the second sample tissue image are not the small images X1 and X2 obtained through data augmentation, but are respectively small Figure X and small image Y. Figure X

[0106] ​In this embodiment, the second sample tissue image is used as the negative sample in contrastive learning. Contrastive learning aims to reduce the distance between the anchor image and the positive sample, and increase the distance between the anchor image and the negative sample.

[0107] With reference to Figure 5 , image X is the first sample tissue image, and the sub-container of the negative sample is the container that holds multiple feature vectors of multiple second sample tissue images.

[0108] Step 402: Perform data augmentation on the first sample tissue image to obtain a first image; input the first image into a first image encoder to obtain a first feature vector; the first image is the positive sample in contrastive learning;

[0109] Data augmentation, also known as data augmentation, aims to generate more data from limited data without substantially increasing the data. In one embodiment, the methods of data augmentation include at least one of the following:

[0110] · Rotation / reflection transformation: Randomly rotate the image by a certain angle to change the orientation of the image content;

[0111] · Flip transformation: Flip the image along the horizontal or vertical direction;

[0112] · Scaling transformation: Enlarge or reduce the image by a certain ratio;

[0113] · Translation transformation: Translate the image on the image plane in a certain way;

[0114] · Specify the translation range and translation step length in a random or artificially defined manner, and translate along the horizontal or vertical direction to change the position of the image;

[0115] · Scale transformation: Enlarge or reduce the image according to the specified scale factor; or refer to the SIFT feature extraction idea, use the specified scale factor to filter the image to construct a scale space, and change the size or blur degree of the image content;

[0116] · Contrast transformation: In the HSV color space of the image, change the saturation S and V brightness components, keep the hue H unchanged, perform an exponential operation (the exponential factor is between 0.25 and 4) on the S and V components of each pixel to increase the illumination change;

[0117] · Noise perturbation: Randomly perturb each pixel RGB of the image. The commonly used noise patterns are salt-and-pepper noise and Gaussian noise;

[0118] · Color change: Add random perturbation to the image channels;

[0119] · Randomly select a region of the input image and blacken it.

[0120] In this embodiment, the first sample tissue image is subjected to data augmentation to obtain a first image, and the first image is used as a positive sample in contrastive learning.

[0121] With reference to Figure 5 , the image X is subjected to data augmentation to obtain the image X k , and then the encoder f is applied to the image X k to transform it into a high-level semantic space i.e., the first feature vector f is obtained k .

[0122] Step 403: The first sample tissue image is subjected to data augmentation to obtain a second image; the second image is input into a second image encoder to obtain a second feature vector; the second image is the anchor image in contrastive learning;

[0123] In this embodiment, the first sample tissue image is subjected to data augmentation to obtain a second image, and the second image is used as the anchor image in contrastive learning.

[0124] In one embodiment, the second image is input into a second image encoder to obtain a first intermediate feature vector; the first intermediate feature vector is input into a first MLP (Multilayer Perceptron) to obtain a second feature vector. Among them, the first MLP plays a transitional role and is used to improve the expressive ability of the second image.

[0125] With reference to Figure 5 , the image X is subjected to data augmentation to obtain the image X p , and then the encoder h is applied to the image X p to transform it into a high-level semantic space i.e., the first intermediate feature vector h is obtained p , and the first intermediate feature vector h p is input into the first MLP to obtain a second feature vector g p2 .

[0126] Step 404: Multiple second sample tissue images are input into the first image encoder to obtain multiple feature vectors of the multiple second sample tissue images; the multiple feature vectors are clustered to obtain multiple cluster centers; based on the similarity values between the multiple cluster centers and the first feature vector, multiple weights are generated;

[0127] In this embodiment, the second sample tissue images are negative samples in contrastive learning. The multiple feature vectors of the multiple second sample tissue images are clustered, and multiple weights are assigned to the multiple feature vectors according to the similarity values between the multiple cluster centers and the first feature vector.

[0128] With reference to Figure 5, in the sub - container of negative samples, there are multiple feature vectors of multiple second - sample tissue images. After multiple second - sample tissue images pass through the encoder f, they are pushed into the storage queue through a stack operation. In the storage queue, the multiple feature vectors in the queue are clustered into Q categories by K - means clustering, and then Q sub - queues are constructed. The clustering center of each sub - queue is denoted as c j (j = 1, …, Q). After that, calculate the similarity score between each clustering center and the first feature vector f k to judge potential incorrect negative samples. Finally, obtain the weight of each feature vector in the storage queue The calculation is as follows:

[0129]

[0130] where δ() is a discriminant function. If the two inputs are the same, δ() outputs 1; otherwise, δ() outputs 0. In this embodiment, δ() is used to judge whether the clustering center c j of the j - th class is similar to f k , w is the assigned weight, and w ∈ [0, 1].

[0131] In one embodiment, the magnitude of the weights of multiple clustering centers is negatively correlated with the similarity value between the clustering center and the first feature vector. For the j - th clustering center among multiple clustering centers, the feature vectors included in the category to which the j - th clustering center belongs correspond to the same weight. In formula (1), for the multiple feature vectors corresponding to the category of the clustering center that is more similar to f k , a smaller weight value w is assigned, and for the multiple feature vectors corresponding to the category of the clustering center that is less similar to f k , a larger weight value w is assigned.

[0132] Schematically, after clustering the multiple feature vectors of multiple second - sample tissue images, 3 categories are obtained, and the clustering centers are c1, c2, and c3 respectively. Among them, the category to which the clustering center c1 belongs includes feature vectors 1, 2, and 3; the category to which the clustering center c2 belongs includes feature vectors 4, 5, and 6; the category to which the clustering center c3 belongs includes feature vectors 7, 8, and 9.

[0133] If the similarity values of the clustering centers c1, c2, and c3 and f k are arranged from large to small, then the weights corresponding to the categories to which the clustering centers c1, c2, and c3 belong are arranged from small to large. Moreover, feature vectors 1, 2, and 3 correspond to the same weight, feature vectors 4, 5, and 6 correspond to the same weight, and feature vectors 7, 8, and 9 correspond to the same weight.

[0134] In one embodiment, when the first sample tissue image belongs to the first sample tissue image in the first training batch, multiple feature vectors of multiple second sample tissue images are clustered to obtain multiple cluster centers of the first training batch.

[0135] In another embodiment, when the first sample tissue image belongs to the first sample tissue image in the nth training batch, the multiple cluster centers corresponding to the (n - 1)th training batch are updated to the multiple cluster centers corresponding to the nth training batch, where n is a positive integer greater than 1.

[0136] Optionally, for the jth cluster center among the multiple cluster centers of the (n - 1)th training batch, based on the first sample tissue images belonging to the jth category in the nth training batch, the jth cluster center of the (n - 1)th training batch is updated to obtain the jth cluster center of the nth training batch, where i is a positive integer.

[0137] With reference to Figure 5 , according to the jth cluster center c j of the (n - 1)th training batch, the jth cluster center c j* of the nth training batch is updated, and the formula is as follows:

[0138]

[0139] Among them, c j* represents the jth cluster center of the nth training batch after update; m c represents the weight used for update, and m c ∈ [0, 1]; represents the feature set belonging to the jth category within the multiple first feature vectors (multiple f k ) of multiple first sample tissue images (multiple images X) in the nth training batch. represents the ith feature vector within the multiple first feature vectors (multiple f k ) belonging to the jth category in the nth training batch. is used to calculate the feature mean of the multiple first feature vectors (multiple f k ) belonging to the jth category in the nth training batch.

[0140] In one embodiment, within each training cycle, all cluster centers will be updated by reclustering all negative sample feature vectors in the repository.

[0141] It can be understood that updating the multiple cluster centers of the (n - 1)th training batch to the multiple cluster centers of the nth training batch aims to prevent the distance between the negative sample feature vectors in the negative sample container and the input first sample tissue image from getting farther and farther.

[0142] As the image encoder is continuously trained, the effect of the image encoder in pulling away the anchor image and negative samples is getting better and better. Suppose the image encoder pulls the image X of the previous training batch and the negative samples to a first distance, the image encoder pulls the image X of the current training batch and the negative samples to a second distance, and the second distance is greater than the first distance. The image encoder pulls the image X of the subsequent training batch and the negative samples to a third distance, and the third distance is greater than the second distance. However, if the negative sample images (i.e., the cluster centers) are not updated, the increase between the third distance and the second distance will be less than the increase between the second distance and the first distance, and the training effect of the image encoder will gradually deteriorate. If the negative sample images (i.e., the cluster centers) are updated, the distance between the updated negative sample images and the image X will be appropriately shortened, balancing the gradually improving pulling-away effect of the image encoder, enabling the image encoder to maintain long-term and frequent training, and ultimately the trained image encoder can also have a better effect.

[0143] Step 405: Based on the first feature vector and the second feature vector, generate a first sub-function for characterizing the error between the anchor image and the positive sample;

[0144] In this embodiment, according to the first feature vector and the second feature vector, a first sub-function is generated, and the first sub-function is used to characterize the error between the anchor image and the positive sample.

[0145] Combined reference Figure 5 The first sub-function can be expressed as exp(g p2 ·f k / τ). It can be seen that the first sub-function is composed of the first feature vector f k and the second feature vector g p2 .

[0146] Step 406: Based on the second feature vector and multiple feature vectors, combine multiple weights to generate a second sub-function for characterizing the error between the anchor image and the negative sample;

[0147] In this embodiment, according to the second feature vector and the multiple feature vectors of multiple second-sample tissue images, multiple weights are combined to generate a second sub-function, and the second sub-function is used to characterize the error between the anchor image and the negative sample.

[0148] Combined reference Figure 5 The second sub-function can be expressed as where represents the weight of the i-th negative sample feature vector (i.e., the feature vector of the second-sample tissue image), represents the i-th negative sample feature vector, and there are a total of K negative sample feature vectors in the negative sample container. g p2 represents the feature vector of the anchor image (i.e., the second feature vector).

[0149] Step 407: Generate a first weight loss function based on the first sub - function and the second sub - function;

[0150] Combined with reference Figure 5 , the first weight loss function can be expressed as:

[0151]

[0152] where, represents the first weight loss function, and log represents the logarithmic operation.

[0153] Step 408: Train the first image encoder and the second image encoder based on the first weight loss function;

[0154] Train the first image encoder and the second image encoder according to the first weight loss function.

[0155] Step 409: Update the first image encoder based on the second image encoder.

[0156] Update the first image encoder based on the second image encoder. Optionally, update the parameters of the first image encoder in a weighted manner according to the parameters of the second image encoder.

[0157] Schematically, the formula for updating the parameters of the first image encoder is as follows:

[0158] θ′ = m·θ′+(1 - m)·θ; (4)

[0159] where, θ′ on the left side of formula (4) represents the parameters of the updated first image encoder, θ′ on the right side of formula (4) represents the parameters of the first image encoder before update, θ represents the parameters of the second image encoder, and m is a constant. Optionally, m = 0.99.

[0160] In summary, by assigning weights to the negative samples identified in the related technology, further distinguishing the "negative degree" of the negative samples among the negative samples, the loss function used in contrast learning (also known as the contrast learning paradigm) can more precisely pull apart the anchor image and the negative samples, reducing the influence of potential false negative samples, and thus can better train the image encoder. The trained image encoder can better distinguish the different features between the anchor image and the negative samples.

[0161] The above Figure 3 and Figure 4It shows that the first image encoder is trained with a sample triplet, where the sample triplet includes (anchor image, positive sample, negative sample). In another embodiment, it is also possible to train the first image encoder with multiple sample triplets simultaneously. Hereinafter, the training of the first image encoder with two sample triplets (anchor image 1, positive sample, negative sample) and (anchor image 2, positive sample, negative sample) will be introduced. Anchor image 1 and anchor image 2 are images obtained by performing data augmentation on the same small image respectively. It should be noted that the present application does not limit the specific number of sample triplets constructed.

[0162] Related content of the second weight loss function - 1-1-2:

[0163] Figure 6 It shows the training framework of an image encoder provided by an exemplary embodiment, taking the application of this framework to Figure 1 the illustrated image encoder training device 21 as an example for illustration.

[0164] Figure 6 It shows that: multiple second sample tissue images 301 pass through the first image encoder 305 to generate multiple feature vectors 307; the first sample tissue image 302 undergoes data augmentation to obtain the first image 303, and the first image 303 passes through the first image encoder 305 to generate the first feature vector 308; the first sample tissue image 302 undergoes data augmentation to obtain the second image 304, and the second image 304 passes through the second image encoder 306 to generate the second feature vector 309; based on the first feature vector 308 and the second feature vector 309, a first sub-function 310 is generated; based on the multiple feature vectors 307 and the second feature vector 309, a second sub-function 311 is generated; based on the first sub-function 310 and the second sub-function 311, a first weight loss function 312 is generated.

[0165] And Figure 3 The difference from the Figure 6 illustrated training framework is that

[0166] it also shows that: the first sample tissue image 302 undergoes data augmentation to obtain the third image 313, and the third image 313 passes through the third image encoder 314 to obtain the third feature vector 315; the third feature vector 315 and the first feature vector 308 generate a third sub-function 316; the third feature vector 315 and the multiple feature vectors 307 generate a fourth sub-function 317; the third sub-function 316 and the fourth sub-function 317 generate a second weight loss function 318.

[0167] Based on Figure 4 the illustrated image encoder training method, Figure 7 In Figure 4On the basis of the method steps, steps 410 to 414 are further provided to Figure 7 Apply the method shown in Figure 6 For example, the training framework of the image encoder shown in, the method includes:

[0168] Step 410: Perform data augmentation on the first sample tissue image to obtain a third image; input the third image into a third image encoder to obtain a third feature vector; the third image is the anchor image in contrastive learning.

[0169] In this embodiment, the first sample tissue image is subjected to data augmentation to obtain a third image, and the third image is used as the anchor image in contrastive learning.

[0170] In one embodiment, input the third image into a third image encoder to obtain a second intermediate feature vector; input the second intermediate feature vector into a second MLP to obtain a third feature vector. Among them, the second MLP plays a transitional role to improve the expression ability of the third image.

[0171] Combined with reference Figure 5 , the image X is subjected to data augmentation to obtain the image X q , and then the encoder h is applied to the image X q Convert to the high-level semantic space That is, the second intermediate feature vector h is obtained q , input the second intermediate feature vector h q Input into the second MLP to obtain a third feature vector g q1 .

[0172] Step 411: Generate a third sub-function for characterizing the error between the anchor image and the positive sample based on the first feature vector and the third feature vector;

[0173] In this embodiment, a third sub-function is generated according to the first feature vector and the third feature vector, and the third sub-function is used to characterize the error between the anchor image and the positive sample.

[0174] Combined with reference Figure 5 , the third sub-function can be expressed as exp(g q1 ·f k / τ), it can be seen that the third sub-function is composed of a feature vector f k And the third feature vector g q1 .

[0175] Step 412: Generate a fourth sub-function for characterizing the error between the anchor image and the negative sample based on the third feature vector and multiple feature vectors, combined with multiple weights;

[0176] In this embodiment, a fourth sub-function is generated according to the third eigenvector and the eigenvectors of multiple second sample tissue images, in combination with multiple weights. The fourth sub-function is used to characterize the error between the anchor image and the negative samples.

[0177] Combined with the reference Figure 5 , the fourth sub-function can be expressed as where represents the weight of the i-th negative sample eigenvector (i.e., the eigenvector of the second sample tissue image), represents the i-th negative sample eigenvector. There are a total of K negative sample eigenvectors in the negative sample container, and g q1 represents the eigenvector of the anchor image (i.e., the third eigenvector).

[0178] Step 413: Generate a second weight loss function based on the third sub-function and the fourth sub-function;

[0179] Combined with the reference Figure 5 , the second weight loss function can be expressed as:

[0180]

[0181] where represents the second weight loss function, and log represents the logarithmic operation.

[0182] Step 414: Train the first image encoder and the third image encoder based on the second weight loss function.

[0183] Train the first image encoder and the third image encoder according to the second weight loss function.

[0184] In one embodiment, in combination with the first weight loss function obtained in the above step 308, a complete weight loss function can be constructed:

[0185]

[0186] where is the complete weight loss function. Train the first image encoder, the second image encoder, and the third image encoder according to the complete weight loss function.

[0187] Optionally, in the above step 409, "update the first image encoder based on the second image encoder" can be replaced with "update the parameters of the first image encoder in a weighted manner according to the parameters shared between the second image encoder and the third image encoder", that is, θ in formula (4) in step 409 represents the parameters shared between the second image encoder and the third image encoder, and the first image encoder is slowly updated through the parameters shared between the second image encoder and the third image encoder.

[0188] To summarize, the above scheme constructs two sample triplets (first image, second image, multiple second sample tissue images) and (third image, second image, multiple second sample tissue images), wherein the first image is anchor image 1 and the third image is anchor image 2, which further improves the encoding effect of the trained image encoder, and the constructed complete weight loss function will be more robust than the first weight loss function or the second weight loss function.

[0189] In the above, the content of training the image encoder based on the weight loss function has been fully introduced, wherein the image encoder includes a first image encoder, a second image encoder and a third image encoder. In the following, the training of the image encoder based on the group loss function will also be introduced.

[0190] Related content of the first group loss function——1-2-1:

[0191] Figure 8 A training framework for an image encoder provided by an exemplary embodiment is shown, and the framework is applied to Figure 1 The image encoder training device 21 is shown for illustration.

[0192] Figure 8 It shows that: the first sample tissue image 801 obtains the second image 802 through data enhancement, the second image 802 obtains the fourth eigenvector 806 through the second image encoder 804, and when multiple first sample tissue images 801 are input at the same time, the multiple fourth eigenvectors will be distinguished into positive sample vectors 807 in the multiple fourth eigenvectors and negative sample vectors 808 in the multiple fourth eigenvectors; the first sample tissue image 801 obtains the third image 803 through data enhancement, and the third image 803 obtains the fifth eigenvector 809 through the third image encoder 805; based on the positive sample vectors 807 and the fifth eigenvector 809 in the multiple fourth eigenvectors, a fifth sub-function 810 is generated; based on the negative sample vectors 808 and the fifth eigenvector 809 in the multiple fourth eigenvectors, a sixth sub-function 811 is generated; based on the fifth sub-function 810 and the sixth sub-function 811, a first group loss function 812 is constructed;

[0193] The first group loss function 812 is used to shorten the distance between the anchor image and the positive sample.

[0194] Figure 9 A flowchart of a training method for an image encoder provided by an exemplary embodiment is shown, wherein the method is applied to Figure 8 The training framework of the image encoder shown is illustrated as an example, and the method includes:

[0195] Step 901: Obtain the first sample tissue image;

[0196] The first sample tissue image, in this application, refers to the image used for training the image encoder, that is, the local region image (small image) within the WSI.

[0197] Combined reference Figure 10 , Image X is the first sample tissue image.

[0198] Step 902: Perform data augmentation on the first sample tissue image to obtain a second image; Input the second image into the second image encoder to obtain a fourth feature vector;

[0199] In this embodiment, perform data augmentation on the first sample tissue image to obtain a second image, and perform feature extraction on the second image through the second image encoder to obtain a fourth feature vector.

[0200] In one embodiment, input the second image into the second image encoder to obtain a first intermediate feature vector; Input the first intermediate feature vector into the third MLP to obtain a fourth feature vector. Among them, the third MLP plays a transitional role to improve the expression ability of the second image.

[0201] Combined reference Figure 10 , Image X is subjected to data augmentation to obtain Image X p , Then apply the encoder h to Image X p to transform it into the high-level semantic space that is, obtain the first intermediate feature vector h p , Input the first intermediate feature vector h p into the third MLP to obtain the second feature vector g p1 .

[0202] Step 903: Perform data augmentation on the first sample tissue image to obtain a third image; Input the third image into the third image encoder to obtain a fifth feature vector;

[0203] In this embodiment, perform data augmentation on the first sample tissue image to obtain a second image, and perform feature extraction on the second image through the second image encoder to obtain a fourth feature vector.

[0204] In one embodiment, input the third image into the third image encoder to obtain a second intermediate feature vector; Input the second intermediate feature vector into the fourth MLP to obtain a fifth feature vector. Among them, the fourth MLP plays a transitional role to improve the expression ability of the third image.

[0205] Combined reference Figure 10 , Image X is subjected to data augmentation to obtain Image X q , Then apply the encoder h to Image X qConvert to the high-level semantic space That is, the second intermediate feature vector h is obtained q , and the second intermediate feature vector h q is input into the fourth MLP to obtain the fifth feature vector g q2 .

[0206] Step 904: Determine the fourth feature vector as the contrast vector for contrast learning, and determine the fifth feature vector as the anchor vector for contrast learning;

[0207] In this embodiment, the fourth feature vector is determined as the contrast vector for contrast learning, and the fifth feature vector is determined as the anchor vector for contrast learning. The contrast vector for contrast learning can be a positive sample vector or a negative sample vector.

[0208] Step 905: Cluster multiple fourth feature vectors of different first sample tissue images to obtain multiple first cluster centers;

[0209] In one embodiment, multiple different first sample tissue images are input simultaneously, and multiple fourth feature vectors of the multiple first sample tissue images are clustered to obtain multiple first cluster centers. Optionally, the multiple different first sample tissue images are sample tissue images of the same training batch.

[0210] In one embodiment, the fourth feature vectors of different first sample tissue images will be clustered into S categories, and the S first cluster centers of the S categories are denoted as where j ∈ [1, …, S].

[0211] Combined with reference Figure 10 , which shows one of the multiple first cluster centers of the fourth feature vectors of different first sample tissue images.

[0212] Step 906: Determine the feature vector with the largest similarity value between the multiple first cluster centers and the fifth feature vector as the positive sample vector among the multiple fourth feature vectors;

[0213] In one embodiment, the first cluster center closest to the fifth feature vector among the S first cluster centers is used as the positive sample vector, denoted as

[0214] Step 907: Determine the remaining first feature vectors as the negative sample vectors among the multiple fourth feature vectors;

[0215] Among them, the remaining first feature vectors refer to the feature vectors among the multiple fourth feature vectors except the feature vector with the largest similarity value between the multiple fourth feature vectors and the fifth feature vector.

[0216] In one embodiment, the feature vectors of the S first clustering centers except are used as negative sample vectors, denoted as

[0217] Step 908: Generate a fifth sub-function based on the fifth feature vector and the positive sample vectors among the multiple fourth feature vectors;

[0218] In one embodiment, the fifth sub-function is denoted as

[0219] where g q2 serves as the anchor vector in contrastive learning, serves as the positive sample vector in contrastive learning.

[0220] Step 909: Generate a sixth sub-function based on the fifth feature vector and the negative sample vectors among the multiple fourth feature vectors;

[0221] In one embodiment, the sixth sub-function is denoted as

[0222] where the fifth feature vector g q2 serves as the anchor vector in contrastive learning, serves as the negative sample vector in contrastive learning.

[0223] Step 910: Generate a first group loss function based on the fifth sub-function and the sixth sub-function;

[0224] In one embodiment, the first group loss function is denoted as:

[0225]

[0226] where represents the first group loss function, and log represents the logarithm operation.

[0227] Step 911: Train the second image encoder and the third image encoder based on the first group loss function; determine the third image encoder as the finally trained image encoder.

[0228] According to the first group loss function, the second image encoder and the third image encoder can be trained.

[0229] In this embodiment, the third image encoder is determined as the finally trained image encoder.

[0230] In summary, by further distinguishing the positive samples identified in the relevant technology and further distinguishing the "degree of positivity" of the positive samples among the positive samples, the loss function used in contrastive learning (also called contrastive learning paradigm) can more accurately bring the anchor image and the positive sample closer, thereby better training the image encoder. The trained image encoder can better learn the common features between the anchor image and the positive sample.

[0231] Above Figure 8 and Figure 9 It is shown that an image encoder is trained by a feature vector sample triplet, and the comparative learning sample triplet includes (anchor vector, positive vector, negative vector). In another embodiment, it is also possible to train the first image encoder by multiple feature vector sample triplets at the same time. The following will introduce the training of the image encoder by two feature vector sample triplets (anchor vector 1, positive vector 1, negative vector 1) (anchor vector 2, positive vector 2, negative vector 2) at the same time, wherein anchor vector 1 and anchor vector 2 are different vectors obtained by respectively performing data enhancement on the first sample tissue image and respectively passing through different image encoders and different MLPs. It should be noted that the present application does not limit the number of specific constructed feature vector sample triplets.

[0232] Related content of the second group loss function——1-2-2:

[0233] Figure 11 A training framework for an image encoder provided by an exemplary embodiment is shown, and the framework is applied to Figure 1 The image encoder training device 21 is shown for illustration.

[0234] Figure 11 It shows that: the first sample tissue image 801 obtains the second image 802 through data enhancement, and the second image 802 obtains the fourth eigenvector 806 through the second image encoder 804. When multiple first sample tissue images 801 are input at the same time, the multiple fourth eigenvectors will be distinguished into positive sample vectors 807 in multiple fourth eigenvectors and negative sample vectors 808 in multiple fourth eigenvectors; the first sample tissue image 801 obtains the third image 803 through data enhancement, and the third image 803 obtains the fifth eigenvector 809 through the third image encoder 805; based on the positive sample vectors 807 and the fifth eigenvector 809 in the multiple fourth eigenvectors, a fifth sub-function 810 is generated; based on the negative sample vectors 808 and the fifth eigenvector 809 in the multiple fourth eigenvectors, a sixth sub-function 811 is generated; based on the fifth sub-function 810 and the sixth sub-function 811, a first group loss function 812 is constructed.

[0235] and Figure 8 The difference between the training framework shown is that Figure 11It is also shown that when multiple first sample tissue images 801 are input simultaneously, multiple fifth eigenvectors will be distinguished into positive sample vectors 813 in multiple fifth eigenvectors and negative sample vectors 814 in multiple fifth eigenvectors; based on the positive sample vectors 813 in multiple fifth eigenvectors and the fourth eigenvector 806, a seventh sub-function 815 is generated; based on the negative sample vectors 814 in multiple fifth eigenvectors and the fourth eigenvector 806, an eighth sub-function 816 is generated; based on the seventh sub-function 815 and the eighth sub-function 816, a second group loss function 817 is constructed;

[0236] The second group loss function 817 is used to shorten the distance between the anchor image and the positive sample.

[0237] based on Figure 9 The training method of the image encoder shown in Figure 12 exist Figure 8 Based on the method steps, steps 912 to 919 are further provided to Figure 11 The method shown is applied to Figure 10 The training framework of the image encoder shown is illustrated as an example, and the method includes:

[0238] Step 912, determining the fifth eigenvector as a contrast vector for contrastive learning, and determining the fourth eigenvector as an anchor vector for contrastive learning;

[0239] In this embodiment, the fifth eigenvector is determined as a contrast vector for contrastive learning, and the fourth eigenvector is determined as an anchor vector for contrastive learning. The contrast vector for contrastive learning can be a positive sample vector or a negative sample vector.

[0240] Step 913, clustering multiple fifth eigenvectors of different first sample tissue images to obtain multiple second cluster centers;

[0241] In one embodiment, multiple different first sample tissue images are input simultaneously, and multiple fifth eigenvectors of the multiple first sample tissue images are clustered to obtain multiple second cluster centers. Optionally, the multiple different first sample tissue images are sample tissue images of the same training batch.

[0242] In one embodiment, the fifth feature vectors of different first sample tissue images are clustered into S categories, and the S second cluster centers of the S categories are expressed as where j∈[1,…,S].

[0243] Combined with reference Figure 10 , which shows one second cluster center among a plurality of second cluster centers of the fifth eigenvectors of different first sample tissue images.

[0244] Step 914: Determine the eigenvector with the largest similarity value between the multiple second cluster centers and the fourth eigenvector as the positive sample vector among the multiple fifth eigenvectors;

[0245] In one embodiment, the second cluster center closest to the fourth eigenvector among the S second cluster centers is used as the positive sample vector, denoted as

[0246] Step 915: Determine the remaining second eigenvectors as the negative sample vectors among the multiple fifth eigenvectors;

[0247] Among them, the remaining second eigenvectors refer to the eigenvectors among the multiple fifth eigenvectors except the eigenvector with the largest similarity value between the fourth eigenvector.

[0248] In one embodiment, the eigenvectors among the S second cluster centers except are used as the negative sample vectors, denoted as

[0249] Step 916: Generate a seventh sub-function based on the fourth eigenvector and the positive sample vector among the multiple fifth eigenvectors;

[0250] In one embodiment, the seventh sub-function is denoted as

[0251] Step 917: Generate an eighth sub-function based on the fourth eigenvector and the negative sample vector among the multiple fifth eigenvectors;

[0252] In one embodiment, the eighth sub-function is denoted as

[0253] Among them, the fourth eigenvector g p1 serves as the anchor vector in contrastive learning, and serves as the negative sample vector in contrastive learning.

[0254] Step 918: Generate a second group loss function based on the seventh sub-function and the eighth sub-function;

[0255] In one embodiment, the second group loss function is denoted as:

[0256]

[0257] Among them, represents the second group loss function, and log represents the logarithm operation.

[0258] Step 919: Train the second image encoder and the third image encoder based on the second group loss function; Determine the second image encoder as the finally trained image encoder.

[0259] Train a second image encoder and a third image encoder according to the second group loss function; determine the second image encoder as the finally trained image encoder.

[0260] In one embodiment, a complete group loss function can be constructed by combining the first group loss function obtained in step 910 above;

[0261]

[0262] wherein, is the complete group loss function. Train the second image encoder and the third image encoder according to the complete group loss function. Determine the second image encoder and the third image encoder as the finally trained image encoders.

[0263] Optionally, after step 919, step 920 is further included, and the parameters of the first image encoder are updated in a weighted manner according to the parameters shared between the second image encoder and the third image encoder

[0264] Schematically, the formula for updating the parameters of the first image encoder is as follows:

[0265] θ′ = m·θ′+(1 - m)·θ; (10)

[0266] wherein, θ′ on the left side of formula (10) represents the parameters of the updated first image encoder, θ′ on the right side of formula (10) represents the parameters of the first image encoder before update, θ represents the parameters shared by the second image encoder and the third image encoder, and m is a constant. Optionally, m is 0.99.

[0267] In summary, by constructing two feature vector sample triples (the fifth feature vector, the positive vector among multiple fourth feature vectors, the negative vector among multiple fourth feature vectors), (the fourth feature vector, the positive vector among multiple fifth feature vectors, the negative vector among multiple fifth feature vectors), the encoding effect of the finally trained image encoder is further improved, and the constructed complete group loss function is more robust than the first group loss function or the second group loss function.

[0268] Related content of the complete loss function - 1 - 3:

[0269] From the above Figures 3 to 7 , it is possible to train the first image encoder through the weighted loss function; from the above Figures 8 to 12 , it is possible to train the first image encoder through the group loss function.

[0270] In an alternative embodiment, the first image encoder can be jointly trained by a weight loss function and a group loss function. Please refer to Figure 13 , which shows a schematic diagram of the training architecture of the first image encoder provided by an exemplary embodiment of the present application.

[0271] Relevant part of the weight loss function:

[0272] The image X is data-augmented to obtain the image X k , the image X k passes through the encoder f to obtain the first feature vector f k ; the image X is data-augmented to obtain the image X p , the image X p passes through the encoder h to obtain the first intermediate feature vector h p , the first intermediate feature vector h p passes through the first MLP to obtain the second feature vector g p2 ; the image X is data-augmented to obtain the image X q , the image X q passes through the encoder h to obtain the second intermediate feature vector h q , the first intermediate feature vector h q passes through the second MLP to obtain the third feature vector g p1 ;

[0273] Multiple second sample tissue images are input into the encoder f and placed in the storage queue through a push operation. In the storage queue, the negative sample feature vectors in the queue are clustered into Q categories by K-means clustering, and then Q sub-queues are constructed. Based on the similarity value between each cluster center and f k , weights are assigned to each cluster center;

[0274] Based on the Q cluster centers and the second feature vector g p2 a sub-function for characterizing negative samples and anchor images is constructed; based on the first feature vector f k and the second feature vector g p2 a sub-function for characterizing positive samples and anchor images is constructed; the two sub-functions are combined to form the first weight loss function;

[0275] Based on the Q cluster centers and the third feature vector g p1 a sub-function for characterizing negative samples and anchor images is constructed; based on the first feature vector f k and the third feature vector g p1 a sub-function for characterizing positive samples and anchor images is constructed; the two sub-functions are combined to form the second weight loss function;

[0276] Train the first image encoder, the second image encoder, and the third image encoder using a weight loss function obtained by combining a first weight loss function and a second weight loss function, and slowly update the parameters of the first image encoder through the parameters shared by the second image encoder and the third image encoder.

[0277] Relevant part of the group loss function:

[0278] Image X is data-augmented to obtain image X p , image X p passes through the encoder h to obtain the first intermediate feature vector h p , the first intermediate feature vector h p passes through the third MLP to obtain the fourth feature vector g p1 ; Image X is data-augmented to obtain image X q , image X q passes through the encoder h to obtain the second intermediate feature vector h q , the first intermediate feature vector h q passes through the fourth MLP to obtain the fifth feature vector g q2 ;

[0279] In the same training batch, aggregate the multiple fourth feature vectors g of multiple first sample tissue images p1 , to obtain multiple first cluster centers; determine the first cluster center closest to the fifth feature vector g of a first sample tissue image among the multiple first cluster centers as the positive sample vector; determine the remaining feature vectors of the multiple first cluster centers as negative sample vectors; construct a sub-function for characterizing the error between the positive sample vector and the anchor vector based on the positive sample vector and the fifth feature vector g q2 ; construct a sub-function for characterizing the error between the negative sample vector and the anchor vector based on the negative sample vector and the fifth feature vector g q2 ; combine the two sub-functions to form the first group loss function; q2 ; construct a sub-function for characterizing the error between the negative sample vector and the anchor vector based on the negative sample vector and the fifth feature vector g

[0280] In the same training batch, aggregate the multiple fifth feature vectors g of multiple first sample tissue images q2 , to obtain multiple second cluster centers; determine the second cluster center closest to the fourth feature vector g of a first sample tissue image among the multiple second cluster centers as the positive sample vector; determine the remaining feature vectors of the multiple second cluster centers as negative sample vectors; construct a sub-function for characterizing the error between the positive sample vector and the anchor vector based on the positive sample vector and the fourth feature vector g p1 ; construct a sub-function for characterizing the error between the negative sample vector and the anchor vector based on the negative sample vector and the fourth feature vector g p1 ; combine the two sub-functions to form the second group loss function; p1 ; construct a sub-function for characterizing the error between the negative sample vector and the anchor vector based on the negative sample vector and the fourth feature vector g

[0281] Train a second image encoder and a third image encoder using a group loss function obtained by combining a first group loss function and a second group loss function.

[0282] Combine relevant parts of a weight loss function and a group loss function:

[0283] It can be understood that training the image encoder based on the weight loss function and the group loss function, both of which determine similarity values based on clustering and reassign positive and negative sample hypotheses. The above weight loss function is used to correct the positive and negative sample hypotheses of negative samples in related technologies, and the above group loss function is used to correct the positive and negative sample hypotheses of positive samples in related technologies.

[0284] At Figure 13 In the training architecture shown, the weight loss function and the group loss function are combined through a hyperparameter, expressed as:

[0285]

[0286] Among them, the on the left side of formula (11) is the final loss function, is the weight loss function, is the group loss function, and λ is a hyperparameter that adjusts the contributions of the two loss functions.

[0287] In summary, the final loss function is jointly constructed by the weight loss function and the group loss function. Compared with a single weight loss function or a single group loss function, the final loss function will be more robust, and the finally trained image encoder will have a better encoding effect. The features of the small images obtained by feature extraction using the image encoder can better represent the small images.

[0288] Usage stage of the image encoder - 2:

[0289] The training stage of the image encoder has been introduced above. In the following, the usage stage of the image encoder will be introduced. In an embodiment provided by the present application, the image encoder will be used in the scenario of WSI image search. Figure 14 Shows a flowchart of a method for searching for whole-slide pathology sections provided by an exemplary embodiment of the present application. Taking the method applied to Figure 1 the usage device 22 of the image encoder shown as an example, at this time, the usage device 22 of the image encoder can also be called a search device for whole-slide pathology sections.

[0290] Step 1401, obtain a whole-slide pathology section and crop the whole-slide pathology section into multiple tissue images;

[0291] Whole-slide pathology image (WSI). WSI is obtained by scanning traditional pathology slides using a digital scanner to collect high-resolution images, and then seamlessly stitching the fragmented images collected by a computer to produce a visualized digital image. In this application, WSI is often referred to as a large image.

[0292] Tissue image, which refers to a local tissue area within a WSI. In this application, tissue images are often referred to as small images.

[0293] In one embodiment, in the preprocessing stage of the WSI, the foreground tissue area within the WSI is extracted through thresholding techniques, and then the foreground tissue area of the WSI is cropped into multiple tissue images based on the sliding window technique.

[0294] Step 1402, generate multiple image feature vectors of multiple tissue images through an image encoder;

[0295] In one embodiment, through the first image encoder trained by the method embodiment shown above Figure 4 generate multiple image feature vectors of multiple tissue images; at this time, the first image encoder is trained based on the first weight loss function. Or,

[0296] through the first image encoder trained by the method embodiment shown above Figure 7 generate multiple image feature vectors of multiple tissue images; at this time, the first image encoder is trained based on the first weight loss function and the second weight loss function; or,

[0297] through the first image encoder trained by the method embodiment shown above Figure 9 generate multiple image feature vectors of multiple tissue images; at this time, the third image encoder is trained based on the first group loss function; or,

[0298] through the first image encoder trained by the method embodiment shown above Figure 12 generate multiple image feature vectors of multiple tissue images; at this time, both the second image encoder and the third image encoder are trained based on the first group loss function and the second group loss function; or,

[0299] through the first image encoder trained by the embodiment shown above Figure 13 generate multiple image feature vectors of multiple tissue images; at this time, the first image encoder is trained based on the weight loss function and the group loss function.

[0300] Step 1403, determine multiple key images from multiple tissue images by clustering multiple image feature vectors;

[0301] In one embodiment, multiple image feature vectors of multiple tissue images are clustered to obtain multiple first - type clusters; multiple clustering centers of the multiple first - type clusters are respectively determined as multiple image feature vectors of multiple key images, that is, multiple key images are determined from multiple tissue images.

[0302] In another embodiment, multiple image feature vectors of multiple tissue images are clustered to obtain multiple first - type clusters. After that, re - clustering is performed. For a target first - type cluster among the multiple first - type clusters, based on the position features of the multiple tissue images corresponding to the target first - type cluster in their respective whole - slide pathology sections, multiple second - type clusters are obtained by clustering; multiple clustering centers corresponding to the multiple second - type clusters included in the target first - type cluster are determined as the image feature vectors of the key images; where the target first - type cluster is any one of the multiple first - type clusters.

[0303] Schematically, the K - means clustering method is used for clustering. When clustering for the first time, multiple image feature vectors f all The clustering results in K1 different categories, denoted as F i , i = 1, 2, …, K1. When clustering for the second time, within each cluster F i , using the spatial coordinate information of multiple tissue images as features, further clustering is performed into K2 categories, where K2 = round(R·N), R is a proportionality parameter, optionally, R is 20%; N is the number of small images in the cluster F i . Based on the above two - stage clustering, finally, K1*K2 clustering centers are obtained. The tissue images corresponding to the K1*K2 clustering centers are used as K1*K2 key images, and the K1*K2 key images are used as the global representation of the WSI. In some embodiments, the key images are often referred to as mosaic images.

[0304] Step 1404: Based on the image feature vectors of multiple key images, multiple candidate image packs are retrieved from the database. The multiple candidate image packs correspond one - to - one with the multiple key images, and any one candidate image pack contains at least one candidate tissue image;

[0305] From the above step 1404, it can be obtained that WSI = {P1, P2, …, P i , …, P k}, where P i and k respectively represent the feature vector of the i - th key image and the total number of key images in the WSI. Both i and k are positive integers. When searching for the WSI, each key image will be used as a query image one by one to generate candidate image packs, and a total of k candidate image packs are generated, denoted as where the i - th candidate image pack b ij and t respectively represent the j - th candidate tissue image and The total number of candidate tissue images inside, where j is a positive integer.

[0306] Step 1405, screen multiple candidate image packs according to the attributes of the candidate image packs to obtain multiple target image packs;

[0307] As can be seen from the above step 1405, a total of k candidate image packs are generated. To improve the search speed of the WSI and optimize the final search result, it is also necessary to screen the k candidate image packs. In one embodiment, k candidate image packs are screened according to the similarity between the candidate image pack and the WSI and / or the diagnostic categories included in the candidate image pack to obtain multiple target image packs. The specific screening steps will be introduced in detail below.

[0308] Step 1406, determine the whole-slide pathology sections to which the multiple target tissue images in the multiple target image packs belong as the final search result.

[0309] After screening out multiple target image packs, determine the whole-slide pathology sections to which the multiple target tissue images in the target image packs belong as the final search result. Optionally, the multiple target tissue images in the target image pack may come from the same whole-slide pathology section or from multiple different whole-slide pathology sections.

[0310] In summary, first, the WSI is cropped to obtain multiple small images, and the multiple small images are passed through an image encoder to obtain multiple image feature vectors of the multiple small images; then, the multiple image feature vectors are clustered, and the small images corresponding to the cluster centers are used as key images; then, each key image is queried to obtain candidate image packs; then, the candidate image packs are screened to obtain target image packs; finally, the WSI corresponding to at least one small image in the candidate image pack is used as the final search result; this method provides a way to search for a WSI (large image) with a WSI (large image), and moreover, the clustering step and screening step mentioned therein can greatly reduce the amount of data processed and improve the search efficiency. And, the method of searching for a WSI (large image) with a WSI (large image) provided in this embodiment does not require a training process and can achieve fast search matching.

[0311] In the related art, the method of using small images to represent large images often adopts the method of manual selection. Pathologists select core small images according to the color and texture features of each small image in the WSI (such as histogram statistical information from various color spaces). Then, the features of these core small images are accumulated as the global representation of the WSI. Next, a support vector machine (SVM) is used to classify the WSI global representations of multiple WSIs into two main disease types. In the search stage, once the disease type of the WSI to be searched is determined, image search can be performed in the WSI library with the same disease type.

[0312] Based on Figure 14 the optional embodiment shown, step 1405 may be replaced with 1405-1.

[0313] 1405-1. Screen multiple candidate image packs according to the number of diagnostic categories in the candidate image pack, and obtain multiple target image packs.

[0314] In one embodiment, for the first candidate image pack among multiple candidate image packs, based on the cosine similarity between at least one candidate tissue image in the first candidate image pack and a key image, the occurrence probability of at least one diagnostic category in the database, and the diagnostic category of at least one candidate tissue image, calculate the entropy value of the candidate image pack. The entropy value is used to measure the number of diagnostic categories corresponding to the first candidate image pack, and the first candidate image pack is any one of the multiple candidate image packs;

[0315] Finally, screen multiple candidate image packs to obtain multiple target image packs with entropy values lower than the entropy threshold.

[0316] Schematically, the calculation formula of the entropy value is as follows:

[0317]

[0318] where Ent i represents the entropy value of the i-th candidate image pack, u i represents the total number of diagnostic categories in the i-th candidate image pack, p m represents the probability that the m-th diagnostic type occurs in the i-th candidate image pack, and m is a positive integer.

[0319] It can be understood that the entropy value is used to represent the uncertainty of the i-th candidate image pack. The larger the entropy value, the higher the uncertainty of the i-th candidate image pack, the more disordered the distribution of the candidate tissue images in the i-th candidate image pack in the dimension of diagnostic categories, that is, the higher the uncertainty of the i-th key image, and the less the i-th key image can be used to characterize the WSI. If multiple candidate tissue images in the i-th candidate image pack have the same diagnostic result, the entropy value of the candidate image pack will be 0, and the i-th key image has the best effect in characterizing the WSI.

[0320] In formula (12), the calculation method of p m is as follows:

[0321]

[0322] where y jrepresents the diagnostic category of the j-th candidate tissue image in the i-th candidate image package; δ() is a discriminant function used to determine whether the diagnostic category of the j-th candidate tissue image is the same as the m-th diagnostic category. If they are the same, it outputs 1; otherwise, it outputs 0; is the weight of the j-th candidate tissue image, which is calculated based on the occurrence probability of at least one diagnostic category in the database; d j represents the cosine similarity between the j-th candidate tissue image and the i-th key image in the i-th candidate package, (d j +1) / 2 is used to ensure that the value range is between 0 and 1.

[0323] For ease of understanding, formula (13) can regard as a weight score v j , which is used to characterize the j-th candidate tissue image in the i-th candidate image package. The denominator of formula (13) represents the total score of the i-th candidate image package, and the numerator of formula (13) represents the sum of the scores of the m-th diagnostic category in the i-th candidate image package.

[0324] Through the above formula (12) and formula (13), multiple candidate image packages can be screened, and the candidate image packages with entropy values higher than the preset entropy threshold can be excluded. Multiple target image packages can be screened out from multiple candidate image packages, denoted as where k′ is the number of multiple target image packages, and k′ is a positive integer.

[0325] In summary, by excluding the candidate image packages with entropy values higher than the preset entropy threshold, candidate image packages with higher stability are screened out, further reducing the amount of data processed in the process of searching WSI by WSI, and improving the search efficiency.

[0326] Based on Figure 14 the optional embodiment shown, step 1405 can be replaced by 1405-2.

[0327] 1405-2. Screen multiple candidate image packages according to the similarity between multiple candidate tissue images and the key image to obtain multiple target image packages.

[0328] In one embodiment, for a first candidate image package among multiple candidate image packages, at least one candidate tissue image in the first candidate image package is arranged in descending order of cosine similarity with a key image; the first m candidate tissue images of the first candidate image package are obtained; m cosine similarities corresponding to the first m candidate tissue images are calculated; wherein, the first candidate image package is any one of the multiple candidate image packages; the average value of the m cosine similarities of the first m candidate tissue images of the multiple candidate image packages is determined as the first average value; a candidate image package whose average value of the cosine similarities of the included at least one candidate tissue image is greater than the first average value is determined as a target image package, and multiple target image packages are obtained, where m is a positive integer.

[0329] Schematically, the multiple candidate image packages are represented as The candidate tissue images in each candidate image package are arranged in descending order of cosine similarity, and the first average value can be expressed as:

[0330]

[0331] Wherein, i and k respectively represent the i-th candidate image package and the total number of the multiple candidate image packages, AveTop represents the average value of the first m cosine similarities in the i-th candidate image package, η is the first average value, and η is used as an evaluation criterion to delete candidate image packages with an average cosine similarity less than η, and then multiple target image packages can be obtained. The multiple target image packages are represented as: i″ and k″ respectively represent the i-th target image package and the total number of the multiple target image packages, and k″ is a positive integer.

[0332] In summary, by eliminating candidate image packages with a similarity to the key image lower than the first average value, that is, screening out candidate image packages with a higher similarity between candidate tissue images and the key image, the amount of data processed in the process of searching WSI by WSI is further reduced, and the search efficiency can be improved.

[0333] It should be noted that the above 1405-1 and 1405-2 can perform the step of screening multiple candidate image packages separately, or can perform the step of screening multiple candidate image packages jointly. At this time, either 1405-1 can be executed first and then 1405-2, or 1405-2 can be executed first and then 1405-1. This application does not make any restrictions on this.

[0334] Based on Figure 14 In the method embodiment shown, in step 1404, querying candidate image packages through a database is involved. Next, the construction process of the database will be introduced. Please refer to Figure 15 , which shows a schematic diagram of the construction framework of a database provided by an exemplary embodiment of the present application.

[0335] Taking a WSI as an example for introduction:

[0336] First, multiple tissue images 1502 are obtained by cropping the WSI 1501; optionally, the cropping method includes: in the preprocessing stage of the WSI, the foreground tissue region within the WSI is extracted through threshold technology, and then the foreground tissue region of the WSI is cropped into multiple tissue images based on the sliding window technology.

[0337] Then, the multiple tissue images 1502 are input into the image encoder 1503 to extract features of the multiple tissue images, obtaining multiple image feature vectors 1505 of the multiple tissue images;

[0338] Finally, based on the multiple image feature vectors 1505 of the multiple tissue images, selection of the multiple tissue images 1502 is performed (i.e., selection of small images 1506). Optionally, the selection of small images 1506 includes two - stage clustering. The first - stage clustering is feature - based clustering 1506 - 1, and the second - stage clustering is coordinate - based clustering 1506 - 2.

[0339] - In the feature - based clustering 1506 - 1, the K - means clustering is used to cluster the multiple image feature vectors 1505 of the multiple tissue images into K1 categories, corresponding to obtaining K1 cluster centers, Figure 15 A small image corresponding to one of the cluster centers is shown;

[0340] - In the feature - based clustering 1506 - 2, for any one of the K1 categories, the K - means clustering is used to cluster the multiple feature vectors included in this category into K2 categories, corresponding to obtaining K2 cluster centers, Figure 15 A small image corresponding to one of the cluster centers is shown;

[0341] - The small images corresponding to the K1*K2 cluster centers obtained through the two - stage clustering are used as representative small images 1506 - 3, Figure 15 A small image corresponding to one of the cluster centers is shown;

[0342] - All the representative small images are used as the small images of the WSI to represent the WSI. Based on this, multiple small images of one WSI are constructed.

[0343] In summary, the process of constructing the database and searching for WSI by WSI is relatively similar. Its purpose is to determine multiple small images for representing one WSI to support matching large images by matching small images during the search process.

[0344] In an alternative embodiment, the training concept of the above image encoder can also be applied to other image domains. Through sample starfield images (small images), the starfield images are derived from starry sky images (large images), and the starfield images indicate local regions in the starry sky images. For example, the starry sky image is an image of the starry sky within a first range, and the starfield image is an image of a sub-range within the first range.

[0345] The training stage of the image encoder includes:

[0346] Obtain a first sample starfield image and multiple second sample starfield images, where the second sample starfield images are negative samples in contrastive learning; perform data augmentation on the first sample starfield image to obtain a first image; input the first image into a first image encoder to obtain a first feature vector; the first image is a positive sample in contrastive learning; perform data augmentation on the first sample starfield image to obtain a second image; input the second image into a second image encoder to obtain a second feature vector; the second image is an anchor image in contrastive learning; input the multiple second sample starfield images into the first image encoder to obtain multiple feature vectors of the multiple second sample starfield images; cluster the multiple feature vectors to obtain multiple cluster centers; generate multiple weights based on the similarity values between the multiple cluster centers and the first feature vector; generate a first sub-function for characterizing the error between the anchor image and the positive sample based on the first feature vector and the second feature vector; generate a second sub-function for characterizing the error between the anchor image and the negative samples by combining the second feature vector and the multiple feature vectors with the multiple weights; generate a first weight loss function based on the first sub-function and the second sub-function; train the first image encoder and the second image encoder based on the first weight loss function; update the first image encoder based on the second image encoder.

[0347] Similarly, the image encoder for starfield images can also adopt other training methods similar to the image encoder for the above sample tissue images, which will not be elaborated here.

[0348] The usage stage of the image encoder includes:

[0349] Obtain a starry sky image and crop the starry sky image into multiple starfield images; generate multiple image feature vectors of the multiple starfield images through the image encoder; determine multiple key images from the multiple starfield images by clustering the multiple image feature vectors; query multiple candidate image packs from a database based on the image feature vectors of the multiple key images, where the multiple candidate image packs correspond to the multiple key images one by one, and any one candidate image pack contains at least one candidate starfield image; screen the multiple candidate image packs according to the attributes of the candidate image packs to obtain multiple target image packs; determine the starry sky images to which the multiple target starfield images in the multiple target image packs belong as the final search result.

[0350] In another alternative embodiment, the training idea of the above image encoder can also be applied to the field of geographical images. The image encoder is trained with sample topographic images (small images), where the topographic images are derived from geomorphic images (large images), and the topographic images indicate local regions in the geomorphic images. For example, the geomorphic image is an image of the geomorphology within a second range captured by a satellite, and the topographic image is an image of a sub-range within the second range.

[0351] The training stage of the image encoder includes:

[0352] Obtain a first sample topographic image and multiple second sample topographic images, where the second sample topographic images are negative samples in contrastive learning; perform data augmentation on the first sample topographic image to obtain a first image; input the first image into a first image encoder to obtain a first feature vector; the first image is a positive sample in contrastive learning; perform data augmentation on the first sample topographic image to obtain a second image; input the second image into a second image encoder to obtain a second feature vector; the second image is an anchor image in contrastive learning; input the multiple second sample topographic images into the first image encoder to obtain multiple feature vectors of the multiple second sample topographic images; cluster the multiple feature vectors to obtain multiple cluster centers; generate multiple weights based on the similarity values between the multiple cluster centers and the first feature vector; generate a first sub-function for characterizing the error between the anchor image and the positive sample based on the first feature vector and the second feature vector; generate a second sub-function for characterizing the error between the anchor image and the negative sample by combining the second feature vector and the multiple feature vectors with the multiple weights; generate a first weight loss function based on the first sub-function and the second sub-function; train the first image encoder and the second image encoder based on the first weight loss function; update the first image encoder based on the second image encoder.

[0353] Similarly, the image encoder for topographic images can also adopt other training methods similar to those of the image encoder for the above sample-organized images, which will not be elaborated here.

[0354] The usage stage of the image encoder includes:

[0355] Obtain a geomorphic image and crop the geomorphic image into multiple topographic images; generate multiple image feature vectors of the multiple topographic images through the image encoder; determine multiple key images from the multiple topographic images by clustering the multiple image feature vectors; query multiple candidate image packs from a database based on the image feature vectors of the multiple key images, where the multiple candidate image packs correspond to the multiple key images one by one, and any one candidate image pack contains at least one candidate topographic image; screen the multiple candidate image packs according to the attributes of the candidate image packs to obtain multiple target image packs; determine the geomorphic image to which the multiple target topographic images in the multiple target image packs belong as the final search result.

[0356] Figure 16 It is a structural block diagram of a training device for an image encoder provided by an exemplary embodiment of the present application. The device includes:

[0357] An acquisition module 1601, configured to acquire a first sample tissue image and multiple second sample tissue images, where the second sample tissue images are negative samples in contrastive learning;

[0358] A processing module 1602, configured to perform data augmentation on the first sample tissue image to obtain a first image; input the first image into a first image encoder to obtain a first feature vector; the first image is a positive sample in contrastive learning;

[0359] The processing module 1602 is further configured to perform data augmentation on the first sample tissue image to obtain a second image; input the second image into a second image encoder to obtain a second feature vector; the second image is an anchor image in contrastive learning;

[0360] The processing module 1602 is further configured to input the multiple second sample tissue images into the first image encoder to obtain multiple feature vectors of the multiple second sample tissue images; cluster the multiple feature vectors to obtain multiple cluster centers; generate multiple weights based on the similarity values between the multiple cluster centers and the first feature vector;

[0361] A generation module 1603, configured to generate a first sub-function for characterizing the error between the anchor image and the positive sample based on the first feature vector and the second feature vector; generate a second sub-function for characterizing the error between the anchor image and the negative sample by combining the second feature vector and the multiple feature vectors and the multiple weights; generate a first weight loss function based on the first sub-function and the second sub-function;

[0362] A training module 1604, configured to train the first image encoder and the second image encoder based on the first weight loss function; update the first image encoder based on the second image encoder.

[0363] In an optional embodiment, the processing module 1602 is further configured to cluster the multiple feature vectors of the multiple second sample tissue images to obtain multiple cluster centers of the first training batch when the first sample tissue image belongs to the first sample tissue image in the first training batch.

[0364] In an optional embodiment, the processing module 1602 is further configured to update the multiple cluster centers corresponding to the (n - 1)th training batch to the multiple cluster centers corresponding to the nth training batch when the first sample tissue image belongs to the first sample tissue image in the nth training batch, where n is a positive integer greater than 1.

[0365] In an optional embodiment, the processing module 1602 is further configured to update the j-th cluster center of the n-1 training batches based on the first sample organization images belonging to the j-th category in the n-th training batch for the j-th cluster center among the multiple cluster centers of the n-1 training batches, so as to obtain the j-th cluster center of the n-th training batch, where i is a positive integer.

[0366] In an optional embodiment, the magnitude of the weight is negatively correlated with the similarity value between the cluster center and the first feature vector; for the j-th cluster center among the multiple cluster centers, the feature vectors included in the category to which the j-th cluster center belongs correspond to the same weight.

[0367] In an optional embodiment, the processing module 1602 is further configured to input the second image into the second image encoder to obtain a first intermediate feature vector; and input the first intermediate feature vector into the first multi-layer perceptron (MLP) to obtain a second feature vector.

[0368] In an optional embodiment, the training module 1604 is further configured to update the parameters of the first image encoder in a weighted manner according to the parameters of the second image encoder.

[0369] In an optional embodiment, the processing module 1602 is further configured to perform data augmentation on the first sample organization image to obtain a third image; input the third image into the third image encoder to obtain a third feature vector; and the third image is the anchor image in contrastive learning.

[0370] In an optional embodiment, the generation module 1603 is further configured to generate a third sub-function for characterizing the error between the anchor image and the positive sample based on the first feature vector and the third feature vector; generate a fourth sub-function for characterizing the error between the anchor image and the negative sample based on the third feature vector and multiple feature vectors in combination with multiple weights; and generate a second weight loss function based on the third sub-function and the fourth sub-function.

[0371] In an optional embodiment, the training module 1604 is further configured to train the first image encoder and the third image encoder based on the second weight loss function.

[0372] In an optional embodiment, the processing module 1602 is further configured to input the third image into the third image encoder to obtain a second intermediate feature vector; and input the second intermediate feature vector into the second MLP to obtain a third feature vector.

[0373] In an optional embodiment, the training module 1604 is further configured to update the parameters of the first image encoder in a weighted manner according to the parameters shared between the second image encoder and the third image encoder.

[0374] In summary, by assigning weights to the negative samples identified in the related art and further distinguishing the "negative degree" of the negative samples among the negative samples, the loss function used in contrastive learning (also known as the contrastive learning paradigm) can more precisely separate the anchor image from the negative samples, reducing the impact of potential false negative samples. Furthermore, it can better train the image encoder. The trained image encoder can better distinguish the different features between the anchor image and the negative samples. The features of the small images extracted by the image encoder can better represent the small images.

[0375] Figure 17 FIG. 4 is a structural block diagram of a training device for an image encoder provided by an exemplary embodiment of the present application. The device includes:

[0376] An acquisition module 1701, configured to acquire a first sample tissue image;

[0377] A processing module 1702, configured to perform data augmentation on the first sample tissue image to obtain a second image; input the second image into a second image encoder to obtain a fourth feature vector;

[0378] The processing module 1702 is further configured to perform data augmentation on the first sample tissue image to obtain a third image; input the third image into a third image encoder to obtain a fifth feature vector;

[0379] A determination module 1703, configured to determine the fourth feature vector as a contrast vector for contrastive learning, and determine the fifth feature vector as an anchor vector for contrastive learning;

[0380] A clustering module 1704, configured to cluster multiple fourth feature vectors of different first sample tissue images to obtain multiple first cluster centers; determine the feature vector with the largest similarity value between the multiple first cluster centers and the fifth feature vector as the positive sample vector among the multiple fourth feature vectors; determine the remaining first feature vectors as the negative sample vectors among the multiple fourth feature vectors, where the remaining first feature vectors refer to the feature vectors among the multiple fourth feature vectors other than the feature vector with the largest similarity value between the multiple first cluster centers and the fifth feature vector;

[0381] A generation module 1705, configured to generate a fifth sub-function based on the fifth feature vector and the positive sample vector among the multiple fourth feature vectors; generate a sixth sub-function based on the fifth feature vector and the negative sample vector among the multiple fourth feature vectors; generate a first group loss function based on the fifth sub-function and the sixth sub-function;

[0382] A training module 1706, configured to train the second image encoder and the third image encoder based on the first group loss function; determine the third image encoder as the finally trained image encoder.

[0383] In an optional embodiment, the processing module 1702 is further configured to input the second image into a second image encoder to obtain a first intermediate feature vector; and input the first intermediate feature vector into a third MLP to obtain a fourth feature vector.

[0384] In an optional embodiment, the processing module 1702 is further configured to input the third image into a third image encoder to obtain a second intermediate feature vector; and input the second intermediate feature vector into a fourth MLP to obtain a fifth feature vector.

[0385] In an optional embodiment, the determination module 1703 is further configured to determine the fifth eigenvector as a contrast vector for contrastive learning, and determine the fourth eigenvector as an anchor vector for contrastive learning.

[0386] In an optional embodiment, the clustering module 1704 is further used to cluster multiple fifth eigenvectors of different first sample tissue images to obtain multiple second cluster centers; determine the eigenvector with the largest similarity value with the fourth eigenvector in the multiple second cluster centers as the positive sample vector in the multiple fifth eigenvectors; determine the second remaining eigenvectors as the negative sample vectors in the multiple fifth eigenvectors, wherein the second remaining eigenvectors refer to the eigenvectors in the multiple fifth eigenvectors except the eigenvector with the largest similarity value with the fourth eigenvector.

[0387] In an optional embodiment, the generation module 1705 is also used to generate a seventh sub-function based on the fourth eigenvector and the positive sample vector in multiple fifth eigenvectors; generate an eighth sub-function based on the fourth eigenvector and the negative sample vector in multiple fifth eigenvectors; and generate a second group loss function based on the seventh sub-function and the eighth sub-function.

[0388] In an optional embodiment, the training module 1706 is further used to train the second image encoder and the third image encoder based on the second group loss function; and determine the second image encoder as the image encoder finally obtained by training.

[0389] In an optional embodiment, the training module 1706 is further configured to update the parameters of the first image encoder in a weighted manner according to the parameters shared between the second image encoder and the third image encoder.

[0390] In summary, by further distinguishing the positive samples identified in the relevant technology and further distinguishing the "degree of positivity" of the positive samples among the positive samples, the loss function used in contrastive learning (also called contrastive learning paradigm) can more accurately bring the anchor image and the positive sample closer, thereby better training the image encoder. The trained image encoder can better learn the common features between the anchor image and the positive sample.

[0391] Figure 18 It is a structural block diagram of a search device for whole - slide pathology sections provided by an exemplary embodiment of the present application. The device includes:

[0392] An acquisition module 1801, configured to acquire whole - slide pathology sections and crop the whole - slide pathology sections into multiple tissue images;

[0393] A generation module 1802, configured to generate multiple image feature vectors of multiple tissue images through an image encoder;

[0394] A clustering module 1803, configured to determine multiple key images from multiple tissue images by clustering multiple image feature vectors;

[0395] A query module 1804, configured to query multiple candidate image packages from a database based on the image feature vectors of multiple key images. Multiple candidate image packages correspond to multiple key images one by one, and any one candidate image package contains at least one candidate tissue image;

[0396] A screening module 1805, configured to screen multiple candidate image packages according to the attributes of the candidate image packages to obtain multiple target image packages;

[0397] A determination module 1806, configured to determine the whole - slide pathology sections to which multiple target tissue images in multiple target image packages belong as the final search result.

[0398] In an optional embodiment, the clustering module 1803 is further configured to cluster multiple image feature vectors of multiple tissue images to obtain multiple first - type clusters; and determine the multiple clustering centers of multiple first - type clusters as the multiple image feature vectors of multiple key images respectively.

[0399] In an optional embodiment, the clustering module 1803 is further configured to, for a target first - type cluster among multiple first - type clusters, cluster to obtain multiple second - type clusters based on the position features of multiple tissue images corresponding to the target first - type cluster in their respective whole - slide pathology sections; and for a target first - type cluster among multiple first - type clusters, determine the multiple clustering centers corresponding to multiple second - type clusters included in the target first - type cluster as the image feature vectors of key images; where the target first - type cluster is any one of multiple first - type clusters.

[0400] In an optional embodiment, the screening module 1805 is further configured to screen multiple candidate image packages according to the number of diagnostic categories of the candidate image packages to obtain multiple target image packages.

[0401] In an alternative embodiment, the screening module 1805 is further configured to calculate the entropy value of a candidate image packet for a first candidate image packet among a plurality of candidate image packets, based on the cosine similarity between at least one candidate tissue image in the first candidate image packet and a key image, the occurrence probability of at least one diagnostic category in a database, and the diagnostic category of at least one candidate tissue image; wherein the entropy value is used to measure the number of diagnostic categories corresponding to the first candidate image packet, and the first candidate image packet is any one of the plurality of candidate image packets; and screen the plurality of candidate image packets to obtain a plurality of target image packets with an entropy value lower than an entropy value threshold.

[0402] In an alternative embodiment, the screening module 1805 is further configured to screen a plurality of candidate image packets according to the similarity between multiple candidate tissue images and a key image, to obtain a plurality of target image packets.

[0403] In an alternative embodiment, the screening module 1805 is further configured to, for a first candidate image packet among a plurality of candidate image packets, arrange at least one candidate tissue image in the first candidate image packet in descending order of cosine similarity to the key image; obtain the first m candidate tissue images of the first candidate image packet; calculate the m cosine similarities corresponding to the first m candidate tissue images; determine the average value of the m cosine similarities of the first m candidate tissue images of the plurality of candidate image packets as a first average value; and determine a candidate image packet whose average value of the cosine similarities of the included at least one candidate tissue image is greater than the first average value as a target image packet, to obtain a plurality of target image packets; wherein the first candidate image packet is any one of the plurality of candidate image packets.

[0404] In summary, first, the WSI is cropped to obtain multiple small images, and the multiple small images are passed through an image encoder to obtain multiple image feature vectors of the multiple small images; then, the multiple image feature vectors are clustered, and the small images corresponding to the cluster centers are used as key images; next, each key image is queried to obtain candidate image packets; then, the candidate image packets are screened to obtain target image packets; finally, the WSI corresponding to at least one small image in the candidate image packet is used as the final search result; the device supports the method of searching for a WSI (large image) by a WSI (large image), and the clustering module and screening module mentioned therein can greatly reduce the amount of data processed and improve the search efficiency. Moreover, the device for searching for a WSI (large image) by a WSI (large image) provided in this embodiment does not require a training process and can achieve fast search matching.

[0405] Figure 19 is a schematic structural diagram of a computer device shown according to an exemplary embodiment. The computer device 1900 may be Figure 2 the training device 21 of the image encoder, or may be Figure 2The device 22 using the middle image encoder. The computer device 1900 includes a Central Processing Unit (CPU) 1901, a system memory 1904 including a Random Access Memory (RAM) 1902 and a Read-Only Memory (ROM) 1903, and a system bus 1905 connecting the system memory 1904 and the central processing unit 1901. The computer device 1900 further includes a basic Input / Output (I / O) system 1906 for facilitating the transfer of information between various components within the computer device, and a mass storage device 1907 for storing an operating system 1913, application programs 1914, and other program modules 1915.

[0406] The basic Input / Output system 1906 includes a display 1908 for displaying information and input devices 1909 such as a mouse and a keyboard for user input of information. The display 1908 and the input devices 1909 are both connected to the central processing unit 1901 through an input / output controller 1910 connected to the system bus 1905. The basic Input / Output system 1906 may further include an input / output controller 1910 for receiving and processing inputs from a plurality of other devices such as a keyboard, a mouse, or an electronic stylus. Similarly, the input / output controller 1910 also provides outputs to a display screen, a printer, or other types of output devices.

[0407] The mass storage device 1907 is connected to the central processing unit 1901 through a mass storage controller (not shown) connected to the system bus 1905. The mass storage device 1907 and its associated computer-readable medium provide non-volatile storage for the computer device 1900. That is, the mass storage device 1907 may include computer-readable media (not shown) such as a hard disk or a Compact Disc Read-Only Memory (CD-ROM) drive.

[0408] Without loss of generality, the computer device-readable medium may include a computer device storage medium and a communication medium. The computer device storage medium includes volatile and non-volatile, removable and non-removable media implemented by any method or technology for storing information such as computer device-readable instructions, data structures, program modules, or other data. The computer device storage medium includes RAM, ROM, erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), CD-ROM, digital video disc (DVD), or other optical storage, magnetic tape cartridges, tapes, disk storage, or other magnetic storage devices. Of course, those skilled in the art will understand that the computer device storage medium is not limited to the above several types. The above-mentioned system memory 1904 and the mass storage device 1907 can be collectively referred to as the memory.

[0409] According to various embodiments of the present disclosure, the computer device 1900 can also run by connecting to a remote computer device on a network such as the Internet. That is, the computer device 1900 can be connected to the network 1911 through the network interface unit 1912 connected to the system bus 1905. Or rather, the network interface unit 1912 can also be used to connect to other types of networks or remote computer device systems (not shown).

[0410] The memory further includes one or more programs. The one or more programs are stored in the memory, and the central processing unit 1901 implements all or part of the steps of the above-mentioned training method of the image encoder by executing the one or more programs.

[0411] This application also provides a computer-readable storage medium. At least one instruction, at least one segment of program, code set, or instruction set is stored in the storage medium. The at least one instruction, the at least one segment of program, the code set, or the instruction set is loaded and executed by a processor to implement the training method of the image encoder provided by the above method embodiment.

[0412] This application provides a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the training method of the image encoder provided by the above method embodiment.

[0413] The serial numbers of the embodiments of the present application above are only for description and do not represent the superiority or inferiority of the embodiments.

[0414] Those of ordinary skill in the art can understand that all or part of the steps to implement the above embodiments can be completed by hardware, or can be completed by instructing relevant hardware through a program. The program can be stored in a computer-readable storage medium. The storage medium mentioned above can be a read-only memory, a magnetic disk, an optical disc, or the like.

[0415] The above are only optional embodiments of the present application and are not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application shall be included within the protection scope of the present application.

Claims

1. A training method for an image encoder, characterized in that The method includes: Obtaining a first sample tissue image and multiple second sample tissue images, where the second sample tissue images are negative samples in contrastive learning; Performing data augmentation on the first sample tissue image to obtain a first image; inputting the first image into a first image encoder to obtain a first feature vector; the first image is a positive sample in the contrastive learning; Performing data augmentation on the first sample tissue image to obtain a second image; inputting the second image into a second image encoder to obtain a second feature vector; the second image is an anchor image in the contrastive learning; Inputting the multiple second sample tissue images into the first image encoder to obtain multiple feature vectors of the multiple second sample tissue images; clustering the multiple feature vectors to obtain multiple cluster centers; generating multiple weights based on the similarity values between the multiple cluster centers and the first feature vector, and the magnitude of the weights is negatively correlated with the similarity values between the cluster centers and the first feature vector; Generating a first sub-function for characterizing the error between the anchor image and the positive sample based on the first feature vector and the second feature vector; generating a second sub-function for characterizing the error between the anchor image and the negative samples by combining the second feature vector and the multiple feature vectors and the multiple weights; generating a first weight loss function based on the first sub-function and the second sub-function; Training the first image encoder and the second image encoder based on the first weight loss function; updating the first image encoder based on the second image encoder.

2. The method according to claim 1, wherein The clustering the multiple feature vectors to obtain multiple cluster centers includes: When the first sample tissue image belongs to the first sample tissue image in the first training batch, clustering the multiple feature vectors of the multiple second sample tissue images to obtain multiple cluster centers of the first training batch; When the first sample tissue image belongs to the first sample tissue image in the nth training batch, updating the multiple cluster centers corresponding to the (n - 1)th training batch to the multiple cluster centers corresponding to the nth training batch, where n is a positive integer greater than 1.

3. The method according to claim 2, characterized in that, The updating the multiple cluster centers corresponding to the (n - 1)th training batch to the multiple cluster centers corresponding to the nth training batch includes: For the jth cluster center among the multiple cluster centers of the (n - 1)th training batch, updating the jth cluster center of the (n - 1)th training batch based on the first sample tissue images belonging to the jth category in the nth training batch to obtain the jth cluster center of the nth training batch, where j is a positive integer.

4. The method according to any one of claims 1 to 3, characterized in that For the jth cluster center among the multiple cluster centers, the feature vectors included in the category to which the jth cluster center belongs correspond to the same weight.

5. The method according to any one of claims 1 to 3, characterized in that The inputting the second image into the second image encoder to obtain a second feature vector includes: Inputting the second image into the second image encoder to obtain a first intermediate feature vector; Input the first intermediate feature vector into a first multi-layer perceptron (MLP) to obtain the second feature vector.

6. The method according to any one of claims 1 to 3, characterized in that Updating the first image encoder based on the second image encoder includes: Updating the parameters of the first image encoder in a weighted manner according to the parameters of the second image encoder.

7. The method according to any one of claims 1 to 3, characterized in that, The method further includes: Performing data augmentation on the first sample tissue image to obtain a third image; inputting the third image into a third image encoder to obtain a third feature vector; the third image is the anchor image in the contrastive learning. Generating a third sub-function for characterizing the error between the anchor image and the positive sample based on the first feature vector and the third feature vector; generating a fourth sub-function for characterizing the error between the anchor image and the negative samples by combining the third feature vector and the multiple feature vectors with the multiple weights; generating a second weight loss function based on the third sub-function and the fourth sub-function. Training the first image encoder and the third image encoder based on the second weight loss function.

8. The method according to claim 7, characterized in that, The step of inputting the third image into the third image encoder to obtain a third feature vector includes: Inputting the third image into the third image encoder to obtain a second intermediate feature vector. Inputting the second intermediate feature vector into a second MLP to obtain the third feature vector.

9. The method according to claim 7, wherein Updating the first image encoder based on the second image encoder includes: Updating the parameters of the first image encoder in a weighted manner according to the parameters shared between the second image encoder and the third image encoder.

10. A search method for a whole-field pathological section, characterized in that, The method is executed by a computer device that runs an image encoder trained by any one of the methods recited in claims 1 to 9. The method includes: Obtaining a whole-slide pathology section and cropping the whole-slide pathology section into multiple tissue images. Generating multiple image feature vectors of the multiple tissue images through the image encoder. Determining multiple key images from the multiple tissue images by clustering the multiple image feature vectors. Querying a database based on the image feature vectors of the multiple key images to obtain multiple candidate image packs, where the multiple candidate image packs correspond to the multiple key images one by one, and any one of the candidate image packs contains at least one candidate tissue image. Filtering the multiple candidate image packs according to the attributes of the candidate image packs to obtain multiple target image packs. Determining the whole-slide pathology sections to which the multiple target tissue images in the multiple target image packs belong as the final search result.

11. The method according to claim 10, wherein, The step of determining multiple key images from the multiple tissue images by clustering the multiple image feature vectors includes: Clustering the multiple image feature vectors of the multiple tissue images to obtain multiple first clusters. Determining the multiple clustering centers of the multiple first clusters as the multiple image feature vectors of the multiple key images respectively.

12. The method according to claim 11, wherein The method further includes: For a target first - type cluster among the multiple first - type clusters, based on the position features of the multiple tissue images corresponding to the target first - type cluster in the whole - slide pathology sections to which they belong, multiple second - type clusters are obtained by clustering; The step of respectively determining the multiple clustering centers of the multiple first - type clusters as the multiple image feature vectors of the multiple key images includes: For a target first - type cluster among the multiple first - type clusters, determining the multiple clustering centers corresponding to the multiple second - type clusters included in the target first - type cluster as the image feature vectors of the key images; Wherein, the target first - type cluster is any one of the multiple first - type clusters.

13. The method according to any one of claims 10 to 12, characterized in that The step of screening the multiple candidate image packs according to the attributes of the candidate image packs to obtain multiple target image packs includes: Screening the multiple candidate image packs according to the number of diagnostic categories of the candidate image packs to obtain the multiple target image packs with entropy values lower than the entropy threshold.

14. The method according to claim 13, wherein The step of screening the multiple candidate image packs according to the number of diagnostic categories of the candidate image packs to obtain the multiple target image packs with entropy values lower than the entropy threshold includes: For a first candidate image pack among the multiple candidate image packs, calculating the entropy value of the first candidate image pack based on the cosine similarity between at least one candidate tissue image in the first candidate image pack and the key image, the occurrence probability of at least one diagnostic category in the database, and the diagnostic category of the at least one candidate tissue image; wherein, the entropy value is used to measure the number of diagnostic categories corresponding to the first candidate image pack, and the first candidate image pack is any one of the multiple candidate image packs; Screening the multiple candidate image packs to obtain the multiple target image packs with entropy values lower than the entropy threshold; Wherein, the step of calculating the entropy value of the first candidate image pack based on the cosine similarity between at least one candidate tissue image in the first candidate image pack and the key image, the occurrence probability of at least one diagnostic category in the database, and the diagnostic category of the at least one candidate tissue image includes: Determining the occurrence probability of the m - th diagnostic category in the first candidate image pack according to the output result of the discriminant function, the weight of the j - th candidate tissue image in the first candidate image pack, and the cosine similarity between the j - th candidate tissue image and the i - th key image; the discriminant function is used to determine whether the diagnostic category of the j - th candidate tissue image is consistent with the m - th diagnostic category, and the weight of the j - th candidate tissue image is calculated according to the occurrence probability of at least one diagnostic category in the database; Calculating the entropy value of the first candidate image pack based on the total number of diagnostic categories in the first candidate image pack and the occurrence probability of the m - th diagnostic category in the first candidate image pack; m is a positive integer.

15. The method according to any one of claims 10 to 12, characterized in that The step of screening the multiple candidate image packs according to the attributes of the candidate image packs to obtain multiple target image packs includes: Screening the multiple candidate image packs according to the similarity between multiple candidate tissue images and the key image to obtain the multiple target image packs.

16. The method according to claim 15, characterized in that, Screening the multiple candidate image packs according to the similarity between multiple candidate tissue images and the key image to obtain the multiple target image packs includes: For a first candidate image pack among the multiple candidate image packs, arranging at least one candidate tissue image in the first candidate image pack in descending order of cosine similarity to the key image; obtaining the first m candidate tissue images of the first candidate image pack; calculating m cosine similarities corresponding to the first m candidate tissue images; Determining the average value of the m cosine similarities of the first m candidate tissue images of the multiple candidate image packs as a first average value; Determining candidate image packs whose average value of the cosine similarities of the included at least one candidate tissue image is greater than the first average value as the target image packs to obtain the multiple target image packs; Wherein, the first candidate image pack is any one of the multiple candidate image packs, and m is a positive integer.

17. A training device for an image encoder, characterized in that, The device includes: An acquisition module, configured to acquire a first sample tissue image and multiple second sample tissue images, where the second sample tissue images are negative samples in contrast learning; A processing module, configured to perform data augmentation on the first sample tissue image to obtain a first image; input the first image into a first image encoder to obtain a first feature vector; the first image is a positive sample in the contrast learning; The processing module is further configured to perform data augmentation on the first sample tissue image to obtain a second image; input the second image into a second image encoder to obtain a second feature vector; the second image is an anchor image in the contrast learning; The processing module is further configured to input the multiple second sample tissue images into the first image encoder to obtain multiple feature vectors of the multiple second sample tissue images; cluster the multiple feature vectors to obtain multiple cluster centers; generate multiple weights based on the similarity values between the multiple cluster centers and the first feature vector, and the magnitude of the weights is negatively correlated with the similarity values between the cluster centers and the first feature vector; A generation module, configured to generate a first sub-function for characterizing the error between the anchor image and the positive sample based on the first feature vector and the second feature vector; generate a second sub-function for characterizing the error between the anchor image and the negative sample by combining the second feature vector and the multiple feature vectors and the multiple weights; generate a first weight loss function based on the first sub-function and the second sub-function; A training module, configured to train the first image encoder and the second image encoder based on the first weight loss function; update the first image encoder based on the second image encoder.

18. A search device for a whole-field pathological section, characterized in that, The device runs an image encoder trained by any method of claims 1 to 9, and the device includes: An acquisition module, configured to acquire a whole-slide pathology section and crop the whole-slide pathology section into multiple tissue images; A generation module, configured to generate multiple image feature vectors of the multiple tissue images through the image encoder; A clustering module, configured to determine multiple key images from the multiple tissue images by clustering the multiple image feature vectors; A query module, configured to query, from a database, multiple candidate image packs based on the image feature vectors of the multiple key images, where the multiple candidate image packs correspond to the multiple key images one by one, and any one of the candidate image packs contains at least one candidate tissue image; A screening module, configured to screen the multiple candidate image packs according to the attributes of the candidate image packs to obtain multiple target image packs; A determination module, configured to determine the whole-slide pathology sections to which the multiple target tissue images in the multiple target image packs belong as the final search result.

19. A computer device, characterized in that, The computer device includes: a processor and a memory, where the memory stores a computer program, and the computer program is loaded and executed by the processor to implement the training method of the image encoder according to any one of claims 1 to 9, or, the search method of the whole-slide pathology section according to any one of claims 10 to 16.

20. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, and the computer program is loaded and executed by a processor to implement the training method of the image encoder according to any one of claims 1 to 9, or, the search method of the whole-slide pathology section according to any one of claims 10 to 16.

Citation Information

Patent Citations

  • Retrieval method and device for pulmonary nodule image block based on convolution neural networks

    CN106874489A

  • Supervised learning method and device for image features, equipment and storage medium

    CN113822325A