Training Method, Device, Equipment and Medium of Image Encoder

By using clustering to determine more accurate positive sample vectors and negative sample vectors in image encoder training, sub-functions and group loss functions are generated, the problem of excessively broad positive sample definition in the prior art is solved, and the encoding effect of image encoder is improved.

CN115115856BActive Publication Date: 2025-06-27TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210531185.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-05-16
Publication Date
2025-06-27
Estimated Expiration
2042-05-16

AI Technical Summary

Technical Problem

In the prior art, when training image encoders, the definition of positive samples is too broad, resulting in limited encoding effects. How to set more accurate positive samples assumptions has become a technical problem that needs to be solved urgently.

Method used

By acquiring the first sample organization image, data augmentation is performed to obtain the first image and the second image, and inputting the first image encoder and the second image encoder respectively to generate the first feature vector and the second feature vector. The first eigenvector is determined as the contrast vector learned by comparison, and the second eigenvector is determined as the anchor vector. Then, the positive sample vector and the negative sample vector in the plurality of first cluster centers are determined by clustering, subfunctions are generated based on these vectors, and a group loss function is constructed for training the image encoder.

Benefits of technology

Through more accurate positive sample assumptions, the loss function in contrast learning is improved, and the image encoder is better trained, which enhances the encoding effect, so that the image encoder can better learn the common features between the anchor image and the positive sample.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115115856B_ABST
    Figure CN115115856B_ABST
Patent Text Reader

Abstract

The present application discloses a training method, device, equipment and medium of an image encoder, belonging to the field of artificial intelligence. The method includes: obtaining a first sample tissue image; respectively performing data augmentation on the first sample tissue image to obtain a first image and a second image; inputting the first image into a first image encoder to obtain a first feature vector; inputting the second image into a second image encoder to obtain a second feature vector; clustering multiple first feature vectors of different first sample tissue images to obtain multiple first clustering centers; determining the feature vector with the largest similarity value to the second feature vector among the multiple first clustering centers as a positive sample vector; determining the remaining first feature vectors as negative sample vectors; generating a first group loss function based on the second feature vector, the positive sample vector and the negative sample vector; and training the first image encoder and the second image encoder based on the first group loss function.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence, and particularly to a method, apparatus, device, and medium for training an image encoder. Background Art

[0002] In the medical field, there is a scenario of searching for whole slide images (WSIs) similar to a given whole slide image. Each whole slide image (large image) includes a huge number of histopathological images (small images).

[0003] In the related art, the most representative small image within the large image is used to represent the entire large image. Then, based on the feature vector of the small image, the most similar target small image is searched for in the database, and the large image corresponding to the target small image is used as the final search result. The above process requires an image encoder to extract the feature vector of the small image. In the related art, contrastive learning is used to train the image encoder. Contrastive learning aims to learn the common features of the anchor image and the positive sample, and distinguish the different features between the anchor image and the negative sample (often simply referred to as bringing the anchor image closer to the positive sample and pulling the anchor image away from the negative sample).

[0004] When the related art trains the image encoder using contrastive learning, for image X, the images X1 and X2 obtained by performing data augmentation on image X twice are used as a pair of positive samples. The definition of positive samples in the related art is too broad, and there may be a large difference in the similarity degree between image X1 and X2 and the anchor image. The encoding effect of the image encoder trained using the related art will be limited by the broad assumption of positive samples. How to set a more accurate positive sample assumption in contrastive learning has become an urgent technical problem to be solved. Summary of the Invention

[0005] This application provides a method, apparatus, device, and medium for training an image encoder, which can improve the encoding effect of the image encoder. The technical solution is as follows:

[0006] According to one aspect of this application, a method for training an image encoder is provided. The method includes:

[0007] Obtain a first sample tissue image;

[0008] Perform data augmentation on the first sample tissue image to obtain a first image; input the first image into a first image encoder to obtain a first feature vector;

[0009] Perform data augmentation on the first sample tissue image to obtain a second image; input the second image into a second image encoder to obtain a second feature vector;

[0010] Determine the first feature vector as the contrast vector for contrast learning, and determine the second feature vector as the anchor vector for contrast learning;

[0011] Cluster the multiple first feature vectors of different first sample tissue images to obtain multiple first cluster centers; determine the feature vector with the largest similarity value between the multiple first cluster centers and the second feature vector as the positive sample vector among the multiple first feature vectors; determine the remaining first feature vectors as the negative sample vectors among the multiple first feature vectors, where the remaining first feature vectors refer to the feature vectors among the multiple first feature vectors other than the feature vector with the largest similarity value between it and the second feature vector;

[0012] Generate a first sub-function based on the second feature vector and the positive sample vector among the multiple first feature vectors; generate a second sub-function based on the second feature vector and the negative sample vectors among the multiple first feature vectors; generate a first group loss function based on the first sub-function and the second sub-function;

[0013] Train a first image encoder and a second image encoder based on the first group loss function; determine the second image encoder as the finally trained image encoder.

[0014] According to another aspect of the present application, there is provided a method for training an image encoder, the method comprising:

[0015] Obtain a first sample tissue image and multiple second sample tissue images, and the second sample tissue images are negative samples in contrast learning;

[0016] Perform data augmentation on the first sample tissue image to obtain a third image; input the third image into a third image encoder to obtain a third feature vector; the third image is a positive sample in contrast learning;

[0017] Perform data augmentation on the first sample tissue image to obtain a first image; input the first image into a first image encoder to obtain a fourth feature vector; the first image is an anchor image in contrast learning;

[0018] Input the multiple second sample tissue images into a third image encoder to obtain multiple feature vectors of the multiple second sample tissue images; cluster the multiple feature vectors to obtain multiple cluster centers; generate multiple weights based on the similarity values between the multiple cluster centers and the third feature vector;

[0019] Generate a fifth sub-function for characterizing the error between the anchor image and the positive sample based on the fourth feature vector and the third feature vector; generate a sixth sub-function for characterizing the error between the anchor image and the negative sample based on the fourth feature vector and the multiple feature vectors in combination with the multiple weights; generate a first weight loss function based on the fifth sub-function and the sixth sub-function;

[0020] Based on the first weight loss function, train the third image encoder and the first image encoder; update the third image encoder based on the first image encoder.

[0021] According to another aspect of the present application, a search method for whole-slide pathology sections is provided, and the method includes:

[0022] Obtain a whole-slide pathology section, and crop the whole-slide pathology section into multiple tissue images;

[0023] Generate multiple image feature vectors of the multiple tissue images through an image encoder;

[0024] Determine multiple key images from the multiple tissue images by clustering the multiple image feature vectors;

[0025] Based on the image feature vectors of the multiple key images, query multiple candidate image packages from a database. The multiple candidate image packages correspond to the multiple key images one by one, and any one candidate image package contains at least one candidate tissue image;

[0026] Filter the multiple candidate image packages according to the attributes of the candidate image packages to obtain multiple target image packages;

[0027] Determine the whole-slide pathology sections to which the multiple target tissue images in the multiple target image packages belong as the final search result.

[0028] According to another aspect of the present application, a training device for an image encoder is provided, and the device includes:

[0029] An acquisition module for acquiring a first sample tissue image;

[0030] A processing module for performing data augmentation on the first sample tissue image to obtain a first image; inputting the first image into the first image encoder to obtain a first feature vector;

[0031] The processing module is further configured to perform data augmentation on the first sample tissue image to obtain a second image; input the second image into the second image encoder to obtain a second feature vector;

[0032] A determination module for determining the first feature vector as a contrast vector for contrast learning and determining the second feature vector as an anchor vector for contrast learning;

[0033] A clustering module, configured to cluster the first feature vectors of different first sample tissue images to obtain multiple first cluster centers; determine the feature vector with the largest similarity value between the multiple first cluster centers and the second feature vector as the positive sample vector among the multiple first feature vectors; determine the remaining feature vectors of the multiple first cluster centers as the negative sample vectors among the multiple first feature vectors;

[0034] A generation module, configured to generate a first sub-function based on the second feature vector and the positive sample vector among the multiple first feature vectors; generate a second sub-function based on the second feature vector and the negative sample vectors among the multiple first feature vectors; generate a first group loss function based on the first sub-function and the second sub-function;

[0035] A training module, configured to train a first image encoder and a second image encoder based on the first group loss function; determine the second image encoder as the finally trained image encoder.

[0036] According to another aspect of the present application, there is provided a training device for an image encoder, the device including:

[0037] An acquisition module, configured to acquire first sample tissue images and multiple second sample tissue images, and the second sample tissue images are negative samples in contrastive learning;

[0038] A processing module, configured to perform data augmentation on the first sample tissue image to obtain a third image; input the third image into a third image encoder to obtain a third feature vector; the third image is a positive sample in contrastive learning;

[0039] The processing module is further configured to perform data augmentation on the first sample tissue image to obtain a first image; input the first image into a first image encoder to obtain a fourth feature vector; the first image is an anchor image in contrastive learning;

[0040] The processing module is further configured to input the multiple second sample tissue images into the third image encoder to obtain multiple feature vectors of the multiple second sample tissue images; cluster the multiple feature vectors to obtain multiple cluster centers; generate multiple weights based on the similarity values between the multiple cluster centers and the third feature vector;

[0041] A generation module, configured to generate a fifth sub-function for characterizing the error between the anchor image and the positive sample based on the fourth feature vector and the third feature vector; generate a sixth sub-function for characterizing the error between the anchor image and the negative samples by combining the fourth feature vector and the multiple feature vectors and the multiple weights; generate a first weight loss function based on the fifth sub-function and the sixth sub-function;

[0042] A training module, configured to train a third image encoder and a first image encoder based on a first weight loss function; and update the third image encoder based on the first image encoder.

[0043] According to another aspect of the present application, there is provided a search device for whole-slide pathology sections, the device comprising:

[0044] An acquisition module, configured to acquire whole-slide pathology sections and crop the whole-slide pathology sections into multiple tissue images;

[0045] A generation module, configured to generate multiple image feature vectors of the multiple tissue images through an image encoder;

[0046] A clustering module, configured to determine multiple key images from the multiple tissue images by clustering the multiple image feature vectors;

[0047] A query module, configured to query multiple candidate image packages from a database based on the image feature vectors of the multiple key images, where the multiple candidate image packages correspond to the multiple key images one by one, and any one of the candidate image packages contains at least one candidate tissue image;

[0048] A screening module, configured to screen the multiple candidate image packages according to the attributes of the candidate image packages to obtain multiple target image packages;

[0049] A determination module, configured to determine the whole-slide pathology sections to which the multiple target tissue images in the multiple target image packages belong as the final search result.

[0050] According to one aspect of the present application, there is provided a computer device, which includes a processor and a memory. The memory stores a computer program, and the computer program is loaded and executed by the processor to implement the above method for training an image encoder or the method for searching whole-slide pathology sections.

[0051] According to another aspect of the present application, there is provided a computer-readable storage medium, which stores a computer program, and the computer program is loaded and executed by the processor to implement the above method for training an image encoder or the method for searching whole-slide pathology sections.

[0052] According to another aspect of the present application, there is provided a computer program product or a computer program, which includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the above method for training an image encoder or the method for searching whole-slide pathology sections.

[0053] The beneficial effects brought by the technical solution provided by the embodiments of the present application at least include:

[0054] By further distinguishing the positive samples recognized in the related art and further differentiating the "degree of positivity" of the positive samples among the positive samples, the loss function used in contrastive learning (also known as the contrastive learning paradigm) can more precisely bring the anchor image closer to the positive samples, and thus can better train the image encoder. The trained image encoder can better learn the common features between the anchor image and the positive samples. BRIEF DESCRIPTION OF THE DRAWINGS

[0055] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0056] Figure 1 It is a schematic diagram of the related introduction of contrastive learning provided by an embodiment of the present application;

[0057] Figure 2 It is a schematic diagram of a computer system provided by an embodiment of the present application;

[0058] Figure 3 It is a schematic diagram of the training architecture of an image encoder provided by an embodiment of the present application;

[0059] Figure 4 It is a flowchart of the training method of an image encoder provided by an embodiment of the present application;

[0060] Figure 5 It is a schematic diagram of the training architecture of an image encoder provided by another embodiment of the present application;

[0061] Figure 6 It is a schematic diagram of the training architecture of an image encoder provided by another embodiment of the present application;

[0062] Figure 7 It is a flowchart of the training method of an image encoder provided by another embodiment of the present application;

[0063] Figure 8 It is a schematic diagram of the training architecture of an image encoder provided by another embodiment of the present application;

[0064] Figure 9 It is a flowchart of the training method of an image encoder provided by another embodiment of the present application;

[0065] Figure 10 It is a schematic diagram of the training architecture of an image encoder provided by another embodiment of the present application;

[0066] Figure 11 It is a schematic diagram of the training architecture of an image encoder provided by another embodiment of the present application;

[0067] Figure 12 It is a flowchart of a method for training an image encoder provided by another embodiment of the present application;

[0068] Figure 13 It is a schematic diagram of the training architecture of an image encoder provided by another embodiment of the present application;

[0069] Figure 14 It is a flowchart of a method for searching a whole - field pathological section provided by an embodiment of the present application;

[0070] Figure 15 It is a schematic diagram of the construction architecture of a database provided by an embodiment of the present application;

[0071] Figure 16 It is a block diagram of the structure of a training device for an image encoder provided by an embodiment of the present application;

[0072] Figure 17 It is a block diagram of the structure of a training device for an image encoder provided by an embodiment of the present application;

[0073] Figure 18 It is a block diagram of the structure of a searching device for a whole - field pathological section provided by an embodiment of the present application;

[0074] Figure 19 It is a block diagram of the structure of a computer device provided by an embodiment of the present application. Detailed implementation manners

[0075] To make the objectives, technical solutions, and advantages of the present application clearer, the following will further describe the embodiments of the present application in detail with reference to the accompanying drawings.

[0076] First, briefly introduce the nouns involved in the embodiments of the present application:

[0077] Whole Slide Image (WSI): WSI uses a digital scanner to scan traditional pathology slides, collects high-resolution images, and then uses a computer to seamlessly stitch the collected fragmented images to produce a visual digital image. WSI can be enlarged and reduced in any proportion, moved and browsed in any direction, etc. by using specific software. Usually the data size of a WSI is between hundreds of megabytes (MB) and even several gigabytes (GB). In this application, WSI is often referred to as a large picture. The relevant technology focuses on the selection and analysis of local tissue areas within the WSI when processing WSI. In this application, the local tissue areas within the WSI are often referred to as small pictures.

[0078] Contrastive learning (also known as contrastive learning): please refer to Figure 1 , deep learning can be divided into supervised learning and unsupervised learning according to whether the data is labeled. Supervised learning requires labeling of massive amounts of data, while unsupervised learning allows autonomous discovery of potential structures in the data. Unsupervised learning can be further divided into generative learning and contrastive learning. Generative learning is represented by methods such as autoencoders (such as GAN, VAE, etc.), which generate data from data to make it similar to the training data in terms of overall or high-level semantics. For example, multiple horse images in the training set are used to learn the characteristics of horses through a generative model, and then new horse images can be generated.

[0079] Contrastive learning focuses on learning the common features between samples of the same type and distinguishing the different features between samples of different types. In contrastive learning, the encoder is often trained through sample triplets (anchor image, negative sample, positive sample). Figure 1 As shown, circle A is an anchor image in contrastive learning, circle A1 is a positive sample in contrastive learning, and square B is a negative sample in contrastive learning. The contrastive learning aims to shorten the distance between circle A and circle A1 and to lengthen the distance between circle A and square B through the trained encoder. That is, the trained encoder supports similar encoding of similar data and makes the encoding results of different types of data as different as possible. In this application, a method for training an image encoder through contrastive learning will be introduced.

[0080] Next, the implementation environment of this application is introduced.

[0081] Figure 2 FIG. 1 is a schematic diagram of a computer system according to an exemplary embodiment. Figure 2 As shown, the image encoder training device 21 is used to train the image encoder, and then the image encoder training device 21 sends the image encoder to the image encoder using device 22, and the image encoder using device 22 uses the image encoder to search for full-field pathological slices.

[0082] During the training phase of the image encoder, as Figure 2 shown, the image encoder is trained by contrastive learning. The distance between the anchor image 210 and the positive samples is less than the distance between the anchor image 210 and the negative samples. In Figure 2 , the positive samples include the positive sample clusters 211 and 212 obtained through clustering, and the negative samples include the negative sample clusters 213 and 214 obtained through clustering. The distance between the clustering center of the positive sample cluster 211 and the anchor image 210 is L1, the distance between the clustering center of the positive sample cluster 212 and the anchor image 210 is L2, the distance between the clustering center of the negative sample cluster 213 and the anchor image 210 is L3, and the distance between the clustering center of the negative sample cluster 214 and the anchor image 210 is L4.

[0083] In this application, after clustering multiple positive samples, multiple positive sample clusters are obtained. The distance between the clustering center of the cluster most similar to the anchor image and the anchor image is set as L2, and the distances between the other positive samples in the multiple positive samples and the anchor image are set as L1 (note: Figure 2 The L2 shown is only the distance between the clustering center of the positive sample cluster 212 and the anchor image, and the distances between the other positive samples in the positive sample cluster 212 and the anchor image are L1). According to the redefined distances between the multiple positive samples and the anchor image, the anchor image and the multiple positive samples are brought closer. In the related art, it is considered that the distances between all positive samples and the anchor image are the same.

[0084] In this application, after clustering multiple negative samples, multiple negative sample clusters are obtained. Weights are assigned to each cluster based on the similarity between the clustering center of each cluster and the anchor image, and the anchor image and the negative samples are pulled apart according to the cluster weights. Figure 2 The distances L3 and L4 shown are the distances after weighting. In the related art, it is considered that the distances between all negative samples and the anchor image are the same.

[0085] During the usage phase of the image encoder, as Figure 2 shown, in this application, the usage phase of the image encoder is the search process for whole-slide pathology sections.

[0086] First, a whole-slide image (WSI) is cropped to obtain multiple tissue images (small images); then, the multiple tissue images are clustered to obtain multiple key images, and the multiple key images are jointly used to represent a WSI. Next, for one of the key images (small image A), the small image A is input into the image encoder to obtain the image feature vector of the small image A; finally, according to the image feature vector of the small image A, the database is queried to obtain small images A1 to AN, and the WSIs corresponding to the small images A1 to AN are used as the search results. All the multiple key images are used as query images to determine the WSIs from the database.

[0087] Optionally, the training device 21 of the above image encoder and the using device 22 of the image encoder may be computer devices with machine learning capabilities. For example, the computer device may be a terminal or a server.

[0088] Optionally, the training device 21 of the above image encoder and the using device 22 of the image encoder may be the same computer device, or the training device 21 of the image encoder and the using device 22 of the image encoder may be different computer devices. Moreover, when the training device 21 of the image encoder and the using device 22 of the image encoder are different devices, the training device 21 of the image encoder and the using device 22 of the image encoder may be devices of the same type. For example, the training device 21 of the image encoder and the using device 22 of the image encoder may both be servers; or, the training device 21 of the image encoder and the using device 22 of the image encoder may also be devices of different types. The above server may be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms. The above terminal may be a smart phone, in-vehicle terminal, smart TV, wearable device, tablet computer, notebook computer, desktop computer, smart speaker, smart watch, etc., but is not limited thereto. The terminal and the server may be directly or indirectly connected through wired or wireless communication methods, and this application does not make any restrictions here.

[0089] The following will be introduced in detail in the following order.

[0090] Training stage of the image encoder - 1;

[0091] - Relevant content of pulling the anchor image closer to the positive sample - 1-1;

[0092] - Relevant content of the first group loss function - 1-1-1;

[0093] - Relevant content of the second group loss function - 1-1-2;

[0094] - Relevant content of pulling the anchor image away from the negative sample - 1-2;

[0095] - Relevant content of the first weight loss function - 1-2-1;

[0096] - Relevant content of the second weight loss function - 1-2-2;

[0097] - Complete loss function related content - 1-3;

[0098] The use phase of the image encoder (the search process of the full-field pathology slice) – 2;

[0099] Related content of the first group loss function——1-1-1:

[0100] Figure 3 A training framework for an image encoder provided by an exemplary embodiment is shown, and the framework is applied to Figure 1 The image encoder training device 21 is shown for illustration.

[0101] Figure 3 It shows that: the first sample tissue image 301 obtains the first image 302 through data enhancement, the first image 302 obtains the first feature vector 306 through the first image encoder 304, and when multiple first sample tissue images 301 are input at the same time, the multiple first feature vectors will be distinguished into positive sample vectors 307 in the multiple first feature vectors and negative sample vectors 308 in the multiple first feature vectors; the first sample tissue image 301 obtains the second image 303 through data enhancement, and the second image 303 obtains the second feature vector 309 through the second image encoder 305; based on the positive sample vectors 307 and the second feature vector 309 in the multiple first feature vectors, a first sub-function 310 is generated; based on the negative sample vectors 308 and the second feature vector 309 in the multiple first feature vectors, a second sub-function 311 is generated; based on the first sub-function 310 and the second sub-function 311, a first group loss function 312 is constructed;

[0102] The first group loss function 312 is used to shorten the distance between the anchor image and the positive sample.

[0103] Figure 4 A flowchart of a training method for an image encoder provided by an exemplary embodiment is shown, wherein the method is applied to Figure 3 The training framework of the image encoder shown is illustrated as an example, and the method includes:

[0104] Step 401, obtaining a first sample tissue image;

[0105] The first sample tissue image, in this application, refers to an image used to train an image encoder, that is, a local area image (small image) within the WSI.

[0106] Combined with reference Figure 5 , image X is the first sample tissue image.

[0107] Step 402, performing data enhancement on the first sample tissue image to obtain a first image; inputting the first image into a first image encoder to obtain a first feature vector;

[0108] Data augmentation, also known as data amplification, aims to generate more data from limited data without substantially increasing the data. In one embodiment, the data augmentation method includes at least one of the following:

[0109] · Rotation / reflection transformation: Randomly rotate the image by a certain angle to change the orientation of the image content;

[0110] · Flip transformation: Flip the image along the horizontal or vertical direction;

[0111] · Scaling transformation: Enlarge or reduce the image by a certain ratio;

[0112] · Translation transformation: Translate the image in a certain way on the image plane;

[0113] · Specify the translation range and translation step length in a random or artificially defined manner, and translate along the horizontal or vertical direction to change the position of the image;

[0114] · Scale transformation: Enlarge or reduce the image according to the specified scale factor; or refer to the SIFT feature extraction idea, use the specified scale factor to filter the image to construct a scale space, and change the size or blur degree of the image content;

[0115] · Contrast transformation: In the HSV color space of the image, change the saturation S and V brightness components, keep the hue H unchanged, perform an exponential operation (the exponential factor is between 0.25 and 4) on the S and V components of each pixel to increase the illumination change;

[0116] · Noise perturbation: Randomly perturb each pixel RGB of the image, and the commonly used noise patterns are salt-and-pepper noise and Gaussian noise;

[0117] · Color change: Add random perturbations to the image channels;

[0118] · Randomly select a region of the input image and blacken it.

[0119] In this embodiment, the first sample tissue image is subjected to data augmentation to obtain a first image, and the first image encoder is used to extract features from the first image to obtain a first feature vector.

[0120] In one embodiment, the first image is input into the first image encoder to obtain a first intermediate feature vector; the first intermediate feature vector is input into the first MLP to obtain a first feature vector. Among them, the first MLP plays a transitional role and is used to improve the expression ability of the first image.

[0121] With reference to Figure 5 , the image X is subjected to data augmentation to obtain the image X p, and then apply the encoder h to the image X p to transform it into a high-level semantic space i.e., obtain the first intermediate feature vector h p , take the first intermediate feature vector h p and input it into the first MLP to obtain the first feature vector g p1 .

[0122] Step 403: Perform data augmentation on the first sample tissue image to obtain a second image; input the second image into a second image encoder to obtain a second feature vector;

[0123] In this embodiment, perform data augmentation on the first sample tissue image to obtain a second image, and extract features from the second image through a second image encoder to obtain a second feature vector.

[0124] In one embodiment, input the second image into a second image encoder to obtain a second intermediate feature vector; input the second intermediate feature vector into a second MLP to obtain a second feature vector. Among them, the second MLP plays a transitional role to improve the expression ability of the second image.

[0125] Combined with reference Figure 5 , perform data augmentation on the image X to obtain the image X q , and then apply the encoder h to the image X q to transform it into a high-level semantic space i.e., obtain the second intermediate feature vector h q , take the second intermediate feature vector h q and input it into the second MLP to obtain the second feature vector g q2 .

[0126] Step 404: Determine the first feature vector as the contrast vector for contrast learning, and determine the second feature vector as the anchor vector for contrast learning;

[0127] In this embodiment, determine the first feature vector as the contrast vector for contrast learning, and determine the second feature vector as the anchor vector for contrast learning. The contrast vector for contrast learning can be a positive sample vector or a negative sample vector.

[0128] Step 405: Cluster multiple first feature vectors of different first sample tissue images to obtain multiple first cluster centers;

[0129] In one embodiment, input multiple different first sample tissue images simultaneously, and cluster multiple first feature vectors of the multiple first sample tissue images to obtain multiple first cluster centers. Optionally, the multiple different first sample tissue images are sample tissue images of the same training batch.

[0130] In one embodiment, the first feature vectors of different first sample tissue images will be clustered into S categories, and the S first clustering centers of the S categories are denoted as where j ∈ [1, …, S].

[0131] With reference to Figure 5 , which shows one of the first clustering centers among the multiple first clustering centers of the first feature vectors of different first sample tissue images.

[0132] Step 406, determine the feature vector with the largest similarity value between the multiple first clustering centers and the second feature vector as the positive sample vector among the multiple first feature vectors;

[0133] In one embodiment, use the first clustering center closest to the second feature vector among the S first clustering centers as the positive sample vector, denoted as

[0134] Step 407, determine the remaining first feature vectors as the negative sample vectors among the multiple first feature vectors;

[0135] In one embodiment, use the feature vectors among the S first clustering centers except as the negative sample vectors, denoted as

[0136] where the remaining first feature vectors refer to the feature vectors among the multiple first feature vectors except the feature vector with the largest similarity value between it and the second feature vector.

[0137] Step 408, generate a first sub-function based on the second feature vector and the positive sample vector among the multiple first feature vectors;

[0138] In one embodiment, the first sub-function is denoted as

[0139] where g q2 serves as the anchor vector in contrastive learning, serves as the positive sample vector in contrastive learning.

[0140] Step 409, generate a second sub-function based on the second feature vector and the negative sample vectors among the multiple first feature vectors;

[0141] In one embodiment, the second sub-function is denoted as

[0142] where the second feature vector g q2 serves as the anchor vector in contrastive learning, serves as the negative sample vector in contrastive learning.

[0143] Step 410, generating a first group loss function based on the first sub-function and the second sub-function;

[0144] In one embodiment, the first group loss function is expressed as:

[0145]

[0146] in, Represents the first group loss function, and log represents the logarithmic operation.

[0147] Step 411: train the first image encoder and the second image encoder based on the first group loss function; and determine the second image encoder as the image encoder finally obtained by training.

[0148] According to the first group loss function, the first image encoder and the second image encoder can be trained.

[0149] In this embodiment, the second image encoder is determined as the image encoder finally obtained by training.

[0150] In summary, by further distinguishing the positive samples identified in the relevant technology and further distinguishing the "degree of positivity" of the positive samples among the positive samples, the loss function used in contrastive learning (also called contrastive learning paradigm) can more accurately bring the anchor image and the positive sample closer, thereby better training the image encoder. The trained image encoder can better learn the common features between the anchor image and the positive sample.

[0151] Above Figure 3 and Figure 4 It is shown that an image encoder is trained by a feature vector sample triplet, and the comparative learning sample triplet includes (anchor vector, positive vector, negative vector). In another embodiment, it is also possible to train the image encoder by multiple feature vector sample triplets at the same time. The following will introduce the training of the image encoder by two feature vector sample triplets (anchor vector 1, positive vector 1, negative vector 1) (anchor vector 2, positive vector 2, negative vector 2) at the same time, wherein anchor vector 1 and anchor vector 2 are different vectors obtained by respectively performing data enhancement on the first sample tissue image and respectively passing through different image encoders and different MLPs. It should be noted that the present application does not limit the number of feature vector sample triplets to be constructed.

[0152] Related content of the second group loss function——1-1-2:

[0153] Figure 6 A training framework for an image encoder provided by an exemplary embodiment is shown, and the framework is applied to Figure 1 The image encoder training device 21 is shown for illustration.

[0154] Figure 6 It is shown that: the first sample tissue image 301 obtains the first image 302 through data enhancement, the first image 302 obtains the first feature vector 306 through the first image encoder 304, and when multiple first sample tissue images 301 are input at the same time, the multiple first feature vectors will be distinguished into positive sample vectors 307 in the multiple first feature vectors and negative sample vectors 308 in the multiple first feature vectors; the first sample tissue image 301 obtains the second image 303 through data enhancement, and the second image 303 obtains the second feature vector 309 through the second image encoder 305; based on the positive sample vectors 307 and the second feature vector 309 in the multiple first feature vectors, a first sub-function 310 is generated; based on the negative sample vectors 308 and the second feature vector 309 in the multiple first feature vectors, a second sub-function 311 is generated; based on the first sub-function 310 and the second sub-function 311, a first group loss function 312 is constructed.

[0155] and Figure 3 The difference between the training framework shown is that Figure 6 It also shows that when multiple first sample tissue images 301 are input simultaneously, multiple second eigenvectors will be distinguished into positive sample vectors 313 among multiple second eigenvectors and negative sample vectors 314 among multiple second eigenvectors; based on the positive sample vectors 313 among multiple second eigenvectors and the first eigenvector 306, a third sub-function 315 is generated; based on the negative sample vectors 314 among multiple second eigenvectors and the first eigenvector 306, a fourth sub-function 316 is generated; based on the third sub-function 315 and the fourth sub-function 316, a second group loss function 317 is constructed; wherein the second group loss function 317 is used to shorten the distance between the anchor image and the positive sample.

[0156] based on Figure 4 The training method of the image encoder shown in Figure 7 exist Figure 4 Based on the method steps, steps 412 to 419 are further provided to Figure 7 The method shown is applied to Figure 6 The training framework of the image encoder shown is illustrated as an example, and the method includes:

[0157] Step 412, determining the second feature vector as a contrast vector for contrastive learning, and determining the first feature vector as an anchor vector for contrastive learning;

[0158] In this embodiment, the second feature vector is determined as a contrast vector for contrastive learning, and the first feature vector is determined as an anchor vector for contrastive learning. The contrast vector for contrastive learning can be a positive sample vector or a negative sample vector.

[0159] Step 413: Cluster the multiple second feature vectors of different first sample tissue images to obtain multiple second cluster centers;

[0160] In one embodiment, multiple different first sample tissue images are input simultaneously, and the multiple second feature vectors of the multiple first sample tissue images are clustered to obtain multiple second cluster centers. Optionally, the multiple different first sample tissue images are sample tissue images of the same training batch.

[0161] In one embodiment, the second feature vectors of different first sample tissue images will be clustered into S categories, and the S second cluster centers of the S categories are denoted as where j ∈ [1, …, S].

[0162] With reference to Figure 5 , which shows one of the multiple second cluster centers of the second feature vectors of different first sample tissue images.

[0163] Step 414: Determine the feature vector with the largest similarity value between the multiple second cluster centers and the first feature vector as the positive sample vector among the multiple second feature vectors;

[0164] In one embodiment, the second cluster center closest to the first feature vector among the S second cluster centers is used as the positive sample vector, denoted as

[0165] Step 415: Determine the remaining second feature vectors as the negative sample vectors among the multiple second feature vectors;

[0166] In one embodiment, the feature vectors other than among the S second cluster centers are used as the negative sample vectors, denoted as

[0167] where the remaining second feature vectors refer to the feature vectors among the multiple second feature vectors other than the feature vector with the largest similarity value between the multiple second feature vectors and the first feature vector.

[0168] Step 416: Generate a third sub-function based on the first feature vector and the positive sample vector among the multiple second feature vectors;

[0169] In one embodiment, the third sub-function is denoted as

[0170] Step 417: Generate a fourth sub-function based on the first feature vector and the negative sample vectors among the multiple second feature vectors;

[0171] In one embodiment, the fourth sub-function is denoted as

[0172] Among them, the first eigenvector g p1 serves as the anchor vector in contrastive learning, and serves as the negative sample vector in contrastive learning.

[0173] Step 418: Generate a second group loss function based on the third sub-function and the fourth sub-function;

[0174] In one embodiment, the second group loss function is expressed as:

[0175]

[0176] Among them, represents the second group loss function, and log represents the logarithm operation.

[0177] Step 419: Train the first image encoder and the second image encoder based on the second group loss function; determine the first image encoder as the finally trained image encoder.

[0178] Train the first image encoder and the second image encoder according to the second group loss function; determine the first image encoder as the finally trained image encoder.

[0179] In one embodiment, by combining the first group loss function obtained in the above step 410 and the second group loss function obtained in step 418, a complete group loss function can be constructed:

[0180]

[0181] Among them, is the complete group loss function. Train the first image encoder and the second image encoder according to the complete group loss function. Determine the first image encoder and the second image encoder as the finally trained image encoders.

[0182] Optionally, after step 419, there is further step 420 of updating the parameters of the third image encoder in a weighted manner according to the parameters shared between the first image encoder and the second image encoder.

[0183] Schematically, the formula for updating the parameters of the third image encoder is as follows:

[0184] θ′ = m·θ′ + (1 - m)·θ; (4)

[0185] Among them, θ' on the left side of formula (4) represents the parameters of the updated third image encoder, and θ' on the right side of formula (4) represents the parameters of the third image encoder before update. θ represents the parameters shared by the first image encoder and the second image encoder, and m is a constant. Optionally, m is 0.99.

[0186] In summary, by constructing two feature vector sample triples (the second feature vector, the positive vector among multiple first feature vectors, the negative vector among multiple first feature vectors), (the first feature vector, the positive vector among multiple second feature vectors, the negative vector among multiple second feature vectors), the encoding effect of the trained image encoder is further improved, and the constructed complete group loss function is more robust than the first group loss function or the second group loss function.

[0187] The content of training the image encoder based on the group loss function has been fully introduced above. Among them, the image encoder includes a first image encoder, a second image encoder, and a third image encoder. In the following, training the image encoder based on the weight loss function will also be introduced.

[0188] Relevant content of the first weight loss function - 1 - 2 - 1:

[0189] Figure 8 Shows the training framework of the image encoder provided by an exemplary embodiment. Taking this framework applied to Figure 2 the training device 21 of the image encoder shown as an example.

[0190] Figure 8 Shows that: multiple second sample tissue images 801 pass through the third image encoder 805 to generate multiple feature vectors 807; the first sample tissue image 802 undergoes data augmentation to obtain the third image 803, and the third image 803 passes through the third image encoder 805 to generate the third feature vector 808; the first sample tissue image 802 undergoes data augmentation to obtain the first image 804, and the first image 804 passes through the first image encoder 806 to generate the fourth feature vector 809; based on the third feature vector 808 and the fourth feature vector 809, the fifth sub - function 810 is generated; based on the multiple feature vectors 807 and the fourth feature vector 809, the sixth sub - function 811 is generated; based on the fifth sub - function 810 and the sixth sub - function 811, the first weight loss function 812 is generated.

[0191] Among them, the first weight loss function 812 is used to widen the distance between the anchor image and the negative sample.

[0192] Figure 9 Shows the flowchart of the training method of the image encoder provided by an exemplary embodiment. Taking this method applied to Figure 8An example is given for the training framework of the image encoder shown, and the method includes:

[0193] Step 901: Obtain a first sample tissue image and multiple second sample tissue images, where the second sample tissue images are negative samples in contrastive learning;

[0194] The first sample tissue image refers to the image used to train the image encoder in this application; the second sample tissue image refers to the image used to train the image encoder in this application. Among them, the first sample tissue image and the second sample tissue image are different small images, that is, the first sample tissue image and the second sample tissue image are not small images X1 and X2 obtained through data augmentation, but are respectively small Figure X and small image Y. Figure X

[0195] In this embodiment, the second sample tissue image is used as the negative sample in contrastive learning. Contrastive learning aims to shorten the distance between the anchor image and the positive sample and widen the distance between the anchor image and the negative sample.

[0196] Combined with reference Figure 10 , image X is the first sample tissue image, and the sub-container of the negative samples is the container that holds multiple feature vectors of multiple second sample tissue images.

[0197] Step 902: Perform data augmentation on the first sample tissue image to obtain a third image; input the third image into a third image encoder to obtain a third feature vector; the third image is the positive sample in contrastive learning;

[0198] In this embodiment, perform data augmentation on the first sample tissue image to obtain a third image, and use the third image as the positive sample in contrastive learning.

[0199] Combined with reference Figure 10 , image X is augmented to obtain image X k , and then the encoder f is applied to transform image X k to the high-level semantic space that is, the third feature vector f k is obtained.

[0200] Step 903: Perform data augmentation on the first sample tissue image to obtain a first image; input the first image into a first image encoder to obtain a fourth feature vector; the first image is the anchor image in contrastive learning;

[0201] In this embodiment, perform data augmentation on the first sample tissue image to obtain a first image, and use the first image as the anchor image in contrastive learning.

[0202] ​In one embodiment, the first image is input into the first image encoder to obtain a first intermediate feature vector; the first intermediate feature vector is input into a third MLP (Multilayer Perceptron) to obtain a fourth feature vector. Among them, the third MLP plays a transitional role and is used to improve the expression ability of the first image.

[0203] Combined with reference Figure 10 , the image X is subjected to data augmentation to obtain the image X p , and then the encoder h is applied to the image X p to transform it into a high-level semantic space that is, the first intermediate feature vector h is obtained p , and the first intermediate feature vector h p is input into the third MLP to obtain the fourth feature vector g p2 .

[0204] Step 904, input multiple second sample tissue images into a third image encoder to obtain multiple feature vectors of the multiple second sample tissue images; cluster the multiple feature vectors to obtain multiple cluster centers; generate multiple weights based on the similarity values between the multiple cluster centers and the third feature vector;

[0205] In this embodiment, the second sample tissue images are negative samples in contrast learning. The multiple feature vectors of the multiple second sample tissue images are clustered, and multiple weights are assigned to the multiple feature vectors according to the similarity values between the multiple cluster centers and the third feature vector.

[0206] Combined with reference Figure 10 , multiple feature vectors of multiple second sample tissue images are stored in the sub-container of the negative sample. After the multiple second sample tissue images pass through the encoder f, they are pushed onto the stack and placed in the storage queue. In the storage queue, the multiple feature vectors in the queue are clustered into Q categories by K-means clustering, and then Q sub-queues are constructed. The cluster center of each sub-queue is denoted as c j (j = 1,..., Q). Then, calculate the similarity score between each cluster center and the third feature vector f k to determine potential incorrect negative samples. Finally, the weight of each feature vector in the storage queue is obtained The calculation is as follows:

[0207]

[0208] where δ() is a discriminant function. If the two inputs are the same, δ() outputs 1; otherwise, δ() outputs 0. In this embodiment, δ() is used to determine whether the cluster center c j of the j-th class is similar to f k , w is the assigned weight, and w ∈ [0, 1].

[0209] In one embodiment, the magnitude of the weights of multiple cluster centers is negatively correlated with the similarity value between the cluster center and the third feature vector; for the j-th cluster center among the multiple cluster centers, the feature vectors included in the category to which the j-th cluster center belongs correspond to the same weight. In formula (1), for the multiple feature vectors corresponding to the category to which the cluster center that is j more similar to f k is, a smaller weight value w is assigned, and for the multiple feature vectors corresponding to the category to which the cluster center that is

[0210] Schematically, after clustering the multiple feature vectors of multiple second sample tissue images, 3 categories are obtained, and the cluster centers are c1, c2, and c3 respectively. Among them, the category to which the cluster center c1 belongs includes feature vector 1, feature vector 2, and feature vector 3; the category to which the cluster center c2 belongs includes feature vector 4, feature vector 5, and feature vector 6; the category to which the cluster center c3 belongs includes feature vector 7, feature vector 8, and feature vector 9.

[0211] If the similarity values of the cluster centers c1, c2, and c3 with f k are arranged in descending order, then the weights corresponding to the categories to which the cluster centers c1, c2, and c3 belong are arranged in ascending order. Moreover, feature vector 1, 2, and 3 correspond to the same weight, feature vector 4, 5, and 6 correspond to the same weight, and feature vector 7, 8, and 9 correspond to the same weight.

[0212] In one embodiment, when the first sample tissue image belongs to the first sample tissue image in the first training batch, the multiple feature vectors of multiple second sample tissue images are clustered to obtain multiple cluster centers of the first training batch.

[0213] In another embodiment, when the first sample tissue image belongs to the first sample tissue image in the n-th training batch, the multiple cluster centers corresponding to the (n - 1)-th training batch are updated to the multiple cluster centers corresponding to the n-th training batch, where n is a positive integer greater than 1.

[0214] Optionally, for the j-th cluster center among the multiple cluster centers of the (n - 1)-th training batch, based on the first sample tissue images belonging to the j-th category in the n-th training batch, the j-th cluster center of the (n - 1)-th training batch is updated to obtain the j-th cluster center of the n-th training batch, where i is a positive integer.

[0215] With reference to Figure 10 , according to the j-th cluster center c j of the (n - 1)-th training batch, the j-th cluster center c j* of the n-th training batch is updated, and the formula is as follows:

[0216]

[0217] where c j* represents the j-th cluster center of the n-th training batch after update; m c represents the weight used for update, m c ∈[0,1]; represents the set of features belonging to the j-th class within multiple third feature vectors (multiple f k ) of multiple first sample tissue images (multiple images X) of the n-th training batch. represents the i-th feature vector within multiple third feature vectors (multiple f k ) of the n-th training batch belonging to the j-th class. is used to calculate the mean value of the features of multiple third feature vectors (multiple f k ) of the n-th training batch belonging to the j-th class.

[0218] In one embodiment, within each training cycle, all cluster centers will be updated by reclustering all negative sample feature vectors in the repository.

[0219] It can be understood that updating the multiple cluster centers of the (n - 1)-th training batch to the multiple cluster centers of the n-th training batch aims to prevent the distance between the negative sample feature vectors in the negative sample container and the input first sample tissue image from getting farther and farther.

[0220] As the image encoder is continuously trained, the image encoder's effect of pulling the anchor image and negative samples farther and farther is getting better. Suppose the image encoder pulls the image X of the previous training batch and the negative samples to a first distance, the image encoder pulls the image X of the current training batch and the negative samples to a second distance, the second distance is greater than the first distance, and the image encoder pulls the image X of the subsequent training batch and the negative samples to a third distance, the third distance is greater than the second distance. However, if the negative sample images are not updated (i.e., the cluster centers are updated), the increase in the third distance compared to the second distance will be less than the increase in the second distance compared to the first distance, and the training effect of the image encoder will gradually deteriorate. If the negative sample images are updated (i.e., the cluster centers are updated), then the distance between the updated negative sample images and the image X will be appropriately shortened, balancing the gradually improving pulling-away effect of the image encoder, enabling the image encoder to maintain long-term and frequent training, and ultimately the finally trained image encoder can also have a better effect.

[0221] Step 905, generate a fifth sub-function for characterizing the error between the anchor image and the positive sample based on the third feature vector and the fourth feature vector;

[0222] In this embodiment, a fifth sub-function is generated according to the third eigenvector and the fourth eigenvector, and the fifth sub-function is used to characterize the error between the anchor image and the positive sample.

[0223] Combined with reference Figure 10 , the fifth sub-function can be expressed as exp(g p2 ·f k / τ). It can be seen that the fifth sub-function is composed of the third eigenvector f k and the fourth eigenvector g p2 .

[0224] Step 906: Based on the fourth eigenvector and multiple eigenvectors, combine multiple weights to generate a sixth sub-function for characterizing the error between the anchor image and the negative sample;

[0225] In this embodiment, according to the fourth eigenvector and the multiple eigenvectors of multiple second sample tissue images, combine multiple weights to generate a sixth sub-function, and the sixth sub-function is used to characterize the error between the anchor image and the negative sample.

[0226] Combined with reference Figure 10 , the sixth sub-function can be expressed as where represents the weight of the i-th negative sample eigenvector (i.e., the eigenvector of the second sample tissue image), represents the i-th negative sample eigenvector, and there are a total of K negative sample eigenvectors in the negative sample container, and g p2 represents the eigenvector of the anchor image (i.e., the fourth eigenvector).

[0227] Step 907: Based on the fifth sub-function and the sixth sub-function, generate a first weight loss function;

[0228] Combined with reference Figure 10 , the first weight loss function can be expressed as:

[0229]

[0230] where represents the first weight loss function, and log represents the logarithm operation.

[0231] Step 908: Based on the first weight loss function, train the first image encoder and the third image encoder;

[0232] Train the first image encoder and the third image encoder according to the first weight loss function.

[0233] Step 909: Based on the first image encoder, update the third image encoder.

[0234] Update the third image encoder based on the first image encoder. Optionally, update the parameters of the third image encoder in a weighted manner according to the parameters of the first image encoder.

[0235] Schematically, the formula for updating the parameters of the third image encoder is as follows:

[0236] θ′ = m·θ′+(1 - m)·θ; (8)

[0237] Among them, θ′ on the left side of formula (8) represents the parameters of the updated third image encoder, θ′ on the right side of formula (8) represents the parameters of the third image encoder before update, θ represents the parameters of the first image encoder, and m is a constant. Optionally, m = 0.99.

[0238] In summary, by assigning weights to the negative samples identified in the related art and further distinguishing the "negative degree" of the negative samples among the negative samples, the loss function used in contrastive learning (also known as the contrastive learning paradigm) can more precisely separate the anchor image from the negative samples, reduce the influence of potential false negative samples, and thus can better train the image encoder. The trained image encoder can better distinguish the different features between the anchor image and the negative samples.

[0239] The above Figure 8 and Figure 9 show the training of the third image encoder through a sample triple, where the sample triple includes (anchor image, positive sample, negative sample). In another embodiment, it is also possible to train the third image encoder through multiple sample triples at the same time. Hereinafter, the training of the third image encoder through two sample triples (anchor image 1, positive sample, negative sample) (anchor image 2, positive sample, negative sample) will be introduced. Anchor image 1 and anchor image 2 are images obtained by performing data augmentation on the same small image respectively. It should be noted that the present application does not limit the specific number of sample triples constructed.

[0240] Relevant content of the second weight loss function - 1 - 2 - 2:

[0241] Figure 11 shows the training framework of the image encoder provided by an exemplary embodiment, taking the application of this framework to the Figure 1 shown image encoder training device 21 as an example for illustration.

[0242] Figure 11It shows that: multiple second sample tissue images 801 pass through a third image encoder 805 to generate multiple feature vectors 807; a first sample tissue image 802 undergoes data augmentation to obtain a third image 803, and the third image 803 passes through the third image encoder 805 to generate a third feature vector 808; the first sample tissue image 802 undergoes data augmentation to obtain a first image 804, and the first image 804 passes through a first image encoder 806 to generate a fourth feature vector 809; based on the fourth feature vector 809 and the third feature vector 808, a fifth sub-function 810 is generated; based on the multiple feature vectors 807 and the fourth feature vector 809, a sixth sub-function 811 is generated; based on the fifth sub-function 810 and the sixth sub-function 811, a first weight loss function 812 is generated.

[0243] and Figure 8 the difference from the training framework shown is that Figure 11 it also shows that: the first sample tissue image 802 undergoes data augmentation to obtain a second image 813, and the second image 813 passes through a second image encoder 814 to obtain a fifth feature vector 815; the fifth feature vector 815 and the third feature vector 808 generate a seventh sub-function 816; the fifth feature vector 815 and the multiple feature vectors 807 generate an eighth sub-function 817; the seventh sub-function 816 and the eighth sub-function 817 generate a second weight loss function 818.

[0244] Among them, the second weight loss function 818 is used to increase the distance between the anchor image and the negative samples.

[0245] Based on Figure 9 the training method of the image encoder shown, Figure 12 on the basis of Figure 9 the method steps, steps 910 to 914 are further provided to Figure 12 illustrate by way of example the application of the method shown to Figure 11 the training framework of the image encoder shown. The method includes:

[0246] Step 910: Perform data augmentation on the first sample tissue image to obtain a second image; input the second image into the second image encoder to obtain a fifth feature vector; the second image is the anchor image in contrastive learning;

[0247] In this embodiment, the first sample tissue image is subjected to data augmentation to obtain a second image, and the second image is used as the anchor image in contrastive learning.

[0248] In one embodiment, the second image is input into the second image encoder to obtain a second intermediate feature vector; the second intermediate feature vector is input into a fourth MLP to obtain a fifth feature vector. Among them, the fourth MLP plays a transitional role and is used to improve the expression ability of the second image.

[0249] Combined reference Figure 10 , the image X is subjected to data augmentation to obtain the image X q , and then the encoder h is applied to the image X q to transform it into the high-level semantic space i.e., the second intermediate feature vector h is obtained q , the second intermediate feature vector h q is input into the fourth MLP to obtain the fifth feature vector g q1 .

[0250] Step 911, based on the third feature vector and the fifth feature vector, generate a seventh sub-function for characterizing the error between the anchor image and the positive sample;

[0251] In this embodiment, according to the third feature vector and the fifth feature vector, a seventh sub-function is generated, and the seventh sub-function is used to characterize the error between the anchor image and the positive sample.

[0252] Combined reference Figure 10 , the seventh sub-function can be expressed as exp(g q1 ·f k / τ), and it can be seen that the seventh sub-function is composed of the three-feature vector f k and the fifth feature vector g q1 .

[0253] Step 912, based on the fifth feature vector and multiple feature vectors, combined with multiple weights, generate an eighth sub-function for characterizing the error between the anchor image and the negative sample;

[0254] In this embodiment, according to the fifth feature vector and the multiple feature vectors of multiple second sample tissue images, an eighth sub-function is generated by combining multiple weights, and the eighth sub-function is used to characterize the error between the anchor image and the negative sample.

[0255] Combined reference Figure 10 , the eighth sub-function can be expressed as where represents the weight of the i-th negative sample feature vector (i.e., the feature vector of the second sample tissue image), represents the i-th negative sample feature vector, and there are a total of K negative sample feature vectors in the negative sample container, and g q1 represents the feature vector of the anchor image (i.e., the fifth feature vector).

[0256] Step 913, based on the seventh sub-function and the eighth sub-function, generate a second weight loss function;

[0257] Combined reference Figure 10 , the second weight loss function can be expressed as:

[0258]

[0259] Among them, represents the second weight loss function, and log represents the logarithm operation.

[0260] Step 914: Based on the second weight loss function, train the second image encoder and the third image encoder.

[0261] Train the second image encoder and the third image encoder according to the second weight loss function.

[0262] In one embodiment, combining with the first weight loss function obtained in the above step 908, a complete weight loss function can be constructed as follows:

[0263]

[0264] Among them, is the complete weight loss function. Train the first image encoder, the second image encoder, and the third image encoder according to the complete weight loss function.

[0265] Optionally, in the above step 909, "updating the third image encoder based on the first image encoder" can be replaced with "updating the parameters of the third image encoder in a weighted manner according to the parameters shared between the first image encoder and the second image encoder", that is, θ in formula (8) in step 909 represents the parameters shared between the first image encoder and the second image encoder, and the third image encoder is slowly updated through the parameters shared between the first image encoder and the second image encoder.

[0266] In summary, the above solution constructs two sample triplets (the first image, the third image, multiple second sample tissue images), (the second image, the third image, multiple second sample tissue images), where the first image is the anchor image 1 and the second image is the anchor image 2, further improving the encoding effect of the trained image encoder, and the constructed complete weight loss function is more robust than the first weight loss function or the second weight loss function.

[0267] Related content of the complete loss function - 1 - 3:

[0268] From the above Figures 3 to 7 , it is possible to train the third image encoder through the group loss function; from the above Figures 8 to 12 , it is possible to train the third image encoder through the weight loss function.

[0269] In an alternative embodiment, the third image encoder can be trained jointly by the group loss function and the weight loss function. Please refer to Figure 13, which shows a schematic diagram of the training architecture of the first image encoder provided by an exemplary embodiment of the present application.

[0270] Relevant part of the group loss function:

[0271] The image X is subjected to data augmentation to obtain the image X p , the image X p passes through the encoder h to obtain the first intermediate feature vector h p , the first intermediate feature vector h p passes through the first MLP to obtain the first feature vector g p1 ; the image X is subjected to data augmentation to obtain the image X q , the image X q passes through the encoder h to obtain the second intermediate feature vector h q , the first intermediate feature vector h q passes through the second MLP to obtain the second feature vector g q2 ;

[0272] In the same training batch, aggregate the multiple first feature vectors g of multiple first sample tissue images p1 , to obtain multiple first cluster centers; determine the first cluster center closest to the second feature vector g of a first sample tissue image among the multiple first cluster centers as the positive sample vector; determine the remaining feature vectors of the multiple first cluster centers as negative sample vectors; based on the positive sample vector and the second feature vector g q2 construct a sub-function for characterizing the error between the positive sample vector and the anchor vector; based on the negative sample vector and the second feature vector g q2 construct a sub-function for characterizing the error between the negative sample vector and the anchor vector; combine the two sub-functions to form the first group loss function; q2 construct a sub-function for characterizing the error between the negative sample vector and the anchor vector; combine the two sub-functions to form the first group loss function;

[0273] In the same training batch, aggregate the multiple second feature vectors g of multiple first sample tissue images q2 , to obtain multiple second cluster centers; determine the second cluster center closest to the first feature vector g of a first sample tissue image among the multiple second cluster centers as the positive sample vector; determine the remaining feature vectors of the multiple second cluster centers as negative sample vectors; based on the positive sample vector and the first feature vector g p1 construct a sub-function for characterizing the error between the positive sample vector and the anchor vector; based on the negative sample vector and the first feature vector g p1 construct a sub-function for characterizing the error between the negative sample vector and the anchor vector; combine the two sub-functions to form the second group loss function; p1 construct a sub-function for characterizing the error between the negative sample vector and the anchor vector; combine the two sub-functions to form the second group loss function;

[0274] Train the first image encoder and the second image encoder using a group loss function obtained by combining a first group loss function and a second group loss function. Update the third image encoder according to the first image encoder and the second image encoder.

[0275] Relevant part of the weight loss function:

[0276] The image X is data-augmented to obtain the image X k , the image X k passes through the encoder f to obtain the third feature vector f k ; the image X is data-augmented to obtain the image X p , the image X p passes through the encoder h to obtain the first intermediate feature vector h p , the first intermediate feature vector h p passes through the third MLP to obtain the fourth feature vector g p2 ; the image X is data-augmented to obtain the image X q , the image X q passes through the encoder h to obtain the second intermediate feature vector h q , the first intermediate feature vector h q passes through the fourth MLP to obtain the fifth feature vector g p1 ;

[0277] Multiple second sample tissue images are input into the encoder f and placed in the storage queue through a push operation. In the storage queue, the negative sample feature vectors in the queue are clustered into Q categories by K-means clustering, and then Q sub-queues are constructed. Based on the similarity value between each cluster center and f k , assign weights to each cluster center;

[0278] Based on the Q cluster centers and the fourth feature vector g p2 construct a sub-function for characterizing negative samples and anchor images; based on the third feature vector f k and the fourth feature vector g p2 construct a sub-function for characterizing positive samples and anchor images; combine the two sub-functions to form the first weight loss function;

[0279] Based on the Q cluster centers and the fifth feature vector g p1 construct a sub-function for characterizing negative samples and anchor images; based on the third feature vector f k and the fifth feature vector g p1 construct a sub-function for characterizing positive samples and anchor images; combine the two sub-functions to form the second weight loss function;

[0280] Based on the weight loss function obtained by combining the first weight loss function and the second weight loss function, train the first image encoder, the second image encoder, and the third image encoder, and slowly update the parameters of the third image encoder through the parameters shared by the first image encoder and the second image encoder.

[0281] Combined with the relevant part of the group loss function for the weight loss function:

[0282] It can be understood that training the image encoder based on the weight loss function and based on the group loss function, both of which determine the similarity value based on clustering and reassign the positive and negative sample hypotheses. The above weight loss function is used to correct the positive and negative sample hypotheses of the negative samples in the related technology, and the above group loss function is used to correct the positive and negative sample hypotheses of the positive samples in the related technology.

[0283] In Figure 13 In the shown training architecture, the weight loss function and the group loss function are combined through hyperparameters, expressed as:

[0284]

[0285] Among them, the on the left side of formula (11) is the final loss function, is the weight loss function, is the group loss function, and λ is a hyperparameter that adjusts the contributions of the two loss functions.

[0286] In summary, the final loss function is jointly constructed by the weight loss function and the group loss function. Compared with a single weight loss function or a single group loss function, the final loss function will be more robust, and the image encoder finally trained will have a better encoding effect. The features of the small images obtained by feature extraction using the image encoder can better represent the small images.

[0287] Usage stage of the image encoder - 2:

[0288] The training stage of the image encoder has been introduced above. In the following, the usage stage of the image encoder will be introduced. In an embodiment provided by the present application, the image encoder will be used in the scenario of WSI image search. Figure 14 Shows the flowchart of the search method for whole-slide pathology sections provided by an exemplary embodiment of the present application. Taking the method applied to Figure 1 the usage device 22 of the shown image encoder as an example, at this time, the usage device 22 of the image encoder can also be called the search device for whole-slide pathology sections.

[0289] Step 1401, obtain the whole-slide pathology section and crop the whole-slide pathology section into multiple tissue images;

[0290] Whole-slide pathology image (WSI). WSI is obtained by scanning traditional pathology slides with a digital scanner, collecting high-resolution images, and then seamlessly stitching the fragmented images collected by a computer to produce a visualized digital image. In this application, WSI is often referred to as a large image.

[0291] Tissue image, which refers to a local tissue area within the WSI. In this application, tissue images are often referred to as small images.

[0292] In one embodiment, in the preprocessing stage of the WSI, the foreground tissue area within the WSI is extracted through threshold technology, and then the foreground tissue area of the WSI is cropped into multiple tissue images based on the sliding window technology.

[0293] Step 1402: Generate multiple image feature vectors of multiple tissue images through an image encoder;

[0294] In one embodiment, generate multiple image feature vectors of multiple tissue images through the second image encoder trained by the method embodiment shown above; at this time, the second image encoder is trained based on the first group loss function. Or, Figure 4 Generate multiple image feature vectors of multiple tissue images through the first image encoder (or the second image encoder) trained by the method embodiment shown above; at this time, the first image encoder (or the second image encoder) is trained based on the first group loss function and the second group loss function; or,

[0295] Generate multiple image feature vectors of multiple tissue images through the third image encoder trained by the method embodiment shown above; at this time, the third image encoder is trained based on the first weight loss function; or, Figure 7 Generate multiple image feature vectors of multiple tissue images through the third image encoder trained by the method embodiment shown above; at this time, the third image encoder is trained based on the first weight loss function and the second weight loss function; or,

[0296] Generate multiple image feature vectors of multiple tissue images through the third image encoder trained by the embodiment shown above; at this time, the third image encoder is trained based on the group loss function and the weight loss function. Figure 9 Generate multiple image feature vectors of multiple tissue images through the third image encoder trained by the embodiment shown above; at this time, the third image encoder is trained based on the group loss function and the weight loss function.

[0297] Generate multiple image feature vectors of multiple tissue images through the third image encoder trained by the embodiment shown above; at this time, the third image encoder is trained based on the group loss function and the weight loss function. Figure 12 Generate multiple image feature vectors of multiple tissue images through the third image encoder trained by the embodiment shown above; at this time, the third image encoder is trained based on the group loss function and the weight loss function.

[0298] Generate multiple image feature vectors of multiple tissue images through the third image encoder trained by the embodiment shown above; at this time, the third image encoder is trained based on the group loss function and the weight loss function. Figure 13 Generate multiple image feature vectors of multiple tissue images through the third image encoder trained by the embodiment shown above; at this time, the third image encoder is trained based on the group loss function and the weight loss function.

[0299] Step 1403: Determine multiple key images from multiple tissue images by clustering multiple image feature vectors;

[0300] In one embodiment, multiple image feature vectors of multiple tissue images are clustered to obtain multiple first - type clusters; the multiple cluster centers of the multiple first - type clusters are respectively determined as the multiple image feature vectors of multiple key images, that is, multiple key images are determined from multiple tissue images.

[0301] In another embodiment, multiple image feature vectors of multiple tissue images are clustered to obtain multiple first - type clusters, and then, reclustering is performed. For a target first - type cluster among the multiple first - type clusters, based on the position features of the multiple tissue images corresponding to the target first - type cluster in their respective whole - slide pathology sections, multiple second - type clusters are obtained by clustering; the multiple cluster centers corresponding to the multiple second - type clusters included in the target first - type cluster are determined as the image feature vectors of the key images; wherein, the target first - type cluster is any one of the multiple first - type clusters.

[0302] Schematically, the K - means clustering method is used for clustering. In the first clustering, multiple image feature vectors f all are clustered into K1 different categories, denoted as F i , i = 1, 2, …, K1. In the second clustering, within each cluster F i , taking the spatial coordinate information of multiple tissue images as features, it is further clustered into K2 categories, where K2 = round(R·N), R is a proportionality parameter, optionally, R is 20%; N is the number of small images in the cluster F i . Based on the above two - layer clustering, finally, K1*K2 cluster centers will be obtained, and the tissue images corresponding to the K1*K2 cluster centers are used as K1*K2 key images, and the K1*K2 key images are used as the global representation of the WSI. In some embodiments, the key images are often referred to as mosaic images.

[0303] Step 1404: Based on the image feature vectors of multiple key images, multiple candidate image packs are retrieved from the database. The multiple candidate image packs correspond one - to - one with the multiple key images, and any one candidate image pack contains at least one candidate tissue image;

[0304] From the above step 1404, it can be obtained that WSI = {P1, P2, …, P i , …, P k}, where P i and k respectively represent the feature vector of the i - th key image and the total number of key images in the WSI, and both i and k are positive integers. When searching the WSI, each key image will be used as a query image one by one to generate candidate image packs, and a total of k candidate image packs are generated, denoted as where the i - th candidate image pack b ijAnd $t$ respectively represent the $j$-th candidate tissue image and the total number of candidate tissue images inside, where $j$ is a positive integer.

[0305] Step 1405: Screen multiple candidate image packs according to the attributes of the candidate image packs to obtain multiple target image packs;

[0306] As can be seen from the above step 1405, a total of $k$ candidate image packs are generated. To improve the search speed of the WSI and optimize the final search result, it is also necessary to screen the $k$ candidate image packs. In one embodiment, the $k$ candidate image packs are screened according to the similarity between the candidate image packs and the WSI and / or the diagnostic categories in the candidate image packs to obtain multiple target image packs. The specific screening steps will be introduced in detail below.

[0307] Step 1406: Determine the whole-slide pathology sections to which the multiple target tissue images in the multiple target image packs belong as the final search result.

[0308] After screening out the multiple target image packs, the whole-slide pathology sections to which the multiple target tissue images in the target image packs belong are determined as the final search result. Optionally, the multiple target tissue images in the target image pack may come from the same whole-slide pathology section or multiple different whole-slide pathology sections.

[0309] To sum up, first, the WSI is cropped to obtain multiple small images, and the multiple small images are passed through an image encoder to obtain multiple image feature vectors of the multiple small images; then, the multiple image feature vectors are clustered, and the small images corresponding to the cluster centers are used as key images; then, each key image is queried to obtain candidate image packs; then, the candidate image packs are screened to obtain target image packs; finally, the WSI corresponding to at least one small image in the candidate image pack is used as the final search result; this method provides a way to search for a WSI (large image) with a WSI (large image), and the clustering step and screening step mentioned therein can greatly reduce the amount of data processed and improve the search efficiency. Moreover, the method of searching for a WSI (large image) with a WSI (large image) provided in this embodiment does not require a training process and can achieve fast search matching.

[0310] In the related art, the method of using small images to represent large images often adopts the method of manual selection. Pathologists select core small images according to the color and texture features of each small image in the WSI (for example, histogram statistical information from various color spaces). Then, the features of these core small images are accumulated as the global representation of the WSI. Next, a Support Vector Machine (SVM) is used to classify the WSI global representations of multiple WSIs into two main disease types. In the search stage, once the disease type of the WSI to be searched is determined, image search can be performed in the WSI library with the same disease type.

[0311] Based on Figure 14 In the optional embodiment shown, step 1405 can be replaced by 1405-1.

[0312] 1405-1. Screen multiple candidate image packages according to the number of diagnostic categories of the candidate image packages to obtain multiple target image packages.

[0313] In one embodiment, for the first candidate image package among multiple candidate image packages, based on the cosine similarity between at least one candidate tissue image in the first candidate image package and the key image, the occurrence probability of at least one diagnostic category in the database, and the diagnostic category of at least one candidate tissue image, calculate the entropy value of the candidate image package. The entropy value is used to measure the number of diagnostic categories corresponding to the first candidate image package, and the first candidate image package is any one of the multiple candidate image packages;

[0314] Finally, screen multiple candidate image packages to obtain multiple target image packages with an entropy value lower than the entropy threshold.

[0315] Schematically, the calculation formula of the entropy value is as follows:

[0316]

[0317] Among them, Ent i represents the entropy value of the i-th candidate image package, u i represents the total number of diagnostic categories in the i-th candidate image package, p m represents the probability of the m-th diagnostic type occurring in the i-th candidate image package, and m is a positive integer.

[0318] It is understandable that the entropy value is used to represent the uncertainty of the i-th candidate image packet. The larger the entropy value, the higher the uncertainty of the i-th candidate image packet, the more disordered the distribution of the candidate tissue images in the i-th candidate image packet in the diagnostic category dimension, that is, the higher the uncertainty of the i-th key image, and the less the i-th key image can be used to characterize the WSI. If multiple candidate tissue images in the i-th candidate image packet have the same diagnostic result, the entropy value of the candidate image packet will be 0, and the i-th key image has the best effect in characterizing the WSI.

[0319] In formula (12), p m is calculated as follows:

[0320]

[0321] where y j represents the diagnostic category of the j-th candidate tissue image in the i-th candidate image packet; δ() is a discriminant function used to determine whether the diagnostic category of the j-th candidate tissue image is the same as the m-th diagnostic category. If they are the same, it outputs 1, otherwise it outputs 0; is the weight of the j-th candidate tissue image, is calculated based on the occurrence probability of at least one diagnostic category in the database; d j represents the cosine similarity between the j-th candidate tissue image and the i-th key image in the i-th candidate packet, and (d j +1) / 2 is used to ensure that the value range is between 0 and 1.

[0322] For ease of understanding, formula (13) can regard as a weight score v j to represent the j-th candidate tissue image in the i-th candidate image packet. The denominator of formula (13) represents the total score of the i-th candidate image packet, and the numerator of formula (13) represents the sum of the scores of the m-th diagnostic category in the i-th candidate image packet.

[0323] Through the above formula (12) and formula (13), multiple candidate image packets can be screened, and the candidate image packets with entropy values lower than the preset entropy value threshold are excluded. Multiple candidate image packets can be screened out multiple target image packets, denoted as where k′ is the number of multiple target image packets, and k′ is a positive integer.

[0324] To sum up, excluding the candidate image packets with entropy values lower than the preset entropy value threshold, that is, screening out the candidate image packets with higher stability, further reducing the amount of data processed in the process of searching WSI by WSI, and improving the search efficiency.

[0325] Based on Figure 14In the optional embodiment shown, step 1405 can be replaced by 1405-2.

[0326] 1405-2. Screen multiple candidate image packages according to the similarity between multiple candidate tissue images and the key image to obtain multiple target image packages.

[0327] In one embodiment, for the first candidate image package among multiple candidate image packages, at least one candidate tissue image in the first candidate image package is arranged in descending order of cosine similarity with the key image; the first m candidate tissue images of the first candidate image package are obtained; the m cosine similarities corresponding to the first m candidate tissue images are calculated; wherein, the first candidate image package is any one of the multiple candidate image packages; the average value of the m cosine similarities of the first m candidate tissue images of the multiple candidate image packages is determined as the first average value; the candidate image packages whose average value of the cosine similarities of the included at least one candidate tissue image is greater than the first average value are determined as target image packages to obtain multiple target image packages, and m is a positive integer.

[0328] Schematically, multiple candidate image packages are represented as The candidate tissue images in each candidate image package are arranged in descending order of cosine similarity, and the first average value can be expressed as:

[0329]

[0330] Wherein, i and k respectively represent the i-th candidate image package and the total number of multiple candidate image packages, AveTop represents the average value of the first m cosine similarities in the i-th candidate image package, η is the first average value, and η is used as an evaluation criterion to delete the candidate image packages with an average cosine similarity less than η, and then multiple target image packages can be obtained. The multiple target image packages are represented as: i″ and k″ respectively represent the i-th target image package and the total number of multiple target image packages, and k″ is a positive integer.

[0331] In summary, by eliminating the candidate image packages with a similarity to the key image lower than the first average value, that is, screening out the candidate image packages with a higher similarity between the candidate tissue images and the key image, the amount of data processed in the process of searching WSI by WSI is further reduced, and the search efficiency can be improved.

[0332] It should be noted that the above 1405-1 and 1405-2 can perform the step of screening multiple candidate image packages separately, or can perform the step of screening multiple candidate image packages jointly. At this time, 1405-1 can be executed first and then 1405-2, or 1405-2 can be executed first and then 1405-1. The present application does not limit this.

[0333] Based onFigure 14 In the method embodiment shown, in step 1404, querying candidate image packages through a database is involved. Next, the construction process of the database will be introduced. Please refer to Figure 15 , which shows a schematic diagram of the construction framework of the database provided by an exemplary embodiment of the present application.

[0334] Taking a WSI as an example for introduction:

[0335] First, a plurality of tissue images 1502 are obtained by cropping the WSI 1501; optionally, the cropping method includes: in the preprocessing stage of the WSI, the foreground tissue region in the WSI is extracted through threshold technology, and then the foreground tissue region of the WSI is cropped into a plurality of tissue images based on the sliding window technology.

[0336] Then, the plurality of tissue images 1502 are input into the image encoder 1503 to perform feature extraction on the plurality of tissue images 1502, and a plurality of image feature vectors 1505 of the plurality of tissue images are obtained;

[0337] Finally, based on the plurality of image feature vectors 1505 of the plurality of tissue images, selection of the plurality of tissue images 1502 is performed (i.e., selection of small images 1506). Optionally, the selection of small images 1506 includes two - stage clustering. The first - stage clustering is feature - based clustering 1506 - 1, and the second - stage clustering is coordinate - based clustering 1506 - 2.

[0338] - In the feature - based clustering 1506 - 1, the K - means clustering is used to cluster the plurality of image feature vectors 1505 of the plurality of tissue images into K1 categories, corresponding to obtaining K1 cluster centers, Figure 15 shows a small image corresponding to one of the cluster centers;

[0339] - In the feature - based clustering 1506 - 2, for any one of the K1 categories, the K - means clustering is used to cluster the plurality of feature vectors included in this category into K2 categories, corresponding to obtaining K2 cluster centers, Figure 15 shows a small image corresponding to one of the cluster centers;

[0340] - The small images corresponding to the K1*K2 cluster centers obtained through the two - stage clustering are used as representative small images 1506 - 3, Figure 15 shows a small image corresponding to one of the cluster centers;

[0341] - All the representative small images are used as the small images of the WSI to represent the WSI. Based on this, multiple small images of one WSI are constructed.

[0342] In summary, the process of constructing a database and searching for WSI by WSI is quite similar. Its purpose is to determine multiple small images for representing a WSI to support matching large images by matching small images during the search process.

[0343] In an optional embodiment, the training idea of the above image encoder can also be applied to other image fields. Through sample starfield images (small images), the starfield images are from a starry sky image (large image), and the starfield images indicate local regions in the starry sky image. For example, the starry sky image is an image of the starry sky in a first range, and the starfield image is an image of a sub - range within the first range.

[0344] The training stage of the image encoder includes:

[0345] Obtain a first sample starfield image; perform data augmentation on the first sample starfield image to obtain a first image; input the first image into a first image encoder to obtain a first feature vector; perform data augmentation on the first sample starfield image to obtain a second image; input the second image into a second image encoder to obtain a second feature vector; determine the first feature vector as a contrast vector for contrast learning, and determine the second feature vector as an anchor vector for contrast learning; cluster multiple first feature vectors of different first sample starfield images to obtain multiple first cluster centers; determine the feature vector with the largest similarity value between the multiple first cluster centers and the second feature vector as the positive sample vector among the multiple first feature vectors; determine the remaining first feature vectors as the negative sample vectors among the multiple first feature vectors, where the remaining first feature vectors refer to the feature vectors among the multiple first feature vectors except the feature vector with the largest similarity value between it and the second feature vector; generate a first sub - function based on the second feature vector and the positive sample vector among the multiple first feature vectors; generate a second sub - function based on the second feature vector and the negative sample vector among the multiple first feature vectors; generate a first group loss function based on the first sub - function and the second sub - function; train the first image encoder and the second image encoder based on the first group loss function; determine the second image encoder as the finally trained image encoder.

[0346] Similarly, the image encoder for starfield images can also adopt other training methods similar to the image encoder for the above sample tissue images, which will not be elaborated here.

[0347] The usage stage of the image encoder includes:

[0348] Obtain a starry sky image and crop the starry sky image into multiple star field images; generate multiple image feature vectors of the multiple star field images through an image encoder; determine multiple key images from the multiple star field images by clustering the multiple image feature vectors; query multiple candidate image packs from a database based on the image feature vectors of the multiple key images, where the multiple candidate image packs correspond to the multiple key images one by one, and any one candidate image pack contains at least one candidate star field image; filter the multiple candidate image packs according to the attributes of the candidate image packs to obtain multiple target image packs; determine the starry sky image to which the multiple target star field images in the multiple target image packs belong as the final search result.

[0349] In another alternative embodiment, the training idea of the above image encoder can also be applied to the field of geographical images. The image encoder is trained through sample terrain images (small images), and the terrain images are from a geomorphic image (large image). The terrain image indicates a local area in the geomorphic image. For example, the geomorphic image is an image of the geomorphology in a second range captured by a satellite, and the terrain image is an image of a sub-range within the second range.

[0350] The training stage of the image encoder includes:

[0351] Obtain a first sample terrain image; perform data augmentation on the first sample terrain image to obtain a first image; input the first image into a first image encoder to obtain a first feature vector; perform data augmentation on the first sample terrain image to obtain a second image; input the second image into a second image encoder to obtain a second feature vector; determine the first feature vector as a contrast vector for contrast learning, and determine the second feature vector as an anchor vector for contrast learning; cluster the multiple first feature vectors of different first sample terrain images to obtain multiple first cluster centers; determine the feature vector with the largest similarity value between the multiple first cluster centers and the second feature vector as the positive sample vector among the multiple first feature vectors; determine the remaining first feature vectors as the negative sample vectors among the multiple first feature vectors, where the remaining first feature vectors refer to the feature vectors among the multiple first feature vectors except the feature vector with the largest similarity value between the multiple first feature vectors and the second feature vector; generate a first sub-function based on the second feature vector and the positive sample vector among the multiple first feature vectors; generate a second sub-function based on the second feature vector and the negative sample vector among the multiple first feature vectors; generate a first group loss function based on the first sub-function and the second sub-function; train the first image encoder and the second image encoder based on the first group loss function; determine the second image encoder as the finally trained image encoder.

[0352] Similarly, the image encoder for terrain images can also adopt other training methods similar to the image encoder for the above sample organization images, which will not be elaborated here.

[0353] The usage phase of the image encoder includes:

[0354] Obtain a geomorphic image and crop the geomorphic image into multiple terrain images; generate multiple image feature vectors of the multiple terrain images through the image encoder; determine multiple key images from the multiple terrain images by clustering the multiple image feature vectors; query multiple candidate image packages from the database based on the image feature vectors of the multiple key images, where the multiple candidate image packages correspond to the multiple key images one by one, and any one candidate image package contains at least one candidate terrain image; screen the multiple candidate image packages according to the attributes of the candidate image packages to obtain multiple target image packages; determine the geomorphic image to which the multiple target terrain images in the multiple target image packages belong as the final search result.

[0355] Figure 16 It is a structural block diagram of a training device for an image encoder provided by an exemplary embodiment of the present application. The device includes:

[0356] An acquisition module 1601, configured to acquire a first sample tissue image;

[0357] A processing module 1602, configured to perform data augmentation on the first sample tissue image to obtain a first image; input the first image into a first image encoder to obtain a first feature vector;

[0358] The processing module 1602 is further configured to perform data augmentation on the first sample tissue image to obtain a second image; input the second image into a second image encoder to obtain a second feature vector;

[0359] A determination module 1603, configured to determine the first feature vector as a contrast vector for contrast learning, and determine the second feature vector as an anchor vector for contrast learning;

[0360] A clustering module 1604, configured to cluster multiple first feature vectors of different first sample tissue images to obtain multiple first clustering centers; determine the feature vector with the largest similarity value between the multiple first clustering centers and the second feature vector as the positive sample vector among the multiple first feature vectors; determine the remaining first feature vectors as the negative sample vectors among the multiple first feature vectors, where the remaining first feature vectors refer to the feature vectors among the multiple first feature vectors except the feature vector with the largest similarity value between the multiple first feature vectors and the second feature vector;

[0361] A generation module 1605, configured to generate a first sub-function based on the second feature vector and the positive sample vector among the multiple first feature vectors; generate a second sub-function based on the second feature vector and the negative sample vector among the multiple first feature vectors; generate a first group loss function based on the first sub-function and the second sub-function;

[0362] A training module 1606, configured to train a first image encoder and a second image encoder based on a first group loss function; and determine the second image encoder as the finally trained image encoder.

[0363] In an optional embodiment, the processing module 1602 is further configured to input the first image into the first image encoder to obtain a first intermediate feature vector; and input the first intermediate feature vector into the first MLP to obtain a first feature vector.

[0364] In an optional embodiment, the processing module 1602 is further configured to input the second image into the second image encoder to obtain a second intermediate feature vector; and input the second intermediate feature vector into the second MLP to obtain a second feature vector.

[0365] In an optional embodiment, the determining module 1603 is further configured to determine the second feature vector as a contrast vector for contrast learning, and determine the first feature vector as an anchor vector for contrast learning.

[0366] In an optional embodiment, the clustering module 1604 is further configured to cluster multiple second feature vectors of different first sample organization images to obtain multiple second cluster centers; determine the feature vector with the largest similarity value between the multiple second cluster centers and the first feature vector as the positive sample vector among the multiple second feature vectors; and determine the remaining second feature vectors as the negative sample vectors among the multiple second feature vectors, where the remaining second feature vectors refer to the feature vectors among the multiple second feature vectors except the feature vector with the largest similarity value between the multiple second cluster centers and the first feature vector.

[0367] In an optional embodiment, the generating module 1605 is further configured to generate a third sub-function based on the first feature vector and the positive sample vector among the multiple second feature vectors; generate a fourth sub-function based on the first feature vector and the negative sample vectors among the multiple second feature vectors; and generate a second group loss function based on the third sub-function and the fourth sub-function.

[0368] In an optional embodiment, the training module 1606 is further configured to train the first image encoder and the second image encoder based on the second group loss function; and determine the first image encoder as the finally trained image encoder.

[0369] In an optional embodiment, the training module 1606 is further configured to update the parameters of the third image encoder in a weighted manner according to the parameters shared between the first image encoder and the second image encoder.

[0370] In summary, by further distinguishing the positive samples identified in the relevant technology and further distinguishing the "positivity" of the positive samples among the positive samples, the loss function used in contrastive learning (also called contrastive learning paradigm) can more accurately bring the anchor image and the positive sample closer, thereby better training the image encoder. The trained image encoder can better learn the common features between the anchor image and the positive sample, and the features of the small image obtained by feature extraction of the image encoder can better represent the small image.

[0371] Figure 17 : is a structural block diagram of a training device for an image encoder provided by an exemplary embodiment of the present application, the device comprising:

[0372] An acquisition module 1701 is used to acquire a first sample tissue image and a plurality of second sample tissue images, where the second sample tissue images are negative samples in contrast learning;

[0373] The processing module 1702 is used to perform data enhancement on the first sample tissue image to obtain a third image; input the third image into a third image encoder to obtain a third feature vector; the third image is a positive sample in contrastive learning;

[0374] The processing module 1702 is further used to perform data enhancement on the first sample tissue image to obtain a first image; input the first image into a first image encoder to obtain a fourth feature vector; the first image is an anchor image in contrastive learning;

[0375] The processing module 1702 is further configured to input the plurality of second sample tissue images into the third image encoder to obtain a plurality of feature vectors of the plurality of second sample tissue images; cluster the plurality of feature vectors to obtain a plurality of cluster centers; and generate a plurality of weights based on similarity values ​​between the plurality of cluster centers and the third feature vector;

[0376] A generating module 1703 is used to generate a fifth sub-function for characterizing the error between the anchor image and the positive sample based on the fourth eigenvector and the third eigenvector; generate a sixth sub-function for characterizing the error between the anchor image and the negative sample based on the fourth eigenvector and the plurality of eigenvectors in combination with a plurality of weights; and generate a first weight loss function based on the fifth sub-function and the sixth sub-function;

[0377] The training module 1704 is used to train the first image encoder and the third image encoder based on the first weight loss function; and update the third image encoder based on the first image encoder.

[0378] In an optional embodiment, the processing module 1702 is further used to cluster multiple feature vectors of multiple second sample tissue images to obtain multiple cluster centers of the first training batch when the first sample tissue image belongs to the first sample tissue image in the first training batch.

[0379] In an alternative embodiment, the processing module 1702 is further configured to, when the first sample tissue image belongs to the first sample tissue image of the nth training batch, update the multiple cluster centers corresponding to the (n-1)th training batch to the multiple cluster centers corresponding to the nth training batch, where n is a positive integer greater than 1.

[0380] In an alternative embodiment, the processing module 1702 is further configured to, for the jth cluster center among the multiple cluster centers of the (n-1)th training batch, update the jth cluster center of the (n-1)th training batch based on the first sample tissue images belonging to the jth category in the nth training batch, to obtain the jth cluster center of the nth training batch, where i is a positive integer.

[0381] In an alternative embodiment, the magnitude of the weight is negatively correlated with the similarity value between the cluster center and the third feature vector; for the jth cluster center among the multiple cluster centers, the feature vectors included in the category to which the jth cluster center belongs correspond to the same weight.

[0382] In an alternative embodiment, the processing module 1702 is further configured to input the first image into the first image encoder to obtain a first intermediate feature vector; and input the first intermediate feature vector into the third multi-layer perceptron (MLP) to obtain a fourth feature vector.

[0383] In an alternative embodiment, the training module 1704 is further configured to update the parameters of the third image encoder in a weighted manner according to the parameters of the first image encoder.

[0384] In an alternative embodiment, the processing module 1702 is further configured to perform data augmentation on the first sample tissue image to obtain a second image; input the second image into the second image encoder to obtain a fifth feature vector; and the second image is the anchor image in contrastive learning.

[0385] In an alternative embodiment, the generation module 1703 is further configured to generate a seventh sub-function for characterizing the error between the anchor image and the positive sample based on the fifth feature vector and the third feature vector; generate an eighth sub-function for characterizing the error between the anchor image and the negative sample based on the fifth feature vector, the multiple feature vectors, and the multiple weights; and generate a second weight loss function based on the seventh sub-function and the eighth sub-function.

[0386] In an alternative embodiment, the training module 1704 is further configured to train the second image encoder and the third image encoder based on the second weight loss function.

[0387] In an alternative embodiment, the processing module 1702 is further configured to input the second image into the second image encoder to obtain a second intermediate feature vector; and input the second intermediate feature vector into the fourth MLP to obtain a fifth feature vector.

[0388] In an alternative embodiment, the training module 1704 is further configured to update the parameters of the third image encoder in a weighted manner according to the parameters shared between the first image encoder and the second image encoder.

[0389] In summary, by assigning weights to the negative samples identified in the related art and further distinguishing the "negative degree" of the negative samples among the negative samples, the loss function used in contrast learning (also known as the contrast learning paradigm) can more accurately separate the anchor image from the negative samples, reducing the influence of potential false negative samples. Furthermore, it can better train the image encoder. The trained image encoder can better distinguish the different features between the anchor image and the negative samples, and the features of the small images obtained by feature extraction using the image encoder can better represent the small images.

[0390] Figure 18 FIG. 10 is a structural block diagram of a search device for a whole-slide pathological section provided by an exemplary embodiment of the present application. The device includes:

[0391] An acquisition module 1801, configured to acquire a whole-slide pathological section and crop the whole-slide pathological section into multiple tissue images;

[0392] A generation module 1802, configured to generate multiple image feature vectors of the multiple tissue images through an image encoder;

[0393] A clustering module 1803, configured to determine multiple key images from the multiple tissue images by clustering the multiple image feature vectors;

[0394] A query module 1804, configured to query multiple candidate image packages from a database based on the image feature vectors of the multiple key images. The multiple candidate image packages correspond to the multiple key images one by one, and any one of the candidate image packages contains at least one candidate tissue image;

[0395] A screening module 1805, configured to screen the multiple candidate image packages according to the attributes of the candidate image packages to obtain multiple target image packages;

[0396] A determination module 1806, configured to determine the whole-slide pathological section to which the multiple target tissue images in the multiple target image packages belong as the final search result.

[0397] In an alternative embodiment, the clustering module 1803 is further configured to cluster multiple image feature vectors of multiple tissue images to obtain multiple first-class clusters; and determine the multiple clustering centers of the multiple first-class clusters as the multiple image feature vectors of multiple key images, respectively.

[0398] In an alternative embodiment, the clustering module 1803 is further configured to, for a target first-class cluster among the multiple first-class clusters, cluster to obtain multiple second-class clusters based on the position features of the multiple tissue images corresponding to the target first-class cluster in their respective whole-slide pathology sections; and for the target first-class cluster among the multiple first-class clusters, determine the multiple clustering centers corresponding to the multiple second-class clusters included in the target first-class cluster as the image feature vectors of the key image; wherein the target first-class cluster is any one of the multiple first-class clusters.

[0399] In an alternative embodiment, the screening module 1805 is further configured to screen multiple candidate image packages according to the number of diagnostic categories of the candidate image packages to obtain multiple target image packages.

[0400] In an alternative embodiment, the screening module 1805 is further configured to, for a first candidate image package among the multiple candidate image packages, calculate the entropy value of the candidate image package based on the cosine similarity between at least one candidate tissue image in the first candidate image package and the key image, the occurrence probability of at least one diagnostic category in the database, and the diagnostic category of at least one candidate tissue image; wherein the entropy value is used to measure the number of diagnostic categories corresponding to the first candidate image package, and the first candidate image package is any one of the multiple candidate image packages; and screen multiple candidate image packages to obtain multiple target image packages with entropy values lower than the entropy threshold.

[0401] In an alternative embodiment, the screening module 1805 is further configured to screen multiple candidate image packages according to the similarity between multiple candidate tissue images and the key image to obtain multiple target image packages.

[0402] In an alternative embodiment, the screening module 1805 is further configured to, for a first candidate image package among the multiple candidate image packages, arrange at least one candidate tissue image in the first candidate image package in descending order of cosine similarity with the key image; obtain the first m candidate tissue images of the first candidate image package; calculate the m cosine similarities corresponding to the first m candidate tissue images; determine the average value of the m cosine similarities of the first m candidate tissue images of the multiple candidate image packages as the first average value; and determine the candidate image packages whose average value of the cosine similarities of the included at least one candidate tissue image is greater than the first average value as the target image packages to obtain multiple target image packages; wherein the first candidate image package is any one of the multiple candidate image packages.

[0403] In summary, first, the WSI is cropped to obtain multiple small images, and the multiple small images are passed through an image encoder to obtain multiple image feature vectors of the multiple small images; then, the multiple image feature vectors are clustered, and the small images corresponding to the cluster centers are used as key images; then, each key image is queried to obtain a candidate image package; then, the candidate image package is filtered to obtain a target image package; finally, the WSI corresponding to at least one small image in the candidate image package is used as the final search result; the device supports the method of searching for a WSI (large image) by a WSI (large image), and moreover, the clustering module and the filtering module mentioned therein can greatly reduce the amount of data processed and improve the search efficiency. Moreover, the device for searching for a WSI (large image) by a WSI (large image) provided in this embodiment does not require a training process and can achieve fast search and matching.

[0404] Figure 19 is a schematic structural diagram of a computer device shown according to an exemplary embodiment. The computer device 1900 may be Figure 2 the training device 21 of the image encoder in Figure 2 or the using device 22 of the image encoder in

[0405] The computer device 1900 includes a central processing unit (CPU) 1901, a system memory 1904 including a random access memory (RAM) 1902 and a read-only memory (ROM) 1903, and a system bus 1905 connecting the system memory 1904 and the central processing unit 1901. The computer device 1900 also includes a basic input / output system (Input / Output, I / O system) 1906 for facilitating the transfer of information between various components within the computer device, and a mass storage device 1907 for storing an operating system 1913, application programs 1914, and other program modules 1915.

[0406] The large-capacity storage device 1907 is connected to the central processing unit 1901 through a large-capacity storage controller (not shown) connected to the system bus 1905. The large-capacity storage device 1907 and its associated computer-readable medium provide non-volatile storage for the computer device 1900. That is, the large-capacity storage device 1907 may include computer-readable media (not shown) such as a hard disk or a compact disc read-only memory (CD-ROM) drive.

[0407] Without loss of generality, the computer-readable medium may include computer storage media and communication media. Computer storage media includes volatile and non-volatile, removable and non-removable media implemented by any method or technology for storing information such as computer-readable instructions, data structures, program modules, or other data. Computer storage media includes RAM, ROM, erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), CD-ROM, digital video disc (DVD), or other optical storage, magnetic tape cartridges, tapes, disk storage, or other magnetic storage devices. Of course, those skilled in the art will appreciate that the computer storage media is not limited to the above several types. The above system memory 1904 and large-capacity storage device 1907 may be collectively referred to as memory.

[0408] According to various embodiments of the present disclosure, the computer device 1900 may also operate by connecting to a remote computer device on a network such as the Internet. That is, the computer device 1900 may be connected to the network 1911 through a network interface unit 1912 connected to the system bus 1905, or, in other words, the network interface unit 1912 may also be used to connect to other types of networks or remote computer device systems (not shown).

[0409] The memory further includes one or more programs, and the one or more programs are stored in the memory. The central processing unit 1901 implements all or part of the steps of the above image encoder training method by executing the one or more programs.

[0410] The present application also provides a computer-readable storage medium, in which at least one instruction, at least one program, a code set or an instruction set is stored, and the at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by a processor to implement the training method of the image encoder provided in the above method embodiment.

[0411] The present application provides a computer program product or a computer program, the computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the training method of the image encoder provided in the above method embodiment.

[0412] The serial numbers of the embodiments of the present application above are only for description and do not represent the advantages and disadvantages of the embodiments.

[0413] Those of ordinary skill in the art can understand that all or part of the steps for implementing the above embodiments can be completed by hardware, or can be completed by a program instructing relevant hardware. The program can be stored in a computer-readable storage medium, and the storage medium mentioned above can be a read-only memory, a magnetic disk or an optical disc, etc.

[0414] The above are only optional embodiments of the present application and are not intended to limit the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.

Claims

1. A training method for an image encoder, characterized in that, The method includes: Obtain a first sample tissue image; Perform data augmentation on the first sample tissue image to obtain a first image; input the first image into a first image encoder to obtain a first feature vector; Perform data augmentation on the first sample tissue image to obtain a second image; input the second image into a second image encoder to obtain a second feature vector; Determine the first feature vector as a contrast vector for contrast learning, and determine the second feature vector as an anchor vector for contrast learning; Cluster multiple first feature vectors of different first sample tissue images to obtain multiple first cluster centers; determine the feature vector with the largest similarity value between the multiple first cluster centers and the second feature vector as the positive sample vector among the multiple first feature vectors; determine the first remaining feature vectors as the negative sample vectors among the multiple first feature vectors, where the first remaining feature vectors refer to the feature vectors among the multiple first feature vectors other than the feature vector with the largest similarity value between the multiple first cluster centers and the second feature vector; Generate a first sub-function based on the second feature vector and the positive sample vector among the multiple first feature vectors; generate a second sub-function based on the second feature vector and the negative sample vectors among the multiple first feature vectors; generate a first group loss function based on the first sub-function and the second sub-function; Train the first image encoder and the second image encoder based on the first group loss function; determine the second image encoder as the finally trained image encoder.

2. The method according to claim 1, wherein The step of inputting the first image into the first image encoder to obtain a first feature vector includes: Input the first image into the first image encoder to obtain a first intermediate feature vector; input the first intermediate feature vector into a first MLP to obtain the first feature vector; The step of inputting the second image into the second image encoder to obtain a second feature vector includes: Input the second image into the second image encoder to obtain a second intermediate feature vector; input the second intermediate feature vector into a second MLP to obtain the second feature vector.

3. The method according to claim 1, wherein The method further includes: Determine the second feature vector as a contrast vector for contrast learning, and determine the first feature vector as an anchor vector for contrast learning; Cluster multiple second feature vectors of different first sample tissue images to obtain multiple second cluster centers; determine the feature vector with the largest similarity value between the multiple second cluster centers and the first feature vector as the positive sample vector among the multiple second feature vectors; determine the second remaining feature vectors as the negative sample vectors among the multiple second feature vectors, where the second remaining feature vectors refer to the feature vectors among the multiple second feature vectors other than the feature vector with the largest similarity value between the multiple second cluster centers and the first feature vector; Generate a third sub-function based on the first feature vector and the positive sample vectors among the multiple second feature vectors; generate a fourth sub-function based on the first feature vector and the negative sample vectors among the multiple second feature vectors; generate a second group loss function based on the third sub-function and the fourth sub-function; Train the first image encoder and the second image encoder based on the second group loss function; determine the first image encoder as the finally trained image encoder.

4. The method according to claim 1, wherein The method further includes: Obtain multiple second sample tissue images, which are negative samples in contrast learning; Perform data augmentation on the first sample tissue image to obtain a third image; input the third image into a third image encoder to obtain a third feature vector, where the third image is a positive sample in contrast learning; Input the first image into the first image encoder to obtain a fourth feature vector, where the first image is an anchor image in contrast learning; Input the multiple second sample tissue images into the third image encoder to obtain multiple feature vectors of the multiple second sample tissue images; cluster the multiple feature vectors to obtain multiple cluster centers; generate multiple weights based on the similarity values between the multiple cluster centers and the third feature vector; Generate a fifth sub-function for characterizing the error between the anchor image and the positive sample based on the third feature vector and the fourth feature vector; generate a sixth sub-function for characterizing the error between the anchor image and the negative sample by combining the multiple weights based on the fourth feature vector and the multiple feature vectors; generate a first weight loss function based on the fifth sub-function and the sixth sub-function; Train the first image encoder and the third image encoder based on the first weight loss function; update the third image encoder based on the first image encoder; determine the third image encoder as the finally trained image encoder.

5. A search method for a whole-field pathological section, characterized in that, The method is executed by a computer device, and the computer device runs an image encoder trained by any of the methods in claims 1 to 4. The method includes: Obtain a whole-slide pathology section and crop the whole-slide pathology section into multiple tissue images; Generate multiple image feature vectors of the multiple tissue images through the image encoder; Determine multiple key images from the multiple tissue images by clustering the multiple image feature vectors; Query multiple candidate image packs from a database based on the image feature vectors of the multiple key images. The multiple candidate image packs correspond to the multiple key images one by one, and any one of the candidate image packs contains at least one candidate tissue image; Filter the multiple candidate image packs according to the attributes of the candidate image packs to obtain multiple target image packs; Determine the whole-slide pathology sections to which the multiple target tissue images in the multiple target image packs belong as the final search result.

6. The method according to claim 5, wherein The determining multiple key images from the multiple tissue images by clustering the multiple image feature vectors includes: Cluster the multiple image feature vectors of the multiple tissue images to obtain multiple first - type clusters; Respectively determine the multiple clustering centers of the multiple first - type clusters as the multiple image feature vectors of the multiple key images.

7. The method according to claim 6, wherein The method further includes: For a target first - type cluster among the multiple first - type clusters, cluster based on the position features of the multiple tissue images corresponding to the target first - type cluster in their respective whole - slide pathology sections to obtain multiple second - type clusters; The step of respectively determining the multiple clustering centers of the multiple first - type clusters as the multiple image feature vectors of the multiple key images includes: For a target first - type cluster among the multiple first - type clusters, determine the multiple clustering centers corresponding to the multiple second - type clusters included in the target first - type cluster as the image feature vectors of the key images; Wherein, the target first - type cluster is any one of the multiple first - type clusters.

8. The method according to any one of claims 5 to 7, characterized in that The step of screening the multiple candidate image packs according to the attributes of the candidate image packs to obtain multiple target image packs includes: Screen the multiple candidate image packs according to the number of diagnostic categories of the candidate image packs to obtain the multiple target image packs.

9. The method according to claim 8, wherein The step of screening the multiple candidate image packs according to the number of diagnostic categories of the candidate image packs to obtain multiple target image packs includes: For a first candidate image pack among the multiple candidate image packs, calculate the entropy value of the candidate image pack based on the cosine similarity between at least one candidate tissue image in the first candidate image pack and the key image, the occurrence probability of at least one diagnostic category in the database, and the diagnostic category of the at least one candidate tissue image; wherein, the entropy value is used to measure the number of diagnostic categories corresponding to the first candidate image pack, and the first candidate image pack is any one of the multiple candidate image packs; Screen the multiple candidate image packs to obtain the multiple target image packs with entropy values lower than the entropy threshold.

10. The method according to any one of claims 5 to 7, characterized in that The step of screening the multiple candidate image packs according to the attributes of the candidate image packs to obtain multiple target image packs includes: Screen the multiple candidate image packs according to the similarity between the multiple candidate tissue images and the key image to obtain the multiple target image packs.

11. The method according to claim 10, wherein The step of screening the multiple candidate image packs according to the similarity between the multiple candidate tissue images and the key image to obtain the multiple target image packs includes: For a first candidate image pack among the multiple candidate image packs, arrange at least one candidate tissue image in the first candidate image pack in descending order of cosine similarity with the key image; obtain the first m candidate tissue images of the first candidate image pack; calculate the m cosine similarities corresponding to the first m candidate tissue images; Determine the average value of the m cosine similarities of the first m candidate tissue images of the multiple candidate image packs as the first average value; Determine the candidate image packs whose average value of the cosine similarities of the included at least one candidate tissue image is greater than the first average value as the target image packs to obtain the multiple target image packs; Wherein, the first candidate image packet is any one of the multiple candidate image packets, and m is a positive integer.

12. A training device for an image encoder, characterized in that, The apparatus includes: An acquisition module, configured to acquire a first sample tissue image; A processing module, configured to perform data augmentation on the first sample tissue image to obtain a first image; input the first image into a first image encoder to obtain a first feature vector; The processing module is further configured to perform data augmentation on the first sample tissue image to obtain a second image; input the second image into a second image encoder to obtain a second feature vector; A determination module, configured to determine the first feature vector as a contrast vector for contrast learning, and determine the second feature vector as an anchor vector for contrast learning; A clustering module, configured to cluster the first feature vectors of different first sample tissue images to obtain multiple first clustering centers; determine the feature vector with the largest similarity value between the multiple first clustering centers and the second feature vector as the positive sample vector among the multiple first feature vectors; determine the remaining feature vectors as the negative sample vectors among the multiple first feature vectors, where the remaining feature vectors refer to the first feature vectors of different first sample tissue images except the feature vector with the largest similarity value between the multiple first clustering centers and the second feature vector; A generation module, configured to generate a first sub-function based on the second feature vector and the positive sample vector among the multiple first feature vectors; generate a second sub-function based on the second feature vector and the negative sample vector among the multiple first feature vectors; generate a first group loss function based on the first sub-function and the second sub-function; A training module, configured to train the first image encoder and the second image encoder based on the first group loss function; determine the second image encoder as the finally trained image encoder.

13. A search device for a whole-field pathological section, characterized in that, The apparatus runs an image encoder trained by any one of the methods recited in claims 1 to 4. The apparatus includes: An acquisition module, configured to acquire a whole-slide pathology section and crop the whole-slide pathology section into multiple tissue images; A generation module, configured to generate multiple image feature vectors of the multiple tissue images through the image encoder; A clustering module, configured to determine multiple key images from the multiple tissue images by clustering the multiple image feature vectors; A query module, configured to query multiple candidate image packets from a database based on the image feature vectors of the multiple key images, where the multiple candidate image packets correspond to the multiple key images one by one, and any one of the candidate image packets contains at least one candidate tissue image; A screening module, configured to screen the multiple candidate image packets according to the attributes of the candidate image packets to obtain multiple target image packets; A determination module, configured to determine the whole-slide pathology section to which the multiple target tissue images in the multiple target image packets belong as the final search result.

14. A computer device, characterized in that, The computer device includes: a processor and a memory, where the memory stores a computer program, and the computer program is loaded and executed by the processor to implement the training method of the image encoder according to any one of claims 1 to 4, or, the search method of the whole-field pathological section according to any one of claims 5 to 11.

15. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, and the computer program is loaded and executed by a processor to implement the training method of the image encoder according to any one of claims 1 to 4, or, the search method of the whole-field pathological section according to any one of claims 5 to 11.

16. A computer program product, characterized in that, The computer program product stores a computer program, and the computer program is loaded and executed by a processor to implement the training method of the image encoder according to any one of claims 1 to 4, or, the search method of the whole-field pathological section according to any one of claims 5 to 11.

Citation Information

Patent Citations

  • Analysis of histopathology samples

    GB202106397D0

  • Multiple instance learning method

    US20210334994A1