Full-slice image processing method and device, electronic equipment and storage medium
By introducing position enhancement, two-dimensional rotation embedding and attention modules in the full-slice image processing, the problem of inaccurate feature information in the prior art is solved, and the accuracy of image feature extraction and the effect of biometric prediction are improved.
Patent Information
- Application Number
- CN202510308246.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-14
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2045-03-14
AI Technical Summary
When the existing pathological feature aggregation method processes full-section images, it is difficult to accurately model the morphological variability of tissue regions, resulting in inaccurate feature information and ineffective processing of clinical tasks.
By introducing position enhancement module, two-dimensional rotation embedding module and attention module into the feature extraction model, the degree of attention and stability of instance positions during the instance aggregation process is enhanced, and the accuracy of image feature information is improved.
It improves the accuracy of feature extraction of full-slice image, enhances the ability to generalize spatial position information, stabilizes attention entropy, and improves the accuracy of biometric information prediction.
Smart Images

Figure CN120147804A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of computer technology, and in particular, to a method, apparatus, electronic device, and storage medium for processing whole slide images. Background Art
[0002] Computer-aided analysis of whole slide images (WSIs) has been widely used in clinical tasks, such as cancer subtype classification, prognosis analysis, and biomarker prediction. However, the gigapixel resolution of WSIs poses a significant challenge to achieving effective representation through end-to-end training. Currently, the standard process for embedding WSIs mainly includes two stages: instance-level embedding and aggregating the embedding results to adapt to clinical tasks. Although domain-specific visual encoders have made progress driven by large-scale datasets and state-of-the-art self-supervised learning, existing pathological feature aggregation methods mainly rely on attention mechanisms and diverse spatial location embedding techniques to model spatial relationships and capture dependencies between instances. However, due to differences in clinical procedures and tumor characteristics, the morphological variability of tissue regions (foregrounds in WSIs) leads to challenges in the aggregation strategy, such as a decline in the ability to model the dependency relationships of key instances and uneven spatial location distributions, resulting in inaccurate feature information of the whole slide image and an inability to accurately handle clinical tasks. Summary of the Invention
[0003] In view of this, the present disclosure provides a method, apparatus, electronic device, and storage medium for processing whole slide images, aiming to improve the accuracy of the image feature information obtained by extracting features from whole slide images by enhancing the attention degree of instance positions in the instance aggregation process and the stability of the instance aggregation process.
[0004] According to a first aspect of the present disclosure, there is provided a method for processing whole slide images, the method comprising:
[0005] Determine a plurality of instance region images included in the whole slide image, and coordinate information corresponding to each instance region image, the coordinate information being used to characterize the position of the corresponding instance region image in the whole slide image;
[0006] Input the plurality of instance region images and the corresponding coordinate information into a trained feature extraction model, and output corresponding image feature information;
[0007] Wherein, the feature extraction model includes a position enhancement module, a two-dimensional rotation embedding module, and an attention module;
[0008] The inputting the plurality of instance region images and the corresponding coordinate information into a trained feature extraction model and outputting the image feature information corresponding to the whole slide image includes:
[0009] Input the multiple pieces of coordinate information into the position enhancement module for position enhancement to obtain the enhanced positions corresponding to each piece of coordinate information;
[0010] Perform instance fusion with redundancy removal based on the multiple instance region images and the corresponding enhanced positions to obtain multiple instance encoding information and multiple enhanced positions after redundancy removal;
[0011] Input the multiple pieces of instance encoding information and the multiple enhanced positions after passing through the two-dimensional rotation embedding module into the attention module for information interaction and instance aggregation to obtain the image feature information corresponding to the full-slice image.
[0012] In a possible implementation manner, the feature extraction model further includes an instance fusion module. The performing instance fusion with redundancy removal based on the multiple instance region images and the corresponding enhanced positions to obtain multiple instance encoding information and multiple enhanced positions after redundancy removal includes:
[0013] Perform image encoding on each instance region image to obtain the corresponding instance encoding information;
[0014] Input the multiple pieces of instance encoding information and the corresponding enhanced positions into the instance fusion module for instance fusion with redundancy removal, and output the remaining multiple pieces of instance encoding information and the corresponding enhanced positions after redundancy removal. The number of output instance encoding information and enhanced positions is less than the number of input instance encoding information and enhanced positions.
[0015] In a possible implementation manner, the inputting the multiple pieces of instance encoding information and the multiple enhanced positions after passing through the two-dimensional rotation embedding module into the attention module for information interaction and instance aggregation to obtain the image feature information corresponding to the full-slice image includes:
[0016] Input the multiple pieces of instance encoding information and the multiple enhanced positions after passing through the two-dimensional rotation embedding module into the attention module for semantic information extraction to obtain the semantic information corresponding to each piece of instance encoding information. The semantic information includes key information and value information;
[0017] Perform instance aggregation on the multiple pieces of instance encoding information based on the corresponding semantic information to obtain an information aggregation result;
[0018] For each piece of instance encoding information, search for the corresponding neighboring instance encoding information in the information aggregation result, and perform self-attention operation based on the information aggregation result to obtain the image feature information corresponding to the full-slice image.
[0019] In a possible implementation, the step of inputting the multiple pieces of coordinate information into the position enhancement module for position enhancement to obtain the enhanced position corresponding to each piece of coordinate information includes:
[0020] Inputting the multiple pieces of coordinate information into the position enhancement module for random rotation and random projection to obtain the enhanced position corresponding to each piece of coordinate information.
[0021] In a possible implementation, the step of performing instance aggregation on the multiple instance encoding information based on the corresponding semantic information to obtain an information aggregation result includes:
[0022] Performing information aggregation through the formula to obtain the information aggregation result, where Z g is the information aggregation result, Q s is the matrix composed of the key information corresponding to each piece of instance encoding information, is the transpose of the matrix composed of the value information corresponding to each piece of instance encoding information, d is the dimension of the instance encoding information, and Z′ is the matrix composed of the multiple pieces of instance encoding information.
[0023] In a possible implementation, for each piece of instance encoding information, searching for the corresponding neighborhood instance encoding information in the information aggregation result and performing self-attention operation based on the information aggregation result to obtain the image feature information corresponding to the whole-slide image includes:
[0024] Performing self-attention operation based on the information aggregation result through the formula to obtain the image feature information corresponding to the whole-slide image, where z is an instance, z out is the image feature information, Z nei is the matrix composed of the neighborhood instance encoding information corresponding to each piece of instance encoding information, Z g is the information aggregation result, 2D-RoPE is the two-dimensional rotation position embedding function, p z is the enhanced position of the instance encoding information z, P is the enhanced position corresponding to Z nei and Z g , and W Q and W K are weight vectors.
[0025] In a possible implementation, the whole-slide image is obtained by image acquisition of a target object, and the method further includes:
[0026] Performing feature prediction based on the image feature information corresponding to the whole-slide image to obtain the biometric information of the target object.
[0027] According to a second aspect of the present disclosure, there is provided a whole-slide image processing apparatus, the apparatus comprising:
[0028] An information determination module, configured to determine a plurality of instance region images included in the whole-slide image, and coordinate information corresponding to each of the instance region images, the coordinate information being used to characterize the position of the corresponding instance region image in the whole-slide image;
[0029] A feature extraction module, configured to input the plurality of instance region images and the corresponding coordinate information into a trained feature extraction model, and output corresponding image feature information;
[0030] Wherein, the feature extraction model includes a position enhancement module, a two-dimensional rotation embedding module, and an attention module; the feature extraction module is further configured to:
[0031] Input the plurality of coordinate information into the position enhancement module for position enhancement to obtain an enhanced position corresponding to each of the coordinate information;
[0032] Perform instance fusion with redundancy removal based on the plurality of instance region images and the corresponding enhanced positions to obtain a plurality of instance coding information after redundancy removal and a plurality of enhanced positions;
[0033] Input the plurality of instance coding information and the plurality of enhanced positions after passing through the two-dimensional rotation embedding module into the attention module for information interaction and instance aggregation to obtain the image feature information corresponding to the whole-slide image.
[0034] In a possible implementation manner, the feature extraction model further includes an instance fusion module, and the feature extraction module is configured to:
[0035] Perform image coding on each of the instance region images to obtain corresponding instance coding information;
[0036] Input the plurality of instance coding information and the corresponding enhanced positions into the instance fusion module for instance fusion with redundancy removal, and output the remaining plurality of instance coding information and the corresponding enhanced positions after redundancy removal, and the number of output instance coding information and enhanced positions is less than the number of input instance coding information and enhanced positions.
[0037] In a possible implementation manner, the feature extraction module is further configured to:
[0038] Input the plurality of instance coding information and the plurality of enhanced positions after passing through the two-dimensional rotation embedding module into the attention module for semantic information extraction to obtain semantic information corresponding to each of the instance coding information, and the semantic information includes key information and value information;
[0039] Performing instance aggregation on the multiple instance encoding information based on the corresponding semantic information to obtain an information aggregation result;
[0040] For each of the instance encoding information, searching for corresponding neighborhood instance encoding information in the information aggregation result, and performing self-attention operation according to the information aggregation result to obtain the image feature information corresponding to the whole-slide image.
[0041] In a possible implementation manner, the feature extraction module is further configured to:
[0042] Inputting the multiple coordinate information into the position enhancement module for random projection and random rotation to obtain the enhanced position corresponding to each coordinate information.
[0043] In a possible implementation manner, the feature extraction module is further configured to:
[0044] Through the formula Performing information aggregation to obtain an information aggregation result, where Z g is the information aggregation result, Q s is the matrix composed of the key information corresponding to each instance encoding information, is the transpose of the matrix composed of the value information corresponding to each instance encoding information, d is the dimension of the instance encoding information, and Z′ is the matrix composed of multiple instance encoding information.
[0045] In a possible implementation manner, the feature extraction module is further configured to:
[0046] Through the formula Performing self-attention operation based on the information aggregation result to obtain the image feature information corresponding to the whole-slide image, where z is an instance, z out is the image feature information, Z nei is the matrix composed of the neighborhood instance encoding information corresponding to each instance encoding information, Z g is the information aggregation result, 2D-RoPE is the two-dimensional rotation position embedding function, p z is the enhanced position of the instance encoding information z, P is the enhanced position corresponding to Z nei and Z g corresponding, W Q and W K are weight vectors.
[0047] In a possible implementation manner, the whole-slide image is obtained by image acquisition of a target object, and the device further includes:
[0048] A feature prediction module, configured to perform feature prediction based on the image feature information corresponding to the whole slide image, so as to obtain the biometric information of the target object.
[0049] According to a third aspect of the present disclosure, there is provided an electronic device, including: a processor; a memory for storing processor-executable instructions; wherein, the processor is configured to implement the above method when executing the instructions stored in the memory.
[0050] According to a fourth aspect of the present disclosure, there is provided a non-volatile computer-readable storage medium, on which computer program instructions are stored, wherein the computer program instructions implement the above method when executed by a processor.
[0051] According to a fifth aspect of the present disclosure, there is provided a computer program product, including computer-readable code, or a non-volatile computer-readable storage medium carrying the computer-readable code, when the computer-readable code runs in a processor of an electronic device, the processor in the electronic device executes the above method.
[0052] In the embodiments of the present disclosure, the method inputs multiple coordinate information into a position enhancement module for position enhancement to obtain the enhanced positions of each coordinate information. Instance fusion for redundancy removal is performed based on multiple instance region images and the corresponding enhanced positions, obtaining multiple instance coding information after redundancy removal and multiple enhanced positions, and inputting the multiple instance coding information and the multiple enhanced positions after passing through a two-dimensional rotation embedding module into an attention module for information interaction and instance aggregation, obtaining the image feature information of the whole slide image. The present disclosure enhances the attention degree of the algorithm to the positions of each instance in the whole slide image during the instance aggregation process by setting a position enhancement module and a two-dimensional rotation embedding module, and improves the stability of the instance aggregation process through the setting of the attention module, finally obtaining more accurate image feature information, and further being able to improve the accuracy of biometric information prediction based on the image feature information.
[0053] According to the following detailed description of exemplary embodiments with reference to the accompanying drawings, other features and aspects of the present disclosure will become clear. Description of the Drawings
[0054] The drawings included in the specification and constituting a part of the specification, together with the specification, illustrate the exemplary embodiments, features, and aspects of the present disclosure, and are used to explain the principles of the present disclosure.
[0055] Figure 1 A flowchart showing a method for processing a whole slide image according to an embodiment of the present disclosure;
[0056] Figure 2 A schematic diagram showing a process of processing a whole slide image according to an embodiment of the present disclosure;
[0057] Figure 3 Shows a schematic structural diagram of an attention module according to an embodiment of the present disclosure;
[0058] Figure 4 Shows a schematic structural diagram of a position enhancement module according to an embodiment of the present disclosure;
[0059] Figure 5 Shows a schematic diagram of an application process according to an embodiment of the present disclosure;
[0060] Figure 6 Shows a schematic diagram of a whole-slide image processing device according to an embodiment of the present disclosure;
[0061] Figure 7 Shows a schematic diagram of an electronic device according to an embodiment of the present disclosure. Detailed implementation manners
[0062] The following will detail various exemplary embodiments, features, and aspects of the present disclosure with reference to the accompanying drawings. The same reference numerals in the drawings denote elements having the same or similar functions. Although various aspects of the embodiments are shown in the drawings, the drawings are not necessarily drawn to scale unless otherwise specified.
[0063] The specific term "exemplary" herein means "serving as an example, embodiment, or illustration". Any embodiment described herein as "exemplary" is not necessarily to be construed as superior or better than other embodiments.
[0064] In addition, for a better description of the present disclosure, numerous specific details are given in the following detailed implementation manners. Those skilled in the art should understand that the present disclosure can be implemented without some specific details. In some instances, methods, means, elements, and circuits well-known to those skilled in the art are not described in detail so as to highlight the gist of the present disclosure.
[0065] The whole-slide image processing method according to the embodiments of the present disclosure can be executed by an electronic device such as a terminal device or a server. Among them, the terminal device can be any fixed or mobile terminal such as a user equipment (UE), a mobile device, a user terminal, a terminal, a cellular phone, a cordless phone, a personal digital assistant (PDA), a handheld device, a computing device, a vehicle-mounted device, a wearable device, etc. The server can be a single server or a server cluster composed of multiple servers. Any electronic device can implement the whole-slide image processing method according to the embodiments of the present disclosure by a processor calling computer-readable instructions stored in a memory.
[0066] In the embodiments of the present disclosure, the whole-slide image processing method can be used in any application scenarios of whole-slide images in medical imaging, such as in the fields of pathological diagnosis, research, and education.
[0067] Exemplarily, the whole-slide image processing method of the embodiments of the present disclosure can perform tumor evaluation such as tumor detection, grading, and typing on the image feature information obtained by processing the whole-slide image; or perform pathological feature analysis of a large number of tissue samples on the image feature information obtained by processing the whole-slide image to assist in disease research; or perform disease diagnosis on the image feature information obtained by processing the whole-slide image and a deep neural network structure or diagnose the user's physiological parameters; or automatically segment the tumor region, evaluate biomarker expression, and even predict molecular changes on the image feature information obtained by processing the whole-slide image. In short, the whole-slide image technology has broad application prospects in medical imaging, not only improving the diagnostic efficiency and accuracy, but also promoting the development of medical research and education.
[0068] Any of the above application scenarios depends on the process of processing the whole-slide image to obtain image feature information that can accurately reflect the features of the whole-slide image. Therefore, the technical problem to be solved by the embodiments of the present disclosure is: how to improve the accuracy of the image feature extraction result of the whole-slide image through image processing. Based on this technical problem, the embodiments of the present disclosure enhance the attention degree of the instance position in the instance aggregation process by setting a position enhancement module in the feature extraction model, and improve the stability of the instance aggregation process by setting an attention module in the feature extraction model, and finally obtain more accurate image feature information that can reflect the features of the whole-slide image.
[0069] Figure 1 The flowchart of a whole-slide image processing method according to an embodiment of the present disclosure is shown. As Figure 1 shown, the whole-slide image processing method of the embodiments of the present disclosure can include the following steps S10-S20.
[0070] The following description is made with an electronic device as the execution subject. It can be understood that the electronic device is not limited to the electronic device itself, but also includes modules such as a processor or a processing chip in the electronic device that can execute computer tasks.
[0071] Step S10: The electronic device determines a plurality of instance region images included in the whole-slide image, and coordinate information corresponding to each instance region image.
[0072] In a possible implementation, the electronic device in the embodiments of the present application can obtain a whole-slide image. Among them, a whole-slide image refers to a complete and magnifiable digital image generated by scanning a pathological section row by row at a high resolution through a professional scanning device. In the embodiments of the present application, the electronic device can obtain a whole-slide image by performing image acquisition on a target object, or receive a whole-slide image transmitted after another device performs image acquisition on the target object. The target object is a living object such as a human or an animal, or a local organ such as the heart, lungs, liver, etc. on a living creature such as a human or an animal.
[0073] After the electronic device obtains the whole-slide image, it can further determine a plurality of instance region images from the obtained whole-slide image and determine the coordinate information corresponding to each instance region image. Optionally, the electronic device can perform instance segmentation on the whole-slide image by means of image segmentation, obtain the region where each instance is located as an instance region image, and determine the position of each instance region image in the whole-slide image as coordinate information. Among them, the electronic device can perform instance segmentation on the whole-slide image through a pre-trained visual encoder.
[0074] In some optional embodiments, each instance region image is an image block of the same size. A plurality of instance region images can form an image block sequence, and the coordinate information corresponding to the plurality of instance region images can form a coordinate information sequence.
[0075] Step S20: The electronic device inputs the plurality of instance region images and the corresponding coordinate information into the trained feature extraction model and outputs the corresponding image feature information.
[0076] In a possible implementation, in the embodiments of the present application, after the electronic device obtains the whole-slide image, performs segmentation on the whole-slide image to obtain a plurality of instance region images, and the coordinate information corresponding to each instance region image, it can extract the features of the whole-slide image based on the plurality of instance region images and the coordinate information corresponding to each instance region image according to the pre-trained feature extraction model, and obtain the image feature information corresponding to the whole-slide image.
[0077] Optionally, the feature extraction model in the embodiments of the present application may include a position enhancement module, a two-dimensional rotation embedding module, and an attention module. The position enhancement (Random Projection) module is used to map the coordinate information corresponding to each instance region image into a low-dimensional space while preserving the structure and distance information of the original data as much as possible. The attention module is an important mechanism in deep learning, which is used to help the model process input data more effectively, especially when dealing with data with complex structures or long-distance dependencies. The core role of the attention module is to enable the model to dynamically focus on the most important parts of the input data, thereby improving the performance and efficiency of the model.
[0078] In some possible implementation manners, the attention module in the embodiments of the present application may be an EsGLABlock (Efficient Spatial Grouped Local Attention Block). This type of attention module is suitable for tasks that require efficient processing of high-resolution images, such as object detection, image segmentation, and video analysis. While maintaining performance, it significantly reduces the consumption of computing resources. Applying this type of attention module in the embodiments of the present application can stabilize the attention entropy by designing global tokens based on semantic information and applying self-attention between each token and its neighborhood and global tokens, ensuring consistent and stable aggregation of WSIs with different instance numbers.
[0079] Exemplarily, based on the settings of the position enhancement module, the two-dimensional rotation embedding module, and the attention module, after the electronic device in the embodiments of the present application inputs multiple instance region images and corresponding coordinate information into the trained feature extraction model, it inputs the multiple coordinate information into the position enhancement module for position enhancement to obtain the enhanced positions corresponding to each coordinate information. Instance fusion with redundancy removal is performed based on the multiple instance region images and corresponding enhanced positions to obtain multiple instance coding information and multiple enhanced positions after removing redundancy. The multiple instance coding information and the multiple enhanced positions after passing through the two-dimensional rotation embedding module are input into the attention module for information interaction and instance aggregation to obtain the image feature information corresponding to the whole-slide image.
[0080] In some possible implementation manners, the way for the position enhancement module in the embodiments of the present application to perform position embedding on the coordinate information may be to perform random projection and random rotation on each coordinate information.
[0081] In some embodiments, the process of random projection can be represented by the formula f x (p x ) = α x p x + β x , f y (p y ) = αy p y + β y is implemented to randomly map each coordinate information included in the set P of coordinate information with an upper bound P max to a specified range. Among them, the coordinate sequence p x can be used to represent the sequence composed of the x coordinates in each coordinate information, and the coordinate sequence p y can be used to represent the sequence composed of the y coordinates in each coordinate information. f x and f y are pre-constructed random mapping functions, and α x , α y , β x , β y are randomly generated parameters to ensure that the coordinate information after random projection covers a wider spatial range.
[0082] In some other embodiments, due to the irregular shape of the tissue region of the whole slide image, the distribution of coordinate information may vary in different directions. The random rotation operation in the embodiments of the present application is used to reduce the influence of this irregularity on the random projection process. Therefore, the random rotation operation can be introduced before or after the random projection. Optionally, the process of random rotation can be implemented by the formula p' x = p x cos(θ) - p y sin(θ), p' y = p x sin(θ) + p y cos(θ), where θ is a randomly generated rotation angle. Through random rotation and random projection, the embodiments of the present application can obtain an enhanced position after position embedding of the coordinate information, effectively enhancing the generalization ability of the feature extraction model for spatial position information, and at the same time avoiding the problem of attention score explosion caused by uneven coordinate distribution.
[0083] Optionally, the feature extraction module in the embodiments of the present application may further include an instance fusion module. While the electronic device performs position enhancement on multiple coordinate information, it can also perform image encoding on each instance region image to obtain corresponding instance encoding information. And through instance fusion of inputting multiple instance encoding information and corresponding enhanced positions into the instance fusion module for redundancy removal processing, the remaining multiple instance encoding information and corresponding enhanced positions after removing redundancy are output. Among them, the number of output instance encoding information and enhanced positions is less than the number of input instance encoding information and enhanced positions. Optionally, the instance fusion module may be composed of a 2×2 pooling layer, and the neighboring instance encoding information and enhanced positions are respectively fused through average pooling operations to reduce the redundancy.
[0084] In some embodiments, the image encoding module can be an image block encoder (patch-level Encoder), which encodes the full-slice image in the instance dimension, that is, it can encode each instance region image separately to obtain corresponding instance encoding information. After obtaining the instance encoding information of multiple instance region images and the enhanced positions after enhancing each coordinate information position, the multiple instance encoding information and the corresponding enhanced positions can be input into the Instance Merge module for instance fusion to obtain the remaining multiple instance encoding information and the corresponding enhanced positions after removing redundancy. Exemplarily, the way the instance fusion module fuses the instance region images based on the enhanced positions can be average pooling in a 2×2 spatial region. This fusion process is used to reduce the redundancy of the instance encoding information corresponding to the current multiple instance region images, and obtain a sequence Z′∈R N′×d Output as the instance encoding information sequence, where N represents the number of instance region images input to the model, N′ is the number of remaining instance encoding information after removing redundancy, and d represents the dimension of the instance encoding information corresponding to each instance region image.
[0085] In a possible implementation, the feature extraction model of the embodiments of the present application includes a two-dimensional rotation embedding module, which is used to randomly map the input information to further project the enhanced positions obtained after enhancing the coordinate information into an under-learned space while maintaining the relative position relationship between different coordinate information. The setting of this two-dimensional rotation embedding module is used to enhance the generalization ability of the feature extraction model in the embodiments of the present disclosure for spatial position information.
[0086] Based on the setting of the two-dimensional rotation embedding module in the above feature extraction model, the process of information interaction and instance aggregation between instances through the attention module in the embodiments of the present application can include: inputting the multiple instance encoding information and the multiple enhanced positions after passing through the two-dimensional rotation embedding module into the attention module for semantic information extraction to obtain the semantic information corresponding to each instance encoding information, where the semantic information includes key information and value information. Then, based on the semantic information corresponding to each instance encoding information, instance aggregation is performed on the multiple instance encoding information to obtain an information aggregation result. Finally, for each instance encoding information, the corresponding neighboring instance encoding information is searched in the information aggregation result, and a self-attention operation is performed according to the information aggregation result to obtain the image feature information corresponding to the full-slice image.
[0087] Optionally, in the embodiments of the present application, the process of inputting the encoded information of each instance into the attention module to extract semantic information, and obtaining the semantic information corresponding to the encoded information of each instance may be as follows: First, semantic extraction is performed on the encoded information of each embodiment to obtain corresponding candidate information, which also includes key information and value information. Further, the enhanced position is further embedded into the key information and value information in the candidate information through a two-dimensional rotation embedding module to obtain semantic information.
[0088] In some embodiments, the process of obtaining candidate information by the above semantic information extraction in the embodiments of the present application may be implemented by the formula Q s = Z'W Q , K s = Z'W K where Q s is the key information in the candidate information, K s is the value information in the candidate information, Z' is a sequence composed of multiple instance encoded information remaining after removing redundancy, and W Q and W K are pre-learned weight matrices.
[0089] In some embodiments, the above two-dimensional rotation embedding module may, based on the enhanced position obtained after position embedding of the coordinate information, embed the position information into the key information and value information of the encoded information of each instance through matrix multiplication to obtain semantic information. Optionally, the two-dimensional rotation embedding module may be a two-dimensional rotation position encoding module (2D Rotary Position Embedding, 2D-RoPE), and the above position information embedding is realized through a position encoding method extended from RoPE (Rotary Position Embedding). The core idea of this encoding method is to rotate the vector through a rotation matrix to retain the relative position information.
[0090] Exemplarily, in the embodiments of the present application, the two-dimensional embedding module may perform position encoding on the instance encoded information corresponding to each instance in the full-slice image through the formulas Q = 2D-RoPE(Q s , P') and K = 2D-RoPE(K s , P'). Where P' is the enhanced position, that is, the coordinate information after position embedding processing, and Q and K respectively represent the key information and value information corresponding to the instance encoded information after position embedding by the two-dimensional embedding module. Q s and K s are the original key information and value information corresponding to the instance encoded information without position embedding.
[0091] In some possible implementation manners, after obtaining the enhanced positions through the two-dimensional rotation embedding module based on the coordinate information for position embedding, and after embedding the position information into the key information and value information of each instance encoding information to obtain semantic information, the embodiments of the present application perform instance aggregation on multiple instance encoding information based on the obtained semantic information of each instance encoding information to obtain an information aggregation result. Among them, the process of this instance aggregation can be implemented through a global query, that is, the information aggregation can be performed through the global query formula to obtain the information aggregation result, where Z g is the information aggregation result, Q s is the matrix composed of the key information corresponding to each instance encoding information, is the transpose of the matrix composed of the value information corresponding to each instance encoding information, d is the dimension of the instance encoding information, and Z′ is the sequence composed of the remaining multiple instance encoding information after removing redundancy.
[0092] In some embodiments, after the embodiments of the present application obtain the information aggregation result through instance aggregation, for each instance encoding information, the corresponding neighborhood instance encoding information is searched in the information aggregation result, and a self-attention operation is performed based on the information aggregation result to obtain the image feature information corresponding to the whole-slide image. Optionally, for each instance encoding information, the embodiments of the present application may search for the corresponding neighborhood instance encoding information through a tree search algorithm, and then perform a self-attention operation in combination with the information aggregation result representing the global instance features to obtain the image feature information corresponding to the whole-slide image.
[0093] Optionally, the manner in which the embodiments of the present application determine the image feature information may be through the following formula: Performing a self-attention operation based on the information aggregation result to obtain the image feature information corresponding to the whole-slide image, where z is an instance, and z out is the image feature information, Z nei is the matrix composed of the neighborhood instance encoding information corresponding to each instance encoding information, Z g is the information aggregation result, 2D-RoPE is the two-dimensional rotation position embedding function, p z is the enhanced position of the instance encoding information z, P is the enhanced position corresponding to Z nei and Z g , and W Q and W K are weight vectors. Thus, the attention module of the embodiments of the present application can effectively suppress the explosion of attention entropy through the global-local attention mechanism and maintain the relative stability of attention entropy during the training and inference processes.
[0094] In a possible implementation, after obtaining the image feature information, the attention module can also perform a non-linear mapping on the obtained image feature information through a single-layer MLP (Multilayer Perceptron), and the mapping formula can be Z out = MLP(Z′ + z out ), where Z′ is a sequence composed of multiple instance encoding information remaining after removing redundancy, z out is the image feature information, and Z out is the image feature information after non-linear mapping.
[0095] In some other embodiments, after performing non-linear mapping on the image feature information, the embodiments of the present application can also perform average pooling on the image feature information obtained after deep interaction between instances. The specific formula can be Z WSI = MeanPool(Z out ).
[0096] Figure 2 FIG. shows a schematic diagram of a full-slice image processing process according to an embodiment of the present disclosure. As Figure 2 shown, in the embodiments of the present application, after the electronic device obtains the full-slice image, it can obtain multiple instance region images and the coordinate information corresponding to each instance region image through image segmentation, and then input the multiple instance region images and the coordinate information corresponding to each instance region image into the feature extraction model to extract the image feature information of the full-slice image. Among them, the image coding module can perform image coding on each instance region image to obtain coded image information; the position enhancement module can perform position enhancement processing on each coordinate information to obtain enhanced positions. Further, the instance fusion module performs redundancy-removing instance fusion on the multiple coded image information according to the enhanced positions to obtain multiple instance coding information after removing redundancy and multiple enhanced positions.
[0097] Then, the attention module performs semantic extraction on the multiple instance coding information after removing redundancy to obtain candidate information, and then the two-dimensional rotation embedding module embeds the enhanced positions into the candidate information to obtain text information. Then, instance aggregation and information interaction between instances are performed based on the text information to obtain image feature information.
[0098] Figure 3 FIG. shows a schematic diagram of the structure of an attention module according to an embodiment of the present disclosure. As Figure 3As shown in the figure, in the embodiment of the present application, the attention module can search for the corresponding neighborhood instance coding information in the information aggregation result for each instance coding information remaining after removing redundant processing, and perform self-attention operation according to the information aggregation result to obtain the image feature information corresponding to the whole-slide image. Through the structure of this attention module, the embodiment of the present application proposes a global-local attention mechanism. This mechanism balances the stability of attention entropy and the ability to capture global information through global semantic aggregation and local neighborhood interaction. Due to the ultra-high resolution and tissue continuity of WSIs, adjacent instances often contain similar information. Thus, it is possible to ensure consistent and stable aggregation of WSIs with different numbers of instances.
[0099] Figure 4 FIG. shows a schematic structural diagram of a position enhancement module according to an embodiment of the present disclosure. As Figure 4 shown, in the embodiment of the present application, the position enhancement module can perform position embedding through random projection and random rotation. Without destroying the relative position relationship of coordinates, the coordinates are randomly projected into an under-learned coordinate space, enabling the model to fully learn the un-distributed coordinates. Combined with the two-dimensional rotation embedding module, it enhances the utilization ability of spatial information in the instance aggregation process.
[0100] Figure 5 FIG. shows a schematic diagram of an application process according to an embodiment of the present disclosure. As Figure 5 shown, the electronic device in the embodiment of the present application can obtain a whole-slide image by performing image acquisition on a target object, and then determine multiple instance region images and corresponding coordinate information based on the whole-slide image. The multiple instance region images and the coordinate information corresponding to each instance region image are input into a feature extraction model to extract the image feature information of the whole-slide image. Further, the electronic device performs feature prediction based on the image feature information corresponding to the whole-slide image to obtain the biometric information of the target object.
[0101] Optionally, the biometric information in the embodiment of the present application can be any biometric related to the target object, such as cancer subtype classification, prognosis analysis, and biomarker prediction, or some clinical parameters required clinically.
[0102] Based on the above technical features, the embodiment of the present application can improve the stability and generalization ability of the model in processing variable-length instance sequences and spatial position embedding by setting the position enhancement module and the two-dimensional rotation embedding module, thereby enhancing the attention degree of the instance position in the instance aggregation process. And through the setting of the global-local attention mechanism in the attention module, it stabilizes the attention entropy and enhances the spatial position embedding ability, improving the stability of the instance aggregation process, and finally obtaining more accurate image feature information. Further, when performing biometric prediction based on the image feature information extracted in this way, it can more accurately predict biometric features.
[0103] Figure 6 A schematic diagram showing a whole-slide image processing device according to an embodiment of the present disclosure. As Figure 6 shown, a whole-slide image processing device according to an embodiment of the present application may include:
[0104] An information determination module 60, configured to determine a plurality of instance region images included in the whole-slide image, and coordinate information corresponding to each of the instance region images, where the coordinate information is used to characterize the position of the corresponding instance region image in the whole-slide image;
[0105] A feature extraction module 61, configured to input the plurality of instance region images and the corresponding coordinate information into a trained feature extraction model, and output corresponding image feature information;
[0106] Wherein, the feature extraction model includes a position enhancement module, a two-dimensional rotation embedding module, and an attention module; the feature extraction module 61 is further configured to:
[0107] Input the plurality of coordinate information into the position enhancement module for position enhancement to obtain an enhanced position corresponding to each coordinate information;
[0108] Perform instance fusion with redundancy removal based on the plurality of instance region images and the corresponding enhanced positions to obtain a plurality of instance coding information with redundancy removed and a plurality of enhanced positions;
[0109] Input the plurality of instance coding information and the plurality of enhanced positions after passing through the two-dimensional rotation embedding module into the attention module for information interaction and instance aggregation to obtain the image feature information corresponding to the whole-slide image.
[0110] In a possible implementation manner, the feature extraction model further includes an instance fusion module, and the feature extraction module 61 is further configured to:
[0111] Perform image coding on each of the instance region images to obtain corresponding instance coding information;
[0112] Input the plurality of instance coding information and the corresponding enhanced positions into the instance fusion module for instance fusion with redundancy removal, and output the remaining plurality of instance coding information with redundancy removed and the corresponding enhanced positions, where the number of output instance coding information and enhanced positions is less than the number of input instance coding information and enhanced positions.
[0113] In a possible implementation manner, the feature extraction module 61 is further configured to:
[0114] Encode the multiple instance encoding information and the multiple enhanced positions after passing through the two-dimensional rotation embedding module into the attention module for semantic information extraction, to obtain the semantic information corresponding to each instance encoding information, where the semantic information includes key information and value information;
[0115] Perform instance aggregation on the multiple instance encoding information based on the corresponding semantic information to obtain an information aggregation result;
[0116] For each instance encoding information, search for the corresponding neighboring instance encoding information in the information aggregation result, and perform self-attention operation based on the information aggregation result to obtain the image feature information corresponding to the full-slice image.
[0117] In a possible implementation, the feature extraction module 61 is further configured to:
[0118] Input the multiple coordinate information into the position enhancement module for random projection and random rotation to obtain the enhanced position corresponding to each coordinate information.
[0119] In a possible implementation, the feature extraction module 61 is further configured to:
[0120] Perform information aggregation through the formula to obtain the information aggregation result, where Z g is the information aggregation result, Q s is the matrix composed of the key information corresponding to each instance encoding information, is the transpose of the matrix composed of the value information corresponding to each instance encoding information, d is the dimension of the instance encoding information, and Z′ is the matrix composed of multiple instance encoding information.
[0121] In a possible implementation, the feature extraction module 61 is further configured to:
[0122] Perform self-attention operation based on the information aggregation result through the formula to obtain the image feature information corresponding to the full-slice image, where z is an instance, z out is the image feature information, Z nei is the matrix composed of the neighboring instance encoding information corresponding to each instance encoding information, Z g is the information aggregation result, 2D-RoPE is the two-dimensional rotation position embedding function, p z is the enhanced position of the instance encoding information z, P is the enhanced position corresponding to Z nei and Z g , and W Q and W K are weight vectors.
[0123] In a possible implementation, the whole slide image is obtained by collecting an image of a target object, and the apparatus further includes:
[0124] A feature prediction module, configured to perform feature prediction according to the image feature information corresponding to the whole slide image to obtain the biometric information of the target object.
[0125] In some embodiments, the functions or modules included in the apparatus provided in the embodiments of the present disclosure may be used to execute the methods described in the above method embodiments. The specific implementation may refer to the description of the above method embodiments. For the sake of brevity, it will not be repeated here.
[0126] The embodiments of the present disclosure also propose a computer-readable storage medium, on which computer program instructions are stored. When the computer program instructions are executed by a processor, the above methods are implemented. The computer-readable storage medium may be a volatile or non-volatile computer-readable storage medium.
[0127] The embodiments of the present disclosure also propose an electronic device, including: a processor; a memory for storing instructions executable by the processor; wherein the processor is configured to implement the above methods when executing the instructions stored in the memory.
[0128] The embodiments of the present disclosure also provide a computer program product, including computer-readable code, or a non-volatile computer-readable storage medium carrying the computer-readable code. When the computer-readable code runs in a processor of an electronic device, the processor in the electronic device executes the above methods.
[0129] Figure 7 FIG. shows a schematic diagram of an electronic device 1900 according to an embodiment of the present disclosure. For example, the electronic device 1900 may be provided as a server or a terminal device. Referring to Figure 7 , the electronic device 1900 includes a processing component 1922, which further includes one or more processors, and memory resources represented by a memory 1932 for storing instructions executable by the processing component 1922, such as application programs. The application programs stored in the memory 1932 may include one or more modules each corresponding to a set of instructions. In addition, the processing component 1922 is configured to execute instructions to perform the above methods.
[0130] The electronic device 1900 may further include a power supply component 1926 configured to perform power management of the electronic device 1900, a wired or wireless network interface 1950 configured to connect the electronic device 1900 to a network, and an input / output interface 1958 (I / O interface). The electronic device 1900 may operate based on an operating system stored in the memory 1932, such as Windows Server TM , Mac OS X TM , Unix TM , Linux TM , FreeBSD TM or the like.
[0131] In an exemplary embodiment, there is also provided a non-transitory computer-readable storage medium, such as the memory 1932 including computer program instructions, and the above computer program instructions can be executed by the processing component 1922 of the electronic device 1900 to complete the above method.
[0132] The present disclosure may be a system, a method, and / or a computer program product. The computer program product may include a computer-readable storage medium having thereon computer-readable program instructions for causing a processor to implement various aspects of the present disclosure.
[0133] A computer-readable storage medium may be a tangible device that can retain and store instructions for use by an instruction execution device. A computer-readable storage medium may be, for example, but not limited to, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of the computer-readable storage medium include: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disc (DVD), a memory stick, a floppy disk, a mechanically encoded device such as a punched card or raised structures in grooves having instructions stored thereon, and any suitable combination of the foregoing. The computer-readable storage medium as used herein is not construed as being a transitory signal per se, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (e.g., an optical pulse through an optical fiber cable), or an electrical signal transmitted through a wire.
[0134] The computer-readable program instructions described herein can be downloaded to various computing / processing devices from a computer-readable storage medium or downloaded to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network may include copper transmission cables, optical fiber transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in a computer-readable storage medium in each computing / processing device.
[0135] The computer program instructions for performing the operations of the present disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state-setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages such as Smalltalk, C++, etc., and conventional procedural programming languages such as the "C" language or similar programming languages. The computer-readable program instructions may be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on the remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider). In some embodiments, by using the state information of the computer-readable program instructions to customize an electronic circuit, such as a programmable logic circuit, a field-programmable gate array (FPGA), or a programmable logic array (PLA), the electronic circuit can execute the computer-readable program instructions to implement various aspects of the present disclosure.
[0136] Aspects of the present disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the disclosure. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.
[0137] These computer-readable program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus to produce a machine such that the instructions, when executed by the processor of the computer or other programmable data processing apparatus, create an apparatus for implementing the functions / acts specified in one or more boxes of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, a programmable data processing apparatus, and / or other devices to function in a particular manner, such that the computer-readable medium storing the instructions comprises a manufacture including instructions for implementing various aspects of the functions / acts specified in one or more boxes of the flowchart and / or block diagram.
[0138] The computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other devices so that a series of operational steps are performed on the computer, other programmable data processing apparatus, or other devices to produce a computer-implemented process such that the instructions executed on the computer, other programmable data processing apparatus, or other devices implement the functions / acts specified in one or more boxes of the flowchart and / or block diagram.
[0139] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram may represent a module, a segment of code, or a portion of an instruction, and the module, segment of code, or portion of an instruction includes one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions noted in the boxes may occur out of the order noted in the figures. For example, two consecutive boxes may in fact be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each box in the block diagrams and / or flowcharts, and combinations of boxes in the block diagrams and / or flowcharts, can be implemented by a dedicated hardware-based system for performing the specified functions or acts, or by a combination of dedicated hardware and computer instructions.
[0140] The embodiments of the present disclosure have been described above. The above description is exemplary, not exhaustive, and is not limited to the disclosed embodiments. Many modifications and variations will be apparent to those of ordinary skill in the art in the field without departing from the scope and spirit of the described embodiments. The choice of terms used herein is intended to best explain the principles of the embodiments, the practical application, or improvements made to the technology in the market, or to enable other ordinary skilled persons in the art in the field to understand the embodiments disclosed herein.
Claims
1. A full-slice image processing method, characterized in that: The method comprises: Determine a plurality of instance region images included in the full-slice image, and coordinate information corresponding to each of the instance region images, wherein the coordinate information is used to characterize a position of the corresponding instance region image in the full-slice image; Inputting the plurality of instance region images and corresponding coordinate information into the trained feature extraction model, and outputting corresponding image feature information; Wherein, the feature extraction model includes a position enhancement module, a two-dimensional rotation embedding module and an attention module; The step of inputting the plurality of instance region images and the corresponding coordinate information into the trained feature extraction model and outputting the image feature information corresponding to the full slice image comprises: Inputting the plurality of coordinate information into the position enhancement module for position enhancement, and obtaining an enhanced position corresponding to each coordinate information; Instance fusion based on the multiple instance region images and the corresponding enhanced positions is performed to remove redundancy, thereby obtaining multiple instance encoding information and multiple enhanced positions after redundancy removal; The multiple instance encoding information and the multiple enhanced positions after passing through the two-dimensional rotation embedding module are input into the attention module for information interaction and instance aggregation to obtain image feature information corresponding to the full slice image.
2. The method according to claim 1, characterized in that The feature extraction model further includes an instance fusion module, wherein the instance fusion based on the multiple instance region images and the corresponding enhanced positions is subjected to redundancy removal processing to obtain multiple instance encoding information and multiple enhanced positions after redundancy removal, including: Performing image coding on each of the instance region images to obtain corresponding instance coding information; The multiple instance coding information and the corresponding enhanced positions are input into the instance fusion module for instance fusion with redundancy removal, and the multiple instance coding information and the corresponding enhanced positions remaining after the redundancy is removed are output, and the number of the output instance coding information and the number of the enhanced positions are less than the number of the input instance coding information and the number of the enhanced positions.
3. The method according to claim 1, characterized in that The step of inputting the plurality of instance encoding information and the plurality of enhanced positions after passing through the two-dimensional rotation embedding module into the attention module for information interaction and instance aggregation to obtain image feature information corresponding to the full slice image includes: Input the multiple instance encoding information and the multiple enhanced positions after passing through the two-dimensional rotation embedding module into the attention module to extract semantic information, and obtain semantic information corresponding to each instance encoding information, wherein the semantic information includes key information and value information; Performing instance aggregation on the plurality of instance encoding information based on the corresponding semantic information to obtain an information aggregation result; For each instance encoding information, the corresponding neighborhood instance encoding information is searched in the information aggregation result, and a self-attention operation is performed according to the information aggregation result to obtain image feature information corresponding to the full slice image.
4. The method according to claim 1, characterized in that: The step of inputting the plurality of coordinate information into the position enhancement module for position enhancement to obtain an enhanced position corresponding to each coordinate information includes: The plurality of coordinate information are input into the position enhancement module for random rotation and random projection to obtain an enhanced position corresponding to each coordinate information.
5. The method according to claim 3, characterized in that: The performing instance aggregation on the plurality of instance encoding information based on the corresponding semantic information to obtain an information aggregation result includes: By formula Perform information aggregation to obtain the information aggregation result, where Z g is the information aggregation result, Q s A matrix consisting of key information corresponding to each instance encoding information, The value information corresponding to each instance encoding information is the transpose of the matrix, d is the dimension of the instance encoding information, Z ′ A sequence consisting of encoding information for the plurality of instances.
6. The method according to claim 3, characterized in that: For each of the instance encoding information, searching for the corresponding neighborhood instance encoding information in the information aggregation result, and performing a self-attention operation according to the information aggregation result to obtain image feature information corresponding to the full slice image, including: By formula Based on the information aggregation result, a self-attention operation is performed to obtain image feature information corresponding to the full slice image, where z is an instance, z out is the image feature information, Z nei is a matrix consisting of the encoding information of each instance corresponding to the encoding information of the neighboring instances, Z g is the information aggregation result, the two-dimensional rotation position embedding function described in 2D-RoPE, p z is the enhanced position of the instance encoding information z, and P is Z nei and Z g The corresponding enhanced position, W Q and W K is the weight vector.
7. The method according to claim 1, characterized in that The full-slice image is obtained by acquiring an image of the target object, and the method further comprises: Feature prediction is performed based on the image feature information corresponding to the full-slice image to obtain the biological feature information of the target object.
8. A whole-slice image processing device, characterized in that: The device comprises: An information determination module, used to determine a plurality of instance region images included in the full slice image, and coordinate information corresponding to each of the instance region images, wherein the coordinate information is used to characterize a position of the corresponding instance region image in the full slice image; A feature extraction module, used for inputting the plurality of instance region images and corresponding coordinate information into a trained feature extraction model, and outputting corresponding image feature information; Wherein, the feature extraction model includes a position enhancement module, a two-dimensional rotation embedding module and an attention module; The feature extraction model is further used to: Inputting the plurality of coordinate information into the position enhancement module for position enhancement, and obtaining an enhanced position corresponding to each coordinate information; Instance fusion based on the multiple instance region images and the corresponding enhanced positions is performed to remove redundancy, thereby obtaining multiple instance encoding information and multiple enhanced positions after redundancy removal; The multiple instance encoding information and the multiple enhanced positions after passing through the two-dimensional rotation embedding module are input into the attention module for information interaction and instance aggregation to obtain image feature information corresponding to the full slice image.
9. An electronic device, characterized in that: include: processor; a memory for storing processor-executable instructions; Wherein, the processor is configured to implement the method described in any one of claims 1 to 9 when executing the instructions stored in the memory.
10. A non-volatile computer-readable storage medium having computer program instructions stored thereon, characterized in that: When the computer program instructions are executed by a processor, the method according to any one of claims 1 to 9 is implemented.
Citation Information
Patent Citations
Method and device for segmenting cancerization region of breast tissue slice
CN115439493A
Biological context for analyzing full slide images
CN118076970A
Breast cancer full-slice pathological image weak supervision classification method based on neighborhood aggregation graph network
CN119494981A
Histopathology full-slice image classification method and system based on position coding converter
CN119600598A
Biological context for analyzing whole slide images
US20240265541A1