Cell nucleus segmentation method and device based on multi-modal information and computer equipment

Through the nuclear segmentation method of multimodal information fusion, U-Net structure and multi-scale feature extraction technology are used to solve the accuracy and efficiency of nuclear segmentation in the existing technology, and high-precision segmentation in complex morphology and noise environments are achieved.

CN120340023APending Publication Date: 2025-07-18SOUTH CHINA NORMAL UNIV
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510209503.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-25
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The existing threshold segmentation technology based on pixels or regions is difficult to accurately segment the nucleus when facing nucleus images with complex morphology, low contrast or high noise, and is sensitive to image noise and light changes, so it is impossible to effectively deal with irregular shapes and overlapping nuclei.

Method used

The nuclear segmentation method based on multimodal information is adopted, and multi-scale features of pathological images and text prompt information are extracted, feature fusion and decoding are performed, and nuclear segmentation is performed using a U-Net structure model. Combined with image feature extraction, text feature extraction, feature fusion and decoding modules, the segmentation accuracy is improved.

Benefits of technology

It improves the accuracy and efficiency of cell nucleus segmentation, especially in complex morphology and noise environments, which can more accurately segment the cell nucleus.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120340023A_ABST
    Figure CN120340023A_ABST
Patent Text Reader

Abstract

The invention relates to the field of medical image segmentation, in particular to a cell nucleus segmentation method and device based on multi-modal information, computer equipment and a storage medium, and aims to extract different levels of image feature information of a pathological image in the multi-modal information and semantic feature information in text prompt information to perform feature fusion so as to obtain an image feature fusion result. And feature decoding is carried out according to the obtained multi-scale multi-modal fusion feature information and the image feature information, so that cell nucleus segmentation is carried out, and the accuracy and efficiency of cell nucleus segmentation are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of medical image segmentation, and particularly to a method, device, computer device and storage medium for nucleus segmentation based on multi-modal information. Background Art

[0002] Nucleus segmentation is one of the important tasks in pathological image analysis, and has important applications especially in the fields of cell counting, tissue area statistics, tumor diagnosis, etc.

[0003] Currently, nucleus classification mainly relies on pixel- or region-based threshold segmentation techniques. The core idea of such methods is to divide the pixels in the image into different categories, usually background and foreground (nucleus), by setting one or more thresholds. Common methods include the Otsu algorithm, histogram-based thresholding, entropy-based thresholding, etc. In the nucleus segmentation task, the advantage of the threshold method is its simple implementation and high computational efficiency, so it can achieve good results when dealing with nucleus images with obvious contrast between background and foreground and simple shapes. However, the disadvantages of these methods are also obvious. Especially when facing nucleus images with complex morphology, low contrast or more noise, the selection of the threshold becomes inaccurate. The threshold segmentation method is very sensitive to image noise and illumination changes, and also cannot provide an effective solution when dealing with irregularly shaped and overlapping nuclei. Summary of the Invention

[0004] Based on this, the object of the present invention is to provide a method, device, equipment and storage medium for nucleus segmentation based on multi-modal information, which extracts different levels of image feature information of pathological images and semantic feature information in text prompt information from multi-modal information, performs feature fusion, and decodes features according to the obtained multi-modal fusion feature information and image feature information at multiple scales for nucleus segmentation, thereby improving the accuracy and efficiency of nucleus segmentation.

[0005] In the first aspect, an embodiment of the present application provides a method for nucleus segmentation based on multi-modal information, including the following steps:

[0006] Obtain target multi-modal information and a nucleus segmentation model, where the target multi-modal information includes a pathological image to be segmented and corresponding text prompt information; the nucleus segmentation model includes an image feature extraction module, a text feature extraction module, a feature fusion module, a feature decoding module and a class prediction module; the text feature extraction module includes a text encoding unit and a semantic feature extraction unit;

[0007] Input the target multi-modal information into the nucleus segmentation model, and perform multi-scale image feature extraction according to the pathological image to be segmented and the image feature extraction module to obtain image feature representations at several scales;

[0008] Perform encoding processing according to the text prompt information and the text encoding unit to obtain a text encoding representation; perform semantic feature extraction according to the text encoding representation and the semantic feature extraction unit to obtain a text feature representation;

[0009] Perform feature fusion according to the image feature representations at several scales, the text feature representation, and the feature fusion module to obtain multi-modal feature fusion representations at several scales;

[0010] Perform feature decoding according to the multi-modal feature fusion representations at several scales, the image feature representation at the last scale, and the feature decoding module to obtain a feature decoding representation;

[0011] Perform nucleus category prediction according to the feature decoding representation, the pathological image to be segmented, the multi-modal feature fusion representation at the first scale, and the category prediction module to obtain nucleus category prediction data, and obtain the nucleus segmentation result of the target multi-modal text and image information according to the nucleus category prediction data.

[0012] In a second aspect, an embodiment of the present application provides a nucleus segmentation device based on multi-modal information, including:

[0013] A data acquisition module, configured to acquire target multi-modal information and a nucleus segmentation model, where the target multi-modal information includes a pathological image to be segmented and corresponding text prompt information; the nucleus segmentation model includes an image feature extraction module, a text feature extraction module, a feature fusion module, a feature decoding module, and a category prediction module; the text feature extraction module includes a text encoding unit and a semantic feature extraction unit;

[0014] An image processing module, configured to input the target multi-modal information into the nucleus segmentation model, and perform multi-scale image feature extraction according to the pathological image to be segmented and the image feature extraction module to obtain image feature representations at several scales;

[0015] A text processing module, configured to perform encoding processing according to the text prompt information and the text encoding unit to obtain a text encoding representation; perform semantic feature extraction according to the text encoding representation and the semantic feature extraction unit to obtain a text feature representation;

[0016] The first feature processing module is used to perform feature fusion according to the image feature representations, text feature representations of the several scales, and the feature fusion module to obtain multi-modal feature fusion representations of several scales;

[0017] The second feature processing module is used to perform feature decoding according to the multi-modal feature fusion representations of the several scales, the image feature representation of the last scale, and the feature decoding module to obtain a feature decoding representation;

[0018] The cell nucleus category prediction module is used to perform cell nucleus category prediction according to the feature decoding representation, the pathological image to be segmented, the multi-modal feature fusion representation of the first scale, and the category prediction module to obtain cell nucleus category prediction data, and obtain the cell nucleus segmentation result of the target multi-modal text and image information according to the cell nucleus category prediction data.

[0019] In a third aspect, an embodiment of the present application provides a computer device, including: a processor, a memory, and a computer program stored on the memory and executable on the processor; when the computer program is executed by the processor, the steps of the cell nucleus segmentation method based on multi-modal information as described in the first aspect are implemented.

[0020] In a fourth aspect, an embodiment of the present application provides a storage medium, the storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the cell nucleus segmentation method based on multi-modal information as described in the first aspect are implemented.

[0021] In an embodiment of the present application, a cell nucleus segmentation method, device, computer device, and storage medium based on multi-modal information are provided. Different levels of image feature information of a pathological image and semantic feature information in text prompt information in multi-modal information are extracted, feature fusion is performed, and feature decoding is performed according to the obtained multi-modal fusion feature information and image feature information of multiple scales for cell nucleus segmentation, improving the accuracy and efficiency of cell nucleus segmentation.

[0022] For better understanding and implementation, the present invention will be described in detail below with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] Figure 1 It is a schematic flowchart of a cell nucleus segmentation method based on multi-modal information provided by an embodiment of the present application;

[0024] Figure 2 It is a schematic flowchart of S2 in the cell nucleus segmentation method based on multi-modal information provided by an embodiment of the present application;

[0025] Figure 3Schematic flowchart of S3 in the nucleus segmentation method based on multimodal information provided by an embodiment of the present application;

[0026] Figure 4 Schematic flowchart of S4 in the nucleus segmentation method based on multimodal information provided by an embodiment of the present application;

[0027] Figure 5 Schematic flowchart of S5 in the nucleus segmentation method based on multimodal information provided by an embodiment of the present application;

[0028] Figure 6 Schematic flowchart of S6 in the nucleus segmentation method based on multimodal information provided by an embodiment of the present application;

[0029] Figure 7 Schematic flowchart of S7 in the nucleus segmentation method based on multimodal information provided by another embodiment of the present application;

[0030] Figure 8 Schematic structural diagram of the nucleus segmentation device based on multimodal information provided by an embodiment of the present application;

[0031] Figure 9 Schematic structural diagram of a computer device provided by an embodiment of the present application. Detailed implementation manners

[0032] Here, exemplary embodiments will be described in detail, and examples thereof are shown in the accompanying drawings. When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with the present application. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims.

[0033] The terms used in the present application are for the purpose of describing specific embodiments only and are not intended to limit the present application. The singular forms "a", "the", and "said" used in the present application and the appended claims are also intended to include the plural forms unless the context clearly dictates otherwise. It should also be understood that the term "and / or" as used herein refers to and encompasses any and all possible combinations of one or more of the associated listed items.

[0034] It should be understood that although terms such as first, second, and third may be used in this application to describe various information, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from each other. For example, without departing from the scope of this application, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the word "if" as used herein may be interpreted as "when" or "while" or "in response to determining".

[0035] Please refer to Figure 1 , Figure 1 which is a schematic flowchart of a multi-modal information-based cell nucleus segmentation method provided for an embodiment of this application. The method includes the following steps:

[0036] S1: Obtain target multi-modal information and a cell nucleus segmentation model.

[0037] The execution subject of the multi-modal information-based cell nucleus segmentation method is a prediction device for the multi-modal information-based cell nucleus segmentation method (hereinafter referred to as the prediction device). In an optional embodiment, the prediction device may be a computer device, which may be a server, or a server cluster formed by combining multiple computer devices.

[0038] In this embodiment, the prediction device can obtain the target multi-modal information input by the user. Among them, the target multi-modal information includes the pathological image to be segmented and the corresponding text prompt information. Specifically, the generation of the text prompt is based on the characteristics of the cell nuclei in the pathological image, such as morphology, size, position, and density, and constructs the corresponding text description through specific sentence patterns and filling methods. In an optional embodiment, the prediction device is based on a fixed sentence pattern framework, and embeds the obtained feature information of the morphology, size, position, and density of the cell nuclei as filling items into the sentence respectively, so as to automatically generate personalized text sentences for each cell nucleus category as the text prompt information.

[0039] The prediction device obtains a cell nucleus segmentation model. Among them, the cell nucleus segmentation model is a model based on the U-Net structure, which is composed of multiple modules, including an image feature extraction module, a text feature extraction module, a feature fusion module, a feature decoding module, and a class prediction module; the text feature extraction module includes a text encoding unit and a semantic feature extraction unit.

[0040] S2: Input the target multi-modal information into the cell nucleus segmentation model, and perform multi-scale image feature extraction according to the pathological image to be segmented and the image feature extraction module, and obtain image feature representations of several scales.

[0041] In this embodiment, the prediction device inputs the target multimodal information into the cell nucleus segmentation model, and performs multi-scale image feature extraction based on the pathological image to be segmented and the image feature extraction module, obtaining image feature representations at several scales.

[0042] The image feature extraction module includes an embedding unit and a downsampling unit, and the downsampling unit includes several layers of downsampling layers connected in sequence. Please refer to Figure 2 , Figure 2 which is a schematic flowchart of S2 in the cell nucleus segmentation method based on multimodal information provided by an embodiment of the present application, including steps S21 to S22, specifically as follows:

[0043] S21: Input the pathological image to be segmented into the embedding unit, and obtain an embedding feature representation according to a preset embedding feature extraction algorithm.

[0044] The embedding unit is composed of a series of convolutional layers, normalization layers, and ReLU activation functions. The convolutional layers are responsible for extracting local features. The normalization layers accelerate the training process and improve stability through standardization, while the ReLU activation functions effectively increase the non-linear expression ability of the network. Through the preliminary processing of the image, this layer enhances the representation ability of the image features, enabling the subsequent downsampling layers to better understand and analyze the important information in the image. In this embodiment, the prediction device inputs the pathological image to be segmented into the embedding unit, and obtains an embedding feature representation according to a preset embedding feature extraction algorithm, where the embedding feature extraction algorithm is:

[0045] f inc = ReLu(Conv(I img ))

[0046] In the formula, f inc is the embedding feature representation, ReLu(·) is the non-linear activation function, Conv(·) is the convolutional function, and I img is the pathological image to be segmented.

[0047] S22: Use the embedding feature representation as the input representation of the first downsampling layer of the downsampling unit, perform downsampling feature extraction according to a preset downsampling algorithm, obtain the downsampling feature representation output by the first downsampling layer, use the downsampling feature representation as the input representation of the next downsampling layer, repeat the downsampling feature extraction, and obtain the downsampling feature representations output by several layers of downsampling layers of the downsampling unit as the image feature representations at several scales.

[0048] The downsampling unit adopts the classic UNet structure, where each layer of the downsampling layer includes a pooling layer and a convolutional layer. The pooling layer reduces the spatial dimension of the image through downsampling operations, helps the model focus on important regions in the image, and suppresses unnecessary noise.

[0049] In this embodiment, the prediction device uses the embedded feature representation as the input representation of the first downsampling layer of the downsampling unit, and performs downsampling feature extraction according to a preset downsampling algorithm to obtain the downsampling feature representation output by the first downsampling layer. Among them, the downsampling algorithm is:

[0050]

[0051] f down =MaxPool(Conv(Conv(f inc )))

[0052] In the formula, f i is the downsampling feature representation output by the i-th layer of the downsampling layer, and MaxPool(·) is the max pooling function.

[0053] The prediction device uses the downsampling feature representation as the input representation of the next downsampling layer, repeats the downsampling feature extraction, and obtains the downsampling feature representations output by several layers of the downsampling unit as the image feature representations of several scales. Through this structure of layer-by-layer downsampling and feature extraction, local and global features in the image can be effectively captured, providing important spatial and semantic information for subsequent image decoding and nucleus segmentation tasks.

[0054] S3: Perform encoding processing according to the text prompt information and the text encoding unit to obtain a text encoding representation; perform semantic feature extraction according to the text encoding representation and the semantic feature extraction unit to obtain a text feature representation.

[0055] In this embodiment, the prediction device performs encoding processing based on the text prompt information and the text encoding unit to obtain a text encoding representation. Specifically, the prediction device performs normalization processing on the text prompt information. The normalization processing includes text cleaning and encoding conversion. The text cleaning includes removing irrelevant symbols, special characters, and redundant spaces to ensure the consistency and standardization of the text prompt information. The encoding conversion includes using a regularization processing method to split the text prompt information into several string units, constructing a list of string units, and performing BPE (Byte Pair Encoding) encoding processing on the list of string units. BPE is a common sub-word level tokenization method that can effectively solve the problem of words not seen in the vocabulary or rare words. BPE gradually compresses the entire vocabulary into smaller units by continuously merging the most frequently occurring character pairs, making the processing more flexible and efficient, and obtaining the text encoding representation of the list of string units. The text encoding representation includes text encoding vectors of several string units.

[0056] The prediction device performs semantic feature extraction based on the text encoding representation and the semantic feature extraction unit to obtain a text feature representation, which further enriches the semantic information of the model and improves the accuracy of nucleus segmentation. The semantic feature extraction unit uses a common Transformer structure. To improve the stability of the model during training, we load and freeze the parameters in the pre-trained CLIP pre-trained model, enabling the text encoder to better capture the semantic information in the medical field while ensuring precise processing of nucleus segmentation in subsequent tasks.

[0057] The semantic feature extraction unit includes several sequentially connected semantic feature extraction sub-units. The semantic feature extraction sub-unit includes a first multi-head self-attention layer and a feed-forward layer; please refer to Figure 3 , Figure 3 which is a schematic flowchart of S3 in the nucleus segmentation method based on multi-modal information provided by an embodiment of this application, including steps S31 to S32, specifically as follows:

[0058] S31: Use the text encoding representation as the input representation of the first multi-head self-attention layer of the first semantic feature extraction sub-unit, perform a linear transformation on the input representation, and obtain a first multi-head self-attention feature representation based on the matrix vector obtained from the linear transformation and the preset first multi-head self-attention feature extraction algorithm.

[0059] In this embodiment, the prediction device uses the encoded representation of the text as the input representation of the first multi-head self-attention layer of the first semantic feature extraction sub-unit, linearly transforms the input representation, and obtains a first multi-head self-attention feature representation according to the matrix vector obtained from the linear transformation and a preset first multi-head self-attention feature extraction algorithm, where the matrix vector includes a query vector, a key vector, and a value vector; the multi-head self-attention feature extraction algorithm is:

[0060] f t = LN(MLP(A))

[0061]

[0062] In the formula, f t is the first multi-head self-attention feature representation, LN(·) is a normalization function, MLP(·) is a multi-layer perceptron function, A is the self-attention weight of the multi-head self-attention layer, Q1 is the query vector of the multi-head self-attention layer, K1 is the key vector of the multi-head self-attention layer, V1 is the value vector of the multi-head self-attention layer, T is the transpose symbol, C is the dimension parameter, and softmax(·) is a normalization exponential function.

[0063] By multiplying the input representation by different weight matrices to obtain a query vector, a key vector, and a value vector, multiplying the query vector by the key vector and dividing by the square root of the dimension to obtain a self-attention score, and then multiplying by the value vector through a normalization exponential function to generate a self-attention weight, which represents the influence degree of the current Token on other Tokens. After calculating the output of self-attention, through the methods of residual connection and layer normalization, the output of each layer is made more stable, avoiding the phenomenon of gradient disappearance or explosion.

[0064] S32: Input the first multi-head self-attention feature representation into the forward propagation layer for dimensionality reduction processing, obtain the multi-head self-attention feature representation after dimensionality reduction processing as the output representation of the first semantic feature extraction sub-unit, use the output representation of the semantic feature extraction sub-unit as the input representation of the multi-head self-attention module layer of the next semantic feature extraction sub-unit, repeat the semantic feature extraction, and obtain the output representation of the last semantic feature extraction sub-unit as the text feature representation.

[0065] In this embodiment, the prediction device inputs the first multi-head self-attention feature representation into the forward propagation layer for dimensionality reduction processing, and obtains the multi-head self-attention feature representation after dimensionality reduction processing as the output representation of the first semantic feature extraction subunit. Specifically, the forward propagation layer consists of two fully connected networks and a non-linear activation function (such as ReLU or GELU), further enhancing the non-linear representation ability of the model. After the processing of the forward propagation layer, the output passes through a mapping layer for dimensionality reduction operation to ensure that its dimension matches the requirements of subsequent tasks.

[0066] Use the output representation of the semantic feature extraction subunit as the input representation of the multi-head self-attention module layer of the next semantic feature extraction subunit, repeat the semantic feature extraction, and obtain the output representation of the last semantic feature extraction subunit as the text feature representation. After multi-level feature processing, the text prompt information is fully encoded, and the final text feature representation is generated. This text feature representation not only contains the semantic information of the text, but also through the processing of multiple semantic feature extraction subunits, can capture the context relationship, detailed features and global semantics in the text, helping to improve the accuracy of the cell nucleus segmentation task.

[0067] S4: Perform feature fusion based on the image feature representations of the several scales, the text feature representation, and the feature fusion module to obtain multi-modal feature fusion representations of the several scales.

[0068] In this embodiment, the prediction device performs feature fusion based on the image feature representations of the several scales, the text feature representation, and the feature fusion module to obtain multi-modal feature fusion representations of the several scales.

[0069] The feature fusion module includes several image-text feature channel cross self-attention units; the image-text feature channel cross self-attention unit includes a second multi-head self-attention layer and a dimensionality reconstruction layer; please refer to Figure 4 , Figure 4 is a schematic flowchart of S4 in the cell nucleus segmentation method based on multi-modal information provided by an embodiment of the present application, including steps S41 to S43, specifically as follows:

[0070] S41: Perform dimensionality mapping on the image feature representations of the several scales to obtain image feature representations of the same dimension corresponding to the several scales, splice the image feature representations of the same dimension corresponding to the several scales with the text feature representation to obtain a spliced feature representation, and respectively combine the spliced feature representation and the image feature representations of the same dimension corresponding to the several scales to obtain feature combinations of the several scales.

[0071] In this embodiment, the prediction device performs dimensional mapping on the image feature representations of the several scales to obtain image feature representations of the same dimension corresponding to the several scales, so that feature maps with different sizes are unified to the same dimension. This process ensures that features from different levels can be effectively compared and combined in the subsequent fusion process, while avoiding computational problems caused by inconsistent sizes.

[0072] The prediction device splices the image feature representations of the same dimension corresponding to the several scales with the text feature representations to obtain spliced feature representations. The combination of the image features and the text features not only provides visual information for the image segmentation task, but also enables the model to enhance the nuclei in the image at the semantic level according to the text prompts.

[0073] The prediction device combines the spliced feature representations and the image feature representations of the same dimension corresponding to the several scales respectively to obtain feature combinations of the several scales.

[0074] S42: Input the feature combinations of the several scales into the several image-text feature channel cross-attention units respectively. Take the spliced feature representations in the feature combinations as the key vector and the value vector of the second multi-head self-attention layer respectively, and take the image feature representations of the same dimension as the query vector of the second multi-head self-attention layer. According to the second multi-head self-attention feature extraction algorithm, obtain the second multi-head self-attention feature representations output by the several second multi-head self-attention layers.

[0075] The second multi-head self-attention feature extraction algorithm is:

[0076]

[0077] In the formula, O i is the second multi-head self-attention feature representation of the i-th image-text feature channel cross-attention unit, A′ i is the self-attention weight of the i-th image-text feature channel cross-attention unit, N is the number of channels, e i is the image feature representation of the same dimension corresponding to the i-th scale, Q2 is the query vector of the second multi-head self-attention layer, K2 is the key vector of the second multi-head self-attention layer, and V2 is the value vector of the second multi-head self-attention layer.

[0078] In this embodiment, the prediction device inputs the feature combinations of several scales into several of the image-text feature channel cross-attention units respectively, uses the concatenated feature representations in the feature combinations as the key vectors and value vectors of the second multi-head self-attention layer respectively, uses the image feature representations of the same dimension as the query vectors of the second multi-head self-attention layer, and obtains the second multi-head self-attention feature representations output by several of the second multi-head self-attention layers according to the second multi-head self-attention feature extraction algorithm. The image and text features are processed in parallel by multiple attention heads, and each attention head can capture different association methods between the image and the text. The multi-head attention mechanism can parallelly focus on different hierarchical relationships between the image and text features in different subspaces, enhancing the model's multi-dimensional understanding of the nucleus morphology, position, and semantics, thereby providing a more comprehensive and accurate feature representation for fine nucleus segmentation. The features are further non-linearly transformed through a multi-layer perceptron (MLP) layer to enhance the expression ability of the features.

[0079] S43: Input the second multi-head self-attention feature representations into the corresponding dimension reconstruction layers, and obtain the dimension reconstruction feature representations output by several of the dimension reconstruction layers according to the preset dimension reconstruction algorithm. Concatenate the dimension reconstruction feature representations output by several of the dimension reconstruction layers with the image feature representations of the corresponding scales to obtain the multi-modal feature fusion representations of several scales.

[0080] In this embodiment, the prediction device inputs the second multi-head self-attention feature representations into the corresponding dimension reconstruction layers, and obtains the dimension reconstruction feature representations output by several of the dimension reconstruction layers according to the preset dimension reconstruction algorithm, where the dimension reconstruction algorithm is:

[0081] S i = ReLu(Conv(Up(O i )))

[0082] In the formula, S i is the dimension reconstruction feature representation of the i-th scale, Up(·) is the upsampling function, which is implemented by an upsampling layer, a convolutional layer, a batch normalization layer, and a ReLU activation function. The upsampling layer helps to increase the spatial resolution of the feature map to the required size, the convolutional layer is used to further extract spatial features, and the batch normalization layer and the ReLU activation function improve the stability and non-linear expression ability of the network.

[0083] The prediction device splices the dimension reconstruction feature representations output by several of the dimension reconstruction layers with the image feature representations at corresponding scales to obtain multi-modal feature fusion representations at several scales, which can deeply fuse multi-modal information from images and texts, capture details while maintaining multi-scale information, make the image features and text prompts complement each other, and thus provide rich and accurate feature representations for the final nucleus segmentation task.

[0084] S5: Perform feature decoding based on the multi-modal feature fusion representations at several scales, the image feature representations at the last scale, and the feature decoding module to obtain a feature decoding representation.

[0085] In this embodiment, the prediction device performs feature decoding based on the multi-modal feature fusion representations at several scales, the image feature representations at the last scale, and the feature decoding module to obtain a feature decoding representation, which can effectively fuse multi-modal information from images and texts and obtain the final pre-segmentation map through detailed feature transformation.

[0086] The feature decoding module includes several sequentially connected upsampling units; please refer to Figure 5 , Figure 5 FIG.

[0087] S51: Use the image feature representations at the last scale as the output representations of the first upsampling unit, use the multi-modal feature fusion representations at the last scale and the output representations of the first upsampling unit as the input representations of the next upsampling unit, and perform nucleus category segmentation according to a preset segmentation algorithm to obtain the nucleus category segmentation map output by the current upsampling unit.

[0088] The segmentation algorithm is:

[0089] U i =ReLu(Conv(Concat(M i ⊙X i ,X i ))), i = 1, 2,..., N - 1

[0090]

[0091] In the formula, U i is the nucleus category segmentation map output by the i-th upsampling unit, M i is the intermediate feature representation of the i-th upsampling unit, and X iis the corresponding multimodal feature fusion representation of the i-th scale input to the i-th upsampling unit, Concat(·) is the concatenation function, N is the number of cutting units, Sigmoid(·) is the activation function, AvgPool(·) is the average pooling function, and Up(·) is the upsampling function.

[0092] In this embodiment, the prediction device uses the image feature representation of the last scale as the output representation of the first upsampling unit, and uses the multimodal feature fusion representation of the last scale and the output representation of the first upsampling unit as the input representation of the next upsampling unit. According to the preset segmentation algorithm, nuclear category segmentation is performed to obtain the nuclear category segmentation map output by the current upsampling unit, and various types of information are integrated to ensure the full integration of details and global semantics.

[0093] S52: Use the nuclear category segmentation map output by the current upsampling unit and the multimodal feature fusion representation of the next scale as the input representation of the next upsampling unit, and repeat the upsampling process to obtain the upsampling feature representation output by the last upsampling unit as the feature decoding representation.

[0094] In this embodiment, the prediction device uses the nuclear category segmentation map output by the current upsampling unit and the multimodal feature fusion representation of the next scale as the input representation of the next upsampling unit, and repeats the upsampling process to obtain the upsampling feature representation output by the last upsampling unit as the feature decoding representation. Through the combination of the channel self-attention mechanism and non-linear transformation, the model's ability to capture features such as the morphology and position of the nucleus is effectively improved, thereby improving the accuracy of nuclear segmentation.

[0095] S6: Perform nuclear category prediction based on the feature decoding representation, the pathological image to be segmented, the multimodal feature fusion representation of the first scale, and the category prediction module to obtain nuclear category prediction data, and obtain the nuclear segmentation result of the target multimodal text and image information based on the nuclear category prediction data.

[0096] In this embodiment, the prediction device performs nuclear category prediction based on the feature decoding representation, the pathological image to be segmented, the multimodal feature fusion representation of the first scale, and the category prediction module to obtain nuclear category prediction data, where the nuclear category prediction data includes prediction probability data of several nuclear categories.

[0097] The prediction device confirms the nuclear category prediction probability data of each pixel point in the pathological image to be segmented according to the nuclear category prediction data, confirms the category of the nucleus of each pixel point in the image, and obtains the nuclear segmentation result of the target multimodal text and image information.

[0098] Extract the image feature information of different levels of the pathological image in the multimodal information and the semantic feature information in the text prompt information, perform feature fusion, and perform feature decoding based on the obtained multimodal fusion feature information and image feature information of multiple scales for nucleus segmentation, which improves the accuracy and efficiency of nucleus segmentation.

[0099] Please refer to Figure 6 , Figure 6 which is a schematic flowchart of S6 in the nucleus segmentation method based on multimodal information provided by an embodiment of the present application, including steps S61 to S62, specifically as follows:

[0100] S61: Obtain a first cross-feature map attention matrix according to the multimodal feature fusion representation of the first scale, the pathological image to be segmented, and a preset first cross-feature map attention matrix calculation algorithm.

[0101] The first cross-feature map attention matrix calculation algorithm is:

[0102]

[0103] In the formula, A img is the first cross-feature map attention matrix, is the matrix multiplication symbol, P(·) is the chunking function, Emb(·) is the embedding function, I img is the pathological image to be segmented, and X1 is the multimodal feature fusion representation of the first scale.

[0104] In this embodiment, the prediction device obtains a first cross-feature map attention matrix according to the multimodal feature fusion representation of the first scale, the pathological image to be segmented, and a preset first cross-feature map attention matrix calculation algorithm, which can measure the similarity between the image and the feature map, quantify the relationship between the input image and the output feature map of the intermediate encoding layer, and improve the accuracy in the nucleus segmentation task. Among them, the first cross-feature map attention matrix includes a first attention matrix of several nucleus categories.

[0105] S62: Obtain nucleus category prediction data according to the first cross-feature map attention matrix, the feature decoding representation, and a preset category prediction algorithm.

[0106] The category prediction algorithm is:

[0107]

[0108] In the formula, G is the nucleus category prediction data, and U1 is the feature decoding representation.

[0109] In this embodiment, the prediction device obtains nuclear category prediction data according to the first cross-feature map attention matrix, the feature decoding representation, and a preset category prediction algorithm.

[0110] In an alternative embodiment, it further includes step S7: training the nuclear segmentation model. Please refer to Figure 7 , Figure 7 which is a schematic flowchart of S7 in the nuclear segmentation method based on multi-modal information provided by another embodiment of this application, including steps S71 to S73, specifically as follows:

[0111] S71: Obtain sample multi-modal information, input the sample multi-modal information set into the nuclear segmentation model, and obtain the first cross-feature map attention matrix and nuclear category prediction data corresponding to a number of sample multi-modal information.

[0112] In this embodiment, the prediction device obtains sample multi-modal information, inputs the sample multi-modal information set into the nuclear segmentation model, and obtains the first cross-feature map attention matrix and nuclear category prediction data corresponding to a number of sample multi-modal information. Among them, the sample multi-modal information set includes a number of sample multi-modal information, and the sample multi-modal information includes sample pathological images and corresponding text prompt information.

[0113] Specifically, the sample pathological images in the sample multi-modal information can be obtained from a preset CoNSeP dataset, MoNuSAC dataset, and Lizard dataset. These datasets cover a variety of different types of pathological images, providing rich samples for the training and evaluation of the model.

[0114] S72: Obtain the true label images corresponding to a number of sample multi-modal information, calculate the second cross-feature map attention matrix corresponding to a number of sample multi-modal information according to the true label images corresponding to a number of sample multi-modal information and a preset second cross-feature map attention matrix calculation algorithm, and obtain a first loss value according to the first cross-feature map attention matrix, the second cross-feature map attention matrix, and a preset first loss function corresponding to a number of sample multi-modal information.

[0115] In this embodiment, the prediction device obtains the true label images corresponding to a number of sample multi-modal information, and calculates the second cross-feature map attention matrix corresponding to a number of sample multi-modal information according to the true label images corresponding to a number of sample multi-modal information and a preset second cross-feature map attention matrix calculation algorithm. Among them, the second cross-feature map attention matrix includes second attention matrices of a number of nuclear categories, and the second cross-feature map attention matrix calculation algorithm is:

[0116]

[0117] Wherein, A gt is the second cross-feature map attention matrix, and I gt is the ground truth label image;

[0118] The prediction device obtains a first loss value by calculating the cross-feature map attention loss according to the first cross-feature map attention matrix, the second cross-feature map attention matrix corresponding to a plurality of sample multi-modal information, and a preset first loss function, wherein the first loss function is:

[0119]

[0120] Wherein, Loss1 is the first loss value, N c is the number of nucleus categories, k is the k-th nucleus category, A' img is the first cross-feature map attention matrix corresponding to the sample multi-modal information, and A gt is the second cross-feature map attention matrix. By optimizing the representation of the feature map, it helps the model to better capture the spatial relationship and feature similarity between nuclei in the segmentation task, thereby improving the accuracy in the nucleus segmentation task, especially showing stronger robustness and accuracy in the segmentation of small samples and nuclei with complex morphologies.

[0121] S73: Obtain the ground truth data of the nucleus category. According to the ground truth data of the nucleus category, the predicted data of the nucleus category corresponding to a plurality of sample multi-modal information, and a preset second loss function, obtain a second loss value. Accumulate the first loss value and the second loss value to obtain a total loss value. Train the nucleus segmentation model according to the total loss value.

[0122] In this embodiment, the prediction device obtains the ground truth data of the nucleus category, wherein the ground truth data of the nucleus category includes the true probability data of a plurality of nucleus categories.

[0123] According to the ground truth data of the nucleus category, the predicted data of the nucleus category corresponding to a plurality of sample multi-modal information, and a preset second loss function, obtain a second loss value by calculating the Dice coefficient loss and the segmentation image loss, wherein the second loss function is:

[0124]

[0125] Where Loss2 is the second loss value, TP is the number of pixel points corresponding to the cell nuclei in the sample pathological image that are correctly identified, FP is the number of pixel points corresponding to the cell nuclei in the sample pathological image where the predicted result is a cell nucleus but the actual result is not a cell nucleus, FN is the number of pixel points corresponding to the cell nuclei in the sample pathological image where the predicted result is not a cell nucleus but the actual result is a cell nucleus, and w k is the weight parameter of the k-th cell nucleus category, y k is the true probability data of the k-th cell nucleus category, and p k is the predicted probability data of the k-th cell nucleus category. Specifically, the weight parameter w k is as follows:

[0126]

[0127] Where P is the total number of pixels in the image, and P c is the number of pixels occupied by the current category.

[0128] The prediction device accumulates the first loss value and the second loss value to obtain the total loss value, and trains the cell nucleus segmentation model according to the total loss value. Specifically, the prediction device uses a cosine annealing function to dynamically adjust the learning rate. The cosine annealing learning rate scheduling strategy can gradually reduce the learning rate according to the progress of training, which helps the model to optimize more carefully in the later stage of training, avoid overfitting and improve the final segmentation accuracy. Through these training strategies, overfitting can be effectively avoided, the robustness of the model can be improved, and the efficiency and stability during training can be ensured.

[0129] Please refer to Figure 8 , Figure 8 which is a schematic structural diagram of a cell nucleus segmentation device based on multi-modal information provided by an embodiment of the present application. The device can implement all or part of the cell nucleus segmentation device based on multi-modal information through software, hardware, or a combination of both. The device 8 includes:

[0130] A data acquisition module 81, configured to acquire target multi-modal information and a cell nucleus segmentation model, where the target multi-modal information includes a pathological image to be segmented and corresponding text prompt information; the cell nucleus segmentation model includes an image feature extraction module, a text feature extraction module, a feature fusion module, a feature decoding module, and a category prediction module; the text feature extraction module includes a text encoding unit and a semantic feature extraction unit;

[0131] An image processing module 82, configured to input the target multi-modal information into the cell nucleus segmentation model, and perform multi-scale image feature extraction according to the pathological image to be segmented and the image feature extraction module to obtain image feature representations at several scales;

[0132] A text processing module 83, configured to perform encoding processing according to the text prompt information and the text encoding unit to obtain a text encoding representation; perform semantic feature extraction according to the text encoding representation and the semantic feature extraction unit to obtain a text feature representation;

[0133] A first feature processing module 84, configured to perform feature fusion according to the image feature representations of the plurality of scales, the text feature representation, and the feature fusion module to obtain multi-modal feature fusion representations of the plurality of scales;

[0134] A second feature processing module 85, configured to perform feature decoding according to the multi-modal feature fusion representations of the plurality of scales, the image feature representation of the last scale, and the feature decoding module to obtain a feature decoding representation;

[0135] A nucleus category prediction module 86, configured to perform nucleus category prediction according to the feature decoding representation, the pathological image to be segmented, the multi-modal feature fusion representation of the first scale, and the category prediction module to obtain nucleus category prediction data, and obtain the nucleus segmentation result of the target multi-modal graphic and text information according to the nucleus category prediction data.

[0136] In the embodiment of the present application, through a data acquisition module, target multimodal information and a cell nucleus segmentation model are obtained. Among them, the target multimodal information includes a pathological image to be segmented and corresponding text prompt information; the cell nucleus segmentation model includes an image feature extraction module, a text feature extraction module, a feature fusion module, a feature decoding module, and a class prediction module; the text feature extraction module includes a text encoding unit and a semantic feature extraction unit; through an image processing module, the target multimodal information is input into the cell nucleus segmentation model, and according to the pathological image to be segmented and the image feature extraction module, multi-scale image feature extraction is performed to obtain image feature representations at several scales; through a text processing module, encoding processing is performed according to the text prompt information and the text encoding unit to obtain a text encoding representation; according to the text encoding representation and the semantic feature extraction unit, semantic feature extraction is performed to obtain a text feature representation; through a first feature processing module, according to the image feature representations at several scales, the text feature representation, and the feature fusion module, feature fusion is performed to obtain multi-modal feature fusion representations at several scales; through a second feature processing module, according to the multi-modal feature fusion representations at several scales, the image feature representation of the last scale, and the feature decoding module, feature decoding is performed to obtain a feature decoding representation; through a cell nucleus class prediction module, according to the feature decoding representation, the pathological image to be segmented, the multi-modal feature fusion representation of the first scale, and the class prediction module, cell nucleus class prediction is performed to obtain cell nucleus class prediction data, and according to the cell nucleus class prediction data, the cell nucleus segmentation result of the target multimodal graphic information is obtained. By extracting the image feature information at different levels of the pathological image in the multimodal information and the semantic feature information in the text prompt information, performing feature fusion, and performing feature decoding according to the obtained multi-modal fusion feature information and image feature information at multiple scales for cell nucleus segmentation, the accuracy and efficiency of cell nucleus segmentation are improved.

[0137] Please refer to Figure 9 , Figure 9 which is a schematic structural diagram of a computer device provided by an embodiment of the present application. The computer device 8 includes: a processor 91, a memory 92, and a computer program 93 stored on the memory 92 and executable on the processor 91; the computer device can store multiple instructions, and the instructions are suitable for being loaded and executed by the processor 91 to perform the method steps of the above Figures 1 to 7 shown embodiment. The specific execution process can refer to the specific description of the Figures 1 to 7 shown embodiment and will not be elaborated here.

[0138] Among them, the processor 91 may include one or more processing cores. The processor 91 uses various interfaces and circuits to connect various parts within the server, and by running or executing instructions, programs, code sets, or instruction sets stored in the memory 92, and by invoking the data in the memory 92, it executes various functions of the cell nucleus segmentation device 7 based on multi-modal information and processes data. Optionally, the processor 91 may be implemented in at least one hardware form of digital signal processing (DSP), field-programmable gate array (FPGA), or programmable logic array (PLA). The processor 91 may integrate one or a combination of a central processing unit 91 (CPU), a graphics processing unit 91 (GPU), and a modem. Among them, the CPU mainly processes the operating system, user interface, application programs, etc.; the GPU is responsible for rendering and drawing the content to be displayed on the touch display screen; the modem is used to process wireless communication. It can be understood that the above-mentioned modem may not be integrated into the processor 91 and may be implemented separately by a single chip.

[0139] Among them, the memory 92 may include a random access memory 92 (RAM), and may also include a read-only memory 92 (ROM). Optionally, the memory 92 includes a non-transitory computer-readable storage medium. The memory 92 can be used to store instructions, programs, code, code sets, or instruction sets. The memory 92 may include a program storage area and a data storage area. Among them, the program storage area may store instructions for implementing the operating system, instructions for at least one function (such as touch instructions, etc.), instructions for implementing the above-mentioned method embodiments, etc.; the data storage area may store the data involved in the above-mentioned method embodiments. Optionally, the memory 92 may also be at least one storage device located far from the aforementioned processor 91.

[0140] The embodiment of the present application also provides a storage medium, which can store multiple instructions, and the instructions are suitable for being loaded and executed by a processor to perform the Figures 1 to 7 method steps of the embodiments shown above, and the specific execution process can be referred to the Figures 1 to 7 specific description of the embodiments shown above, and details will not be repeated here.

[0141] Those skilled in the art can clearly understand that, for the convenience and conciseness of description, only the above-mentioned division of each functional unit and module is used as an example for illustration. In actual applications, the above functions can be assigned to different functional units and modules according to needs, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiments can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit. In addition, the specific names of each functional unit and module are only for the convenience of mutual distinction and do not limit the protection scope of this application. The specific working processes of the units and modules in the above system can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.

[0142] In the above embodiments, the descriptions of the respective embodiments have their own emphases. For parts not detailed or recorded in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0143] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or by a combination of computer software and electronic hardware. Whether these functions are executed in hardware or software depends on the specific application of the technical solution and the design constraint algorithm. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.

[0144] In the embodiments provided by the present invention, it should be understood that the disclosed device / terminal device and method can be implemented in other ways. For example, the device / terminal device embodiments described above are only illustrative. For example, the division of the modules or units is only a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces. The indirect coupling or communication connection of the device or unit can be in electrical, mechanical or other forms.

[0145] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0146] In addition, in each embodiment of the present invention, each functional unit can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit.

[0147] If the integrated module / unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the above-described embodiment methods of the present invention can also be completed by instructing relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above-described method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file, or some intermediate form, etc.

[0148] The present invention is not limited to the above embodiments. If various modifications or deformations of the present invention do not depart from the spirit and scope of the present invention, and if these modifications and deformations are within the scope of the claims of the present invention and equivalent technical scope, then the present invention also intends to include these modifications and deformations.

Claims

1. A method for nuclear segmentation based on multi-modal information, characterized in that Including the following steps: Obtain target multi-modal information and a nucleus segmentation model, where the target multi-modal information includes a pathological image to be segmented and corresponding text prompt information; the nucleus segmentation model includes an image feature extraction module, a text feature extraction module, a feature fusion module, a feature decoding module, and a class prediction module; the text feature extraction module includes a text encoding unit and a semantic feature extraction unit; Input the target multi-modal information into the nucleus segmentation model, and perform multi-scale image feature extraction according to the pathological image to be segmented and the image feature extraction module to obtain image feature representations of several scales; Perform encoding processing according to the text prompt information and the text encoding unit to obtain a text encoding representation; perform semantic feature extraction according to the text encoding representation and the semantic feature extraction unit to obtain a text feature representation; Perform feature fusion according to the image feature representations of several scales, the text feature representation, and the feature fusion module to obtain multi-modal feature fusion representations of several scales; Perform feature decoding according to the multi-modal feature fusion representations of several scales, the image feature representation of the last scale, and the feature decoding module to obtain a feature decoding representation; Perform nucleus class prediction according to the feature decoding representation, the pathological image to be segmented, the multi-modal feature fusion representation of the first scale, and the class prediction module to obtain nucleus class prediction data, and obtain the nucleus segmentation result of the target multi-modal graphic and text information according to the nucleus class prediction data.

2. The method for nucleus segmentation based on multimodal information according to claim 1, wherein The image feature extraction module includes an embedding unit and a downsampling unit, and the downsampling unit includes several layers of downsampling layers connected in sequence; The step of inputting the target multi-modal information into the nucleus segmentation model and performing multi-scale image feature extraction according to the pathological image to be segmented and the image feature extraction module to obtain image feature representations of several scales includes the steps: Input the pathological image to be segmented into the embedding unit, and obtain an embedding feature representation according to a preset embedding feature extraction algorithm; Use the embedding feature representation as the input representation of the first layer of the downsampling unit, perform downsampling feature extraction according to a preset downsampling algorithm to obtain the downsampling feature representation output by the first layer of the downsampling unit, use the downsampling feature representation as the input representation of the next layer of the downsampling layer, and repeat the downsampling feature extraction to obtain the downsampling feature representations output by several layers of the downsampling unit as the image feature representations of several scales.

3. The method for nuclear segmentation based on multi-modal information according to claim 2, wherein: The semantic feature extraction unit includes several sequentially connected semantic feature extraction sub-units, and the semantic feature extraction sub-unit includes a first multi-head self-attention layer and a feed-forward layer; The step of performing semantic feature extraction according to the text encoding representation and the semantic feature extraction unit to obtain a text feature representation includes the steps: Encode the text representation as the input representation of the first multi-head self-attention layer of the first semantic feature extraction sub-unit, linearly transform the input representation, and obtain the first multi-head self-attention feature representation according to the matrix vector obtained from the linear transformation and the preset first multi-head self-attention feature extraction algorithm; Input the first multi-head self-attention feature representation into the forward propagation layer for dimensionality reduction processing to obtain the dimensionality-reduced multi-head self-attention feature representation as the output representation of the first semantic feature extraction sub-unit. Use the output representation of the semantic feature extraction sub-unit as the input representation of the multi-head self-attention module layer of the next semantic feature extraction sub-unit, and repeat the semantic feature extraction to obtain the output representation of the last semantic feature extraction sub-unit as the text feature representation.

4. The method for nucleus segmentation based on multi-modal information according to claim 3, characterized in that: The feature fusion module includes a number of image-text feature channel cross-attention units; the image-text feature channel cross-attention unit includes a second multi-head self-attention layer and a dimensionality reconstruction layer; The feature fusion of the image feature representations, text feature representations of several scales and the feature fusion module to obtain multi-modal feature fusion representations of several scales includes the steps: Perform dimensionality mapping on the image feature representations of several scales to obtain image feature representations of the same dimension corresponding to several scales, splice the image feature representations of the same dimension corresponding to several scales with the text feature representation to obtain a spliced feature representation, and combine the spliced feature representation and the image feature representations of the same dimension corresponding to several scales respectively to obtain feature combinations of several scales; Input the feature combinations of several scales into several of the image-text feature channel cross-attention units respectively. Use the spliced feature representation in the feature combination as the key vector and value vector of the second multi-head self-attention layer respectively, and use the image feature representation of the same dimension as the query vector of the second multi-head self-attention layer. According to the second multi-head self-attention feature extraction algorithm, obtain the second multi-head self-attention feature representations output by several of the second multi-head self-attention layers; Input the second multi-head self-attention feature representations into the corresponding dimensionality reconstruction layers, and obtain the dimensionality reconstruction feature representations output by several of the dimensionality reconstruction layers according to the preset dimensionality reconstruction algorithm. Splice the dimensionality reconstruction feature representations output by several of the dimensionality reconstruction layers with the image feature representations of the corresponding scales to obtain multi-modal feature fusion representations of several scales.

5. The method for nuclear segmentation based on multi-modal information according to claim 4, wherein: The feature decoding module includes a number of upsampling units connected in sequence; The feature decoding of the multi-modal feature fusion representations of several scales, the image feature representation of the last scale and the feature decoding module to obtain the feature decoding representation includes the steps: Use the image feature representation of the last scale as the output representation of the first upsampling unit, use the multimodal feature fusion representation of the last scale and the output representation of the first upsampling unit as the input representation of the next upsampling unit, and perform nucleus category segmentation according to a preset segmentation algorithm to obtain the nucleus category segmentation map output by the current upsampling unit; Use the nucleus category segmentation map output by the current upsampling unit and the multimodal feature fusion representation of the next scale as the input representation of the next upsampling unit, repeat the upsampling process, and obtain the upsampling feature representation output by the last upsampling unit as the feature decoding representation.

6. The method for nuclear segmentation based on multimodal information according to claim 5, wherein The nucleus category prediction is performed according to the feature decoding representation, the pathological image to be segmented, the multimodal feature fusion representation of the first scale, and the category prediction module to obtain nucleus category prediction data, including the steps of: Obtain a first cross-feature map attention matrix according to the multimodal feature fusion representation of the first scale, the pathological image to be segmented, and a preset first cross-feature map attention matrix calculation algorithm; Obtain nucleus category prediction data according to the first cross-feature map attention matrix, the feature decoding representation, and a preset category prediction algorithm.

7. The method for nuclear segmentation based on multimodal information according to claim 6, wherein It further includes the step of training the nucleus segmentation model; The training of the nucleus segmentation model includes the steps of: Obtain sample multimodal information, input the sample multimodal information set into the nucleus segmentation model to obtain the first cross-feature map attention matrix and nucleus category prediction data corresponding to a number of sample multimodal information, wherein the sample multimodal information set includes a number of sample multimodal information, and the sample multimodal information includes a sample pathological image and corresponding text prompt information; Obtain the true label images corresponding to a number of sample multimodal information, obtain the second cross-feature map attention matrix corresponding to a number of sample multimodal information according to the true label images corresponding to a number of sample multimodal information and a preset second cross-feature map attention matrix calculation algorithm, and obtain a first loss value according to the first cross-feature map attention matrix, the second cross-feature map attention matrix corresponding to a number of sample multimodal information, and a preset first loss function; Obtain the true data of the nucleus category, obtain a second loss value according to the true data of the nucleus category, the nucleus category prediction data corresponding to a number of sample multimodal information, and a preset second loss function, accumulate the first loss value and the second loss value to obtain a total loss value, and train the nucleus segmentation model according to the total loss value.

8. A nucleus segmentation device based on multimodal information, characterized in that It includes: A data acquisition module for acquiring target multimodal information and a nucleus segmentation model, wherein the target multimodal information includes a pathological image to be segmented and corresponding text prompt information; the nucleus segmentation model includes an image feature extraction module, a text feature extraction module, a feature fusion module, a feature decoding module, and a category prediction module; the text feature extraction module includes a text encoding unit and a semantic feature extraction unit; An image processing module, configured to input the target multi-modal information into the cell nucleus segmentation model, perform multi-scale image feature extraction based on the pathological image to be segmented and the image feature extraction module, and obtain image feature representations at several scales; A text processing module, configured to perform encoding processing based on the text prompt information and the text encoding unit to obtain a text encoding representation; perform semantic feature extraction based on the text encoding representation and the semantic feature extraction unit to obtain a text feature representation; A first feature processing module, configured to perform feature fusion based on the image feature representations at several scales, the text feature representation, and the feature fusion module to obtain multi-modal feature fusion representations at several scales; A second feature processing module, configured to perform feature decoding based on the multi-modal feature fusion representations at several scales, the image feature representation at the last scale, and the feature decoding module to obtain a feature decoding representation; A cell nucleus category prediction module, configured to perform cell nucleus category prediction based on the feature decoding representation, the pathological image to be segmented, the multi-modal feature fusion representation at the first scale, and the category prediction module to obtain cell nucleus category prediction data, and obtain the cell nucleus segmentation result of the target multi-modal text and image information based on the cell nucleus category prediction data.

9. A computer device, characterized in that, Comprising: A processor, a memory, and a computer program stored on the memory and executable on the processor; when the computer program is executed by the processor, the steps of the cell nucleus segmentation method based on multi-modal information according to any one of claims 1 to 7 are implemented.

10. A storage medium, characterized in that: The storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the cell nucleus segmentation method based on multi-modal information according to any one of claims 1 to 7 are implemented.

Citation Information

Cited By

  • Text labeling method and device for cell image, electronic equipment and program product

    CN120510612A

  • Rectal cancer medical image focus segmentation method and device based on deep learning

    CN120599274A