Cell image segmentation method, device and electronic equipment

By introducing a multi-scale feature fusion module and a bottleneck residual module into the cell image segmentation model, combined with a hybrid attention mechanism, the problem of insufficient feature map fusion is solved, and the accuracy of cell image segmentation is improved.

CN116630302BActive Publication Date: 2025-11-11INDUSTRIAL AND COMMERCIAL BANK OF CHINA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310788437.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-06-29
Publication Date
2025-11-11
Estimated Expiration
2043-06-29

AI Technical Summary

Technical Problem

Existing cell image segmentation methods suffer from insufficient feature map fusion, resulting in poor model segmentation performance.

Method used

A combined structure of encoding module, multi-scale feature fusion module and decoding module is adopted. The multi-scale feature fusion module replaces the traditional convolutional layer to extract and fuse multi-scale information of the context. Combined with the bottleneck residual module and hybrid attention mechanism, the feature propagation capability and segmentation accuracy are improved.

Benefits of technology

It improves the accuracy of cell image segmentation, solves the problem of poor model segmentation results caused by insufficient feature map fusion, and achieves more accurate cell image segmentation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116630302B_ABST
    Figure CN116630302B_ABST
Patent Text Reader

Abstract

This application discloses a cell image segmentation method, apparatus, and electronic device. Relating to the field of image processing technology, the method includes: acquiring a cell image to be segmented; and using a pre-trained cell image segmentation model to segment the cell image to obtain a segmentation result. The cell image segmentation model includes: an encoding module, a multi-scale feature fusion module, and a decoding module. The output of the encoding module is connected to the input of the multi-scale feature fusion module, and the output of the multi-scale feature fusion module is connected to the input of the decoding module. The multi-scale feature fusion module is used to perform feature fusion processing on the input feature maps. This application solves the problem of insufficient feature map fusion in cell image segmentation in related technologies, which leads to poor model segmentation results.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image processing technology, and more specifically, to a cell image segmentation method, apparatus, and electronic device. Background Technology

[0002] Clinically, pathological slides are the gold standard for cancer diagnosis. Pathologists perform microscopic examination of these slides to complete pathological diagnoses and prognostic assessments, but this process is often time-consuming and labor-intensive. Deep learning-based segmentation methods overcome the drawbacks of manual segmentation, such as its time-consuming and labor-intensive nature, while also addressing the issues of poor accuracy and strong data dependence inherent in traditional automatic segmentation methods like threshold segmentation, region segmentation, and clustering segmentation.

[0003] The accuracy of semantic segmentation of cell images depends on the extraction and processing of image features. Traditional U-Net networks integrate high-level and low-level features in an inefficient stitching manner, resulting in the loss of effective image information, segmentation breaks at weak edges of the image, and limited network depth due to gradient vanishing.

[0004] There is currently no effective solution to the problem of insufficient feature map fusion in cell image segmentation, which leads to poor model segmentation results. Summary of the Invention

[0005] The main objective of this application is to provide a cell image segmentation method, apparatus, and electronic device to solve the problem of insufficient feature map fusion in cell image segmentation in related technologies, which leads to poor model segmentation results.

[0006] To achieve the above objectives, according to one aspect of this application, a cell image segmentation method is provided. The method includes: acquiring a cell image to be segmented; and performing segmentation processing on the cell image to be segmented using a pre-trained cell image segmentation model to obtain an image segmentation result of the cell image to be segmented. The cell image segmentation model includes: an encoding module, a multi-scale feature fusion module, and a decoding module. The output of the encoding module is connected to the input of the multi-scale feature fusion module, and the output of the multi-scale feature fusion module is connected to the input of the decoding module. The multi-scale feature fusion module is used to perform feature fusion processing on the input feature map.

[0007] To achieve the above objectives, according to another aspect of this application, a cell image segmentation apparatus is provided. The apparatus includes: acquiring a cell image to be segmented; and performing segmentation processing on the cell image to be segmented using a pre-trained cell image segmentation model to obtain an image segmentation result of the cell image to be segmented. The cell image segmentation model includes: an encoding module, a multi-scale feature fusion module, and a decoding module. The output of the encoding module is connected to the input of the multi-scale feature fusion module, and the output of the multi-scale feature fusion module is connected to the input of the decoding module. The multi-scale feature fusion module is used to perform feature fusion processing on the input feature map.

[0008] To achieve the above objectives, according to another aspect of this application, an electronic device is also provided, including one or more processors and a memory, wherein the memory is used to store one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors cause the one or more processors to implement any of the above-described cell image segmentation methods.

[0009] This application employs the following steps: acquiring a cell image to be segmented; and using a pre-trained cell image segmentation model to segment the cell image, obtaining the image segmentation result. The cell image segmentation model includes an encoding module, a multi-scale feature fusion module, and a decoding module. The output of the encoding module is connected to the input of the multi-scale feature fusion module, and the output of the multi-scale feature fusion module is connected to the input of the decoding module. The multi-scale feature fusion module performs feature fusion processing on the input feature map, achieving accurate cell image segmentation and solving the problem of insufficient feature map fusion in related technologies, which leads to poor model segmentation results. This improves the accuracy of cell image segmentation. Attached Figure Description

[0010] The accompanying drawings, which form part of this application, are used to provide a further understanding of this application. The illustrative embodiments and descriptions of this application are used to explain this application and do not constitute an undue limitation of this application. In the drawings:

[0011] Figure 1 This is a flowchart of a cell image segmentation method provided according to an embodiment of this application;

[0012] Figure 2 This is a schematic diagram of an optional cell image segmentation method according to an embodiment of this application;

[0013] Figure 3 This is a schematic diagram of an optional bottleneck residual module according to an embodiment of this application;

[0014] Figure 4 This is a schematic diagram of an optional hybrid attention mechanism module according to an embodiment of this application;

[0015] Figure 5 This is a schematic diagram of an optional multi-scale feature fusion module according to an embodiment of this application;

[0016] Figure 6 This is a schematic diagram of a cell image segmentation device provided according to an embodiment of this application;

[0017] Figure 7 This is a schematic diagram of an electronic device provided according to an embodiment of this application. Detailed Implementation

[0018] It should be noted that, unless otherwise specified, the embodiments and features described in this application can be combined with each other. This application will now be described in detail with reference to the accompanying drawings and embodiments.

[0019] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort should fall within the scope of protection of the present application.

[0020] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate for the embodiments of this application described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0021] It should be noted that all information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for display, data used for analysis, etc.) involved in this application are information and data authorized by the user or fully authorized by all parties. For example, if there is an interface between this system and the relevant user or organization, before obtaining the relevant information, it is necessary to send an acquisition request to the aforementioned user or organization through the interface, and obtain the relevant information after receiving the consent information from the aforementioned user or organization.

[0022] Clinically, pathological slides are the gold standard for cancer diagnosis. Pathologists perform microscopic examination of these slides to complete pathological diagnoses and prognostic assessments, but this process is often time-consuming and labor-intensive. With the digitization of pathological slides, artificial intelligence (AI) technology has entered the field of pathology, driving a shift in pathological analysis from qualitative to quantitative methods. This change makes pathological diagnoses more accurate and objective. In particular, AI technologies, represented by deep learning, have achieved remarkable results in pathological analysis, not only making pathological diagnosis more intelligent but also more precise and objective in terms of results.

[0023] Deep learning-based segmentation methods overcome the drawbacks of manual segmentation, such as being time-consuming and labor-intensive, while also addressing the issues of poor accuracy and strong data dependence found in traditional automatic segmentation methods like thresholding, region segmentation, and clustering. Among related technologies, some researchers have proposed fully convolutional networks (FCNs), using convolutional layers instead of fully connected layers in convolutional neural networks (CNNs) to achieve pixel-level classification from image-level classification, laying a crucial foundation for the subsequent application of deep learning in semantic segmentation. Other researchers have improved the FCN network structure, proposing the U-Net network, which combines high-level and shallow image information through skip connections in the decoding layer, preserving information at each level and improving the utilization of feature mappings. Still others have proposed combining segmentation networks with residual networks and introducing summation-based feature layers to achieve more accurate cell segmentation through a deeper network architecture. Other researchers have proposed the Pyramid Pooling Module Network (PSPNet), whose core contribution is the introduction of a global pyramid pooling module. This module can fuse contextual information at different scales, improving the ability to acquire global feature information and increasing the model's expressiveness. Recently, attention mechanisms have also been used in image segmentation to improve segmentation results. Some researchers have combined the Single Shot MultiBox Detector (SSD) model with the joint network U-Net model for instance segmentation of neurons, employing attention mechanisms in both the detection and segmentation modules to focus the model on task-relevant features.

[0024] The accuracy of semantic segmentation of cell images depends on the extraction and processing of image features. Traditional U-Net networks integrate high-level and low-level features in an inefficient stitching manner, resulting in the loss of effective image information, segmentation breaks at weak edges of the image, and limited network depth due to gradient vanishing.

[0025] Based on the above problems, this application will be described below with reference to the preferred implementation steps. Figure 1 This is a flowchart of a cell image segmentation method provided according to an embodiment of this application, such as... Figure 1 As shown, the method includes the following steps:

[0026] Step S101: Obtain the image of the cells to be segmented.

[0027] Optionally, the cell image to be segmented is acquired, and a multi-scale feature fusion module is used to replace the traditional convolutional layer between the encoding and decoding modules to extract multi-scale information of the context for fusion, thereby further optimizing the segmentation effect and improving the segmentation accuracy of the cell image.

[0028] Step S102: The cell image to be segmented is segmented using a pre-trained cell image segmentation model to obtain the image segmentation result of the cell image to be segmented. The cell image segmentation model includes an encoding module, a multi-scale feature fusion module, and a decoding module. The output of the encoding module is connected to the input of the multi-scale feature fusion module, and the output of the multi-scale feature fusion module is connected to the input of the decoding module. The multi-scale feature fusion module is used to perform feature fusion processing on the input feature map.

[0029] Optionally, Figure 2 This is a schematic diagram of an optional cell image segmentation method according to an embodiment of this application, such as... Figure 2 As shown, after the cell image to be segmented enters the cell image segmentation model, it is sequentially input into the encoding module, the multi-scale feature fusion module, and the decoding module to obtain the image segmentation result prediction map. The image segmentation result obtained by using this method is more accurate.

[0030] In one optional embodiment, the above-mentioned segmentation of the cell image to be segmented using a pre-trained cell image segmentation model to obtain the image segmentation result of the cell image to be segmented includes: encoding the cell image to be segmented using the above-mentioned encoding module to obtain a first feature map; performing feature fusion processing on the first feature map using the above-mentioned multi-scale feature fusion module to obtain a second feature map; and decoding the second feature map using the above-mentioned decoding module to obtain the image segmentation result.

[0031] By replacing the traditional convolutional layer with a multi-scale feature fusion module between the encoding and decoding modules, multi-scale contextual information is extracted and fused, further refining the segmentation effect. This cell image segmentation model structure improves upon the shortcomings of traditional models in extracting and utilizing image semantic features, resulting in more accurate image segmentation results.

[0032] In one optional embodiment, the above-mentioned encoding module includes: N bottleneck residual modules and N-1 downsampling modules, wherein any two bottleneck residual modules are connected through one of the N-1 downsampling modules, and N is an integer greater than or equal to 2.

[0033] Alternatively, as Figure 2 As shown, the encoding module includes four bottleneck residual modules and three downsampling modules. Any two bottleneck residual modules are connected by one downsampling module. The input image to be segmented sequentially passes through a bottleneck residual module – a downsampling module (i.e., a pooling layer) – a bottleneck residual module – a downsampling module – a bottleneck residual module – a downsampling module – a bottleneck residual module – a downsampling module – a bottleneck residual module, and finally, the output is the first feature map. The encoding module encodes the cell image to be segmented. The introduction of bottleneck residual modules enhances the feature propagation capability, enabling the extraction of more detailed cell information.

[0034] In one optional embodiment, the decoding module includes: N bottleneck residual modules, N-1 upsampling modules, and N hybrid attention mechanism modules. Any two bottleneck residual modules are connected through one of the N-1 upsampling modules. Each of the N bottleneck residual modules is connected to one of the N hybrid attention mechanism modules. There is a one-to-one correspondence between the N bottleneck residual modules and the N hybrid attention mechanism modules.

[0035] Alternatively, as Figure 2 As shown, the decoding module includes four bottleneck residual modules, three upsampling operation modules (i.e., deconvolution layers), and three hybrid attention mechanism modules. Any two bottleneck residual modules are connected via an upsampling module, and each bottleneck residual module is connected to a hybrid attention mechanism module. The decoding module is used to decode the cell image to be segmented. During the decoding process, the addition of the hybrid attention mechanism modules enables an effective combination of high-level and shallow feature maps.

[0036] In an optional embodiment, the aforementioned N bottleneck residual modules each include: a first convolutional unit, a second convolutional unit, and a third convolutional unit. The output of the first convolutional unit is connected to the input of the second convolutional unit, and the output of the second convolutional unit is connected to the input of the third convolutional unit. The first convolutional unit sequentially includes a 1×1 convolutional layer, a batch normalization layer, and a first activation function layer. The second convolutional unit sequentially includes a 3×3 convolutional layer, a batch normalization layer, and a first activation function layer. The third convolutional unit sequentially includes a 1×1 convolutional layer, a batch normalization layer, and a first activation function layer. The input of the first convolutional unit and the batch normalization layer in the third convolutional unit are connected by a 1×1 convolutional layer and a batch normalization layer.

[0037] Optionally, Figure 3 This is a schematic diagram of an optional bottleneck residual module according to an embodiment of this application, as shown below. Figure 3 As shown, the bottleneck residual module consists of three convolutional units: 1×1, 3×3, and 1×1. The 1×1 convolution can increase or decrease the dimensionality of the channel count, allowing the 3×3 convolution to be performed with a relatively low-dimensional input, thus improving computational efficiency. When performing identity mapping addition, a 1×1 convolution operation is applied to the input to ensure a consistent number of channels during superposition. Each convolutional unit consists of a convolutional layer, a batch normalization layer, and a first activation function layer. The first activation function layer is a ReLU activation function layer. The addition of the batch normalization layer addresses the issue of changes in the distribution of intermediate layer data during training, preventing gradient vanishing or exploding and accelerating training speed.

[0038] In an optional embodiment, the decoding module further includes a 1×1 convolutional layer and a second activation function layer, wherein the output of the last bottleneck residual module among the N bottleneck residual modules is connected to the 1×1 convolutional layer included in the encoding module, the 1×1 convolutional layer included in the encoding module is connected to the second activation function layer included in the encoding module, and the output of the second activation function layer included in the encoding module serves as the output of the encoding module.

[0039] Alternatively, as Figure 2 As shown, the decoding module also includes a 1×1 convolutional layer and a second activation function layer, wherein the second activation function layer is a softmax activation function layer. After the semantic information passes through 4 residual modules and 3 upsampling operations, it passes through a 1×1 convolutional layer and a softmax activation function layer, and finally outputs a segmentation prediction map with the same resolution as the input image.

[0040] In an optional embodiment, the attention mechanism module includes: a first average pooling layer, a first max pooling layer, a tunneling attention module, a third activation function layer, a second average pooling layer, a second max pooling layer, a spatial attention module, and a fourth activation function layer. The outputs of the first average pooling layer and the first max pooling layer are respectively connected to the input of the tunneling attention module. The output of the tunneling attention module is connected to the third activation function layer. The inputs of the second average pooling layer and the second max pooling layer are connected to each other. The inputs of the second average pooling layer and the second max pooling layer are respectively connected to the input of the spatial attention module. The output of the spatial attention module is connected to the fourth activation function layer.

[0041] Optionally, Figure 4 This is a schematic diagram of an optional hybrid attention mechanism module according to an embodiment of this application, such as... Figure 4 As shown, the input feature map is passed through a first average pooling layer and a first max pooling layer to obtain two different feature maps. These are then fed into the tunneling attention module, where the two feature maps are summed point-by-point, and then passed through a sigmoid activation function layer to obtain weight coefficients. The input feature map is then multiplied element-by-element by the weight coefficients to obtain a new input feature map. This process is repeated, passing the input feature map through a second average pooling layer and a second max pooling layer to obtain two different feature maps. These are then fed into the spatial attention module, where the two feature maps are summed point-by-point, and then passed through a sigmoid activation function layer to obtain weight coefficients. The input feature map is then multiplied element-by-element by the weight coefficients to obtain the output feature map. The addition of the hybrid attention mechanism increases the weight of the cell region, enhancing the model's ability to learn the foreground.

[0042] In an optional embodiment, the multi-scale feature fusion module includes: M pooling layers of different scales, M 1×1 convolutional layers, and an upsampling module, wherein the M pooling layers of different scales are each connected to one of the M 1×1 convolutional layers, and there is a one-to-one correspondence between the M pooling layers of different scales and the M 1×1 convolutional layers; the M 1×1 convolutional layers are each connected to the upsampling module, and M is an integer greater than or equal to 2.

[0043] Optionally, Figure 5 This is a schematic diagram of an optional multi-scale feature fusion module according to an embodiment of this application, as shown below. Figure 5As shown, the multi-scale feature fusion module includes four pooling layers at different scales, four 1×1 convolutional layers, and an upsampling module. After the first feature map is input into this module, it passes through the pooling layers, convolutional layers, and upsampling module sequentially to obtain the second feature map. The multi-scale feature fusion module is built based on the Pyramid Scene Parsing Network (PSPNet) model, utilizing pooling and feature fusion to improve the ability to acquire global feature information. The Pyramid Scene Parsing Network (PSPNet) is a deep learning model for scene parsing and semantic segmentation. By introducing a pyramid pooling module, it can extract features and understand scenes from input images at multiple different scales. The PSPNet model has achieved excellent results in image semantic segmentation tasks, especially in handling large-scale scenes and fine-grained objects.

[0044] In an optional embodiment, the above-mentioned multi-scale feature fusion module is used to perform feature fusion processing on the first feature map to obtain the second feature map, including: using M pooling layers of different scales to pool the first feature map respectively to obtain M third feature maps, wherein the scales of the M third feature maps are different; using M 1×1 convolutional layers to perform dimensionality reduction processing on the M third feature maps to obtain M fourth feature maps, wherein there is a one-to-one correspondence between the M 1×1 convolutional layers and the M third feature maps; using the above-mentioned upsampling module to perform upsampling processing on the M fourth feature maps respectively to obtain M fifth feature maps; and concatenating the M fifth feature maps and the first feature map to obtain the second feature map.

[0045] Optionally, the scale of the first feature map can be, but is not limited to, 30×30×320. The first feature map is passed through average pooling operation layers of different scales to obtain four third feature maps of different scales, namely 1×1×320, 2×2×320, 3×3×320 and 6×6×320. Next, a 1×1 convolutional layer is used to reduce the dimensionality of the four different feature maps to obtain four fourth feature maps, namely 1×1×80, 2×2×80, 3×3×80 and 6×6×80. Then, the four feature maps of different scales are upsampled to the size of the input first feature map to obtain four fifth feature maps of different scales, namely 30×30×80, 30×30×80, 30×30×80 and 30×30×80. Finally, the input first feature map and the four upsampled fifth feature maps are concatenated to obtain the second feature map. By employing a multi-scale feature fusion module, we can better extract and fuse multi-scale information from the context, thereby further refining the segmentation effect.

[0046] Through steps S101 to S102 described above, the goal of accurately segmenting cell images can be achieved, solving the problem of insufficient feature map fusion in cell image segmentation in related technologies, which leads to poor model segmentation results. This ultimately improves the accuracy of cell image segmentation.

[0047] Based on the above embodiments and optional embodiments, this application proposes an implementation method for an optional cell image segmentation method, which includes:

[0048] Step S1: Obtain the image of the cells to be segmented.

[0049] In step S2, the cell image to be segmented is input into the encoding module. After passing through 4 bottleneck residual modules and 3 downsampling modules, the encoded cell image, i.e., the first feature map, is obtained.

[0050] Step S3 involves inputting the first feature map into the multi-scale feature fusion module, which specifically includes the following sub-steps:

[0051] Step S31: After four pooling layers, four third feature maps are obtained.

[0052] Step S32: Use four 1×1 convolutional layers to reduce the dimensionality of the four third feature maps to obtain four fourth feature maps.

[0053] Step S33: Use an upsampling module to upsample the four fourth feature maps respectively to obtain four fifth feature maps.

[0054] Step S34: The four fifth feature maps and the first feature map are concatenated to obtain the second feature map.

[0055] In step S4, the second feature map is input into the decoding module. After passing through 4 attention mechanism modules, 4 bottleneck residual modules, and 3 upsampling modules, it passes through a 1×1 convolutional layer and a softmax activation function layer to obtain the image segmentation result prediction map.

[0056] The embodiments of this application can achieve at least the following technical effects: 1) Improves the problem of insufficient extraction and utilization of image semantic features by traditional U-Net networks, making the image segmentation results more accurate; 2) To address the degradation problem of deep networks, a bottleneck residual block module is introduced into U-Net to enhance the propagation ability of features and extract more cell detail information; 3) To address the problem of how the network can fully utilize the contextual information of the image, a multi-scale feature fusion module is used between the encoding and decoding modules to replace the traditional convolutional layer, extracting multi-scale contextual information for fusion, and further refining the segmentation effect; 4) To address the problem that the U-Net network cannot distinguish the effectiveness of information using a cascaded structure, a hybrid attention mechanism is added to the cascaded structure to increase the weight of cell regions, reduce the interference of brightness imbalance and low contrast on the model, and improve the robustness of the model.

[0057] It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowchart, in some cases the steps shown or described may be executed in a different order than that shown here.

[0058] This application also provides a cell image segmentation apparatus. It should be noted that the cell image segmentation apparatus of this application can be used to execute the cell image segmentation method provided in this application. The cell image segmentation apparatus provided in this application is described below.

[0059] Figure 6 This is a schematic diagram of a cell image segmentation apparatus according to an embodiment of this application. Figure 6 As shown, the device includes: a first acquisition module 601 and a first processing module 602, wherein,

[0060] The first acquisition module 601 described above is used to acquire an image of the cells to be segmented;

[0061] The first processing module 602, connected to the first acquisition module 601, is used to segment the cell image to be segmented using a pre-trained cell image segmentation model to obtain the image segmentation result of the cell image to be segmented. The cell image segmentation model includes an encoding module, a multi-scale feature fusion module, and a decoding module. The output of the encoding module is connected to the input of the multi-scale feature fusion module, and the output of the multi-scale feature fusion module is connected to the input of the decoding module. The multi-scale feature fusion module is used to perform feature fusion processing on the input feature map.

[0062] In this application, a first acquisition module 601 is set up to acquire a cell image to be segmented; a first processing module 602 is used to segment the cell image to be segmented using a pre-trained cell image segmentation model to obtain the image segmentation result of the cell image to be segmented. The cell image segmentation model includes an encoding module, a multi-scale feature fusion module, and a decoding module. The output of the encoding module is connected to the input of the multi-scale feature fusion module, and the output of the multi-scale feature fusion module is connected to the input of the decoding module. The multi-scale feature fusion module performs feature fusion processing on the input feature map, achieving accurate segmentation of the cell image and solving the problem of insufficient feature map fusion in related technologies, which leads to poor model segmentation results. This improves the accuracy of cell image segmentation.

[0063] In an optional embodiment, the first processing module includes: a first encoding submodule, used to encode the cell image to be segmented using the encoding module to obtain a first feature map; a first fusion submodule, used to perform feature fusion processing on the first feature map using the multi-scale feature fusion module to obtain a second feature map; and a first decoding submodule, used to decode the second feature map using the decoding module to obtain the image segmentation result.

[0064] In an optional embodiment, the first fusion submodule includes: a first pooling submodule, configured to perform pooling processing on the first feature map using M pooling layers of different scales respectively, to obtain M third feature maps, wherein the scales of the M third feature maps are not the same; a first dimensionality reduction submodule, configured to perform dimensionality reduction processing on the M third feature maps using M 1×1 convolutional layers, to obtain M fourth feature maps, wherein there is a one-to-one correspondence between the M 1×1 convolutional layers and the M third feature maps; a first upsampling submodule, configured to perform upsampling processing on the M fourth feature maps using one upsampling module respectively, to obtain M fifth feature maps; and a first concatenation submodule, configured to concatenate the M fifth feature maps and the first feature map to obtain the second feature map.

[0065] It should be noted that the above modules can be implemented by software or hardware. For example, for the latter, it can be implemented in the following ways: the above modules can be located in the same processor; or the above modules can be located in different processors in any combination.

[0066] It should be noted that the first acquisition module 601 and the first processing module 602 mentioned above correspond to steps S101 to S102 in the embodiments. The instances and application scenarios implemented by the above modules and the corresponding steps are the same, but are not limited to the content disclosed in the above embodiments. It should be noted that the above modules, as part of the device, can run in a computer terminal.

[0067] It should be noted that the optional or preferred implementation methods of this embodiment can be found in the relevant descriptions in the embodiments, and will not be repeated here.

[0068] The aforementioned cell image segmentation device includes a processor and a memory. All of the aforementioned units are stored in the memory as program units, and the processor executes the aforementioned program units stored in the memory to realize the corresponding functions.

[0069] The processor contains a kernel, which retrieves the corresponding program units from memory. One or more kernels can be configured, and their parameters can be adjusted (for the purposes of this application).

[0070] The memory may include non-permanent memory in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM, and the memory includes at least one memory chip.

[0071] This application provides a non-volatile storage medium storing a program that, when executed by a processor, implements the cell image segmentation method described above.

[0072] This application provides a processor for running a program, wherein the program executes the cell image segmentation method described above.

[0073] like Figure 7 As shown in the illustration, this application provides an electronic device 10, which includes a processor, a memory, and a program stored in the memory and executable on the processor. When the processor executes the program, it performs the following steps: acquiring a cell image to be segmented; and segmenting the cell image to be segmented using a pre-trained cell image segmentation model to obtain an image segmentation result. The cell image segmentation model includes an encoding module, a multi-scale feature fusion module, and a decoding module. The output of the encoding module is connected to the input of the multi-scale feature fusion module, and the output of the multi-scale feature fusion module is connected to the input of the decoding module. The multi-scale feature fusion module is used to perform feature fusion processing on the input feature map. The device described herein can be a server, PC, PAD, mobile phone, etc.

[0074] This application also provides a computer program product, which, when executed on a data processing device, is suitable for executing a program with the following initialization steps: acquiring a cell image to be segmented; performing segmentation processing on the cell image to be segmented using a pre-trained cell image segmentation model to obtain an image segmentation result of the cell image to be segmented, wherein the cell image segmentation model includes: an encoding module, a multi-scale feature fusion module, and a decoding module, wherein the output end of the encoding module is connected to the input end of the multi-scale feature fusion module, and the output end of the multi-scale feature fusion module is connected to the input end of the decoding module, wherein the multi-scale feature fusion module is used to perform feature fusion processing on the input feature map.

[0075] Optionally, the above-mentioned computer program product is also suitable for executing an initialization program having the following method steps: using the above-mentioned encoding module to encode the above-mentioned cell image to be segmented to obtain a first feature map; using the above-mentioned multi-scale feature fusion module to perform feature fusion processing on the above-mentioned first feature map to obtain a second feature map; using the above-mentioned decoding module to decode the above-mentioned second feature map to obtain the above-mentioned image segmentation result.

[0076] Optionally, the aforementioned computer program product is also suitable for executing an initialization program with the following steps: pooling the first feature map using M pooling layers of different scales to obtain M third feature maps, wherein the scales of the M third feature maps are different; performing dimensionality reduction processing on the M third feature maps using M 1×1 convolutional layers to obtain M fourth feature maps, wherein there is a one-to-one correspondence between the M 1×1 convolutional layers and the M third feature maps; performing upsampling processing on the M fourth feature maps using the aforementioned upsampling module to obtain M fifth feature maps; and concatenating the M fifth feature maps and the first feature map to obtain the second feature map.

[0077] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0078] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0079] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0080] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0081] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.

[0082] Memory may include non-persistent memory in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.

[0083] Computer-readable media includes both permanent and non-permanent, removable and non-removable media that can store information using any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic magnetic disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.

[0084] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.

[0085] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0086] The above are merely embodiments of this application and are not intended to limit the scope of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of the claims of this application.

Claims

1. A cell image segmentation method, characterized in that, include: Acquire the image of the cells to be segmented; A pre-trained cell image segmentation model is used to segment the cell image to be segmented, resulting in an image segmentation result. The cell image segmentation model includes an encoding module, a multi-scale feature fusion module, and a decoding module. The output of the encoding module is connected to the input of the multi-scale feature fusion module, and the output of the multi-scale feature fusion module is connected to the input of the decoding module. The multi-scale feature fusion module is used to perform feature fusion processing on the input feature map. The decoding module includes: N bottleneck residual modules, N-1 upsampling modules, and N hybrid attention mechanism modules. Any two bottleneck residual modules are connected through one of the N-1 upsampling modules. Each of the N bottleneck residual modules is connected to one of the N hybrid attention mechanism modules. There is a one-to-one correspondence between the N bottleneck residual modules and the N hybrid attention mechanism modules. The hybrid attention mechanism module includes: a first average pooling layer, a first max pooling layer, a tunneling attention module, a third activation function layer, a second average pooling layer, a second max pooling layer, a spatial attention module, and a fourth activation function layer. The outputs of the first average pooling layer and the first max pooling layer are respectively connected to the input of the tunneling attention module. The output of the tunneling attention module is connected to the third activation function layer. The inputs of the second average pooling layer and the second max pooling layer are connected to each other. The inputs of the second average pooling layer and the second max pooling layer are respectively connected to the input of the spatial attention module. The output of the spatial attention module is connected to the fourth activation function layer.

2. The method according to claim 1, characterized in that, The step of segmenting the cell image to be segmented using a pre-trained cell image segmentation model to obtain the image segmentation result of the cell image to be segmented includes: The encoding module is used to encode the cell image to be segmented to obtain a first feature map; The first feature map is processed by the multi-scale feature fusion module to obtain the second feature map; The second feature map is decoded using the decoding module to obtain the image segmentation result.

3. The method according to claim 2, characterized in that, The encoding module includes: N bottleneck residual modules and N-1 downsampling modules. Any two bottleneck residual modules are connected through one of the N-1 downsampling modules, where N is an integer greater than or equal to 2.

4. The method according to claim 3, characterized in that, The N bottleneck residual modules each include: a first convolutional unit, a second convolutional unit, and a third convolutional unit. The output of the first convolutional unit is connected to the input of the second convolutional unit, and the output of the second convolutional unit is connected to the input of the three convolutional units. The first convolutional unit sequentially includes a 1×1 convolutional layer, a batch normalization layer, and a first activation function layer; the second convolutional unit sequentially includes a 3×3 convolutional layer, a batch normalization layer, and a first activation function layer; and the third convolutional unit sequentially includes a 1×1 convolutional layer, a batch normalization layer, and a first activation function layer. The input of the first convolutional unit and the batch normalization layer in the third convolutional unit are connected through a 1×1 convolutional layer and a batch normalization layer.

5. The method according to claim 3, characterized in that, The decoding module further includes: a 1×1 convolutional layer and a second activation function layer, wherein the output of the last bottleneck residual module among the N bottleneck residual modules is connected to the 1×1 convolutional layer included in the encoding module, the 1×1 convolutional layer included in the encoding module is connected to the second activation function layer included in the encoding module, and the output of the second activation function layer included in the encoding module serves as the output of the encoding module.

6. The method according to claim 2, characterized in that, The multi-scale feature fusion module includes: M pooling layers of different scales, M 1×1 convolutional layers, and an upsampling module. Each of the M pooling layers of different scales is connected to one of the M 1×1 convolutional layers, and there is a one-to-one correspondence between the M pooling layers of different scales and the M 1×1 convolutional layers. Each of the M 1×1 convolutional layers is connected to the upsampling module, where M is an integer greater than or equal to 2.

7. The method according to claim 6, characterized in that, The step of performing feature fusion processing on the first feature map using the multi-scale feature fusion module to obtain the second feature map includes: The first feature map is pooled using M pooling layers of different scales to obtain M third feature maps, wherein the scales of the M third feature maps are different. The M 1×1 convolutional layers are used to reduce the dimensionality of the M third feature maps to obtain M fourth feature maps, wherein there is a one-to-one correspondence between the M 1×1 convolutional layers and the M third feature maps; The M fourth feature maps are upsampled using the aforementioned upsampling module to obtain M fifth feature maps; The M fifth feature maps and the first feature map are concatenated to obtain the second feature map.

8. A cell image segmentation device, characterized in that, include: The first acquisition module is used to acquire images of cells to be segmented; The first processing module is used to segment the cell image to be segmented using a pre-trained cell image segmentation model to obtain the image segmentation result of the cell image to be segmented. The cell image segmentation model includes an encoding module, a multi-scale feature fusion module, and a decoding module. The output of the encoding module is connected to the input of the multi-scale feature fusion module, and the output of the multi-scale feature fusion module is connected to the input of the decoding module. The multi-scale feature fusion module is used to perform feature fusion processing on the input feature map. The decoding module includes: N bottleneck residual modules, N-1 upsampling modules, and N hybrid attention mechanism modules. Any two bottleneck residual modules are connected through one of the N-1 upsampling modules. Each of the N bottleneck residual modules is connected to one of the N hybrid attention mechanism modules. There is a one-to-one correspondence between the N bottleneck residual modules and the N hybrid attention mechanism modules. The hybrid attention mechanism module includes: a first average pooling layer, a first max pooling layer, a tunneling attention module, a third activation function layer, a second average pooling layer, a second max pooling layer, a spatial attention module, and a fourth activation function layer. The outputs of the first average pooling layer and the first max pooling layer are respectively connected to the input of the tunneling attention module. The output of the tunneling attention module is connected to the third activation function layer. The inputs of the second average pooling layer and the second max pooling layer are connected to each other. The inputs of the second average pooling layer and the second max pooling layer are respectively connected to the input of the spatial attention module. The output of the spatial attention module is connected to the fourth activation function layer.

9. An electronic device, characterized in that, It includes one or more processors and a memory, the memory being used to store one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors cause the one or more processors to implement the cell image segmentation method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Gland cell image segmentation method and device based on edge sensing network

    CN113034505A

  • Pulmonary nodule image detection method based on deep learning

    CN116188404A