Image segmentation method and device
By adopting Twins Transformer block and multi-scale feature iterative fusion module in medical image segmentation model, combining self-attention and convolution components, the problem that the existing model has not fully utilized its advantages when combining long-term dependencies and convolutional representations is achieved, and image segmentation effect with high precision and strong generalization capabilities is achieved.
Patent Information
- Application Number
- CN202210053264.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-01-18
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2042-01-18
AI Technical Summary
The existing medical image segmentation model fails to fully utilize the advantages of Transformer when combining convolution and self-attention mechanisms, resulting in the inadequate combination of long-term dependencies and convolution representations.
Using the image segmentation method based on the Twins Transformer block and multi-scale feature iterative fusion (MS-FIF) module, a powerful segmentation model is built by iteratively fusing features of different scales, combining self-attention and convolution components.
It achieves good segmentation accuracy and strong generalization capabilities on multi-organ and cardiac segmentation datasets, giving full play to the advantages of Transformer and convolutional components.
Smart Images

Figure CN114419319B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of image processing technologies, and in particular, to an image segmentation method and an image segmentation device. Background Art
[0002] As the core technology of the vast majority of current computer vision systems, the convolutional neural network (CNN) has made great contributions in the field of image classification. Recently, Transformers for processing natural language have also received extensive attention in vision-based research applications. The core idea behind this is to apply the self-attention mechanism to capture long-term dependencies. Compared with CNNs (i.e., convolutional neural networks), Transformers alleviate the inductive bias of locality, making them more capable of handling non-local interactions. Research has also found that the prediction error of Transformers is closer to the human prediction error than that of CNNs.
[0003] Given the advantages of Transformers and CNNs in the field of image processing, there are many methods for building medical image segmentation models based on convolutional neural networks or Transformers. Chen J first proposed TransUNet to explore the potential of Transformers in medical image segmentation. The overall architecture of TransUNet is similar to U-Net, in which the convolutional neural network acts as a feature extractor, while the transformer helps to encode the global context. However, TransUNet and most of its followers only regard the convolutional neural network as the main body and further apply Transformers on top of the main body to capture long-term dependencies. Since convolutional representations usually contain precise spatial information and provide hierarchical concepts, one or two layers of Transformers are not sufficient to combine long-term dependencies with convolutional representations, resulting in the advantages of Transformers not being fully utilized. To solve the above problems, some researchers have started using Transformers as the backbone of the segmentation model. Swin-UNET is based on the Swin Transformer and uses hierarchical transformer blocks to build an encoder and a decoder in a U-Net-like architecture, achieving an improvement over TransUNet, but without exploring how to properly combine convolution and self-attention to build an optimal medical segmentation network. Summary of the Invention
[0004] The present disclosure provides a method and a device for segmenting an image.
[0005] According to a first aspect of the present disclosure, an image segmentation method is provided, including:
[0006] S1: Import the original image to be segmented into an encoder, and obtain N first prediction results after encoding;
[0007] S2: Import N - 1 first prediction results into N - 1 multi-scale feature iterative fusion networks one by one respectively;
[0008] S3: Import the remaining one of the N first prediction results into the (N - 1)-th multi-scale feature iterative fusion network;
[0009] S4: Import the output of the (N - 1)-th multi-scale feature iterative fusion network into the (N - 2)-th multi-scale feature iterative fusion network;
[0010] S5: Repeat step S4 until the output of the second multi-scale feature iterative fusion network is imported into the first multi-scale feature iterative fusion network;
[0011] S6: Import the outputs of the N - 1 multi-scale feature iterative fusion networks and the N-th first prediction result into a decoder respectively to obtain a second prediction result after segmentation.
[0012] Preferably, the S1 includes:
[0013] S101: Import the original image into the first encoder among the N encoders;
[0014] S102: Output the first first prediction result output by the first encoder to the second encoder and the first multi-scale feature iterative fusion network respectively;
[0015] S103: Repeat step S102 in sequence until the (N - 1)-th first prediction result output by the (N - 1)-th encoder is output to the N-th encoder and the (N - 1)-th multi-scale feature iterative fusion network respectively;
[0016] S104: Import the output of the N-th encoder into the decoder.
[0017] Preferably, the S2 to S5 include:
[0018] S201: After downsampling the first first prediction result output by the first encoder, output it to the first multi-channel attention module;
[0019] S202: Output the output of the second multi-scale feature iterative fusion network corresponding to the second encoder to the second multi-channel attention module;
[0020] S203: After connecting the output of the first multi-channel attention module and the output of the second multi-channel attention module, output them to the third multi-channel attention module;
[0021] S204: Use the output of the third multi-channel attention module as the output of the first multi-scale feature iterative fusion network;
[0022] S205: Repeat steps S201 - S204 until the respective outputs of the N - 1 multi-scale feature iterative fusion networks are obtained.
[0023] Preferably, the S6 includes:
[0024] Import the outputs of the N - 1 multi-scale feature iterative fusion networks and the Nth first prediction result into N decoders one by one to obtain N second prediction results after segmentation.
[0025] Preferably, it further includes:
[0026] S7: Perform cascaded feature extraction on the Nth first prediction result to the first first prediction result to obtain a more accurate second prediction result after segmentation.
[0027] According to the second aspect of the present disclosure, there is also provided an image segmentation device, including:
[0028] Encoding module: used to import the original image to be segmented into the encoder, and obtain N first prediction results after encoding;
[0029] First multi-scale feature iterative fusion module: used to import N - 1 first prediction results into N - 1 multi-scale feature iterative fusion networks one by one respectively;
[0030] Second multi-scale feature iterative fusion module: used to import the remaining one first prediction result among the N first prediction results into the (N - 1)th multi-scale feature iterative fusion network;
[0031] Third multi-scale feature iterative fusion module: import the output of the (N - 1)th multi-scale feature iterative fusion network into the (N - 2)th multi-scale feature iterative fusion network;
[0032] First repetition module: used to repeatedly execute the third multi-scale feature iterative fusion module until the output of the second multi-scale feature iterative fusion network is imported into the first multi-scale feature iterative fusion network;
[0033] Decoding module: used to import the outputs of the N - 1 multi-scale feature iterative fusion networks and the Nth first prediction result into the decoder respectively to obtain the second prediction result after segmentation.
[0034] Preferably, the encoding module includes:
[0035] Import module: used to import the original image into the first encoder among N encoders;
[0036] First output module: used to output the first first prediction result output by the first encoder to the second encoder and the first multi-scale feature iterative fusion network respectively;
[0037] Second repetition module: used to sequentially repeat the execution of the first output module until the (N - 1)-th first prediction result output by the (N - 1)-th encoder is output to the N-th encoder and the (N - 1)-th multi-scale feature iterative fusion network respectively:
[0038] Second output module: used to import the output of the N-th encoder into the decoder.
[0039] Preferably, it includes:
[0040] Downsampling module: used to perform downsampling on the first first prediction result output by the first encoder and then output it to the first multi-channel attention module;
[0041] Third output module: output the output of the second multi-scale feature iterative fusion network corresponding to the second encoder to the second multi-channel attention module;
[0042] Connection module: used to connect the output of the first multi-channel attention module and the output of the second multi-channel attention module and then output it to the third multi-channel attention module;
[0043] Fourth output module: use the output of the third multi-channel attention module as the output of the first multi-scale feature iterative fusion network;
[0044] Third repetition module: repeat the execution of the downsampling module, the third output module, the connection module and the fourth output module until the respective outputs of the (N - 1) multi-scale feature iterative fusion networks are obtained.
[0045] Preferably, the decoding module includes:
[0046] Import the outputs of the (N - 1) multi-scale feature iterative fusion networks and the N-th first prediction result into N decoders one by one to obtain N segmented second prediction results.
[0047] Preferably, it further includes:
[0048] Cascade module: used to perform cascade feature extraction on the N-th first prediction result to the first first prediction result to obtain a more accurate segmented second prediction result.
[0049] Beneficial effects:
[0050] (1) Based on the Twins Transformer block and the previously proposed Multi-Scale Feature Iterative Fusion (MS-FIF) module, meaningful information is obtained by iteratively fusing features of different scales.
[0051] (2) A cascaded feature extraction structure is constructed to predict the objects that were not correctly predicted at the previous level layer by layer from high to low, thereby obtaining a more refined target.
[0052] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understandable through the following description. Brief Description of the Drawings
[0053] The drawings are used to better understand the solution and do not constitute a limitation to the present disclosure. Among them:
[0054] Figure 1 is a schematic diagram of the first embodiment of the image segmentation method according to the present disclosure;
[0055] Figure 2 is a schematic diagram of importing the original image to be segmented into the encoder and obtaining N first prediction results after encoding according to the present disclosure;
[0056] Figure 3 is a schematic diagram of the second embodiment of the image segmentation method according to the present disclosure;
[0057] Figure 4 is a schematic diagram of the third embodiment of the image segmentation method according to the present disclosure;
[0058] Figure 5 is a schematic diagram of the first embodiment of the image segmentation device;
[0059] Figure 6 is a schematic diagram of the encoding module;
[0060] Figure 7 is a schematic diagram of the multi-scale feature iterative fusion structure;
[0061] Figure 8 is a schematic diagram of the second embodiment of the image segmentation device;
[0062] Figure 9 is an example of the implementation of the image segmentation method;
[0063] Figure 10 is a block diagram of the electronic device for implementing the image segmentation method of the embodiments of the present disclosure.
[0064] Explanation of the reference numerals in the drawings:
[0065] 5 Image segmentation device
[0066] 501 Encoding module 502 First multi-scale feature iterative fusion module
[0067] 503 Second multi-scale feature iterative fusion module
[0068] 504 Third multi-scale feature iterative fusion module
[0069] 505 First repetition module 506 Decoding module
[0070] 507 Cascade module
[0071] 5011 Import module 5012 First output module
[0072] 5013 Second repetition module 5014 Second output module
[0073] 701 Downsampling module 702 Third output module
[0074] 703 Connection module 704 Fourth output module
[0075] 705 Third repetition module Detailed implementation manners
[0076] The following describes exemplary embodiments of the present disclosure with reference to the accompanying drawings. Various details of the embodiments of the present disclosure are included to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.
[0077] Term explanation:
[0078] Multi-scale feature iterative fusion network: It refers to a network that iteratively fuses multiple feature maps to obtain a multi-scale feature attention map; in the present disclosure, the multi-scale feature iterative fusion network is composed of multiple multi-channel attention modules and a downsampling module; there is a detailed description in the patent application publication number: CN113076960A.
[0079] Multi-channel attention module: A certain input image is divided into two channels for feature extraction, and two branches with different scales are used to extract channel attention weights. One branch extracts the spatial attention of local features, and the other branch uses GAP (Global Avg Pooling) to average and extract the channel attention of global features.
[0080] Downsampling module: It can also be called a subsampling module; its purpose is to make the scales of the Kth feature map and the (K + 1)th feature map input into the multi-scale feature iterative fusion network the same. For example, if the scale of the Kth feature map is (H, W) and the scale of the (K + 1)th feature map is (H / 2, W / 2); after the Kth feature map passes through the downsampling module, the scale of the obtained feature map is (H / 2, W / 2).
[0081] Transformer: It is a method in natural language processing, which is expected to help non-typical convolutional neural networks (convolutional networks) overcome their inherent disadvantages of spatial induction bias; and, it relaxes the induction bias of locality, making it more capable of handling non-local interactions.
[0082] The present invention proposes a powerful segmentation model based on the combination of self-attention and convolution - nnTwins (not-another Twins), which combines convolution and self-attention to give full play to their advantages. A large number of experiments on multi-organ and heart segmentation datasets show that this method has good segmentation accuracy and strong generalization ability.
[0083] As Figure 1 、 Figure 9 shown, according to the first aspect of the present disclosure, there is provided an image segmentation method, including:
[0084] S1: Import the original image to be segmented into the encoder, and obtain N first prediction results after encoding; the encoder in the present disclosure has the same function as a conventional encoder, that is, it encodes the input image to compress the image and reduce its capacity. However, the structure of the encoder in the present disclosure is different, and the output results are also different. In this embodiment, the original image includes medical images and other non-medical images.
[0085] S2: Import N - 1 first prediction results into N - 1 multi-scale feature iterative fusion networks respectively in one-to-one correspondence;
[0086] S3: Import the remaining one first prediction result among the N first prediction results into the (N - 1)th multi-scale feature iterative fusion network;
[0087] S4: Import the output of the (N - 1)-th multi-scale feature iterative fusion network into the (N - 2)-th multi-scale feature iterative fusion network;
[0088] S5: Repeat step S4 until the output of the second multi-scale feature iterative fusion network is imported into the first multi-scale feature iterative fusion network;
[0089] S6: Import the output of the (N - 1) multi-scale feature iterative fusion networks and the N-th first prediction result into the decoder respectively to obtain the segmented second prediction result. The structure of the Transformer block in the decoder is highly symmetric with that of the encoder.
[0090] As Figure 2 shown, preferably, S1 includes:
[0091] S101: Import the original image into the first encoder among the N encoders; each of the N encoders includes: two Transformer encoders, a patch embed, and a positional encoding generator; the positional encoding generator is sandwiched between the two Transformer encoders, and the patch embed is arranged outside the two Transformer encoders; the function of the patch embed is similar to that of the downsampling module, but in addition to downsampling the image, it also includes increasing the number of channels of the image, that is, the input feature (H / 4×W / 4×64) becomes (H / 8×W / 8×128) after passing through the patch embed.
[0092] S102: Output the first first prediction result output by the first encoder to the second encoder and the first multi-scale feature iterative fusion network respectively;
[0093] S103: Repeat step S102 in sequence until the (N - 1)-th first prediction result output by the (N - 1)-th encoder is output to the N-th encoder and the (N - 1)-th multi-scale feature iterative fusion network respectively;
[0094] S104: Import the output of the N-th encoder into the decoder.
[0095] As Figure 3 shown, preferably, S2 to S5 include:
[0096] S201: After downsampling the first first prediction result output by the first encoder, output it to the first multi-channel attention module;
[0097] S202: Output the output of the second multi-scale feature iterative fusion network corresponding to the second encoder to the second multi-channel attention module.
[0098] S203: Connect the output of the first multi-channel attention module and the output of the second multi-channel attention module, and then output the result to the third multi-channel attention module.
[0099] S204: Use the output of the third multi-channel attention module as the output of the first multi-scale feature iterative fusion network.
[0100] S205: Repeat steps S201 - S204 until the respective outputs of the N - 1 multi-scale feature iterative fusion networks are obtained.
[0101] Preferably, S6 includes:
[0102] Import the outputs of the N - 1 multi-scale feature iterative fusion networks and the Nth first prediction result into N decoders one by one to obtain N second prediction results after segmentation.
[0103] As Figure 4 shown, preferably, it further includes:
[0104] S7: Perform cascaded feature extraction on the Nth first prediction result to the first first prediction result to obtain a more accurate second prediction result after segmentation.
[0105] As Figure 5 shown, according to the second aspect of the present disclosure, there is also provided an image segmentation device 5, including:
[0106] Encoding module 501: configured to import the original image to be segmented into an encoder, and obtain N first prediction results after encoding;
[0107] First multi-scale feature iterative fusion module 502: configured to import N - 1 first prediction results into N - 1 multi-scale feature iterative fusion networks one by one respectively;
[0108] Second multi-scale feature iterative fusion module 503: configured to import the remaining one first prediction result among the N first prediction results into the (N - 1)th multi-scale feature iterative fusion network;
[0109] Third multi-scale feature iterative fusion module 504: import the output of the (N - 1)th multi-scale feature iterative fusion network into the (N - 2)th multi-scale feature iterative fusion network;
[0110] First repetition module 505: Used to repeatedly execute the third multi-scale feature iterative fusion module until the output of the second multi-scale feature iterative fusion network is imported into the first multi-scale feature iterative fusion network;
[0111] Decoding module 506: Used to import the outputs of the N - 1 multi-scale feature iterative fusion networks and the Nth first prediction result into the decoder respectively to obtain the segmented second prediction result.
[0112] As Figure 6 shown, preferably, the encoding module 501 includes:
[0113] Import module 5011: Used to import the original image into the first encoder among the N encoders;
[0114] First output module 5012: Used to output the first first prediction result output by the first encoder to the second encoder and the first multi-scale feature iterative fusion network respectively;
[0115] Second repetition module 5013: Used to repeatedly execute the first output module in sequence until the (N - 1)th first prediction result output by the (N - 1)th encoder is output to the Nth encoder and the (N - 1)th multi-scale feature iterative fusion network respectively;
[0116] Second output module 5014: Used to import the output of the Nth encoder into the decoder.
[0117] As Figure 7 shown, preferably, the multi-scale feature iterative fusion structure 7 includes:
[0118] Downsampling module 701: Used to downsample the first first prediction result output by the first encoder and then output it to the first multi-channel attention module;
[0119] Third output module 702: Output the output of the second multi-scale feature iterative fusion network corresponding to the second encoder to the second multi-channel attention module;
[0120] Connection module 703: Used to connect the output of the first multi-channel attention module and the output of the second multi-channel attention module and then output it to the third multi-channel attention module;
[0121] Fourth output module 704: Use the output of the third multi-channel attention module as the output of the first multi-scale feature iterative fusion network;
[0122] Third repetition module 705: Repeatedly execute the downsampling module, the third output module, the connection module, and the fourth output module until the respective outputs of the N-1 multi-scale feature iterative fusion networks are obtained.
[0123] Preferably, the decoding module includes:
[0124] Import the outputs of the N-1 multi-scale feature iterative fusion networks and the Nth first prediction result into N decoders one by one to obtain N segmented second prediction results.
[0125] As Figure 8 shown, preferably, it further includes:
[0126] Cascade module 507: Used to perform cascaded feature extraction on the Nth first prediction result to the first first prediction result to obtain a more accurate segmented second prediction result.
[0127] As Figure 9 shown, the input input is the original image, which is input into 4 encoders (twins encoder). The outputs of the upper 3 encoders among the 4 encoders are respectively imported into the multi-scale feature iterative fusion network (MS-FIF). Moreover, the outputs of each of the upper 3 encoders in sequence are also imported into the next encoder in turn until the output of the third encoder is imported into the fourth encoder; the outputs f1, f2, f3 of the upper three encoders are imported into the multi-scale feature iterative fusion network, and the output f4 of the fourth encoder is directly imported into the decoder (twins decoder). The outputs of the 3 multi-scale feature iterative fusion networks are also imported into 3 decoders. The respective inputs of the 3 multi-scale feature iterative fusion networks include f1, f2, f3, f4, and also include the outputs of the second multi-scale feature iterative fusion network and the third multi-scale feature iterative fusion network. The outputs of the decoders include F1, F2, F3, F4. Perform cascaded feature extraction on F1, F2, F3, F4 to obtain a more accurate segmented second prediction result. The output prediction is the second prediction result; the segmented second prediction result is the final prediction result of the model..
[0128] In the technical solution of the present disclosure, the acquisition, storage, and application of the user's personal information involved all comply with the provisions of relevant laws and regulations and do not violate public order and good customs.
[0129] According to the embodiments of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0130] Figure 10FIG. shows a schematic block diagram of an exemplary electronic device 1000 that can be used to implement embodiments of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, personal digital processors, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely exemplary and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0131] As Figure 10 shown, the device 1000 includes a computing unit 1001 that can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 1002 or a computer program loaded from a storage unit 1008 into a random access memory (RAM) 1003. In the RAM 1003, various programs and data required for the operation of the device 1000 can also be stored. The computing unit 1001, the ROM 1002, and the RAM 1003 are connected to each other via a bus 1004. An input / output (I / O) interface 1005 is also connected to the bus 1004.
[0132] A plurality of components in the device 1000 are connected to the I / O interface 1005, including: an input unit 1006, such as a keyboard, a mouse, etc.; an output unit 1007, such as various types of displays, speakers, etc.; a storage unit 1008, such as a magnetic disk, an optical disk, etc.; and a communication unit 1009, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 1009 allows the device 1000 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0133] The computing unit 1001 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 1001 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 1001 executes the various methods and processes described above.
[0134] The various embodiments of the systems and techniques described above in this specification can be implemented in digital electronic circuitry, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems-on-chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which may be a special-purpose or general-purpose programmable processor that receives data and instructions from, and transmits data and instructions to, a storage system, at least one input device, and at least one output device.
[0135] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general purpose computer, special purpose computer, or other programmable data processing apparatus, such that the program codes, when executed by the processor or controller, cause the functions / operations specified in the flowchart and / or block diagram to be implemented. The program code may execute entirely on the machine, partly on the machine, as a stand-alone software package partly on the machine and partly on a remote machine, or entirely on the remote machine or server.
[0136] In the context of the present disclosure, a machine-readable medium may be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0137] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and a pointing device (e.g., a mouse or a trackball) through which the user can provide input to the computer. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, speech input, or tactile input).
[0138] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected to each other by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), and the Internet.
[0139] A computer system can include a client and a server. The client and the server are generally remote from each other and typically interact through a communication network. The relationship between the client and the server is generated by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, or a server of a distributed system, or a server incorporating a blockchain.
[0140] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps recited in this disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and this is not limited herein.
[0141] The above specific embodiments do not constitute a limitation on the protection scope of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the protection scope of this disclosure.
Claims
1. An image segmentation method, characterized in that, Including: S1: Import the original image to be segmented into the encoder, and obtain N first prediction results after encoding; S2: Import N - 1 first prediction results into N - 1 multi-scale feature iterative fusion networks respectively in one-to-one correspondence; S3: Import the remaining one first prediction result among the N first prediction results into the (N - 1)-th multi-scale feature iterative fusion network; S4: Import the output of the (N - 1)-th multi-scale feature iterative fusion network into the (N - 2)-th multi-scale feature iterative fusion network; S5: Repeat step S4 until the output of the second multi-scale feature iterative fusion network is imported into the first multi-scale feature iterative fusion network; S6: Import the outputs of the N - 1 multi-scale feature iterative fusion networks and the N-th first prediction result into the decoder respectively to obtain the second prediction results after segmentation; S7: Perform cascaded feature extraction on the N-th first prediction result to the first first prediction result to obtain more accurate second prediction results after segmentation.
2. The method according to claim 1, characterized in that The S1 includes: S101: Import the original image into the first encoder among the N encoders; S102: Output the first first prediction result output by the first encoder to the second encoder and the first multi-scale feature iterative fusion network respectively; S103: Repeat step S102 in sequence until the (N - 1)-th first prediction result output by the (N - 1)-th encoder is output to the N-th encoder and the (N - 1)-th multi-scale feature iterative fusion network respectively; S104: Import the output of the N-th encoder into the decoder.
3. The method according to claim 2, characterized in that, The S2 - S5 include: S201: Downsample the first first prediction result output by the first encoder and then output it to the first multi-channel attention module; S202: Output the output of the second multi-scale feature iterative fusion network corresponding to the second encoder to the second multi-channel attention module; S203: Connect the output of the first multi-channel attention module and the output of the second multi-channel attention module, and then output it to the third multi-channel attention module; S204: Use the output of the third multi-channel attention module as the output of the first multi-scale feature iterative fusion network; S205: Repeat steps S201 - S204 until the respective outputs of the N - 1 multi-scale feature iterative fusion networks are obtained.
4. The method according to claim 3, characterized in that, The S6 includes: Import the outputs of the N - 1 multi-scale feature iterative fusion networks and the N-th first prediction result into N decoders respectively in one-to-one correspondence to obtain N second prediction results after segmentation.
5. An image segmentation device, characterized in that, Including: Encoding module: Used to import the original image to be segmented into the encoder, and obtain N first prediction results after encoding; First multi-scale feature iterative fusion module: Used to import N - 1 first prediction results into N - 1 multi-scale feature iterative fusion networks respectively in one-to-one correspondence; Second multi-scale feature iterative fusion module: Used to import the remaining one first prediction result among the N first prediction results into the (N - 1)-th multi-scale feature iterative fusion network; The third multi-scale feature iterative fusion module: Import the output of the (N-1)th multi-scale feature iterative fusion network into the (N-2)th multi-scale feature iterative fusion network; The first repetition module: Used to repeatedly execute the third multi-scale feature iterative fusion module until the output of the second multi-scale feature iterative fusion network is imported into the first multi-scale feature iterative fusion network; The decoding module: Used to import the output of the (N-1) multi-scale feature iterative fusion networks and the Nth first prediction result into the decoder respectively to obtain the second prediction result after segmentation; The cascading module: Used to perform cascaded feature extraction on the Nth first prediction result to the first first prediction result to obtain a more accurate second prediction result after segmentation.
6. The device according to claim 5, characterized in that, The encoding module includes: The import module: Used to import the original image into the first encoder among the N encoders; The first output module: Used to output the first first prediction result output by the first encoder to the second encoder and the first multi-scale feature iterative fusion network respectively; The second repetition module: Used to sequentially repeat the execution of the first output module until the (N-1)th first prediction result output by the (N-1)th encoder is output to the Nth encoder and the (N-1)th multi-scale feature iterative fusion network respectively; The second output module: Used to import the output of the Nth encoder into the decoder.
7. The device according to claim 6, characterized in that, It includes: The downsampling module: Used to downsample the first first prediction result output by the first encoder and then output it to the first multi-channel attention module; The third output module: Output the output of the second multi-scale feature iterative fusion network corresponding to the second encoder to the second multi-channel attention module; The connection module: Used to connect the output of the first multi-channel attention module and the output of the second multi-channel attention module and then output it to the third multi-channel attention module; The fourth output module: Use the output of the third multi-channel attention module as the output of the first multi-scale feature iterative fusion network; The third repetition module: Repeatedly execute the downsampling module, the third output module, the connection module and the fourth output module until the respective outputs of the (N-1) multi-scale feature iterative fusion networks are obtained.
8. The device according to claim 7, characterized in that, The decoding module includes: Import the outputs of the (N-1) multi-scale feature iterative fusion networks and the Nth first prediction result into the N decoders one by one to obtain N second prediction results after segmentation.
Citation Information
Patent Citations
Multi-channel multi-scale parallel encoding and decoding network image segmentation method and system and medium
CN112216371A
Image classification method and device based on multi-scale feature iterative fusion network
CN113076960A