Image segmentation method and device, model training method and device and electronic equipment
Through the encoder and decoder model of the U-net network structure, the dependence relationship of adjacent areas in the image is preserved, and the target feature map is generated for image segmentation, which solves the problem of poor image segmentation effect in the prior art and improves the accuracy and application effect of image segmentation.
Patent Information
- Application Number
- CN202411037993.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-30
- Publication Date
- 2025-07-25
AI Technical Summary
In the prior art, the image segmentation effect is poor.
Using a segmentation model based on U-net network structure, the image is characterized and scanned through an encoder and decoder, the dependence between adjacent areas in the image is preserved, and the target feature map is generated for image segmentation.
The accuracy and effect of image segmentation are improved, especially in applications in the fields of medical image analysis, autonomous driving, drone navigation and security monitoring, and the accuracy of identification and segmentation is enhanced.
Smart Images

Figure CN120374639A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image processing, and in particular, to an image segmentation method, a model training method, an apparatus, and an electronic device. Background Art
[0002] Image segmentation is a core technology in the field of computer vision. In recent years, with the growth of the computing power of computer devices and the improvement of data sets, image segmentation algorithms based on deep learning have developed rapidly and play a crucial role in many fields and application scenarios.
[0003] However, in related technologies, the effect of image segmentation is poor. Summary of the Invention
[0004] This application aims to solve at least one of the technical problems in related technologies to some extent.
[0005] To this end, this application proposes an image segmentation method, a model training method, an apparatus, and an electronic device to improve the effect of image segmentation.
[0006] An embodiment of one aspect of this application proposes an image segmentation method, including:
[0007] Obtain an image to be segmented;
[0008] Use the trained segmentation model to perform feature segmentation and scanning processing on the image to be segmented to obtain the target feature map of the image to be segmented; wherein, the target feature map includes the dependency relationship between adjacent regions in the image to be segmented;
[0009] Perform image segmentation according to the target feature map to obtain an image segmentation result.
[0010] An embodiment of another aspect of this application proposes a model training method, including:
[0011] Obtain a training sample image;
[0012] Use the segmentation model to perform feature segmentation and scanning processing on the training sample image to obtain the target feature map of the training sample image; wherein, the target feature map includes the dependency relationship between adjacent regions in the training sample image;
[0013] Perform image segmentation according to the target feature map to obtain an image predicted segmentation result;
[0014] Determine a loss function according to the difference between the image predicted segmentation result and the true segmentation result corresponding to the training sample image;
[0015] Adjust the parameters of the segmentation model according to the loss function to obtain a trained segmentation model.
[0016] Another embodiment of the present application provides an image segmentation device, including:
[0017] An acquisition module, configured to acquire an image to be segmented;
[0018] A first processing module, configured to perform feature segmentation and scanning processing on the image to be segmented by using a trained segmentation model to obtain a target feature map of the image to be segmented; wherein, the target feature map includes the dependency relationship between adjacent regions in the image to be segmented;
[0019] A second processing module, configured to perform image segmentation according to the target feature map to obtain an image segmentation result.
[0020] Another embodiment of the present application provides a model training device, including:
[0021] An acquisition module, configured to acquire a training sample image;
[0022] A first processing module, configured to perform feature segmentation and scanning processing on the training sample image by using a segmentation model to obtain a target feature map of the training sample image; wherein, the target feature map includes the dependency relationship between adjacent regions in the training sample image;
[0023] A second processing module, configured to perform image segmentation according to the target feature map to obtain an image predicted segmentation result;
[0024] A determination module, configured to determine a loss function according to the difference between the image predicted segmentation result and the true segmentation result corresponding to the training sample image;
[0025] A parameter adjustment module, configured to adjust the parameters of the segmentation model according to the loss function to obtain a trained segmentation model.
[0026] Another embodiment of the present application provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the image segmentation method as described in the foregoing one aspect, and implements the model training method as described in the foregoing other aspect.
[0027] Another embodiment of the present application provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the image segmentation method as described in the foregoing one aspect, and implements the model training method as described in the foregoing other aspect.
[0028] On the other hand, an embodiment of the present application provides a computer program product, on which a computer program is stored. When the program is executed by a processor, it implements the image segmentation method described in the foregoing one aspect, and implements the model training method described in the foregoing other aspect.
[0029] For the image segmentation method, model training method, device and electronic device provided by the present application, a to-be-segmented image is obtained, and the to-be-segmented image is subjected to feature segmentation and scanning processing by using a trained segmentation model to obtain a target feature map of the to-be-segmented image. Wherein, the target feature map includes the dependency relationship between adjacent regions in the to-be-segmented image, and image segmentation is performed according to the target feature map to obtain an image segmentation result. By using the trained segmentation model to perform feature segmentation and scanning processing on the to-be-segmented image, a target feature map including the dependency relationship between adjacent regions in the to-be-segmented image is obtained, and image segmentation based on the target feature map improves the effect of image segmentation.
[0030] Additional aspects and advantages of the present application will be given in part in the following description, become apparent in part from the following description, or be learned through the practice of the present application. Description of the Drawings
[0031] The above-mentioned and / or additional aspects and advantages of the present application will become apparent and easy to understand from the following description of the embodiments in conjunction with the drawings, wherein:
[0032] Figure 1 It is a schematic flowchart of an image segmentation method provided by an embodiment of the present application;
[0033] Figure 2 It is a schematic flowchart of another image segmentation method provided by an embodiment of the present application;
[0034] Figure 3 It is a schematic structural diagram of a segmentation model provided by an embodiment of the present application;
[0035] Figure 4 It is a schematic structural diagram of processing the input feature map and the output feature map of a segmentation scanning module provided by an embodiment of the present application;
[0036] Figure 5 It is a schematic flowchart of another image segmentation method provided by an embodiment of the present application;
[0037] Figure 6 It is a schematic structural diagram of a segmentation scanning module provided by an embodiment of the present application;
[0038] Figure 7 It is a schematic flowchart of a model training method provided by an embodiment of the present application;
[0039] Figure 8Schematic structural diagram of an image segmentation device provided by an embodiment of the present application;
[0040] Figure 9 Schematic structural diagram of a model training device provided by an embodiment of the present application;
[0041] Figure 10 Schematic structural diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners
[0042] The embodiments of the present application will be described in detail below. Examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to explain the present application, and should not be construed as limiting the present application.
[0043] The image segmentation method, model training method, device and electronic device of the embodiments of the present application will be described below with reference to the accompanying drawings.
[0044] Figure 1 Schematic flow diagram of an image segmentation method provided by an embodiment of the present application.
[0045] In the embodiments of the present application, the image segmentation method is configured in an image segmentation device for illustration. The image segmentation device can be applied to any electronic device so that the electronic device can perform the image segmentation function.
[0046] Among them, the electronic device can be any device with computing power. For example, it can be a mobile terminal. The mobile terminal can be a hardware device such as a mobile phone, a tablet computer, a personal digital assistant, a wearable device, etc., which has various operating systems, touch screens, and / or display screens.
[0047] As Figure 1 shown, the method may include the following steps:
[0048] Step 101, obtain the image to be segmented.
[0049] Among them, the image to be segmented can be an image in different scenarios, specifically as follows:
[0050] In one scenario, the image to be segmented is a medical image in a medical image analysis scenario. By performing semantic segmentation and recognition on the medical image, organs, tumors, or other lesion areas can be recognized and separated, which is crucial for early disease diagnosis, surgical planning, treatment monitoring, etc.
[0051] In the second scenario, the image to be segmented is an image of the surrounding environment collected in an autonomous driving scenario. Autonomous vehicles rely on image segmentation to identify roads, pedestrians, vehicles, etc. in order to make safe driving decisions. Through segmentation, the system can accurately distinguish driving lanes, pedestrian walkways, obstacles, etc., improving the safety and reliability of driving.
[0052] In the third scenario, the image to be segmented is an image obtained from aerial photography in a drone scenario. In drone navigation, by performing image segmentation on the image to be segmented, it can help the drone more accurately identify obstacles in the environment, improving the accuracy of navigation. Alternatively, image segmentation helps to identify the health status of farm crops, terrain, disaster areas, etc., providing support for precision agriculture, environmental monitoring, disaster assessment, etc.
[0053] In the fourth scenario, the image to be segmented is an image collected in the field of security and surveillance, by real-time identifying and tracking specific targets, such as pedestrians, etc.
[0054] Step 102: Use the trained segmentation model to perform feature segmentation and scanning processing on the image to be segmented, to obtain the target feature map of the image to be segmented.
[0055] Among them, the target feature map includes the dependency relationship between adjacent regions in the image to be segmented.
[0056] In the embodiments of the present application, the trained segmentation model performs feature segmentation and scanning processing on the feature map of the image to be segmented, converts the feature map into a feature sequence for processing during the segmentation and scanning processing, and retains the dependency relationship between adjacent regions in the image to be segmented during the processing of the feature sequence, so that the features adjacent in the original space remain adjacent during the segmentation and scanning processing, that is, it will not disperse the dependency relationship between the spatial features of the originally adjacent regions, making the modeled target feature map include local features and global features, and the local features in the image include important information about objects and structures in the image. Compared with the related art where local features will be dispersed, the production method of the target feature map of the present application can improve the effect of image segmentation.
[0057] In one implementation of the embodiment of the present application, the segmentation model is based on the network structure of U-net. The segmentation model includes an encoder and a decoder. The encoder is used to perform feature segmentation and scanning processing on the image to be segmented, and a feature map output by the encoder is obtained. Furthermore, the feature map output by the encoder is input into the decoder for feature segmentation and scanning processing to obtain the target feature map of the image to be segmented. In the embodiment of the present application, both the encoder and the decoder in the segmentation model perform segmentation and scanning processing. During the process of converting the feature map into a feature sequence for segmentation and scanning processing, the dependency relationship between adjacent regions in the image to be segmented is retained, so that features adjacent in the original space remain adjacent during the segmentation and scanning processing, that is, the dependency relationship between the spatial features of the originally adjacent regions will not be scattered. As a result, the obtained target feature map includes local features and global features, improving the accuracy of image segmentation and the segmentation effect.
[0058] Step 103, perform image segmentation according to the target feature map to obtain an image segmentation result.
[0059] In the embodiment of the present application, the target feature map includes the dependency relationship between adjacent regions in the image to be segmented. This dependency relationship can identify the spatial adjacency relationship between adjacent regions during image segmentation, and can improve the segmentation effect during image segmentation.
[0060] In the image segmentation method of the embodiment of the present application, an image to be segmented is obtained, and the trained segmentation model is used to perform feature segmentation and scanning processing on the image to be segmented to obtain the target feature map of the image to be segmented, where the target feature map includes the dependency relationship between adjacent regions in the image to be segmented. Image segmentation is performed according to the target feature map to obtain an image segmentation result. By using the trained segmentation model to perform feature segmentation and scanning processing on the image to be segmented, a target feature map including the dependency relationship between adjacent regions in the image to be segmented is obtained, and image segmentation based on the target feature map improves the image segmentation effect.
[0061] Based on the above embodiments, Figure 2 is a schematic flowchart of another image segmentation method provided by the embodiment of the present application. As Figure 2 shown, this method includes the following steps:
[0062] Step 201, obtain an image to be segmented.
[0063] Among them, step 201 can refer to the relevant explanations in the foregoing embodiments. The principles are the same, and details are not described here again.
[0064] In the embodiment of the present application, the segmentation model includes an encoder, the encoder includes L sub-encoders, and each sub-encoder includes a first segmentation and scanning module, where L is a natural number greater than or equal to 1.
[0065] Step 202: For any one of the L sub-encoders, determine the first feature map to be processed as input according to the position of the any one of the L sub-encoders in the L sub-encoders.
[0066] Wherein, the first feature map to be processed is determined according to the feature map of the image to be segmented or the feature map output by the previous sub-encoder of any one of the sub-encoders.
[0067] In the embodiments of the present application, if any one of the sub-encoders is the first sub-encoder among the L sub-encoders, the first feature map to be processed input to the first sub-encoder is determined according to the feature map of the image to be segmented. As an implementation, the feature map of the image to be segmented is first processed through a layer normalization (LN) layer, and the result of the normalization layer processing is passed through a linear transformation and a reshape operation, and then through a depth convolution layer and a Silu activation function to obtain the first feature map to be processed. If any one of the sub-encoders is a non-first sub-encoder among the L sub-encoders, the first feature map to be processed input to the non-first sub-encoder is determined according to the feature map output by the previous sub-encoder of the non-first sub-encoder. As an implementation, the feature map output by the previous sub-encoder of the non-first sub-encoder is first processed through a layer normalization (LN) layer, and the result of the normalization layer processing is passed through a linear transformation and a reshape operation, and then through a depth convolution layer and a Silu activation function to obtain the first feature map to be processed. Except for the first sub-encoder among the L sub-encoders, the output of the previous sub-encoder is the input of the subsequent sub-encoder.
[0068] As an example, Figure 3 is a schematic structural diagram of a segmentation model provided by the embodiments of the present application. As shown in Figure 3 , the image to be segmented is a medical image in a medical image analysis scenario. Among them, the encoder includes 4 sub-encoders, and each sub-encoder includes a first segmentation scanning module. The 4 first segmentation scanning modules are respectively numbered N1, N2, N3, and N4. Among them, the sub-encoder corresponding to N1 is the first sub-encoder, and the sub-encoder corresponding to N4 is the last sub-encoder. The input of N1 is the feature map of the medical image to be segmented. N2, N3, and N4 are cascaded, that is, the feature map output by the previous layer is the input of the subsequent layer. That is to say, the feature map output by N1 is the input of N2, the feature map output by N2 is the input of N3, and the feature map output by N3 is the input of N4. Figure 3 The sub-encoder 1, sub-encoder 4, sub-decoder 1, and sub-decoder 2 are marked in the example only for illustrative purposes. For other sub-encoders and other sub-decoders, reference may be made to the marked sub-encoder 1 and sub-decoder 2.
[0069] Step 203: Use the first segmentation and scanning module of any sub-encoder to perform feature segmentation and scanning on the first feature map to be processed, and obtain the feature map output by the first segmentation and scanning module of any sub-encoder.
[0070] Among them, the first segmentation and scanning module includes a Bidirectional Slice Scan (BSS) module, which divides the first feature map to be processed into segmentation results (slices) in multiple directions, helping to preserve and enhance the spatial continuity of features during the serialization process. Furthermore, each segmentation is scanned bidirectionally, which helps to retain the local information between adjacent segmentations, ensuring that the features still maintain local adjacency after serialization and facilitating the capture of detailed features.
[0071] Specifically, the structure of the first segmentation and scanning module will be described in detail in subsequent embodiments.
[0072] Step 204: Determine the feature map output by the sub-encoder according to the feature map output by the first segmentation and scanning module of the sub-encoder.
[0073] In the embodiments of the present application, the encoder includes multiple sub-encoders. Except for the last sub-encoder among the multiple sub-encoders, each sub-encoder may further include a downsampling module. The following will separately describe the output feature maps according to the positions of the sub-encoders:
[0074] As an implementation manner, in response to the sub-encoder being the last one, since the last sub-encoder is connected to the first sub-decoder of the decoder, the last sub-encoder does not include a downsampling module, and the sub-decoder connected to the last sub-encoder also does not include an upsampling module, reducing the number of modules in the model, reducing unnecessary computational complexity, and improving processing efficiency. Furthermore, the feature map output by the last first segmentation and scanning module is used as the feature map output by the sub-encoder.
[0075] As another implementation manner, in response to the sub-encoder being the mth one, use the downsampling module of the mth sub-encoder to perform downsampling and channel dimension processing on the feature map output by the first segmentation and scanning module of the mth sub-encoder to obtain the feature map output by the mth sub-encoder. Wherein, m is a natural number greater than or equal to 1 and less than L, that is, the mth sub-encoder is not the last sub-encoder.
[0076] As an example, such as Figure 3As shown, m is 1, that is, any one of the sub-encoders is the first sub-encoder among the 4 encoders. The first encoder includes a first segmentation and scanning module and a downsampling module. The processing of the feature map by the first segmentation and scanning module will not change the shape of the features. Among them, the features input to the first segmentation and scanning module are determined according to the feature map of the image to be segmented. As an implementation, as Figure 4 shown Figure 4 is a schematic structural diagram for processing the input feature map and the output feature map of the first segmentation and scanning module provided by an embodiment of the present application. As Figure 4 shown, the feature map of the image to be segmented is first processed by normalization (layer normolization, LN), and the result of the normalization process is passed through a linear transformation (Liner) and a shape adjustment reshape operation, and then passed through a depth convolution layer (DW-Conv) and a Silu activation function to obtain the first feature map to be processed by the first segmentation and scanning module. The first feature map to be processed is processed by the first segmentation and scanning module to obtain the feature map output by the first segmentation and scanning module of the sub-encoder. Among them, after the feature map of the image to be segmented passes through the normalization LN, it is then divided into two parts along the channel dimension. One part is passed through a linear transformation and a Silu activation process, and then multiplied by the feature map obtained after the feature map output by the first segmentation and scanning module of the sub-encoder passes through the normalization process again. The result of the multiplication is passed through a linear transformation and then added to the feature map of the original input image to be segmented to obtain the target feature map output by the first segmentation and scanning module of the sub-encoder. Furthermore, the downsampling module, that is, the Patch Merging module, is used to halve the spatial size of the target feature map output by the first segmentation and scanning module of the sub-encoder, and at the same time double the number of channels. For example, the original size of a medical image is H×W×3, where H and W are the height and width of the medical image respectively, and 3 refers to the number of channels. The number of channels of the image is 3, that is, the RGB three channels. The original image is processed by the feature embedding layer Patchembeding, where the convolution kernel size K of the embedding layer is 4 and the stride S is 4. The feature map to be input to the first segmentation and scanning module obtained by processing each 4*4 pixel area in the original image is to realize mapping the number of channels from 3 to C. The feature map obtained after the feature map is input to the first segmentation and scanning module for processing is still of size Then, the feature map of size is input to the downsampling module in the first encoder for downsampling, and the obtained image size is At the same time, the number of channels is increased to obtain a channel number of 2C, so that the obtained feature map is While achieving a reduced spatial size, Patch Merging will correspondingly increase the number of channels (channel dimension) of the feature map. The purpose of this is to keep the computational complexity of the model within a reasonable range while increasing the model's ability to represent high-level features. Among them, the increase in the number of channels usually follows a predetermined ratio to ensure the balance between the depth and width of the model. As Figure 3 shown, for the second and third sub-encoders among the 4 sub-encoders, the processing method of the first sub-encoder can be referred to, and details will not be elaborated here. In this application, feature maps of different feature scales are extracted through multiple sub-encoders, enabling the segmentation model to learn features of different scales, facilitating the segmentation model to learn rich feature representations, which is crucial for capturing local details and global structures in the image and improving the segmentation effect of the segmentation model.
[0077] Step 205, for the m-th sub-decoder among the L sub-decoders, determine the second feature map to be processed corresponding to the second segmentation scanning module of the m-th sub-decoder according to the position of the m-th sub-decoder among the L sub-decoders.
[0078] In the embodiments of this application, as Figure 3 shown, the segmentation model includes an encoder and a decoder. The encoder includes L cascaded sub-encoders, and the decoder includes L cascaded sub-decoders. There are skip connections between the outputs of the L sub-encoders in the encoder and the outputs of the L sub-decoders in the decoder to strengthen the feature aggregation between low-level features and high-level features and improve the accuracy of feature extraction. Among them, L is a natural number greater than or equal to 1. Taking the encoder as an example, when the encoder includes L sub-encoders, the numbers of the sub-encoders from low level to high level are 1, 2, 3, 4, ··· L. Taking the decoder as an example, when the decoder includes L sub-decoders, the numbers of the sub-decoders from low level to high level are 1, 2, 3, 4, ··· L.
[0079] Among them, for identification, the segmentation scanning module in the encoder is called the first segmentation scanning module, and the segmentation scanning module in the decoder is called the second segmentation scanning module. The structures of the first segmentation scanning module and the second segmentation scanning module are the same, and the functions implemented are also the same. The second segmentation scanning module can refer to the relevant explanations of the first segmentation scanning module, and the principle is the same, so details will not be elaborated here.
[0080] Among them, the second feature map to be processed corresponding to the second segmentation scanning module of the m-th sub-decoder is determined according to the feature map output by the (L - m + 1)-th sub-encoder, or is determined according to the fused feature map obtained by fusing the feature map output by the (L - m + 1)-th sub-encoder and the feature map output by the previous sub-decoder of the m-th sub-decoder.
[0081] In one scenario, if the m-th sub-decoder is the first sub-decoder among the L sub-encoders, i.e., m = 1, then the second feature map to be processed input to the second segmentation and scanning module of the first sub-decoder is determined based on the feature map output by the L-th sub-encoder. That is to say, the output of the last sub-encoder is the input of the first sub-decoder. As an implementation, as Figure 4 shown, the feature map output by the L-th sub-encoder is first processed through a layer normalization (LN) layer, and the result of the normalization layer processing is passed through a linear transformation and a reshape operation, and then through a depth convolution layer and a Silu activation function to obtain the second feature map to be processed.
[0082] In the second scenario, if the m-th sub-decoder is a non-first sub-decoder among the L sub-decoders, then the second feature map to be processed input to the m-th sub-decoder is determined based on the fused feature map obtained by fusing the feature map output by the (L - m + 1)-th sub-encoder and the feature map output by the previous sub-decoder of the m-th sub-decoder. This realizes the enhancement of multi-layer features generated by the encoder and decoder through the skip connection mechanism. Through the fusion of multi-scale features, local and global features are simultaneously captured at different stages and scales of the feature map, increasing the amount of information included in the fused feature map input to each sub-decoder. Among them, for the determination method of the second feature map to be processed, as an implementation, as Figure 4 shown, the fused feature map is first processed through a layer normalization (LN) layer, and the result of the normalization layer processing is passed through a linear transformation and a reshape operation, and then through a depth convolution layer and a Silu activation function to obtain the second feature map to be processed. Among the L sub-decoders, except for the first sub-decoder, the output of the previous sub-decoder is the input of the subsequent sub-decoder.
[0083] As an example, as Figure 3 shown, taking m = 2 as an example, that is, for the second sub-decoder, the second sub-decoder includes an upsampling module and a second segmentation and scanning module S2, and the feature map output by the first sub-decoder is The feature map output by the fourth sub-encoder is By adding the feature map output by the first sub-decoder and the feature map output by the fourth sub-encoder, the obtained fused feature map is Through the skip connection, the model can fuse features at different levels. Even on feature maps with a fixed size and number of channels, it can utilize multi-level details and context information, realizing an increase in the amount of information carried by the fused feature map without increasing the size and number of channels of the feature map.
[0084] In the embodiment of the present application, the first sub-decoder module includes a second segmentation and scanning module, and the sub-decoders other than the first sub-decoder include a second segmentation and scanning module and an upsampling module. Among them, the first sub-decoder module does not include an upsampling module. For the specific reason, reference can be made to the relevant explanation in the foregoing embodiment about the last sub-encoder not including a downsampling module. The principle is the same and will not be elaborated here.
[0085] As an example, Figure 3 as shown, taking m = 2 as an example, that is, for the second sub-decoder, the second sub-decoder includes an upsampling module and a second segmentation and scanning module S2. The input of the upsampling module of the second sub-decoder is Through upsampling and channel dimension processing, the input feature map is changed from to The size is increased and the number of channels is reduced. The feature map output by the upsampling module of the second sub-decoder is It should be noted that among them, upsampling and downsampling in the encoder are inverse operations, that is, the upsampling Patch Expanding operation in the decoder reverses the transformation performed by the downsampling Patch Merging module in the encoder.
[0086] Step 206: Use the second segmentation and scanning module of the m-th sub-decoder to perform feature segmentation and scanning processing on the second to-be-processed feature map, and determine the feature map output by the m-th sub-decoder.
[0087] In the embodiment of the present application, for the processing method of the feature map input and output by the second segmentation and scanning module, in the case of L = 4, taking m = 1 as an example, that is, the second to-be-processed feature map input by the second segmentation and scanning module of the first sub-decoder is determined according to the feature map output by the fourth sub-encoder, as Figure 4As shown in the figure, the feature map output by the fourth sub-encoder is first processed by layer normalization (LN). After the result of the normalization process undergoes a linear transformation (Liner) and a reshape operation, it passes through a depthwise convolution layer (DW-Conv) and a Silu activation function to obtain the second feature map to be processed of the second segmentation scanning module. After the second feature map to be processed is processed by the second segmentation scanning module, the feature map output by the second segmentation scanning module of this sub-decoder is obtained. Among them, after the output of the fourth sub-encoder undergoes LN normalization, it is divided into two parts along the channel dimension. One part undergoes a linear transformation and Silu activation processing, and then is multiplied by the feature map obtained after the feature map output by the second segmentation scanning module of this sub-decoder undergoes normalization again. The result of the multiplication undergoes a linear transformation and then is added to the output of the original input fourth sub-encoder to obtain the feature map output by this sub-decoder, improving the efficiency and accuracy of determining the sub-decoder feature map.
[0088] Among them, the second segmentation scanning module includes a bidirectional slice scan (BSS) module, which divides the second feature map to be processed into segmentation results (slices) in multiple directions, helping to preserve and enhance the spatial continuity of features during the serialization process. Furthermore, a bidirectional scan is performed on each segmentation, which helps to retain the local information between adjacent segmentations, ensuring that the features still maintain local adjacency after serialization processing and is conducive to capturing detailed features.
[0089] Specifically, the structure of the second segmentation scan is the same as that of the first segmentation scan module, which will be described in detail in subsequent embodiments.
[0090] Step 207: Determine the target feature map according to the feature map of the image to be segmented and the feature map output by the last sub-decoder.
[0091] In the embodiment of the present application, an skip connection method between the encoder and the decoder is adopted, and the result of adding the feature map of the image to be segmented and the feature map output by the last sub-decoder is used as the target feature map output by the decoder. For the beneficial effects of the skip connection, reference can be made to the relevant explanations in the foregoing embodiments, which will not be elaborated here.
[0092] Step 208: Perform image segmentation according to the target feature map to obtain an image segmentation result.
[0093] In the embodiment of the present application, the target feature map is mapped through a final mapping layer to implement the segmentation processing of the image and obtain an image segmentation result, such as Figure 3 the black and white binary map of the medical image in, and the accuracy of the image segmentation result is relatively high.
[0094] In the image segmentation method according to the embodiments of the present application, the segmentation model includes an encoder and a decoder. Each sub-encoder in the encoder and each sub-decoder in the decoder are connected in a skip layer manner. Each sub-encoder and each sub-decoder include a segmentation scanning module, which are referred to as the first segmentation scanning module and the second segmentation scanning module. Both the first segmentation scanning module and the second segmentation scanning module adopt the method of dividing the input feature map into segmentation results (slices) in multiple directions, which helps to preserve and enhance the spatial continuity of features during the serialization process. Furthermore, a two-way scan is performed on each segmentation, which helps to preserve the local information between adjacent segmentations, ensures that the features still maintain local adjacency after serialization processing, is conducive to capturing detailed features, and facilitates improving the segmentation effect during subsequent image segmentation.
[0095] Based on the above embodiments, Figure 5 It is a schematic flowchart of another image segmentation method provided by the embodiments of the present application, which specifically illustrates how to perform feature segmentation and scanning processing on image features so that the generated target feature map includes local features for image segmentation. Local features are crucial for image segmentation because they provide important information about the objects and structures in the image. Therefore, effectively modeling the image features for image segmentation can improve the accuracy of image segmentation. In the embodiments of the present application, both the encoder and the decoder include a segmentation scanning module, which are referred to as the first segmentation scanning module and the second segmentation scanning module for easy distinction. Among them, the structures of the first segmentation scanning module and the second segmentation scanning module are the same. In this embodiment, any one of the first segmentation scanning modules in the encoder is taken as an example for illustration. As Figure 5 shown, step 203 includes the following steps:
[0096] Step 501, use the feature segmentation module in the first segmentation scanning module of any sub-encoder to perform feature segmentation operations on the first feature map to be processed in multiple set directions, and obtain segmentation feature sequences in each set direction.
[0097] As an example, Figure 6 It is a schematic structural diagram of a segmentation scanning module provided by the embodiments of the present application. As Figure 6 shown, the first segmentation scanning module includes a feature segmentation module, a feature scanning module, a feature reconstruction module, and a feature fusion module.
[0098] In the embodiments of the present application, a feature segmentation module is used to perform feature segmentation operations on the feature map output by the previous sub-encoding module in a set direction to obtain a segmentation feature sequence in the set direction. Since an image includes multiple pixels, processing features in units of pixels involves a large amount of computation. Through the feature segmentation operation, the image is divided into multiple pixel units. A pixel unit includes multiple adjacent pixels, and a pixel unit corresponds to a segmentation feature region. The multiple segmentation feature regions obtained by segmentation are combined into a segmentation feature sequence in the set direction. Processing based on the segmentation feature sequence can reduce the amount of computation and improve processing efficiency.
[0099] In one implementation manner of the embodiments of the present application, the multiple set directions include the horizontal direction and the vertical direction. That is to say, the feature map is segmented in the horizontal direction and the vertical direction. The first segmentation sub-module of the feature segmentation module is used to perform feature segmentation operations on the first feature map to be processed in the horizontal direction to obtain a segmentation feature sequence in the horizontal direction, and the second segmentation sub-module of the feature segmentation module is used to perform feature segmentation operations on the first feature map to be processed in the vertical direction to obtain a segmentation feature sequence in the vertical direction. Performing feature segmentation in multiple directions helps to retain and enhance the spatial continuity of features.
[0100] Step 502: The feature scanning module in the first segmentation scanning module of any sub-encoder is used to perform feature aggregation on the segmentation feature sequences in each set direction using the corresponding scanning method to obtain the scanning feature sequences in each set direction.
[0101] In one implementation manner of the embodiments of the present application, for the segmentation feature sequences in different directions, corresponding scanning methods are used. That is to say, different scanning mechanisms are applied to features of different shapes. Although the segmentation feature sequences in different directions are obtained by segmenting the feature map, their continuity will be retained during the scanning stage, and feature expansion operations are performed through scanning. Among them, for the segmentation feature sequence in the horizontal direction, the first sub-scanning module of the feature scanning module performs bottom-up scanning and top-down scanning on the segmentation feature sequence in the horizontal direction in the vertical direction respectively to aggregate and obtain the first sub-scanning feature sequence and the second sub-scanning feature sequence. Features adjacent in the original space remain adjacent in the scanning sequence, which is convenient for local feature modeling; for the segmentation feature sequence in the vertical direction, the second sub-scanning module of the feature scanning module performs left-to-right scanning and right-to-left scanning in the horizontal direction respectively to aggregate and obtain the third sub-scanning feature sequence and the fourth sub-scanning feature sequence. Features adjacent in the original space remain adjacent in the scanning sequence, which is convenient for local feature modeling. In contrast, the traditional Mamba-based method will increase the distance between features in the sequence, which will disperse the originally adjacent spatial features and is not conducive to local feature modeling.
[0102] In the embodiments of the present application, since images do not have the context association like text and speech and cannot obtain ordered information, after converting images into feature sequences for processing, it is easy to increase the distance of features in space. The mamba model in the related art does not design a sequence scanning method for the features of image segmentation. However, the present application designs a scanning method suitable for image segmentation for the features of image segmentation, which can effectively extract detailed features and improve the segmentation effect. Through the feature scanning operation of the present application, feature associations are generated between adjacent regions, and operations are performed on the segmentation feature sequences to achieve feature aggregation with a relatively low computational cost. The first sub-scanning feature sequence and the second sub-scanning feature sequence obtained by vertical scanning aggregation, and the third sub-scanning feature sequence and the fourth sub-scanning feature sequence obtained by horizontal scanning aggregation are used to achieve aggregation in different aggregation directions. Since the aggregation results in different directions are different, in order to reduce information loss and make feature aggregation more stable, the effect of feature aggregation is improved by performing aggregation in different directions.
[0103] Step 503: Use the feature reconstruction module in the first segmentation scanning module of any sub-encoder to perform temporal feature extraction on the scanning feature sequences in each set direction to obtain the temporal feature sequences in each set direction.
[0104] In the embodiments of the present application, temporal feature extraction is respectively performed on the first sub-scanning feature sequence and the second sub-scanning feature sequence, and the third sub-scanning feature sequence and the fourth sub-scanning feature sequence to obtain two temporal feature sequences in the horizontal direction and two temporal feature sequences in the vertical direction, so as to obtain the temporal information between features.
[0105] Step 504: Use the feature fusion module in the first segmentation scanning module of any sub-encoder to perform feature fusion and restoration on the temporal feature sequences in multiple set directions to obtain the feature map output by the first segmentation scanning module of any sub-encoder.
[0106] In the embodiments of the present application, a feature map is converted into temporal feature sequences in multiple set directions through operations such as segmentation and scanning, and then through a restoration operation, that is, the features at corresponding positions of the temporal feature sequences in multiple set directions are merged to restore the shape and size of the original feature map, realizing the effectiveness in simultaneously modeling local and global features.
[0107] In the image segmentation method of the embodiments of the present application, a feature map is converted into temporal feature sequences in multiple set directions through operations such as segmentation and scanning, and then through a restoration operation, that is, the features at corresponding positions of the temporal feature sequences in multiple set directions are merged to restore the shape and size of the original feature map, realizing the effectiveness in simultaneously modeling local and global features. Local features are extremely important in the process of feature segmentation and can improve the accuracy of image segmentation for images.
[0108] Among them, on mobile phones, semantic segmentation technology has achieved some practical applications. For example, face swapping, automatic matting, intelligent background replacement, and adjustment of specific regions based on semantic tags on the mobile phone side all use semantic segmentation methods to finely segment the subject and the background. With the development of mobile phone hardware, more and more artificial intelligence (AI) algorithms, including semantic segmentation algorithms, are used on the ISP side of the processor, so as to realize AI intervention during the photo-taking and video-recording stages, further improving the image quality and meeting the diverse needs of users.
[0109] Based on the above embodiments, Figure 7 is a schematic flowchart of a model training method provided by an embodiment of the present application. As Figure 7 shown, this method includes the following steps:
[0110] Step 701, obtain training sample images.
[0111] The training sample images can be images describing real scenes in the dataset, and the image format can be RGB or grayscale.
[0112] Step 702, use the segmentation model to perform feature segmentation and scanning processing on the training sample images to obtain the target feature maps of the training sample images.
[0113] Among them, the target feature maps include the dependency relationships between adjacent regions in the training sample images.
[0114] Step 703, perform image segmentation according to the target feature maps to obtain the image prediction segmentation results.
[0115] Among them, the relevant explanations in the foregoing embodiments also apply to steps 701 to 703 in this embodiment. The principles are the same and will not be elaborated here.
[0116] Step 704, determine the loss function according to the difference between the image prediction segmentation results and the ground-truth segmentation results corresponding to the training sample images.
[0117] Among them, the ground-truth segmentation results refer to the annotated images, that is, a single-channel image with the same length and width as the original input sample training image, and each pixel point corresponds to the class label of that point.
[0118] As an example, the loss function is determined by the following formula:
[0119] loss = -(ylog(y')+(1 - y)log(1 - y'));
[0120] y = Net(x);
[0121] Among them, Net is a segmentation model, x is the input training sample image, y is the predicted segmentation result of the image obtained after segmentation by the segmentation model, and y' is the segmentation label, which is the true segmentation result.
[0122] Step 705: Adjust the parameters of the segmentation model according to the loss function to obtain the trained segmentation model.
[0123] In the embodiment of the present application, the foregoing steps 701 to 705 need to be repeatedly executed multiple times, and different training sample images can be used each time. Thus, training is stopped when the loss function is less than the threshold, or training is stopped when the number of repeated executions is greater than the threshold. The segmentation model obtained after the last adjustment of the model parameters is used as the trained segmentation model.
[0124] In the model training method of the embodiment of the present application, by using different training sample images, training is stopped when the loss function is less than the threshold, or training is stopped when the number of repeated executions is greater than the threshold. The segmentation model obtained after the last adjustment of the model parameters is used as the trained segmentation model, so that the segmentation model can include the dependency relationship between adjacent regions in the image during the image segmentation process to improve the image segmentation effect.
[0125] To implement the above embodiment, the embodiment of the present application also proposes an image segmentation device.
[0126] Figure 8 It is a schematic structural diagram of an image segmentation device provided by the embodiment of the present application.
[0127] As Figure 8 shown, the device may include:
[0128] An acquisition module 81, configured to acquire an image to be segmented;
[0129] A first processing module 82, configured to perform feature segmentation and scanning processing on the image to be segmented by using the trained segmentation model to obtain a target feature map of the image to be segmented; wherein, the target feature map includes the dependency relationship between adjacent regions in the image to be segmented;
[0130] A second processing module 83, configured to perform image segmentation according to the target feature map to obtain an image segmentation result.
[0131] Further, in an implementation manner of the embodiment of the present application, the segmentation model includes an encoder and a decoder, and the first processing module 82 is further configured to:
[0132] Perform feature segmentation and scanning processing on the image to be segmented by using the encoder to obtain the feature map output by the encoder;
[0133] Input the feature map output by the encoder into the decoder for feature segmentation and scanning processing to obtain the target feature map of the image to be segmented.
[0134] In an implementation manner of the embodiment of the present application, the encoder includes L cascaded sub-encoders, and each sub-encoder includes a first segmentation and scanning module and a first processing module 82, and is further configured to:
[0135] For any one of the L sub-encoders, determine the first feature map to be processed corresponding to the first segmentation and scanning module of the any one of the L sub-encoders according to the position of the any one of the L sub-encoders in the L sub-encoders; wherein, the first feature map to be processed is determined according to the feature map of the image to be segmented or the feature map output by the previous sub-encoder of the any one of the L sub-encoders;
[0136] Use the first segmentation and scanning module of the any one of the L sub-encoders to perform feature segmentation and scanning processing on the first feature map to be processed to obtain the feature map output by the first segmentation and scanning module of the any one of the L sub-encoders;
[0137] Determine the feature map output by the any one of the L sub-encoders according to the feature map output by the first segmentation and scanning module of the any one of the L sub-encoders.
[0138] In an implementation manner of the embodiment of the present application, the decoder includes L cascaded sub-decoders, and each sub-decoder module includes a second segmentation and scanning module. The L sub-encoders in the encoder and the L sub-decoders in the decoder are connected in a skip connection manner, and the first processing module 82 is further configured to:
[0139] For the m-th sub-decoder among the L sub-decoders, determine the second feature map to be processed corresponding to the second segmentation and scanning module of the m-th sub-decoder according to the position of the m-th sub-decoder in the L sub-decoders; wherein, the second feature map to be processed is determined according to the feature map output by the (L - m + 1)-th sub-encoder, or according to the fusion feature map obtained by fusing the feature map output by the (L - m + 1)-th sub-encoder and the feature map output by the previous sub-decoder of the m-th sub-decoder;
[0140] Use the second segmentation and scanning module of the m-th sub-decoder to perform feature segmentation and scanning processing on the second feature map to be processed to determine the feature map output by the m-th sub-decoder;
[0141] Determine the target feature map according to the feature map of the image to be segmented and the feature map output by the last sub-decoder.
[0142] In one implementation manner of the embodiment of the present application, the first segmentation and scanning module includes a feature segmentation module, a feature scanning module, a feature reconstruction module, and a feature fusion module. The first processing module 82 is further configured to:
[0143] Use the feature segmentation module to perform feature segmentation operations on the first feature map to be processed in multiple set directions respectively, and obtain segmentation feature sequences in each set direction;
[0144] Use the feature scanning module to perform feature aggregation on the segmentation feature sequences in each set direction by using corresponding scanning methods, and obtain scanning feature sequences in each set direction;
[0145] Use the feature reconstruction module to perform temporal feature extraction on the scanning feature sequences in each set direction, and obtain temporal feature sequences in each set direction;
[0146] Use the feature fusion module to perform feature fusion and restoration on the temporal feature sequences in multiple set directions, and obtain the feature map output by the first segmentation and scanning module of any sub-encoder.
[0147] In one implementation manner of the embodiment of the present application, the multiple set directions include the horizontal direction and the vertical direction. The first processing module 82 is further configured to:
[0148] Use the first segmentation sub-module of the feature segmentation module to perform feature segmentation operations on the first feature map to be processed in the horizontal direction, and obtain a segmentation feature sequence in the horizontal direction;
[0149] Use the second segmentation sub-module of the feature segmentation module to perform feature segmentation operations on the first feature map to be processed in the vertical direction, and obtain a segmentation feature sequence in the vertical direction.
[0150] In one implementation manner of the embodiment of the present application, the first processing module 82 is further configured to:
[0151] For the segmentation feature sequence in the horizontal direction, use the first sub-scanning module of the feature scanning module to perform scans from bottom to top and from top to bottom respectively in the vertical direction, and aggregate to obtain a first sub-scanning feature sequence and a second sub-scanning feature sequence;
[0152] For the segmentation feature sequence in the vertical direction, use the second sub-scanning module of the feature scanning module to perform scans from left to right and from right to left respectively in the horizontal direction, and aggregate to obtain a third sub-scanning feature sequence and a fourth sub-scanning feature sequence.
[0153] It should be noted that the foregoing explanation of the method embodiment is also applicable to the device of this embodiment, and will not be elaborated here.
[0154] The image segmentation device proposed in this application obtains the image to be segmented, and uses the trained segmentation model to perform feature segmentation and scanning processing on the image to be segmented, obtaining the target feature map of the image to be segmented. Among them, the target feature map includes the dependency relationship between adjacent regions in the image to be segmented. Image segmentation is performed based on the target feature map to obtain the image segmentation result. By using the trained segmentation model to perform feature segmentation and scanning processing on the image to be segmented, a target feature map including the dependency relationship between adjacent regions in the image to be segmented is obtained, and image segmentation based on the target feature map improves the effect of image segmentation.
[0155] To implement the above embodiments, an embodiment of this application also proposes a model training device.
[0156] Figure 9 It is a schematic structural diagram of a model training device provided by an embodiment of this application.
[0157] As Figure 9 shown, the device may include:
[0158] An acquisition module 91, configured to acquire a training sample image;
[0159] A first processing module 92, configured to use the segmentation model to perform feature segmentation and scanning processing on the training sample image, obtaining the target feature map of the training sample image; among them, the target feature map includes the dependency relationship between adjacent regions in the training sample image;
[0160] A second processing module 93, configured to perform image segmentation based on the target feature map to obtain an image predicted segmentation result;
[0161] A determination module 94, configured to determine a loss function according to the difference between the image predicted segmentation result and the true value segmentation result corresponding to the training sample image;
[0162] A parameter adjustment module 95, configured to adjust the parameters of the segmentation model according to the loss function to obtain a trained segmentation model.
[0163] It should be noted that the foregoing explanation of the method embodiment also applies to the device of this embodiment, and will not be elaborated here.
[0164] In the model training device of the embodiment of this application, by using different training sample images, training is stopped when the loss function is less than a threshold, or training is stopped when the number of repeated executions is greater than a threshold. The segmentation model obtained after the last model parameter adjustment is used as the trained segmentation model, so that the segmentation model can include the dependency relationship between adjacent regions in the image during the image segmentation process, thereby improving the image segmentation effect.
[0165] To implement the above embodiments, the present application further provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the method described in the foregoing method embodiments is implemented.
[0166] To implement the above embodiments, the present application further provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, the method described in the foregoing method embodiments is implemented.
[0167] To implement the above embodiments, the present application further provides a computer program product, on which a computer program is stored. When the computer program is executed by a processor, the method described in the foregoing method embodiments is implemented.
[0168] Figure 10 FIG. is a schematic structural diagram of an electronic device provided by an embodiment of the present application. For example, the electronic device 800 may be a mobile phone, a computer, a digital broadcast terminal, a messaging device, a game console, a tablet device, a medical device, a fitness device, a personal digital assistant, etc.
[0169] Referring to Figure 10 , the electronic device 800 may include one or more of the following components: a processing component 802, a memory 804, a power component 806, a multimedia component 808, an audio component 810, an input / output (I / O) interface 812, a sensor component 814, and a communication component 816.
[0170] The processing component 802 generally controls the overall operation of the electronic device 800, such as operations associated with display, telephone calls, data communication, camera operations, and recording operations. The processing component 802 may include one or more processors 820 to execute instructions to complete all or part of the steps of the above method. In addition, the processing component 802 may include one or more modules to facilitate the interaction between the processing component 802 and other components. For example, the processing component 802 may include a multimedia module to facilitate the interaction between the multimedia component 808 and the processing component 802.
[0171] The memory 804 is configured to store various types of data to support the operation of the electronic device 800. Examples of such data include instructions for any application or method operating on the electronic device 800, contact data, phone book data, messages, pictures, videos, and the like. The memory 804 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk, or optical disk.
[0172] The power component 806 provides power to various components of the electronic device 800. The power component 806 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power for the electronic device 800.
[0173] The multimedia component 808 includes a screen that provides an output interface between the electronic device 800 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen can be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors can sense not only the boundaries of the touch or swipe actions but also detect the duration and pressure associated with the touch or swipe operation. In some embodiments, the multimedia component 808 includes a front camera and / or a rear camera. When the electronic device 800 is in an operating mode, such as a shooting mode or a video mode, the front camera and / or the rear camera can receive external multimedia data. Each of the front camera and the rear camera can be a fixed optical lens system or have a focal length and optical zoom capabilities.
[0174] The audio component 810 is configured to output and / or input audio signals. For example, the audio component 810 includes a microphone (MIC) that is configured to receive external audio signals when the electronic device 800 is in an operating mode, such as a call mode, a recording mode, and a voice recognition mode. The received audio signals can be further stored in the memory 804 or transmitted via the communication component 816. In some embodiments, the audio component 810 further includes a speaker for outputting audio signals.
[0175] The I / O interface 812 provides an interface between the processing component 802 and a peripheral interface module, which can be a keyboard, click wheel, buttons, etc. These buttons can include, but are not limited to: a home button, volume buttons, a power button, and a lock button.
[0176] The sensor assembly 814 includes one or more sensors for providing an assessment of various aspects of the status of the electronic device 800. For example, the sensor assembly 814 can detect the on / off state of the electronic device 800, the relative positioning of components, such as the display and keypad of the electronic device 800. The sensor assembly 814 can also detect a change in the position of the electronic device 800 or a component of the electronic device 800, the presence or absence of user contact with the electronic device 800, the orientation or acceleration / deceleration of the electronic device 800, and the temperature change of the electronic device 800. The sensor assembly 814 can include a proximity sensor configured to detect the presence of nearby objects without any physical contact. The sensor assembly 814 can also include a light sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, the sensor assembly 814 can also include an acceleration sensor, a gyroscope sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.
[0177] The communication component 816 is configured to facilitate communication between the electronic device 800 and other devices in a wired or wireless manner. The electronic device 800 can access a wireless network based on communication standards, such as WiFi, 4G, or 5G, or a combination thereof. In an exemplary embodiment, the communication component 816 receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 816 further includes a near field communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on radio frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.
[0178] In an exemplary embodiment, the electronic device 800 can be implemented by one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components for performing the above method.
[0179] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as a memory 804 including instructions that can be executed by a processor 820 of the electronic device 800 to complete the above method. For example, the non-transitory computer-readable storage medium can be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device, etc.
[0180] In the description of this specification, the description referring to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples", etc. means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of this application. In this specification, the schematic expressions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.
[0181] In addition, the terms "first" and "second" are used only for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one of such features. In the description of this application, "a plurality of" means at least two, such as two, three, etc., unless otherwise specifically defined.
[0182] Any process or method description shown in the flowchart or described in other ways herein can be understood as representing a module, segment, or part of code including one or more executable instructions for implementing a customized logical function or process, and the scope of the preferred embodiments of this application includes additional implementations, where the functions can be executed in a substantially simultaneous manner or in the reverse order according to the functions involved, rather than in the order shown or discussed, which should be understood by those skilled in the art to which the embodiments of this application belong.
[0183] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a definitional sequence list of executable instructions for implementing logical functions, which can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, apparatus, or device), or in conjunction with these instruction execution systems, apparatus, or devices. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device. More specific examples (non-exhaustive list) of computer-readable media include the following: an electrical connection portion with one or more wirings (electronic device), a portable computer diskette (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). Additionally, the computer-readable medium can even be paper or other suitable media on which the program can be printed, because the program can be obtained electronically, for example, by optically scanning the paper or other media, followed by editing, interpretation, or other suitable processing as necessary, and then stored in a computer memory.
[0184] It should be understood that various parts of the present application can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), etc.
[0185] Those of ordinary skill in the art of this technology can understand that all or part of the steps carried by the method of the above embodiments can be completed by instructing relevant hardware through a program, and the program can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiments.
[0186] In addition, each functional unit in various embodiments of the present application may be integrated into one processing module, may exist separately physically for each unit, or two or more units may be integrated into one module. The above-mentioned integrated module may be implemented in the form of hardware or in the form of a software functional module. When the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it may also be stored in a computer-readable storage medium.
[0187] The above-mentioned storage medium may be a read-only memory, a magnetic disk or an optical disc, etc. Although the embodiments of the present application have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present application. Those of ordinary skill in the art can make changes, modifications, substitutions and variations to the above embodiments within the scope of the present application.
Claims
1. An image segmentation method, characterized in that, Including: Obtain the image to be segmented; Use the trained segmentation model to perform segmentation and scanning processing on the features of the image to be segmented, and obtain the target feature map of the image to be segmented; wherein, the target feature map includes the dependency relationship between adjacent regions in the image to be segmented; Perform image segmentation according to the target feature map to obtain an image segmentation result.
2. The method according to claim 1, wherein The segmentation model includes an encoder and a decoder. The step of using the trained segmentation model to perform segmentation and scanning processing on the features of the image to be segmented to obtain the target feature map of the image to be segmented includes: Use the encoder to perform segmentation and scanning processing on the features of the image to be segmented to obtain the feature map output by the encoder; Input the feature map output by the encoder into the decoder for segmentation and scanning processing of the features to obtain the target feature map of the image to be segmented.
3. The method according to claim 2, wherein The encoder includes L cascaded sub-encoders, and each sub-encoder includes a first segmentation and scanning module. The step of using the encoder to perform segmentation and scanning processing on the features of the image to be segmented to obtain the feature map output by the encoder includes: For any one of the L sub-encoders, determine the first feature map to be processed corresponding to the first segmentation and scanning module of the any one of the sub-encoders according to the position of the any one of the sub-encoders in the L sub-encoders; wherein, the first feature map to be processed is determined according to the feature map of the image to be segmented or the feature map output by the previous sub-encoder of the any one of the sub-encoders; Use the first segmentation and scanning module of the any one of the sub-encoders to perform segmentation and scanning processing on the first feature map to be processed to obtain the feature map output by the first segmentation and scanning module of the any one of the sub-encoders; Determine the feature map output by the any one of the sub-encoders according to the feature map output by the first segmentation and scanning module of the any one of the sub-encoders.
4. The method according to claim 3, wherein The decoder includes L cascaded sub-decoders, and each sub-decoder module includes a second segmentation and scanning module. The L sub-encoders in the encoder and the L sub-decoders in the decoder are connected in a skip connection. The step of inputting the feature map output by the encoder into the decoder for segmentation and scanning processing of the features to obtain the target feature map of the image to be segmented includes: For the m-th sub-decoder among the L sub-decoders, determine the second feature map to be processed corresponding to the second segmentation and scanning module of the m-th sub-decoder according to the position of the m-th sub-decoder in the L sub-decoders; wherein, the second feature map to be processed is determined according to the feature map output by the (L - m + 1)-th sub-encoder, or a fused feature map obtained by fusing the feature map output by the (L - m + 1)-th sub-encoder and the feature map output by the previous sub-decoder of the m-th sub-decoder; Use the second segmentation and scanning module of the m-th sub-decoder to perform segmentation and scanning processing on the second feature map to be processed to determine the feature map output by the m-th sub-decoder; Determine the target feature map according to the feature map of the image to be segmented and the feature map output by the last sub-decoder.
5. The method according to claim 3, wherein The first segmentation and scanning module of any one of the sub-encoders includes a feature segmentation module, a feature scanning module, a feature reconstruction module, and a feature fusion module. Using the first segmentation and scanning module of any one of the sub-encoders to perform feature segmentation and scanning processing on the first feature map to be processed, and obtaining the feature map output by the first segmentation and scanning module of any one of the sub-encoders includes: Using the feature segmentation module to perform feature segmentation operations on the first feature map to be processed in multiple preset directions respectively, and obtaining segmentation feature sequences in each preset direction; Using the feature scanning module to perform feature aggregation on the segmentation feature sequences in each preset direction by using corresponding scanning methods, and obtaining scanning feature sequences in each preset direction; Using the feature reconstruction module to perform temporal feature extraction on the scanning feature sequences in each preset direction, and obtaining temporal feature sequences in each preset direction; Using the feature fusion module to perform feature fusion and restoration on the temporal feature sequences in multiple preset directions, and obtaining the feature map output by the first segmentation and scanning module of any one of the sub-encoders.
6. The method according to claim 5, characterized in that, The multiple preset directions include the horizontal direction and the vertical direction. Using the feature segmentation module to perform feature segmentation operations on the first feature map to be processed in multiple preset directions respectively, and obtaining segmentation feature sequences in each preset direction includes: Using the first segmentation sub-module of the feature segmentation module to perform feature segmentation operations on the first feature map to be processed in the horizontal direction, and obtaining a segmentation feature sequence in the horizontal direction; Using the second segmentation sub-module of the feature segmentation module to perform feature segmentation operations on the first feature map to be processed in the vertical direction, and obtaining a segmentation feature sequence in the vertical direction.
7. The method according to claim 5, wherein Using the feature scanning module to perform feature aggregation on the segmentation feature sequences in each preset direction by using corresponding scanning methods, and obtaining scanning feature sequences in each preset direction includes: For the segmentation feature sequence in the horizontal direction, using the first sub-scanning module of the feature scanning module to perform bottom-up scanning and top-down scanning in the vertical direction respectively, and aggregating to obtain a first sub-scanning feature sequence and a second sub-scanning feature sequence; For the segmentation feature sequence in the vertical direction, using the second sub-scanning module of the feature scanning module to perform left-to-right scanning and right-to-left scanning in the horizontal direction respectively, and aggregating to obtain a third sub-scanning feature sequence and a fourth sub-scanning feature sequence.
8. A model training method, characterized in that, Includes: Obtaining training sample images; Using a segmentation model to perform feature segmentation and scanning processing on the training sample images, and obtaining the target feature maps of the training sample images; wherein, the target feature maps include the dependency relationships between adjacent regions in the training sample images; Performing image segmentation according to the target feature maps to obtain an image prediction segmentation result; Determining a loss function according to the difference between the image prediction segmentation result and the true segmentation result corresponding to the training sample images; Adjusting the parameters of the segmentation model according to the loss function to obtain a trained segmentation model.
9. An image segmentation device, characterized in that, Includes: An acquisition module for acquiring an image to be segmented; A first processing module, configured to perform segmentation and scanning processing on features of the image to be segmented by using a trained segmentation model, so as to obtain a target feature map of the image to be segmented; wherein, the target feature map includes the dependency relationship between adjacent regions in the image to be segmented; A second processing module, configured to perform image segmentation according to the target feature map to obtain an image segmentation result.
10. A model training device, characterized in that, Comprising: An acquisition module, configured to acquire a training sample image; A first processing module, configured to perform segmentation and scanning processing on features of the training sample image by using a segmentation model, so as to obtain a target feature map of the training sample image; wherein, the target feature map includes the dependency relationship between adjacent regions in the training sample image; A second processing module, configured to perform image segmentation according to the target feature map to obtain an image predicted segmentation result; A determination module, configured to determine a loss function according to the difference between the image predicted segmentation result and the true value segmentation result corresponding to the training sample image; A parameter adjustment module, configured to adjust parameters of the segmentation model according to the loss function, so as to obtain a trained segmentation model.
11. An electronic device, characterized in that, Comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, when the processor executes the program, implementing the method according to any one of claims 1-7, or implementing the method according to claim 8.
12. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the method according to any one of claims 1-7, or implements the method according to claim 8.
13. A computer program product, characterized in that, Comprising a computer program, when the computer program is executed by the processor, implementing the method according to any one of claims 1-7, or implementing the method according to claim 8.