Image segmentation method, device, storage medium, and computer program
By designing an image segmentation network, including a global feature extraction module, a local feature extraction module and a feature fusion module, the problem of difficulty in obtaining global and local feature information at the same time in the prior art is solved, and effective image segmentation and feature extraction are achieved.
Patent Information
- Application Number
- CN202111163525.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-09-30
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2041-09-30
AI Technical Summary
In remote sensing segmentation tasks, it is difficult for the prior art to obtain the global feature information and local feature information of the image at the same time, resulting in the loss of detailed information of key objects, and it is impossible to focus on key objects in the image and obtain global information.
An image segmentation method is designed by inputting the image to be processed into an image segmentation network, which includes a global feature extraction module, a local feature extraction module and a feature fusion module. The global feature extraction module adopts a neural network, the local feature extraction module adopts a full attention network, and the feature fusion module adopts a cross network, through which features are extracted and fused, generate semantic segmentation results containing global and local feature information.
It realizes the acquisition of global feature information and local feature information of the image simultaneously, reduces the loss of detailed information of key objects, and can effectively focus key objects in the image and obtain global information.
Smart Images

Figure CN114299073B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing, and in particular, to an image segmentation method, apparatus, storage medium, and computer program. Background Art
[0002] In remote sensing segmentation tasks, it is necessary to extract the feature information of an image and perform object segmentation. However, the image segmentation networks in related technologies are usually network structures that reduce the image resolution, and it is difficult to better implement remote sensing segmentation tasks.
[0003] To solve the above problems, image segmentation networks that can obtain features with higher resolution have emerged in related technologies. However, they can only obtain local information and cannot focus on key objects in the image and obtain global information. Therefore, in remote sensing segmentation tasks, how to reduce the loss of detailed information of key objects while retaining global information is an urgent problem to be solved.
[0004] For the above problems, no effective solutions have been proposed yet. Summary of the Invention
[0005] Embodiments of the present invention provide an image segmentation method, apparatus, storage medium, and computer program to at least solve the technical problem in related technologies that it is difficult to simultaneously obtain the global feature information and local feature information of an image in an image segmentation task.
[0006] According to one aspect of the embodiments of the present invention, an image segmentation method is provided, including: inputting an image to be processed into an image segmentation network, where the image segmentation network includes a global feature extraction module, a local feature extraction module, and a feature fusion module commonly connected to the output ends of the global feature extraction module and the local feature extraction module; processing the image to be processed through the image segmentation network to obtain a semantic segmentation result of the image to be processed, where the semantic segmentation result includes global feature information and local feature information.
[0007] According to another aspect of the embodiments of the present invention, an image segmentation method is further provided, including: inputting a remote sensing image into an image segmentation network, where the image segmentation network includes a global feature extraction module, a local feature extraction module, and a feature fusion module commonly connected to the output ends of the global feature extraction module and the local feature extraction module; processing the remote sensing image through the image segmentation network to obtain a semantic segmentation result of the remote sensing image, where the semantic segmentation result includes the global feature information and local feature information of each ground object.
[0008] According to another aspect of the embodiments of the present invention, there is also provided an image segmentation method, including: inputting a building image into an image segmentation network, where the image segmentation network includes a global feature extraction module and a local feature extraction module, and a feature fusion module commonly connected to the output ends of the global feature extraction module and the local feature extraction module; processing the building image through the image segmentation network to obtain a semantic segmentation result of the building image, where the semantic segmentation result includes global feature information of each building and local feature information of each building.
[0009] According to another aspect of the embodiments of the present invention, there is also provided an image processing apparatus, including: an input unit configured to input an image to be processed into an image segmentation network, where the image segmentation network includes a global feature extraction module and a local feature extraction module, and a feature fusion module commonly connected to the output ends of the global feature extraction module and the local feature extraction module; a processing unit configured to process the image to be processed through the image segmentation network to obtain a semantic segmentation result of the image to be processed, where the semantic segmentation result includes global feature information and local feature information.
[0010] According to another aspect of the embodiments of the present invention, there is also provided a storage medium, where the storage medium includes a stored program, and when the program runs, it controls the device where the storage medium is located to execute the image segmentation method of any one of the above.
[0011] According to another aspect of the embodiments of the present invention, there is also provided a computer program, where when the computer program runs, it executes the image segmentation method of any one of the above.
[0012] In the embodiments of the present invention, an image to be processed is input into an image segmentation network, where the image segmentation network includes a global feature extraction module and a local feature extraction module, and a feature fusion module commonly connected to the output ends of the global feature extraction module and the local feature extraction module; the image to be processed is processed through the image segmentation network to obtain a semantic segmentation result of the image to be processed, where the semantic segmentation result includes global feature information and local feature information. By extracting the features of the image to be processed through the global feature extraction module and the local feature extraction module, and fusing the extracted features through the feature fusion module, the purpose that the semantic segmentation result output by the image segmentation network includes global feature information and local feature information is achieved, thereby realizing the technical effect of simultaneously obtaining the global feature information and local feature information of the image, and further solving the technical problem in the related art that it is difficult to simultaneously obtain the global feature information and local feature information in the image segmentation task. Description of the Drawings
[0013] The accompanying drawings described herein are used to provide a further understanding of the present invention and form a part of this application. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and shall not unduly limit the present invention. In the drawings:
[0014] Figure 1 is a block diagram of the hardware structure of a computer terminal according to an embodiment of the present invention;
[0015] Figure 2 is a flowchart of an image segmentation method provided in Embodiment 1 of the present invention;
[0016] Figure 3 is a schematic diagram of an optional image segmentation method provided in Embodiment 1 of the present invention;
[0017] Figure 4 is a flowchart of an image segmentation method provided in Embodiment 2 of the present invention;
[0018] Figure 5 is a flowchart of an image segmentation method provided in Embodiment 3 of the present invention;
[0019] Figure 6 is a schematic diagram of an image segmentation device provided in Embodiment 4 of the present invention;
[0020] Figure 7 is a schematic diagram of an image segmentation device provided in Embodiment 5 of the present invention;
[0021] Figure 8 is a schematic diagram of an image segmentation device provided in Embodiment 6 of the present invention;
[0022] Figure 9 is a block diagram of an optional computer terminal according to an embodiment of the present invention. Detailed Embodiments
[0023] In order to enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0024] It should be noted that the terms "first", "second", etc. in the description, claims and above-mentioned drawings of the present invention are used to distinguish similar objects, and do not necessarily have to be used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present invention described here can be implemented in an order different from those illustrated or described here. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0025] Embodiment 1
[0026] According to an embodiment of the present invention, an embodiment of an image segmentation method is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that here.
[0027] The method embodiment provided by the first embodiment of the present application can be executed in a mobile terminal, a computer terminal or a similar computing device. Figure 1 The hardware structure block diagram of a computer terminal (or mobile device) for implementing the image segmentation method is shown. As Figure 1 shown, the computer terminal 101 (or mobile device 101) may include one or more processors 102 (illustrated by 102a, 102b,..., 102n in the figure) (the processor 102 may include, but is not limited to, a processing device such as a microprocessor MCU or a programmable logic device FPGA), a memory 104 for storing data, and a transmission device for communication functions. In addition to this, it may further include: a display, an input / output interface (I / O interface), a universal serial bus (USB) port (which may be included as one of the ports of the I / O interface), a network interface, a power supply and / or a camera. Those of ordinary skill in the art can understand that Figure 1 the structure shown is only schematic and does not limit the structure of the above-mentioned electronic device. For example, the computer terminal 101 may further include more or fewer components than those Figure 1 shown, or have a different configuration from that Figure 1 shown.
[0028] It should be noted that the above one or more processors 102 and / or other data processing circuits can generally be referred to as "data processing circuits" herein. The data processing circuit can be embodied in software, hardware, firmware, or any combination thereof, in whole or in part. In addition, the data processing circuit can be a single independent processing module, or be incorporated in whole or in part into any one of other elements in the computer terminal 101 (or mobile device). As involved in the embodiments of the present application, the data processing circuit is a kind of processor control (such as the selection of a variable resistance terminal path connected to an interface).
[0029] The memory 104 can be used to store software programs and modules of application software, such as the program instructions / data storage devices corresponding to the () method in the embodiments of the present invention. The processor 102 executes various functional applications and data processing by running the software programs and modules stored in the memory 104, that is, implements the () method of the above application program. The memory 104 can include a high-speed random access memory, and can also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memories. In some instances, the memory 104 can further include a memory remotely disposed relative to the processor 102, and these remote memories can be connected to the computer terminal 101 through a network. Examples of the above network include but are not limited to the Internet, intranet, local area network, mobile communication network, and combinations thereof.
[0030] The transmission device 106 is used to receive or send data via a network. Specific examples of the above network can include the wireless network provided by the communication provider of the computer terminal 101. In one instance, the transmission device 106 includes a network adapter (Network Interface Controller, NIC), which can be connected to other network devices through a base station and thus communicate with the Internet. In one instance, the transmission device 106 can be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.
[0031] The display can be, for example, a touch-screen liquid crystal display (LCD), which enables the user to interact with the user interface of the computer terminal 101 (or mobile device).
[0032] Under the above operating environment, the present application provides an image segmentation method as shown in Figure 2 shown below. Figure 2 It is a flowchart of the image segmentation method according to Embodiment 1 of the present invention.
[0033] S21. Input the image to be processed into an image segmentation network. The image segmentation network includes a global feature extraction module, a local feature extraction module, and a feature fusion module that is commonly connected to the output ends of the global feature extraction module and the local feature extraction module.
[0034] Specifically, the image to be processed can be a map in a remote sensing segmentation task. For example, an agricultural and forestry map or an urban landform map.
[0035] The structure adopted by the global feature extraction module can be a neural network. Through the neural network, local features of the image can be obtained. The structure adopted by the local feature extraction module is a full attention network. Through the full attention network, global features of the image can be obtained and important objects in the image can be focused. The structure adopted by the feature fusion module can be a cross network, which can perform cross combination on features.
[0036] S22. Process the image to be processed through the image segmentation network to obtain the semantic segmentation result of the image to be processed. The semantic segmentation result contains global feature information and local feature information.
[0037] Specifically, the image segmentation network is a pre-trained model. The model parameters of the global feature extraction module, the local feature extraction module, and the feature fusion module have been determined. Inputting the image to be processed into the image segmentation network can obtain the semantic segmentation result.
[0038] In an optional implementation manner, processing the image to be processed through the image segmentation network to obtain the semantic segmentation result of the image to be processed includes: processing the image to be processed through the global feature extraction module to obtain the global feature information of the image to be processed; processing the image to be processed through the local feature extraction module to obtain the local feature information of the image to be processed; and fusing the global feature information and the local feature information through the feature fusion module to obtain the semantic segmentation result of the image to be processed.
[0039] Specifically, the image to be processed can be processed into images with different resolutions. The global feature extraction module processes the low-resolution image to be processed to efficiently extract the global feature information of the image to be processed, and the local feature extraction module processes the high-resolution image to be processed to extract the local feature information of the image to be processed.
[0040] For example, when the image to be processed is a map, the local feature extraction module can identify the detailed feature information of important objects such as small buildings and vehicles from the map, and the global feature extraction module can obtain the global feature information such as large areas of forests, lakes, and lawns.
[0041] Further, the feature fusion module cross - combines the global feature information and local features of the image to be processed, and the obtained semantic segmentation result contains both global feature information and local feature information.
[0042] In an embodiment of the present invention, the image to be processed is input into an image segmentation network. The image segmentation network includes a global feature extraction module, a local feature extraction module, and a feature fusion module jointly connected to the output ends of the global feature extraction module and the local feature extraction module. The image to be processed is processed by the image segmentation network to obtain a semantic segmentation result of the image to be processed, where the semantic segmentation result contains global feature information and local feature information. By extracting the features of the image to be processed through the global feature extraction module and the local feature extraction module, and fusing the extracted features through the feature fusion module, the purpose that the semantic segmentation result output by the image segmentation network contains global feature information and local feature information is achieved, thereby realizing the technical effect of simultaneously obtaining the global feature information and local feature information of the image, and further solving the technical problem in the related art that it is difficult to simultaneously obtain the global feature information and local feature information in the image segmentation task.
[0043] In an alternative embodiment, before inputting the image to be processed into the image segmentation network, the method further includes: constructing an image segmentation network, where the image segmentation network includes a global feature extraction module, a local feature extraction module, and a feature fusion module jointly connected to the output ends of the global feature extraction module and the local feature extraction module.
[0044] Specifically, the global feature extraction module and the local feature extraction module are parallel feature extraction modules, which respectively extract features from the image to be processed, and then perform feature fusion through the feature fusion module.
[0045] Among them, the structure adopted by the global feature extraction module can be a neural network. For example, the neural network is a convolutional neural network or a recurrent neural network. Both the convolutional neural network and the recurrent neural network can be used to extract the local information of the image to be processed, thereby avoiding the loss of detailed information.
[0046] Among them, the structure adopted by the local feature extraction module can be a full attention network (transformer). The full attention network includes a self-attention network (Self-Attention) and a feedforward neural network connected in sequence. The self-attention network focuses on important objects in the image to be processed through the self-attention mechanism. Specifically, it can be an efficient self-attention network (Efficient-Self-Attention), so as to more efficiently focus on important objects. The classification ability and pattern recognition ability of the feedforward neural network are superior to those of the feedback neural network, which improves the performance of feature extraction. Specifically, an N×N convolutional layer and a multi-layer perceptron can be added to the feedforward neural network, so as to improve the feature recognition rate and classification speed. In addition, the full attention network can also include a component for implementing the operation of reducing the size of the feature map, so as to reduce the size of the feature map and increase the number of channels of the feature map.
[0047] Among them, the feature fusion module can be a cross network, which cross-combines the input features through the cross network to achieve feature fusion. For example, the cross network can be a deep cross network.
[0048] In an optional implementation manner, the global feature extraction module includes a plurality of globally connected global feature extraction components in series, and the local feature extraction module includes a plurality of locally connected local feature extraction components in series. The output features of the last component of each global feature extraction module and the last component of each local feature extraction module are jointly used as the input features of the feature fusion module.
[0049] Specifically, both the global feature extraction module and the local feature extraction module can be one or more. Each feature extraction module can extract the features of a certain resolution of the image to be processed. Each feature extraction module is connected in parallel, and the internal of the feature extraction module includes a plurality of feature extraction components, and the plurality of feature extraction components inside the feature extraction module are connected in series.
[0050] For example, Figure 3 is a schematic diagram of an optional image segmentation method provided according to Embodiment 1 of the present invention, as Figure 3As shown, the image segmentation network may include 4 feature extraction modules. Among them, the first feature extraction module and the second feature extraction module are both local feature extraction components, and the third feature extraction module and the fourth feature extraction module are both global feature extraction components. The resolution of the features extracted by the first feature extraction module is the same as that of the original image. The resolution of the features extracted by the second feature extraction module is one-fourth of the original image. The resolution of the features extracted by the third feature extraction module is one-eighth of the original image. The resolution of the features extracted by the fourth feature extraction module is one-thirty-second of the original image, so as to obtain features of various resolutions of the image to be processed.
[0051] Furthermore, input the features of various resolutions of the image to be processed into the feature fusion component. The feature fusion component can be a deep cross network, which can cross-combine the features of various resolutions and output the processed image. The processed image contains the fused features of each resolution of the image to be processed.
[0052] It should be noted that in order to fuse features during the feature extraction process, the output features of each feature extraction component can be used not only as the input of the next feature extraction component of this module but also as the input of the feature extraction components of other modules.
[0053] In an alternative embodiment, the input of the first feature extraction component in the first feature extraction module is the image to be processed. The input of the first feature extraction component in other feature extraction modules includes the image to be processed and the output features of a feature extraction component in the previous feature extraction module. At the same time, there is a cross relationship between the feature extraction components other than the first feature extraction component in different feature extraction modules. Specifically, the output features of the target feature extraction component in a feature extraction module are used not only as the input features of the next feature extraction component in this module but also as the input features of the next feature extraction component in other modules. Among them, the position of the target feature extraction component can be flexibly set in the feature extraction module.
[0054] For example, as Figure 3As shown, the image segmentation network may include four feature extraction modules. The input of the first feature extraction component in the first feature extraction module is the image to be processed, and the input of other feature extraction modules includes at least the output features of the previous feature extraction component. The input of the first feature extraction component in the second feature extraction module includes the image to be processed and the output features of the fifth feature extraction component in the first feature extraction module, and the input of other feature extraction components includes at least the output features of the previous feature extraction component. The input of the first feature extraction component in the third feature extraction module includes the image to be processed and the output features of the fifth feature extraction component in the second feature extraction module, and the input of other feature extraction components includes at least the output features of the previous feature extraction component. The input of the first feature extraction component in the fourth feature extraction module includes the image to be processed and the output features of the fifth feature extraction component in the third feature extraction module, and the input of other feature extraction components includes at least the output features of the previous feature extraction component.
[0055] Through this embodiment, each feature extraction module can not only extract feature information from the original image to be processed, but also further extract features from the features extracted by other feature extraction modules, improving the efficiency of feature acquisition without losing detailed information.
[0056] In an optional implementation manner, in order to intuitively know the segmentation situation of the image to be processed, after the image to be processed is processed by the image segmentation network to obtain the semantic segmentation result of the image to be processed, the method further includes: displaying the semantic segmentation result of the image to be processed, and marking the global image area corresponding to the global feature and the local image area corresponding to the local feature.
[0057] Specifically, the semantic segmentation result of the image to be processed is the classification situation of the high-dimensional features. The classification situation of the high-dimensional features can be converted into a visual semantic segmentation result, and the global image area corresponding to the global feature and the local image area corresponding to the local feature are marked in the visual semantic segmentation result, so that the user can intuitively know the segmentation situation of the image to be processed.
[0058] The user can also adjust the semantic segmentation result. In an optional implementation manner, after displaying the semantic segmentation result of the image to be processed and marking the global image area corresponding to the global feature and the local image area corresponding to the local feature, the method further includes: obtaining adjustment information for the global image area and / or the local image area; adjusting the global image area and / or the local image area according to the adjustment information to obtain an updated semantic segmentation result.
[0059] Specifically, there are cases where the semantic segmentation result of the image segmentation network does not meet the user's requirements. After converting the semantic segmentation result into a visual semantic segmentation result, the user can adjust the marked global image area and local image area to obtain an updated visual semantic segmentation result, and perform image segmentation based on the updated visual semantic segmentation result.
[0060] In order to enable the user to know the changes in the key area, in an optional implementation manner, after processing the image to be processed through the image segmentation network to obtain the semantic segmentation result of the image to be processed, the method further includes: during the change process of the image to be processed, correspondingly display the target area marked in the semantic segmentation result.
[0061] Specifically, the target area can be the key area or object that the user is concerned about. For example, when the target area is a flying object, the flying object can be marked in the visual semantic segmentation result, and during the change process of the image to be processed, dynamically display the semantic segmentation results corresponding to the image to be processed before and after the change, so that the user can know the changes in the flying object.
[0062] Embodiment 2
[0063] According to an embodiment of the present invention, an embodiment of an image segmentation method is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.
[0064] This application provides an image segmentation method as Figure 4 shown. Figure 4 is a flowchart of the image segmentation method according to the second embodiment of the present invention.
[0065] S41, input the remote sensing image into the image segmentation network, where the image segmentation network includes a global feature extraction module, a local feature extraction module, and a feature fusion module commonly connected to the output ends of the global feature extraction module and the local feature extraction module.
[0066] It should be noted that the remote sensing image is an aerial photograph and a satellite photo. Specifically, it can be a meteorological forecast image, a natural resource image, a water conservancy, agriculture and forestry image, and a building image.
[0067] The structure adopted by the global feature extraction module in the image segmentation network can be a neural network. Through the neural network, local features of the image can be obtained. The structure adopted by the local feature extraction module is a full attention network. Through the full attention network, global features of the image can be obtained and important objects in the image can be focused. The structure adopted by the feature fusion module can be an intersection network, which can perform cross-combination on the features.
[0068] S42. Process the remote sensing image through the image segmentation network to obtain the semantic segmentation result of the remote sensing image. Among them, the semantic segmentation result contains the global feature information and local feature information of each ground object.
[0069] Specifically, the image to be processed can be processed into images with different resolutions. The global feature extraction module processes the low-resolution image to be processed to efficiently extract the global feature information of the image to be processed. The local feature extraction module processes the high-resolution image to be processed to extract the local feature information of the image to be processed.
[0070] For example, in the case where the remote sensing image is a map of urban landforms, the local feature extraction module can identify the detailed feature information of important objects such as small buildings and vehicles from the map, and the global feature extraction module can obtain the global feature information such as lakes and lawns.
[0071] In the embodiment of the present invention, the remote sensing image is input into the image segmentation network. Among them, the image segmentation network includes a global feature extraction module, a local feature extraction module, and a feature fusion module jointly connected to the output ends of the global feature extraction module and the local feature extraction module. The remote sensing image is processed through the image segmentation network to obtain the semantic segmentation result of the remote sensing image. Among them, the semantic segmentation result contains the global feature information and local feature information of each ground object. The features of the image to be processed are extracted through the global feature extraction module and the local feature extraction module, and the extracted features are fused through the feature fusion module, achieving the purpose that the semantic segmentation result output by the image segmentation network contains global feature information and local feature information, thereby realizing the technical effect of simultaneously obtaining the global feature information and local feature information of the image, and further solving the technical problem in the related art that it is difficult to simultaneously obtain the global feature information and local feature information in the image segmentation task.
[0072] In an alternative embodiment, after processing the remote sensing image through the image segmentation network to obtain the semantic segmentation result of the remote sensing image, the method further includes: displaying the regional map of each ground object in the semantic segmentation result of the remote sensing image under the first receptive field; displaying the local information of the target ground object in the semantic segmentation result of the remote sensing image under the second receptive field, where the second receptive field is larger than the first receptive field.
[0073] It should be noted that the semantic segmentation result of the image to be processed is the classification of high-dimensional features. The classification of high-dimensional features can be converted into a visual semantic segmentation result, and the global image area corresponding to the global feature and the local image area corresponding to the local feature can be marked in the visual semantic segmentation result, so that the user can intuitively know the segmentation situation of the image to be processed.
[0074] Specifically, after converting the classification of high-dimensional features into a visual semantic segmentation result, the global information and local information in the semantic segmentation result are displayed under different receptive fields. The global information of the remote sensing image is displayed under a smaller first receptive field, that is, the regional map of each ground object in the remote sensing image is displayed. The local information of the remote sensing image is displayed under a larger second receptive field, that is, the local information of the target ground object in the remote sensing image is displayed, so that the user can intuitively know the global information and local information of the processed image.
[0075] Embodiment 3
[0076] According to an embodiment of the present invention, an embodiment of an image segmentation method is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. And, although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.
[0077] This application provides an Figure 5 image segmentation method as shown. Figure 5 It is a flowchart of the image segmentation method according to Embodiment 3 of the present invention.
[0078] S51, input the building image into the image segmentation network, where the image segmentation network includes a global feature extraction module, a local feature extraction module, and a feature fusion module jointly connected to the output ends of the global feature extraction module and the local feature extraction module.
[0079] Specifically, the building image can be an urban building map, a garden map, or a village map. The structure adopted by the global feature extraction module in the image segmentation network can be a neural network, and the local features of the image can be obtained through the neural network. The structure adopted by the local feature extraction module is a full attention network, and the global features of the image can be obtained through the full attention network and the important objects in the image can be focused. The structure adopted by the feature fusion module can be a cross network, which can perform cross combination on the features.
[0080] S52, process the building image through the image segmentation network to obtain the semantic segmentation result of the building image, where the semantic segmentation result includes the global feature information of each building and the local feature information of each building.
[0081] Specifically, the image to be processed can be processed into images with different resolutions. The global feature extraction module processes the low-resolution image to be processed to efficiently extract the global feature information of the image to be processed, and the local feature extraction module processes the high-resolution image to be processed, thereby extracting the local feature information of the image to be processed.
[0082] For example, when the building image is a garden map, the local feature extraction module can identify the detailed feature information of important objects such as pavilions and towers from the map, and the global feature extraction module can obtain the global feature information such as lakes and vegetation.
[0083] In the embodiment of the present invention, the building image is input into an image segmentation network. The image segmentation network includes a global feature extraction module, a local feature extraction module, and a feature fusion module jointly connected to the output ends of the global feature extraction module and the local feature extraction module. The building image is processed by the image segmentation network to obtain the semantic segmentation result of the building image, where the semantic segmentation result contains the global feature information of each building and the local feature information of each building. The features of the image to be processed are extracted by the global feature extraction module and the local feature extraction module, and the extracted features are fused by the feature fusion module, achieving the purpose that the semantic segmentation result output by the image segmentation network contains global feature information and local feature information, thereby realizing the technical effect of simultaneously obtaining the global feature information and local feature information of the image, and further solving the technical problem in the related art that it is difficult to simultaneously obtain the global feature information and local feature information in the image segmentation task.
[0084] It should be noted that for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the present invention is not limited by the described action sequence, because according to the present invention, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the present invention.
[0085] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases the former is a better implementation. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions for causing a terminal device (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods of the various embodiments of the present invention.
[0086] Embodiment 4
[0087] According to an embodiment of the present invention, there is also provided an image segmentation device for implementing the above image segmentation method, as Figure 6 shown. The device includes:
[0088] An input unit 10 for inputting an image to be processed into an image segmentation network, where the image segmentation network includes a global feature extraction module and a local feature extraction module, and a feature fusion module commonly connected to the output ends of the global feature extraction module and the local feature extraction module.
[0089] A processing unit 20 for processing the image to be processed through the image segmentation network to obtain a semantic segmentation result of the image to be processed, where the semantic segmentation result includes global feature information and local feature information.
[0090] In the embodiment of the present invention, the input unit 10 inputs the image to be processed into the image segmentation network, where the image segmentation network includes a global feature extraction module and a local feature extraction module, and a feature fusion module commonly connected to the output ends of the global feature extraction module and the local feature extraction module; the processing unit 20 processes the image to be processed through the image segmentation network to obtain a semantic segmentation result of the image to be processed, where the semantic segmentation result includes global feature information and local feature information. By extracting the features of the image to be processed through the global feature extraction module and the local feature extraction module, and fusing the extracted features through the feature fusion module, the purpose that the semantic segmentation result output by the image segmentation network includes global feature information and local feature information is achieved, thereby realizing the technical effect of simultaneously obtaining the global feature information and local feature information of the image, and further solving the technical problem in the related art that it is difficult to simultaneously obtain the global feature information and local feature information in the image segmentation task.
[0091] In an alternative embodiment, the processing unit 20 includes: a processing module for processing the image to be processed through a global feature extraction module to obtain the global feature information of the image to be processed; a processing module for processing the image to be processed through a local feature extraction module to obtain the local feature information of the image to be processed; and a fusion module for fusing the global feature information and the local feature information through a feature fusion module to obtain the semantic segmentation result of the image to be processed.
[0092] In an alternative embodiment, the apparatus further includes: a construction unit for constructing an image segmentation network, where the image segmentation network includes a global feature extraction module and a local feature extraction module, and a feature fusion module jointly connected to the output ends of the global feature extraction module and the local feature extraction module.
[0093] In an alternative embodiment, the global feature extraction module includes a plurality of globally connected global feature extraction components in series, and the local feature extraction module includes a plurality of locally connected local feature extraction components in series. The output features of the last component of each global feature extraction module and the last component of each local feature extraction module are jointly used as the input features of the feature fusion module.
[0094] In an alternative embodiment, the apparatus further includes: a display unit for displaying the semantic segmentation result of the image to be processed, and annotating the global image region corresponding to the global feature and the local image region corresponding to the local feature.
[0095] In an alternative embodiment, the apparatus further includes: an adjustment unit, which includes: an acquisition module for acquiring adjustment information for the global image region and / or the local image region; and an adjustment module for adjusting the global image region and / or the local image region according to the adjustment information to obtain an updated semantic segmentation result.
[0096] In an alternative embodiment, the apparatus further includes: a display unit for correspondingly displaying the target region annotated in the semantic segmentation result during the change process of the image to be processed.
[0097] It should be noted here that the above units correspond to the steps in Embodiment 1. The examples and application scenarios implemented by the two units and the corresponding steps are the same, but are not limited to the content disclosed in Embodiment 1 above. It should be noted that the above modules, as part of the apparatus, can run in the computer terminal provided in Embodiment 1.
[0098] Embodiment 5
[0099] According to an embodiment of the present invention, there is also provided an image segmentation apparatus for implementing the above image segmentation method, as Figure 7 shown, the apparatus includes:
[0100] A segmentation unit 30 for inputting a remote sensing image into an image segmentation network, where the image segmentation network includes a global feature extraction module, a local feature extraction module, and a feature fusion module commonly connected to the output ends of the global feature extraction module and the local feature extraction module.
[0101] A processing unit 40 for processing the remote sensing image through the image segmentation network to obtain a semantic segmentation result of the remote sensing image, where the semantic segmentation result includes global feature information and local feature information of each ground object.
[0102] In an embodiment of the present invention, the remote sensing image is input into the image segmentation network through the segmentation unit 30, where the image segmentation network includes a global feature extraction module, a local feature extraction module, and a feature fusion module commonly connected to the output ends of the global feature extraction module and the local feature extraction module; the processing unit 40 processes the remote sensing image through the image segmentation network to obtain a semantic segmentation result of the remote sensing image, where the semantic segmentation result includes global feature information and local feature information of each ground object. By extracting the features of the image to be processed through the global feature extraction module and the local feature extraction module, and fusing the extracted features through the feature fusion module, the purpose that the semantic segmentation result output by the image segmentation network includes global feature information and local feature information is achieved, thereby realizing the technical effect of simultaneously obtaining the global feature information and local feature information of the image, and further solving the technical problem in the related art that it is difficult to simultaneously obtain the global feature information and local feature information in the image segmentation task.
[0103] In an optional implementation manner, the device further includes: a display unit, which includes: a display module for displaying a regional map of each ground object in the semantic segmentation result of the remote sensing image under a first receptive field; a display module for displaying local information of a target ground object in the semantic segmentation result of the remote sensing image under a second receptive field, where the second receptive field is larger than the first receptive field.
[0104] It should be noted here that the above units correspond to the steps in Embodiment 2, and the two units have the same implemented examples and application scenarios as the corresponding steps, but are not limited to the content disclosed in the above Embodiment 1. It should be noted that the above modules, as part of the device, can run in the computer terminal provided in Embodiment 2.
[0105] Embodiment 6
[0106] According to an embodiment of the present invention, there is also provided an image segmentation device for implementing the above image segmentation method, as Figure 8 shown, the device includes:
[0107] A segmentation unit 50 is configured to input a building image into an image segmentation network. The image segmentation network includes a global feature extraction module, a local feature extraction module, and a feature fusion module jointly connected to the output ends of the global feature extraction module and the local feature extraction module.
[0108] A processing unit 60 is configured to process the building image through the image segmentation network to obtain a semantic segmentation result of the building image. The semantic segmentation result includes global feature information of each building and local feature information of each building.
[0109] In an embodiment of the present invention, a segmentation unit 50 is configured to input a building image into an image segmentation network. The image segmentation network includes a global feature extraction module, a local feature extraction module, and a feature fusion module jointly connected to the output ends of the global feature extraction module and the local feature extraction module. A processing unit 60 is configured to process the building image through the image segmentation network to obtain a semantic segmentation result of the building image. The semantic segmentation result includes global feature information of each building and local feature information of each building. By extracting the features of the image to be processed through the global feature extraction module and the local feature extraction module, and fusing the extracted features through the feature fusion module, the purpose that the semantic segmentation result output by the image segmentation network includes global feature information and local feature information is achieved, thereby realizing the technical effect of simultaneously obtaining the global feature information and local feature information of the image, and further solving the technical problem in the related art that it is difficult to simultaneously obtain the global feature information and local feature information in the image segmentation task.
[0110] It should be noted here that the above units correspond to the steps in Embodiment 3. The examples and application scenarios implemented by the two units and the corresponding steps are the same, but are not limited to the content disclosed in the above Embodiment 1. It should be noted that the above modules, as a part of the device, can run in the computer terminal provided in Embodiment 3.
[0111] Embodiment 7
[0112] An embodiment of the present invention can provide a computer terminal, which can be any computer terminal device in a computer terminal group. Optionally, in this embodiment, the above computer terminal can also be replaced with a terminal device such as a mobile terminal.
[0113] Optionally, in this embodiment, the above computer terminal can be located in at least one of multiple network devices in a computer network.
[0114] In this embodiment, the above computer terminal may execute the program code of the following steps in the image segmentation method of the application program: input the image to be processed into the image segmentation network, where the image segmentation network includes a global feature extraction module and a local feature extraction module, and a feature fusion module jointly connected to the output ends of the global feature extraction module and the local feature extraction module; process the image to be processed through the image segmentation network to obtain the semantic segmentation result of the image to be processed, where the semantic segmentation result includes global feature information and local feature information.
[0115] Optionally, Figure 9 is a structural block diagram of a computer terminal according to an embodiment of the present invention. As Figure 9 shown, the computer terminal may include: one or more (only one is shown in the figure) processors and a memory.
[0116] Among them, the memory can be used to store software programs and modules, such as the program instructions / modules corresponding to the image segmentation method and device in the embodiment of the present invention. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory, that is, implements the above image segmentation method. The memory may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memories. In some instances, the memory may further include a memory remotely provided relative to the processor, and these remote memories can be connected to the terminal through a network. Examples of the above network include but are not limited to the Internet, enterprise intranet, local area network, mobile communication network, and combinations thereof.
[0117] The processor can call the information and application program stored in the memory through the transmission device to execute the following steps: input the image to be processed into the image segmentation network, where the image segmentation network includes a global feature extraction module and a local feature extraction module, and a feature fusion module jointly connected to the output ends of the global feature extraction module and the local feature extraction module; process the image to be processed through the image segmentation network to obtain the semantic segmentation result of the image to be processed, where the semantic segmentation result includes global feature information and local feature information.
[0118] Optionally, the above processor may also execute the program code of the following steps: process the image to be processed through the global feature extraction module to obtain the global feature information of the image to be processed; process the image to be processed through the local feature extraction module to obtain the local feature information of the image to be processed; fuse the global feature information and the local feature information through the feature fusion module to obtain the semantic segmentation result of the image to be processed.
[0119] Optionally, the above-mentioned processor may also execute the program code of the following steps: Before inputting the image to be processed into the image segmentation network, the method further includes: constructing the image segmentation network, where the image segmentation network includes the global feature extraction module and the local feature extraction module, and the feature fusion module jointly connected to the output ends of the global feature extraction module and the local feature extraction module.
[0120] Optionally, the above-mentioned processor may also execute the program code of the following steps: The global feature extraction module includes a plurality of globally connected global feature extraction components in series, and the local feature extraction module includes a plurality of locally connected local feature extraction components in series. The output features of the last component of each global feature extraction module and the last component of each local feature extraction module are jointly used as the input features of the feature fusion module.
[0121] Optionally, the above-mentioned processor may also execute the program code of the following steps: After processing the image to be processed through the image segmentation network to obtain the semantic segmentation result of the image to be processed, the method further includes: displaying the semantic segmentation result of the image to be processed, and annotating the global image area corresponding to the global feature and the local image area corresponding to the local feature.
[0122] Optionally, the above-mentioned processor may also execute the program code of the following steps: After displaying the semantic segmentation result of the image to be processed and annotating the global image area corresponding to the global feature and the local image area corresponding to the local feature, the method further includes: obtaining adjustment information for the global image area and / or the local image area; adjusting the global image area and / or the local image area according to the adjustment information to obtain an updated semantic segmentation result.
[0123] Optionally, the above-mentioned processor may also execute the program code of the following steps: During the change of the image to be processed, the target area marked in the semantic segmentation result is correspondingly displayed.
[0124] An embodiment of the present invention provides a solution for an image segmentation method. By inputting an image to be processed into an image segmentation network, where the image segmentation network includes a global feature extraction module, a local feature extraction module, and a feature fusion module commonly connected to the output ends of the global feature extraction module and the local feature extraction module; the image to be processed is processed by the image segmentation network to obtain a semantic segmentation result of the image to be processed, where the semantic segmentation result includes global feature information and local feature information. By extracting the features of the image to be processed through the global feature extraction module and the local feature extraction module, and fusing the extracted features through the feature fusion module, the purpose that the semantic segmentation result output by the image segmentation network includes global feature information and local feature information is achieved, thereby realizing the technical effect of simultaneously obtaining the global feature information and local feature information of the image, and further solving the technical problem in the related art that it is difficult to simultaneously obtain the global feature information and local feature information in the image segmentation task.
[0125] Those of ordinary skill in the art can understand that Figure 9 the structure shown is only schematic, and the computer terminal can also be a smart phone (such as an Android phone, an iOS phone, etc.), a tablet computer, a handheld computer, and a mobile Internet device (Mobile Internet Devices, MID), a PAD and other terminal devices. Figure 9 It does not limit the structure of the above electronic device. For example, the computer terminal 101 may further include more or fewer components (such as a network interface, a display device, etc.) than those shown in Figure 9 the figure, or have a different configuration from that shown in Figure 9 the figure.
[0126] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing the hardware related to the terminal device through a program, and the program can be stored in a computer-readable storage medium. The storage medium may include: a flash drive, a read-only memory (Read-Only Memory, ROM), a random access memory (Random Access Memory, RAM), a magnetic disk or an optical disc, etc.
[0127] Embodiment 8
[0128] An embodiment of the present invention further provides a storage medium. Optionally, in this embodiment, the above storage medium can be used to save the program code executed by the image segmentation method provided in the above Embodiment 1.
[0129] Optionally, in this embodiment, the above storage medium can be located in any computer terminal in a computer terminal group in a computer network, or in any mobile terminal in a mobile terminal group.
[0130] Optionally, in this embodiment, the storage medium is configured to store program code for performing the following steps: input the image to be processed into an image segmentation network, where the image segmentation network includes a global feature extraction module, a local feature extraction module, and a feature fusion module commonly connected to the output ends of the global feature extraction module and the local feature extraction module; process the image to be processed through the image segmentation network to obtain a semantic segmentation result of the image to be processed, where the semantic segmentation result includes global feature information and local feature information.
[0131] The serial numbers of the above embodiments of the present invention are only for description and do not represent the advantages or disadvantages of the embodiments.
[0132] In the above embodiments of the present invention, the descriptions of the various embodiments have their own emphases. For the parts not detailed in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.
[0133] In several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of the units or modules can be in an electrical or other form.
[0134] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0135] In addition, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.
[0136] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes: various media that can store program codes, such as USB flash drives, read-only memories (ROMs), random access memories (RAMs), mobile hard disks, magnetic disks, or optical discs.
[0137] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.
Claims
1. An image segmentation method, characterized in that, comprising: Inputting the image to be processed into an image segmentation network, wherein the image segmentation network includes a global feature extraction module and a local feature extraction module, and a feature fusion module jointly connected to the output ends of the global feature extraction module and the local feature extraction module; Processing the image to be processed through the image segmentation network to obtain a semantic segmentation result of the image to be processed, wherein the semantic segmentation result contains global feature information and local feature information; Among them, different feature extraction modules are in parallel. One feature extraction module extracts features of one resolution of the image to be processed. One feature extraction module includes serially connected feature extraction components inside. The output feature of the target feature extraction component in one feature extraction module is used as the input feature of the next feature extraction component in the current module, and is also used as the input feature of the next feature extraction component of other modules outside the current module. Among them, the feature extraction module includes the global feature extraction module and the local feature extraction module, and both the global feature extraction module and the global feature extraction module are at least two.
2. The image segmentation method according to claim 1, characterized in that, Processing the image to be processed through the image segmentation network to obtain the semantic segmentation result of the image to be processed includes: Processing the image to be processed through the global feature extraction module to obtain the global feature information of the image to be processed; Processing the image to be processed through the local feature extraction module to obtain the local feature information of the image to be processed; Fusing the global feature information and the local feature information through the feature fusion module to obtain the semantic segmentation result of the image to be processed.
3. The image segmentation method according to claim 1, characterized in that, Before inputting the image to be processed into the image segmentation network, the method further includes: constructing the image segmentation network, wherein the image segmentation network includes the global feature extraction module and the local feature extraction module, and the feature fusion module jointly connected to the output ends of the global feature extraction module and the local feature extraction module.
4. The image segmentation method according to claim 3, characterized in that, The global feature extraction module includes a plurality of serially connected global feature extraction components, the local feature extraction module includes a plurality of serially connected local feature extraction components, and the output features of the last component of each global feature extraction module and the last component of each local feature extraction module are jointly used as the input feature of the feature fusion module.
5. The image segmentation method according to claim 1, characterized in that, After processing the image to be processed through the image segmentation network to obtain the semantic segmentation result of the image to be processed, the method further includes: Displaying the semantic segmentation result of the image to be processed, and labeling the global image area corresponding to the global feature and the local image area corresponding to the local feature.
6. The image segmentation method according to claim 5, wherein, after presenting the semantic segmentation result of the image to be processed and annotating the global image region corresponding to the global feature and the local image region corresponding to the local feature, the method further includes: obtaining adjustment information for the global image region and / or the local image region; adjusting the global image region and / or the local image region according to the adjustment information to obtain an updated semantic segmentation result.
7. The image segmentation method according to claim 1, wherein, after processing the image to be processed by the image segmentation network to obtain the semantic segmentation result of the image to be processed, the method further includes: during the change of the image to be processed, correspondingly presenting the target region annotated in the semantic segmentation result.
8. An image segmentation method, wherein, it includes: inputting a remote sensing image into an image segmentation network, where the image segmentation network includes a global feature extraction module and a local feature extraction module, and a feature fusion module commonly connected to the output ends of the global feature extraction module and the local feature extraction module; processing the remote sensing image through the image segmentation network to obtain the semantic segmentation result of the remote sensing image, where the semantic segmentation result contains the global feature information and local feature information of each ground object; wherein, different feature extraction modules are in parallel connection, one feature extraction module extracts the features of one resolution of the remote sensing image, one feature extraction module internally includes serially connected feature extraction components, the output feature of the target feature extraction component in one feature extraction module is used as the input feature of the next feature extraction component in the current module, and is also used as the input feature of the next feature extraction component of other modules outside the current module, where the feature extraction module includes the global feature extraction module and the local feature extraction module, and both the global feature extraction module and the global feature extraction module are at least two.
9. The image segmentation method according to claim 8, wherein, after processing the remote sensing image through the image segmentation network to obtain the semantic segmentation result of the remote sensing image, the method further includes: presenting the regional map of each ground object in the semantic segmentation result of the remote sensing image under the first receptive field; presenting the local information of the target ground object in the semantic segmentation result of the remote sensing image under the second receptive field, where the second receptive field is larger than the first receptive field.
10. An image segmentation method, wherein, it includes: inputting a building image into an image segmentation network, where the image segmentation network includes a global feature extraction module and a local feature extraction module, and a feature fusion module commonly connected to the output ends of the global feature extraction module and the local feature extraction module; Process the building image through the image segmentation network to obtain the semantic segmentation result of the building image, where the semantic segmentation result includes the global feature information of each building and the local feature information of each building; Among them, different feature extraction modules are connected in parallel. One feature extraction module extracts the features of a certain resolution of the building image. One feature extraction module internally includes serially connected feature extraction components. The output features of the target feature extraction component in one feature extraction module serve as the input features of the next feature extraction component in the current module and also serve as the input features of the next feature extraction component of other modules outside the current module. Among them, the feature extraction module includes the global feature extraction module and the local feature extraction module, and both the global feature extraction module and the global feature extraction module are at least two.
11. An image segmentation device, Characterized in that, Comprising: An input unit for inputting the image to be processed into the image segmentation network, where the image segmentation network includes a global feature extraction module and a local feature extraction module, and a feature fusion module commonly connected to the output ends of the global feature extraction module and the local feature extraction module; A processing unit for processing the image to be processed through the image segmentation network to obtain the semantic segmentation result of the image to be processed, where the semantic segmentation result includes global feature information and local feature information; Among them, different feature extraction modules are connected in parallel. One feature extraction module extracts the features of a certain resolution of the image to be processed. One feature extraction module internally includes serially connected feature extraction components. The output features of the target feature extraction component in one feature extraction module serve as the input features of the next feature extraction component in the current module and also serve as the input features of the next feature extraction component of other modules outside the current module. Among them, the feature extraction module includes the global feature extraction module and the local feature extraction module, and both the global feature extraction module and the global feature extraction module are at least two.
12. A storage medium, Characterized in that, The storage medium includes a stored program, where when the program runs, it controls the device where the storage medium is located to execute the image segmentation method according to any one of claims 1 to 7.
13. A computer program product, Characterized in that, The computer program product includes a computer program, and when the computer program is executed, it implements the image segmentation method according to any one of claims 1 to 7.
Citation Information
Patent Citations
An image semantic segmentation method based on a multi-layer information fusion full convolutional neural network
CN109902748A
Medical image labeling method and device for deep learning
CN110993064A