An image scene segmentation method, device, equipment and storage medium
By initially segmenting and fusion of images, detecting and correcting segmentation blocks, the problems of inaccurate and fragmentation of multi-category scene segmentation in image scene segmentation in the prior art are solved, and higher segmentation accuracy and uniformity are achieved.
Patent Information
- Application Number
- CN202210074188.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-01-21
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2042-01-21
AI Technical Summary
The prior art is difficult to achieve accurate segmentation of multiple categories of scenes in image scene segmentation, which often leads to fragmented segmentation results and affects the execution effect of downstream services.
By performing initial scene segmentation and initial fusion processing on the target image, an intermediate scene segmentation diagram is obtained, from which segmentation blocks are detected, and these blocks are segmented and corrected, and finally the target scene segmentation diagram of the target image is obtained.
This method effectively reduces the fragmentation of segmentation blocks, improves the accuracy of segmentation results, and ensures the unified segmentation of image content under the same scene category in the target image.
Smart Images

Figure CN114419070B_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present disclosure relate to the field of image processing technologies, and in particular, to an image scene segmentation method, apparatus, device, and storage medium. Background Art
[0002] Image scene segmentation, as one of the research directions in image processing, is mainly used to separate the scenes included in an image according to scene categories. At present, significant breakthroughs have been made in scene and object segmentation based on deep learning technologies in recent years.
[0003] Existing deep learning networks for image scene segmentation have obvious effects in segmenting single-category or few-category scenes and objects, and the technology is relatively mature. However, for images with multi-category scenes, accurate segmentation cannot be achieved, and fragmented segmentation results are often produced. If the segmentation results are directly applied to downstream business implementation, the execution effect of downstream services will be affected.
[0004] Existing improvement methods mainly consider directly optimizing the deep learning network to optimize the scene segmentation results. However, the deep learning network is overly dependent on the training dataset. Since there are often ambiguities between many scene categories, it is impossible to provide accurate sample data for network training. In addition, more stringent requirements are imposed on the learning ability of the more refined deep learning network itself and the computing power of the device, and it is difficult to achieve a balance between calculation and accuracy. Summary of the Invention
[0005] Embodiments of the present disclosure provide an image scene segmentation method, apparatus, device, and storage medium to optimize the processing of scene segmentation results and reduce the fragmentation of scene segmentation results.
[0006] In a first aspect, embodiments of the present disclosure provide an image scene segmentation method, the method
[0007] obtains an intermediate scene segmentation map by performing initial scene segmentation and initial scene fusion processing on the acquired target image;
[0008] detects to-be-processed segmentation blocks from the intermediate scene segmentation map;
[0009] obtains a target scene segmentation map of the target image by performing segmentation correction on each of the to-be-processed segmentation blocks.
[0010] In a second aspect, embodiments of the present disclosure further provide an image scene segmentation apparatus, the apparatus includes:
[0011] An initial processing module, configured to obtain an intermediate scene segmentation map by performing initial scene segmentation and initial scene fusion processing on the acquired target image;
[0012] An information determination module, configured to detect a to-be-processed segmentation block from the intermediate scene segmentation map;
[0013] A segmentation correction module, configured to obtain a target scene segmentation map of the target image by performing segmentation correction on each of the to-be-processed segmentation blocks.
[0014] In a third aspect, an embodiment of the present disclosure further provides an electronic device, which includes:
[0015] One or more processors;
[0016] A storage device, configured to store one or more programs,
[0017] When the one or more programs are executed by the one or more processors, the one or more processors implement the image scene segmentation method provided in any embodiment of the present disclosure.
[0018] In a fourth aspect, an embodiment of the present disclosure further provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the image scene segmentation method provided in any embodiment of the present disclosure is implemented.
[0019] The technical solution of the embodiment of the present disclosure first performs processing of scene initial segmentation and scene initial fusion on the acquired target image to obtain an intermediate scene segmentation map; then, a segmentation block to be optimized can be determined from the intermediate scene segmentation map, and finally, segmentation correction can be performed on each to-be-processed segmentation block, so as to obtain a target scene segmentation map of the target image. The above technical solution solves the problem that the existing image scene segmentation method cannot achieve accurate segmentation and generates more fragmented segmentation results. Different from the traditional improvement solutions, the key of the solution provided in this embodiment lies in detecting fragmentation of the segmentation result after image scene segmentation, detecting fragmented segmentation blocks and performing segmentation correction, and the corrected segmentation result realizes unified segmentation of the image content under the same scene category in the target image, reduces the fragmentation of the segmentation blocks, and achieves the beneficial effect of effectively improving the accuracy of the segmentation result. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] In order to more clearly illustrate the technical solutions of the exemplary embodiments of the present disclosure, the following briefly introduces the drawings required for describing the embodiments. Obviously, the introduced drawings are only the drawings of a part of the embodiments to be described in the present invention, rather than all the drawings. For those of ordinary skill in the art, other drawings can be obtained according to these drawings without creative efforts.
[0021] Figure 1 It is a schematic flowchart of an image scene segmentation method provided in Embodiment 1 of the present disclosure;
[0022] Figure 2 It is a flowchart showing a method for image scene segmentation provided in the second embodiment of the present disclosure;
[0023] Figure 2a It gives a structural schematic diagram of the scene segmentation network model used in the initial scene segmentation in a method for image scene segmentation provided in the second embodiment of the present disclosure;
[0024] Figure 2b It is a flowchart showing the implementation of image fusion processing in the method for image scene segmentation provided in the second embodiment of the present disclosure;
[0025] Figure 2c It gives an effect display diagram of the determined intermediate scene segmentation map in the method for image scene segmentation provided in this embodiment;
[0026] Figure 2d It gives a flowchart showing the implementation of determining the to-be-processed segmentation blocks in the method for image scene segmentation provided in the second embodiment of this disclosure;
[0027] Figure 2e It gives an example diagram showing the effect of displaying the determined to-be-processed segmentation blocks in the same image in this embodiment;
[0028] Figure 2f It gives a flowchart showing the implementation of determining the segmentation layer to which the to-be-processed segmentation blocks belong in the method for image scene segmentation provided in the second embodiment of this disclosure;
[0029] Figure 2g It gives an effect display diagram of the target scene segmentation map in the method for image scene segmentation provided in this embodiment;
[0030] Figure 3 It is a structural schematic diagram of an image scene segmentation device provided in the third embodiment of the present disclosure;
[0031] Figure 4 It is a structural schematic diagram of an electronic device provided in the seventh embodiment of the present disclosure. Detailed implementation manners
[0032] Next, the embodiments of the present disclosure will be described in more detail with reference to the accompanying drawings. Although some embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. On the contrary, these embodiments are provided to more thoroughly and completely understand the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not used to limit the protection scope of the present disclosure.
[0033] It should be understood that the various steps described in the method embodiments of the present disclosure may be executed in different orders and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this regard.
[0034] As used herein, the term "comprising" and its variations are open-ended, that is, "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". The relevant definitions of other terms will be given in the following description.
[0035] It should be noted that the concepts such as "first", "second", etc. mentioned in the present disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units. It should be noted that the modifications of "one" and "multiple" mentioned in the present disclosure are illustrative rather than restrictive. Those skilled in the art should understand that unless clearly stated otherwise in the context, it should be understood as "one or more".
[0036] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only for illustrative purposes and are not used to limit the scope of these messages or information.
[0037] Embodiment 1
[0038] Figure 1 As shown in the flowchart of an image scene segmentation method provided in Embodiment 1 of the present disclosure, this embodiment is applicable to the case of performing image segmentation on the acquired image. This method can be executed by an image scene segmentation device, which can be implemented by software and / or hardware, and can be configured in a terminal and / or a server to implement the image scene segmentation method in the embodiments of the present disclosure.
[0039] As Figure 1 shown, a specific image scene segmentation method provided in Embodiment 1 may specifically include:
[0040] S101. Obtain an intermediate scene segmentation map by performing initial scene segmentation and initial scene fusion processing on the acquired target image.
[0041] In this embodiment, the target image can be specifically understood as an image to be subjected to image scene segmentation processing. It can be a scene image captured in real time or an image frame intercepted from a captured video stream. In this step, the target image can first be subjected to initial scene segmentation. It should be noted that for the scene segmentation of an image, it is equivalent to segmenting each image content included in the image according to the scene category to which it belongs, so that the image content of the same scene category is segmented within the same scene layer. Exemplarily, for example, the doors or windows appearing in the image can be segmented into the segmentation layer with the scene category of doors and windows, and the vehicles appearing in the image can be segmented into the segmentation layer with the scene category of vehicles.
[0042] In this embodiment, a pre-constructed scene segmentation network model can be used to perform initial scene segmentation on the target image. It should be noted that the pre-constructed scene segmentation network model can be regarded as a general scene segmentation model, which can be used for scene segmentation of multiple different scene categories, but there may not be a suitable setting for the coarse-grained or fine-grained scene categories that can be segmented, nor is it specifically limited to the applicable application scenarios. Therefore, the initial scene segmentation result obtained through the scene segmentation network model may not be the scene segmentation result required by downstream business applications.
[0043] Exemplarily, assume that the object to be processed by downstream business applications is a building group in the image. Before this, a scene segmentation map containing only the building group needs to be obtained. However, the scene segmentation result after the initial scene segmentation in this step also includes other scene segmentation blocks, or fragmented segmentation blocks with a small image area. For example, the doors and windows on the building may be independently segmented and not segmented into the same scene as the building, and an accurate segmentation map of the building group cannot be obtained. Therefore, if only the initial scene segmentation of the target image is performed, downstream business applications cannot obtain effective image information.
[0044] Based on this, after performing initial scene segmentation on the target image in this step, it is also necessary to perform scene fusion of the image on the obtained initial scene segmentation result. This scene fusion can be regarded as the initial scene fusion in this embodiment, and the scene segmentation result after the initial scene fusion is recorded as the intermediate scene segmentation map. In this embodiment, the initial fusion of the image can be achieved by adopting a certain fusion rule for the segmentation layers under each scene category included in the initial scene segmentation map. Among them, the adopted fusion rule can be to fuse the segmentation layer with a smaller scene category range into the segmentation layer corresponding to a larger scene category range.
[0045] Exemplarily, after performing initial scene segmentation, an initial scene segmentation map can be obtained. From this, the scene labels corresponding to each segmentation layer included in the initial scene segmentation map can be obtained. Subsequently, it can be analyzed whether there is an attribution association between the scene labels, and the scene segmentation maps with an attribution association can be fused. For example, for the segmented floor segmentation layer, its scene label can be building floor, and for the segmented door and window segmentation layer, its scene label can be doors and windows. By analyzing the attribution association between the building floor and the doors and windows, it can be found that the doors and windows often depend on the building, that is, the category range of the doors and windows is smaller than that of the building floor. Thus, the door and window segmentation layer and the building floor segmentation layer can be fused to form a new building segmentation layer.
[0046] In this embodiment, compared with the further processing of the scene segmentation map in the subsequent steps, the scene fusion performed on the initial scene segmentation result in this step can be regarded as an initial scene fusion process of the scene segmentation result. Moreover, each segmentation layer newly formed after the scene fusion can constitute a new scene segmentation map. In this embodiment, this scene segmentation map is denoted as the intermediate scene segmentation map, and it is relative to the processing of the scene segmentation map in the subsequent steps.
[0047] S102. Detect the segmentation blocks to be processed from the intermediate scene segmentation map.
[0048] In this embodiment, the intermediate scene segmentation map in this step can be regarded as the scene segmentation result after the initial fusion process of the initial segmentation result corresponding to the target image, which mainly includes segmentation layers that segment the image content according to scene categories. That is, it can be considered that the image content included in each segmentation layer belongs to the same scene category, and the scene category to which it belongs can be considered to have a relatively large scene category division range.
[0049] It can be known that the scene segmentation algorithm used for the initial scene segmentation of the target image cannot guarantee the accuracy of the scene segmentation. Thus, there is a situation where the image content is segmented into the wrong scene category, and through the above initial fusion process, the wrong segmentation of the scene category to which the image content belongs cannot be eliminated.
[0050] For the segmentation layers included in the intermediate scene segmentation map, in the case of correct scene segmentation, the image content area it has should be a connected area with a relatively large area; if there are isolated other image content areas in the connected area with a relatively large area, then the isolated other image content area is likely to be an area with abnormal scene segmentation, that is, there is a wrong scene segmentation in this segmentation layer.
[0051] In this embodiment, the above-mentioned incorrectly segmented area of the scene can be recorded as a segmentation block to be processed, and the detection of the segmentation block to be detected and processed can be determined by detecting the connected areas of each segmentation layer in the intermediate scene segmentation map. Exemplarily, in the segmentation layer containing the image content of the same scene category, by scanning each pixel point in the segmentation layer, the detection of the connected area of the image can be realized, and the area of each connected area can be determined. If there is a connected area with an area smaller than a certain threshold, this embodiment can use this connected area as a segmentation block to be processed.
[0052] S103. Obtain the target scene segmentation map of the target image by performing segmentation correction on each of the segmentation blocks to be processed.
[0053] In this embodiment, the above-mentioned detected segmentation block to be processed is equivalent to the incorrectly segmented block of the scene segmentation. Through this step, the segmentation correction can be performed on the segmentation block to be processed to determine the correct segmentation layer to which the segmentation block to be processed should belong, and the segmentation block to be processed can be fused into the correct segmentation layer. When all the segmentation blocks to be processed are fused to the correct scene segmentation map through the above logic, the obtained segmentation layers constitute the target scene segmentation map of the target image.
[0054] Exemplarily, one implementation method for performing segmentation correction on the segmentation block to be processed and determining the segmentation layer to which the segmentation block to be processed actually should belong can be described as follows: perform region expansion on the segmentation block to be processed to obtain the segmentation expansion region of the segmentation block to be processed. There are overlapping regions in the segmentation expansion region that overlap with other already determined connected regions on each segmentation layer; this embodiment can determine which connected region the segmentation block to be processed should belong to through the overlapping ratio of other already determined connected regions in the overlapping region, and then can determine the segmentation layer where the connected region to which it belongs is located, which can be used as the segmentation layer to which the segmentation block to be processed actually should belong.
[0055] The image scene segmentation method provided in this embodiment 1 first processes the initial segmentation set and initial fusion of the scene of the acquired target image to obtain an intermediate scene segmentation map; then, the segmentation blocks to be optimized and processed can be determined from the intermediate scene segmentation map, and finally, segmentation correction can be performed on each segmentation block to be processed, so as to obtain the target scene segmentation map of the target image. The above technical solution solves the problem that the existing image scene segmentation method cannot achieve accurate segmentation and produces more fragmented segmentation results. Different from the traditional improvement solutions, the key of the solution provided in this embodiment lies in detecting the fragmentation of the segmentation result after image scene segmentation, detecting the fragmented segmentation blocks and performing segmentation correction. The corrected segmentation result realizes the unified segmentation of the image content under the same scene category in the target image, reduces the fragmentation of the segmentation blocks, and achieves the beneficial effect of effectively improving the accuracy of the segmentation result.
[0056] Example 2
[0057] Figure 2 As shown in the flowchart of an image scene segmentation method provided in Example 2 of the present disclosure. Based on any optional technical solution in the embodiments of the present disclosure, optionally, the process of obtaining an intermediate scene segmentation map by performing initial scene segmentation and initial scene fusion processing on the acquired target image can be specifically optimized as follows: taking the acquired target image as input data and inputting it into a preset scene segmentation network model to obtain an output initial scene segmentation map, where the initial scene segmentation map includes at least one initial segmentation layer; based on the content labels corresponding to each initial segmentation layer, performing initial scene fusion on each initial segmentation layer to obtain an intermediate scene segmentation map.
[0058] Meanwhile, in this embodiment, the process of detecting the segmentation blocks to be processed from the intermediate scene segmentation map can be specifically optimized as follows: extracting each intermediate segmentation layer included in the intermediate scene segmentation map; by performing connected component detection on each intermediate segmentation layer, determining the segmentation blocks to be processed in the intermediate scene segmentation map.
[0059] In addition, in this embodiment, the process of obtaining the target scene segmentation map of the target image by correcting the segmentation results of each segmentation block to be processed can be specifically optimized as follows: for each segmentation block to be processed, performing regional dilation processing on the segmentation block to be processed according to a set dilation coefficient to obtain a corresponding dilation region; based on the dilation region, determining the target segmentation layer to which the segmentation block to be processed belongs from the intermediate scene segmentation map; performing image fusion on the segmentation block to be processed and the target segmentation layer; using the intermediate scene segmentation map after the fusion processing as the target scene segmentation map of the target image.
[0060] As Figure 2 shown, an image scene segmentation method provided in Example 2 of the present disclosure may specifically include the following steps:
[0061] S201. Taking the acquired target image as input data and inputting it into a preset scene segmentation network model to obtain an output initial scene segmentation map, where the initial scene segmentation map includes at least one initial segmentation layer.
[0062] In this embodiment, the logical implementation of the initial scene segmentation is given in this step. Specifically, this step mainly performs the initial scene segmentation through a given scene segmentation network model. Among them, the target image can be directly used as input data to input into the scene segmentation network model, and the scene segmentation network model can be considered as a neural network model with a specific network structure that is pre-constructed. After iteratively learning and training the neural network model with a pre-set training sample set, the scene segmentation network model adopted in this step can be formed. The scene segmentation network model performs feature extraction and arithmetic processing based on network parameters on the input target image, and can output an initial scene segmentation map containing at least one initial segmentation layer.
[0063] It can be known that the image content belonging to the same scene category is included in the initial segmentation layer in the initial scene segmentation map. To better distinguish each initial segmentation layer included in the initial scene segmentation map, different color assignments can be made for different segmentation layers.
[0064] In this embodiment, the scene segmentation network model can be regarded as a general scene segmentation model, that is, it can be applied to various application scenarios that appear in business applications. In addition to the input layer and the output layer, the scene segmentation network model also includes hidden layers that actually participate in the scene segmentation process. Optionally, the hidden layer of the scene segmentation network model includes a set number of residual sub-network models; the residual sub-network models are sequentially connected in a hierarchical order, and there is also a residual connection from one residual sub-network model to another non-adjacent residual sub-network model; each residual sub-network model is composed of a convolutional layer, a batch normalization layer, and a non-linear activation function layer.
[0065] In this embodiment, the convolutional kernel adopted by the convolutional layer in the residual sub-network model can be a 3*3 convolutional kernel; the non-linear activation function adopted can be a ReLU function; at the same time, in addition to the sequential connection between the residual sub-network models, there is also a residual connection. When the network structure of the scene segmentation network model is relatively deep, through the above connections, it is more conducive to the training of the network model.
[0066] Exemplarily, Figure 2a The structural schematic diagram of the scene segmentation network model adopted in the initial scene segmentation in an image scene segmentation method provided in the second embodiment of the present disclosure is given. As Figure 2aAs shown in the figure, the scene segmentation network model includes several basic ResNet units of the residual network. Each basic ResNet unit is composed of a convolutional layer with a 3X3 convolutional kernel, a batch normalization (BN batchnorm) layer, and a ReLU (a non-linear activation function) layer. There is a direct connection path 21 between each basic ResNet unit. In addition, there is an additional residual connection path 22. In this embodiment, the network model composed of these basic ResNet units is equivalent to the backbone part of image scene segmentation, and the scene segmentation map in the target image can be calculated through feature extraction.
[0067] S202. Based on the content labels corresponding to each initial segmentation layer, perform initial scene fusion on each initial segmentation layer to obtain an intermediate scene segmentation map.
[0068] In this embodiment, this step gives the logical implementation of initial scene fusion. Among them, the initial segmentation layer can be considered as the segmentation layer in the initial scene segmentation map obtained in S201 above; each of the initial segmentation layers contains image content in the same scene category; the content label can be regarded as the scene category label of the initial segmentation layer, which is used to identify the scene category of the image content included in the scene segmentation image; this content label can be obtained together when the initial scene segmentation map is obtained.
[0069] Based on the above analysis of this embodiment, it can be seen that the scene categories that can be segmented in the initial scene segmentation map are relatively diverse, the thickness and granularity of the scene categories are not the same, and there is a situation where a certain scene category actually belongs to another scene category. And the over-fine division of the scene category may not match the application scenario corresponding to the image scene segmentation, thus unable to ensure the effectiveness of the obtained segmentation result.
[0070] Exemplarily, assume that the actual application segmentation scenario in the business application is to segment the building complex from the ground and the sky, and there are segmentation layers with content labels of flowers, grass, and trees in the obtained initial scene segmentation map. At this time, it is equivalent to that the segmentation result does not match the required application segmentation scenario. Further analysis shows that flowers, grass, and trees are actually plants growing on the ground and should belong to a part of the ground. To obtain a more matching segmentation result, it is necessary to perform scene fusion processing through this step.
[0071] The scene fusion processing in this step can be implemented based on the content labels of each initial segmentation layer. Specifically, in the execution logic, corresponding scene category fusion rules can be set for the application scenario. Then, multiple content labels that meet the scene category fusion rules can be determined, and the segmentation layers corresponding to them can be fused to form new segmentation layers. After completing the scene fusion, the intermediate scene segmentation map can be formed based on the segmentation layers formed after the fusion processing.
[0072] Optionally, Figure 2b This is the implementation flowchart of image fusion processing in the image scene segmentation method provided in the second embodiment of the present disclosure. As Figure 2b shown, on the basis of the above embodiment, in this embodiment, further scene initial fusion of each initial segmentation layer is performed based on the content tags corresponding to each initial segmentation layer to obtain an intermediate scene segmentation map, which is specifically implemented as the following steps:
[0073] S2021. Obtain the content tags of each of the initial segmentation layers.
[0074] In this embodiment, the content tags of each initial segmentation layer can be extracted from the obtained initial scene segmentation map.
[0075] S2022. Search the pre-set tag category association table to determine the scene branch to which each of the content tags belongs.
[0076] In this embodiment, the tag category association table is a pre-set information rule table, which can be specifically set depending on the current application scenario. Through analysis of the application scenario requirements, relevant technicians can determine multiple scene branches that match the application scenario, and there can be multiple content tags with attribution relationships or parallel relationships under different scene branches.
[0077] Exemplarily, assume that a scene branch is the ground. In an application scenario, it can be considered that the content tags associated with this scene branch at least include: ground, flowers, grass, and trees, etc. Thus, in the tag category association table set for this application scenario, one record can be expressed as that content tags such as flowers, grass, trees, and the ground respectively belong to the ground scene branch. In this embodiment, after obtaining the content tags of each initial segmentation layer, through the search of the tag category association table in this step, the scene branch associated with each content tag can be obtained.
[0078] S2023. Perform image content fusion on the initial segmentation layers belonging to the same scene branch to obtain a fused intermediate scene segmentation map.
[0079] Continuing with the above example description, assume that it is determined that the ground, flowers, grass, and trees all belong to the ground scene branch. Then, the initial segmentation layers corresponding to the ground, flowers, grass, and trees in the initial scene segmentation map can be subjected to scene initial fusion, and finally, an intermediate scene segmentation map can be obtained through this step.
[0080] In the following S203 and S204 of this embodiment, the logical implementation of detecting the segmentation block to be processed is given.
[0081] Exemplarily, Figure 2cThe figure shows the effect display diagram of the intermediate scene segmentation map determined in the image scene segmentation method provided in this embodiment. As Figure 2c shown, for better understanding of the details of the intermediate scene segmentation map, Figure 2c it shows the intermediate scene segmentation map 23, and also specifically shows each intermediate segmentation layer included in the intermediate scene segmentation map 23. It can be seen that the main content presented in the first layer 231 shown is the building complex; the main content presented in the second layer 232 shown is the ground, and the main content presented in the third layer 233 shown is the sky.
[0082] S203. Extract each intermediate segmentation layer included in the intermediate scene segmentation map.
[0083] After obtaining the intermediate scene segmentation map through the above S202, it is equivalent to knowing the included intermediate segmentation layers, and this step extracts each intermediate segmentation layer.
[0084] S204. Determine the to-be-processed segmentation blocks of the intermediate scene segmentation map by performing connected component detection on each intermediate segmentation layer.
[0085] In this embodiment, the connected component detection in this step can be implemented through a set connected component detection algorithm. Among them, the core of the connected component detection algorithm can be to scan pixel points of the binarized image to determine whether the pixel points are in the same area, and then the connected areas in the intermediate segmentation layer can be determined; then the to-be-processed segmentation blocks with abnormal segmentation can be found according to the areas of the connected areas.
[0086] Specifically, Figure 2d the figure shows the implementation flowchart of determining the to-be-processed segmentation blocks in the image scene segmentation method provided in the second embodiment of this invention. As Figure 2d shown, on the basis of the above embodiment, in this embodiment, further determining the to-be-processed segmentation blocks of the intermediate scene segmentation map by performing connected component detection on each of the intermediate segmentation layers is specifically preferably the following steps:
[0087] S2041. Perform binarization processing on each of the intermediate segmentation layers to obtain corresponding binarized segmentation layers.
[0088] Exemplarily, the binarization processing can be to assign pixel values of 0 or 1 to each pixel point in the intermediate segmentation layer.
[0089] S2042. For each binarized segmentation layer, scan the pixel values of the binarized segmentation layer in a set scanning order.
[0090] Exemplarily, the scanning order of pixel points can be from left to right and from top to bottom; through this scanning step, the pixel values of each pixel point can be determined.
[0091] S2043. Determine each connected region included in the binary segmentation layer according to the pixel value scanning results of each one.
[0092] Exemplarily, the process of detecting connected regions based on pixel values in this embodiment can be performed in real time during the pixel value scanning process. Specifically, the detection of connected regions can be described as follows: If the pixel value of the currently scanned pixel point is 0, move to the next pixel point in the scanning order; if the pixel value of the currently scanned pixel point is 1, detect the two adjacent pixel points on the left and above of the current pixel point, and then, according to the pixel values and detection marks of these two adjacent pixel points, consider the following 4 situations:
[0093] 1) The pixel values of both adjacent pixel points are 0. At this time, give a new mark to the current pixel point (indicating the start of a new connected domain).
[0094] 2) Only one of the pixel values of the two adjacent pixel points is 1. At this time, the mark of the current pixel point is the same as the mark of the adjacent pixel point with a pixel value of 1.
[0095] 3) The pixel values of both adjacent pixel points are 1 and the marks are the same. At this time, the mark of the current pixel point is also this mark.
[0096] 4) The pixel values of both adjacent pixel points are 1 and the marks are different. Assign the smaller mark among the marks corresponding to the two adjacent pixel points to the current pixel point.
[0097] Following the above description, after the pixel point scanning is completed, through the marks corresponding to each pixel point, the regions with the same marks can be regarded as a connected region. Through the above operations in this embodiment, each included connected region can be determined.
[0098] S2044. Take the connected regions with an area smaller than the set area threshold as the to-be-processed segmentation blocks.
[0099] In this embodiment, the area of each connected region can be determined, and this area can be characterized by the number of pixel values. Through the above description, in scene segmentation, the segmentation blocks with abnormal segmentation often appear as independent segmentation blocks with a smaller area. Therefore, in this step, the connected regions with an area smaller than the set area threshold can be taken as the to-be-processed segmentation blocks.
[0100] Exemplarily, following the above Figure 2c description, the to-be-processed segmentation blocks are also determined through connected domain detection in each intermediate segmentation layer shown. For example, the connected region within the first rectangular frame 234 in the first layer 231; the connected region within the second rectangular frame 235 in the second image 232 can both be regarded as the determined to-be-processed segmentation blocks.
[0101] Meanwhile,Figure 2e Figure 10 shows an example of the effect of displaying each to-be-processed segmentation block determined in the same image in this embodiment; as Figure 2e shown, Figure 2e image 24 in FIG. 10 includes each to-be-processed segmentation block detected from the corresponding intermediate scene segmentation map 23 above. To facilitate better identification of each to-be-processed segmentation block, different color values can be used to fill the pixel points in the to-be-processed segmentation block. Figure 2c In this embodiment, the following S205 and S206 give the specific implementation of segmenting and correcting the to-be-processed segmentation block.
[0102] S205. For each to-be-processed segmentation block, perform regional dilation processing on the to-be-processed segmentation block according to the set dilation coefficient to obtain the corresponding segmentation dilation region.
[0103] In this embodiment, the set dilation coefficient can be a 3*3 all-1 matrix, and the to-be-processed segmentation block participating in the dilation is the dilation center. Then, the to-be-processed segmentation block is dilated around it with a 3*3 all-1 matrix. In this embodiment, the dilated region can be denoted as the segmentation dilation region. Among them, the segmentation dilation region can only be the peripheral dilation region that dilates around and does not include the to-be-processed segmentation block; it can also be the fusion of the to-be-processed segmentation block and the peripheral dilation region.
[0104] S206. Based on the segmentation dilation region, determine the target segmentation layer to which the to-be-processed segmentation block belongs from the intermediate scene segmentation map.
[0105] It should be understood that for the detected to-be-processed segmentation block, the corresponding segmentation dilation region thereof may overlap with any intermediate segmentation layer in the intermediate scene segmentation map. In this step, based on the overlapping ratio of the segmentation dilation region and any intermediate segmentation layer, it can be determined which intermediate segmentation layer the to-be-processed segmentation block belongs to.
[0106] Specifically,
[0107] Figure 11 shows a flowchart of the implementation of determining the segmentation layer to which the to-be-processed segmentation block belongs in the image scene segmentation method provided in the second embodiment of this application. As Figure 2f shown, on the basis of the above embodiment, this embodiment further specifies the step of determining the target segmentation layer to which the to-be-processed segmentation block belongs from the intermediate scene segmentation map based on the segmentation dilation region as the following steps: Figure 2f S2061. Obtain each intermediate segmentation layer included in the intermediate scene segmentation map, and determine the candidate segmentation layers that overlap with the segmentation dilation region.
[0108]
[0109] Exemplarily, by dividing the pixel positions of each pixel point included in the dilated region and the pixel positions of the image content included in each intermediate division layer, it is possible to determine which intermediate division layers overlap with the dilated region, and the overlapping intermediate division layers are recorded as candidate division layers.
[0110] S2062. Count the number of pixel points in the region overlapping with each of the candidate division layers.
[0111] S2063. Take the candidate division layer corresponding to the maximum number of pixel points as the target division layer to which the to-be-processed division block belongs.
[0112] In this embodiment, the maximum number of pixel points is equivalent to the largest number of overlapping pixel points of the dilated region in the target division layer.
[0113] S207. Perform image fusion on the to-be-processed division block and the target division layer.
[0114] It can be known that in this embodiment, it is preferable that the pixel values of the pixel points corresponding to the image content in the same division layer are the same. Exemplarily, one image fusion method can be described as making the pixel values of each pixel point in the to-be-processed division block equal to the pixel values of the pixel points in the target division layer.
[0115] S208. Take the intermediate scene division map after the fusion process as the target scene division map of the target image.
[0116] In this embodiment, the fusion process in this step is equivalent to the scene fusion of the to-be-processed division block with the target division layer during the division correction. Thus, the abnormal division repair of each to-be-processed division block is achieved, and the number of fragmented division blocks on each division layer in the finally obtained target scene division map is significantly reduced.
[0117] Exemplarily, Figure 2g shows the effect display diagram of the target scene division map in the image scene division method provided in this embodiment. As Figure 2g shown, for better understanding of the details of the target scene division map, the presented effect diagram corresponds to the above Figure 2c mutually, where Figure 2g shows the target scene division map 25 and also shows each target division layer included in the target scene division map 25. It can be seen that the main buildings are presented in the fourth layer 251 shown; the ground is mainly presented in the fifth layer 252 shown, and the sky is mainly presented in the sixth layer 253 shown.
[0118] Compare Figure 2g with Figure 2c and it can be found that Figure 2cThe fragmented segmentation blocks within the second rectangular box 235 in the presented second layer 232 are finally fused through segmentation correction into Figure 2g the presented fourth layer 251, thereby realizing the completion of the building complex scene, and further realizing Figure 2g the accuracy of the ground scene within the presented fifth layer 235.
[0119] An image scene segmentation method provided in the second embodiment gives the initial scene segmentation of the image through the scene segmentation network module and the processing of the first segmentation result through scene initial fusion; at the same time, it also gives the specific implementation of detecting the segmentation blocks to be processed, and also gives the specific implementation of segmenting and correcting the segmentation blocks to be processed. Through the method provided in this embodiment, the problem that the existing image scene segmentation method cannot achieve accurate segmentation and produces many fragmented segmentation results is solved. Different from the traditional improvement solutions, the key of the solution provided in this embodiment lies in detecting the fragmentation of the segmentation result after image scene segmentation, detecting the fragmented segmentation blocks and performing segmentation correction, and the corrected segmentation result realizes the unified segmentation of the image content under the same scene category in the target image, reduces the fragmentation of the segmentation blocks, and achieves the beneficial effect of effectively improving the accuracy of the segmentation result.
[0120] Embodiment Three
[0121] Figure 3 The following is a schematic structural diagram of an image scene segmentation device provided in the third embodiment of the present disclosure. This embodiment is applicable to the situation of image segmentation of the acquired image. The device can be implemented by software and / or hardware, and can be configured in a terminal and / or a server to implement the image scene segmentation method in the embodiments of the present disclosure. The device specifically includes: an initial processing module 31, an information determination module 32, and a segmentation correction module 33.
[0122] Among them, the initial processing module 31 is used to obtain an intermediate scene segmentation map through initial scene segmentation and initial scene fusion processing of the acquired target image;
[0123] The information determination module 32 is used to detect the segmentation blocks to be processed from the intermediate scene segmentation map;
[0124] The segmentation correction module 33 is used to obtain the target scene segmentation map of the target image by performing segmentation correction on each of the segmentation blocks to be processed.
[0125] An image scene segmentation device provided in Embodiment 3 solves the problem that the existing image scene segmentation method cannot achieve accurate segmentation and generates many fragmented segmentation results. Different from traditional improvement schemes, the key to the solution provided in this embodiment lies in detecting fragmentation of the segmentation result after image scene segmentation, detecting fragmented segmentation blocks and performing segmentation correction. The corrected segmentation result realizes unified segmentation of image content under the same scene category in the target image, reduces the fragmentation of segmentation blocks, and achieves the beneficial effect of effectively improving the accuracy of the segmentation result.
[0126] Based on any optional technical solution in the embodiments of the present disclosure, optionally, the initial processing module 31 includes:
[0127] An initial segmentation unit, configured to use the obtained target image as input data, input it into a preset scene segmentation network model, and obtain an output initial scene segmentation map, where the initial scene segmentation map includes at least one initial segmentation layer;
[0128] An initial fusion unit, configured to perform initial scene fusion on each of the initial segmentation layers based on the content labels corresponding to the initial segmentation layers, and obtain an intermediate scene segmentation map.
[0129] Based on any optional technical solution in the embodiments of the present disclosure, optionally, the initial fusion unit may specifically be configured to:
[0130] Obtain the content labels of each of the initial segmentation layers; search a preset label category association table to determine the scene branches to which the content labels belong; perform image content fusion on the initial segmentation layers belonging to the same scene branch to obtain a fused intermediate scene segmentation map.
[0131] Based on any optional technical solution in the embodiments of the present disclosure, optionally, the hidden layer of the scene segmentation network model includes a set number of residual sub-network models; the residual sub-network models are sequentially connected in a hierarchical order, and there is also a residual connection from one residual sub-network model to another non-adjacent residual sub-network model; each residual sub-network model is composed of a convolutional layer, a batch normalization layer, and a non-linear activation function layer.
[0132] Based on any optional technical solution in the embodiments of the present disclosure, optionally, the information determination module 32 may specifically include:
[0133] An information extraction unit, configured to extract each intermediate segmentation layer included in the intermediate scene segmentation map;
[0134] An information determination unit, configured to determine the segmentation blocks to be processed in the intermediate scene segmentation map by performing connected component detection on each of the intermediate segmentation layers.
[0135] Based on any optional technical solution in the embodiments of the present disclosure, optionally, the information determination unit may specifically be configured to: perform binarization processing on each of the intermediate segmentation layers to obtain corresponding binarized segmentation layers; for each binarized segmentation layer, scan the pixel values of the binarized segmentation layer in a set scanning order; determine each connected region included in the binarized segmentation layer according to the pixel value scanning results; and use the connected regions with an area smaller than a set area threshold as the segmentation blocks to be processed.
[0136] Based on any optional technical solution in the embodiments of the present disclosure, optionally, the segmentation correction module may specifically include:
[0137] a region determination unit, configured to perform region dilation processing on each segmentation block to be processed according to a set dilation coefficient to obtain a corresponding dilated segmentation region;
[0138] a first correction unit, configured to determine, based on the dilated segmentation region, a target segmentation layer to which the segmentation block to be processed belongs in the intermediate scene segmentation map;
[0139] a second correction unit, configured to perform image fusion on the segmentation block to be processed and the target segmentation layer;
[0140] a target determination unit, configured to use the intermediate scene segmentation map after the fusion processing as the target scene segmentation map of the target image.
[0141] Based on any optional technical solution in the embodiments of the present disclosure, optionally, the second correction unit may specifically be configured to:
[0142] obtain each intermediate segmentation layer included in the intermediate scene segmentation map, determine candidate segmentation layers that overlap with the dilated segmentation region; count the number of pixel points in the regions overlapping with each candidate segmentation layer; and use the candidate segmentation layer corresponding to the largest number of pixel points as the target segmentation layer to which the segmentation block to be processed belongs.
[0143] The above device can execute the method provided in any embodiment of the present disclosure, and has corresponding functional modules and beneficial effects for executing the method.
[0144] It should be noted that the various units and modules included in the above device are only divided according to functional logic, but are not limited to the above division, as long as the corresponding functions can be achieved; in addition, the specific names of the functional units are only for the convenience of mutual distinction and do not limit the protection scope of the embodiments of the present disclosure.
[0145] Embodiment 4
[0146] Figure 4The following is a schematic structural diagram of an electronic device provided in Embodiment 7 of the present disclosure. Refer to Figure 4 , which shows a schematic structural diagram of an electronic device (such as a Figure 4 terminal device or a server in) 40 suitable for implementing the embodiments of the present disclosure. The terminal device in the embodiments of the present disclosure may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Tablet Computers), PMPs (Portable Multimedia Players), in-vehicle terminals (such as in-vehicle navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 4 The electronic device shown is only an example and should not impose any limitation on the functions and usage scope of the embodiments of the present disclosure.
[0147] As Figure 4 shown, the electronic device 40 may include a processing device (such as a central processing unit, a graphics processing unit, etc.) 41, which may perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 42 or a program loaded from a storage device 48 into a random access memory (RAM) 43. In the RAM 43, various programs and data required for the operation of the electronic device 40 are also stored. The processing device 41, the ROM 42, and the RAM 43 are connected to each other through a bus 45. An editing / output (I / O) interface 44 is also connected to the bus 45.
[0148] Generally, the following devices may be connected to the I / O interface 44: an input device 46 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 47 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 48 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 49. The communication device 49 may allow the electronic device 40 to communicate with other devices wirelessly or wiredly to exchange data. Although Figure 4 the electronic device 40 with various devices is shown, it should be understood that it is not required to implement or have all the shown devices. More or fewer devices may be alternatively implemented or had.
[0149] In particular, according to the embodiments of the present disclosure, the process described above with reference to the flowchart may be implemented as a computer software program. For example, the embodiments of the present disclosure include a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program includes program codes for performing the methods shown in the flowchart. In such an embodiment, the computer program may be downloaded and installed from the network through the communication device 49, or installed from the storage device 48, or installed from the ROM 42. When the computer program is executed by the processing device 41, the above-mentioned functions defined in the methods of the embodiments of the present disclosure are executed.
[0150] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are for illustrative purposes only and are not used to limit the scope of these messages or information.
[0151] The electronic device provided in the embodiments of the present disclosure and the image scene segmentation method provided in the above embodiments belong to the same inventive concept. Technical details not described in detail in this embodiment can be found in the above embodiments, and this embodiment has the same beneficial effects as the above embodiments.
[0152] Embodiment Five
[0153] The embodiments of the present disclosure provide a computer storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the image scene segmentation method provided in the above embodiments.
[0154] It should be noted that the computer-readable medium in the present disclosure may be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. A computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, a computer-readable storage medium may be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, apparatus, or device. In the present disclosure, a computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, in which computer-readable program code is carried. Such a propagated data signal may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium, and this computer-readable signal medium can send, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device. The program code contained on a computer-readable medium can be transmitted by any appropriate medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.
[0155] In some embodiments, the client and the server can communicate using any currently known or future-developed network protocol such as HTTP (HyperText Transfer Protocol), and can be interconnected with digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include local area networks ("LANs"), wide area networks ("WANs"), the Internet (e.g., the Internet), and end-to-end networks (e.g., ad hoc end-to-end networks), as well as any currently known or future-developed networks.
[0156] The above computer-readable medium can be included in the above electronic device; or can exist separately without being assembled into the electronic device.
[0157] The above computer-readable medium carries one or more programs, and when the one or more programs are executed by the electronic device, the electronic device is caused to:
[0158] Computer program code for performing the operations of the present disclosure can be written in one or more programming languages or combinations thereof. The programming languages include, but are not limited to, object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computer (e.g., by using an Internet service provider to connect through the Internet).
[0159] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a portion of code that contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions noted in the blocks may occur in a different order than noted in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or by a combination of dedicated hardware and computer instructions.
[0160] The units involved in the embodiments described in the present disclosure can be implemented in software or in hardware. Among them, the name of the unit does not constitute a limitation on the unit itself in some cases. For example, the first acquisition unit can also be described as "the unit for acquiring at least two Internet protocol addresses".
[0161] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that can be used include: Field Programmable Gate Arrays (FPGA), Application Specific Integrated Circuits (ASIC), Application Specific Standard Products (ASSP), System on a Chip (SOC), Complex Programmable Logic Devices (CPLD), and so on.
[0162] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a Random Access Memory (RAM), a Read-Only Memory (ROM), an Erasable Programmable Read-Only Memory (EPROM or Flash memory), an optical fiber, a portable Compact Disc Read-Only Memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0163] According to one or more embodiments of the present disclosure, [Example 1] provides an image scene segmentation method, which includes: obtaining an intermediate scene segmentation map by performing initial scene segmentation and initial scene fusion processing on the acquired target image; detecting to-be-processed segmentation blocks from the intermediate scene segmentation map; and obtaining a target scene segmentation map of the target image by performing segmentation correction on each of the to-be-processed segmentation blocks.
[0164] According to one or more embodiments of the present disclosure, [Example 2] provides an image scene segmentation method. The step in this method: obtaining an intermediate scene segmentation map by performing initial scene segmentation and initial scene fusion processing on the acquired target image may preferably include: taking the acquired target image as input data and inputting it into a preset scene segmentation network model to obtain an output initial scene segmentation map, where the initial scene segmentation map includes at least one initial segmentation layer; and performing initial scene fusion on each of the initial segmentation layers based on the content labels corresponding to each of the initial segmentation layers to obtain an intermediate scene segmentation map.
[0165] According to one or more embodiments of the present disclosure, [Example 3] provides an image scene segmentation method. The step in this method: performing initial scene fusion on each of the initial segmentation layers based on the content labels corresponding to each of the initial segmentation layers to obtain an intermediate scene segmentation map may preferably include: obtaining the content labels of each of the initial segmentation layers; looking up a preset label category association table to determine the scene branches to which each of the content labels belong; and performing image content fusion on the initial segmentation layers belonging to the same scene branch to obtain a fused intermediate scene segmentation map.
[0166] According to one or more embodiments of the present disclosure, [Example 4] provides an image scene segmentation method. The hidden layer of the preferably used scene segmentation network model in this method includes a set number of residual sub-network models; the residual sub-network models are sequentially connected in a hierarchical order, and there is also a residual connection from one residual sub-network model to another non-adjacent residual sub-network model; each residual sub-network model is composed of a convolutional layer, a batch normalization layer, and a non-linear activation function layer.
[0167] According to one or more embodiments of the present disclosure, [Example 5] provides an image scene segmentation method. The step in this method: detecting to-be-processed segmentation blocks from the intermediate scene segmentation map may preferably include: extracting each intermediate segmentation layer included in the intermediate scene segmentation map; and determining the to-be-processed segmentation blocks of the intermediate scene segmentation map by performing connected component detection on each of the intermediate segmentation layers.
[0168] According to one or more embodiments of the present disclosure, [Example Six] provides an image scene segmentation method. The steps in this method: By performing connected component detection on each of the intermediate segmentation layers to determine the segmentation blocks to be processed in the intermediate scene segmentation map, specifically, it may include: performing binarization processing on each of the intermediate segmentation layers to obtain corresponding binarized segmentation layers; for each binarized segmentation layer, scanning the pixel values of the binarized segmentation layer in a set scanning order; according to the pixel value scanning results, determining each connected region included in the binarized segmentation layer; taking the connected regions with an area smaller than the set area threshold as the segmentation blocks to be processed.
[0169] According to one or more embodiments of the present disclosure, [Example Seven] provides an image scene segmentation method. The steps in this method: By performing segmentation result correction on each of the segmentation blocks to be processed to obtain the target scene segmentation map of the target image, it can be specifically optimized as: for each segmentation block to be processed, performing region dilation processing on the segmentation block to be processed according to the set dilation coefficient to obtain the corresponding dilated segmentation region; based on the dilated segmentation region, determining the target segmentation layer to which the segmentation block to be processed belongs from the intermediate scene segmentation map; performing image fusion on the segmentation block to be processed and the target segmentation layer; taking the intermediate scene segmentation map after the fusion process as the target scene segmentation map of the target image.
[0170] According to one or more embodiments of the present disclosure, [Example Eight] provides an image scene segmentation method. The steps in this method: Based on the dilated segmentation region, determining the target segmentation layer to which the segmentation block to be processed belongs from the intermediate scene segmentation map, which can be specifically optimized to include: obtaining each intermediate segmentation layer included in the intermediate scene segmentation map, and determining the candidate segmentation layers that overlap with the dilated segmentation region; counting the number of pixel points in the regions overlapping with each of the candidate segmentation layers; taking the candidate segmentation layer corresponding to the largest number of pixel points as the target segmentation layer to which the segmentation block to be processed belongs.
[0171] The above description is only for the preferred embodiments of the present disclosure and the explanation of the applied technical principles. Those skilled in the art should understand that the scope of disclosure involved in the present disclosure is not limited to the technical solutions formed by the specific combination of the above technical features, and should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above disclosure concept. For example, the technical solutions formed by mutually replacing the above features with the (but not limited to) technical features with similar functions disclosed in the present disclosure.
[0172] In addition, although the operations are depicted in a particular order, this should not be construed as requiring that the operations be performed in the particular order shown or in sequential order. In certain environments, multitasking and parallel processing may be advantageous. Similarly, although some specific implementation details are included in the foregoing discussion, these should not be construed as limitations on the scope of the present disclosure. Certain features that are described in the context of separate embodiments may also be implemented in combination in a single embodiment. Conversely, the various features that are described in the context of a single embodiment may also be implemented separately or in any suitable sub-combination in multiple embodiments.
[0173] Although the subject matter has been described in language specific to structural features and / or methodological acts, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are merely example forms of implementing the claims.
Claims
1. An image scene segmentation method, characterized in that, Including: Performing initial scene segmentation and initial scene fusion processing on the acquired target image to obtain an intermediate scene segmentation map; Detecting a segmentation block to be processed from the intermediate scene segmentation map, where the intermediate scene segmentation map includes multiple segmentation layers obtained by segmenting image content according to scene categories, and the segmentation block to be processed is a fragmented segmentation block with incorrect segmentation of the scene category to which the image content in the segmentation layer belongs; Performing segmentation correction on each of the segmentation blocks to be processed to obtain a target scene segmentation map of the target image; The performing segmentation correction on each of the segmentation blocks to be processed to obtain a target scene segmentation map of the target image includes: Determining a target segmentation layer to which the segmentation block to be processed correctly belongs, and fusing the segmentation block to be processed into the target segmentation layer; Based on each of the fused target segmentation layers, constituting a target scene segmentation map of the target image, where the target segmentation layer includes only image content belonging to the same scene category.
2. The method according to claim 1, wherein The performing initial scene segmentation and initial scene fusion processing on the acquired target image to obtain an intermediate scene segmentation map includes: Taking the acquired target image as input data and inputting it into a preset scene segmentation network model to obtain an output initial scene segmentation map, where the initial scene segmentation map includes at least one initial segmentation layer; Performing initial scene fusion on each of the initial segmentation layers based on the content labels corresponding to each of the initial segmentation layers to obtain an intermediate scene segmentation map.
3. The method according to claim 2, wherein The performing initial scene fusion on each of the initial segmentation layers based on the content labels corresponding to each of the initial segmentation layers to obtain an intermediate scene segmentation map includes: Obtaining the content labels of each of the initial segmentation layers; Searching a pre-set label category association table to determine the scene branches to which each of the content labels belongs; Performing image content fusion on the initial segmentation layers belonging to the same scene branch to obtain a fused intermediate scene segmentation map.
4. The method according to claim 2, wherein The hidden layer of the scene segmentation network model includes a set number of residual sub-network models; Each of the residual sub-network models is connected in sequence according to the hierarchical order, and there is also a residual connection from one residual sub-network model to another non-adjacent residual sub-network model; Each residual sub-network model is composed of a convolutional layer, a batch normalization layer, and a non-linear activation function layer.
5. The method according to claim 1, wherein The detecting a segmentation block to be processed from the intermediate scene segmentation map includes: Extracting each intermediate segmentation layer included in the intermediate scene segmentation map; Determining the segmentation block to be processed in the intermediate scene segmentation map by performing connected component detection on each of the intermediate segmentation layers.
6. The method according to claim 5, wherein The determining the segmentation block to be processed in the intermediate scene segmentation map by performing connected component detection on each of the intermediate segmentation layers includes: Performing binarization processing on each of the intermediate segmentation layers to obtain corresponding binarized segmentation layers; For each binarized segmentation layer, scanning the pixel values of the binarized segmentation layer in a set scanning order; Determining each connected region included in the binarized segmentation layer according to the pixel value scanning results; Taking the connected regions with an area smaller than a set area threshold as the segmentation blocks to be processed.
7. The method according to claim 1, wherein Obtaining the target scene segmentation map of the target image by performing segmentation result correction on each of the to-be-processed segmentation blocks includes: For each to-be-processed segmentation block, performing regional dilation processing on the to-be-processed segmentation block according to a set dilation coefficient to obtain a corresponding segmentation dilation region; Based on the segmentation dilation region, determining the target segmentation layer to which the to-be-processed segmentation block belongs from the intermediate scene segmentation map; Performing image fusion on the to-be-processed segmentation block and the target segmentation layer; Taking the intermediate scene segmentation map after the fusion processing as the target scene segmentation map of the target image.
8. The method according to claim 7, characterized in that The determining, based on the segmentation dilation region, the target segmentation layer to which the to-be-processed segmentation block belongs from the intermediate scene segmentation map includes: Obtaining each intermediate segmentation layer included in the intermediate scene segmentation map and determining candidate segmentation layers that overlap with the segmentation dilation region; Counting the number of pixel points in the regions overlapping with each of the candidate segmentation layers; Taking the candidate segmentation layer corresponding to the largest number of pixel points as the target segmentation layer to which the to-be-processed segmentation block belongs.
9. An image scene segmentation device, characterized in that, Includes: An initial processing module for obtaining an intermediate scene segmentation map by performing initial scene segmentation and initial scene fusion processing on the acquired target image; An information determination module for detecting to-be-processed segmentation blocks from the intermediate scene segmentation map, where the intermediate scene segmentation map includes multiple segmentation layers obtained by segmenting image content according to scene categories, and the to-be-processed segmentation blocks are fragmented segmentation blocks with incorrect segmentation of the scene categories to which the image content in the segmentation layers belongs; A segmentation correction module for obtaining the target scene segmentation map of the target image by performing segmentation correction on each of the to-be-processed segmentation blocks; Specifically, the segmentation correction module is configured to: determine the target segmentation layer to which the to-be-processed segmentation block correctly belongs and fuse the to-be-processed segmentation block into the target segmentation layer; based on each of the target segmentation layers after fusion, constitute the target scene segmentation map of the target image, and the target segmentation layer includes image content belonging only to the same scene category.
10. An electronic device, characterized in that, The electronic device includes: One or more processors; A storage device for storing one or more programs, When the one or more programs are executed by the one or more processors, enabling the one or more processors to implement the image scene segmentation method according to any one of claims 1-8.
11. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the image scene segmentation method according to any one of claims 1-8.
Citation Information
Patent Citations
Scene segmentation method, device and equipment and computer readable storage medium
CN113470048A
Tooth image segmentation method and device thereof
CN113506301A