Image segmentation method, electronic device, and storage medium

CN122820747APending Publication Date: 2026-09-25ZHEJIANG DAHUA TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611290750.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-08-25
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

然而,该自注意力机制在应用中,其计算复杂度与像素数量的平方成正比,在处理高分辨率图像时导致巨大的计算与内存开销,严重限制了实时应用与硬件部署

Benefits of technology

[0009]上述方案,获取待分割图像的第一特征图以及待分割图像中各分割类别的初始分割概率;根据第一特征图和初始分割概率进行分割概率更新处理,得到待分割图像中各分割类别的目标分割概率;根据目标分割概率进行图像分割处理,得到图像分割结果,由此通过第一特征图对初始分割概率进行更新,使得分割概率更新过程能够自适应地捕捉像素间的语义关联,从而便于提高分割边界的清晰度,进而在保持分割精度的同时降低计算复杂度,实现低算力开销下的图像分割。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122820747A_ABST
    Figure CN122820747A_ABST
Patent Text Reader

Abstract

The application discloses an image segmentation method, an electronic device and a storage medium. The image segmentation method comprises the following steps: obtaining a first feature map of an image to be segmented and initial segmentation probabilities of each segmentation category in the image to be segmented, wherein the initial segmentation probabilities are obtained based on the first feature map; performing segmentation probability updating processing according to the first feature map and the initial segmentation probabilities to obtain target segmentation probabilities of each segmentation category in the image to be segmented; and performing image segmentation processing according to the target segmentation probabilities to obtain an image segmentation result. According to the above scheme, the calculation complexity can be reduced while the segmentation accuracy is maintained, and the image segmentation under low calculation power consumption can be realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image processing, and in particular to an image segmentation method, an electronic device, and a storage medium. Background Technology

[0002] With the rapid development of artificial intelligence technology, in the field of image segmentation, the introduction of the visual Transformer model with self-attention mechanism at its core, which utilizes the establishment of long-range dependencies between pixels, can significantly improve the accuracy of image segmentation results. However, in application, the computational complexity of this self-attention mechanism is proportional to the square of the number of pixels, resulting in huge computational and memory overhead when processing high-resolution images, which severely limits real-time applications and hardware deployment.

[0003] Therefore, there is an urgent need for a method that can achieve effective image segmentation with low overhead. Summary of the Invention

[0004] This application provides at least one image segmentation method, electronic device, and storage medium, which can reduce computational complexity while maintaining segmentation accuracy, and achieve real-time image segmentation with low computational overhead.

[0005] This application provides an image segmentation method, which includes: obtaining a first feature map of an image to be segmented and an initial segmentation probability of each segmentation category in the image to be segmented, wherein the initial segmentation probability is obtained based on the first feature map; performing segmentation probability update processing based on the first feature map and the initial segmentation probability to obtain a target segmentation probability of each segmentation category in the image to be segmented; and performing image segmentation processing based on the target segmentation probability to obtain an image segmentation result.

[0006] This application provides an image segmentation apparatus, comprising: an acquisition module, an update module, and a segmentation module; the acquisition module is used to acquire a first feature map of an image to be segmented and an initial segmentation probability of each segmentation category in the image to be segmented, wherein the initial segmentation probability is obtained based on the first feature map; the update module is used to perform segmentation probability update processing based on the first feature map and the initial segmentation probability to obtain a target segmentation probability of each segmentation category in the image to be segmented; the segmentation module is used to perform image segmentation processing based on the target segmentation probability to obtain an image segmentation result.

[0007] This application provides an electronic device, including a memory and a processor, wherein the processor is used to execute program instructions stored in the memory to implement the above-described image segmentation method.

[0008] This application provides a computer-readable storage medium storing program instructions thereon, which, when executed by a processor, implement the above-described image segmentation method.

[0009] The above scheme obtains the first feature map of the image to be segmented and the initial segmentation probability of each segmentation category in the image to be segmented; performs segmentation probability update processing based on the first feature map and the initial segmentation probability to obtain the target segmentation probability of each segmentation category in the image to be segmented; performs image segmentation processing based on the target segmentation probability to obtain the image segmentation result. Thus, the initial segmentation probability is updated through the first feature map, so that the segmentation probability update process can adaptively capture the semantic relationship between pixels, thereby improving the clarity of the segmentation boundary, and reducing the computational complexity while maintaining the segmentation accuracy, achieving image segmentation with low computational overhead.

[0010] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this application. Attached Figure Description

[0011] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the specification, serve to explain the technical solutions of this application.

[0012] Figure 1 This is a first flowchart illustrating an exemplary embodiment of the image segmentation method of this application; Figure 2 This is a schematic diagram of the second process of an exemplary embodiment of the image segmentation method of this application; Figure 3 This is a schematic diagram of the framework for obtaining the target segmentation probability in an exemplary embodiment of the image segmentation method of this application; Figure 4 This is a schematic diagram of the third process of an exemplary embodiment of the image segmentation method of this application; Figure 5 This is a schematic diagram of the framework for obtaining the updated probability of the current update round in an exemplary embodiment of the image segmentation method of this application; Figure 6 This is a schematic diagram of the fourth process of an exemplary embodiment of the image segmentation method of this application; Figure 7 This is another schematic diagram of the framework for obtaining the updated probability of the current update round in an exemplary embodiment of the image segmentation method of this application; Figure 8 This is a schematic diagram of the structure of an embodiment of the image segmentation apparatus of this application; Figure 9 This is a schematic diagram of the structure of an embodiment of the electronic device of this application; Figure 10 This is a schematic diagram of the structure of an embodiment of the computer-readable storage medium of this application. Detailed Implementation

[0013] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. It is understood that the specific embodiments described herein are only for explaining this application and not for limiting it. Furthermore, it should be noted that, for ease of description, only the parts related to this application are shown in the accompanying drawings, not all structures. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are within the scope of protection of this application.

[0014] In the following description, specific details such as particular system architectures, interfaces, and technologies are presented for illustrative purposes rather than for limiting purposes, in order to provide a thorough understanding of this application.

[0015] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that includes a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus. The term "and / or" is merely a description of the association of related objects, indicating that three relationships can exist; for example, A and / or B can represent: A alone, A and B simultaneously, and B alone. Additionally, the character " / " in this document generally indicates that the preceding and following related objects are in an "or" relationship. Furthermore, "many" in this document means two or more. Furthermore, the term "at least one" in this document means any combination of at least two of any one or more of a plurality of elements, for example, including at least one of A, B, and C, which can mean including any one or more elements selected from the set consisting of A, B, and C. Additionally, the term "several" in this document means one or more.

[0016] This application provides several image segmentation methods and apparatuses. The application scenarios of these image segmentation methods include, but are not limited to, image segmentation scenarios. The image segmentation methods of this application can involve various business domains when performing image segmentation tasks. For example, these business domains may include, but are not limited to, object recognition (such as vehicle and obstacle recognition), animal recognition (such as cat and dog recognition), etc. The executing entity of the image segmentation method can be an image segmentation apparatus. For example, the image segmentation apparatus can be located in a terminal device, server, or other processing device. The terminal device can be a user equipment (UE), mobile device, user terminal, terminal, cellular phone, cordless phone, personal digital assistant (PDA), handheld device, computing device, in-vehicle device, etc. In some possible implementations, the image segmentation method can be implemented by a processor calling computer-readable instructions stored in memory.

[0017] Please see Figure 1 , Figure 1 This is a first flowchart illustrating an exemplary embodiment of the image segmentation method of this application. Specifically, the image segmentation method may include the following steps: S11: Obtain the first feature map of the image to be segmented and the initial segmentation probability of each segmentation category in the image to be segmented.

[0018] The image to be segmented represents the image that requires segmentation processing. For example, a frame from a data stream or video stream acquired by an image acquisition device is used for image segmentation. The first feature map represents the feature map obtained by feature extraction from the image to be segmented. The initial segmentation probability is based on the first feature map. The segmentation category represents the target type for image segmentation, and the segmentation category is set according to the business domain of image segmentation. For example, in the business domain of object recognition (such as image segmentation of traffic scenes), several segmentation categories include, but are not limited to, vehicles, obstacles, plants, traffic lights, etc.

[0019] The method for obtaining the first feature map of the image to be segmented can be as follows: first, obtain the image to be segmented; then, perform preset feature extraction on the image to be segmented to obtain the first feature map of the image to be segmented, including: inputting the image to be segmented into an image feature extraction module, and obtaining the first feature map of the image to be segmented output by the image feature extraction module. The image feature extraction module can be equipped with a preset feature extraction network. Specifically, the preset feature extraction network can be a convolutional neural network, a recurrent neural network, a long short-term memory network, a gated recurrent unit, a BiGRU neural network, or other feature extraction networks. In some application scenarios, the preset feature extraction network can be a preset perceptron network. The preset perceptron network can be obtained from at least one perceptron module. Each of the at least one perceptron module can be a multilayer perceptron (MLP).

[0020] In some application scenarios, the initial segmentation probabilities of each segmentation category in the image to be segmented can be obtained by: mapping the first feature map to obtain the initial segmentation probabilities of each segmentation category in the image to be segmented; or, performing feature extraction on the first feature map to obtain the second feature map; and then mapping the second feature map to obtain the initial segmentation probabilities of each segmentation category in the image to be segmented.

[0021] In other application scenarios, the initial segmentation probabilities of each segmentation category in the image to be segmented can be obtained as follows: Input the image to be segmented into an image feature extraction module to obtain a second feature map of the image to be segmented, output by the image feature extraction module; perform mapping processing on the first feature map to obtain a first segmentation probability of each segmentation category in the image to be segmented; perform mapping processing on the second feature map to obtain the initial segmentation probability of each segmentation category in the image to be segmented, including: inputting the second feature map into a mapping module to obtain a second segmentation probability of each segmentation category in the image to be segmented, output by the mapping module; and combining the first and second segmentation probabilities to obtain the initial segmentation probability of each segmentation category in the image to be segmented. The settings of the image feature extraction module can refer to the above content. The mapping module is equipped with a preset activation function or a fully connected layer. The preset activation function can be ReLU, Sigmoid, Softmax, etc. For example, taking a mapping module equipped with a Sigmoid activation function as an example. The second feature map and the first feature map have different scales. For example, the width of the second feature map is smaller than the width of the first feature map or the image to be segmented, and the height of the second feature map is smaller than the height of the first feature map or the image to be segmented.

[0022] In other application scenarios, S11 above can also be: performing multi-scale feature extraction on the image to be segmented to obtain a first feature map and a second feature map; and mapping the second feature map to obtain the initial segmentation probabilities of each segmentation category in the image to be segmented. Here, the second feature map has a different scale than the first feature map. The mapping process for the second feature map can be referenced above and will not be repeated here. Multi-scale feature extraction can be performed using at least two cascaded convolutional neural networks. These at least two convolutional neural networks can use the same convolutional kernel, but with different strides to make the scales of the first and second feature maps different.

[0023] S12: Perform segmentation probability update processing based on the first feature map and the initial segmentation probability to obtain the target segmentation probability of each segmentation category in the image to be segmented.

[0024] The target segmentation probability represents the initial segmentation probability after the segmentation probability update process.

[0025] In some application scenarios, the initial segmentation probability is obtained by mapping the second feature map. S12 above can be: inputting the first feature map into the mapping module to obtain the advanced segmentation probability of each segmentation category in the image to be segmented output by the mapping module; combining the initial segmentation probability of the corresponding segmentation category and the advanced segmentation probability of the corresponding segmentation category to obtain the target segmentation probability of the corresponding segmentation category. For example, the initial segmentation probability and the advanced segmentation probability are weighted and fused to obtain the target segmentation probability of the corresponding segmentation category.

[0026] In other application scenarios, S12 above can be: calculating the segmentation probability weights of each segmentation category based on each sub-feature map in the first feature map, including: traversing each sub-feature map in the first feature map, taking the sub-feature maps in the neighborhood of the currently traversed sub-feature map as candidate sub-feature maps; mapping the convolution result between the current sub-feature map and the candidate sub-feature maps to the sub-segmentation probability weights of each segmentation category in the current sub-feature map; until all sub-feature maps in the first feature map have been traversed, obtaining the sub-segmentation probability weights of each segmentation category to which each sub-feature map in the first feature map belongs; for each segmentation category, determining the segmentation probability weight of that segmentation category based on the statistical value of the sub-segmentation probability weights of each sub-feature map in that segmentation category. The statistical value can be the average, weighted average, maximum, mode, median, etc. For each segmentation category, adjusting the initial segmentation probability of that segmentation category based on the segmentation probability weight of that segmentation category to obtain the target segmentation probability of that segmentation category.

[0027] S13: Perform image segmentation processing based on the target segmentation probability to obtain the image segmentation result.

[0028] The image segmentation result represents the predicted segmentation category of each pixel in the image to be segmented, and the predicted segmentation category is one of the various segmentation categories.

[0029] In some application scenarios, S13 can be: for each pixel in the image to be segmented, directly use the segmentation category with the highest target segmentation probability as the predicted segmentation category for that pixel. Alternatively, use the segmentation category with the highest target segmentation probability for each pixel as the candidate segmentation category for that pixel, and adjust the candidate segmentation categories for each pixel according to a segmentation category smoothing strategy to obtain the predicted segmentation category for each pixel in the image to be segmented. The segmentation category smoothing strategy is used to ensure that the segmentation categories of adjacent pixels in the image to be segmented remain consistent, making the segmentation result smoother. In other application scenarios, S13 can be: substitute the target segmentation probability into the energy function, and determine the predicted segmentation category for each pixel in the image to be segmented based on the energy function value.

[0030] The above scheme obtains the first feature map of the image to be segmented and the initial segmentation probability of each segmentation category in the image to be segmented; performs segmentation probability update processing based on the first feature map and the initial segmentation probability to obtain the target segmentation probability of each segmentation category in the image to be segmented; performs image segmentation processing based on the target segmentation probability to obtain the image segmentation result. Thus, the initial segmentation probability is updated through the first feature map, so that the segmentation probability update process can adaptively capture the semantic relationship between pixels, thereby improving the clarity of the segmentation boundary, and reducing the computational complexity while maintaining the segmentation accuracy, achieving image segmentation with low computational overhead.

[0031] Specifically, S12 above can be: updating the initial segmentation probability at least once based on the first feature map to obtain the target segmentation probability. The at least once update can be one round or multiple rounds, with the total number of update rounds dynamically set according to the image segmentation accuracy. The input for each update round is the current probability to be updated, which is either the initial segmentation probability or the updated probability from the previous update round.

[0032] In some application scenarios, S12 above can also be: For each round of updating the current probability to be updated, the following steps are performed: The logarithmic result of the current probability to be updated is used as the current probability base for the current update round. Based on the deviation between the current probability base and the obtained preset probability, the updated probability for the current update round is determined, including: using the difference between the current probability base and the preset probability deviation as the target difference; substituting the target difference into a preset exponential function to obtain the target function value; normalizing the target function value to obtain the updated probability for the current update round. In response to the updated probability for the current update round satisfying the preset update condition, the updated probability for the current update round is used as the target segmentation probability. The preset probability deviation can be the historical probability deviation of historical frame images, where the acquisition time of the historical frame images is earlier than the acquisition time of the image to be segmented, and the historical frame images and the image to be segmented are images acquired by the image acquisition device from the same acquisition area. The historical probability deviation of historical frame images is constructed as follows: based on the first feature map of the historical frame images and the current probability to be updated of the historical frame images, the historical probability deviation of the historical frame images is obtained. The first feature map of the historical frame image is similar to the first feature map of the image to be segmented. Preset feature extraction is performed on the historical frame image to obtain its first feature map. The current probability to be updated in the historical frame image is similar to the current probability to be updated in the image to be segmented, being either the initial segmentation probability of the historical frame image under the same update rounds or the updated probability obtained by updating the initial segmentation probability of the historical frame image. Multi-scale feature extraction is performed on the historical frame image to obtain first feature maps and second feature maps of the historical frame image at different scales. The preset probability deviation characterizes the difference between the current sub-probabilities of each sub-feature map in the second feature map of the historical frame image.

[0033] In some embodiments, S12 may include the following steps: when the current update round is the first update round, use the initial segmentation probability as the current probability to be updated. Alternatively, when the current update round is not the first update round, use the updated probability of the previous update round as the current probability to be updated. For each round of updating the current probability to be updated, perform the following steps: use the logarithmic result of the current probability to be updated as the current probability base for the current update round. Determine the updated probability of the current update round based on the current probability base, the first feature map, and the current probability to be updated. In response to the updated probability of the current update round satisfying a preset update condition, use the updated probability of the current update round as the target segmentation probability.

[0034] The current probability basis represents the probability basis of the current update round, used to represent the logarithmic result of the current probability to be updated. The updated probability of the current update round represents the update result of one round of segmentation probability update processing on the current probability to be updated in the current update round. The preset update condition represents the update termination condition for at least one round of updates to the initial segmentation probability.

[0035] Specifically, the method for determining the current probability base for the current update round can refer to the following formula (1): Formula (1); in, The current probability base represents the current update round, specifically including the segmentation probability base for each segmentation category in the current probability to be updated. This represents the current probability to be updated in the current update round, specifically including the segmentation probability of each segmentation category within the current probability to be updated. The segmentation probability base of each segmentation category is positively correlated with the segmentation probability of that category; the larger the segm value, the larger the base, while a negative base represents the energy in the energy function. Conversely, by optimizing to reduce the energy contained in the negative base, a larger segm value is obtained, resulting in higher accuracy in classification using the segm value, and thus clearer segmentation boundaries.

[0036] In some application scenarios, the steps described above for determining the updated probability of the current update round based on the current probability base, the first feature map, and the current probability to be updated include: performing convolution processing on the first feature map to obtain the convolution result of the first feature map; adjusting the obtained preset probability deviation based on the convolution result of the first feature map to obtain the current probability deviation of the current update round; and determining the updated probability of the current update round based on the current probability base and the current probability deviation, including: using the difference between the current probability base and the current probability deviation as the target difference; substituting the target difference into a preset exponential function to obtain the target function value; and normalizing the target function value to obtain the updated probability of the current update round.

[0037] The preset update condition can be that the number of updates to which the updated probability belongs in the current update round reaches a threshold, or the update round to which the current update round belongs reaches a round threshold, or the updated probability of the current update round reaches a probability threshold. For example, in response to the number of updates to which the updated probability belongs in the current update round reaches a threshold, the updated probability of the current update round is used as the target segmentation probability.

[0038] For example, if the updated probability in the current update round does not meet the preset update conditions, the updated probability in the current update round is used as the current probability to be updated in the next update round. If the updated probability in the current update round meets the preset update conditions, the overall update of the initial segmentation probability ends, and the updated probability in the current update round is used as the target segmentation probability.

[0039] Please see Figure 2 , Figure 2 This is a schematic diagram of the second process of an exemplary embodiment of the image segmentation method of this application.

[0040] In some embodiments, the current probability to be updated includes the current sub-probabilities to which several sub-feature maps in the second feature map belong. The second feature map is obtained by feature extraction from the first feature map. The scale of the first feature map is different from that of the second feature map. The step of determining the updated probability of the current update round based on the current probability base, the first feature map, and the current probability to be updated may include S21 and S22.

[0041] Specifically, S11 above can also include: performing multi-scale feature extraction on the image to be segmented to obtain a first feature map and a second feature map, including: performing first feature extraction on the image to be segmented to obtain a first feature map of the image to be segmented; performing second feature extraction on the first feature map of the image to be segmented to obtain a second feature map of the image to be segmented. The second feature map is then mapped to obtain the initial segmentation probabilities for each segmentation category in the image to be segmented. The scale of the first feature map is different from that of the second feature map. The first feature extraction and the second feature extraction are used to perform feature extraction at different scales on the image to be segmented. The sub-feature maps in the second feature map represent local feature blocks in the second feature map corresponding to different spatial locations or different semantic channels.

[0042] S21: Determine the current probability deviation for the current update round based on the first feature map and the current probability to be updated.

[0043] The current probability bias characterizes the difference between the current sub-probabilities of each sub-feature map in the second feature map. For example, the current probability bias characterizes the distribution or degree of conflict of segmentation probabilities of the same segmentation category in a local region, such as at image edges or in regions with complex textures. Sub-feature maps at different locations in the second feature map may have similar semantic features but belong to different segmentation categories, or the distribution of segmentation probabilities of sub-feature maps at adjacent locations may fluctuate significantly.

[0044] In some embodiments, the current sub-probability of a sub-feature map includes the current sub-class probability of the sub-feature map in each segmentation category. S21 described above may include the following steps: performing convolution processing on the first feature map to obtain a first target convolution result of the first feature map; obtaining the current probability deviation based on the first target convolution result and the current probability to be updated; or, obtaining the current probability deviation based on the first target convolution result, the current probability to be updated, and the preset association weights between several obtained segmentation categories.

[0045] The first target convolution result of the first feature map represents the convolution result obtained by performing at least one convolution process on the first feature map. Each convolution process can be implemented by convolution kernels of the same or different scales.

[0046] In some embodiments, the step of performing convolution processing on the first feature map to obtain a first target convolution result of the first feature map may include the following steps: performing a convolution operation on the first feature map using at least one Gaussian kernel and at least one bilateral kernel to obtain a first candidate convolution result of the first feature map. Alternatively, performing a convolution operation on the first feature map using at least one Gaussian kernel to obtain a second candidate convolution result of the first feature map. Alternatively, performing a convolution operation on the first feature map using a preset convolution kernel to obtain a third candidate convolution result of the first feature map. One of the first candidate convolution result, the second candidate convolution result, and the third candidate convolution result is used as the first target convolution result.

[0047] Gaussian kernels and bilateral kernels are used to perform convolution processing on each sub-feature map in the first feature map to calculate the feature difference between the sub-feature maps in the first feature map. It is understood that in the bilateral kernel, this application uses features to calculate weights, rather than traditional image brightness. The Gaussian kernel and the bilateral kernel have different scales. At least one Gaussian kernel can be one or more Gaussian kernels; for example, the number of Gaussian kernels can be three, and the scales of the Gaussian kernels are different. At least one bilateral kernel can be one or more bilateral kernels; for example, the number of bilateral kernels can be three, and the scales of the bilateral kernels are different. The number of Gaussian kernels and the number of bilateral kernels can be the same or different. The preset convolution kernel can be used for depthwise convolution; for example, the convolution type of the preset convolution kernel can be a 3×3 standard convolution. In other application scenarios, the convolution type of the preset convolution kernel can also be other convolution types, such as grouped convolution, depthwise separable convolution, dilated convolution, transposed convolution, etc. The first candidate convolution result represents the convolution result obtained by convolution operation of the first feature map using at least one Gaussian kernel and at least one bilateral kernel. The second candidate convolution result represents the convolution result obtained by performing a convolution operation on the first feature map using at least one Gaussian kernel. The third candidate convolution result represents the convolution result obtained by performing a convolution operation on the first feature map using a preset convolution kernel. The computational resources consumed in obtaining the first, second, and third candidate convolution results from the first feature map, in descending order, are: obtaining the first candidate convolution result through convolution operation, obtaining the second candidate convolution result through convolution operation, and obtaining the third candidate convolution result through convolution operation.

[0048] In some application scenarios, the first target convolution result of the first feature map can be obtained by performing a convolution operation on the first feature map using at least one Gaussian kernel and at least one bilateral kernel to obtain a first candidate convolution result of the first feature map, and then using the first candidate convolution result as the first target convolution result. The Gaussian kernels and bilateral kernels can be cascaded, and their processing order is not limited. One of the Gaussian kernels or bilateral kernels is taken as the first processing kernel. The first feature map is input into the convolutional network corresponding to the first processing kernel to obtain the convolution result processed by the first processing kernel. The convolution result processed by the first processing kernel is then input into the next processing kernel cascaded with the first processing kernel, until the convolution result processed by the last processing kernel is obtained, and the convolution result processed by the last processing kernel is used as the first candidate convolution result.

[0049] In other application scenarios, the first target convolution result of the first feature map can be obtained by performing a convolution operation on the first feature map using at least one Gaussian kernel to obtain a second candidate convolution result of the first feature map, and using the second candidate convolution result as the first target convolution result. The Gaussian kernels can be cascaded, and their processing order is not limited. The first feature map is input into the first Gaussian kernel to obtain the convolution result output by the first Gaussian kernel, and the convolution result output by the first Gaussian kernel is input into the next Gaussian kernel cascaded with the first Gaussian kernel, until the convolution result output by the last Gaussian kernel is obtained, and the convolution result of the last Gaussian kernel is used as the second candidate convolution result.

[0050] In other application scenarios, the first target convolution result of the first feature map can be obtained by performing a convolution operation on the first feature map using a preset convolution kernel to obtain the third candidate convolution result of the first feature map, and then using the third candidate convolution result as the first target convolution result.

[0051] In other application scenarios, a target convolution kernel is selected from the first, second, and third convolution kernels according to the kernel selection instruction; the first feature map is then convolved using the target convolution kernel to obtain the first target convolution result. The first convolution kernel includes at least one Gaussian kernel and at least one bilateral kernel; the second convolution kernel includes at least one Gaussian kernel; and the third convolution kernel is a preset convolution kernel.

[0052] Specifically, the process of performing convolution on the first feature map using a bilateral kernel can be referred to in the following formula (2), and the process of performing convolution on the first feature map using a Gaussian kernel can be referred to in the following formula (3): Formula (2); Formula (3); in, This represents the position of any sub-feature map in the first feature map. Except for the first feature graph The positions of other sub-feature maps besides the first feature map, for example, representing the positions in the first feature map. The location within the neighborhood of the first target convolution result includes the convolution results of each sub-feature map in the first feature map. Representing the first feature map The corresponding first target convolution result. The position in the first feature map is Sub-feature map. The position in the first feature map is The sub-feature map. Fea is the feature extracted from the network. It represents the product of two vectors, that is, the corresponding elements of the vectors are multiplied and summed. express In the position of The neighborhood range is defined by the center. For example, with three bilateral kernels and three Gaussian kernels, the neighborhood ranges within the depth are set to 11, 17, and 23 respectively. A larger neighborhood results in clearer segmentation boundaries, but also consumes more computational resources. The neighborhood range is dynamically set based on the image segmentation accuracy. With three bilateral kernels and three Gaussian kernels... and There are three parameters, obtained through deep learning. Formula (2) represents performing a convolution operation on the first feature map using a two-sided kernel. This represents the learning parameters corresponding to the two-sided kernel. Formula (3) represents performing a convolution operation on the first feature map using a Gaussian kernel. This represents the learning parameters corresponding to the Gaussian kernel.

[0053] It is understandable that if multiple bilateral kernels and multiple Gaussian kernels, or multiple Gaussian kernels, are used to perform convolution operations on the first feature map to obtain the first target convolution result, the input of each current processing kernel is the first feature map or the convolution result after being processed by the previous processing kernel. The position of the input in the current processing core is Sub-feature map. The position of the input in the current processing core is Sub-feature map.

[0054] In some application scenarios, the step of obtaining the current probability deviation based on the first target convolution result and the current probability to be updated can include: using the fusion result between the first target convolution result and the current probability to be updated as the current probability deviation. For example, the first target convolution result and the current probability to be updated can be input into a fusion module for feature fusion processing to obtain the current probability deviation. The fusion module is equipped with a preset feature fusion network, which may include, but is not limited to, a feature concatenation network, an attention-based fusion network, or a bilinear pooling-based fusion network.

[0055] In some embodiments, the step of obtaining the current probability deviation based on the first target convolution result and the current probability to be updated may include the following steps: performing convolution processing on the current probability to be updated to obtain a second target convolution result for the current probability to be updated; and obtaining the current probability deviation based on the first target convolution result, the second target convolution result, and the current probability to be updated.

[0056] The second target convolution result represents the convolution result obtained by performing a convolution operation on the current probability to be updated. For example, a second target convolution result is obtained by performing a convolution operation on the current probability to be updated using a preset convolution kernel. The preset convolution kernel can be a convolution kernel used for depthwise convolution. For example, the convolution type of the preset convolution kernel can be a 3×3 standard convolution.

[0057] In some applications, the product of the first target convolution result, the second target convolution result, and the current probability to be updated is directly used as the current probability bias. Alternatively, the first target convolution result, the second target convolution result, and the current probability to be updated are weighted and fused to obtain the current probability bias.

[0058] For example, from the perspective of saving computing resources, the above S21 may include: performing a convolution operation on the first feature map through a preset convolution kernel to obtain a third candidate convolution result of the first feature map, and using the third candidate convolution result as the first target convolution result; performing convolution processing on the current probability to be updated to obtain a second target convolution result of the current probability to be updated; and obtaining the current probability deviation based on the first target convolution result, the second target convolution result, and the current probability to be updated. Specifically, the process of obtaining the current probability deviation through the third candidate convolution result, the second target convolution result, and the current probability to be updated can be referred to the following formula (4): Formula (4); in, This represents the current probability deviation. This represents the current probability that needs to be updated. This represents the first target convolution result obtained by convolving the first feature map. This represents the second target convolution result obtained by performing a convolution operation on the current probability to be updated. This represents multiplication, which means multiplying corresponding elements of two vectors and then summing them. For example, multiplying corresponding elements of two vectors... and Multiply corresponding elements of two vectors and sum them.

[0059] In some embodiments, the preset association weights include sub-association weights between any two segmentation categories. The sub-association weights correspond to the probability that the two segmentation categories will appear simultaneously in the same image region. For example, the sub-association weights are inversely proportional to the probability that the two segmentation categories to which the sub-association weights belong will appear simultaneously in the same image region.

[0060] Specifically, before the steps of obtaining the first feature map of the image to be segmented and the initial segmentation probabilities of each segmentation category in the image to be segmented, the image segmentation method may further include the following steps: obtaining the actual segmentation results of several preset image regions in the sample segmentation image; determining the statistical frequency of several co-occurring segmentation category pairs based on the common segmentation categories in the actual segmentation results of each preset image region; and constructing the sub-association weights between the two segmentation categories to which the corresponding co-occurring segmentation category pairs belong based on the statistical frequency of each co-occurring segmentation category pair.

[0061] The sample segmentation image can be an image of the same type as the image to be segmented, or an image acquired by the image acquisition device of the image to be segmented. The sample segmentation image represents the labeled image after segmentation result annotation. The number of sample segmentation images can be one frame or multiple frames. Several preset image regions are image regions obtained by dividing the sample segmentation image into regions. The actual segmentation result of the preset image region represents the annotation result of the actual segmentation category existing in that preset image region.

[0062] Co-occurrence segmentation category pairs include combinations of any two segmentation categories from a set of segmentation categories. The statistical frequency of a co-occurrence segmentation category pair represents the total number of times the two segmentation categories in that pair appear together in the actual segmentation results of the same preset image region across all sample segmented images.

[0063] Specifically, the above-mentioned determination of the statistical frequency of several co-occurring segmentation category pairs based on the common segmentation categories in the actual segmentation results of each preset image region includes: for each sample segmented image, performing the following steps: traversing each preset image region in the sample segmented image; in response to the number of actual segmentation results of the current image region being greater than or equal to two, combining multiple labeled segmentation categories in the actual segmentation results of the current image region pairwise to obtain at least one set of candidate segmentation category pairs; for each candidate segmentation category pair, incrementing the current statistical frequency of the candidate segmentation category pair by one to obtain a new statistical frequency for each candidate segmentation category pair; until all preset image regions in all sample images have been traversed, obtaining the final statistical frequency of each candidate segmentation category pair. In some application scenarios, for each candidate segmentation category pair, the final statistical frequency of the candidate segmentation category pair is directly used as the statistical frequency of the co-occurring segmentation category pair to which the candidate segmentation category pair belongs. In other application scenarios, for each candidate segmentation category pair, the correction coefficient corresponding to the sample segmentation image is determined according to the segmentation scale when dividing the sample segmentation image into regions. The segmentation scale represents the number of blocks in which the sample segmentation image is divided into regions. Different segmentation scales correspond to different correction coefficients. The more blocks in the segmentation scale, the smaller the correction coefficient, and vice versa. After traversing each preset image region in each sample segmentation image, the statistical frequency of at least one candidate segmentation category corresponding to the sample segmentation image is obtained. The product between the correction coefficient corresponding to the segmentation scale of the sample segmentation image and the statistical frequency of at least one candidate segmentation category corresponding to the sample segmentation image is used as the target statistical frequency of at least one candidate segmentation category corresponding to the sample segmentation image. For all sample segmentation images, the target statistical frequencies of the same candidate segmentation category in each sample segmentation image are accumulated to obtain the final statistical frequency of each candidate segmentation category.

[0064] Based on the statistical frequency of each co-occurring segmentation pair, a sub-association weight is constructed between the two segmentation categories to which the corresponding co-occurring segmentation pair belongs. The sub-association weight represents the penalty weight between the two segmentation categories to which the corresponding co-occurring segmentation pair belongs. The higher the probability that the two segmentation categories to which the co-occurring segmentation pair belongs co-occur, the smaller the sub-association weight.

[0065] The more times a co-occurrence segmentation pair is statistically counted, the more times the two segmentation categories within that pair co-occur, indicating that the pair is more common in natural scenes and less likely to conflict or mis-segment. Therefore, a smaller penalty weight or a higher affinity is assigned, corresponding to the aforementioned sub-association weight. Conversely, co-occurrence segmentation pairs with very few or even zero co-occurrence counts are assigned a larger penalty weight to impose strong constraints during segmentation. For example, the sub-association weight can be set to be inversely proportional to the statistical count, or the statistical count can be mapped to a specific numerical range through normalization.

[0066] In some application scenarios, the statistical frequency of each co-occurrence segmentation pair is sorted to obtain the ranking of the corresponding co-occurrence segmentation pair; based on the ranking of each co-occurrence segmentation pair, the target penalty weight of the corresponding co-occurrence segmentation pair is selected from several preset penalty weights; for each co-occurrence segmentation pair, the target penalty weight of the co-occurrence segmentation pair is directly used as the sub-association weight between the two segmentation categories to which the co-occurrence segmentation pair belongs; or, for each co-occurrence segmentation pair, the target penalty weight of the co-occurrence segmentation pair is mapped to obtain the sub-association weight between the two segmentation categories to which the co-occurrence segmentation pair belongs.

[0067] In other application scenarios, for each co-occurring segmentation category pair, the sub-association weight between the two segmentation categories to which the co-occurring segmentation category pair belongs is determined based on the statistical frequency of the co-occurring segmentation category pair and the statistical frequency of all co-occurring segmentation category pairs.

[0068] For example, this application generates association weights between segmentation categories based on considerations of practical problems, i.e., a probability deviation penalty table for each segmentation category. For instance, if a train appears on railway tracks, we consider it normal for a train and tracks to be together, and no penalty is imposed. However, if a train appears on a road in the captured image content, we consider it abnormal for a train and road to be together, and a penalty is required. Therefore, by performing domain association statistics on the training data, such a probability deviation penalty table for each segmentation category can be generated, thereby obtaining the association weights between each segmentation category. First, a two-dimensional statistical table of co-occurring segmentation category pairs (referred to as the SPT table) is generated. The SPT table includes several combinations of any two segmentation categories, i.e., the SPT table includes several co-occurring segmentation category pairs, and the SPT table also includes the statistical frequency of each co-occurring segmentation category pair. The SPT table is initialized, and the statistical frequency of each co-occurring segmentation category pair is initialized to 0. Simultaneously, this application divides the training sample segmentation images into regions, cutting them into 8×8 image blocks. If the last few columns and rows of the sample segmentation image are not large enough for an 8×8 block during region division, the blocks are cut according to their actual size. Within each block, the true segmentation categories appearing in the actual segmentation results are counted. For example, if five true segmentation categories appear in an 8×8 block: animal, car, tree, grass, and road, the relationships between the categories themselves are removed, as well as symmetrical relationships like (animal, car) and (car, animal). This leaves 10 co-occurring segmentation category pairs. At this point, the co-occurring segmentation category pairs are: (animal, car), (animal, tree), (animal, grass), (animal, road), (car, tree), (car, grass), (car, road), (tree, grass), (tree, road), (tree, road), (grass, road). This information is updated in the sPT table. For example, the update for (animal, car) can be expressed by the formula: new sPT(animal, car) = sPT(animal, car) + 1. The update calculation method for the other 9 cases of co-occurrence segmentation category pairs is the same. After all blocks in the sample segmentation image are statistically analyzed using the above formula, an initial segmentation probability deviation penalty table for each category can be generated. Then, the maximum value of the initial segmentation probability deviation penalty table for each category is calculated. For example, the maximum value obtained by the maximum number of statistical counts for one of the co-occurrence segmentation category pairs is maxPT. MaxPT is used to construct the sub-association weights of the two segmentation categories in each co-occurrence segmentation category pair. Specifically, the process of determining the sub-association weights of the two segmentation categories in any co-occurrence segmentation category pair can refer to the following formula (5): Formula (5); in, This represents the sub-association weights of the two segmentation categories in any co-occurring segmentation category pair, i.e., the sub-association weights of any co-occurring segmentation category pair. This represents the maximum number of times any co-occurring segmentation class pair is found in the segmented images of each sample. This represents the sum of the statistical occurrences of all co-occurring segmentation category pairs, or the sum of the statistical occurrences of any one co-occurring segmentation category pair across all sample segmented images. B represents a preset value, for example, 0.0001. It can be assumed that the more frequently a pair appears, the more likely it is to occur, and therefore the smaller the penalty weight, i.e., the smaller the sub-association weight between the two segmentation categories. A probability bias penalty table (PT table) for each segmentation category is constructed using the sub-association weights of each co-occurring segmentation category pair. In practical applications, to achieve fast table lookup, the sPT table can be symmetrically copied. For example, to query the value of sPT(car, animal), if only the case of sPT(animal, car) is counted, the constructed sPT table and / or PT table can be simply symmetrically copied, such that the statistical occurrences sPT(car, animal) are equal to sPT(animal, car), and the sub-association weights PT(car, animal) and PT(animal, car) are also the same.

[0069] In some embodiments, the current probability bias includes the current sub-probability bias of each sub-feature map in the second feature map. The current sub-probability bias of a sub-feature map includes the current sub-class probability bias of the sub-feature map within several segmentation categories. The current sub-probability bias of a sub-feature map is a one-dimensional vector, where each element represents the current sub-class probability bias of the sub-feature map within its corresponding segmentation category. For the same sub-feature map, the current sub-class probability bias of the sub-feature map is used to update the current sub-class probability of the corresponding segmentation category within the current sub-probability of that sub-feature map. The preset association weight includes the sub-association weight between any two segmentation categories. For example, the sub-association weight between any two segmentation categories represents the penalty weight between the two corresponding segmentation categories; the more times the two corresponding segmentation categories co-occur, the smaller the penalty weight between the two segmentation categories.

[0070] Specifically, the step of obtaining the current probability deviation based on the first target convolution result, the current probability to be updated, and the preset association weights between several segmentation categories includes: traversing the current subclass probability of each sub-feature map in the second feature map to each segmentation category, and taking the segmentation category to which the current sub-subclass probability of the current sub-feature map belongs as the current segmentation category; taking at least some of the neighboring sub-feature maps associated with the current sub-feature map in the second feature map as each candidate sub-feature map; for each candidate sub-feature map, obtaining the feature map association weight between the candidate sub-feature map and the current sub-feature map under the current segmentation category based on the sub-association weights between the current segmentation category and each segmentation category and the current subclass probability of the candidate sub-feature map to each segmentation category; obtaining the current subclass probability deviation of the current sub-feature map to the current segmentation category based on the first target convolution result and the feature map association weights between all candidate sub-feature maps and the current sub-feature map under the current segmentation category; until all the current subclass probability deviations of the sub-feature maps in the second feature map to each segmentation category have been traversed, the current sub-probability deviation of each sub-feature map in the second feature map is obtained.

[0071] The current sub-probability of the current sub-feature map is a one-dimensional vector. This vector represents the mapping result of the current sub-feature map in each segmentation category. Each element in the vector represents the mapping result of the current sub-feature map under a specific segmentation category (i.e., the sub-class probability).

[0072] The current sub-class probability of the current sub-feature map represents the mapping result or segmentation probability of the current sub-feature map under the currently traversed segmentation category. The current sub-feature map is the sub-feature map currently traversed in the second feature map. The current segmentation category represents the segmentation category to which the current sub-class probability corresponding to the current sub-feature map belongs. The candidate segmentation categories include other segmentation categories besides the current segmentation category. Specifically, when the number of segmentation categories is two, the number of candidate segmentation categories is one; or, when the number of segmentation categories is at least three, the number of candidate segmentation categories is at least two.

[0073] A neighborhood sub-feature map represents a sub-feature map in the second feature map that falls within the neighborhood of the current sub-feature map. At least some neighborhood sub-feature maps represent all sub-feature maps in the second feature map that fall within the neighborhood of the current sub-feature map, or sub-feature maps that have not been traversed. For example, to save computational resources, untraversed sub-feature maps in the neighborhood sub-feature maps associated with the current sub-feature map can be used as candidate sub-feature maps. Each candidate sub-feature map represents at least some of the neighborhood sub-feature maps associated with the current sub-feature map.

[0074] The current subprobability of the candidate sub-feature map is a one-dimensional vector that represents the mapping result of the candidate sub-feature map in each segmentation category. Each element in the vector represents the mapping result of the candidate sub-feature map under a specific segmentation category (i.e., the sub-class probability).

[0075] The feature map association weight represents the association between two sub-feature maps in the second feature map under the current segmentation category, that is, it represents the association between the current sub-feature map and the candidate sub-feature map under the current segmentation category.

[0076] The feature map association weights are obtained by fusing the current sub-probabilities of candidate sub-feature maps under each segmentation category and the sub-association weights between the current segmentation category and each segmentation category. In some application scenarios, for each segmentation category, the product of the current sub-probability of the candidate sub-feature map under that segmentation category and the sub-association weights between the current segmentation category and that segmentation category is used as the feature map sub-association weight between the current sub-feature map and the candidate sub-feature map under that segmentation category. The feature map sub-association weights between the current sub-feature map and the candidate sub-feature map under the current segmentation category and each segmentation category are then aggregated to obtain the feature map association weight between the candidate sub-feature map and the current sub-feature map under the current segmentation category.

[0077] In other application scenarios, the probability of each sub-feature map in the second feature map belonging to the current subclass of each segmentation category is traversed. The segmentation category to which the current sub-feature map belongs is taken as the current segmentation category, and at least one segmentation category among several segmentation categories is taken as a candidate segmentation category. For example, the candidate segmentation category can be other segmentation categories besides the current segmentation category among several segmentation categories. At least some of the neighboring sub-feature maps associated with the current sub-feature map are taken as candidate sub-feature maps. For each candidate sub-feature map, the feature map association weight between the candidate sub-feature map and the current sub-feature map under the current segmentation category is obtained based on the sub-association weight between the current segmentation category and the candidate segmentation category and the probability of the candidate sub-feature map in the current subclass of the candidate segmentation category. Based on the first target convolution result and the feature map association weight between all candidate sub-feature maps and the current sub-feature map under the current segmentation category, the current subclass probability deviation of the current sub-feature map in the current segmentation category is obtained. This process continues until the current subclass probability deviations of all sub-feature maps in the second feature map belonging to each segmentation category are traversed, resulting in the current sub-probability deviation of each sub-feature map in the second feature map.

[0078] Understandably, the feature map association weight between the candidate sub-feature map and the current sub-feature map under the current segmentation category can be the sub-association weight between the current segmentation category and the target segmentation category, plus the probability of the candidate sub-feature map belonging to the current sub-class of the target segmentation category. Here, the target segmentation category represents at least one segmentation category, and at least one segmentation category includes one or more segmentation categories. For example, at least one segmentation category can be all segmentation categories among several segmentation categories, or at least one segmentation category can be any segmentation category other than the current segmentation category among several segmentation categories.

[0079] In some application scenarios, the step of obtaining the probability deviation of the current sub-feature map in the current sub-class of the current segmentation category based on the first target convolution result and the feature map association weights between all candidate sub-feature maps and the current sub-feature map in the current segmentation category includes: taking the statistical value of the first target convolution result as the first statistical value; taking the statistical value between the new feature map association weights between all candidate sub-feature maps and the current sub-feature map in the current segmentation category as the second statistical value; taking the sum of the first statistical value and the second statistical value as the target statistical value; and normalizing the target statistical value to obtain the probability deviation of the current sub-feature map in the current sub-class of the current segmentation category, wherein the statistical value can be the mean or median, etc.

[0080] In some embodiments, the first feature map includes a first sub-feature map corresponding to the current sub-feature map position and a second sub-feature map corresponding to the candidate sub-feature map position. The second feature map is obtained by feature extraction from the first feature map. The first and second feature maps have different scales, but each includes a corresponding sub-feature map. The sub-feature map corresponding to the current sub-feature map position is used as the first sub-feature map in the first feature map; the sub-feature map corresponding to the candidate sub-feature map position is used as the second sub-feature map in the first feature map. The sub-convolution result is an intermediate feature obtained by performing local convolution operations between the first sub-feature map and each of the second sub-feature maps. The first target convolution result includes the sub-convolution results between each sub-feature map in the first feature map. The first target convolution result includes the sub-convolution results between the first sub-feature map and the second sub-feature map corresponding to the candidate sub-feature map.

[0081] Specifically, the step of obtaining the probability deviation of the current sub-feature map in the current segmentation category based on the first target convolution result and the feature map association weights between all candidate sub-feature maps and the current sub-feature map in the current segmentation category includes: for each candidate sub-feature map, determining the inter-feature probability deviation between the candidate sub-feature map and the current sub-feature map in the current segmentation category based on the sub-convolution result between the first sub-feature map and the corresponding second sub-feature map of the candidate sub-feature map and the feature map association weights between the candidate sub-feature map and the current sub-feature map in the current segmentation category; and obtaining the probability deviation of the current sub-feature map in the current segmentation category based on the inter-feature probability deviation between each candidate sub-feature map and the current sub-feature map in the current segmentation category.

[0082] The probability bias between features refers to the probability correction caused by the feature interaction between the current sub-feature map and the candidate sub-feature map under the current segmentation category.

[0083] Specifically, for each candidate sub-feature map, the following steps are performed: The product of the sub-convolution result between the first sub-feature map and the corresponding second sub-feature map of the candidate sub-feature map, and the feature map association weights between the candidate sub-feature map and the current sub-feature map under the current segmentation category, is used as the candidate product of the association between the candidate sub-feature map and the current sub-feature map; the candidate product of the association between the candidate sub-feature map and the current sub-feature map is directly used as the inter-feature probability deviation between the candidate sub-feature map and the current sub-feature map in the current segmentation category; or, the product of the candidate product of the association between the candidate sub-feature map and the current sub-feature map and the preset weights is used as the inter-feature probability deviation between the candidate sub-feature map and the current sub-feature map in the current segmentation category. The positional distance between the candidate sub-feature map and the current sub-feature map is related to the preset weights; the smaller the distance between the center point of the candidate sub-feature map and the center point of the current sub-feature map in the second feature map, the larger the preset weights.

[0084] Specifically, the above-mentioned method of obtaining the probability deviation of the current sub-feature map in the current segmentation category based on the inter-feature probability deviations between each candidate sub-feature map and the current sub-feature map in the current segmentation category includes: using the sum of the inter-feature probability deviations between each candidate sub-feature map and the current sub-feature map in the current segmentation category as the probability deviation of the current sub-feature map in the current segmentation category; or, performing weighted fusion on the inter-feature probability deviations between each candidate sub-feature map and the current sub-feature map in the current segmentation category to obtain the probability deviation of the current sub-feature map in the current segmentation category. For example, the positional distance between the candidate sub-feature map and the current sub-feature map is related to the weight in the weighted fusion process. The smaller the distance between the center point of the candidate sub-feature map and the center point of the current sub-feature map in the second feature map, the greater the weight of the candidate sub-feature map.

[0085] Specifically, the process of determining the current probability deviation in any update round can be referred to the following formulas (6), (7), and (8): Formula (6); Formula (7); Formula (8); in, This represents the current probability deviation, specifically, This represents the probability deviation of the current sub-feature map in the current subclass to which the current segmentation category belongs. This represents the position of any sub-feature map in the second feature map, i.e., the current sub-feature map in the second feature map. The initial segmentation probability can be represented as segm0, where the initial segmentation probability size is... L represents the number of segmentation categories, and sgH and sgW represent the height and width of segm0, respectively. and This indicates the index of the probability to be updated in the L direction. Represents the current segmentation category among several segmentation categories. This represents any one of several segmentation categories (i.e., the target segmentation category mentioned above). For example, It can be the same as or different from the current segmentation category. For example, it can represent other segmentation categories besides the current segmentation category (i.e., the above candidate segmentation categories). Except for the second feature graph The locations of other sub-feature maps besides, for example, Representing the second feature graph In The location within the neighborhood of the first feature map specifically represents the location of the candidate sub-feature map in the second feature map. It is understood that the first and second feature maps have different scales, but each includes a corresponding sub-feature map; for formulas (2) and (3) as well as With respect to formulas (6) to (7) here as well as The difference is that the same symbols are used to represent the same positions corresponding to the first and second feature maps. This represents the sub-association weight between the current segmentation category and any other segmentation category. For example, it represents the sub-association weight between the current segmentation category and the target segmentation category. It represents the probability of a candidate sub-feature map belonging to the current subclass of any segmentation category. The first target convolution result represents the sub-convolution result between the first sub-feature map corresponding to the current sub-feature map position and the second sub-feature map corresponding to the candidate sub-feature map in the first target convolution result. The first target convolution result can be obtained by convolving the first feature map with at least one Gaussian kernel and at least one bilateral kernel. This represents multi-scale fusion, which can use a size of This is implemented using standard depthwise convolutions, where M is 6, corresponding to 6 layers. Parameters. Fea represents the first feature map. The first target convolution result is obtained by convolving the first feature map. Specifically, the convolution process can be performed on the first feature map using at least one Gaussian kernel. In formulas (6) and (8), segm represents the current probability to be updated.

[0086] It is understandable that in any update round, the process of obtaining the current probability deviation using formula (6) is the first probability deviation construction strategy, the process of obtaining the current probability deviation using formulas (7) and (8) is the second probability deviation construction strategy, and the process of obtaining the current probability deviation using formula (4) is the third probability deviation construction strategy. In the first probability deviation construction strategy, the first target convolution result is obtained through at least one Gaussian kernel and at least one bilateral kernel, the correlation weight between each segmentation category is introduced, and the current probability deviation is obtained by multi-scale fusion through a standard depth convolution. In the second probability deviation construction strategy, the first target convolution result is obtained through at least one Gaussian kernel, the correlation weight between each segmentation category is introduced, and the current probability deviation is obtained by multi-scale fusion through a standard depth convolution. In the third probability deviation construction strategy, the current probability to be updated, the first feature map after convolution processing, and the current probability to be updated after convolution processing are fused to obtain the current probability deviation. Of these three probability deviation construction strategies, the first probability deviation construction strategy has the best segmentation effect, but consumes the most computational resources. The third probability deviation construction strategy has the worst segmentation effect, but consumes the least computational resources. The second probability deviation construction strategy is a compromise. In practical applications, the specific probability deviation construction strategy used in each update round can be selected based on the positioning or needs of the actual product. Therefore, in the process of updating the initial segmentation probability in multiple rounds to obtain the target segmentation probability, the specific probability deviation construction strategy used in each update round is not limited. The probability deviation construction strategies used between each update round can be the same or different, and this application does not impose any restrictions.

[0087] S22: Determine the updated probability for the current update round based on the current probability base and the current probability deviation.

[0088] In some application scenarios, S22 above can be: taking the difference between the current probability base and the current probability deviation as the target difference; substituting the target difference into a preset exponential function to obtain the target function value; and directly taking the target function value as the updated probability of the current update round.

[0089] In some embodiments, S22 may include the following steps: using the difference between the current probability base and the current probability deviation as the target difference; substituting the target difference into a preset exponential function to obtain the target function value; and normalizing the target function value to obtain the updated probability of the current update round.

[0090] The current probability base represents the logarithmic result of the probability to be updated. Specifically, the current probability base includes the current sub-probability bases to which each sub-feature map in the second feature map belongs. The target difference represents the difference between the current probability base and the current probability deviation. The target difference includes the sub-target difference of each sub-feature map in the second feature map. The sub-target difference of the sub-feature maps in the second feature map represents the difference between the current sub-probability base and the current sub-probability deviation to which the same sub-feature map in the second feature map belongs.

[0091] Specifically, the process of determining the updated probability of the current update round based on the current probability base and the current probability deviation is described in formulas (9) and (10) below: Formula (9); Formula (10); in, This represents the current probability base to be updated. S represents the sequence number of the current update round. S+1 represents the sequence number of the next update round after the current update round. The current probability deviation in the current update round can be calculated based on the above formula (4), or the above formula (6), or the above formula (7) and the above formula (8). This represents the default exponential function. This represents the objective function value in the current update round. Represents the Sigmoid function, used to... Normalize.

[0092] For example, such as Figure 3 As shown, the image to be segmented undergoes an initial segmentation generation stage, resulting in the initial segmentation probabilities output by this stage. Pre-defined depthwise segmentation networks, such as VGGnet, ResNet, and EfficientNet, can be used to construct this initial segmentation generation stage. The pre-defined extraction modules are configured with sequentially executed standard convolution operations, normalization processes, and activation operations. The initial segmentation generation stage includes several cascaded pre-defined extraction modules and a mapping module. The number of these pre-defined extraction modules is not limited. These modules can be... Figure 3The diagram includes C11, C12, C21, C22, C31, C32, and other preset extraction modules (not shown). The stride of the standard convolution operation in each preset extraction module can be the same or different. For example, the stride of the standard convolution operation in C11 and C12 can be the same. The stride of the standard convolution operation in C21 and C31 can be different from the stride of the standard convolution operation in C11. C21 and C31 are based on C11, but with a stride of 2. The mapping module has preset activation functions or fully connected layers. The preset activation functions can be ReLU, Sigmoid, Softmax, etc. For example, let's take a mapping module with a Sigmoid activation function as an example. Figure 3 The mapping module in the code maps the data output from the preset extraction module corresponding to C32 to probabilities. The input image to be segmented is denoted as I, and its size is... Where 3 represents the RGB three channels, and IH and IW represent the height and width of the image to be segmented, respectively. The output of the mapping module is the initial segmentation probability, which can be represented as segm0. The size of the initial segmentation probability is... Where L represents the number of segmentation categories, and sgH and sgW represent the height and width of segm0, respectively. Generally, sgH is less than IH, and sgW is less than IW, in a power of 2. In practical applications, considering the depth receptive field and computing power, this application can use an 8-fold relationship, or the ratio can be adjusted according to the product's positioning. For example... as well as .

[0093] The target segmentation probability is obtained by updating the initial segmentation probability through multiple segmentation boundary clarification stages. Each segmentation boundary clarification stage corresponds to one round of updating the current probability to be updated. The current probability to be updated can be represented by `segm`. In the case that the current update round is the first round, the current probability to be updated `segm` is specifically the initial segmentation probability `segm0`. The data at a certain position in `segm`, for example... , can be directly represented as Here, l corresponds to the index in the segmentation category direction, ranging from 0 to L-1; x corresponds to the index in the height direction; and y corresponds to the index in the width direction. After the network outputs segm, the image segmentation result at that location is determined, i.e., for that location... Use its corresponding L values ​​to determine Predicting the segmentation category based on location. For example, for target segmentation probability, first extract the location. Data x corresponds to the index in the height direction, and y corresponds to the index in the width direction. Let L be a one-dimensional vector containing L values, each representing the probability of a segmentation category. The segmentation category corresponding to the largest value among the L values ​​is taken as... Predicted segmentation category for location.

[0094] like Figure 3 As shown, Fea represents the first feature map of the image to be segmented. PT represents the association weights between each segmentation category. In some application scenarios, each segmentation boundary sharpening stage corresponds to an update process of the current probability to be updated. The current probability to be updated can be updated using the current probability to be updated, the first feature map, and the association weights between each segmentation category, thus obtaining the updated probability for the current update round. Specifically, as... Figure 4 As shown, any update round can execute the following steps: S41: Use the logarithmic result of the current probability to be updated as the current probability basis for the current update round. S42: Perform a convolution operation on the first feature map using at least one Gaussian kernel and at least one bilateral kernel to obtain the first candidate convolution result of the first feature map. S43: Perform a convolution operation on the first feature map using at least one Gaussian kernel to obtain the second candidate convolution result of the first feature map. S44: Use the first candidate convolution result of the first feature map, or the second candidate convolution result of the first feature map, as the first target convolution result of the first feature map. S45: Obtain the current probability deviation based on the first target convolution result, the current probability to be updated, and the preset association weights between several segmentation categories. S46: Determine the updated probability of the current update round based on the current probability basis and the current probability deviation. For example, as shown... Figure 5 As shown, Fea represents the first feature map of the image to be segmented. PT represents the association weights between each segmentation category. In S45, the current probability bias is constructed through scale filtering and multi-scale fusion. Then, in S46, the updated probability of the current update round is determined using the current probability base and the current probability bias.

[0095] In other application scenarios, each segmentation boundary sharpening stage corresponds to an update process of the current probability to be updated. The update of the current probability to be updated can be achieved using only the current probability to be updated and the first feature map, thus obtaining the updated probability for the current update round. Specifically, such as... Figure 6As shown, any update round can also execute the following steps: S61: Use the logarithmic result of the current probability to be updated as the current probability basis for the current update round. S62: Perform a convolution operation on the first feature map using a preset convolution kernel to obtain the third candidate convolution result of the first feature map, and use the third candidate convolution result as the first target convolution result of the first feature map. S63: Perform convolution processing on the current probability to be updated to obtain the second target convolution result of the current probability to be updated. S64: Obtain the current probability deviation based on the first target convolution result, the second target convolution result, and the current probability to be updated. S65: Determine the updated probability of the current update round based on the current probability basis and the current probability deviation. For example, as shown... Figure 7 As shown, Fea represents the first feature map of the image to be segmented. Convolutional processing is performed directly on the first feature map and the current probability to be updated, respectively. Then, the current probability to be updated, the first feature map after convolution, and the current probability to be updated after convolution are fused to obtain the current probability bias.

[0096] This application can be considered as re-representing the post-processing algorithm using the language of deep learning and proposing a method for constructing various segmentation probability deviation penalty tables. It introduces correlation weights between segmentation categories, thereby achieving better segmentation boundary clarity and improving the accuracy of image segmentation results. Furthermore, the above image segmentation method can be applied to image segmentation models. By implementing steps S11 to S13 and their specific steps through these models, this application achieves unified deep training by re-representing the post-processing algorithm using the language of deep learning, thus maximizing the role of the initial segmentation network and the post-processing algorithm and achieving better segmentation boundary clarity. This application designs several simplified versions of various segmentation probability deviation generation methods to achieve segmentation boundary clarity with low computational power. Moreover, based on practical applications, the proposed method for constructing various segmentation probability deviation penalty tables involves the co-occurrence probability of each segmentation category, effectively simplifying the cost of deep network training and significantly improving the clarity of segmentation boundaries.

[0097] Please see Figure 8 , Figure 8 This is a schematic diagram of an embodiment of the image segmentation apparatus of this application. The image segmentation apparatus 80 includes an acquisition module 81, an update module 82, and a segmentation module 83; the acquisition module 81 is used to acquire a first feature map of the image to be segmented and an initial segmentation probability of each segmentation category in the image to be segmented, wherein the initial segmentation probability is obtained based on the first feature map; the update module 82 is used to perform segmentation probability update processing based on the first feature map and the initial segmentation probability to obtain a target segmentation probability of each segmentation category in the image to be segmented; the segmentation module 83 is used to perform image segmentation processing based on the target segmentation probability to obtain an image segmentation result.

[0098] Please refer to the image segmentation method for the functions performed by each module; they will not be repeated here.

[0099] The above scheme obtains the first feature map of the image to be segmented and the initial segmentation probability of each segmentation category in the image to be segmented; performs segmentation probability update processing based on the first feature map and the initial segmentation probability to obtain the target segmentation probability of each segmentation category in the image to be segmented; performs image segmentation processing based on the target segmentation probability to obtain the image segmentation result. Thus, the initial segmentation probability is updated through the first feature map, so that the segmentation probability update process can adaptively capture the semantic relationship between pixels, thereby improving the clarity of the segmentation boundary, and reducing the computational complexity while maintaining the segmentation accuracy, achieving image segmentation with low computational overhead.

[0100] Please see Figure 9 , Figure 9 This is a schematic diagram of the structure of an embodiment of the electronic device of this application. The electronic device 90 includes a memory 91 and a processor 92. The processor 92 is used to execute program instructions stored in the memory 91 to implement the steps in the above-described image segmentation method embodiment. In a specific implementation scenario, the electronic device 90 may include, but is not limited to, a microcomputer or a server. In addition, the electronic device 90 may also include mobile devices such as laptops and tablets, which are not limited here.

[0101] Specifically, processor 92 controls itself and memory 91 to implement the steps in the above-described image segmentation method embodiments. Processor 92 can also be referred to as a CPU (Central Processing Unit). Processor 92 may be an integrated circuit chip with signal processing capabilities. Processor 92 can also be a general-purpose processor, digital signal processor (DSP), application-specific integrated circuit (ASIC), field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. A general-purpose processor can be a microprocessor or any conventional processor. Furthermore, processor 92 can be implemented using integrated circuit chips.

[0102] Please see Figure 10 , Figure 10This is a schematic diagram of a computer-readable storage medium according to an embodiment of the present application. The computer-readable storage medium 100 stores program instructions 1001 thereon, which, when executed by a processor, implement the steps in any of the above-described image segmentation method embodiments.

[0103] In some embodiments, the functions or modules of the apparatus provided in this disclosure can be used to perform the methods described in the above method embodiments. The specific implementation can be referred to the description of the above method embodiments, and for the sake of brevity, it will not be repeated here.

[0104] The description of the various embodiments above tends to emphasize the differences between the various embodiments. The similarities or similarities between them can be referred to, and for the sake of brevity, they will not be repeated here.

[0105] In the several embodiments provided in this application, it should be understood that the disclosed methods and apparatus can be implemented in other ways. For example, the apparatus implementations described above are merely illustrative. For instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the mutual coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection of devices or units may be electrical, mechanical, or other forms.

[0106] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0107] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) or processor to execute all or part of the steps of the methods of various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0108] If the technical solution of this application involves personal information, the product using this technical solution has clearly informed the user of the personal information processing rules and obtained the user's voluntary consent before processing the personal information. If the technical solution of this application involves sensitive personal information, the product using this technical solution has obtained the user's separate consent before processing the sensitive personal information, and also meets the requirement of "express consent". For example, at personal information collection devices such as cameras, clear and prominent signs are set up to inform users that they have entered the scope of personal information collection and that personal information will be collected. If an individual voluntarily enters the collection scope, it is deemed that they have agreed to the collection of their personal information; or on the personal information processing device, with clear signs / information informing users of the personal information processing rules, authorization is obtained from the individual through pop-up information or by asking the individual to upload their personal information; wherein, the personal information processing rules may include information such as the personal information processor, the purpose of personal information processing, the processing method, and the types of personal information processed.

Claims

1. An image segmentation method, characterized in that, The image segmentation method includes: Obtain a first feature map of the image to be segmented and an initial segmentation probability for each segmentation category in the image to be segmented, wherein the initial segmentation probability is obtained based on the first feature map; The segmentation probability is updated based on the first feature map and the initial segmentation probability to obtain the target segmentation probability for each segmentation category in the image to be segmented. This includes: when the current update round is the first update round, using the initial segmentation probability as the current probability to be updated; or, when the current update round is not the first update round, using the updated probability of the previous update round as the current probability to be updated. For each update round of the current probability to be updated, the following steps are performed: using the logarithmic result of the current probability to be updated as the current probability base for the current update round; determining the updated probability of the current update round based on the current probability base, the first feature map, and the current probability to be updated; and, in response to the updated probability of the current update round satisfying a preset update condition, using the updated probability of the current update round as the target segmentation probability. The total number of update rounds for the segmentation probability update process is multiple rounds. Image segmentation is performed based on the target segmentation probability to obtain the image segmentation result.

2. The image segmentation method according to claim 1, characterized in that, The current probability to be updated includes the current sub-probabilities of several sub-feature maps in the second feature map. The second feature map is obtained by feature extraction from the first feature map, and the scale of the first feature map is different from that of the second feature map. The step of determining the updated probability of the current update round based on the current probability base, the first feature map, and the current probability to be updated includes: Based on the first feature map and the current probability to be updated, the current probability deviation of the current update round is determined, and the current probability deviation represents the difference between the current sub-probabilities of each sub-feature map in the second feature map; The updated probability of the current update round is determined based on the current probability base and the current probability deviation.

3. The image segmentation method according to claim 2, characterized in that, The current sub-probability to which the sub-feature map belongs includes the probability of the current sub-class to which the sub-feature map belongs in each segmentation category; The step of determining the current probability deviation of the current update round based on the first feature map and the current probability to be updated includes: The first feature map is convolved to obtain the first target convolution result of the first feature map. The current probability bias is obtained based on the first target convolution result and the current probability to be updated; or, The current probability deviation is obtained based on the first target convolution result, the current probability to be updated, and the preset association weights between several segmentation categories.

4. The image segmentation method according to claim 3, characterized in that, The current probability deviation includes the current sub-probability deviation of each sub-feature map in the second feature map, and the current sub-probability deviation of the sub-feature map includes the current sub-class probability deviation of the sub-feature map in several segmentation categories, and the preset association weight includes the sub-association weight between any two segmentation categories. The step of obtaining the current probability deviation based on the first target convolution result, the current probability to be updated, and the preset association weights between several segmentation categories includes: Iterate through the probability of each sub-feature map in the second feature map to the current sub-class of each segmentation category, and take the segmentation category to which the current sub-feature map belongs as the current segmentation category; At least a portion of the neighboring sub-feature maps associated with the current sub-feature map in the second feature map are respectively used as candidate sub-feature maps; For each candidate sub-feature map, the feature map association weight between the candidate sub-feature map and the current sub-feature map under the current segmentation category is obtained based on the sub-association weight between the current segmentation category and each segmentation category and the probability of the candidate sub-feature map in the current sub-class to which each segmentation category belongs. Based on the first target convolution result and the feature map association weights between all candidate sub-feature maps and the current sub-feature map under the current segmentation category, the probability deviation of the current sub-feature map in the current sub-class to which the current segmentation category belongs is obtained; The process continues until all sub-feature maps in the second feature map have been traversed, and the probability deviations of the current sub-classes belonging to each segmentation category are obtained, thus obtaining the current sub-probability deviations of each sub-feature map in the second feature map.

5. The image segmentation method according to claim 4, characterized in that, The first feature map includes a first sub-feature map corresponding to the current sub-feature map position and a second sub-feature map corresponding to the candidate sub-feature map position. The first target convolution result includes the sub-convolution result between the first sub-feature map and the second sub-feature map corresponding to the candidate sub-feature map. The step of obtaining the probability deviation of the current sub-feature map in the current sub-class of the current segmentation category based on the first target convolution result and the feature map association weights between all candidate sub-feature maps and the current sub-feature map in the current segmentation category includes: For each candidate sub-feature map, the probability deviation between the candidate sub-feature map and the current sub-feature map in the current segmentation category is determined based on the sub-convolution result between the first sub-feature map and the second sub-feature map corresponding to the candidate sub-feature map and the feature map association weight between the candidate sub-feature map and the current sub-feature map in the current segmentation category. Based on the probability deviation between each candidate sub-feature map and the current sub-feature map in the current segmentation category, the probability deviation of the current sub-feature map in the current sub-class to which the current segmentation category belongs is obtained.

6. The image segmentation method according to claim 3, characterized in that, The step of performing convolution processing on the first feature map to obtain the first target convolution result of the first feature map includes: The first feature map is convolved using at least one Gaussian kernel and at least one bilateral kernel to obtain a first candidate convolution result for the first feature map; or... The first feature map is convolved using at least one Gaussian kernel to obtain a second candidate convolution result for the first feature map; or... The first feature map is convolved by a preset convolution kernel to obtain the third candidate convolution result of the first feature map; One of the first candidate convolution result, the second candidate convolution result, and the third candidate convolution result is taken as the first target convolution result.

7. The image segmentation method according to claim 3, characterized in that, The preset association weights include sub-association weights between any two segmentation categories; Before the steps of obtaining the first feature map of the image to be segmented and the initial segmentation probabilities of each segmentation category in the image to be segmented, the image segmentation method further includes: Obtain the true segmentation results of several preset image regions in the sample segmentation image; Based on the common segmentation categories in the actual segmentation results of each preset image region, determine the statistical frequency of several co-occurring segmentation category pairs; Based on the statistical frequency of each co-occurrence segmentation pair, construct the sub-association weights between the two segmentation categories to which the corresponding co-occurrence segmentation pair belongs.

8. The image segmentation method according to claim 3, characterized in that, The step of obtaining the current probability deviation based on the first target convolution result and the current probability to be updated includes: Perform convolution processing on the current probability to be updated to obtain the second target convolution result of the current probability to be updated; The current probability bias is obtained based on the first target convolution result, the second target convolution result, and the current probability to be updated.

9. The image segmentation method according to claim 2, characterized in that, The step of determining the updated probability of the current update round based on the current probability base and the current probability deviation includes: The difference between the current probability base and the current probability deviation is taken as the target difference; Substitute the target difference into a preset exponential function to obtain the target function value; The objective function value is normalized to obtain the updated probability of the current update round.

10. An electronic device, characterized in that, include: A memory and a processor, wherein the memory stores program instructions, and the processor retrieves the program instructions from the memory to perform the image segmentation method as described in any one of claims 1-9.

11. A computer-readable storage medium having program instructions stored thereon, characterized in that, When the program instructions are executed by the processor, they are used to implement the image segmentation method as described in any one of claims 1-9.