Image semantic segmentation method and system based on dynamic convolution attention
By introducing a dynamic convolutional attention mechanism in image semantic segmentation technology, combining object overlap coefficients and local detail coefficients, dynamically adjusting the image magnification, the problem of difficult processing of small targets and overlapping areas is solved, and a more efficient image segmentation effect is achieved.
Patent Information
- Application Number
- CN202510630017.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-16
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2045-05-16
AI Technical Summary
Existing image semantic segmentation technology has the problem of difficulty in accurately segmenting and processing when dealing with small targets and overlapping areas, resulting in poor segmentation effect.
The image semantic segmentation method based on dynamic convolution attention is adopted, and the object overlap coefficient and local detail enrichment coefficient are calculated, and the image magnification is dynamically adjusted to achieve refined processing of overlapping areas.
It improves the recognition and segmentation accuracy of small targets, effectively processes overlapping areas, and improves the quality and detailed richness of image segmentation.
Smart Images

Figure CN120147652A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing, and particularly relates to an image semantic segmentation method and system based on dynamic convolution attention. Background Art
[0002] When existing image semantic segmentation technologies identify and segment small targets, due to the small size of small targets, traditional segmentation algorithms often have difficulty accurately distinguishing and processing them, resulting in poor segmentation effects; in addition, the overlapping problem between small targets also greatly affects the segmentation quality. Especially when performing image segmentation in natural scenes and complex backgrounds, the mutual occlusion and overlap between small targets become more complex, and it is difficult for existing technologies to effectively solve this problem; Based on the above technical problems, existing image segmentation methods have certain limitations when dealing with overlapping regions; improving the detail information of local regions can help improve the segmentation effect, but most of these methods rely on a fixed magnification factor and lack personalized adjustment for the precise magnification requirements of specific regions; therefore, when existing image magnification technologies deal with overlapping regions, they often cannot balance detail richness and target clarity, resulting in poor final segmentation image effects; In summary, existing image semantic segmentation technologies have obvious defects in dealing with small targets and overlapping regions, and there is an urgent need for an image segmentation method that can accurately segment small targets and flexibly handle overlapping regions. Summary of the Invention
[0003] The purpose of the present invention is to solve the problems in the background art, and to propose an image semantic segmentation method and system based on dynamic convolution attention that can define the maximum magnification upper limit formula according to the object overlap coefficient and local detail richness coefficient; then set correction and suppression terms based on the above coefficients, and determine the dynamic adjustment strategy of the magnification factor according to the threshold comparison.
[0004] To achieve the above purpose, the present invention adopts the following technical solutions. An image semantic segmentation method based on dynamic convolution attention includes the following steps: Collect the image to be segmented, preprocess the image to be segmented to obtain multiple to-be-segmented images of different scales; input the multiple to-be-segmented images of different scales into a preset image semantic segmentation model for image segmentation to obtain an image segmentation result; Collect the morphological feature information of the small target image, perform data processing on it to obtain the object overlap coefficient, compare the object overlap coefficient of each small target image with the overlap threshold, and obtain the to-be-magnified object overlap region image according to the comparison result; Based on the overlapping region images of the objects to be enlarged corresponding to each small target image and their corresponding texture feature information, analyze and calculate each texture feature information to obtain the local detail enrichment coefficient, and construct an image magnification detection model in combination with the object overlapping coefficient. Obtain the magnification factor corresponding to each overlapping region image of the object to be enlarged through the image magnification detection model; Apply the magnification factor obtained according to the image magnification detection model to the overlapping region image of the object to be enlarged for magnification to obtain the magnification result of the overlapping region image of the object, and judge and iterate whether the magnification result of the overlapping region image of the object meets the preset conditions; if it meets, end the iteration and perform image segmentation to obtain the final image segmentation result. If it does not meet, continue the iteration and return to the step of determining the magnification factor.
[0005] In a preferred embodiment, the image segmentation result is the large target image and the small target image segmented from the image to be segmented.
[0006] In a preferred embodiment, the morphological feature information includes the boundary distance and the contour intersection area; the acquisition logic of the boundary distance is to record the minimum boundary distance as D a , according to the formula ; where D a represents the minimum boundary distance between the contour C i of each small target image and another contour C j , a represents the total number of small target images, i represents the index of the contour of the i-th small target image and , represents the contour point set of the i-th small target image, is the contour point set adjacent to the j-th contour, p and q respectively represent a point in the contour C i and another contour C j , represents the square of the coordinate difference between two points p and q on the x-axis , represents the square of the coordinate difference between two points p and q on the y-axis; The acquisition logic of the contour intersection area is defined by performing a per-pixel logical AND operation and combining the Hadamard product that is mathematically equivalent to a matrix to obtain the overlapping region of the two entity contours ; where represents the pixel value of the i-th small target image at the coordinate (x,y), M j (x,y) represents the pixel value of the j-th small target image adjacent to the i-th small target image at the coordinate (x,y); Obtain the judgment function through a per-pixel logical AND operation and combining the Hadamard product that is mathematically equivalent to a matrix ; where 0 indicates that the pixel does not belong to the contour region of the small target image, and 1 indicates that the pixel belongs to the contour region of the small target image; Traverse and sum the above contour overlapping regions to obtain the contour intersection area ; where represents the intersection area of the contour overlapping region between the i-th small target image and the adjacent small target image, W represents the number of pixels of the image in the x direction, and H represents the number of pixels of the image in the y direction; Based on the boundary distance D i a and the contour intersection area B i a Perform weighting to obtain the object overlap coefficient ; where represents the contour C of the i-th small target image i and the adjacent contour C j The minimum boundary distance between them, represents the intersection area of the contour overlapping region between the i-th small target image and the adjacent small target image, ω 1 、ω 2 respectively represent the weights corresponding to the boundary distance and the contour intersection area.
[0007] In a preferred embodiment, comparing the object overlap coefficient of each small target image with the overlap threshold, and obtaining the image of the object overlap region to be enlarged according to the comparison result: comparing the object overlap coefficient OC i a of each small target image with the overlap threshold P; When the object overlap coefficient , then there is an overlap region that needs further analysis in the small target image, which is marked as the image of the object overlap region to be enlarged; When the object overlap coefficient , mark the corresponding small target image as the normal segmentation result.
[0008] In a preferred embodiment, the texture feature information includes the gradient amplitude corresponding to the image of the object overlap region to be enlarged; the acquisition logic of the local detail richness coefficient is: For each object overlap region Re to be enlarged, calculate its local detail richness coefficient LDR e , defined as the weighted average of the gradient amplitudes within the region: ; where, is the number of pixels of each object overlap region Re to be enlarged, is the gradient amplitude of each pixel point; e is the index of the object overlap region to be enlarged and e = 1, 2,..., u; u is the number of object overlap regions to be enlarged.
[0009] In a preferred embodiment, the object overlap coefficient OC corresponding to the overlapping area image of each object to be magnified is i a Recorded as OC e And combined with its corresponding local detail rich coefficient LDR e The process of building an image magnification detection model is as follows: S1: Define the upper limit formula of the maximum magnification factor based on the object overlap coefficient and the local detail richness coefficient; The maximum magnification M max Constrained by the overall image characteristics, it is defined as a balance function between detail and overlap coefficients: ;in, is the global average detail coefficient; is the global average overlap coefficient; where e is the index of the overlapping area of the object to be magnified and e=1,2,...,u; u is the number of overlapping areas of the object to be magnified, LDR e is the local detail richness coefficient of the e-th area to be enlarged ,OC e is the object overlap coefficient of the e-th area to be magnified, ɑ is the empirical weight, and ϵ is a minimum constant; S2: Determine the optimal magnification based on the maximum magnification and local characteristics through dynamic adjustment; The local characteristic is dynamically adjusted to a ratio of a local detail richness coefficient to a global detail richness coefficient; ; Where γ is the preset adaptive coefficient, λ is the overlap suppression weight, θ is the preset overlap threshold, and A1 represents the object overlap coefficient When the maximum magnification is reached, the dynamic adjustment coefficient corresponding to the maximum magnification is A2, which represents the object overlap coefficient. When , the dynamic adjustment coefficient corresponding to the maximum magnification; when When the overlapping area of the object to be magnified is selected Determine the optimal magnification; when When the overlapping area of the object to be magnified is selected Determine the optimal magnification; S3: Integrate the dynamic adjustment process of calculating the upper limit of the optimal magnification in step S1 and determining the optimal magnification in step S2 to form a global constraint mechanism and a local adaptive mechanism, and construct an image magnification detection model.
[0010] The present invention provides a system for implementing an image semantic segmentation method based on dynamic convolution attention, including: an image preprocessing and initial segmentation module, an object overlapping region determination module, a magnification decision module, and an image magnification and iteration judgment module; The image preprocessing and initial segmentation module is used to collect the image to be segmented, preprocess the image to be segmented to obtain multiple images to be segmented with different scales; input the multiple images to be segmented with different scales into a preset image semantic segmentation model for image segmentation to obtain an image segmentation result; The object overlapping region determination module is used to collect the morphological feature information of the small target image, perform data processing on it to obtain an object overlapping coefficient, compare the object overlapping coefficients of each small target image with an overlapping threshold, and obtain an image of the object overlapping region to be magnified according to the comparison result; The magnification decision module is used to analyze and calculate the local detail enrichment coefficient based on the image of the object overlapping region to be magnified corresponding to each small target image and its corresponding texture feature information, construct an image magnification detection model in combination with the object overlapping coefficient, and obtain the magnification multiple corresponding to each image of the object overlapping region to be magnified through the image magnification detection model; The image magnification and iteration judgment module is used to apply the magnification multiple obtained by the image magnification detection model to the image of the object overlapping region to be magnified for magnification to obtain a magnified result of the image of the object overlapping region, and judge whether the magnified result of the image of the object overlapping region after iteration meets a preset condition; if it meets, the iteration ends, and if it does not meet, the iteration continues.
[0011] Compared with the existing technology, the beneficial effects of the present invention are as follows: 1. By collecting the image to be segmented and preprocessing it, and inputting the preprocessed image to be segmented into a preset model for segmentation, a preliminary segmentation result is obtained; collecting the morphological features of small targets, obtaining the object overlapping coefficient through data processing, and screening the image of the object overlapping region to be magnified; through segmenting images of different scales and specifically processing the object overlapping region to be magnified, not only can the image details and targets of different sizes be better captured, the image of the object overlapping region to be magnified be clarified, but also the subsequent processing can focus on the key parts in the image that are prone to segmentation difficulties due to overlap, avoiding non-discriminatory processing of the entire image, and improving the processing pertinence and resource utilization efficiency; 2. The present invention defines an upper limit formula for the maximum magnification factor by considering the image of the area to be magnified and its texture features, based on the object overlap coefficient and the local detail enrichment coefficient. Based on the maximum magnification factor, combined with the data of the object overlap coefficient and the local detail enrichment coefficient, a local detail correction term and an overlap suppression term are set, and a dynamic adjustment strategy for the magnification factor is obtained through comparison with a threshold. Then, the image of the area to be magnified is magnified according to the magnification factor, and it is determined whether the result meets the preset conditions. If it meets, the final segmentation result is obtained; if not, the magnification factor is recalculated and returned. It can not only divide different overlapping areas through the threshold, dynamically combine local details and global overlap suppression, but also avoid the distortion diffusion in high-overlap areas while ensuring detail enhancement, achieving refined magnification. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] Figure 1 is a flowchart of an image semantic segmentation method based on dynamic convolutional attention proposed by the present invention; Figure 2 is an overall structural block diagram of an image semantic segmentation system based on dynamic convolutional attention proposed by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0013] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments.
[0014] In the description of the present invention, it should be understood that the orientation or positional relationship indicated by the terms "upper", "lower", "front", "rear", "left", "right", "top", "bottom", "inner", "outer", etc. is based on the orientation or positional relationship shown in the accompanying drawings, and is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of the present invention.
[0015] Refer to Figure 1 - Figure 2 An image semantic segmentation method based on dynamic convolutional attention includes the following steps: Step 1: Collect the image to be segmented, preprocess the image to be segmented to obtain multiple images to be segmented at different scales; input the multiple images to be segmented at different scales into a preset image semantic segmentation model for image segmentation to obtain an image segmentation result; It should be noted that the preset image semantic segmentation model described in this application is used to segment large target images and small target images from the image to be segmented to achieve target segmentation. Among them, the image to be segmented is a high-resolution remote sensing image, and the large target image and the small target image represent large-scale ground object targets and small-scale ground object targets in the image to be segmented; Specifically, the process of obtaining the image segmentation result is as follows: S1: Input and feature extraction; The preprocessing of the image to be segmented is specifically to extract the geographic metadata of the image to be segmented (such as projection parameters and affine transformation matrix in the GeoTIFF format), store this information, and determine the target coordinate system; Input the preprocessed images to be segmented at multiple different scales into a preset image semantic segmentation model established based on GCM-LANet; In the feature extraction stage of the preset image semantic segmentation model, first, the image features are initially extracted through ResNet50, then the local context information is enhanced through PAM, high-level semantic information is embedded through AEM, and finally, the multi-scale information is fused through the GCM module and the receptive field is further enlarged, so as to extract the features in the image more comprehensively; S2: Multi-scale prediction and weighted fusion; Multi-scale prediction: Based on the input images at different scales, the model makes predictions respectively to obtain the prediction results at different scales; since the images at different scales contain target information of different sizes, multi-scale prediction can capture richer target features; Weighted fusion: Perform weighted averaging on the prediction results at different scales to obtain the final image segmentation result. The determination of the weights can be adjusted according to factors such as experimental experience and the performance of the model at different scales; for example, the weights can be determined according to the accuracy rate or other evaluation indicators of the model on the validation set at each scale, so that the scales with better performance have larger weights; S3: Output the result; Finally, output the image segmentation result, which includes large target images and small target images, corresponding to large-size ground object targets and small-size ground object targets in the image to be segmented respectively; after the model outputs the segmentation result, extract the boundary pixel coordinates (x, y) of each target through contour detection (such as findContours in OpenCV); It should be noted that the above preset image semantic segmentation model is established based on GCM-LANet, including the basic network LANet (Local Attention Network) based on the preliminary feature extraction network (ResNet50), patch attention module (PAM), and attention embedding module (AEM); and a global convolution module (GCM) and activation function (FReLU) are added to form a complete image semantic segmentation model; Model training process: Divide the preprocessed multi-scale images to be segmented into a training set, a validation set, and a test set according to a certain ratio (such as 6:2:2), and input the multi-scale images to be segmented in the training set into the preset image semantic segmentation model established based on GCM-LANet; During the forward propagation of the model, the image is processed through various modules (including ResNet50, PAM, AEM, GCM, etc.) to obtain the predicted segmentation result; the loss value is calculated based on the predicted result and the ground truth label, and then the gradient is calculated through the backpropagation algorithm, and the parameters of the model are updated according to the optimization algorithm; Through the above training process, the image semantic segmentation model is obtained; It should be noted that the establishment process of the above preset image semantic segmentation model is prior art and will not be elaborated in this embodiment; in this embodiment, the size of the small target image is defined as 512×512, with the upper left corner of the small target image as the origin (0,0), the horizontal direction as the x-axis (column coordinate), and the vertical direction as the y-axis (row coordinate).
[0016] Step 2: Collect the morphological feature information of the small target image, perform data processing on it to obtain the object overlap coefficient, compare the object overlap coefficient of each small target image with the overlap threshold, and obtain the image of the object overlap area to be enlarged according to the comparison result; Among them, the morphological feature information includes the boundary distance and the contour intersection area; The acquisition logic of the boundary distance is to calculate the minimum distance between any two different contours in the contours of all small target images, and take the global minimum value as the boundary distance , according to the formula: ; where represents the contour C of the i-th small target image i and the adjacent contour C j The minimum boundary distance between them, a represents the total number of small target images, i represents the index of the contour of the i-th small target image and , represents the contour point set of the i-th small target image, is the contour point set adjacent to the j-th contour, p and q respectively represent the contour C i and the adjacent contour C j One point in, represents the square of the coordinate difference of two points p and q on the x-axis, represents the square of the coordinate difference of two points p and q on the y-axis; The acquisition logic of the contour intersection area is to define the overlapping area of the two entity contours to obtain ; where represents the pixel value of the i-th small target image at the coordinate (x,y), M j(x, y) represents the pixel value at the coordinate (x, y) of the j-th small target image adjacent to the i-th small target image; The judgment function is obtained through a per-pixel logical AND operation combined with the Hadamard product that is mathematically equivalent to a matrix ; where 0 indicates that the pixel does not belong to the contour area of the small target image, and 1 indicates that the pixel belongs to the contour area of the small target image; It should be noted that the above formula means that only when the pixel belongs to the contour areas of both objects at the same time, the value of the overlapping area mask at this coordinate is 1, otherwise it is 0; in this way, the pixel distribution of the overlapping area in the image is logically determined; Traverse and sum the above contour overlapping areas to obtain the contour intersection area ; among them, represents the intersection area of the contour overlapping area between the i-th small target image and the adjacent small target image, W represents the number of pixels in the x direction of the image, that is, the image width, H represents the number of pixels in the y direction of the image, that is, the image height, and W and H determine the range of traversing and summing the image; Based on the boundary distance D i a and the contour intersection area B i a perform weighting to obtain the object overlapping coefficient Among them, the contour C of the i-th small target image i and the adjacent contour C j the minimum boundary distance between them, represents the intersection area of the contour overlapping area between the i-th small target image and the adjacent small target image, ω 1 、ω 2 respectively represent the weights corresponding to the boundary distance and the contour intersection area, and their specific values are set by the researchers according to actual needs; Compare the object overlapping coefficient OC i a corresponding to each small target image with the overlapping threshold P; among them, the specific value of the overlapping threshold P is set by the researchers according to actual needs; When the object overlapping coefficient , then there is an overlapping area that needs further analysis in this small target image, which is marked as the image of the object overlapping area to be enlarged; When the object overlapping coefficient , then the overlapping situation is within the acceptable range, and no further processing is required. The corresponding small target image is directly marked as the normal segmentation result.
[0017] Step 3: Collect the images of the overlapping regions of the objects to be magnified corresponding to each small target image and their corresponding texture feature information, analyze and calculate each texture feature information to obtain the local detail enrichment coefficient, and construct an image magnification detection model in combination with the object overlapping coefficient. Obtain the magnification factor corresponding to each image of the overlapping region of the object to be magnified through the image magnification detection model; Among them, the texture feature information includes the gradient amplitude corresponding to the image of the overlapping region of the object to be magnified; The acquisition logic of the above local detail enrichment coefficient is as follows: For each overlapping region Re of the object to be magnified, calculate its local detail enrichment coefficient LDR e , which is defined as the weighted average of the gradient amplitudes within the region: ; where is the number of pixels in each overlapping region Re of the object to be magnified, is the gradient amplitude of the pixel point (x, y), and e is the number of overlapping regions of the object to be magnified; It should be noted that the specific value of the above gradient amplitude can be directly obtained, and the specific process is the prior art and will not be elaborated too much; Denote the object overlapping coefficient OC i a of the image corresponding to each overlapping region of the object to be magnified as OC e and combine its corresponding local detail enrichment coefficient LDR e The process of constructing the image magnification detection model is as follows: S1: Define the maximum magnification upper limit formula based on the object overlapping coefficient and the local detail enrichment coefficient; The maximum magnification M max is constrained by the overall image characteristics and is defined as a balance function of details and overlapping coefficients: ; where is the global average detail coefficient; is the global average overlapping coefficient; where e is the index of the overlapping region of the object to be magnified and e = 1, 2,..., u; u is the number of overlapping regions of the object to be magnified, LDR e is the local detail enrichment coefficient of the e-th region to be magnified, OC e is the object overlapping coefficient of the e-th region to be magnified, ɑ is the empirical weight, and ϵ is a very small constant; S2: Dynamically adjust based on the maximum magnification and local characteristics to determine the optimal magnification; The dynamic adjustment of the local characteristics is the ratio of the local detail enrichment coefficient to the global detail enrichment coefficient; ; where γ is the preset adaptive coefficient, λ is the overlapping suppression weight, θ is the preset overlapping threshold, and A1 represents the object overlapping coefficient When, the dynamic adjustment coefficient corresponding to the maximum magnification, and A2 represents the object overlap coefficient When, the dynamic adjustment coefficient corresponding to the maximum magnification; When When, the overlapping area selection of the object to be magnified is performed to determine the optimal magnification; When When, the overlapping area selection of the object to be magnified is performed to determine the optimal magnification; Specifically, when When, it represents that the overlapping area of the object to be magnified is a high-overlap area, and the overlapping damage in this area is relatively strong. It is necessary to suppress its influence, and the effect of OC is restricted by a fixed threshold θ e function;
[0018] S3: Integrate the above processes of calculating the upper limit of the maximum magnification and the pixel-level optimal magnification to form a global constraint mechanism and a local adaptive mechanism, and construct an image magnification detection model. Specifically, input the object overlap coefficient and local detail richness coefficient data of the small target image, and obtain the upper limit constraint of the magnification of the entire small target image through the formula for calculating the upper limit of the maximum magnification; based on this, combined with the local characteristics of each pixel, determine the magnification of each pixel in the image through the optimal magnification calculation process, so as to construct a complete image magnification detection model; this model can adaptively determine the reasonable magnification of each pixel according to the local characteristics of the image; In this embodiment, by defining the formula for the upper limit of the maximum magnification based on the object overlap coefficient and the local detail richness coefficient; and based on the maximum magnification, combined with the object overlap coefficient and local detail richness coefficient data, setting the local detail correction term and the overlap suppression term and combining the comparison of thresholds to obtain the magnification dynamic adjustment strategy, it can not only divide different overlap areas through thresholds, realize the dynamic combination of local details and global overlap suppression, but also avoid the distortion diffusion in the high-overlap area while ensuring detail enhancement, and achieve refined magnification.
[0019] Step Four: Apply the magnification obtained from the image magnification detection model to the image of the object overlapping area to be magnified for magnification, obtain the magnification result of the object overlapping area image, and judge and iterate whether the magnification result of the object overlapping area image meets the preset conditions; if it meets, end the iteration, if not, continue the iteration; including: According to the magnification calculated for each object overlapping area image in Step Three, perform image magnification processing on these areas using nearest neighbor interpolation or bilinear interpolation; The above-mentioned and judging whether the iterative complemented segmentation slice meets the preset conditions after each iterative process includes: calculating the structural similarity between the image of the overlapping region of the object to be magnified before magnification and the image of the overlapping region of the object to be magnified after magnification based on the SSIM (Structural Similarity Index Method); and comparing it with a preset threshold; if the calculation result of the structural similarity is less than or equal to the preset threshold, the iteration ends, and the magnified object overlapping region image is transmitted to Step 1 for image segmentation to obtain the final image segmentation result; if the calculation result of the structural similarity is greater than the preset threshold, continue the iteration, that is, continue to perform the determination and adjustment operations of the magnification factor of the object overlapping region image.
[0020] It should be noted that in this embodiment, calculating the structural similarity between the image of the overlapping region of the object to be magnified before magnification and the image of the overlapping region of the object to be magnified after magnification and using the structural similarity as the mark to end the iteration has the advantage that it can more accurately verify and optimize the magnification operation of the image of the overlapping region of the object to be magnified; In this embodiment, by collecting the image to be segmented and preprocessing the image to be segmented, inputting the preprocessed image to be segmented into a preset model for segmentation to obtain a preliminary segmentation result; collecting the morphological features of small targets, obtaining the object overlapping coefficient through data processing, and screening the image of the object overlapping region to be magnified; according to the image of the region to be magnified and its texture features, defining the upper limit formula of the maximum magnification factor based on the object overlapping coefficient and the local detail enrichment coefficient; and based on the maximum magnification factor, combining the data of the object overlapping coefficient and the local detail enrichment coefficient, setting the local detail correction term and the overlap suppression term and obtaining the dynamic adjustment strategy of the magnification factor through the comparison of the threshold, and magnifying the image of the region to be magnified according to the magnification factor, judging whether the result meets the preset conditions, if it meets, obtaining the final segmentation result, if it does not meet, returning to recalculate the magnification factor; it can not only divide different overlapping regions through the threshold, realize the dynamic combination of local details and global overlap suppression, but also avoid the distortion diffusion in the high-overlap region while ensuring the enhancement of details, and achieve refined magnification.
[0021] The above is only a preferred specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution of the present invention and its inventive concept, makes equivalent substitutions or changes, and all should be covered within the protection scope of the present invention.
Claims
1. A method for image semantic segmentation based on dynamic convolutional attention, characterized in that: include: Collecting images to be segmented, and preprocessing the images to be segmented to obtain multiple images to be segmented at different scales; Inputting the plurality of images to be segmented of different scales into a preset image semantic segmentation model for image segmentation to obtain an image segmentation result; The morphological feature information of the small target image is collected, and the object overlap coefficient is obtained by data processing, and the object overlap coefficient of each small target image is compared with the overlap threshold, and the overlap area image of the object to be enlarged is obtained according to the comparison result; Based on the overlapping area images of the objects to be magnified corresponding to each small target image and the corresponding texture feature information, each texture feature information is analyzed and calculated to obtain the local detail richness coefficient, and an image magnification detection model is constructed in combination with the object overlap coefficient. The magnification factor corresponding to the overlapping area images of each object to be magnified is obtained through the image magnification detection model; The magnification factor obtained according to the image magnification detection model is applied to the image of the overlapping area of the object to be magnified to magnify it, and the magnification result of the image of the overlapping area of the object is obtained. It is then judged whether the magnification result of the image of the overlapping area of the object meets the preset conditions. If so, the iteration is terminated to perform image segmentation to obtain the final image segmentation result. If not, the iteration is continued to return to the step of determining the magnification factor.
2. According to claim 1, the image semantic segmentation method based on dynamic convolutional attention is characterized in that: include: The image segmentation result is a large target image and a small target image segmented from the image to be segmented.
3. The image semantic segmentation method based on dynamic convolutional attention according to claim 1, characterized in that: include: The morphological feature information includes boundary distance and contour intersection area; the logic of obtaining the boundary distance is to calculate the minimum distance between any two different contours in the contours of all small target images, and take the global minimum value as the boundary distance , according to the formula: ;in, Represents the contour C of the i-th small target image i With adjacent contour C j The minimum boundary distance between them, a represents the total number of small target images, i represents the index of the contour of the i-th small target image and , represents the contour point set of the i-th small target image, is the set of contour points adjacent to the jth contour, p and q represent contour C respectively. i With adjacent contour C j A point in represents the square of the difference in coordinates of two points p and q on the x-axis, Represents the square of the difference in coordinates of two points p and q on the y-axis; The logic for obtaining the contour intersection area is to define the overlapping area of the two entity contours. ;in, represents the pixel value of the i-th small target image at the coordinate (x, y), M j (x, y) represents the pixel value of the jth small target image adjacent to the i-th small target image at the coordinate (x, y); The judgment function is obtained by performing pixel-by-pixel logical AND operations and combining the mathematical equivalent of the matrix Hadamard product ; Where 0 means that the pixel does not belong to the small target image contour area, and 1 means that the pixel belongs to the small target image contour area; Traverse and sum the overlapping areas of the above contours to obtain the contour intersection area ;in, represents the cross area of the contour overlap area between the i-th small target image and the adjacent small target image, W represents the number of pixels of the image in the x direction, and H represents the number of pixels of the image in the y direction; Based on the boundary distance D i a and the contour intersection area B i a Weighted to get the object overlap coefficient ;in, Represents the contour C of the i-th small target image i With adjacent contour C j The minimum boundary distance between represents the intersection area of the contour overlap region between the i-th small target image and the adjacent small target image, ω1 and ω2 represent the weights corresponding to the boundary distance and the contour intersection area, respectively.
4. The image semantic segmentation method based on dynamic convolutional attention according to claim 3, characterized in that: The object overlap coefficient of each small target image is compared with the overlap threshold, and the image of the overlap area of the object to be enlarged is obtained according to the comparison result: the object overlap coefficient OC corresponding to each small target image is i a Compare with the overlap threshold P; When the object overlap coefficient , then the small target image has an overlapping area that needs further analysis, which is marked as the overlapping area image of the object to be enlarged; When the object overlap coefficient , the corresponding small target image is marked as a normal segmentation result.
5. The image semantic segmentation method based on dynamic convolutional attention according to claim 1, characterized in that: include: The texture feature information includes the gradient amplitude corresponding to the image of the overlapping area of the object to be magnified; the acquisition logic of the local detail richness coefficient is: For each overlapping area Re of the object to be magnified, calculate its local detail richness coefficient LDR e , defined as the weighted average of the gradient amplitudes within the region: ;in, is the number of pixels in the overlapping area Re of each object to be magnified, is the gradient amplitude of each pixel; e is the index of the overlapping area of the object to be magnified and e=1,2,...,u; u is the number of overlapping areas of the object to be magnified.
6. The image semantic segmentation method based on dynamic convolutional attention according to claim 1, characterized in that: include: The object overlap coefficient OC corresponding to the overlapping area image of each object to be magnified i a Recorded as OC e And combined with its corresponding local detail rich coefficient LDR e The process of building an image magnification detection model is as follows: S1: Define the upper limit formula of the maximum magnification factor based on the object overlap coefficient and the local detail richness coefficient; The maximum magnification M max Constrained by the overall image characteristics, it is defined as a balance function between detail and overlap coefficients: ;in, is the global average detail coefficient; is the global average overlap coefficient; where e is the index of the overlapping area of the object to be magnified and e=1,2,...,u; u is the number of overlapping areas of the object to be magnified, LDR e is the local detail richness coefficient of the e-th area to be magnified, OC e is the object overlap coefficient of the e-th area to be magnified, ɑ is the empirical weight, and ϵ is a minimum constant; S2: Determine the optimal magnification based on the maximum magnification and local characteristics through dynamic adjustment; The local characteristic is dynamically adjusted to a ratio of a local detail richness coefficient to a global detail richness coefficient; ; Where γ is the preset adaptive coefficient, λ is the overlap suppression weight, θ is the preset overlap threshold, and A1 represents the object overlap coefficient When the maximum magnification is reached, the dynamic adjustment coefficient corresponding to the maximum magnification is A2, which represents the object overlap coefficient. When , the dynamic adjustment coefficient corresponding to the maximum magnification; when When the overlapping area of the object to be magnified is selected Determine the optimal magnification; when When the overlapping area of the object to be magnified is selected Determine the optimal magnification; S3: Integrate the upper limit of the optimal magnification calculated in step S1 and the dynamic adjustment process of determining the optimal magnification in step S2 to form a global constraint mechanism and a local adaptive mechanism, and construct an image magnification detection model.
7. A system for implementing the image semantic segmentation method based on dynamic convolutional attention as described in any one of claims 1 to 6, characterized in that: include: The image preprocessing and initial segmentation module is used to collect the image to be segmented, preprocess the image to be segmented, and obtain multiple images to be segmented of different scales; Inputting the plurality of images to be segmented of different scales into a preset image semantic segmentation model for image segmentation to obtain an image segmentation result; The object overlap region determination module is used to collect the morphological feature information of the small target image, and perform data processing on it to obtain the object overlap coefficient, and compare the object overlap coefficient of each small target image with the overlap threshold, and obtain the overlap region image of the object to be enlarged according to the comparison result; The magnification factor decision module is used to analyze and calculate the local detail richness coefficient based on the overlapping area image of the object to be magnified corresponding to each small target image and its corresponding texture feature information, and to build an image magnification detection model in combination with the object overlapping coefficient, and to obtain the magnification factor corresponding to the overlapping area image of each object to be magnified through the image magnification detection model; The image magnification and iteration judgment module is used to apply the magnification factor obtained by the image magnification detection model to the image of the overlapping area of the object to be magnified to magnify it, obtain the magnification result of the image of the overlapping area of the object, and judge whether the magnification result of the image of the overlapping area of the iterative object meets the preset conditions; if so, the iteration is terminated, if not, the iteration is continued.
Citation Information
Patent Citations
Blind person intelligent navigation system and method based on image semantic segmentation
CN118298170A
Semantic segmentation method and system for remote sensing image multi-level mask classification optimization
CN119762779A
Method and system for segmenting overlapping cytoplasm in medical image
US20220237782A1
KR20190119261A