Skeleton detection method, system, equipment and medium
Through the dual-branch aggregation network model, combined with Swin-Transformer and cascading aggregation network, the problem of inaccurate detection of large-scale samples in the existing technology is solved, and the accurate skeleton extraction of large-scale samples is achieved.
Patent Information
- Application Number
- CN202510117273.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-24
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2045-01-24
AI Technical Summary
The existing skeleton detection methods cannot accurately extract the skeleton of the target object when the large-scale samples or target objects account for a large proportion of images. The main reason is that the network lacks the receptive field and multi-scale feature extraction capabilities of the entire image.
The dual-branch aggregation network model is adopted, including the backbone network, the main branch network, the auxiliary branch network and the output network. The backbone network extracts features through Swin-Transformer, the main branch network performs multi-scale extraction and decoding, and the auxiliary branch network models the relationship between feature maps of different resolutions through cascading aggregation processing, and finally generates the skeleton detection results through feature fusion.
The semantic spatial information of the network is enriched through the dual-branch structure, and the multi-scale detection capability can be improved, so as to accurately extract the skeleton of large-scale samples or target objects, surpassing the existing CNN and Transformer-based methods.
Smart Images

Figure CN120047405A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of computer vision, and particularly relates to a skeleton detection method, system, device and medium. Background Art
[0002] Skeleton representation describes the properties of complex and variable object contours, connectivity, symmetry, etc. in the form of simple lines. As an important representation means in the development process of computer vision, it has important application value in the fields of computer vision, computer graphics, image processing, etc.
[0003] Currently, most of the existing skeleton detection means are deep learning-based skeleton detection methods; and there are two types of problems with deep learning-based skeleton detection methods; among them, one is that due to the influence of factors such as illumination, contrast, background noise, etc. in natural images, it is impossible to effectively detect the object skeletons of complex samples; the other is that the scale change range of each part of the object in natural images is large, and it is impossible to accurately capture large-scale features.
[0004] With the application of the method based on Transformer (a deep learning model based on self-attention mechanism) in skeleton detection, for natural image samples, it can effectively detect object skeletons; from a qualitative level, it is found that the existing Transformer-based methods can indeed effectively improve the long-distance modeling ability, and to a certain extent solve the problem of incorrect detection of skeleton regions caused by reasons such as cluttered backgrounds and similar foregrounds and backgrounds in traditional CNN models.
[0005] However, in the face of large-scale image samples in natural images, although the existing Transformer-based methods can achieve approximate global feature extraction, their multi-scale feature modeling ability is still insufficient, as shown in the appendix Figure 1 shown; in the appendix Figure 1 shows the problems of broken or inaccurate skeletons predicted by the existing Transformer-based methods on some difficult samples; combined with the appendix Figure 1 it can be seen that for large-scale samples or objects with a large proportion of the target object in the image, neither the traditional CNN method nor the existing Transformer-based method can accurately extract the skeleton of the target object; the main reasons are two: 1) The target object in the image occupies most of the space of the whole image, and the network needs to have the receptive field of the whole image; 2) Detecting large parts of the target object requires the network to have a high multi-scale feature extraction ability. Summary of the Invention
[0006] In view of the technical problems existing in the prior art, the present invention provides a skeleton detection method, system, device and medium to solve the technical problem that the existing skeleton detection methods cannot accurately extract the skeleton of the target object for large-scale samples or objects with a large proportion of the target object in the image.
[0007] To achieve the above object, the technical solution adopted by the present invention is as follows:
[0008] The present invention provides a skeleton detection method, including:
[0009] Obtain the image I to be detected;
[0010] Input the image I to be detected into the dual-branch aggregation network model, and output the skeleton detection result of the image I to be detected; wherein, the dual-branch aggregation network model includes a backbone network, a main branch network, an auxiliary branch network and an output network;
[0011] The backbone network is used to extract features from the image I to be detected, and obtain several levels of feature maps X with different resolutions i ;
[0012] The main branch network is used to aggregate several levels of feature maps X with different resolutions i to obtain an aggregated feature map F X ; and perform pixel regression processing on the aggregated feature map F X to obtain the main branch prediction result Y X ;
[0013] The auxiliary branch network is used to perform cascaded aggregation processing on several levels of feature maps X with different resolutions i to obtain an aggregated feature map F X '; and perform pixel regression processing on the aggregated feature map F X ' to obtain the auxiliary branch prediction result Y X ';
[0014] The output network is used to perform feature fusion on the main branch prediction result Y X and the auxiliary branch prediction result Y X ' to obtain the skeleton detection result of the image I to be detected.
[0015] Further, the backbone network is a Swin-Transformer network structure.
[0016] Further, the main branch network includes a Sem-FPN module, a ShiftASPP module and a main branch classifier module;
[0017] The Sem-FPN module is used to process several levels of feature maps X with different resolutionsi Perform channel alignment convolution processing and perform feature map size restoration operations to obtain a number of feature maps with consistent channels; perform summation processing on the number of feature maps with consistent channels to obtain an aggregation result;
[0018] The ShiftASPP module is used to perform dilated convolution calculations after performing shifted convolution processing on the aggregation result to obtain an aggregated feature map F X ;
[0019] The main branch classifier module is used to perform convolution processing on the aggregated feature map F X to obtain the main branch prediction result Y X 。
[0020] Furthermore, the ShiftASPP module includes four parallel branch modules; among them, each branch module includes a shifted convolution module and a dilated convolution pooling pyramid module;
[0021] The shifted convolution module is used to perform a first convolution operation on the aggregation result to obtain a first convolution feature map; perform a shifting operation on the first convolution feature map to obtain a shifted feature map; perform a second convolution operation on the shifted feature map to obtain a second convolution feature map; perform an inverse shifting operation on the second convolution feature map to obtain a shifted convolution feature map;
[0022] The dilated convolution pooling pyramid module is used to perform dilated convolution calculation processing on the shifted convolution feature map to obtain an aggregated feature map F X 。
[0023] Furthermore, the auxiliary branch network includes a first-level aggregation layer, a feature fusion module, a second-level aggregation layer, and an auxiliary branch classifier module;
[0024] The first-level aggregation layer is used to parallelly input adjacent three-level feature maps with different resolutions and perform window self-attention calculations to obtain several groups of self-attention results A i ;
[0025] The feature fusion module is used to perform addition fusion operations on several groups of self-attention results A i to obtain a fusion result A with several levels of different scale features;
[0026] The second-level aggregation layer is used to perform self-attention calculations on the fusion result A with several levels of different scale features to obtain an aggregation result; sum the residuals of the feature maps X i with different resolutions of several levels and the aggregation result, and perform mapping in the feature space to obtain an aggregated feature map F X ′;
[0027] The auxiliary branch classifier is used to perform convolution processing on the aggregated feature map F X ′ to obtain the auxiliary branch prediction result Y X ′.
[0028] Furthermore, the loss function of the dual-branch aggregation network model is:
[0029] L = L A + αL B
[0030]
[0031] where L is the total loss of the network; L A is the main branch loss; L B is the auxiliary branch loss; α is a hyperparameter used to balance the main branch and the auxiliary branch; W i is the pixel point weight; X i is the feature map of different resolutions; is the proxy label after the processing of the skeleton line neighbor gradient representation; i is each pixel.
[0032] Furthermore, the image I to be detected is a natural image after data augmentation.
[0033] The present invention also provides a skeleton detection system, including:
[0034] An input module for obtaining the image I to be detected;
[0035] A detection output module for inputting the image I to be detected into the dual-branch aggregation network model and outputting the skeleton detection result of the image I to be detected; wherein, the dual-branch aggregation network model includes a backbone network, a main branch network, an auxiliary branch network and an output network;
[0036] The backbone network is used to perform feature extraction on the image I to be detected to obtain several levels of feature maps X i ;
[0037] The main branch network is used to aggregate several levels of feature maps X i to obtain the aggregated feature map F X ; and perform pixel regression processing on the aggregated feature map F X to obtain the main branch prediction result Y X ;
[0038] The auxiliary branch network is used to perform cascaded aggregation processing on several levels of feature maps X i to obtain the aggregated feature map F X ′; and perform convolution processing on the aggregated feature map F XPerform pixel regression processing to obtain the auxiliary branch prediction result Y X ';
[0039] The output network is used to perform feature fusion on the main branch prediction result Y X and the auxiliary branch prediction result Y X ' to obtain the skeleton detection result of the image I to be detected.
[0040] The present invention provides a skeleton detection device, including:
[0041] A processor suitable for executing a computer program;
[0042] A computer-readable storage medium, in which a computer program is stored. When the computer program is executed by the processor, the described skeleton detection method is executed.
[0043] The present invention provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the described skeleton detection method is implemented.
[0044] Compared with the prior art, the beneficial effects of the present invention are:
[0045] The skeleton detection method provided by the present invention extracts features of the image to be detected through a backbone network to obtain feature maps of several levels with different resolutions, uses a main branch network to perform multi-scale extraction and decoding on the feature maps of several levels with different resolutions, uses an auxiliary branch network to model the relationships between the feature maps of several levels with different resolutions, and performs feature fusion on the results of the main branch network and the auxiliary branch network to achieve skeleton extraction of the image to be detected; in the present invention, a dual-branch structure is used to model multi-scale features to enrich the information in the overall semantic space of the network, enhance the expression ability of the network feature space, improve the multi-scale detection ability, and thus meet the skeleton detection requirements of large-scale samples or objects with a large proportion of the target object in the image.
[0046] Furthermore, a shift operation is introduced in the convolution operation of the ShiftASPP module. By changing the relative positions between the pixels of the convolutional feature map, the long-distance interaction ability of the convolutional feature is improved, so as to improve the information exchange ability of the feature map during the convolution process. By performing a shift operation on a small number of pixels, through convolution operation and activation, the effect of enhancing the nonlinear ability of the network is achieved; secondly, the shifted convolution operation is combined with the atrous spatial pyramid pooling. Through the shift operation between the feature maps, the multi-scale feature extraction ability is strengthened.
[0047] Furthermore, a cascaded aggregation network is proposed in the auxiliary branch to perform cascaded self-attention calculations on feature maps of several different resolutions, which can further extract the similarities between features of different scales, realize the modeling of the correlation degree between multi-level and different-scale feature maps of the Transformer, supplement the feature expression of the main branch, enrich the information in the overall feature space of the network, and enhance the expression ability of the network feature space.
[0048] The skeleton detection system, skeleton detection device, and computer-readable storage medium provided by the present invention have all the advantages of the above-mentioned skeleton detection method. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0050] Figure 1 It is a skeleton result graph predicted by the existing Transformer-based method on difficult samples;
[0051] Figure 2 It is a schematic structural diagram of the dual-branch aggregation network model in Embodiment 1;
[0052] Figure 3 It is a schematic structural diagram of the shifted dilated convolution pooling pyramid module in Embodiment 1;
[0053] Figure 4 It is a schematic structural diagram of the auxiliary branch network in Embodiment 1;
[0054] Figure 5 It is a visualization result graph of the two-level output feature maps of the auxiliary branch network in Embodiment 1;
[0055] Figure 6 It is a qualitative analysis graph of the prediction results of the dual-branch aggregation network model in Embodiment 1 and other existing models;
[0056] Figure 7 It is a structural block diagram of the skeleton detection system provided in Embodiment 2;
[0057] Figure 8 It is a structural block diagram of the skeleton detection device provided in Embodiment 3. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0058] In order to make the technical problems, technical solutions, and beneficial effects solved by the present invention clearer and more understandable, the following specific embodiments are used to further elaborate on the present invention. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0059] Embodiment 1
[0060] As shown in the Figure 2 attachment, Embodiment 1 provides a skeleton detection method, including the following steps:
[0061] Step 1: Obtain the image I to be detected. Among them, the image I to be detected is a natural image that has undergone data augmentation; it should be noted that for the data augmentation method of natural images, existing image processing methods are used, such as pixel-level transformation or spatial-level transformation, which are not the improvement points of this application and will not be elaborated here.
[0062] Step 2: Input the image I to be detected into the Dual-branch Cascaded aggregation Transformer (DCFormer) model, and output the skeleton detection result of the image I to be detected.
[0063] In Embodiment 1, the dual-branch cascaded aggregation network model uses two parallel branches to complement and enrich the semantic space features; among them, the auxiliary branch network performs cascaded sub-attention calculations on the four-level different-scale side outputs of the backbone network, which can further extract the similarity between different-scale features; specifically, the dual-branch cascaded aggregation network model includes a backbone network, a main branch network, an auxiliary branch network, and an output network, as shown in the Figure 2 attachment.
[0064] The backbone network is used to extract features from the image I to be detected and obtain several levels of feature maps X with different resolutions i ; the main branch network and the auxiliary branch network, as the network decoders of the dual-branch cascaded aggregation network model, model multi-scale features to enhance the expression ability of the network feature space; the main branch network is used to aggregate several levels of feature maps X with different resolutions i to obtain an aggregated feature map F X ; and perform pixel regression processing on the aggregated feature map F X to obtain the main branch prediction result Y X ; the auxiliary branch network is used to perform cascaded aggregation processing on several levels of feature maps X with different resolutions i to obtain an aggregated feature map F X '; and perform pixel regression processing on the aggregated feature map F X ' to obtain the auxiliary branch prediction result YX '; The output network is used to process the main branch prediction result Y X and the auxiliary branch prediction result Y X ' for feature fusion to obtain the skeleton detection result of the image I to be detected.
[0065] In this Embodiment 1, the backbone network is a Swin-Transformer network structure. Specifically, after inputting the image I to be detected into the Swin-Transformer network structure, four-level side outputs, namely feature maps X of four different resolutions, are obtained through the four-level network of the Swin-Transformer network structure i , where 1 ≤ i ≤ 4. It should be noted that the Swin-Transformer network structure is a vision model based on Transformer, which solves the problems of computational complexity and insufficient local feature extraction by introducing window attention mechanism, hierarchical structure, and patch merging layer.
[0066] The window attention mechanism is specifically as follows: The input image is segmented into non-overlapping windows, and self-attention is calculated within each window. By restricting the range of attention calculation, it can effectively capture the local features of the image while maintaining computational efficiency. The hierarchical structure is specifically as follows: The Swin Transformer adopts a hierarchical structure, and the features are calculated through self-attention within the window at each layer. As the number of layers increases, the size of the window gradually increases, the resolution of the feature map gradually decreases, and the number of channels gradually increases, enabling the model to capture features at different levels and being suitable for multi-scale vision tasks. The patch merging layer is used to reduce the size of the feature map and increase the number of channels. By merging adjacent patches, it increases the number of channels while reducing the resolution, thereby extracting higher-level features. The shifted window mechanism is specifically as follows: The window of the Swin Transformer shifts half of its size every other layer, enabling the information across windows to be fused through multi-layer stacking, achieving a larger receptive field, enhancing the model's global context capture ability, and maintaining computational efficiency at the same time.
[0067] In this Embodiment 1, the main branch network includes a Sem-FPN module, a ShiftASPP module, and a main branch classifier module. The Sem-FPN module is used to perform channel-uniform convolution processing on feature maps X of several different resolutions i and perform feature map size restoration operations to obtain feature maps with consistent channels; perform summation processing on several feature maps with consistent channels to obtain an aggregation result. The ShiftASPP module is used to perform shifted convolution processing on the aggregation result and then perform dilated convolution calculation to obtain an aggregated feature map F X ; The main branch classifier module is used to process the aggregated feature map FX Perform convolution processing to obtain the main branch prediction result Y X .
[0068] In the main branch network, the Sem-FPN module is used as a multi-scale feature extractor to converge the feature maps with different resolutions output from the 1st to 4th levels of the backbone network; the specific operation is as follows: for the feature maps X with different resolutions at each level i Perform channel-uniform convolution and perform an operation to restore the size of the feature map. Finally, the obtained results are integrated by summation, and then connected to the ShiftASPP module and the main branch classifier module at the back; specifically, input the feature map X i into the Sem-FPN module for aggregation to obtain an aggregation result; sum up and send the aggregation result into the ShiftASPP module to obtain the aggregated feature map F X ; then, for the aggregated feature map F X perform pixel regression through a 1×1 convolutional layer to obtain the main branch prediction result Y X .
[0069] As shown in the appendix Figure 3 The ShiftASPP module includes four parallel branch modules; among them, each branch module includes a shifted convolution module and an atrous convolution pooling pyramid module; the shifted convolution module is used to perform a first convolution operation on the aggregation result to obtain a first convolutional feature map; perform a shift operation on the first convolutional feature map to obtain a shifted feature map; perform a second convolution operation on the shifted feature map to obtain a second convolutional feature map; perform an inverse shift operation on the second convolutional feature map to obtain a shifted convolutional feature map; the atrous convolution pooling pyramid module is used to perform atrous convolution calculation on the shifted convolutional feature map to obtain the aggregated feature map F X .
[0070] Specifically, in each branch module, the input feature map undergoes a 3×3 convolution operation, and then the obtained feature map is subjected to a shift operation. The obtained shifted feature map is convolved again with 3×3, and the result is shifted back to restore the feature to its original position, enhancing the element interaction during the convolution process through the shift operation; finally, through the ASPP module, atrous convolution calculation is performed to extract semantic information at different scales for the feature maps at the same scale. In this way, through the four parallel branches, the calculation can be completed.
[0071] It should be noted that the Atrous Spatial Pyramid Pooling (ASPP) is an important means of multi-scale feature extraction and aggregation, which makes an important contribution to the improvement of skeleton detection performance. At the same time, inspired by the shifted window self-attention in Swin Transformer, the shifting operation can strengthen the information interaction within the feature map when the window size is determined, enabling features that were originally not in the same window to move into the same window through the shifting operation, completing the self-attention calculation, and extracting richer semantic information. Therefore, applying the shifting operation to the convolutional operation aims to improve the information interaction of the convolutional feature map, enabling the convolutional operation to indirectly increase the receptive field size. In the ShiftASPP module, the shifted convolutional operation is combined with the atrous spatial pyramid pooling. Through the shifting operation between feature maps, the multi-scale feature extraction ability is strengthened, effectively improving the performance of the network.
[0072] It should also be noted that the purpose of the shifted convolutional operation is to improve the information exchange ability of the feature map during the convolutional process. By performing a shifting operation on a small number of pixels, followed by convolutional operations and activation, the effect of enhancing the non-linear ability of the network is achieved. Specifically, in the shifted convolutional module, the shifting operation is introduced into the convolutional operation, and an inverse transformation operation is performed on the feature map after convolution.
[0073] The working principle of the ShiftASPP module is described in detail as follows by taking a set of feature maps containing a preset number as an example:
[0074] First, set each group to contain 12 feature maps. After convolving the feature maps to increase the number of channels to the maximum number of groups, then perform cyclic left, right, up, and down shifts on the first four groups respectively to obtain pixel-shifted feature maps. Subsequently, perform convolutional operations and activation. Finally, after shifting the shifted feature maps back to the original pixel positions in the reverse direction, perform subsequent atrous convolutional operations.
[0075] Specifically, the feature map input to the ShiftASPP module is the feature map that has undergone multi-scale feature extraction by the Sem-FPN module and contains multi-scale semantic information. Next, it is further decoded by ShiftASPP, that is, dilated convolution is used on the single-scale feature map for different-scale sampling operations, aiming to further mine multi-scale feature information. The feature map first passes through the first convolutional layer in the ShiftConvBlock to increase the number of channels of the feature map to meet the subsequent requirements for grouping the feature map. At the same time, it can also further perform re-projection operations on different-scale features in the feature space to enhance the differentiation of features in each branch and improve the network's expression ability. Subsequently, a partial channel shift operation is performed on each parallel branch. The purpose of the shift operation is to change the pixel order between the original feature maps by performing a cyclic shift operation on some groups. By shuffling the pixel positions of the feature map, elements that could not be filtered simultaneously originally will be allowed to perform the next linear mapping, and the resulting feature map will contain more feature information compared to the unshifted feature map, which improves the interaction ability between elements in the feature map to a certain extent. Since the skeleton detection task is a pixel-level task and the ultimate goal is to classify each pixel point in the feature map, after the second convolution operation, an inverse shift operation needs to be performed on the feature map to restore the initial positions of each element and continue with the subsequent multi-scale feature extraction.
[0076] In this Embodiment 1, a cascaded aggregation layer and a classifier are set in the auxiliary branch network. Among them, the cascaded aggregation layer associates the four-level side outputs of the Transformer. On the basis that the Transformer focuses on the self-attention of each level, the cascaded aggregation network focuses on the self-attention of feature maps at different levels to find more semantic information, thereby enriching the semantic space representation. Specifically, the auxiliary branch network includes a first-level aggregation layer, a feature fusion module, a second-level aggregation layer, and an auxiliary branch classifier module, as shown in the appendix Figure 4 as follows.
[0077] The first-level aggregation layer is used to parallelly input feature maps with different resolutions of adjacent three levels and perform window self-attention calculations to obtain several groups of self-attention results A i ; the feature fusion module is used to perform an addition fusion operation on several groups of self-attention results A i to obtain a fusion result A with several levels of different-scale features; the second-level aggregation layer is used to perform self-attention calculations on the fusion result A with several levels of different-scale features to obtain an aggregation result; the residual of feature maps X i with several levels of different resolutions is summed with the aggregation result and mapped in the feature space to obtain an aggregated feature map F X '; the auxiliary branch classifier is used to classify the aggregated feature map FX Perform convolution processing to obtain the auxiliary branch prediction result Y X '.
[0078] In the auxiliary branch network, by designing a cascaded pooling layer, two-level self-attention calculations are performed on the output features of the backbone network at different scales to refine the relationship degree between feature maps at different scales and enhance the feature representation ability of the overall model in the semantic feature space; among them, the cascaded pooling layer is divided into upper and lower levels, the first level is divided into two modules, which are respectively used to extract feature maps of adjacent three scales; the second level will re-pool the two feature maps at different scales calculated in the first level, and the obtained result is further relationship modeling of the four levels of different scales of the Swin Transformer network structure.
[0079] Specifically, three adjacent levels among the four-level side outputs of the Swin Transformer network structure are respectively input into the cascaded pooling layer, that is, three adjacent levels among the feature maps X i , 1≤i≤4 are respectively input; among them, taking S 1 = {X 1 , X 2 , X 3} and S 2 = {X 2 , X 3 , X 4} as two groups of inputs, and are input into the cascaded pooling layer in parallel; first, through the first-level pooling layer, window self-attention calculations are performed on the inputs S 1 and S 2 to obtain two groups of self-attention results A 1 and A 2 ; subsequently, in the feature fusion module, the two groups of self-attention results A 1 and A 2 are subjected to an addition fusion operation to obtain a fusion result A with four-level different-scale features, and then the fusion result A with four-level different-scale features is input into the second-level pooling layer to perform self-attention calculation, complete the re-extraction and projection operations of the four-level scale features, and then use the residuals of the feature maps X i at four different resolutions to sum with the pooling result and map again in the feature space to complete the entire pooling operation process.
[0080] In this Embodiment 1, both the main branch network and the auxiliary branch network are supervised with the proxy label after the skeleton line near-neighbor gradient representation processing; for the main branch prediction result Y X and the auxiliary branch prediction result Y XIn the process of performing feature fusion to obtain the skeleton detection result of the image I to be detected, the weighted Euclidean loss is adopted, and the main branch prediction result Y X and the auxiliary branch prediction result Y X ′ are connected; among them, the preset hyperparameters are mainly used to balance the contribution degrees of the main branch network and the auxiliary branch network to the feature space. Since the auxiliary branch only models the features between different scales and does not further perform feature decoding, it is used as the auxiliary branch; therefore, the loss function of the double-branch convergence network model is:
[0081] L = L A + αL B
[0082]
[0083] where L is the total loss of the network; L A is the main branch loss; L B is the auxiliary branch loss; L A and L B are both weighted Euclidean losses; α is a hyperparameter used to balance the main branch and the auxiliary branch; W i is the pixel point weight; X i is the feature map of different resolutions; is the proxy label after the processing of the skeleton line neighbor gradient representation; i is each pixel.
[0084] Experimental results and analysis:
[0085] In order to verify the performance of the double-branch convergence network model in skeleton detection, it is trained and evaluated on four preset datasets, and compared with the SkeSwin method and the CNN-based methods. The evaluation metric is F-measure; among the CNN-based methods, DeepFlux and ProMask improved the network proxy label, Hi-Fi and MSB+ adopted the multi-scale image fusion strategy, GeoSkeletonNet obtained geometric perception ability using the Hough distance, and AdaLSN adopted the neural network architecture search method. The above methods are all CNN-based methods; the SkeSwin method is based on Transformer.
[0086] The quantitative comparison results of the DCFormer and the classic skeleton detection methods are shown in Table 1 below; in Table 1 below, methods 1-9 are all CNN-based methods, methods 10-14 are Swin Transformer-based skeleton detection methods SkeSwin, and methods 15-18 are dual-branch cascade convergence networks DCFormer; among them, DCFormer-T, DCFormer-S, DCFormer-B and DCFormer-L represent Tiny, Small, Base and Large variants based on Swin Transformer, respectively.
[0087] Table 1 Quantitative comparison of DCFormer and classic skeleton detection methods
[0088]
[0089]
[0090] It should be noted that some method names in Table 1 above are followed by superscripts, which represent their backbone networks. V stands for VGG16, R stands for ResNet50, and I stands for Inception V3. In the DeepFlux method, different suffixes are used. P means that the network output needs to be post-processed to obtain the predicted skeleton image, and E means that the network itself is an end-to-end network, and the network output result is the predicted skeleton image.
[0091] As can be seen from Table 1 above, the DCFormer surpasses all current CNN methods on four commonly used datasets and reaches the most advanced level (State Of The Art, SOTA); it also verifies that the Transformer-based skeleton detection method is still very promising, and the DCFormer will become a new baseline for skeleton detection tasks; among them, DCFormer-L is quantitatively compared with the most advanced CNN method, and the F-measure on the SK-LARGE dataset is 0.794, surpassing the AdaLSN method using neural network structure search; the F-measure on the WH-SYMMAX dataset is 0.903, and the best method also surpasses AdaLSN; the F-measure on the SYM-PASCAL dataset is 0.607, compared with the best CNN method DeepFlux R -E improved by 1.4%; F-measure was 0.550 on the SYMMAX300 dataset, compared with the best CNN method DeepFlux V-E increased by 1.9%; It can be seen that the DCFormer using the Large variant surpasses all current optimal performances and reaches an advanced level.
[0092] In this Embodiment 1, an auxiliary branch network is used to model the feature maps of different scales at all levels to find more semantic information; among them, the visualization displays of the aggregation results of the previous stage and the subsequent stage in the cascaded aggregation network are respectively shown, as shown in the appendix Figure 5 shown; among them, the appendix Figure 5 a and the appendix Figure 5 b are the feature map outputs of the first-level aggregation layer, which respectively represent the modeling of the set S 1 of the first three side outputs and the set S 2 of the last three side outputs of the backbone network respectively; it can be found that for the feature map A 1 after aggregation of S 1 fuses the semantic information of the low-level feature map and contains more low-level semantic information: the blurred outline of the object and a small amount of skeleton information; while for the feature set A 2 after aggregation of S 2 fuses the semantic information of the high-level feature map and contains more high-level semantic contents: a small amount of contour features and skeleton features. This further shows that the aggregation layer can have the ability of multi-scale feature fusion; the appendix Figure 5 c shows the output result of the subsequent stage of the second cascaded aggregation layer. It can be clearly seen that the cascaded aggregation network fuses the multi-scale semantic information of different levels, improves the representation ability of the feature space of the network model, can accurately locate the object and distinguish the background, and secondly fuses the low-level contour information and the high-level skeleton information, and finally presents a state of activation in the skeleton area and inhibition near the skeleton; from the content shown in the appendix Figure 5 it is verified that the idea of modeling the relationship between different-scale features by the cascaded aggregation network can enrich the content of the network semantic feature space, thereby helping the network to better predict the skeleton position.
[0093] In this Embodiment 1, large-scale samples and difficult samples are respectively selected from four datasets, and the prediction images generated by DeepFlux V -P, ProMask, AdaLSN, SkeSwin-L, and DCFormer are respectively used for qualitative analysis, and the qualitative analysis is as shown in the appendix Figure 6 shown; in the appendix Figure 6Among them, the first row shows a difficult sample in the SK-LARGE dataset where the background and the object have similar colors. None of the existing CNN methods can detect the deer's legs. The SkeSwin-L method can roughly detect the skeletal structure of the object's legs, but limited by its multi-scale extraction ability, the detected skeleton is blurred. While the DCFormer-L can distinguish the foreground from the background more excellently and extract the skeleton more accurately. In the second row, there is one of the few difficult samples in WH-SYMMAX. In this sample, the color of the horse's legs blends into the background, making it difficult to distinguish. During the detection process, CNN-based methods tend to overlook the horse's legs, while the DCFormer can distinguish the object from the background excellently. The third row shows an image of a dog in SYM-PASCAL. The scale of the object's body part is large, and classic CNN methods and SkeSwin are unable to effectively detect the object's skeleton. The DCFormer can extract the object's skeleton more accurately. The last row shows the complex sample dataset SYMMAX300, which is an image of a tiger walking in the jungle. The true label of its skeleton is the central axis of each line of the tiger's body, and the detection difficulty is high. However, the DCFormer can effectively capture low-level features such as stripes, thus completing the detection of the stripe central axis.
[0094] For the skeleton detection method described in Embodiment 1, a dual-branch aggregation network model is used to extract the skeleton of the image to be detected. Through experimental verification, it is found that the DCFormer has a better detection effect on large-scale samples and samples with large object-scale changes, thus improving the skeleton detection performance. Specifically, through quantitative experiments, it is verified that the DCFormer exceeds the existing optimal CNN methods on all common datasets and reaches the state-of-the-art level (SOTA). From a qualitative perspective, for complex natural images with similar backgrounds and foregrounds, the DCFormer uses its dual-branch structure to supplement the semantic feature space and solves this problem well to a certain extent. At the same time, the DCFormer inherits the ability of the Transformer to capture long-distance features and better solves the problem of inaccurate skeleton prediction caused by large object scales.
[0095] Embodiment 2
[0096] As shown in the appendix Figure 7 This Embodiment 2 provides a skeleton detection system, including an input module and a detection output module.
[0097] An input module for obtaining an image I to be detected; a detection output module for inputting the image I to be detected into a dual-branch aggregation network model and outputting a skeleton detection result of the image I to be detected; wherein the dual-branch aggregation network model includes a backbone network, a main branch network, an auxiliary branch network, and an output network; the backbone network is used for extracting features from the image I to be detected to obtain several levels of feature maps X with different resolutions i ; the main branch network is used for aggregating several levels of feature maps X i to obtain an aggregated feature map F X ; and performing pixel regression processing on the aggregated feature map F X to obtain a main branch prediction result Y X ; the auxiliary branch network is used for performing cascaded aggregation processing on several levels of feature maps X i to obtain an aggregated feature map F X '; and performing pixel regression processing on the aggregated feature map F X ' to obtain an auxiliary branch prediction result Y X '; the output network is used for performing feature fusion on the main branch prediction result Y X and the auxiliary branch prediction result Y X ' to obtain a skeleton detection result of the image I to be detected.
[0098] Example 3
[0099] As shown in the appendix Figure 8 This Example 3 provides a skeleton detection device, including: a memory for storing a computer program; a processor for implementing the steps of the skeleton detection method when executing the computer program.
[0100] Obtain an image I to be detected; input the image I to be detected into a dual-branch aggregation network model and output a skeleton detection result of the image I to be detected; wherein the dual-branch aggregation network model includes a backbone network, a main branch network, an auxiliary branch network, and an output network; the backbone network is used for extracting features from the image I to be detected to obtain several levels of feature maps X with different resolutions i ; the main branch network is used for aggregating several levels of feature maps X i to obtain an aggregated feature map F X ; and performing pixel regression processing on the aggregated feature map F X to obtain a main branch prediction result Y X ; the auxiliary branch network is used for performing cascaded aggregation processing on several levels of feature maps X i to obtain an aggregated feature map F X '; and performing pixel regression processing on the aggregated feature map FX Perform pixel regression processing to obtain the auxiliary branch prediction result Y X ′; The output network is used to perform feature fusion on the main branch prediction result Y X and the auxiliary branch prediction result Y X ′ to obtain the skeleton detection result of the image I to be detected.
[0101] Alternatively, when the processor executes the computer program, it implements the functions of each module in the above skeleton detection system. For example:
[0102] The input module is used to obtain the image I to be detected; the detection output module is used to input the image I to be detected into the dual-branch aggregation network model and output the skeleton detection result of the image I to be detected. Among them, the dual-branch aggregation network model includes a backbone network, a main branch network, an auxiliary branch network, and an output network; the backbone network is used to perform feature extraction on the image I to be detected to obtain several levels of feature maps X with different resolutions i ; The main branch network is used to aggregate several levels of feature maps X with different resolutions i to obtain an aggregated feature map F X ; and perform pixel regression processing on the aggregated feature map F X to obtain the main branch prediction result Y X ; The auxiliary branch network is used to perform cascade aggregation processing on several levels of feature maps X with different resolutions i to obtain an aggregated feature map F X ′; and perform pixel regression processing on the aggregated feature map F X ′ to obtain the auxiliary branch prediction result Y X ′; The output network is used to perform feature fusion on the main branch prediction result Y X and the auxiliary branch prediction result Y X ′ to obtain the skeleton detection result of the image I to be detected.
[0103] Exemplarily, the computer program can be divided into one or more modules / units. The one or more modules / units are stored in the memory and executed by the processor to complete the present invention. The one or more modules / units can be a series of computer program instruction segments capable of completing preset functions, and the instruction segments are used to describe the execution process of the computer program in the skeleton detection device.
[0104] The skeleton detection device may be a computing device such as a desktop computer, a notebook, a palm computer, or a cloud server. The skeleton detection device may include, but is not limited to, a processor and a memory. Those skilled in the art can understand that the above are examples of the skeleton detection device and do not constitute a limitation on the skeleton detection device. It may include more components than the above, or combine certain components, or different components. For example, the skeleton detection device may also include input / output devices, network access devices, buses, etc.
[0105] The so-called processor may be a central processing unit (CPU), or may also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), off-the-shelf programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor, or the processor may also be any conventional processor, etc. The processor is the control center of the skeleton detection device, and uses various interfaces and lines to connect all parts of the entire skeleton detection device.
[0106] The memory can be used to store the computer program and / or module. The processor realizes various functions of the skeleton detection device by running or executing the computer program and / or module stored in the memory, and by calling the data stored in the memory.
[0107] The memory may mainly include a program storage area and a data storage area. Among them, the program storage area may store an operating system, application programs required for at least one function (such as a sound playback function, an image playback function, etc.); the data storage area may store data created according to the use of the mobile phone (such as audio data, phone book, etc.). In addition, the memory may include high-speed random access memory, and may also include non-volatile memory, such as a hard disk, a memory, a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, at least one magnetic disk storage device, a flash memory device, or other volatile solid-state storage devices.
[0108] Embodiment 4
[0109] Embodiment 4 of the present invention also provides a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the steps of the skeleton detection method described above are implemented.
[0110] If the modules / units integrated in the skeleton detection system are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium.
[0111] Based on such understanding, all or part of the processes in the above-mentioned skeleton detection method of the present invention can also be completed by instructing relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above-mentioned skeleton detection method can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file or preset intermediate form, etc.
[0112] The computer-readable storage medium may include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium, etc.
[0113] It should be noted that the content included in the computer-readable storage medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the computer-readable storage medium does not include electrical carrier signals and telecommunication signals.
[0114] In the present invention, the main branch network is used to perform multi-scale extraction and decoding of skeleton features, and the cascaded aggregation network of the auxiliary branch network is used to model the relationship between feature maps of different resolutions; the overall semantic space information of the network is enriched through the dual-branch structure, and the multi-scale detection ability is improved; specifically, based on Swin Transformer as the backbone network, the powerful multi-scale feature aggregation ability of Sem-FPN is used to decode features, and a parallel self-attention aggregation branch is proposed to find the correlation between different-scale features of the Transformer as a supplement to the semantic space, improving the multi-scale extraction ability of the network; in addition, a shifted convolution operation is also proposed and applied to the atrous spatial pyramid pooling module to strengthen the multi-scale feature extraction ability through the shifting operation between feature maps; through experimental verification, it is found that the method of the present invention can effectively improve the detection accuracy of large-scale samples and reach the state-of-the-art level on four natural scene datasets.
[0115] The above embodiments are merely one of the implementation manners capable of implementing the technical solution of the present invention. The scope of protection required by the present invention is not limited only by this embodiment, but also includes any changes, substitutions, and other implementation manners that are easily conceivable by those skilled in the art within the technical scope disclosed by the present invention.
Claims
1. A skeleton detection method, characterized in that: include: Obtain an image I to be detected; Input the image to be detected I into the dual-branch convergence network model, and output the skeleton detection result of the image to be detected I; wherein the dual-branch convergence network model includes a backbone network, a main branch network, an auxiliary branch network and an output network; The backbone network is used to extract features from the image to be detected I to obtain feature maps X at several levels of different resolutions. i ; The main branch network is used to analyze several levels of feature maps X with different resolutions. i Aggregate and get the aggregate feature map F X ; and the aggregate feature map F X Perform pixel regression processing to obtain the main branch prediction result Y X ; The auxiliary branch network is used to analyze several levels of feature maps X with different resolutions. i Perform cascade aggregation processing to obtain the aggregate feature map F X '; and the aggregate feature map F X ′Perform pixel regression processing to obtain the auxiliary branch prediction result Y X ′; The output network is used to predict the result Y of the main branch X and the auxiliary branch prediction result Y X ′Perform feature fusion to obtain the skeleton detection result of the image I to be detected.
2. A skeleton detection method according to claim 1, characterized in that: The backbone network is a Swin-Transformer network structure.
3. A skeleton detection method according to claim 1, characterized in that: The main branch network includes a Sem-FPN module, a ShiftASPP module and a main branch classifier module; The Sem-FPN module is used to generate feature maps X at different resolutions. i Perform channel-consistent convolution processing and perform feature map size recovery operations to obtain several channel-consistent feature maps; sum up several channel-consistent feature maps to obtain an aggregated result; The ShiftASPP module is used to perform shift convolution processing on the aggregation result and then perform hole convolution calculation to obtain the aggregation feature map F X ; The main branch classifier module is used to classify the aggregate feature map F X Perform convolution processing to obtain the main branch prediction result Y X .
4. A skeleton detection method according to claim 1, characterized in that: The ShiftASPP module includes four parallel branch modules; each branch module includes a shift convolution module and a dilated convolution pooling pyramid module; The shift convolution module is used to perform a first convolution operation on the aggregation result to obtain a first convolution feature map; perform a shift operation on the first convolution feature map to obtain a shift feature map; perform a second convolution operation on the shift feature map to obtain a second convolution feature map; perform a reverse shift operation on the second convolution feature map to obtain a shift convolution feature map; The atrous convolution pooling pyramid module is used to perform atrous convolution calculation processing on the shifted convolution feature map to obtain an aggregated feature map F X .
5. A skeleton detection method according to claim 1, characterized in that: The auxiliary branch network includes a first-level convergence layer, a feature fusion module, a second-level convergence layer and an auxiliary branch classifier module; The first-level convergence layer is used to input the feature maps of three adjacent levels with different resolutions in parallel and perform window self-attention calculation to obtain several groups of self-attention results A i ; The feature fusion module is used to combine several sets of self-attention results A i Perform an additive fusion operation to obtain a fusion result A with features of several levels and different scales; The second-level pooling layer is used to perform self-attention calculation on the fusion result A with several levels of different scale features to obtain the pooling result; the feature maps X with several levels of different resolutions are i The residual of is summed with the aggregation result and mapped in the feature space to obtain the aggregate feature map F X ′; The auxiliary branch classifier is used to classify the aggregate feature map F X 'Convolution processing is performed to obtain the auxiliary branch prediction result Y X ′.
6. A skeleton detection method according to claim 1, characterized in that: The loss function of the dual-branch convergence network model is: L=L A +αL B Where L is the total loss of the network; L A is the main branch loss; L B is the auxiliary branch loss; α is a hyperparameter used to balance the main branch and the auxiliary branch; W i is the pixel weight; X i are feature maps of different resolutions; is the proxy label after the skeleton line neighbor gradient representation processing; i is each pixel.
7. A skeleton detection method according to claim 1, characterized in that: The image to be detected I is a natural image that has been data enhanced.
8. A skeleton detection system, characterized in that: include: An input module is used to obtain an image I to be detected; A detection output module, used to input the image to be detected I into the dual-branch convergence network model, and output the skeleton detection result of the image to be detected I; wherein the dual-branch convergence network model includes a backbone network, a main branch network, an auxiliary branch network and an output network; The backbone network is used to extract features from the image to be detected I to obtain feature maps X at several levels of different resolutions. i ; The main branch network is used to analyze several levels of feature maps X with different resolutions. i Aggregate and get the aggregate feature map F X ; and the aggregate feature map F X Perform pixel regression processing to obtain the main branch prediction result Y X ; The auxiliary branch network is used to analyze several levels of feature maps X with different resolutions. i Perform cascade aggregation processing to obtain the aggregate feature map F X '; and the aggregate feature map F X ′Perform pixel regression processing to obtain the auxiliary branch prediction result Y X ′; The output network is used to predict the result Y of the main branch X and the auxiliary branch prediction result Y X ′Perform feature fusion to obtain the skeleton detection result of the image I to be detected.
9. A skeleton detection device, characterized in that: include: a processor suitable for executing a computer program; A computer-readable storage medium, wherein a computer program is stored in the computer-readable storage medium, and when the computer program is executed by the processor, the skeleton detection method according to any one of claims 1 to 7 is executed.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the skeleton detection method according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Multi-scale sparse convolution change detection method for high-resolution remote sensing image
CN112348814A
Weak and small target detection method and device based on infrared and visible light feature fusion
CN116958782A
Lane line identification method and system based on inspection image of unmanned aerial vehicle
CN117593716A