Linear building group identification method and system fusing dynamic snakelike convolution and YOLO11
By integrating dynamic serpentine convolution with C3k2 modules in the YOLO11 model, the YOLO11-DSC model is solved, and the problem of difficult to identify complex linear building groups in the prior art is achieved, achieving higher recognition accuracy and recall rate.
Patent Information
- Application Number
- CN202510142069.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-10
- Publication Date
- 2025-06-03
AI Technical Summary
It is difficult to effectively identify complex linear building groups, especially in the presence of building interference and complex scenarios.
Fusion of dynamic snake convolution with YOLO11 C3k2 module to form the C3k2_DySnakeConv module, and replace the original C3k2 module in YOLO11 backbone network and neck network to obtain the YOLO11-DSC model.
The recognition accuracy and recall rate of linear building groups are improved, the recognition ability of linear building groups of different sizes and shapes is enhanced, and the detection performance is improved.
Smart Images

Figure CN120088614A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of cartographic generalization, and particularly to a method and system for identifying linear building groups by integrating dynamic snake-shaped convolution and YOLO11. Background Art
[0002] Maps are important carriers for expressing and transmitting geographical information, which have developed from traditional paper maps to digital maps. An important step in cartography is to replace the original elements with fewer elements while retaining their spatial distribution characteristics. However, in real life, building groups have various distribution characteristics, and the linearly distributed building groups are the most common and basic distribution types. Identifying linear building groups is one of the key points in cartographic generalization.
[0003] However, in the existing methods for identifying linear building groups, whether it is a geometric method (such as a proximity graph) or an intelligent algorithm (such as a graph convolutional neural network method), it is necessary to rely on the geometric features of the building group itself to establish a graph structure and judge by describing the geometric relationship between buildings. However, the situation of building groups is complex, and its own judgment is a fuzzy process, which is difficult to express in a relatively unified form. Therefore, it is necessary to find a method that can identify the whole of linear building groups. Summary of the Invention
[0004] In order to at least partially solve the problem that the situation of linear building groups is complex and it is difficult to identify the whole of linear building groups, the present invention provides a method and system for identifying linear building groups by integrating dynamic snake-shaped convolution and YOLO11. The present invention integrates dynamic snake-shaped convolution with the C3k2 module in YOLO11 to obtain the C3k2_DySnakeConv module, and uses this module to replace the original C3k2 module in the backbone network and the neck network, and finally obtains the YOLO11 network structure YOLO11-DSC integrated with dynamic snake-shaped convolution. By constructing a training sample library for linear building groups and using the training sample library to train the YOLOv11-DSC object detection model, the corresponding linear building group recognition model is finally obtained, realizing the whole of linear building groups.
[0005] In order to achieve the above object, the technical solution of the present invention is:
[0006] The first aspect of the present invention proposes a method for identifying linear building groups by integrating dynamic snake-shaped convolution and YOLO11, including:
[0007] Step 1: Collect data of linear building groups, and construct a training data set according to the collected data of linear building groups to facilitate training the model;
[0008] Step 2: Improve the YOLO11 model to obtain the YOLO11-DSC model, which is convenient for accurately identifying linear building groups;
[0009] Step 3: Train and optimize the YOLO11-DSC model according to the training dataset to obtain the optimal YOLO11-DSC model;
[0010] Step 4: Input the target linear building group data into the optimal YOLO11-DSC model to obtain the recognition result of the linear building group.
[0011] Furthermore, the training dataset includes single building group training data, surrounding building interference training data, and complex scene training data;
[0012] The single building group training data includes that there is only one linear building group in the selected area, and there are no other buildings around this group;
[0013] The surrounding building interference training data includes that there is one or more linear building groups in the selected area, and there is interference from buildings that do not form a linear pattern around them;
[0014] The complex scene training data includes that the range involved in the selected area and the total number of buildings are larger than those in the surrounding building interference training data. The total number of buildings includes linear building groups and interfering buildings, and the number of linear building groups and interfering buildings is larger than that of linear building groups and interfering buildings in the surrounding building interference training data.
[0015] Furthermore, the improvement of the YOLO11 model to obtain the YOLO11-DSC model specifically includes:
[0016] Fuse the dynamic snake-shaped convolution with the C3k2 module in the YOLO11 model to obtain the C3k2_DySnakeConv module, and use this module to replace the original C3k2 module in the backbone network and neck network of the YOLO11 model, and finally obtain the YOLO11-DSC model.
[0017] Furthermore, the change of the dynamic snake-shaped convolution in the x-axis direction is expressed by the following formula:
[0018]
[0019] where K i±c is the coordinate offset range of the dynamic snake-shaped convolution in the x-axis direction, x i+c and x i-c are the offset ranges in the x-axis direction respectively, y i+c and yi-c are the offset ranges in the y-axis direction, c is the horizontal distance from the center coordinates, Δy is the offset in the x-axis direction, x i and y i are the x-axis coordinate and y-axis coordinate of the i-th center coordinate respectively;
[0020] The change of the dynamic snake-shaped convolution in the x-axis direction is expressed by the following formula:
[0021]
[0022] where K j±c is the coordinate offset range of the dynamic snake-shaped convolution in the y-axis direction, and Δx is the offset in the y-axis direction.
[0023] Furthermore, the fractional part value of the dynamic snake-shaped convolution coordinates is calculated according to the following formula:
[0024] K = ∑ K′ B(K′, K)·K′
[0025] B(K, K′) = b(K x , K′ x )·b(K y , K′ y )
[0026] where K is the fractional part value of the coordinates of K i±c and K j±c in the convolution grid, K′ is all the integer part values of the original convolution kernel, B is the bilinear interpolation kernel, b is a one-dimensional kernel, K x and K y are K i±c and K j±c respectively the fractional part values of the x-axis and y-axis coordinates in the convolution grid, and K′ x and K′ y are the integer part values on the x-axis and y-axis of the original convolution kernel respectively.
[0027] Furthermore, the C3k2_DySnakeConv module is divided into two structures, specifically as follows:
[0028] When the parameter is set to False, the Bottleneck_DySnakeConv module is used to replace the Bottleneck module in the original C3k2 module; the Bottleneck_DySnakeConv module includes two dynamic snake-shaped convolutions and a concat function; among them, the two dynamic snake-shaped convolutions and the concat function are connected in sequence, and the input data of the first dynamic snake-shaped convolution is also input into the concat function;
[0029] When the parameter is set to True, the Bottleneck_DySnakeConv module is used to replace the Bottleneck module in the original C3k module.
[0030] In a second aspect of the present invention, a linear building group recognition system integrating dynamic snake convolution and YOLO11 is proposed, including:
[0031] A collection module, configured to collect linear building group data and construct a training dataset based on the collected linear building group data for facilitating the training of the model;
[0032] A construction module, configured to improve the YOLO11 model to obtain a YOLO11-DSC model for accurately recognizing linear building groups;
[0033] A training module, configured to train and optimize the YOLO11-DSC model according to the training dataset to obtain an optimal YOLO11-DSC model;
[0034] An identification module, configured to input target linear building group data into the optimal YOLO11-DSC model to obtain a linear building group recognition result.
[0035] In a third aspect of the present invention, an electronic device is proposed, including a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor. When the processor executes the computer program, the method for recognizing a linear building group integrating dynamic snake convolution and YOLO11 as described in the first aspect above is implemented.
[0036] In a fourth aspect of the present invention, a computer-readable storage medium is proposed. The storage medium includes a stored computer program. When the computer program runs, the device where the storage medium is located is controlled to execute the method for recognizing a linear building group integrating dynamic snake convolution and YOLO11 as described in the first aspect above.
[0037] Advantages of the present invention:
[0038] In the backbone network of the YOLO11 object detection model, the present invention introduces a Dynamic Snake Convolution (DSC) module, fuses the DSC module with the C3k2 module in YOLO11 to obtain the C3k2_DySnakeConv module, and uses this module to replace the original C3k2 module in the backbone network and the neck network. Finally, the YOLO11 network structure integrated with DSC, namely the YOLO11-DSC model, is obtained. By integrating the DSC module, the present invention can increase the number of detected objects, improve the recall rate. The recognition accuracy of linear building groups of the YOLO11-DSC model has been improved compared with the prior art, and the recognition ability for linear building groups of different sizes has been enhanced, with stronger detection performance. Description of the Drawings
[0039] Figure 1 It is a flowchart of the method for identifying linear building groups by integrating dynamic snake convolution and YOLO11 provided by an embodiment of the present invention.
[0040] Figure 2 It is a schematic diagram of a linear building group provided by an embodiment of the present invention.
[0041] Figure 3 It is a schematic diagram of the distribution characteristics of buildings in different regions provided by an embodiment of the present invention.
[0042] Figure 4 It is a schematic diagram of an example of a training data set provided by an embodiment of the present invention.
[0043] Figure 5 It is a schematic diagram of the YOLO11 network structure provided by an embodiment of the present invention.
[0044] Figure 6 It is a schematic diagram of the improvement of the YOLO11 detection head provided by an embodiment of the present invention.
[0045] Figure 7 It is a schematic diagram of the comparison between dynamic snake convolution and common convolution provided by an embodiment of the present invention.
[0046] Figure 8 It is a schematic diagram of the calculation process of the convolution kernel coordinates and the corresponding receptive field in the dynamic snake convolution provided by an embodiment of the present invention.
[0047] Figure 9 It is a schematic diagram of the YOLO11-DSC model provided by an embodiment of the present invention.
[0048] Figure 10 It is a schematic diagram of the structure of the C3k2_DySnakeCov module provided by an embodiment of the present invention.
[0049] Figure 11Schematic diagram of dynamic snake-shaped convolution feature extraction provided by an embodiment of the present invention.
[0050] Figure 12 Schematic diagram of comparison of prediction results of test dataset 1 provided by an embodiment of the present invention.
[0051] Figure 13 Schematic diagram of comparison of prediction results of test dataset 2 provided by an embodiment of the present invention.
[0052] Figure 14 Schematic diagram of comparison of prediction results of test dataset 3 provided by an embodiment of the present invention.
[0053] Figure 15 Architecture diagram of a linear building group recognition system that integrates dynamic snake-shaped convolution and YOLO11 provided by an embodiment of the present invention. Detailed implementation manners
[0054] To make the objectives, technical solutions, and advantages of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0055] Embodiment 1
[0056] As Figure 1 shown, a method for recognizing a linear building group that integrates dynamic snake-shaped convolution and YOLO11 includes:
[0057] S101: Collect data of linear building groups, and construct a training dataset according to the collected data of linear building groups.
[0058] Specifically, a linear building group refers to a group of buildings arranged in a linear pattern, and its distribution presents linear geometric morphological characteristics. The distribution pattern of linear building groups is a common regular pattern in large-scale maps, and effective recognition of its type is the premise and basis for subsequent building generalization. According to specific forms, the distribution pattern of linear building groups can be further refined into three different morphological patterns: linear, oblique, and curved, as Figure 2 shown.
[0059] (1) Linear: The center line of the buildings in the linear distribution pattern is linear.
[0060] (2) Oblique: The difference between the oblique pattern and the linear pattern is that the facing ratio between the buildings is relatively small.
[0061] (3) Curvilinear: Its centerline extends and varies according to the building direction, presenting a shape similar to a curve.
[0062] Although the linear distribution pattern can be further refined, the present invention does not make a subdivision. The reasons are as follows: First, the two refinement modes of linear and curvilinear are not significantly different in some scenarios, and the judgment between them is relatively vague without a clear division standard. Moreover, through the analysis of a large amount of data, it is found that building groups exist more in the linear mode, while the curvilinear and oblique line modes are relatively less, and it is not easy to obtain a sufficient number of samples. The sample imbalance will affect the learning effect of the deep convolutional neural network. Second, regardless of the linear distribution mode, a typicalization operator is subsequently used to perform comprehensive operations on the building group. Due to the above two reasons, the linear distribution mode is not further refined, and these three specific forms are uniformly detected as the linear mode.
[0063] Since there is no publicly available dataset for the distribution pattern of building groups at present, a self-made dataset is created manually from the urban building dataset provided by the open-source OSM (OpenStreetMap) website. Since linear building groups mainly present a discrete distribution pattern and are mainly distributed along roads, as Figure 3 shown, in the central area of the city, buildings are mostly adjacent, and show a block distribution characteristic along with the blocks, with a relatively high distribution density and relatively complex building shapes. Because they are mostly public places such as shopping malls and office buildings, it is relatively difficult to form linear building groups; in the suburban and rural areas, there are mostly residential areas, and the buildings are mostly in simple rectangular or quasi-rectangular shapes and are distributed along roads, so they present an obvious linear distribution characteristic. Therefore, when selecting the research area, the suburban and rural areas are mainly used as the research scope, and the building data of the suburban and rural areas are selected. Specifically, the suburbs and surrounding villages of three cities A, B, and C are selected as the research area, and the building data are segmented and divided using the vehicle road layer. When selecting the vehicle road, the roads with the six attributes of "high-speed, first-class, second-class, third-class, residential area, unclassified" are selected as the judgment basis for the vehicle road, that is, the road layer attribute field fclass = motorway, primary, secondary, tertiary, trunk, residential is selected. The selected area has more building groups distributed in a linear shape.
[0064] Since the linear building group distribution pattern is composed of multiple buildings, and linear group distribution patterns composed of buildings with different forms, sizes, and orientations are all possible, it is necessary to take into account various types of buildings. And in order for the model to fully learn the distribution patterns in different scenarios, it is necessary to provide training data with different scales, sizes, and complexities for the model to train and learn.
[0065] Therefore, the training data set is constructed from the following three levels:
[0066] (1) Level 1: Training data for individual building groups. It means that the selected area contains only one linear building group, and there are no other buildings around this group to avoid interference. Its purpose is to enable the model to fully learn the characteristics of linear-distributed building groups.
[0067] (2) Level 2: Training data for interference from surrounding buildings: It means that the selected area contains one or more linear building groups, and there are interferences from buildings that do not form a linear pattern around them. Its purpose is to enable the model to distinguish linear building groups from interfering buildings and train the model's discrimination ability so that it can correctly identify linear building groups in data with interfering buildings.
[0068] (3) Level 3: Training data for complex scenarios: The scope of the selected area and the total number of buildings are further increased. It contains multiple linear building groups, and the number of interfering buildings is also further increased. Its purpose is to further enhance the model's resolution ability so that it can effectively identify linear building groups in a larger range and area.
[0069] Based on the above ideas for constructing the data set, a total of 350 pictures were intercepted in the study area as the training data set. Example pictures of data sets at different levels are as Figure 4 shown. Through the method of manual judgment and manual annotation, the labeling tool Labelme was used to label each target one by one. Since the main purpose is to detect and identify linear-distributed building groups, only the building groups showing a linear distribution in the training data set are labeled, and other categories are not labeled.
[0070] S102: Improve the YOLO11 model to obtain the YOLO11-DSC model.
[0071] S103: Train and optimize the YOLO11-DSC model according to the training data set to obtain the optimal YOLO11-DSC model.
[0072] S104: Input the target linear building group data into the optimal YOLO11-DSC model to obtain the recognition result of the linear building group.
[0073] The present invention constructs a training data set containing three levels based on the collected data of linear building groups. Then, by integrating dynamic snake-shaped convolution into the YOLO11 model, the YOLO11-DSC model is obtained. And according to the constructed training data set, the YOLO11-DSC model is optimized to obtain the optimal YOLO11-DSC model, which enhances the detection performance of the model, can increase the number of detected targets, and improve the recall rate. According to the optimal YOLO11-DSC model, the detection of the target linear building group data can be completed. Compared with the existing methods, the present invention has improved the recognition accuracy of the linear building group and the recognition ability of linear building groups of different sizes.
[0074] Embodiment 2
[0075] On the basis of the above embodiment, the present invention provides the structure of the YOLO11-DSC model, which specifically includes:
[0076] The present invention uses the YOLOv11 model as the basic object detection model for linear building group recognition. YOLO11 is the latest version of the YOLO series for real-time object detection and can reach the forefront level in different tasks. Compared with the previous version (YOLOv8), YOLO11 adopts an improved backbone and neck architecture, which enhances the ability to extract features from images and improves the accuracy of object detection and the performance of complex tasks. Figure 5 The network structure of YOLO11 is shown. The model is mainly composed of three parts: the backbone network, the neck network, and the head network. Compared with YOLOv8, the main improvements and innovations of YOLO11 are:
[0077] The C3k2 mechanism is proposed: on the basis of YOLOv8, the C2f module is replaced by the C3k2 module. The C3k2 module has two structures. When the parameter is set to False, it is the original C2f module, and its Bottleneck is an ordinary Bottleneck; on the contrary, when it is set to True, the Bottleneck module is replaced by the C3k module. The difference between the C3k module and the C3 module lies in the number of Bottlenecks. C3k has 2 Bottlenecks, and C3 defaults to 1, and changes according to the model scaling factor and the number of repetitions.
[0078] Proposed C2PSA mechanism: YOLO11 adds a C2PSA module after the SPPF layer. The C2PSA module combines the PSA (Pointwise Spatial Attention) Block to enhance feature extraction and the attention mechanism. By introducing the PSA Block into the standard C2f module, C2PSA achieves a more powerful attention mechanism, thereby further improving the model's ability to capture important features.
[0079] Improved detection head: Compared with YOLOv8, YOLO11 adds two depthwise convolution structures DWConv (Depthwise Convolution) to the classification detection head in the decoupled head, as Figure 6 shown. DWConv is a commonly used efficient convolution operation. Different from standard convolution, DWConv processes each channel of the input separately, that is, each channel has a separate convolution kernel for convolution operation and does not interact with other channels. Therefore, the number of parameters is greatly reduced, and since DWConv only processes the convolution in the spatial dimension and no longer processes the convolution between channels, the computational amount can be significantly reduced.
[0080] However, there will be some problems and deficiencies when the original YOLO11 model is used to detect linear building group data. The problems and deficiencies are as follows: For linearly distributed building groups, especially some groups composed of small-area buildings, the targets to be detected account for a relatively small proportion of the entire map sheet, while the receptive field of the original YOLO11 model is much larger than the detected targets. Therefore, it is easy to cause the targets to be easily ignored or regarded as the background. Therefore, it is necessary to improve the original YOLO11 model according to the long-strip distribution characteristics of linear buildings to improve the detection effect of linear building groups.
[0081] Standard convolution has a fixed receptive field ( Figure 7 (a)), which will bring certain difficulties to detection when facing a dataset with diverse features, especially tubular structure targets with slender features. In response to this problem of standard convolution, although some new convolution strategies such as dilated convolution and deformable convolution have been proposed, however, operations such as dilated convolution cannot adjust the attention area according to the features of tubular structure targets ( Figure 7 (b)), and although deformable convolution can adaptively learn the area of interest according to the features, for tubular targets, deformable convolution cannot limit the connectivity of the attention area ( Figure 7(c)). Therefore, some scholars have proposed a dynamic snake-shaped convolution, whose core idea is to use a convolution kernel that can dynamically change its shape to enhance the feature perception ability. This convolution method can focus on the slender and tortuous local structures in the tubular structure by adaptively changing the shape of the convolution kernel during the feature learning process, accurately capture the features of the tubular structure, and effectively improve the detection and segmentation effects of targets with tubular structures ( Figure 7 (d)).
[0082] On the one hand, the dynamic snake-shaped convolution can freely fit the tubular structure to learn features, and on the other hand, it can not deviate too far from the target structure under constraints, that is, the convolution kernel can dynamically change its shape like a snake. In order to make the size of the receptive field more suitable for the detection target, the dynamic snake-shaped convolution introduces a deformation offset, and realizes the learning of the deformation offset through an iterative strategy, greatly improving the flexibility of the convolution kernel in the two-dimensional convolution operation. In order to make the receptive field size more suitable for the target, the dynamic snake-shaped convolution adopts an iterative strategy to ensure that only one target is processed at a time and maintain the continuity of the detection results. The specific principle of the dynamic snake-shaped convolution is as follows:
[0083] First, for a standard two-dimensional convolution (N×N) with coordinates K, its center coordinates are denoted as K i =(x i ,y i ), i∈[0,N]. For a 3×3 convolution kernel with a dilation rate of 1, its coordinates are as follows:
[0084] K ={(x - 1,y - 1),(x - 1,y),(x - 1,y + 1),(x,y - 1),(x,y),(x,y + 1),(x + 1,y - 1),(x
[0085] + 1,y),(x + 1,y + 1)}
[0086] where K is the standard two-dimensional convolution coordinate, and x and y are the x-axis coordinate and y-axis coordinate of the standard two-dimensional convolution coordinate respectively.
[0087] To make the convolution kernel more flexible to focus on the geometric features of complex targets, the dynamic snake-shaped convolution introduces a deformation offset Δ. However, if the model freely learns the deformation offset, its perception field may deviate from the target when processing slender tubular structure targets. Therefore, in order to control the deviation, an iterative strategy is adopted for observation to ensure the continuity of the perception field and not spread the perception range too far due to large deformation deviations.
[0088] The dynamic snake-shaped convolution makes a linear adjustment to the convolution kernel in the x-axis and y-axis directions. Taking the x-axis direction as an example for a convolution kernel of size 9, the specific position of each grid in K can be expressed as: K i±c =(xi±c , y i±c ), where c is the horizontal distance from the center, and the value range is (0, 4). Each grid position K in the convolution kernel K i±c The selection is an accumulative process. Starting from the center K i , the positions away from the center grid depend on the previous grid position, that is, K i+1 relative to K i has an offset Δ added. Therefore, the offset needs to be accumulated to ensure that the convolution kernel conforms to a linear morphological structure.
[0089] The change of the dynamic snake convolution in the x-axis direction is expressed by the following formula:
[0090]
[0091] where, K i±c is the coordinate offset range of the dynamic snake convolution in the x-axis direction, x i+c and x i-c are the offset ranges in the x-axis direction respectively, y i+c and y i-c are the offset ranges in the y-axis direction respectively, c is the horizontal distance from the center coordinate, Δy is the offset in the x-axis direction, x i and y i are the x-axis coordinate and y-axis coordinate of the i-th center coordinate respectively.
[0092] The change of the dynamic snake convolution in the x-axis direction is expressed by the following formula:
[0093]
[0094] where, K j±c is the coordinate offset range of the dynamic snake convolution in the y-axis direction, and Δx is the offset in the y-axis direction.
[0095] Considering that the offset Δ is usually a decimal, the dynamic snake convolution uses bilinear interpolation for rounding, which is specifically expressed by the following formula:
[0096] K = ∑ K′ B(K′, K) · K′
[0097] where, K is the numerical value at the decimal position of the coordinates of K i±c and K j±c in the convolution grid, K′ is the numerical value at all integer positions of the original convolution kernel, and B is the bilinear interpolation kernel.
[0098] The bilinear interpolation kernel can be decomposed into two one-dimensional kernels, which are specifically shown by the following formula:
[0099] B(K, K′) = b(K x, K′ x )·b(K y , K′ y )
[0100] where b is a one-dimensional kernel, and K x and K y are respectively the fractional part values of the x-axis and y-axis coordinates of K i±c and K j±c in the convolutional grid, and K′ x and K′ y are respectively the integer part values on the x-axis and y-axis of the original convolutional kernel.
[0101] Due to the changes in the x and y directions, the dynamic snake-shaped convolutional kernel covers a 9×9 receptive field selectable range during the deformation process, can better adapt to slender tubular structures based on the dynamic structure, and can better perceive key features. Figure 8 It reflects the calculation process of the convolutional kernel coordinates and the corresponding receptive field in the dynamic snake-shaped convolution.
[0102] To improve the detection accuracy of linear building groups, the present invention proposes a YOLO11 network model integrating dynamic snake-shaped convolution (DySnakeConv, DSC). By integrating the dynamic snake-shaped convolution with the C3k2 module in YOLO11, the C3k2_DySnakeConv module is obtained, and the original C3k2 module in the backbone network is replaced with this module, and finally the YOLO11 network structure integrating DSC is obtained ( Figure 9 ).
[0103] In the C3k2_DySnakeConv module, specifically, DySnakeConv is integrated into the original Bottleneck module in the C3k2 module, and the original standard convolution in the original Bottleneck module is replaced with DySnakeConv to form the Bottleneck_DySnakeConv module, and then the original Bottleneck module in the original C3K2 model is replaced with the Bottleneck_DySnakeConv module to form the C3K2_DySnakeConv module, as Figure 10 shown.
[0104] The C3k2_DySnakeConv module is divided into two structures, specifically as follows:
[0105] When the parameter is set to False, the Bottleneck_DySnakeConv module is used to replace the Bottleneck module in the original C3k2 module. The Bottleneck_DySnakeConv module includes two dynamic snake-shaped convolutions and a concat function. Among them, the two dynamic snake-shaped convolutions and the concat function are connected in sequence, and the input data of the first dynamic snake-shaped convolution is also input into the concat function.
[0106] When the parameter is set to True, the Bottleneck_DySnakeConv module is used to replace the Bottleneck module in the original C3k module. With the help of the adaptive focusing ability of DySnakeConv for slender and curved structures, it can effectively improve the sensitivity of the network to linear targets, enhance the detection ability for linear building groups of different sizes and shapes, and improve the model detection accuracy.
[0107] Standard convolution works well when dealing with regular data, but there will be certain limitations in the case of large changes in the target shape. The advantage of dynamic snake-shaped convolution compared to standard convolution is that by introducing a dynamic and deformable convolution kernel, it can adaptively adjust the shape and size of the convolution kernel, capture target features more effectively, and better adapt to the specific shape of the target. The specific process is as Figure 11 shown.
[0108] Embodiment 3
[0109] Based on the above embodiments, the present invention provides a verification process for the YOLO11-DSC model, which specifically includes:
[0110] The experimental operating system of the present invention is the Windows11 64-bit operating system, with the CPU version being Inter(R) Core(TM) i9-14900HX CPU@2.20GHz, 32GB RAM, and the GPU being NVIDIA GeForce RTX 4090 laptop GPU.
[0111] To verify the practicability and effectiveness of the present invention, the original YOLO11 model and two improved models YOLO11-CBAM (CBAM attention mechanism) and YOLO11-AKConv (changeable convolution kernel) are selected for comparison in the test dataset. Table 1 shows the evaluation index results of the present invention and the above three methods. It can be seen that the present invention is the best in both accuracy and recall. Figures 12 - 14 Figure 4 shows the recognition result examples of the four methods in different test datasets. It can be seen that the present invention can identify more linear building groups and has higher recognition accuracy.
[0112] Table 1 Overall evaluation index results of the test data sets of different YOLO models
[0113]
[0114] Example 4
[0115] Based on the above embodiments, as Figure 15 shown, the embodiment of the present invention provides a linear building group recognition system integrating dynamic snake-shaped convolution and YOLO11, including:
[0116] A collection module, configured to collect linear building group data and construct a training data set according to the collected linear building group data.
[0117] A construction module, configured to improve the YOLO11 model to obtain a YOLO11-DSC model.
[0118] A training module, configured to train and optimize the YOLO11-DSC model according to the training data set to obtain an optimal YOLO11-DSC model.
[0119] An identification module, configured to input the target linear building group data into the optimal YOLO11-DSC model to obtain a linear building group recognition result.
[0120] It should be noted that the linear building group recognition system integrating dynamic snake-shaped convolution and YOLO11 provided by the embodiment of the present invention is to implement the above-mentioned method for recognizing linear building groups integrating dynamic snake-shaped convolution and YOLO11. For its functions, reference can be specifically made to the above-mentioned method embodiments, and details are not described herein again.
[0121] Example 5
[0122] Based on the above embodiments, the embodiment of the present invention further provides a terminal device, including a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor. When the processor executes the computer program, it implements the method for recognizing linear building groups integrating dynamic snake-shaped convolution and YOLO11 in the above embodiments.
[0123] The present invention also provides a computer-readable storage medium, where the storage medium includes a stored computer program. When the computer program runs, it controls the device where the storage medium is located to execute the method for recognizing linear building groups integrating dynamic snake-shaped convolution and YOLO11 in the above embodiments.
[0124] In summary, the present invention introduces a Dynamic Snake Convolution (DSC) module into the backbone network of the YOLO11 object detection model, fuses the DSC module with the C3k2 module in YOLO11 to obtain the C3k2_DySnakeConv module, and uses this module to replace the original C3k2 modules in the backbone network and the neck network, finally obtaining the YOLO11 network structure YOLO11-DSC model integrated with DSC. By integrating the DSC module, the present invention can increase the number of detected objects and improve the recall rate. The recognition accuracy of the linear building group of the YOLO11-DSC model has been improved compared with the prior art, and the recognition ability for linear building groups of different sizes has been enhanced, with stronger detection performance.
[0125] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A linear building group recognition method integrating dynamic snake convolution and YOLO11, characterized in that: include: Step 1: Collect linear building group data, and construct a training data set based on the collected linear building group data; Step 2: Improve the YOLO11 model to obtain the YOLO11-DSC model; Step 3: Train and optimize the YOLO11-DSC model according to the training data set to obtain the optimal YOLO11-DSC model; Step 4: Input the target linear building group data into the optimal YOLO11-DSC model to obtain the linear building group recognition result.
2. The linear building group recognition method integrating dynamic snake convolution and YOLO11 according to claim 1, characterized in that: The training data set includes individual building group training data, surrounding building interference training data and complex scene training data; The separate building group training data includes a selected area containing only one linear building group, and no other buildings exist around the group; The surrounding building interference training data includes interference of one or more linear building groups in the selected area and buildings that do not form a linear pattern in the surrounding area; The complex scene training data includes that the scope and the total number of buildings involved in the selected area are larger than the scope and the total number of buildings in the surrounding building interference training data, the total number of buildings includes linear building groups and buildings that cause interference, and the number of linear building groups and buildings that cause interference is larger than the linear building groups and buildings that cause interference in the surrounding building interference training data.
3. The linear building group recognition method integrating dynamic snake convolution and YOLO11 according to claim 1, characterized in that: The YOLO11 model is improved to obtain the YOLO11-DSC model, which specifically includes: The dynamic snake convolution is fused with the C3k2 module in the YOLO11 model to obtain the C3k2_DySnakeConv module, and this module is used to replace the original C3k2 modules in the backbone network and neck network in the YOLO11 model, and finally the YOLO11-DSC model is obtained.
4. The linear building group recognition method integrating dynamic snake convolution and YOLO11 according to claim 3, characterized in that: The change of the dynamic snake convolution in the x-axis direction is expressed by the following formula: Among them, K i±c is the coordinate offset range of the dynamic snake convolution coordinates in the x-axis direction, x i+c and x i-c are the offset ranges in the x-axis direction, y i+c and i-c are the offset range in the y-axis direction, c is the horizontal distance from the center coordinate, Δy is the offset in the x-axis direction, and x i and i are the x-axis coordinates and y-axis coordinates of the i-th center coordinate respectively; The change of dynamic snake convolution in the x-axis direction is expressed by the following formula: Among them, K j±c is the coordinate offset range of the dynamic snake convolution coordinate in the y-axis direction, and Δx is the offset in the y-axis direction.
5. The linear building group recognition method integrating dynamic snake convolution and YOLO11 according to claim 4, characterized in that: The decimal position value of the dynamic serpentine convolution coordinate is calculated according to the following formula: K=∑ K′ B(K′,K)·K′ B(K,K′)=b(K x ,K′ x )·b(K y ,K′ y ) Where K is K i±c and K j±c The decimal position value of the coordinates in the convolution grid, K′ is the integer position value of all the original convolution kernels, B is the bilinear interpolation kernel, b is a one-dimensional kernel, K x and K y K i±c and K j±c The decimal position of the x- and y-coordinates in the convolution grid, K′ x and K′ y They are the integer position values on the x-axis and y-axis of the original convolution kernel respectively.
6. The linear building group recognition method integrating dynamic snake convolution and YOLO11 according to claim 3, characterized in that: The C3k2_DySnakeConv module is divided into two structures, as follows: When the parameter is set to False, the Bottleneck_DySnakeConv module is used to replace the Bottleneck module in the original C3k2 module; the Bottleneck_DySnakeConv module includes two dynamic snake convolutions and a concat function; wherein the two dynamic snake convolutions and the concat function are connected in sequence, and the input data of the first dynamic snake convolution is also input into the concat function; When the parameter is set to True, the Bottleneck_DySnakeConv module is used to replace the Bottleneck module in the original C3k module.
7. A linear building group recognition system integrating dynamic snake convolution and YOLO11, characterized by: include: A collection module, used for collecting linear building group data, and constructing a training data set according to the collected linear building group data; A construction module is used to improve the YOLO11 model to obtain the YOLO11-DSC model; The training module is used to train and optimize the YOLO11-DSC model according to the training data set to obtain the optimal YOLO11-DSC model; The recognition module is used to input the target linear building group data into the optimal YOLO11-DSC model to obtain the linear building group recognition result.
8. An electronic device, characterized in that: It comprises a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, wherein when the processor executes the computer program, the linear building group recognition method integrating dynamic snake convolution and YOLO11 as described in any one of claims 1 to 6 is implemented.
9. A computer-readable storage medium, characterized in that: The storage medium includes a stored computer program, wherein when the computer program is running, the device where the storage medium is located is controlled to execute the linear building group recognition method integrating dynamic snake convolution and YOLO11 as described in any one of claims 1 to 6.
Citation Information
Cited By
Small target detection and identification method and system based on improved YOLOv8
CN120451518A
Rapid identification method, system and equipment for direct drainage of pond tail water
CN121354039A
A method, system and device for quickly identifying a pond tail water direct discharge behavior
CN121354039B