Smart city target detection method, device and storage medium based on improved YOLOv5
By combining the large vision model SAM with the improved YOLOv5 model, using the method of parallel connection and neck fusion structure, the problem of insufficient category judgment in complex scenarios is solved, and higher detection accuracy and adaptability are achieved.
Patent Information
- Application Number
- CN202510193272.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-21
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2045-02-21
AI Technical Summary
The existing smart city target detection model lacks adaptability in category judgments in complex scenarios, resulting in a decline in detection performance and the inability to accurately provide target detection classification information, affecting the effective management of smart cities.
The improved YOLOv5 model is adopted to integrate the large vision model SAM into the object detection model. By connecting the YOLOv5 backbone structure and the SAM backbone structure in parallel, the two features are fused using the neck fusion structure to generate a fusion feature map for object detection.
It enhances the full extraction of effective target features, improves the detection accuracy of the model, significantly improves the category judgment ability in complex scenarios, and improves the performance of smart city target detection.
Smart Images

Figure CN119672550B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of smart city target detection, and in particular to a smart city target detection method, device and storage medium based on improved YOLOv5. Background Art
[0002] Smart city cloud AI detection systems usually use the YOLO series of models based on convolutional neural networks (CNN) or the DETR series of models based on Transformer to achieve target detection; they mainly rely on a large number of data samples to train a single or multiple target detection models that meet the characteristics of urban objects to determine the object categories of smart cities, and then implement urban supervision or event judgment based on the object categories. However, since the same object in a smart city usually appears in different categories in different scenes, and is affected by factors such as interference from the complex urban environment, differences in target attributes or data distribution, the existing target detection models lack adaptability to category judgment in complex scenes in smart cities.
[0003] To address the above issues, existing technologies generally use a combination of multiple models for judgment. Although this has improved the performance of urban target detection to a certain extent, it has not fundamentally solved the problem of insufficient feature extraction of target images by the detection model, resulting in a decline in model detection performance and an inability to accurately provide target detection classification information, affecting the effective management of smart cities. Summary of the invention
[0004] The purpose of the present invention is to overcome the shortcomings of the prior art, provide a smart city target detection method, device and storage medium based on the improved YOLOv5, integrate the large visual model SAM with powerful generalization ability into the target detection model YOLOv5, which can enhance the full extraction of effective features of the target and improve the detection accuracy of the model.
[0005] To achieve the above object, the present invention is implemented by adopting the following technical solutions:
[0006] In a first aspect, the present invention provides a smart city target detection method based on improved YOLOv5, the method comprising:
[0007] Collect images of smart city scenes to be detected;
[0008] Input the scene image into a pre-trained SAM-YOLOv5 network model to obtain a target detection result;
[0009] The SAM-YOLOv5 network model includes a backbone network, a neck network and a detection head network connected in sequence, the backbone network includes a YOLOv5 backbone structure and a SAM backbone structure connected in parallel, the neck network includes a YOLOv5 neck structure and a SAM neck structure connected in parallel, and a neck fusion structure connecting the YOLOv5 neck structure and the SAM neck structure;
[0010] The processing process of the SAM-YOLOv5 network model on the scene image includes:
[0011] Input the scene image into the YOLOv5 backbone structure and the SAM backbone structure respectively for feature extraction, and obtain a YOLOv5 backbone feature map and a SAM backbone feature map respectively;
[0012] The YOLOv5 neck structure is used to sample and fuse the YOLOv5 trunk feature map to generate the YOLOv5 neck feature map. At the same time, the SAM neck structure is used to extract and transform the SAM trunk feature map to generate the SAM neck feature map.
[0013] The neck fusion structure is used to fuse the YOLOv5 neck feature map with the SAM neck feature map to generate a fused feature map;
[0014] The fused feature map is input into the detection head network for convolution operation to predict the target category and position, thereby obtaining the target detection result.
[0015] In combination with the first aspect, further, the YOLOv5 backbone structure includes a first convolution module, a second convolution module, a first feature enhancement module, a third convolution module, a second feature enhancement module, a fourth convolution module, a third feature enhancement module, a fifth convolution module, a fourth feature enhancement module and a spatial pyramid pooling module which are sequentially connected in series;
[0016] The method of inputting the scene image into a YOLOv5 backbone structure for feature extraction to obtain a YOLOv5 backbone feature map includes:
[0017] The scene image is sequentially passed through a first convolution module, a second convolution module, and a first feature enhancement module to capture low-level feature information, and then sequentially passed through a third convolution module and a second feature enhancement module to perform a downsampling operation, and output a YOLOv5 backbone low-level feature map;
[0018] The low-level feature map of the YOLOv5 backbone is downsampled through the fourth convolution module and the third feature enhancement module in turn, and the middle-level feature map of the YOLOv5 backbone is output;
[0019] The middle-level feature map of the YOLOv5 backbone is processed by the fifth convolution module, the fourth feature enhancement module and the spatial pyramid pooling module in sequence, and the high-level feature map of the YOLOv5 backbone is output.
[0020] In combination with the first aspect, further, the SAM backbone structure includes a sixth convolution module, a mapping transposition module, a position encoding module, a first fusion module and a visual encoding module;
[0021] The method of inputting the scene image into the SAM backbone structure for feature extraction to obtain a SAM backbone feature map comprises:
[0022] Inputting the scene image into a sixth convolution module for block division and convolution to obtain an image block feature map of a fixed size;
[0023] A mapping transposition module is used to flatten and sort the fixed-size image block feature map to obtain a flattened image block feature map;
[0024] The flattened image block feature map and the position feature map generated by the position encoding module are directly added together through a first fusion module to perform feature fusion, thereby obtaining a preprocessed feature map;
[0025] The preprocessed feature map is input into the visual encoding module for processing, and feature extraction is performed in the ratio of 1:1:3:1. The SAM backbone low-level feature map, SAM backbone middle-level feature map and SAM backbone high-level feature map are output in sequence according to the latter three ratios.
[0026] In combination with the first aspect, further, the YOLOv5 neck structure includes a first neck branch structure, a second neck branch structure, a third neck branch structure and a fourth neck branch structure;
[0027] Among them, the first neck branch structure includes the seventh convolution module, the first upsampling module, the second fusion module, the fifth feature enhancement module, the eighth convolution module, the second upsampling module and the third fusion module; the second neck branch structure includes the sixth feature enhancement module, the ninth convolution module and the fourth fusion module; the third neck branch structure includes the seventh feature enhancement module, the tenth convolution module and the fifth fusion module; the fourth neck branch structure includes the eighth feature enhancement module;
[0028] The method of sampling and fusing a YOLOv5 trunk feature map by using a YOLOv5 neck structure to generate a YOLOv5 neck feature map includes:
[0029] The YOLOv5 trunk high-level feature map is refined through the seventh convolution module to generate the YOLOv5 neck high-level intermediate feature map;
[0030] After the YOLOv5 neck high-level intermediate feature map is upsampled by the first upsampling module, it is fused with the YOLOv5 trunk middle-level feature map by the second fusion module, and then processed by the fifth feature enhancement module and the eighth convolution module in sequence to generate the YOLOv5 neck middle-level intermediate feature map;
[0031] After the YOLOv5 neck middle-level intermediate feature map is upsampled by the second upsampling module, it is fused with the YOLOv5 trunk low-level feature map by the third fusion module to output the YOLOv5 neck low-level feature map;
[0032] The YOLOv5 neck low-level feature map is fused through the neck fusion structure to obtain a fused low-level intermediate feature map, and the fused low-level intermediate feature map is enhanced by the sixth feature enhancement module to output the fused low-level feature map, which is then downsampled by the ninth convolution module and then fused with the YOLOv5 neck middle-level intermediate feature map by the fourth fusion module to output the YOLOv5 neck middle-level feature map;
[0033] The YOLOv5 neck middle-level feature map is fused through the neck fusion structure to obtain a fused middle-level intermediate feature map, and the seventh feature enhancement module is used to enhance the fused middle-level intermediate feature map, and the fused middle-level feature map is output. After downsampling operation is performed by the tenth convolution module, the feature fusion is performed with the YOLOv5 neck high-level intermediate feature map through the fifth fusion module, and the YOLOv5 neck high-level feature map is output;
[0034] The YOLOv5 neck high-level feature map is fused through the neck fusion structure to obtain a fused high-level intermediate feature map. The eighth feature enhancement module is used to enhance the fused high-level intermediate feature map and output the fused high-level feature map.
[0035] In combination with the first aspect, further, the SAM neck structure is a fixed-scale feature pyramid module, including a fourteenth convolution module, a fifteenth convolution module, a sixteenth convolution module, a sixth fusion module, a seventeenth convolution module, an eighteenth convolution module, a seventh fusion module and a nineteenth convolution module;
[0036] The method for extracting and converting a SAM trunk feature map by using a SAM neck structure to generate a SAM neck feature map includes:
[0037] The high-level feature map of the SAM trunk is converted into a spatial state by the fourteenth convolution module to obtain a high-level intermediate feature map of the SAM neck, which is then extracted and fused by the fifteenth convolution module to output a high-level feature map of the SAM neck;
[0038] After the spatial state of the middle-level feature map of the SAM trunk is converted by the sixteenth convolution module, it is directly added and integrated with the high-level intermediate feature map of the SAM neck by the sixth fusion module to obtain the middle-level intermediate feature map of the SAM neck, and then the feature is extracted and fused by the seventeenth convolution module to output the middle-level feature map of the SAM neck;
[0039] After the spatial state of the SAM trunk low-level feature map is converted by the eighteenth convolution module, it is directly added to the SAM neck middle-level intermediate feature map by the seventh fusion module for integration, and then feature extraction and fusion are performed by the nineteenth convolution module to output the SAM neck low-level feature map.
[0040] In combination with the first aspect, further, the neck fusion structure is three parallel large visual feature fusion modules, including a first large visual feature fusion module, a second large visual feature fusion module and a third large visual feature fusion module, and the first large visual feature fusion module, the second large visual feature fusion module and the third large visual feature fusion module all include a branch feature processing structure, a weight branch processing structure and a feature fusion calculation structure;
[0041] Among them, the branch feature processing structure includes a parallel YOLOv5 branch feature processing structure and a SAM branch feature processing structure, the YOLOv5 branch feature processing structure includes a twentieth convolution module, and the SAM branch feature processing structure includes a twenty-first convolution module and a twenty-second convolution module in series; the weight branch processing structure includes an eighth fusion module, an average pooling layer, a twenty-third convolution module, a third upsampling module, a ninth fusion module and a Sigmoid activation function; the feature fusion calculation structure includes a tenth fusion module, an eleventh fusion module, a twelfth fusion module, a twenty-fourth convolution module and a thirteenth fusion module;
[0042] The method of fusing the YOLOv5 neck feature map with the SAM neck feature map by using the neck fusion structure to generate a fused feature map includes:
[0043] The YOLOv5 neck feature map is processed by the 20th convolution module to obtain a preprocessed YOLOv5 branch feature map, and the SAM neck feature map is processed by the 21st convolution module to generate a SAM branch intermediate feature map, which is then extracted by the 22nd convolution module to obtain a preprocessed SAM branch feature map;
[0044] The YOLOv5 neck feature map and the SAM branch intermediate feature map are directly added and fused through the eighth fusion module to obtain a preliminary fusion feature map, which is then processed in turn through the average pooling layer, the twenty-third convolution module and the third upsampling module, and then residually connected with the preliminary fusion feature map through the ninth fusion module to generate weight features;
[0045] Input the weight feature into the Sigmoid activation function for conversion to generate branch weights;
[0046] The pre-processed YOLOv5 branch feature map and the pre-processed SAM branch feature map are directly multiplied by the tenth fusion module and the eleventh fusion module according to the branch weights for fusion, and the two fused branch feature maps are directly added and fine-tuned by the twelfth fusion module to generate a basic fusion feature map;
[0047] After fine-tuning the basic fusion feature map through the twenty-fourth convolution module, the basic fusion feature maps before and after fine-tuning are residually connected through the thirteenth fusion module to output the fusion intermediate feature map;
[0048] After the fused intermediate feature map is enhanced by the feature enhancement module in the YOLOv5 neck structure, the fused feature map is output.
[0049] In combination with the first aspect, further, a channel attention mechanism is introduced into the weight branch processing structure, which is used to dynamically adjust the weight features through calculation to enhance the attention of key features;
[0050] Among them, the expression of the channel attention mechanism is:
[0051]
[0052] in, represents the Sigmoid activation function, and represents the linear convolution kernel, and Represent average pooling and maximum pooling respectively, X Represents the weight feature of the input.
[0053] In combination with the first aspect, further, the detection head network includes an eleventh convolution module, a twelfth convolution module and a thirteenth convolution module connected in parallel;
[0054] The detection head network inputs the fused low-level feature map, the fused mid-level feature map and the fused high-level feature map into the eleventh convolution module, the twelfth convolution module and the thirteenth convolution module respectively for convolution operations to generate target boxes and confidence scores to predict the target category and position, thereby obtaining target detection results.
[0055] In a second aspect, the present invention further provides a computer device, including a storage medium and a processor;
[0056] The storage medium is used to store instructions;
[0057] The processor is used to operate according to the instructions to execute the steps of any method described in the first aspect.
[0058] In a third aspect, the present invention further provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of any one of the methods described in the first aspect.
[0059] Compared with the prior art, the present invention has the following beneficial effects:
[0060] The technical solution provided by the present invention is to detect the target using the SAM-YOLOv5 network model, and integrate the large visual model SAM with strong generalization ability into the target detection model YOLOv5, which can enhance the full extraction of effective features of the target and improve the detection accuracy of the model. Among them, the SAM-YOLOv5 network model is based on the YOLOv5 model to introduce a fixed-scale feature pyramid module to fully extract rich features and enhance the further optimization of the SAM model on the image; the large visual feature fusion module is introduced to seamlessly integrate the features extracted by the SAM model into the YOLOv5 model, significantly improving the feature representation, and improving the expressiveness and generalization performance of the model. BRIEF DESCRIPTION OF THE DRAWINGS
[0061] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0062] Figure 1 It is a flow chart of a smart city target detection method based on improved YOLOv5 provided by an embodiment of the present invention;
[0063] Figure 2 It is a schematic diagram of the network structure of a SAM-YOLOv5 network model provided by an embodiment of the present invention;
[0064] Figure 3 yes Figure 2 A schematic diagram of the structure of a convolution module, a feature enhancement module and a spatial pyramid pooling module;
[0065] Figure 4 yes Figure 2 A schematic diagram of the structure of a visual encoding module;
[0066] Figure 5 yes Figure 2 A structural diagram of a fixed-scale feature pyramid module;
[0067] Figure 6 yes Figure 2 A schematic diagram of the structure of a large visual feature fusion module;
[0068] Figure 7 It is a schematic diagram of the network structure of another SAM-YOLOv5 network model provided by an embodiment of the present invention;
[0069] Figure 8 It is a thermal map generated by actual detection of an application example of the present invention;
[0070] Fig. 9 It is an internal structure diagram of a computer device provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0071] The present invention will be further described below in conjunction with the accompanying drawings. The following embodiments are only used to more clearly illustrate the technical solution of the present invention, and cannot be used to limit the protection scope of the present invention.
[0072] Embodiment 1:
[0073] This embodiment provides a smart city target detection method based on improved YOLOv5. Figure 1 FIG. 1 is a flow chart of the method provided in this embodiment, which includes the following steps:
[0074] Step S1: collecting the smart city scene image to be detected;
[0075] Step S2: Input the scene image into a pre-trained SAM-YOLOv5 network model to obtain target detection results.
[0076] Specifically, Figure 2 As shown, it is a schematic diagram of the network structure of the SAM-YOLOv5 network model provided in this embodiment, including a backbone network, a neck network and a detection head network connected in sequence; wherein, the backbone network includes a YOLOv5 backbone structure and a SAM backbone structure connected in parallel, and the neck network includes a YOLOv5 neck structure and a SAM neck structure connected in parallel and a neck fusion structure connecting the YOLOv5 neck structure and the SAM neck structure.
[0077] Furthermore, the processing of the scene image by the SAM-YOLOv5 network model in step S2 includes the following steps:
[0078] Step S201: inputting the scene image into the YOLOv5 backbone structure and the SAM backbone structure respectively for feature extraction, and obtaining a YOLOv5 backbone feature map and a SAM backbone feature map respectively;
[0079] Step S202: using the YOLOv5 neck structure to sample and fuse the YOLOv5 trunk feature map to generate a YOLOv5 neck feature map, and using the SAM neck structure to extract and transform the SAM trunk feature map to generate a SAM neck feature map;
[0080] Step S203: using the neck fusion structure to fuse the YOLOv5 neck feature map with the SAM neck feature map to generate a fused feature map;
[0081] Step S204: Input the fused feature map into the detection head network for convolution operation to predict the target category and position, thereby obtaining the target detection result.
[0082] It should be noted that the YOLOv5 model is a target detection model in the YOLO series of methods; the large visual model SAM (Segment Anything Model) has zero-sample capability and powerful feature extraction capabilities, and can recognize any object based on text instructions or image recognition. The SAM-YOLOv5 network model provided in this embodiment is an improvement based on the network structure of the YOLOv5 model: the large visual model SAM with powerful generalization ability is integrated into the target detection model YOLOv5, and the performance of target detection is improved from the perspective of the visual basic model, so as to enhance the full extraction of effective target features and improve the detection accuracy of the model.
[0083] refer to Figure 2 ,The left most box in the figure provides a network structure of the SAM-YOLOv5 model backbone network,a dual-branch encoding structure, including a YOLOv5 backbone structure and a SAM backbone structure connected in parallel.
[0084] Specifically, the YOLOv5 backbone structure includes a first convolution module, a second convolution module, a first feature enhancement module, a third convolution module, a second feature enhancement module, a fourth convolution module, a third feature enhancement module, a fifth convolution module, a fourth feature enhancement module and a spatial pyramid pooling module connected in series in sequence; wherein the first convolution module, the second convolution module, the third convolution module, the fourth convolution module and the fifth convolution module are used to implement a downsampling operation to reduce the scale of the feature map; the first feature enhancement module, the second feature enhancement module, the third feature enhancement module and the fourth feature enhancement module are used to perform feature enhancement and extraction on the feature map; the spatial pyramid pooling module is used to improve the feature extraction capability of the model.
[0085] Further, such as Figure 3As shown, it is a structural schematic diagram of the convolution module, feature enhancement module and spatial pyramid pooling module provided in this embodiment. The convolution module includes a two-dimensional convolution layer, a batch normalization layer and a SiLU activation function connected in series. The feature enhancement module includes a twenty-fifth convolution module, N series-connected bottleneck modules, a fourteenth fusion module, a twenty-sixth convolution module and a twenty-seventh convolution module; wherein, the twenty-fifth convolution module is connected in series with the N series-connected bottleneck modules and then connected in parallel with the twenty-seventh convolution module, the output end of the Nth bottleneck module and the output end of the twenty-seventh convolution module are connected to the input end of the fourteenth fusion module, and the output end of the fourteenth fusion module is connected to the input end of the twenty-sixth convolution module. The bottleneck module in the feature enhancement module in the YOLOv5 backbone structure uses a residual structure by default, that is, it includes a twenty-eighth convolution module, a twenty-ninth convolution module and a fifteenth fusion module; wherein, the twenty-eighth convolution module is connected in series with the twenty-ninth convolution module, and the input end of the twenty-eighth convolution module and the output end of the twenty-ninth convolution module are connected to the input end of the fifteenth fusion module for residual connection.
[0086] The spatial pyramid pooling module includes a thirty-second convolution module, a first maximum pooling layer, a second maximum pooling layer, a third maximum pooling layer, a sixteenth fusion module and a thirty-third convolution module; wherein the output end of the thirty-second convolution module is respectively connected to the input ends of the first maximum pooling layer, the second maximum pooling layer, and the third maximum pooling layer, the output ends of the first maximum pooling layer, the second maximum pooling layer, and the third maximum pooling layer and the output end of the thirty-second convolution module are commonly connected to the input end of the sixteenth fusion module, and the output end of the sixteenth fusion module is connected to the input end of the thirty-third convolution module.
[0087] Combine the following Figure 2 , the method of inputting the scene image into the YOLOv5 backbone structure for feature extraction and obtaining the YOLOv5 backbone feature map in step S201 is further described in detail:
[0088] The scene image is sequentially passed through the first convolution module, the second convolution module and the first feature enhancement module to capture low-level feature information, and then sequentially passed through the third convolution module and the second feature enhancement module for downsampling operations to output a YOLOv5 backbone low-level feature map ;
[0089] YOLOv5 backbone low-level feature map The fourth convolution module and the third feature enhancement module are downsampled in turn to output the middle-level feature map of the YOLOv5 backbone. ;
[0090] The hierarchical feature maps in the YOLOv5 backbone It is processed in turn by the fifth convolution module, the fourth feature enhancement module and the spatial pyramid pooling module to output the YOLOv5 backbone high-level feature map. , obtain YOLOv5 backbone feature maps of different scales .
[0091] It should be noted that, for ease of description, this embodiment describes the third layer feature map as a low-level feature map, the fourth layer feature map as a medium-level feature map, and the fifth layer feature map as a high-level feature map. In addition, all convolution modules in the embodiments of the present invention include a two-dimensional convolution layer, a batch normalization layer, and a SiLU activation function connected in series.
[0092] In this embodiment, in order to efficiently extract rich features from the input scene image and improve the generalization ability of the model, this embodiment introduces the SAM backbone structure into the backbone network of the YOLOv5 model to enhance the feature extraction capability.
[0093] like Figure 2 As shown, the SAM backbone structure includes a sixth convolution module, a mapping transposition module, a position encoding module, a first fusion module and a visual encoding module; wherein the output end of the sixth convolution module is connected to the input end of the mapping transposition module, the output end of the mapping transposition module and the output end of the position encoding module are commonly connected to the input end of the first fusion module, and the output end of the first fusion module is connected to the input end of the visual encoding module.
[0094] Specifically, the sixth convolution module is used to divide the feature map into blocks to obtain image blocks of fixed size; the mapping transposition module is used to flatten the image blocks of fixed size and linearly map them to a lower dimensional space; the position encoding module is used to generate position encoding feature information; and the visual encoding module is used to process the input image and perform feature extraction.
[0095] Further, such as Figure 4As shown, it is a structural diagram of the visual encoding module Vision Encoder provided in this embodiment, which is mainly composed of n series-connected transformer modules Transformer Block. The transformer module Transformer Block includes a first normalization layer, a multi-head attention mechanism Multi-Head Attention, a seventeenth fusion module, a second normalization layer, a feed-forward neural network Feed-Forward Network and an eighteenth fusion module; wherein, the output end of the first normalization layer is connected to multiple input ends of the multi-head attention mechanism, the output end of the multi-head attention mechanism and the input end of the first normalization layer are commonly connected to the input end of the seventeenth fusion module, the output end of the seventeenth fusion module is sequentially connected in series with the second normalization layer and the feed-forward neural network, and the output end of the feed-forward neural network and the output end of the seventeenth fusion module are commonly connected to the input end of the eighteenth fusion module.
[0096] The visual encoding module Vision Encoder mainly obtains strong generalization ability by extensive pre-training on the SA-1B dataset, which contains 1 billion masks extracted from 11 million images; then uses the Masked Autoencoder MAE (Masked Autoencoders) technology to pre-train the visual transformer ViT to process the input image; finally, the image is feature extracted in a ratio of 1:1:3:1.
[0097] It should be noted that if the Vision Encoder module is composed of 24 transformer blocks connected in series, the preprocessed image features will be subjected to 24 layers of feature extraction, and the output of the 4th layer is used as Features, the output of layer 8 is Features, the output of the 20th layer is Features, the output of the last layer is Features, take the feature maps of the last three ratio outputs as the SAM backbone feature maps.
[0098] Combine the following Figure 2 , the method of how to input the scene image into the SAM backbone structure for feature extraction and obtain the SAM backbone feature map in step S201 is further described in detail:
[0099] Inputting the scene image into a sixth convolution module for block division and convolution to obtain an image block feature map of a fixed size;
[0100] A mapping transposition module is used to flatten and sort the fixed-size image block feature map to obtain a flattened image block feature map;
[0101] The flattened image block feature map and the position feature map generated by the position encoding module are directly added together through a first fusion module to perform feature fusion, thereby obtaining a preprocessed feature map;
[0102] The preprocessed feature map is input into the visual encoding module for processing, and feature extraction is performed according to the ratio of 1:1:3:1. The low-level feature map of the SAM backbone is output in sequence from the last three ratios. , SAM backbone mid-level feature map And SAM backbone high-level feature map , obtain the SAM backbone feature map of the same scale .
[0103] In the backbone network with a dual-branch encoding structure proposed in this embodiment, the YOLOv5 backbone structure and the SAM backbone structure form a synergistic combination, combining YOLOv5's ability to extract specific features and SAM's powerful generalization ability, enriching the feature representation, thereby enabling the SAM-YOLOv5 network model to achieve excellent performance in complex visual tasks.
[0104] As an example, Figure 2 As shown, a network structure of the neck network is provided in the box in the middle of the SAM-YOLOv5 network model, including a YOLOv5 neck structure and a SAM neck structure connected in parallel and a neck fusion structure connecting the YOLOv5 neck structure and the SAM neck structure.
[0105] Specifically, the YOLOv5 neck structure includes a first neck branch structure, a second neck branch structure, a third neck branch structure and a fourth neck branch structure. The first neck branch structure includes a seventh convolution module, a first upsampling module, a second fusion module, a fifth feature enhancement module, an eighth convolution module, a second upsampling module and a third fusion module; wherein the input end of the seventh convolution module is connected to the output end of the spatial pyramid pooling module in the YOLOv5 trunk structure, the output end of the seventh convolution module is connected to the input end of the first upsampling module, the output end of the first upsampling module and the output end of the third feature enhancement module in the YOLOv5 trunk structure are connected to the input end of the second fusion module, the output end of the second fusion module is connected in series with the fifth feature enhancement module, the eighth convolution module and the second upsampling module, the output end of the second upsampling module and the output end of the second feature enhancement module in the YOLOv5 trunk structure are connected to the input end of the third fusion module, and the output end of the third fusion module is connected to the second input end of the neck fusion structure.
[0106] The second neck branch structure includes a sixth feature enhancement module, a ninth convolution module and a fourth fusion module; wherein the input end of the sixth feature enhancement module is connected to the first output end of the neck fusion structure, the output end of the sixth feature enhancement module is connected to the input end of the ninth convolution module, the output end of the ninth convolution module and the output end of the eighth convolution module in the first neck branch structure are commonly connected to the input end of the fourth fusion module, and the output end of the fourth fusion module is connected to the fourth input end of the neck fusion structure.
[0107] The third neck branch structure includes a seventh feature enhancement module, a tenth convolution module and a fifth fusion module; wherein the input end of the seventh feature enhancement module is connected to the second output end of the neck fusion structure, the output end of the seventh feature enhancement module is connected to the input end of the tenth convolution module, the output end of the tenth convolution module and the output end of the seventh convolution module in the first neck branch structure are commonly connected to the input end of the fifth fusion module, and the output end of the fifth fusion module is connected to the sixth input end of the neck fusion structure.
[0108] The fourth neck branch structure includes an eighth feature enhancement module, and the input end of the eighth feature enhancement module is connected to the third output end of the neck fusion structure.
[0109] Furthermore, the seventh convolution module and the eighth convolution module are used to refine the features of the feature map and reduce the model parameters; the fifth feature enhancement module is used to enhance the features of the feature map and reduce the model parameters; the sixth feature enhancement module, the seventh feature enhancement module and the eighth feature enhancement module are used to enhance and extract the features of the feature map; the ninth convolution module and the tenth convolution module are used to implement downsampling operations to reduce the scale of the feature map.
[0110] It should be noted that the reference Figure 3 Schematic diagram of the structure of the feature enhancement module in the YOLOv5 neck structure. The feature enhancement module in the YOLOv5 trunk structure has the same structure as the feature enhancement module in the YOLOv5 bottleneck structure. The difference is that the bottleneck module does not use the residual structure, that is, it includes the 30th convolution module and the 31st convolution module connected in series.
[0111] refer to Figure 2 In step S202, the YOLOv5 neck structure performs the following processing steps on the YOLOv5 trunk feature map:
[0112] First, the YOLOv5 backbone high-level feature map After feature refinement and conversion calculations are performed by the seventh convolution module, a YOLOv5 neck high-level intermediate feature map is generated;
[0113] The high-level intermediate feature map of the YOLOv5 neck is upsampled by the first upsampling module and compared with the middle-level feature map of the YOLOv5 trunk of the same scale. After the feature fusion is performed by the second fusion module, the feature enhancement calculation is performed by the fifth feature enhancement module and the feature refinement and conversion calculation is performed by the eighth convolution module in turn to generate the YOLOv5 neck middle-level intermediate feature map;
[0114] The middle-level feature map of the YOLOv5 neck is upsampled by the second upsampling module and compared with the low-level feature map of the YOLOv5 backbone of the same scale. The third fusion module performs feature fusion and outputs the YOLOv5 neck low-level feature map ;
[0115] Secondly, the YOLOv5 neck low-level feature map Fusion is performed through the neck fusion structure to obtain a fused low-level intermediate feature map , using the sixth feature enhancement module to fuse the low-level intermediate feature maps Perform feature enhancement calculations and output fused low-level feature maps , and then after downsampling by the ninth convolution module, it is fused with the YOLOv5 neck middle-level intermediate feature map by the fourth fusion module, and the YOLOv5 neck middle-level feature map is output. ;
[0116] Next, the YOLOv5 neck mid-level feature map Fusion is performed through the neck fusion structure to obtain the intermediate feature map of the fusion layer , using the seventh feature enhancement module to fusion the intermediate feature map Perform feature enhancement calculations and output fusion mid-level feature maps , and then downsampled by the tenth convolution module, and then fused with the YOLOv5 neck high-level intermediate feature map by the fifth fusion module, and output the YOLOv5 neck high-level feature map , obtain YOLOv5 neck feature maps of different scales ,
[0117] Finally, the YOLOv5 neck high-level feature map Fusion is performed through the neck fusion structure to obtain a fused high-level intermediate feature map , using the eighth feature enhancement module to fuse high-level intermediate feature maps Perform feature enhancement calculations and output fused high-level feature maps , obtain fusion feature maps of different scales .
[0118] Furthermore, in order to realize the SAM backbone feature map Neck feature map with YOLOv5 The successful fusion of the YOLOv5 model and the YOLOv5 model requires resolving the information differences caused by the inconsistency in spatial standards or feature incompatibility between the two models, and effectively refining the feature representation of the SAM model; therefore, this embodiment introduces a fixed-scale feature pyramid module FRFPN that can perform adaptive feature conversion as the SAM neck structure in the neck network of the YOLOv5 model.
[0119] In this embodiment, if Figure 5 As shown, it is a structural schematic diagram of the fixed-scale feature pyramid module FRFPN provided in this embodiment, and the fixed-scale feature pyramid module FRFPN includes a fourteenth convolution module, a fifteenth convolution module, a sixteenth convolution module, a sixth fusion module, a seventeenth convolution module, an eighteenth convolution module, a seventh fusion module and a nineteenth convolution module; the fourteenth convolution module, the sixteenth convolution module and the eighteenth convolution module are all 1x1 convolution kernels, which are used to convert the spatial state of the feature map and keep the dimension consistent; the fifteenth convolution module, the seventeenth convolution module and the nineteenth convolution module are all 3x3 convolution kernels, which are used to extract features from the feature map.
[0120] Among them, the input end of the fourteenth convolution module is connected to the third output end of the visual coding module in the SAM trunk structure, the output end of the fourteenth convolution module is connected to the input end of the fifteenth convolution module, and the output end of the fifteenth convolution module is connected to the fifth input end of the neck fusion structure; the input end of the sixteenth convolution module is connected to the second output end of the visual coding module in the SAM trunk structure, the output end of the sixteenth convolution module and the output end of the fourteenth convolution module are commonly connected to the input end of the sixth fusion module, the output end of the sixth fusion module is connected to the input end of the seventeenth convolution module, and the output end of the seventeenth convolution module is connected to the third input end of the neck fusion structure; the input end of the eighteenth convolution module is connected to the first output end of the visual coding module in the SAM trunk structure, the output end of the eighteenth convolution module and the output end of the sixth fusion module are commonly connected to the input end of the seventh fusion module, the output end of the seventh fusion module is connected to the input end of the nineteenth convolution module, and the output end of the nineteenth convolution module is connected to the first input end of the neck fusion structure.
[0121] Combine the following Figure 5 , the method of extracting and converting the SAM trunk feature map using the SAM neck structure in step S202 to generate the SAM neck feature map is further described in detail:
[0122] SAM backbone high-level feature map After the spatial state is converted by the fourteenth convolution module and the dimension is kept consistent, the high-level intermediate feature map of the SAM neck is obtained, and then the feature is extracted and fused by the fifteenth convolution module to output the high-level feature map of the SAM neck. ;
[0123] The hierarchical feature maps in the SAM backbone After the spatial state is converted by the sixteenth convolution module and the dimension is kept consistent, it is directly added to the SAM neck high-level intermediate feature map through the sixth fusion module for integration to obtain the SAM neck middle-level intermediate feature map, which is then extracted and fused by the seventeenth convolution module to output the SAM neck middle-level feature map. ;
[0124] The SAM backbone low-level feature map After the spatial state is converted by the 18th convolution module and the dimension is kept consistent, it is directly added to the intermediate feature map of the middle level of the SAM neck through the 7th fusion module for integration, and then the feature is extracted and fused by the 19th convolution module to output the low-level feature map of the SAM neck. , obtain the SAM neck feature map of the same scale .
[0125] It should be noted that the fixed-scale feature pyramid module can be expressed as:
[0126] F i s x = g g x , i ,1 , i ,3 ,if i =5 and x ∈[ C 5 s ] g g x , i +1,1 +g x , i ,1 , i ,3 ,if i =3,4 and x ∈[ C 3 s , C 4 s ]
[0127]
[0128] in, represents the result obtained after the i-th layer feature map passes through the fixed-scale feature pyramid module; i represents the i-th layer feature map index; x represents the input feature map, which belongs to ; Indicates that the input feature map x passes through The result of the convolution calculation; The convolution kernel of the i-th layer feature map is The convolution weights of Represents the convolution corresponding bias of the i-th layer feature map.
[0129] The fixed-scale feature pyramid module FRFPN proposed in this embodiment improves the traditional feature pyramid network FPN by fixing the spatial scale of the feature map in a top-down manner without upsampling; it not only retains the integrity of the SAM backbone feature map, but also enhances the model's perception of local areas, more effectively promotes information exchange and fusion between different levels, and can generate richer and multi-level feature representations.
[0130] As an optional embodiment, in order to convert the SAM neck feature map Neck feature map with YOLOv5 Effective fusion, this embodiment introduces three parallel large visual feature fusion modules LVFF (Large Vision Feature Fusion) as the neck fusion structure in the neck network of the YOLOv5 model, including the first large visual feature fusion module, the second large visual feature fusion module and the third large visual feature fusion module. Among them, the input end of the first large visual feature fusion module is the first input end and the second input end of the neck fusion structure, and the output end is the first output end of the neck fusion structure, which is used to convert the SAM neck low-level feature map Compared with YOLOv5 neck low-level feature map Fusion is performed to generate fused low-level intermediate feature maps The second largest visual feature fusion module has the third input terminal and the fourth input terminal of the neck fusion structure as its input terminals, and the second output terminal of the neck fusion structure as its output terminal, which is used to convert the SAM neck mid-level feature map into Compared with YOLOv5 neck mid-level feature map Perform fusion to generate intermediate feature maps in the fusion layer The input ends of the third visual feature fusion module are the fifth input end and the sixth input end of the neck fusion structure, and the output end is the third output end of the neck fusion structure, which is used to convert the SAM neck high-level feature map High-level feature map of the neck with YOLOv5 Fusion is performed to generate a fused high-level intermediate feature map .
[0131] It should be noted that the SAM neck feature map The same scale size, while the YOLOv5 neck feature map For different scales; in order to effectively fuse the two, a bilinear interpolation algorithm is needed to be used to fusion the SAM neck feature map before entering the large visual feature fusion module LVFF. Perform feature map scaling to make and The same scale, and The same scale, and In addition, the fusion position of the large visual feature fusion module LVFF in this embodiment is after the fusion module in the YOLOv5 neck structure and before the feature enhancement module.
[0132] Further, such as Figure 6 As shown, it is a structural diagram of the large visual feature fusion module LVFF provided in this embodiment. The first large visual feature fusion module, the second large visual feature fusion module and the third large visual feature fusion module all include a branch feature processing structure, a weight branch processing structure and a feature fusion calculation structure.
[0133] Specifically, the branch feature processing structure includes a parallel YOLOv5 branch feature processing structure and a SAM branch feature processing structure. The YOLOv5 branch feature processing structure includes a 20th convolution module, the input end of the 20th convolution module is the second input end, the fourth input end or the sixth input end of the neck fusion structure; the SAM branch feature processing structure includes a 21st convolution module and a 22nd convolution module connected in series, the input end of the 21st convolution module is the first input end, the third input end or the fifth input end of the neck fusion structure.
[0134] The weight branch processing structure includes an eighth fusion module, an average pooling layer, a twenty-third convolution module, a third upsampling module, a ninth fusion module and a Sigmoid activation function. The input end of the eighth fusion module is connected to the input end of the twentieth convolution module and the output end of the twenty-first convolution module in the branch feature processing structure, the output end of the eighth fusion module is connected in series with the average pooling layer, the twenty-third convolution module and the third upsampling module in sequence, the output end of the third upsampling module and the output end of the eighth fusion module are connected to the input end of the ninth fusion module, and the output end of the ninth fusion module is connected to the input end of the Sigmoid activation function.
[0135] The feature fusion calculation structure includes a tenth fusion module, an eleventh fusion module, a twelfth fusion module, a twenty-fourth convolution module and a thirteenth fusion module. Among them, the input end of the tenth fusion module is connected to the output end of the twentieth convolution module in the branch feature processing structure and the first output end of the Sigmoid activation function in the weight branch processing structure, the input end of the eleventh fusion module is connected to the output end of the twenty-second convolution module in the branch feature processing structure and the second output end of the Sigmoid activation function in the weight branch processing structure, the output end of the tenth fusion module and the output end of the eleventh fusion module are jointly connected to the input end of the twelfth fusion module, the output end of the twelfth fusion module is connected to the input end of the twenty-fourth convolution module, and the output end of the twenty-fourth convolution module and the output end of the twelfth fusion module are jointly connected to the input end of the thirteenth fusion module.
[0136] refer to Figure 6 In step S203, the neck fusion structure performs the following processing steps on the YOLOv5 neck feature map and the SAM neck feature map:
[0137] First, the YOLOv5 neck feature map Input into the YOLOv5 branch feature processing structure, and process the proprietary features through the 20th convolution module to obtain the preprocessed YOLOv5 branch feature map , this process can be expressed as:
[0138]
[0139] in, Represents a 3x3 convolution operation;
[0140] At the same time, the SAM neck feature map Input to the SAM branch feature processing structure, processed by the 21st convolution module, reducing the redundancy between channels without changing the resolution of the feature map, and generating the SAM branch intermediate feature map , and then extracted by the twenty-second convolution module to obtain the preprocessed SAM branch feature map , this process can be expressed as:
[0141]
[0142]
[0143] in, represents a 1x1 convolution operation, Represents a 3x3 convolution operation.
[0144] Secondly, the YOLOv5 neck feature map Intermediate feature map with SAM branch The eighth fusion module directly adds and fuses them to obtain the preliminary fusion feature map. , and then processed by the average pooling layer, the twenty-third convolution module and the third upsampling module, and then combined with the preliminary fusion feature map The ninth fusion module performs residual connection to generate weight features ;
[0145] Weight features Input to the Sigmoid activation function and transform it through adaptive calculation to generate branch weights ;
[0146] This process can be expressed as:
[0147]
[0148]
[0149]
[0150] in, represents the tie pooling operation, represents a 3x3 convolution operation, represents the upsampling operation, Represents the Sigmoid activation function.
[0151] Next, the preprocessed YOLOv5 branch feature map With the preprocessed SAM branch feature map According to branch weight The two fused branch feature maps are directly multiplied and fused through the tenth fusion module and the eleventh fusion module respectively, and the two fused branch feature maps are directly added and fine-tuned through the twelfth fusion module to generate the basic fusion feature map , this process can be expressed as:
[0152]
[0153] The basic fusion feature map The 24th convolution module is used to further optimize the fused feature expression and fine-tune the basic fusion feature map. After the information is obtained, the basic fusion feature maps before and after fine-tuning are respectively After the thirteenth fusion module performs residual connection, the fused intermediate feature maps of different scales are output. , to improve gradient propagation and ensure effective training of deep networks;
[0154] This process can be expressed as:
[0155]
[0156] in, Represents a 1x1 convolution operation.
[0157] Finally, the intermediate feature maps are fused After feature enhancement by the feature enhancement module in the YOLOv5 neck structure, fusion feature maps of different scales are output. .
[0158] The large visual feature fusion module LVFF proposed in this embodiment effectively integrates the SAM neck feature map seamlessly into the YOLOv5 neck structure, enhances the feature representation, and improves the expressiveness and generalization performance of the network model.
[0159] In this embodiment, if Figure 2 As shown, a network structure of a detection head network is provided in the rightmost box of the SAM-YOLOv5 network model, and the detection head network includes an eleventh convolution module, a twelfth convolution module and a thirteenth convolution module connected in parallel; wherein, the input end of the eleventh convolution module is connected to the output end of the sixth feature enhancement module, the input end of the twelfth convolution module is connected to the output end of the seventh feature enhancement module, and the input end of the thirteenth convolution module is connected to the output end of the eighth feature enhancement module.
[0160] The detection head network will fuse low-level feature maps , fusion of mid-level feature maps And fusion of high-level feature maps The images are respectively input into the eleventh, twelfth and thirteenth convolution modules for convolution operations to generate target frames and confidence scores to predict the target category and position, thereby obtaining target detection results.
[0161] The technical solution provided in this embodiment is to use the SAM coding structure with strong generalization ability and the YOLOv5 backbone structure with specific feature extraction ability as dual encoders to extract image features respectively. and , and then use the fixed-scale feature pyramid module FRFPN to fully extract rich features to promote the further refinement and conversion of the SAM model; then introduce the large visual feature fusion module LVFF into the neck network of YOLOv5 to extract features from the SAM model Processing features with the YOLOv5 model Effective fusion is performed to significantly improve feature representation and obtain a fused semantic feature map; finally, the enriched semantic feature map is input into the detection head network to generate a target box and a confidence score. The SAM-YOLOv5 network model constructed in this embodiment enhances the full extraction of effective target features and improves the detection accuracy of the model.
[0162] Embodiment 2:
[0163] This embodiment provides a smart city target detection method based on improved YOLOv5, which is different from the first embodiment in that, in order to pay full attention to the feature importance of the channel dimension, this embodiment draws on the idea of CBAM (Convolutional Block Attention Module) and introduces a channel attention mechanism CAM (Channel Attention Module) in the weight branch processing structure of the large visual feature fusion module LVFF. The input end of the channel attention mechanism is connected to the output end of the ninth fusion module, and the output end is connected to the input end of the Sigmoid activation function. By calculating the importance score of each channel, the weight features of each channel can be dynamically adjusted, thereby enhancing the attention of the key feature map channel.
[0164] Among them, the expression of the channel attention mechanism is:
[0165]
[0166] in, represents the Sigmoid activation function, and represents the linear convolution kernel, and They represent average pooling and maximum pooling respectively, and X represents the weighted features of the input.
[0167] This embodiment explores the diversity and effectiveness of the LVFF structure by introducing a channel attention mechanism CAM into the benchmark large visual feature fusion module LVFF.
[0168] Embodiment three:
[0169] This embodiment provides a smart city target detection method based on improved YOLOv5, and further describes in detail how to train the SAM-YOLOv5 network model in step S1 and step S2 of embodiment 1:
[0170] Collect target sample images in complex scenes in smart cities;
[0171] Marking a detection frame of the target and a corresponding category label in the target sample image;
[0172] Constructing a data set using target sample images with the annotations;
[0173] The data set is randomly divided into a training set, a validation set, and a test set in a ratio of 8:1:1;
[0174] The backbone network in the SAM-YOLOv5 network model is frozen, the constructed model is cross-trained using the training set, the trained model is verified using the verification set, the verified model is tested using the test set, and the model with the optimal value of the test result is used as the final trained SAM-YOLOv5 network model.
[0175] It should be noted that the complex scenes collected in smart cities may include scenes of exposed garbage on the road surface, unclean roads, out-of-store operations, damaged roads, overflowing trash cans, peddlers and other unclean roads in the field of urban appearance and sanitation; it may also include scenes of dirty green spaces such as garbage, missing or dead vegetation, and exposed loess in the field of landscaping.
[0176] Embodiment 4:
[0177] This embodiment provides a smart city target detection method based on improved YOLOv5. Figure 7 As shown, it is a schematic diagram of the network structure of another SAM-YOLOv5 network model provided in this embodiment. The difference between it and the first embodiment is that the fusion position of the large visual feature fusion module LVFF in this embodiment is after the YOLOv5 neck structure, that is, after the feature enhancement module and before the detection head network.
[0178] In order to verify the detection effect of the embodiment of the present invention on smart city targets, a comparative experiment was carried out. Figure 8 This is a thermal map generated by actual detection in the application example of the present invention. Figure 8 As shown in Figure 1, this experiment shows the heat maps of the YOLOv5 model and the SAM-YOLOv5 model at different scales in the object detection task, where "n", "s", "m", "l", and "x" represent the very small, small, medium, large, and extra-large model scales, respectively. Figure 8 It can be seen that the heat map generated by the SAM-YOLOv5 model shows a clearer outline of the object than the YOLOv5 model; especially at the n scale, the clarity of the SAM-YOLOv5 model heat map is particularly obvious. It can be seen that compared with the YOLOv5 model, the SAM-YOLOv5 model, which integrates the large visual model SAM with strong generalization ability into the target detection model YOLOv5, has enhanced its ability to accurately capture the feature map of the object, and can focus more on the expression of the feature information of the target itself, so as to achieve higher accuracy and confidence in target detection tasks at different scales.
[0179] In summary, the SAM-YOLOV5 network model constructed in the smart city target detection method based on the improved YOLOv5 proposed in the embodiment of the present invention, with the help of the SAM model with powerful feature extraction capability, enhances the full extraction of effective image features by the YOLOv5 model and can provide rich effective feature information. This model still has high detection accuracy even in complex scenarios, and can be deployed to the smart city cloud server to achieve smart city monitoring, scheduling, analysis and other tasks to meet terminal applications.
[0180] In addition, deploying the SAM-YOLOv5 model to the smart city cloud server can more accurately identify and classify various objects in the urban environment, which helps to achieve efficient traffic management and safety monitoring and improve the professional level of urban management. For example, in terms of urban appearance and sanitation management, it can use high-definition cameras to capture road conditions, automatically detect problems such as exposed garbage and unclean roads, and promptly notify relevant departments to handle them, thereby improving the response speed and service quality of cleaning work; in terms of garden greening maintenance, it can monitor the vegetation status in green spaces, identify garbage or dead plants, etc., assist in formulating scientific and reasonable maintenance plans, and promote the construction of urban ecological environment; in terms of infrastructure inspection, it can conduct regular inspections of key facilities such as bridges, tunnels, and street lights, quickly locate fault points and warn of potential risks, and ensure the convenience and safety of citizens' lives; in terms of abnormal event response, it can combine real-time video streams for analysis, quickly discover abnormal events, provide decision support for emergency command, and speed up rescue response time; therefore, the SAM-YOLOv5 model can reasonably allocate public resources such as police force, cleaning personnel, and transportation through intelligent analysis, reducing waste while improving work efficiency.
[0181] It should be noted that Figure 1 The logical sequence of the method described in the embodiment of the present invention is only shown. In other possible embodiments of the present application, different methods may be used without conflict. Figure 1 The steps shown or described are completed in the order shown. The method provided in the embodiment of the present invention can be applied to a terminal and can be executed by a target detection device, which can execute the method provided in any one of the embodiments 1 to 4, and has the corresponding functional modules and beneficial effects of the execution method, and the device can be implemented by software and / or hardware, and the device can be integrated in a terminal, for example: any computer device with communication function.
[0182] Furthermore, the target detection device can be easily deployed on the cloud server of the smart city, seamlessly connected with the existing IT infrastructure, and easy to promote and use on a large scale; and both PC and mobile devices can be easily and quickly connected to the device to meet the needs of different user groups. In addition, the device also reserves interfaces to accept more advanced algorithms and technical components, ensuring the stability and foresight of the device.
[0183] Specifically, the hardware parts relied on by the method provided in any one of the embodiments 1 to 4 include: a high-definition camera and a computing platform. The high-definition camera may have 2 million pixels ( ) and have ipx6 waterproof function, the distance between its deployment detection area and the camera should be less than 10 meters and greater than 1 meter; the computing platform can be NVIDIA RTX 4090 server, which should have 24G video memory and no less than 32G memory.
[0184] Embodiment five:
[0185] This embodiment also provides a computer device, which may be a server, and its internal structure diagram may be shown in Figure 9. The computer device includes a processor, a memory, an input / output interface (Input / Output, referred to as I / O), and a communication interface. The processor, the memory, and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface.
[0186] The processor of the computer device is used to provide computing and control capabilities; the memory of the computer device includes a non-volatile storage medium and an internal memory, the non-volatile storage medium stores an operating system, a computer program and a database, the internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium, and the database of the computer device is used to store data obtained and generated in the method of the robot autonomously entering the packaging container; the input / output interface of the computer device is used to exchange information between the processor and an external device; the communication interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, the method provided in any one of the above-mentioned embodiments 1 to 4 is implemented.
[0187] Those skilled in the art will understand that Fig. 9 The structure shown in the figure is only a block diagram of a part of the structure related to the present application scheme, and does not constitute a limitation on the computer device to which the present application scheme is applied. The specific computer device may include Fig. 9 More or fewer components may be shown, or certain components may be combined, or may have a different arrangement of components.
[0188] The computer device provided in this embodiment can execute the smart city target detection method based on the improved YOLOv5 provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the execution method.
[0189] Embodiment six:
[0190] This embodiment also provides a computer-readable storage medium on which a computer program is stored. When the program is executed by a processor, the steps of the method described in any one of Embodiments 1 to 4 are implemented.
[0191] The computer-readable storage medium provided in this embodiment can execute the smart city target detection method based on the improved YOLOv5 provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the execution method.
[0192] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the present application may adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program codes.
[0193] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0194] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0195] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.
[0196] The embodiments of the present invention are described above in conjunction with the accompanying drawings, but the present invention is not limited to the above-mentioned specific implementation methods. The above-mentioned specific implementation methods are merely illustrative and not restrictive. Under the enlightenment of the present invention, ordinary technicians in this field can also make many forms without departing from the scope of protection of the purpose of the present invention and the claims, which all fall within the protection of the present invention.
Claims
1. A smart city target detection method based on improved YOLOv5, characterized in that: The method comprises: Collect images of smart city scenes to be detected; Input the scene image into a pre-trained SAM-YOLOv5 network model to obtain a target detection result, wherein the SAM-YOLOv5 network model refers to a Segment Anything Model-YOLOv5 network model; The SAM-YOLOv5 network model includes a backbone network, a neck network and a detection head network connected in sequence, the backbone network includes a YOLOv5 backbone structure and a SAM backbone structure connected in parallel, the neck network includes a YOLOv5 neck structure and a SAM neck structure connected in parallel, and a neck fusion structure connecting the YOLOv5 neck structure and the SAM neck structure; The neck fusion structure is three parallel large visual feature fusion modules, including a first large visual feature fusion module, a second large visual feature fusion module and a third large visual feature fusion module, wherein the first large visual feature fusion module, the second large visual feature fusion module and the third large visual feature fusion module all include a branch feature processing structure, a weight branch processing structure and a feature fusion calculation structure; The processing process of the SAM-YOLOv5 network model on the scene image includes: Inputting the scene image into the YOLOv5 backbone structure and the SAM backbone structure for feature extraction, respectively, to obtain a YOLOv5 backbone feature map and a SAM backbone feature map; The YOLOv5 neck structure is used to sample and fuse the YOLOv5 trunk feature map to generate the YOLOv5 neck feature map. At the same time, the SAM neck structure is used to extract and transform the SAM trunk feature map to generate the SAM neck feature map. The neck fusion structure is used to fuse the YOLOv5 neck feature map with the SAM neck feature map to generate a fused feature map; The fused feature map is input into the detection head network for convolution operation to predict the target category and position, thereby obtaining the target detection result.
2. The smart city target detection method based on improved YOLOv5 according to claim 1, characterized in that: The YOLOv5 backbone structure includes a first convolution module, a second convolution module, a first feature enhancement module, a third convolution module, a second feature enhancement module, a fourth convolution module, a third feature enhancement module, a fifth convolution module, a fourth feature enhancement module and a spatial pyramid pooling module which are sequentially connected in series; The method of inputting the scene image into a YOLOv5 backbone structure for feature extraction to obtain a YOLOv5 backbone feature map includes: The scene image is sequentially passed through a first convolution module, a second convolution module, and a first feature enhancement module to capture low-level feature information, and then sequentially passed through a third convolution module and a second feature enhancement module to perform a downsampling operation, and output a YOLOv5 backbone low-level feature map; The low-level feature map of the YOLOv5 backbone is downsampled through the fourth convolution module and the third feature enhancement module in turn, and the middle-level feature map of the YOLOv5 backbone is output; The middle-level feature map of the YOLOv5 backbone is processed by the fifth convolution module, the fourth feature enhancement module and the spatial pyramid pooling module in sequence, and the high-level feature map of the YOLOv5 backbone is output.
3. The smart city target detection method based on improved YOLOv5 according to claim 1, characterized in that: The SAM backbone structure includes a sixth convolution module, a mapping transposition module, a position encoding module, a first fusion module and a visual encoding module; The method of inputting the scene image into the SAM backbone structure for feature extraction to obtain a SAM backbone feature map comprises: Inputting the scene image into a sixth convolution module for block division and convolution to obtain an image block feature map of a fixed size; A mapping transposition module is used to flatten and sort the fixed-size image block feature map to obtain a flattened image block feature map; The flattened image block feature map and the position feature map generated by the position encoding module are directly added together through a first fusion module to perform feature fusion, thereby obtaining a preprocessed feature map; The preprocessed feature map is input into the visual encoding module for processing, and feature extraction is performed in the ratio of 1:1:3:
1. The SAM backbone low-level feature map, SAM backbone middle-level feature map and SAM backbone high-level feature map are output in sequence according to the latter three ratios.
4. The smart city target detection method based on improved YOLOv5 according to claim 1, characterized in that: The YOLOv5 neck structure includes a first neck branch structure, a second neck branch structure, a third neck branch structure and a fourth neck branch structure; Among them, the first neck branch structure includes the seventh convolution module, the first upsampling module, the second fusion module, the fifth feature enhancement module, the eighth convolution module, the second upsampling module and the third fusion module; the second neck branch structure includes the sixth feature enhancement module, the ninth convolution module and the fourth fusion module; the third neck branch structure includes the seventh feature enhancement module, the tenth convolution module and the fifth fusion module; the fourth neck branch structure includes the eighth feature enhancement module; The method of sampling and fusing a YOLOv5 trunk feature map by using a YOLOv5 neck structure to generate a YOLOv5 neck feature map includes: The YOLOv5 trunk high-level feature map is refined through the seventh convolution module to generate the YOLOv5 neck high-level intermediate feature map; After the YOLOv5 neck high-level intermediate feature map is upsampled by the first upsampling module, it is fused with the YOLOv5 trunk middle-level feature map by the second fusion module, and then processed by the fifth feature enhancement module and the eighth convolution module in sequence to generate the YOLOv5 neck middle-level intermediate feature map; After the YOLOv5 neck middle-level intermediate feature map is upsampled by the second upsampling module, it is fused with the YOLOv5 trunk low-level feature map by the third fusion module to output the YOLOv5 neck low-level feature map; The YOLOv5 neck low-level feature map is fused through the neck fusion structure to obtain a fused low-level intermediate feature map, and the fused low-level intermediate feature map is enhanced by the sixth feature enhancement module to output the fused low-level feature map, which is then downsampled by the ninth convolution module and then fused with the YOLOv5 neck middle-level intermediate feature map by the fourth fusion module to output the YOLOv5 neck middle-level feature map; The YOLOv5 neck middle-level feature map is fused through the neck fusion structure to obtain a fused middle-level intermediate feature map, and the seventh feature enhancement module is used to enhance the fused middle-level intermediate feature map, and the fused middle-level feature map is output. After downsampling operation is performed by the tenth convolution module, the feature fusion is performed with the YOLOv5 neck high-level intermediate feature map through the fifth fusion module, and the YOLOv5 neck high-level feature map is output; The YOLOv5 neck high-level feature map is fused through the neck fusion structure to obtain a fused high-level intermediate feature map. The eighth feature enhancement module is used to enhance the fused high-level intermediate feature map and output the fused high-level feature map.
5. The smart city target detection method based on improved YOLOv5 according to claim 1, characterized in that: The SAM neck structure is a fixed-scale feature pyramid module, including a fourteenth convolution module, a fifteenth convolution module, a sixteenth convolution module, a sixth fusion module, a seventeenth convolution module, an eighteenth convolution module, a seventh fusion module and a nineteenth convolution module; The method for extracting and converting a SAM trunk feature map by using a SAM neck structure to generate a SAM neck feature map includes: The high-level feature map of the SAM trunk is converted into a spatial state by the fourteenth convolution module to obtain a high-level intermediate feature map of the SAM neck, which is then extracted and fused by the fifteenth convolution module to output a high-level feature map of the SAM neck; After the spatial state of the middle-level feature map of the SAM trunk is converted by the sixteenth convolution module, it is directly added and integrated with the high-level intermediate feature map of the SAM neck by the sixth fusion module to obtain the middle-level intermediate feature map of the SAM neck, and then the feature is extracted and fused by the seventeenth convolution module to output the middle-level feature map of the SAM neck; After the spatial state of the SAM trunk low-level feature map is converted by the eighteenth convolution module, it is directly added to the SAM neck middle-level intermediate feature map by the seventh fusion module for integration, and then feature extraction and fusion are performed by the nineteenth convolution module to output the SAM neck low-level feature map.
6. The smart city target detection method based on improved YOLOv5 according to claim 1, characterized in that: in, The branch feature processing structure includes a parallel YOLOv5 branch feature processing structure and a SAM branch feature processing structure, the YOLOv5 branch feature processing structure includes a 20th convolution module, and the SAM branch feature processing structure includes a 21st convolution module and a 22nd convolution module in series; the weight branch processing structure includes an 8th fusion module, an average pooling layer, a 23rd convolution module, a 3rd upsampling module, a 9th fusion module and a Sigmoid activation function; the feature fusion calculation structure includes a 10th fusion module, an 11th fusion module, a 12th fusion module, a 24th convolution module and a 13th fusion module; The method of fusing the YOLOv5 neck feature map with the SAM neck feature map by using the neck fusion structure to generate a fused feature map includes: The YOLOv5 neck feature map is processed by the 20th convolution module to obtain a preprocessed YOLOv5 branch feature map, and the SAM neck feature map is processed by the 21st convolution module to generate a SAM branch intermediate feature map, which is then extracted by the 22nd convolution module to obtain a preprocessed SAM branch feature map; The YOLOv5 neck feature map and the SAM branch intermediate feature map are directly added and fused through the eighth fusion module to obtain a preliminary fusion feature map, which is then processed in turn through the average pooling layer, the twenty-third convolution module and the third upsampling module, and then residually connected with the preliminary fusion feature map through the ninth fusion module to generate weight features; Input the weight feature into the Sigmoid activation function for conversion to generate branch weights; The pre-processed YOLOv5 branch feature map and the pre-processed SAM branch feature map are directly multiplied by the tenth fusion module and the eleventh fusion module according to the branch weights for fusion, and the two fused branch feature maps are directly added and fine-tuned by the twelfth fusion module to generate a basic fusion feature map; After fine-tuning the basic fusion feature map through the twenty-fourth convolution module, the basic fusion feature map before and after fine-tuning is residually connected through the thirteenth fusion module to output the fusion intermediate feature map; After the fused intermediate feature map is enhanced by the feature enhancement module in the YOLOv5 neck structure, the fused feature map is output.
7. The smart city target detection method based on improved YOLOv5 according to claim 6, characterized in that: A channel attention mechanism is introduced into the weight branch processing structure, which is used to dynamically adjust the weight features through calculation to enhance the attention of key features; Among them, the expression of the channel attention mechanism is: , in, represents the Sigmoid activation function, and represents the linear convolution kernel, and Represent average pooling and maximum pooling respectively, X Represents the weight feature of the input.
8. The smart city target detection method based on improved YOLOv5 according to claim 1, characterized in that: The detection head network includes an eleventh convolution module, a twelfth convolution module and a thirteenth convolution module connected in parallel; The detection head network inputs the fused low-level feature map, the fused mid-level feature map and the fused high-level feature map into the eleventh convolution module, the twelfth convolution module and the thirteenth convolution module respectively for convolution operations to generate target boxes and confidence scores to predict the target category and position, thereby obtaining target detection results.
9. A computer device, characterized in that: including storage media and processors; The storage medium is used to store instructions; The processor is configured to operate according to the instructions to execute the steps of the method according to any one of claims 1 to 8.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps of the method according to any one of claims 1 to 8 are implemented.
Citation Information
Patent Citations
Remote sensing image target detection method
CN118314434A
Urban facility anomaly identification method and device, electronic equipment and storage medium
CN118429623A