Pedestrian detection method, device, equipment and medium
By improving the neck network of the YOLOv5 network and adding feature fusion layer, the problem of low detection accuracy of infrared images is solved, and higher detection accuracy and lower error detection rate are achieved.
Patent Information
- Application Number
- CN202111546976.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-16
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2041-12-16
AI Technical Summary
The existing infrared image pedestrian detection methods have low detection accuracy and are prone to misdetection, especially because the large number of convolutional layers in the YOLOv5 network cause the loss of important feature information at low resolution images.
Improve the network structure of the YOLOv5 network, including improving the neck network consisting of three FPN network stacks, splicing deep and shallow features through cross-layer splicing layers, and adding feature fusion layers to fuse multi-scale feature maps.
Through the improved YOLOv5 network, pedestrian characteristic information can be better learned, detection accuracy can be improved, error detection rate can be reduced, and the problem of low detection accuracy of existing pedestrian detection methods can be improved.
Smart Images

Figure CN114220125B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of image processing technology, and in particular to a pedestrian detection method, device, equipment and medium. Background Art
[0002] Infrared thermal imaging technology is a passive, non-contact optical imaging technology with the advantages of strong anti-interference ability, long effective distance and day and night operation. It is widely used in intelligent security systems.
[0003] However, infrared images have low contrast and poor detail resolution, which results in low detection accuracy and easy false detection of pedestrians in infrared surveillance videos. Among them, the YOLOv5 network has gained wide attention for its good performance, but the large number of convolutional layers in its network structure makes it easy for low-resolution images to lose important feature information, resulting in low detection accuracy. Summary of the invention
[0004] The present application provides a pedestrian detection method, device, equipment and medium for improving the technical problem of low detection accuracy in existing pedestrian detection methods.
[0005] In view of this, the first aspect of the present application provides a pedestrian detection method, comprising:
[0006] Improve the network structure of the YOLOv5 network to obtain an improved YOLOv5 network, wherein the improved YOLOv5 network includes a backbone network, an improved neck network, a feature fusion layer and a detection head, wherein the improved neck network is composed of three FPN networks stacked together, wherein the FPN network includes a convolutional layer, an upsampling layer, a cross-layer splicing layer and a CSP module connected in sequence, wherein the other end of the cross-layer splicing layer is connected to the backbone network, and the input end of each feature fusion layer is respectively connected to the convolutional layer and the CSP module in the first FPN network in the improved neck network, the CSP module in the second FPN network and the CSP module in the third FPN network, wherein the feature fusion layer is used to perform feature fusion on the input feature map, and the detection head is used to perform pedestrian detection on the fused feature map output by each feature fusion layer;
[0007] Acquire an infrared pedestrian training set, where the infrared pedestrian training set includes a plurality of infrared pedestrian images;
[0008] The improved YOLOv5 network is trained by using the infrared pedestrian training set until the loss value of the improved YOLOv5 network meets a preset condition, thereby obtaining a pedestrian detection model;
[0009] Pedestrian detection is performed on the infrared pedestrian image to be detected by using the pedestrian detection model to obtain a pedestrian detection result.
[0010] Optionally, the backbone network includes a Focus layer, a convolutional layer, 3 CSP modules and an SPP layer, the other end of the cross-layer splicing layer in the first FPN network is connected to the third CSP module in the backbone network, the other end of the cross-layer splicing layer in the second FPN network is connected to the second CSP module in the backbone network, and the other end of the cross-layer splicing layer in the third FPN network is connected to the first CSP module in the backbone network.
[0011] Optionally, a depth parameter depth_multiple of the backbone network is set to 0.33, and a width parameter width_multiple is set to 0.5.
[0012] Optionally, the number of the feature fusion layers is 4.
[0013] A second aspect of the present application provides a pedestrian detection device, comprising:
[0014] A network construction unit is used to improve the network structure of a YOLOv5 network to obtain an improved YOLOv5 network, wherein the improved YOLOv5 network includes a backbone network, an improved neck network, a feature fusion layer and a detection head, wherein the improved neck network is composed of three FPN networks stacked together, wherein the FPN network includes a convolutional layer, an upsampling layer, a cross-layer splicing layer and a CSP module connected in sequence, wherein the other end of the cross-layer splicing layer is connected to the backbone network, and the input ends of each of the feature fusion layers are respectively connected to the convolutional layer and the CSP module in the first FPN network in the improved neck network, the CSP module in the second FPN network and the CSP module in the third FPN network, wherein the feature fusion layer is used to perform feature fusion on the input feature map, and the detection head is used to perform pedestrian detection on the fused feature map output by each of the feature fusion layers;
[0015] An acquisition unit, used for acquiring an infrared pedestrian training set, wherein the infrared pedestrian training set includes a plurality of infrared pedestrian images;
[0016] A training unit, used for training the improved YOLOv5 network through the infrared pedestrian training set until the loss value of the improved YOLOv5 network meets a preset condition, thereby obtaining a pedestrian detection model;
[0017] The detection unit is used to perform pedestrian detection on the infrared pedestrian image to be detected by using the pedestrian detection model to obtain a pedestrian detection result.
[0018] Optionally, the backbone network includes a Focus layer, a convolutional layer, 3 CSP modules and an SPP layer, the other end of the cross-layer splicing layer in the first FPN network is connected to the third CSP module in the backbone network, the other end of the cross-layer splicing layer in the second FPN network is connected to the second CSP module in the backbone network, and the other end of the cross-layer splicing layer in the third FPN network is connected to the first CSP module in the backbone network.
[0019] Optionally, a depth parameter depth_multiple of the backbone network is set to 0.33, and a width parameter width_multiple is set to 0.5.
[0020] Optionally, the number of the feature fusion layers is 4.
[0021] A third aspect of the present application provides a pedestrian detection device, the device comprising a processor and a memory;
[0022] The memory is used to store program code and transmit the program code to the processor;
[0023] The processor is used to execute any one of the pedestrian detection methods described in the first aspect according to the instructions in the program code.
[0024] A fourth aspect of the present application provides a computer-readable storage medium, which is used to store program code. When the program code is executed by a processor, it implements any pedestrian detection method described in the first aspect.
[0025] It can be seen from the above technical solutions that this application has the following advantages:
[0026] The present application provides a pedestrian detection method, including: improving the network structure of a YOLOv5 network to obtain an improved YOLOv5 network, wherein the improved YOLOv5 network includes a backbone network, an improved neck network, a feature fusion layer and a detection head, wherein the improved neck network is composed of three FPN networks stacked together, wherein the FPN network includes a convolutional layer, an upsampling layer, a cross-layer splicing layer and a CSP module connected in sequence, wherein the other end of the cross-layer splicing layer is connected to the backbone network, and the input ends of each feature fusion layer are respectively connected to the convolutional layer, the CSP module, the second FPN module in the first FPN network in the improved neck network, and the input ends of each feature fusion layer are respectively connected to the convolutional layer, the CSP module, the second FPN module in the second FPN network in the improved neck network, and the input ends of each feature fusion layer are respectively connected to the convolutional layer, the CSP module, the second FPN module in the first FPN network in the improved neck network, and the second FPN module in the second FPN network ... The CSP module in the first FPN network is connected to the CSP module in the third FPN network, the feature fusion layer is used to perform feature fusion on the input feature map, and the detection head is used to perform pedestrian detection on the fused feature maps output by each feature fusion layer; an infrared pedestrian training set is obtained, and the infrared pedestrian training set includes several infrared pedestrian images; an improved YOLOv5 network is trained by using the infrared pedestrian training set until the loss value of the improved YOLOv5 network meets the preset conditions, so as to obtain a pedestrian detection model; pedestrian detection is performed on the infrared pedestrian image to be detected by using the pedestrian detection model, and a pedestrian detection result is obtained.
[0027] In the present application, the neck network in the YOLOv5 network is improved, and the improved YOLOv5 network is trained by using an infrared pedestrian training set to obtain a pedestrian detection model for pedestrian detection. By deleting multiple convolutional layers, the infrared pedestrian image is prevented from losing detail information after a large number of convolution operations, so that the improved YOLOv5 network can better learn pedestrian feature information and improve detection accuracy. In the improved neck network, the deep feature information in the neck network is spliced with the shallow feature information extracted by the backbone network through a cross-layer splicing layer to compensate for the detail information lost by the deep features after multiple convolution operations. A feature fusion layer is added to the improved neck network to perform feature fusion on multiple scale feature maps output by the improved neck network, enhance feature representation, and further improve detection accuracy, thereby improving the technical problem of low detection accuracy in existing pedestrian detection methods. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.
[0029] Figure 1 A flowchart of a pedestrian detection method provided in an embodiment of the present application;
[0030] Figure 2A schematic diagram of the network structure of an improved rear neck network provided in an embodiment of the present application;
[0031] Figure 3 A schematic diagram of the network structure of an improved YOLOv5 network provided in an embodiment of the present application;
[0032] Figure 4 A schematic diagram of the structure of a pedestrian detection device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0033] The present application provides a pedestrian detection method, device, equipment and medium for improving the technical problem of low detection accuracy in existing pedestrian detection methods.
[0034] In order to enable those skilled in the art to better understand the solution of the present application, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0035] For easier understanding, see Figure 1 , the present application embodiment provides a pedestrian detection method, including:
[0036] Step 101: Improve the network structure of the YOLOv5 network to obtain an improved YOLOv5 network.
[0037] The embodiments of the present application take into account that the existing monitoring systems mostly use infrared images for pedestrian detection, but the infrared images have low contrast and poor detail resolution, which results in low detection accuracy of pedestrian detection methods in monitoring videos and prone to false detection. In addition, the existing pedestrian detection methods mostly use the YOLOv5 network for pedestrian detection. If the YOLOv5 network is directly used to detect pedestrians in the collected infrared images, due to the presence of a large number of convolutional layers in the YOLOv5 network structure, the low-resolution infrared images are prone to lose important feature information, which in turn reduces the detection accuracy.
[0038] In order to improve the above problems, the embodiment of the present application improves the network structure of the YOLOv5 network. Specifically, an upsampling layer is added to the original neck network (Neck) of YOLOv5, and the PAN network in the original neck network is removed to obtain an improved neck network, such as Figure 2As shown in the figure, the PAN network is the original neck network. After multiple upsampling, the feature map is downsampled through multiple convolutional layers, which is called the PAN network. By deleting the PAN network, the convolution processing of multiple convolutional layers can be avoided, which causes the low-resolution image to lose important feature information and reduce the detection accuracy. The improved neck network is composed of three FPN networks stacked together. Each FPN network consists of a convolution layer (Conv) with a convolution kernel of 1×1, an upsampling layer (Upsample), a cross-layer splicing layer (Concat) and a CSP module (C3). The feature map extracted by the backbone network is sequentially convolved, upsampled, and spliced with the shallow features in the backbone network after the improved neck network. After feature extraction by the CSP module, a feature map rich in shallow information is obtained. The feature map output by the first FPN network is then sequentially extracted and spliced through the second FPN network and the third FPN network.
[0039] Please refer to Figure 2 In an embodiment of the present application, the other end of the cross-layer splicing layer in the first FPN network is connected to the third CSP module in the backbone network, for splicing the feature map output by the first upsampling layer with the feature map output by the third CSP module, the other end of the cross-layer splicing layer in the second FPN network is connected to the second CSP module in the backbone network, for splicing the feature map output by the second upsampling layer with the feature map output by the second CSP module, and the other end of the cross-layer splicing layer in the third FPN network is connected to the first CSP module in the backbone network, for splicing the feature map output by the third upsampling layer with the feature map output by the first CSP module.
[0040] In the embodiment of the present application, a feature multi-scale fusion method is adopted, and an FPN network is used in the neck network to integrate more shallow information of the backbone network, thereby further improving the network's ability to extract and resolve small target features.
[0041] Furthermore, the embodiment of the present application adds a feature fusion layer to the improved neck network. Preferably, the number of feature fusion layers is 4. The input end of each feature fusion layer is respectively connected to the convolution layer (i.e., the first detection layer) and the CSP module (i.e., the second detection layer) in the first FPN network in the improved neck network, the CSP module (i.e., the third detection layer) in the second FPN network, and the CSP module (i.e., the fourth detection layer) in the third FPN network. The feature fusion layer is used to perform feature fusion on the input feature map. The four feature fusion layers (ASFF) are adaptive spatial feature fusion networks, which perform adaptive linear weighting on the input feature map. When a small target requires more shallow information, the feature fusion layer will assign a larger weight to the feature map rich in shallow information, and vice versa. Each feature fusion layer inputs the output fusion feature into the detection head (Detect) to perform pedestrian detection on the fusion feature map output by each feature fusion layer. The feature fusion layer can be expressed as:
[0042]
[0043] In the formula, α, β, γ, and δ represent the weights of the first detection layer, the second detection layer, the third detection layer, and the fourth detection layer, respectively. They are learnable parameters, and the weights of each detection layer are different. 1→l is the feature map of the first detection layer input to the lth feature fusion layer, and (i, j) is the point coordinate of the feature map. Since the feature maps output by each detection layer need to be linearly weighted, the size and number of channels of each feature map must be the same, so the feature maps of different detection layers need to be upsampled or downsampled to adjust the number of channels and size. For example, to obtain ASFF-4: first adjust the feature maps output by the first, second, and third detection layers to the same number of channels in the fourth detection layer through 1x1 convolution, and then use the interpolation method (upsampling) to adjust the size of the feature maps to be consistent; if ASFF-1 is to be obtained: the second detection layer is convolved with the convolution kernel (3x3, stride = 2) and the third detection layer is convolved with the maximum pooling output of (3x3, stride = 2), and the fourth detection layer is convolved with the maximum pooling output of (3x3, stride = 2) and then stride = 4 to obtain the feature map with the same size as the first detection layer, and then the number of channels is adjusted through 1x1 convolution to perform the addition operation. ASFF-2 and ASFF-3 were obtained in a similar manner.
[0044] In an embodiment of the present application, the output feature map of the detection layer is weightedly fused on the pedestrian detection side to enhance the feature representation, so that the network can better learn the features required by the target and improve the pedestrian detection accuracy.
[0045] Furthermore, in the embodiment of the present application, the backbone network (Backbone) in YOLOv5 is improved. The backbone network includes a Focus layer, a convolution layer, three CSP modules and an SPP layer. The Focus layer is used to perform image slicing operations on the input infrared pedestrian image. The CSP module is used to enhance the learning ability of the network, reduce the amount of calculation while ensuring detection accuracy. The SPP layer uses pooling layers of different sizes to perform pooling operations on the input feature maps, and then perform feature fusion to form pooled features; the width_multiple and depth_multiple parameters are set to 0.5 and 0.33 respectively to control the network depth and width, and the output dimensions of all layers in the network are reduced to 1 / 2 of the original, and the number of stacked CSP Bottleneck-3 layers is reduced to 1 / 3 of the original. The final improved YOLOv5 network structure is as follows: Figure 3 As shown in the figure, the improved YOLOv5 network includes a backbone network, an improved neck network, four feature fusion layers and a detection head, where the parameters in brackets of each layer correspond to: C3(c_in,c_out), Conv(c_in,c_out,kernel_size,stride), SPP(c_in,c_out,[kernel_size1,kernel_size2,kernel_size3]), c_in and c_out are the number of input and output feature map channels, kernel_size is the convolution kernel size, and stride is the stride. C3 is a CSP module, C3*1 means that the CSP module consists of a CSP layer, and the CSP layer is consistent with the CSP layer structure in the existing YOLOv5, which will not be repeated here, and C3*3 means that the CSP module consists of 3 CSP layers stacked to improve the network structure's ability to extract features.
[0046] Step 102: Obtain an infrared pedestrian training set, where the infrared pedestrian training set includes a number of infrared pedestrian images.
[0047] Infrared pedestrian images can be collected through visual acquisition devices such as infrared surveillance cameras, and then preprocessed and labeled to obtain an infrared pedestrian training set.
[0048] Step 103: Train the improved YOLOv5 network using the infrared pedestrian training set until the loss value of the improved YOLOv5 network meets a preset condition, thereby obtaining a pedestrian detection model.
[0049] The improved YOLOv5 network is iteratively trained through the infrared pedestrian training set, and the loss value after each round of training is calculated. If the loss value does not meet the preset conditions, the next round of iterative training is entered until the loss value of the improved YOLOv5 network meets the preset conditions, and the trained improved YOLOv5 network is obtained. The trained improved YOLOv5 network is used as the pedestrian detection model.
[0050] A loss threshold may be set. When the loss value is lower than or equal to the loss threshold, it may be determined that the loss value meets the preset condition. If the loss value is greater than the loss threshold, it may be determined that the loss value does not meet the preset condition.
[0051] Step 104: Perform pedestrian detection on the infrared pedestrian image to be detected using the pedestrian detection model to obtain a pedestrian detection result.
[0052] The infrared pedestrian images to be detected collected in real time can be input into the pedestrian detection model for pedestrian detection, and the pedestrian detection results can be output.
[0053] In an embodiment of the present application, the neck network in the YOLOv5 network is improved, and the improved YOLOv5 network is trained by using an infrared pedestrian training set to obtain a pedestrian detection model for pedestrian detection. By deleting multiple convolutional layers, the infrared pedestrian image is prevented from losing detail information after a large number of convolution operations, so that the improved YOLOv5 network can better learn pedestrian feature information and improve detection accuracy; and in the improved neck network, the deep feature information in the neck network is spliced with the shallow feature information extracted by the backbone network through a cross-layer splicing layer to make up for the detail information lost by the deep features after multiple convolution operations, and a feature fusion layer is added to the improved neck network to perform feature fusion on multiple scale feature maps output by the improved neck network, enhance feature representation, and further improve detection accuracy, thereby improving the technical problem of low detection accuracy in existing pedestrian detection methods.
[0054] The above is an embodiment of a pedestrian detection method provided by the present application, and the following is an embodiment of a pedestrian detection device provided by the present application.
[0055] Please refer to Figure 4 , a pedestrian detection device provided in an embodiment of the present application includes:
[0056] A network construction unit is used to improve the network structure of a YOLOv5 network to obtain an improved YOLOv5 network. The improved YOLOv5 network includes a backbone network, an improved neck network, a feature fusion layer and a detection head. The improved neck network is composed of three FPN networks stacked together. The FPN network includes a convolutional layer, an upsampling layer, a cross-layer splicing layer and a CSP module connected in sequence. The other end of the cross-layer splicing layer is connected to the backbone network. The input ends of each feature fusion layer are respectively connected to the convolutional layer and the CSP module in the first FPN network in the improved neck network, the CSP module in the second FPN network and the CSP module in the third FPN network. The feature fusion layer is used to perform feature fusion on the input feature map. The detection head is used to perform pedestrian detection on the fused feature map output by each feature fusion layer.
[0057] An acquisition unit, used for acquiring an infrared pedestrian training set, where the infrared pedestrian training set includes a plurality of infrared pedestrian images;
[0058] A training unit, used for training the improved YOLOv5 network through the infrared pedestrian training set until the loss value of the improved YOLOv5 network meets a preset condition, thereby obtaining a pedestrian detection model;
[0059] The detection unit is used to perform pedestrian detection on the infrared pedestrian image to be detected through the pedestrian detection model to obtain the pedestrian detection result.
[0060] As a further improvement, the backbone network includes a Focus layer, a convolutional layer, 3 CSP modules and an SPP layer. The other end of the cross-layer splicing layer in the first FPN network is connected to the third CSP module in the backbone network, the other end of the cross-layer splicing layer in the second FPN network is connected to the second CSP module in the backbone network, and the other end of the cross-layer splicing layer in the third FPN network is connected to the first CSP module in the backbone network.
[0061] As a further improvement, the depth parameter depth_multiple of the backbone network is set to 0.33, and the width parameter width_multiple is set to 0.5.
[0062] In an embodiment of the present application, the neck network in the YOLOv5 network is improved, and the improved YOLOv5 network is trained by using an infrared pedestrian training set to obtain a pedestrian detection model for pedestrian detection. By deleting multiple convolutional layers, the infrared pedestrian image is prevented from losing detail information after a large number of convolution operations, so that the improved YOLOv5 network can better learn pedestrian feature information and improve detection accuracy; and in the improved neck network, the deep feature information in the neck network is spliced with the shallow feature information extracted by the backbone network through a cross-layer splicing layer to make up for the detail information lost by the deep features after multiple convolution operations, and a feature fusion layer is added to the improved neck network to perform feature fusion on multiple scale feature maps output by the improved neck network, enhance feature representation, and further improve detection accuracy, thereby improving the technical problem of low detection accuracy in existing pedestrian detection methods.
[0063] The embodiment of the present application also provides a pedestrian detection device, the device comprising a processor and a memory;
[0064] The memory is used to store the program code and transmit the program code to the processor;
[0065] The processor is used to execute the pedestrian detection method in the aforementioned method embodiment according to the instructions in the program code.
[0066] An embodiment of the present application also provides a computer-readable storage medium, which is used to store program code. When the program code is executed by a processor, the pedestrian detection method in the aforementioned method embodiment is implemented.
[0067] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the above-described devices and units can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0068] It should be understood that in the present application, "at least one (item)" means one or more, and "plurality" means two or more. "And / or" is used to describe the association relationship of associated objects, indicating that three relationships may exist. For example, "A and / or B" can mean: only A exists, only B exists, and A and B exist at the same time, where A and B can be singular or plural. The character " / " generally indicates that the objects associated before and after are in an "or" relationship. "At least one of the following" or similar expressions refers to any combination of these items, including any combination of single or plural items. For example, at least one of a, b or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.
[0069] In the several embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0070] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0071] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.
[0072] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for executing all or part of the steps of the method described in each embodiment of the present application through a computer device (which can be a personal computer, a server, or a network device, etc.). The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (full name in English: Read-Only Memory, English abbreviation: ROM), random access memory (full name in English: Random Access Memory, English abbreviation: RAM), disk or optical disk and other media that can store program codes.
[0073] As described above, the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. A pedestrian detection method, It is characterized in that include: Improve the network structure of the YOLOv5 network to obtain an improved YOLOv5 network, wherein the improved YOLOv5 network includes a backbone network, an improved neck network, a feature fusion layer and a detection head, wherein the improved neck network is composed of three FPN networks stacked together, wherein the FPN network includes a convolutional layer, an upsampling layer, a cross-layer splicing layer and a CSP module connected in sequence, wherein the other end of the cross-layer splicing layer is connected to the backbone network, and the input end of each feature fusion layer is respectively connected to the convolutional layer and the CSP module in the first FPN network in the improved neck network, the CSP module in the second FPN network and the CSP module in the third FPN network, wherein the feature fusion layer is used to perform feature fusion on the input feature map, and the detection head is used to perform pedestrian detection on the fused feature map output by each feature fusion layer; Acquire an infrared pedestrian training set, where the infrared pedestrian training set includes a plurality of infrared pedestrian images; The improved YOLOv5 network is trained by using the infrared pedestrian training set until the loss value of the improved YOLOv5 network meets a preset condition, thereby obtaining a pedestrian detection model; Pedestrian detection is performed on the infrared pedestrian image to be detected by using the pedestrian detection model to obtain a pedestrian detection result.
2. The pedestrian detection method according to claim 1, It is characterized in that The backbone network includes a Focus layer, a convolutional layer, three CSP modules and an SPP layer. The other end of the cross-layer splicing layer in the first FPN network is connected to the third CSP module in the backbone network, the other end of the cross-layer splicing layer in the second FPN network is connected to the second CSP module in the backbone network, and the other end of the cross-layer splicing layer in the third FPN network is connected to the first CSP module in the backbone network.
3. The pedestrian detection method according to claim 1, It is characterized in that The depth parameter depth_multiple of the backbone network is set to 0.33, and the width parameter width_multiple is set to 0.
5.
4. The pedestrian detection method according to claim 1, It is characterized in that The number of the feature fusion layers is 4.
5. A pedestrian detection device, It is characterized in that include: A network construction unit is used to improve the network structure of a YOLOv5 network to obtain an improved YOLOv5 network, wherein the improved YOLOv5 network includes a backbone network, an improved neck network, a feature fusion layer and a detection head, wherein the improved neck network is composed of three FPN networks stacked together, wherein the FPN network includes a convolutional layer, an upsampling layer, a cross-layer splicing layer and a CSP module connected in sequence, wherein the other end of the cross-layer splicing layer is connected to the backbone network, and the input ends of each of the feature fusion layers are respectively connected to the convolutional layer and the CSP module in the first FPN network in the improved neck network, the CSP module in the second FPN network and the CSP module in the third FPN network, wherein the feature fusion layer is used to perform feature fusion on the input feature map, and the detection head is used to perform pedestrian detection on the fused feature map output by each of the feature fusion layers; An acquisition unit, used for acquiring an infrared pedestrian training set, wherein the infrared pedestrian training set includes a plurality of infrared pedestrian images; A training unit, used for training the improved YOLOv5 network through the infrared pedestrian training set until the loss value of the improved YOLOv5 network meets a preset condition, thereby obtaining a pedestrian detection model; The detection unit is used to perform pedestrian detection on the infrared pedestrian image to be detected by using the pedestrian detection model to obtain a pedestrian detection result.
6. The pedestrian detection device according to claim 5, It is characterized in that The backbone network includes a Focus layer, a convolutional layer, three CSP modules and an SPP layer. The other end of the cross-layer splicing layer in the first FPN network is connected to the third CSP module in the backbone network, the other end of the cross-layer splicing layer in the second FPN network is connected to the second CSP module in the backbone network, and the other end of the cross-layer splicing layer in the third FPN network is connected to the first CSP module in the backbone network.
7. The pedestrian detection device according to claim 5, It is characterized in that The depth parameter depth_multiple of the backbone network is set to 0.33, and the width parameter width_multiple is set to 0.
5.
8. The pedestrian detection device according to claim 5, It is characterized in that The number of the feature fusion layers is 4.
9. A pedestrian detection device, It is characterized in that The device comprises a processor and a memory; The memory is used to store program codes and transmit the program codes to the processor; The processor is used to execute the pedestrian detection method according to any one of claims 1-4 according to the instructions in the program code.
10. A computer-readable storage medium, It is characterized in that The computer-readable storage medium is used to store program codes, and when the program codes are executed by a processor, the pedestrian detection method according to any one of claims 1 to 4 is implemented.
Citation Information
Patent Citations
Face detection model training method, face detection method and related devices thereof
CN113128413A
Micropapilla detection system based on YOLOv5
CN113344849A