Pedestrian Target Detection Method and System
By introducing a multi-scale feature extraction module and an ExtremeNet detector in pedestrian target detection, the problems of low pedestrian target detection accuracy and occlusion overlap in complex environments are solved, and more accurate detection results are achieved.
Patent Information
- Application Number
- CN202111341424.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-12
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2041-11-12
AI Technical Summary
The prior art has problems such as low accuracy, detection errors caused by occlusion and overlap in complex environments.
The multi-scale feature extraction module is used to enhance the pedestrian re-identification network, and combined with the ExtremeNet detector, the detection accuracy is improved through diversified feature extraction and extreme point detection methods.
It significantly improves the accuracy of pedestrian target detection, effectively solves the occlusion and overlap problems in complex scenarios, and makes the detection more accurate.
Smart Images

Figure CN114332908B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of object re-identification, and particularly to a pedestrian object detection method and system. Background Art
[0002] The statements in this section merely mention the background art related to the present invention and do not necessarily constitute prior art.
[0003] Object detection technology is a comprehensive technology of computer technology, pattern recognition technology, image processing technology, artificial intelligence, etc. It is applied to computer vision fields such as motion analysis, video surveillance, video compression, intelligent transportation, human-computer interaction, and information security. Although object tracking technology has been improved in recent years, due to factors such as complex environmental occlusion, deformation, illumination change, scale change, and fast movement of objects in actual application scenarios, long-term visual tracking is still challenging.
[0004] Achieving efficient and accurate object detection in complex scenes by video detection technology has always been one of the problems to be solved. Although current object detection has achieved good results in terms of accuracy, robustness, and speed, when the object is in a complex scene similar to object interference, a series of problems such as template drift will still occur.
[0005] The inventor found that the technical defects of existing moving object detection methods are:
[0006] 1) Low accuracy of pedestrian detection caused by complex environmental occlusion, deformation, illumination change, and scale change in actual application scenarios;
[0007] 2) Detection errors caused by pedestrian overlap, pedestrian position interchange, position crossing in the detector, and the disappearance and reappearance of pedestrians in the camera lens. Summary of the Invention
[0008] To solve the deficiencies of the prior art, the present invention provides a pedestrian object detection method and system;
[0009] In a first aspect, the present invention provides a pedestrian object detection method;
[0010] The pedestrian object detection method includes:
[0011] Obtain a video to be detected;
[0012] Perform pedestrian object detection on the video to be detected;
[0013] Perform pedestrian re-identification on the detected pedestrian objects to obtain pedestrian object detection results;
[0014] Among them, pedestrian re-identification is implemented using a trained pedestrian re-identification network; the pedestrian re-identification network incorporates a multi-scale feature extraction module to expand the receptive field and improve the ability to extract target features.
[0015] In a second aspect, the present invention provides a pedestrian target detection system;
[0016] The pedestrian target detection system includes:
[0017] An acquisition module, which is configured to: acquire a video to be detected;
[0018] A target detection module, which is configured to: perform pedestrian target detection on the video to be detected;
[0019] A pedestrian re-identification module, which is configured to: perform pedestrian re-identification on the detected pedestrian target to obtain a pedestrian target detection result;
[0020] Among them, pedestrian re-identification is implemented using a trained pedestrian re-identification network; the pedestrian re-identification network incorporates a multi-scale feature extraction module to expand the receptive field and improve the ability to extract target features.
[0021] In a third aspect, the present invention further provides an electronic device, including:
[0022] A memory for non-temporarily storing computer-readable instructions; and
[0023] A processor for running the computer-readable instructions,
[0024] Among them, when the computer-readable instructions are run by the processor, the method described in the first aspect above is executed.
[0025] In a fourth aspect, the present invention further provides a storage medium that non-temporarily stores computer-readable instructions, wherein when the non-temporary computer-readable instructions are executed by a computer, the instructions for executing the method described in the first aspect are executed.
[0026] In a fifth aspect, the present invention further provides a computer program product, including a computer program, where the computer program is used to implement the method described in the first aspect above when running on one or more processors.
[0027] Compared with the prior art, the beneficial effects of the present invention are:
[0028] Selecting ExtremeNet as our target detection network greatly improves the detection accuracy, and improving the pedestrian re-identification network by incorporating a multi-scale extraction module greatly expands the receptive field and improves the ability to extract target features, solving problems such as the target being occluded multiple times and target overlap, thereby making our detection more accurate.
[0029] The present invention aims to propose a pedestrian target detection algorithm based on multi-scale feature extraction. This algorithm uses a multi-scale feature extraction module to extract pedestrian target feature information from top to bottom. For problems such as complex environmental occlusion, deformation, scale change, and pedestrian overlap, an extreme point detection method is adopted for diversified extraction. And on the improved pedestrian re-identification network, we integrate the multi-scale feature extraction module to further enhance the accuracy of pedestrian target detection, enabling more diversified processing of target detection in complex scenarios and improving the accuracy of pedestrian target detection. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] The accompanying drawings forming a part of this invention are used to provide a further understanding of the invention. The schematic embodiments and descriptions thereof of the invention are used to explain the invention and do not constitute an improper limitation of the invention.
[0031] Figure 1 It is the flowchart of the method for the first embodiment;
[0032] Figure 2 It is the component of ExtremeNet for the first embodiment;
[0033] Figure 3 It is the feature extraction and basic block network for the first embodiment;
[0034] Figure 4 It is the network structure of the multi-scale module for the first embodiment;
[0035] Figure 5 It is the schematic diagram of post-detection of target overlap for the first embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0036] It should be noted that the following detailed description is exemplary and is intended to provide further explanation of the present invention. Unless otherwise specified, all technical and scientific terms used in the present invention have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs.
[0037] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present invention. As used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.
[0038] Without conflict, the embodiments in the present invention and the features in the embodiments may be combined with each other.
[0039] For all data acquisition in this embodiment, it is a legal application of data on the basis of compliance with laws, regulations and user consent.
[0040] Embodiment 1
[0041] This embodiment provides a pedestrian target detection method;
[0042] As Figure 1 shown, the pedestrian target detection method includes:
[0043] S101: Obtain the video to be detected;
[0044] S102: Perform pedestrian target detection on the video to be detected;
[0045] S103: Perform pedestrian re-identification on the detected pedestrian targets to obtain pedestrian target detection results;
[0046] Among them, pedestrian re-identification is implemented by using a trained pedestrian re-identification network;
[0047] The pedestrian re-identification network incorporates a multi-scale feature extraction module to expand the receptive field and improve the ability to extract target features.
[0048] The flowchart of the pedestrian target detection system of the present invention is as Figure 1 shown. First, input the pedestrian image sequence. After the pedestrian video sequence is processed by the ExtremeNet detector, the location of the target can be quickly determined. Then, through our improved pedestrian re-detection mechanism, the target features can be efficiently extracted, so that the pedestrian targets in each frame of the image can be detected more, and the detection results are more accurate.
[0049] Further, in S102: performing pedestrian target detection on the video to be detected is implemented by using a trained ExtremeNet detector.
[0050] Further, for the trained ExtremeNet detector, its training steps include:
[0051] Construct a first training set; the first training set includes: images with five extreme points of known pedestrian target detection;
[0052] Input the first training set into the ExtremeNet detector. When the loss value of the loss function no longer decreases, stop training to obtain the trained ExtremeNet detector.
[0053] Further, in step S102, perform pedestrian target detection on the video to be detected, and detect the four extreme points of the target, namely the upper, lower, left, and right extreme points, and the central extreme point of the target respectively;
[0054] After detecting all the extreme points, combine all the extreme points. After obtaining all the combinations, screen the combinations by verifying whether there is a central extreme point in the combination:
[0055] If there is a central extreme point in the current combination, a target bounding box is obtained;
[0056] If there is no central extreme point in the current combination, it means that the current target bounding box is an incorrect box.
[0057] ExtremeNet makes further improvements on the basis of the idea of CornerNet. Instead of detecting the upper left corner point and the lower right corner point of the target, it detects 4 extreme points of the target, namely the uppermost point, the lowermost point, the leftmost point, the rightmost point and a center point. The extreme points of ExtremeNet are on the target, are visually separable, and have consistent local appearance features.
[0058] Therefore, the detection method of ExtremeNet we adopted is end-to-end differentiable, simpler, faster and more accurate compared with the detector based on rectangular bounding boxes. Our model achieves the best trade-off between speed and accuracy.
[0059] As Figure 2 shown, the backbone network of ExtremeNet is the Hourglass Network. The network first detects the 4 extreme points of the upper, lower, left, and right respectively through 4 heatmaps of the upper, lower, left, and right, and detects the central extreme point through the Center heatmap. After detecting all the extreme points, use the matching algorithm to combine all the extreme points. After obtaining all the combinations, screen all the combination examples by verifying whether there is a central extreme point in the combination. If there is a central extreme point, a target is obtained. If there is no central extreme point, the target bounding box is an incorrect box.
[0060] Further, the training steps of the trained person re-identification network include:
[0061] Construct a person re-identification network;
[0062] Construct a second training set; the second training set includes: images with known pedestrian target detection results;
[0063] Input the second training set into the person re-identification network. When the loss value of the loss function no longer decreases, stop training to obtain the trained person re-identification network.
[0064] Further, as Figure 3 shown, the person re-identification network has a network structure including:
[0065] An input layer, a convolutional layer a1, a max pooling layer, a first residual module, a second residual module, a third residual module, a fourth residual module, an average pooling layer, an adder J1, and an output layer connected in sequence;
[0066] Among them, the output end of the second residual module is connected to the input end of the adder J1 through a multi-scale feature extraction module and a convolutional layer a2 connected in sequence.
[0067] Further, the working principle of the person re-identification network includes:
[0068] The convolutional layer a1 extracts the required picture features; the max pooling layer reduces the picture dimension and retains more texture information; the first residual module, the second residual module, the third residual module, and the fourth residual module change the picture size in sequence, accelerate convergence, and stabilize the performance; the average pooling layer reduces the size of the spatial information and improves the operation efficiency.
[0069] The input image first enters a 3×3 convolution and a max pooling layer to reduce the vector dimension, then passes through two residual blocks. When reaching the second residual block, it is divided into two branches. One serves as the input of the multi-scale module and enters a 3×3 convolution, and the other is input into two residual blocks and an average pooling layer again. Then the outputs of the two branches are added, and finally the extracted appearance features are output.
[0070] Further, the internal structures of the first residual module, the second residual module, the third residual module, and the fourth residual module are the same.
[0071] Further, as Figure 3 shown, the first residual module includes:
[0072] An input unit, a convolutional layer b1, a convolutional layer b2, an adder J2, and an activation function layer connected in sequence;
[0073] Among them, the input unit is also connected to the input end of the adder J2.
[0074] Further, the working principle of the first residual module includes:
[0075] The convolutional layer b1 and the convolutional layer b2 change the picture dimension, use a skip connection to add the input and the outputs of the two convolutions to avoid gradient disappearance or gradient explosion, and then activate through the relu activation function.
[0076] Further, as Figure 4 shown, the multi-scale feature extraction module includes:
[0077] The input end is respectively connected to four branches;
[0078] Among them, the first branch includes: a convolutional layer c1, a convolutional layer c2, and a convolutional layer c3 connected in sequence;
[0079] Among them, the second branch includes: a convolutional layer d1, a convolutional layer d2, and a convolutional layer d3 connected in sequence;
[0080] Among them, the third branch includes: a convolutional layer e1, a convolutional layer e2, and a convolutional layer e3 connected in sequence;
[0081] Among them, the fourth branch includes: a convolutional layer f1, a convolutional layer f2, and a convolutional layer f3 connected in sequence;
[0082] Among them, the output ends of the convolutional layer c3 and the convolutional layer d3 are both connected to the input end of the first connector; the output end of the first connector is connected to the input end of the convolutional layer g1; the working principle of the first connector is series connection;
[0083] Among them, the output ends of the convolutional layer e3 and the convolutional layer f3 are both connected to the input end of the second connector; the output end of the second connector is connected to the input end of the convolutional layer g2; the working principle of the second connector is series connection;
[0084] The output end of the convolutional layer g1 is connected to the input end of the adder J3; the output end of the convolutional layer g2 is connected to the input end of the adder J3; the output end of the convolutional layer f1 is connected to the input end of the adder J3 through a shortcut; the output end of the adder J3 is connected to the activation function layer.
[0085] Shortcut is a spanning method to solve the gradient divergence in deep networks.
[0086] Furthermore, as Figure 4 shown, the multi-scale feature extraction module, the working principle includes: using dilated convolution to expand the receptive field while keeping the size of the feature map unchanged, extracting multi-scale information, and making the detection more accurate.
[0087] The network structure of the multi-scale module is as Figure 4As shown in the figure. It uses a 3×3 convolutional layer instead of a 5×5 convolutional layer, and uses 2×2, 4×4, 6×6, and 8×8 convolutional layers. The main purpose is to reduce the computational complexity. The network structure also includes four 1×1 convolutions and four 3×3 dilated convolutions. In a general CNN, the receptive field of each layer is fixed, which will lose some information and the ability to distinguish different fields of view. Therefore, this module uses dilated convolutional kernels. Different rates correspond to different sizes of holes. The larger the rate, the larger the hole size, the farther the sampling point is from the center point, and the larger the receptive field. Finally, the outputs of convolutional layers with different sizes and rates are concatenated to achieve the purpose of fusing different features. And its output is concatenated with the output of the shortcut, and through the activation function, the feature map is output.
[0088] For deep convolutional neural networks, the deeper the network, the better its performance. However, there are also certain problems in training deep convolutional neural networks, such as vanishing gradients and dispersion. Blindly deepening the network structure model does not bring obvious performance improvement, but instead requires a large amount of computing resources to support. Therefore, we combined a multi-scale feature extraction module in this network structure, aiming to enhance the discriminability of features and the robustness of the model, and extract larger regional features while maintaining the same number of parameters.
[0089] In the object detection evaluation, the VOC2012 dataset is used for training, and the present invention uses the VOC2012 trainval dataset. From Figure 5 From the detection results, it can be seen that our improved algorithm can well detect and distinguish occluded and overlapping objects. The commonly used evaluation index for object detection is mAP, which represents the mean average precision. The present invention uses mAP and the single-image detection time index to evaluate the object detection network. From the data in Table 1, it can be seen that our improved detection algorithm is more accurate.
[0090] Table 1 Comparison of experimental results on the VOC2012 dataset
[0091]
[0092] Example 2
[0093] This example provides a pedestrian object detection system;
[0094] The pedestrian object detection system includes:
[0095] An acquisition module, which is configured to: acquire the video to be detected;
[0096] An object detection module, which is configured to: perform pedestrian object detection on the video to be detected;
[0097] A pedestrian re-identification module, which is configured to: perform pedestrian re-identification on the detected pedestrian target to obtain a pedestrian target detection result;
[0098] Among them, pedestrian re-identification is implemented by using a trained pedestrian re-identification network; the pedestrian re-identification network incorporates a multi-scale feature extraction module to expand the receptive field and improve the ability to extract target features.
[0099] It should be noted here that the above acquisition module, target detection module, and pedestrian re-identification module correspond to steps S101 to S103 in Embodiment 1. The examples and application scenarios implemented by the above modules and the corresponding steps are the same, but are not limited to the content disclosed in Embodiment 1 above. It should be noted that the above modules, as part of the system, can be executed in a computer system such as a set of computer-executable instructions.
[0100] In the above embodiments, the descriptions of each embodiment have their own emphases. For parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0101] The proposed system can be implemented in other ways. For example, the system embodiments described above are only illustrative. For example, the above module division is only a logical function division. In actual implementation, there can be other division methods. For example, multiple modules can be combined or integrated into another system, or some features can be ignored or not executed.
[0102] Embodiment 3
[0103] This embodiment also provides an electronic device, including: one or more processors, one or more memories, and one or more computer programs; wherein, the processor is connected to the memory, and the above one or more computer programs are stored in the memory. When the electronic device runs, the processor executes the one or more computer programs stored in the memory so that the electronic device executes the method described in Embodiment 1 above.
[0104] It should be understood that in this embodiment, the processor can be a central processing unit CPU, and the processor can also be other general-purpose processors, digital signal processors DSP, application-specific integrated circuits ASIC, off-the-shelf programmable gate arrays FPGA, or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc.
[0105] The memory can include a read-only memory and a random access memory, and provide instructions and data to the processor. A part of the memory can also include a non-volatile random memory. For example, the memory can also store information about the device type.
[0106] In the implementation process, each step of the above method can be completed by the integrated logic circuit of the hardware in the processor or the instructions in the form of software.
[0107] The method in Embodiment 1 can be directly embodied as being executed and completed by the hardware processor, or executed and completed by the combination of the hardware and software modules in the processor. The software module can be located in a mature storage medium in the art such as random access memory, flash memory, read-only memory, programmable read-only memory, or electrically erasable programmable memory, register, etc. This storage medium is located in the memory, and the processor reads the information in the memory and combines its hardware to complete the steps of the above method. To avoid repetition, it will not be described in detail here.
[0108] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with this embodiment can be implemented by electronic hardware or the combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.
[0109] Embodiment 4
[0110] This embodiment also provides a computer-readable storage medium for storing computer instructions. When the computer instructions are executed by a processor, the method described in Embodiment 1 is completed.
[0111] The above are only the preferred embodiments of the present invention and are not used to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. Pedestrian target detection method, Characterized in that, Comprising: Obtain the video to be detected; Perform pedestrian target detection on the video to be detected; Perform pedestrian re-identification on the detected pedestrian target to obtain the pedestrian target detection result; Among them, pedestrian re-identification is implemented by using a trained pedestrian re-identification network; the pedestrian re-identification network incorporates a multi-scale feature extraction module to expand the receptive field and improve the ability to extract target features; The multi-scale feature extraction module includes: An input end, and the input end is respectively connected to four branches; Among them, the first branch includes: a 1×1 convolutional layer c1, a 2×2 convolutional layer c2, and a 3×3 convolutional layer c3 with a dilation rate of 2 connected in sequence; Among them, the second branch includes: a 1×1 convolutional layer d1, a 4×4 convolutional layer d2, and a 3×3 convolutional layer d3 with a dilation rate of 4 connected in sequence; Among them, the third branch includes: a 1×1 convolutional layer e1, a 6×6 convolutional layer e2, and a 3×3 convolutional layer e3 with a dilation rate of 6 connected in sequence; Among them, the fourth branch includes: a 1×1 convolutional layer f1, an 8×8 convolutional layer f2, and a 3×3 convolutional layer f3 with a dilation rate of 8 connected in sequence; Among them, the output end of convolutional layer c3 and the output end of convolutional layer d3 are both connected to the input end of the first connector; the output end of the first connector is connected to the input end of the 1×1 convolutional layer g1; the working principle of the first connector is series connection; Among them, the output end of convolutional layer e3 and the output end of convolutional layer f3 are both connected to the input end of the second connector; the output end of the second connector is connected to the input end of the 1×1 convolutional layer g2; the working principle of the second connector is series connection; The output end of convolutional layer g1 is connected to the input end of adder J3; the output end of convolutional layer g2 is connected to the input end of adder J3; the output end of convolutional layer f1 is connected to the input end of adder J3 through a shortcut; the output end of adder J3 is connected to the activation function layer; The working principle of the multi-scale feature extraction module includes: using dilated convolution to expand the receptive field while keeping the size of the feature map unchanged, extracting multi-scale information, and making the detection more accurate.
2. The pedestrian target detection method according to claim 1, Characterized in that, Performing pedestrian target detection on the video to be detected specifically includes: Detect the upper, lower, left, and right four extreme points of the target and the central extreme point of the target respectively; After detecting all the extreme points, combine all the extreme points. After obtaining all the combinations, screen the combinations by verifying whether there is a central extreme point in the combination: If there is a central extreme point in the current combination, a target box is obtained; If there is no central extreme point in the current combination, it means that the current target box is an incorrect box.
3. The pedestrian target detection method according to claim 1, Characterized in that, The network structure of the pedestrian re-identification network includes: An input layer, a convolutional layer a1, a max pooling layer, a first residual module, a second residual module, a third residual module, a fourth residual module, an average pooling layer, an adder J1, and an output layer connected in sequence; Among them, the output end of the second residual module is connected to the input end of the adder J1 through a multi-scale feature extraction module and a convolutional layer a2 connected in sequence; The working principle of the pedestrian re-identification network includes: The convolutional layer a1 extracts the required image features; the max pooling layer reduces the image dimension and retains more texture information; the first residual module, the second residual module, the third residual module, and the fourth residual module change the image size in sequence to accelerate convergence and stabilize performance; the average pooling layer reduces the size of spatial information and improves the operation efficiency.
4. The pedestrian target detection method according to claim 3, wherein, The first residual module includes: An input unit, a convolutional layer b1, a convolutional layer b2, an adder J2, and an activation function layer connected in sequence; Among them, the input unit is also connected to the input end of the adder J2; The working principle of the first residual module includes: The convolutional layer b1 and the convolutional layer b2 change the image dimension, add the input and the outputs of the two convolutions with a skip connection to avoid gradient disappearance or gradient explosion, and then activate through the relu activation function.
5. The pedestrian target detection method according to claim 1, wherein, Perform pedestrian target detection on the video to be detected, and use the trained ExtremeNet detector to implement it; The trained ExtremeNet detector, Its training steps include: Construct a first training set; the first training set includes images with five extreme points of known pedestrian target detection; Input the first training set into the ExtremeNet detector, and stop training when the loss value of the loss function no longer decreases to obtain the trained ExtremeNet detector.
6. The pedestrian target detection method according to claim 1, wherein, The training steps of the trained pedestrian re-identification network include: Construct a pedestrian re-identification network; Construct a second training set; the second training set includes images with known pedestrian target detection results; Input the second training set into the pedestrian re-identification network, and stop training when the loss value of the loss function no longer decreases to obtain the trained pedestrian re-identification network.
7. A pedestrian target detection system, wherein, It includes: An acquisition module configured to acquire the video to be detected; A target detection module configured to perform pedestrian target detection on the video to be detected; A pedestrian re-identification module configured to perform pedestrian re-identification on the detected pedestrian target to obtain a pedestrian target detection result; Among them, pedestrian re-identification is implemented using the trained pedestrian re-identification network; the pedestrian re-identification network incorporates a multi-scale feature extraction module to expand the receptive field and improve the ability to extract target features; The multi-scale feature extraction module includes: An input end, and the input end is respectively connected to four branches; Among them, the first branch includes a 1×1 convolutional layer c1, a 2×2 convolutional layer c2, and a 3×3 convolutional layer c3 with a dilation rate of 2 connected in sequence; Among them, the second branch includes a 1×1 convolutional layer d1, a 4×4 convolutional layer d2, and a 3×3 convolutional layer d3 with a dilation rate of 4 connected in sequence; Among them, the third branch includes: a 1×1 convolutional layer e1, a 6×6 convolutional layer e2, and a 3×3 convolutional layer e3 with a dilation rate of 6, which are connected in sequence; Among them, the fourth branch includes: a 1×1 convolutional layer f1, an 8×8 convolutional layer f2, and a 3×3 convolutional layer f3 with a dilation rate of 8, which are connected in sequence; Among them, the output ends of the convolutional layer c3 and the convolutional layer d3 are both connected to the input end of the first connector; the output end of the first connector is connected to the input end of the 1×1 convolutional layer g1; the working principle of the first connector is series connection; Among them, the output ends of the convolutional layer e3 and the convolutional layer f3 are both connected to the input end of the second connector; the output end of the second connector is connected to the input end of the 1×1 convolutional layer g2; the working principle of the second connector is series connection; The output end of the convolutional layer g1 is connected to the input end of the adder J3; the output end of the convolutional layer g2 is connected to the input end of the adder J3; the output end of the convolutional layer f1 is connected to the input end of the adder J3 through a shortcut; the output end of the adder J3 is connected to the activation function layer; The multi-scale feature extraction module has a working principle that includes: using dilated convolution to expand the receptive field without changing the size of the feature map, extracting multi-scale information, and making the detection more accurate.
8. An electronic device, characterized in that it includes: a memory for non-temporarily storing computer-readable instructions; and a processor for running the computer-readable instructions, wherein, when the computer-readable instructions are run by the processor, the method according to any one of claims 1-6 above is executed.
9. A storage medium, characterized in that it non-temporarily stores computer-readable instructions, wherein, when the non-temporary computer-readable instructions are executed by a computer, instructions for executing the method according to any one of claims 1-6 are executed.
Citation Information
Patent Citations
Pedestrian re-identification method based on improved YOLOv3 network and feature fusion
CN111783576A