Lightweight pedestrian detection method
By introducing the C2f_GhostDynamicConv module and GSConv module into the YOLOv8 network, a lightweight pedestrian detection model was built, which solved the problem of large parameters and high computational complexity of the YOLOv8 model, and achieved faster detection speed and lower computational volume.
Patent Information
- Application Number
- CN202510550968.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-29
- Publication Date
- 2025-08-12
AI Technical Summary
The existing YOLOv8 single-stage object detection technology has problems such as large model parameters and high computational complexity in pedestrian detection, which is difficult to meet the end-to-end industrial deployment needs.
The C2f_GhostDynamicConv module is used to replace the C2f module of YOLOv8, and the lightweight convolution method GSConv is introduced to build a lightweight pedestrian detection model, reducing the number of parameters and speeding up the training speed.
While keeping the detection accuracy unchanged, the number of parameters decreased by about 39.1%, and the calculation speed increased by about 66.3%, solving the real-time problem of intensive pedestrian detection.
Smart Images

Figure CN120472424A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of computer vision deep learning and image processing, and in particular to a lightweight pedestrian detection method. Background Art
[0002] Pedestrian detection is an important and actively researched area in computer vision. It primarily uses advanced computer vision algorithms to identify and recognize pedestrians in images or videos. It has important applications in traffic monitoring systems and autonomous driving. As part of a broad range of object detection tasks, pedestrian detection has attracted significant research attention. Pedestrian detection holds great potential in applications such as behavior recognition, person re-identification, and computer vision tracking.
[0003] Pedestrian detection algorithms based on deep learning can generally be divided into two categories based on their algorithmic process: two-stage detection algorithms and single-stage detection algorithms. Two-stage pedestrian detection algorithms first generate a region of interest and then use a convolutional neural network to classify samples. Single-stage pedestrian detection algorithms directly predict the category and bounding box coordinates of each object in a single run. Mainly including the Single Shot MultiBox Detector (SSD) and the You Only LookOnce (YOLO) series. Currently, single-stage object detection technologies, represented by YOLOv8, have achieved significant optimizations in the object extraction network. However, the object fusion network fails to efficiently integrate contextual information, resulting in large model parameters and high computational complexity in pedestrian detection. There is still room for improvement to meet end-to-end industrial deployment requirements.
[0004] Therefore, the present invention designs a lightweight pedestrian detection method, which improves the pedestrian detection performance while making lightweight improvements to meet end-to-end industrial deployment, which is a problem that technical personnel in this field urgently need to solve. Summary of the Invention
[0005] The purpose of this paper is to propose a lightweight pedestrian detection method, which achieves a balance between model lightweight and performance by improving the YOLOv8 structure and improving the accuracy of pedestrian detection results.
[0006] The technical solution of the present invention is to propose a C2f_GhostDynamicConv module with a smaller number of parameters to replace the C2f module of YOLOv8 based on the YOLOv8 network, thereby greatly reducing the number of parameters, reducing triggers and operations, and speeding up the training speed. In addition, a new lightweight convolution method GSConv is proposed to replace part of the Conv module to further construct a lightweight network structure. On this basis, the present invention constructs a lightweight pedestrian detection model based on YOLOv8. Compared with the existing YOLOv8 network, the YOLOv8-GS network proposed in the present invention maintains the same detection accuracy, reduces the number of parameters by about 39.1%, and increases the calculation speed by about 66.3%. The flow chart of the method of the present invention is as follows Figure 1 As shown, the specific implementation steps are as follows:
[0007] The first step is to use the corresponding database to load the pedestrian detection data, the original image data I∈R H×W×3 , H and W are the length and width of the original image respectively, 3 is the number of color channels, the training network batch size is set to 16, a total of 100 iterations are set, the initial number of iterations is 1, the learning rate is 0.01, and the optimizer selects SDG optimizer;
[0008] The second step is to preprocess the image and convert the input image into the YOLO training format. The preprocessing includes 1) image scaling: scaling the image to the model input size; 2) normalization: normalizing the pixel values to the range [0, 1]; 3) channel conversion: converting the image from BGR format to RGB format and dividing it into training and test sets. In this way, the input feature image input data I∈R H×W×3 ;
[0009] The third step is to input the original image data into the feature extraction network, such as Figure 2 As shown:
[0010] (3.1) The input size of the feature extraction network is I∈R H×W×3 The image passes through the first convolutional layer with a kernel size of 3, a stride of 2, and a padding of 1, and then the output size is The feature map with 64 channels is transferred to the second convolutional layer. The convolutional layer parameter settings are the same as the first layer, and the size is obtained. Feature map with 128 channels;
[0011] (3.2) is transferred to the first C2f_GhostDynamicConv module with 2 convolutional layers and n Bottleneck_GhostDynamicConv modules composed of Ghost modules and dynamic convolutions (such as Figure 3As shown in Figure 2.2). The C2f_GhostDynamicConv module first passes the input data through the first convolution layer Conv with a kernel size of 2 and a stride of 1, and then divides the output into two parts. One part is directly passed to the output, and the other part passes through n Bottleneck_GhostDynamicConv modules. Figure 3 As shown in Figure 2, Bottleneck_GhostDynamicConv implemented using dynamic convolution technology can introduce more parameters to a large extent while minimizing the increase in computational complexity. In the Bottleneck_GhostDynamicConv module, the input feature X is first convolved by dynamic convolution to generate an intrinsic feature map Y of a fixed number of intrinsic feature maps X. Assume that the input feature X∈R h×w×3 , where h and w are the length and width of the original image respectively, 3 is the number of color channels, and the intrinsic feature map Y created is expressed as follows:
[0012]
[0013] Among them, F i is the i-th convolution weight tensor, * represents convolution, α i is the corresponding dynamic coefficient, where The input feature map X is converted into a vector using global average pooling “Pool(·)”, and then a two-layer MLP module “MLP(·)” with SoftMax “SoftMax(·)” activation is applied to dynamically generate coefficients.
[0014] Using the created intrinsic feature map Y, the related feature Z is calculated through a series of linear operations. The expression is as follows:
[0015]
[0016] where Y i ′ is the i-th feature of the intrinsic feature map Y, Φ ij It is a linear operation to generate the jth related feature, s and t represent the number of i and j respectively. The feature information Y generated by the linear operation ij Connected to the inherent features Y, Y and Y ij The results of the two parts are concatenated in the channel dimension to obtain feature Z.
[0017] The feature Z and the features directly output by the previous convolutional layer are partially processed by the aggregation module and passed through the second convolutional layer Conv with a convolution kernel size of 2 and a stride of 1 to obtain the final output. The size and number of channels of the obtained feature map remain unchanged and are transmitted to the first GSConv convolutional layer;
[0018] (3.3) The feature map is input into the first GSConv convolution layer, and the output size is The feature map has 256 channels. Figure 5 As shown in the figure, the workflow is as follows: the input of the GSConv module undergoes a 1×1 standard convolution to change the number of channels. The depthwise separable convolution module performs a 1×1 convolution on each input channel separately. The number of channels remains unchanged after processing. The result of the first convolution is then concatenated with the structure after the depthwise separable convolution to fuse the features of different receptive fields and enhance the expression ability. The concatenated result is shuffled (Channel Shuffle) to rearrange different channels and promote cross-channel information interaction. Assuming that the input feature map is X and the output feature map is Y, the expression is as follows:
[0019] Y=ChannelShuffle(Concat(Conv(X),DWConv(X))), (3)
[0020] Among them, "Conv(·)" is a standard convolution, "DWConv(·)" represents a depthwise separable convolution, "Concat(·)" represents a concatenation operation, and "ChannelShuffle(·)" represents a shuffling operation. After the first GSConv convolution layer, the data has two directions: one is transmitted to the feature fusion network, and the other is transmitted to the second C2f_GhostDynamicConv module consistent with step (3.2);
[0021] (3.4) The size of the output of the second C2f_GhostDynamicConv module is The feature map with 256 channels is input to the second GSConv module consistent with step (3.3);
[0022] (3.5) After the second GSConv module, the output size is The feature map with 512 channels is input to the third C2f_GhostDynamicConv module consistent with (2); the output size of the third C2f_GhostDynamicConv module is The feature map with 512 channels also has two directions: on the one hand, it is transmitted to the third GSConv module consistent with step (3.3); on the other hand, it is transmitted to the feature fusion network;
[0023] (3.6) is input into the third GSConv module consistent with step (3.3) and the output size is The feature map with 512 channels is output to the fourth C2f_GhostDynamicConv module;
[0024] (3.7) After passing through the fourth C2f_GhostDynamicConv module, the feature map size and number of channels remain unchanged. It is then processed by the Spatial Pyramid Pooling-Fast (SPPF) module to increase the receptive field, and the result is finally passed to the feature fusion network. The SPPF module pools the feature map to a fixed size to increase the diversity of feature expression.
[0025] The fourth step is to input the feature map into the feature fusion network, such as Figure 2 As shown:
[0026] (4.1) The feature fusion network input is the size of the output of the spatial pyramid fast pooling SPPF module of the feature extraction network. The feature map with a channel number of 512 is input to the first convolutional layer. The convolution kernel size of the convolution layer is 3, the stride is 2, and the boundary expansion is 1. After the first convolutional layer, the output size is The feature map with a channel number of 512 is transmitted to the first upsampling module on the one hand, and to the fourth aggregation module on the other hand;
[0027] (4.2) After the first upsampling module, the output size is The feature map with 512 channels is transmitted to the first aggregation module; the input of the first aggregation module consists of the output of the second C2f_GhostDynamicConv module of the feature extraction network and the output of the first upsampling module of the feature fusion network. The output size after aggregation is The feature map with 512 channels is transmitted to the first aggregation module;
[0028] (4.3) The output size after aggregation is The feature map with 512 channels is input to the first C2f_GhostDynamicConv module of the feature fusion network, which is consistent with step (3.2). The size and number of channels of the feature map after the first C2f_GhostDynamicConv module remain unchanged and are transmitted to the second convolutional layer.
[0029] (4.4) After the second convolutional layer, the feature map has unchanged size and number of channels. It is input to the second upsampling module and transmitted to the third aggregation module.
[0030] (4.5) After the second upsampling module, the output size is The feature map with 512 channels is output to the second aggregation module; the input of the aggregation module consists of the output of the first GSConv module of the feature extraction network and the output of the second upsampling module of the feature fusion network. The output size after aggregation is The feature map with 768 channels is output to the second C2f_GhostDynamicConv module consistent with step (3.2);
[0031] (4.6) After the second C2f_GhostDynamicConv module, the output size is The feature map with 256 channels is output to the detection head network, and the output size is The feature map with 256 channels is output to the second GSConv module consistent with step (3.3);
[0032] (4.7) After the second GSConv module, the output size is The feature map with 256 channels is output to the third aggregation module; the third aggregation module aggregates the first GSConv module and the first convolution module and outputs a size of The feature map with 768 channels is output to the third C2f_GhostDynamicConv module consistent with step (3.2);
[0033] (4.8) After the third C2f_GhostDynamicConv module, the output size is The feature map with 512 channels is output to the detection head network, and the output size is The feature map with 512 channels is output to the third GSConv module consistent with step (3.3);
[0034] (4.9) After the third GSConv module, the output size is The feature map with 512 channels is output to the fourth aggregation module; the third aggregation module aggregates the second GSConv module and the first convolution module and outputs a size of The feature map with 512 channels is output to the fourth C2f_GhostDynamicConv module consistent with step (3.2);
[0035] (4.10) After the fourth C2f_GhostDynamicConv module, the output size is The feature map with 512 channels is output to the detection head network.
[0036] The fifth step is to input the features output by the feature fusion network into the detection head network to output the bounding box, category and confidence of the target; the detection head network receives the feature maps of three different scales from the feature fusion network. They are used to detect small, medium and large objects respectively, and are characterized by:
[0037] The detection head network adopts a decoupled head and anchor-free design. It first processes the multi-scale feature maps from the feature fusion network through a shared convolutional layer, and then separates it into classification and regression branches. The classification branch outputs category probabilities (sigmoid activation, supporting multi-label) and outputs the category probability of each anchor point. The regression branch directly predicts the center offset, width and height of the bounding box and outputs the coordinate parameters of the bounding box. During training, the dynamic label Task-AlignedAssigner is used to assign matching positive samples, and the classification loss, regression loss, and confidence loss are jointly optimized. During inference, the prediction results are filtered by confidence threshold and non-maximum suppression, and the coordinates, category, and confidence of the detection box are finally output.
[0038] Step 6: Obtain the gradient based on the loss, back-propagate and update the optimizer parameters. Repeat steps 4 to 8 until the number of iterations reaches the maximum, and the optimal pedestrian detection model is obtained.
[0039] Step 7: Input the test set into the optimal pedestrian detection model and output the pedestrian detection result.
[0040] This paper proposes a lightweight pedestrian detection method based on an improved YOLOv8. While maintaining a certain level of accuracy and robustness in dense pedestrian scenes, the algorithm's speed is improved, solving the problem of real-time pedestrian detection in dense crowds. For image feature extraction and feature fusion, the algorithm introduces a lightweight feature extraction network and a feature fusion network. Experimental results on a dataset show that compared to the original YOLOv8, the improved model has 39.0% fewer parameters and 40.0% less computation. While maintaining the same detection accuracy as the original model, the detector's speed is significantly improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] Figure 1 This is a flowchart of the lightweight pedestrian detection method and system based on the improved YOLOv8 proposed in the present invention.
[0042] Figure 2 This is the improved YOLOv8 network structure diagram of the present invention.
[0043] Figure 3 This is the structure diagram of the C2f_GhostDynamicConv module.
[0044] Figure 4 This is the structure diagram of the Bottleneck_GhostDynamicConv module.
[0045] Figure 5 This is the GSConv module structure diagram. DETAILED DESCRIPTION
[0046] A specific embodiment of the present invention is described in detail below in conjunction with the technical solution and the accompanying drawings.
[0047] This article uses a Windows 11 64-bit operating system, an NVIDIA GeForce RTX4090 GPU, and trains the model using the PyTorch 3.11 deep learning framework based on CUDA 12.1 and CUDNN 8.9.4. This article uses the YOLOv8 network and optimizes the network. The batch size is set to 16, with a total of 100 iterations, an initial iteration of 1, a learning rate of 0.01, and the SDG optimizer. The input image size is 640*640. Figure 1 As shown in FIG, a lightweight pedestrian detection method based on an improvement of YOLOv8 includes the following steps:
[0048] The first step is to use the corresponding database to load the pedestrian detection data, the original image data I∈R H×W×3 , H and W are the length and width of the original image respectively, and 3 is the number of color channels;
[0049] The second step is to preprocess the image and convert the input image into the YOLO training format. The preprocessing includes 1) image scaling: scaling the image to the model input size; 2) normalization: normalizing the pixel values to the range [0, 1]; 3. channel conversion: converting the image from BGR format to RGB format and dividing it into training and test sets. In this way, the input feature image input data I∈R 640×640×3 To Figure 2 In the detection network shown;
[0050] The third step is to input the original image data into the feature extraction network:
[0051] (3.1) The input size of the feature extraction network is I∈R 640×640×3 After the image passes through the first convolution layer, the output feature map with a size of 320×320 and 64 channels is transmitted to the second convolution layer;
[0052] (3.2) After the second convolutional layer, the output feature map is 160×160 in size and 128 in number of channels, which is passed to the C2f_GhostDynamicConv module. The C2f_GhostDynamicConv module structure is shown in the figure below. Figure 3 As shown in Figure 2; the C2f_GhostDynamicConv module consists of two convolutional layers and n Bottleneck_GhostDynamicConv modules consisting of Ghost modules and dynamic convolutions. The Bottleneck_GhostDynamicConv module structure is shown in Figure 2. Figure 4As shown in Figure 2, the size and number of channels of the feature map after passing through the first C2f_GhostDynamicConv module remain unchanged and are transmitted to the first GSConv convolutional layer;
[0053] (3.3) After the first GSConv convolution layer, the output feature map with a size of 80×80 and a channel number of 256 has two directions: one is transmitted to the feature fusion network, and the other is transmitted to the second C2f_GhostDynamicConv module. The GSConv convolution layer structure is shown in the figure below. Figure 5 As shown;
[0054] (3.4) The feature map output by the second C2f_GhostDynamicConv module with a size of 80×80 and a channel number of 256 is input to the second GSConv module;
[0055] (3.5) After the second GSConv module, the output feature map with a size of 40×40 and a channel number of 512 is input to the third C2f_GhostDynamicConv module; the third C2f_GhostDynamicConv module outputs a feature map with a size of 40×40 and a channel number of 512. It also has two directions: one is transmitted to the third GSConv module, and the other is transmitted to the feature fusion network;
[0056] (3.6) After being input into the third GSConv module, the output feature map has a size of 20×20 and a channel number of 512, which is then output to the fourth C2f_GhostDynamicConv module;
[0057] (3.7) After passing through the fourth C2f_GhostDynamicConv module, the size and number of channels of the feature map remain unchanged. It is then processed by the spatial pyramid fast pooling SPPF module to increase the receptive field. After passing through the feature extraction network, the feature map is output to the feature fusion network.
[0058] The fourth step is to perform feature fusion on the output of the feature extraction network:
[0059] (4.1) Feature Fusion Network Input: The feature map with a size of 20×20 and a number of channels of 512 output by the spatial pyramid fast pooling SPPF module of the feature extraction network is input to the first convolutional layer. After the first convolutional layer, the feature map with a size of 20×20 and a number of channels of 512 is output. It is transmitted to the first upsampling module on the one hand and to the fourth aggregation module on the other hand.
[0060] (4.2) After the first upsampling module, the output feature map is 40×40 in size and 512 in number of channels, which is transmitted to the first aggregation module. The input of the first aggregation module consists of the output of the second C2f_GhostDynamicConv module of the feature extraction network and the output of the first upsampling module of the feature fusion network. After aggregation, the output feature map is 40×40 in size and 512 in number of channels, which is transmitted to the first aggregation module.
[0061] (4.3) After aggregation, the output feature map with a size of 40×40 and a channel number of 512 is input to the first C2f_GhostDynamicConv module of the feature fusion network. The size and channel number of the feature map after the first C2f_GhostDynamicConv module remain unchanged and are transmitted to the second convolutional layer;
[0062] (4.4) After the second convolutional layer, the feature map has unchanged size and number of channels. It is input to the second upsampling module and transmitted to the third aggregation module.
[0063] (4.5) After the second upsampling module, the output feature map with a size of 80×80 and a number of channels of 512 is output to the second aggregation module; the input of the aggregation module consists of the output of the first GSConv module of the feature extraction network and the output of the second upsampling module of the feature fusion network. After aggregation, the output feature map with a size of 80×80 and a number of channels of 768 is output to the second C2f_GhostDynamicConv module;
[0064] (4.6) After the second C2f_GhostDynamicConv module, the feature map with a size of 80×80 and a number of channels of 256 is output to the detection head network, and the feature map with a size of 80*80 and a number of channels of 256 is output to the second GSConv module;
[0065] (4.7) After the second GSConv module outputs a feature map with a size of 40×40 and a number of channels of 256, it is output to the third aggregation module; the third aggregation module aggregates the first GSConv module and the first convolution module and outputs a feature map with a size of 40×40 and a number of channels of 768, which is output to the third C2f_GhostDynamicConv module;
[0066] (4.8) After passing through the third C2f_GhostDynamicConv module, on the one hand, the output feature map with a size of 40×40 and a number of channels of 512 is output to the detection head network, and on the other hand, the output feature map with a size of 40×40 and a number of channels of 512 is output to the third GSConv module;
[0067] (4.9) After the third GSConv module outputs a feature map with a size of 20×20 and a number of channels of 512, it is output to the fourth aggregation module; the third aggregation module aggregates the second GSConv module and the first convolution module and outputs a feature map with a size of 20×20 and a number of channels of 512, which is output to the fourth C2f_GhostDynamicConv module;
[0068] (4.10) After the fourth C2f_GhostDynamicConv module outputs a feature map with a size of 20×20 and a channel number of 512, it is output to the detection head network to enhance the detection capability of multi-scale objects and input to the detection head network;
[0069] In the fifth step, the detection head network outputs the bounding box, category, and confidence of the target; the features output by the feature fusion network are input to the detection head network to output the bounding box, category, and confidence of the target; the detection head network receives three feature maps of different scales (80×80, 40×40, and 20×20) from the feature fusion network, which are used to detect small, medium, and large objects respectively. The characteristics are:
[0070] The detection head network adopts a decoupled head and anchor-free design. It first processes the multi-scale feature maps from the feature fusion network through a shared convolutional layer, and then separates it into classification and regression branches. The classification branch outputs category probabilities (sigmoid activation, supporting multi-label) and outputs the category probability of each anchor point. The regression branch directly predicts the center offset, width and height of the bounding box and outputs the coordinate parameters of the bounding box. During training, the dynamic label Task-AlignedAssigner is used to assign matching positive samples, and the classification loss, regression loss, and confidence loss are jointly optimized. During inference, the prediction results are filtered by confidence threshold and non-maximum suppression, and the coordinates, category, and confidence of the detection box are finally output.
[0071] Step 6: Obtain the gradient based on the loss, back-propagate and update the optimizer parameters. Repeat steps 4 to 8 until the number of iterations reaches the maximum, and the optimal pedestrian detection model is obtained.
[0072] Step 7: Input the test set into the optimal pedestrian detection model and output the pedestrian detection result.
Claims
1. A lightweight pedestrian detection method, characterized by The following steps are involved: The first step is to use the corresponding database to load the pedestrian detection data, the original image data I∈R H×W×3 , H and W are the length and width of the original image respectively, 3 is the number of color channels, and the image is preprocessed to convert the input original image into the YOLO training format; In the second step, the original image data is input into a lightweight feature extraction network consisting of a C2f_GhostDynamicConv module, a Conv module, and a GSConv module. The feature extraction network consists of a sequential stack of a Conv module, another Conv module, and an instance of a C2f_GhostDynamicConv module, as well as three sequential stacks of a GSConv module and a C2f_GhostDynamicConv module, designed to effectively capture image features. The data is then integrated using the Spatial Pyramid Pooling-Fast (SPPF) module. The third step is to input the feature map output by the SPPF module of the feature extraction network into a lightweight feature fusion network composed of a Conv module, a C2f_GhostDynamicConv module, and a GSConv module. Feature maps of different scales are generated through upsampling and aggregation operations. The network contains the same Conv module, C2f_GhostDynamicConv module, and GSConv module as the feature extraction network. After feature fusion processing, the detection head network outputs feature maps of three different scales; In the fourth step, the features output by the feature fusion network are input to the detection head network to output the bounding box, category, and confidence of the target; the detection head network receives three feature maps of different scales from the feature fusion network, which are used to detect small, medium, and large objects respectively; Step 5: Obtain the gradient based on the loss, back-propagate and update the optimizer parameters. Repeat steps 4 to 8 until the number of iterations reaches the maximum, and the optimal pedestrian detection model is obtained. Step 6: Input the test set into the optimal pedestrian detection model and output the pedestrian detection result.
2. A lightweight feature fusion network composed of a Conv module, a C2f_GhostDynamicConv module and a GSConv module in a lightweight pedestrian detection method according to claim 1, characterized in that The following steps are involved: In the first step, the feature extraction network input size is I∈R H×W×3 The image passes through the first convolutional layer with a kernel size of 3, a stride of 2, and a padding of 1, and then the output size is The feature map with 64 channels is transferred to the second convolutional layer. The convolutional layer parameter settings are the same as the first layer, and the size is obtained. Feature map with 128 channels; In the second step, it is transmitted to the first C2f_GhostDynamicConv module with 2 convolutional layers and n Bottleneck_GhostDynamicConvs composed of Ghost modules and dynamic convolutions. The C2f_GhostDynamicConv module first passes the input data through the first convolution layer Conv with a convolution kernel size of 2 and a stride of 1, and then divides the output into two parts. One part is passed directly to the output, and the other part passes through n Bottleneck_GhostDynamicConv modules. Among them, the Bottleneck_GhostDynamicConv implemented by the Bottleneck_GhostDynamicConv module using dynamic convolution technology can introduce more parameters to a large extent while minimizing the increase in computational complexity. In the Bottleneck_GhostDynamicConv module, the input feature X is first convolved by dynamic convolution to generate an intrinsic feature map Y of a fixed number of intrinsic feature maps X. Assume that the input feature X∈R h ×w×3 , where h and w are the length and width of the original image respectively, 3 is the number of color channels, and the intrinsic feature map Y created is expressed as follows: Among them, F i is the i-th convolution weight tensor, * represents convolution, α i is the corresponding dynamic coefficient, where The input feature map X is converted into a vector using global average pooling "Pool(·)", and then a two-layer MLP module "MLP(·)" with SoftMax "SoftMax(·)" activation is applied to dynamically generate coefficients; Using the created intrinsic feature map Y, the related feature Z is calculated through a series of linear operations. The expression is as follows: where Y i ′ is the i-th feature of the intrinsic feature map Y, Φ ij It is a linear operation to generate the jth related feature, s and t represent the number of i and j respectively. The feature information Y generated by the linear operation ij Connected to the inherent features Y, Y and Y ij The results of the two parts are concatenated in the channel dimension to obtain feature Z; The feature Z and the features directly output by the previous convolutional layer are partially processed by the aggregation module and passed through the second convolutional layer Conv with a convolution kernel size of 2 and a stride of 1 to obtain the final output. The size and number of channels of the obtained feature map remain unchanged and are transmitted to the first GSConv convolutional layer; In the third step, the feature map is input into the first GSConv convolution layer, and the output size is The number of channels of the feature map is 256. The GSConv module includes a depthwise separable convolution module. The workflow is as follows: the input of the GSConv module undergoes a 1×1 standard convolution to change the number of channels. The depthwise separable convolution module performs a 1×1 convolution on each input channel separately. The number of channels remains unchanged after processing. The result of the first convolution is then concatenated with the structure after the depthwise separable convolution to fuse features of different receptive fields and enhance expression capabilities. The concatenated result is shuffled (Channel Shuffle) to rearrange different channels and promote cross-channel information interaction. Assuming the input feature map is X and the output feature map is Y, the expression is as follows: Y=ChannelShuffle(Concat(Conv(X),DWConv(X))), (3) Among them, "Conv(·)" is a standard convolution, "DWConv(·)" represents a depth-wise separable convolution, "Concat(·)" represents a concatenation operation, and "ChannelShuffle(·)" represents a shuffling operation. After the first GSConv convolution layer, the data has two directions: one is transmitted to the feature fusion network, and the other is transmitted to the second C2f_GhostDynamicConv module consistent with step 2; In the fourth step, the size of the output of the second C2f_GhostDynamicConv module is The feature map with 256 channels is input to the second GSConv module consistent with step 3; Step 5: After the second GSConv module, the output size is The feature map with 512 channels is input to the third C2f_GhostDynamicConv module consistent with (2); the output size of the third C2f_GhostDynamicConv module is The feature map with 512 channels also has two directions: on the one hand, it is transmitted to the third GSConv module consistent with step 3, and on the other hand, it is transmitted to the feature fusion network; In the sixth step, the output size is input into the third GSConv module consistent with step three. The feature map with 512 channels is output to the fourth C2f_GhostDynamicConv module; In the seventh step, the feature map, after passing through the fourth C2f_GhostDynamicConv module, retains its size and number of channels. It is then processed by the SPPF module to increase the receptive field, and the result is passed to the feature fusion network. The SPPF module pools the feature map to a fixed size to increase the diversity of feature expression.
3. A lightweight feature fusion network composed of a Conv module, a C2f_GhostDynamicConv module and a GSConv module in a lightweight pedestrian detection method according to claim 1, characterized in that The following steps are involved: In the first step, the feature fusion network input is the size of the output of the spatial pyramid fast pooling SPPF module of the feature extraction network. The feature map with a channel number of 512 is input to the first convolutional layer. The convolution kernel size of the convolution layer is 3, the stride is 2, and the boundary expansion is 1. After the first convolutional layer, the output size is The feature map with a channel number of 512 is transmitted to the first upsampling module on the one hand, and to the fourth aggregation module on the other hand; In the second step, the output size after the first upsampling module is The feature map with 512 channels is transmitted to the first aggregation module; the input of the first aggregation module consists of the output of the second C2f_GhostDynamicConv module of the feature extraction network and the output of the first upsampling module of the feature fusion network. The output size after aggregation is The feature map with 512 channels is transmitted to the first aggregation module; In the third step, the output size after aggregation is The feature map with a channel number of 512 is input to the first C2f_GhostDynamicConv module of the feature fusion network consistent with claim 2. The size and channel number of the feature map after passing through the first C2f_GhostDynamicConv module remain unchanged and is transmitted to the second convolutional layer; In the fourth step, the size and number of channels of the feature map after the second convolution layer remain unchanged. On the one hand, it is input into the second upsampling module, and on the other hand, it is transmitted to the third aggregation module; In the fifth step, after the second upsampling module, the output size is The feature map with a channel number of 512 is output to the second aggregation module; The input of the aggregation module consists of the output of the first GSConv module of the feature extraction network and the output of the second upsampling module of the feature fusion network. The output size after aggregation is The feature map with a channel number of 768 is output to the second C2f_GhostDynamicConv module consistent with the step claim 2; In the sixth step, after the second C2f_GhostDynamicConv module, the output size is The feature map with 256 channels is output to the detection head network, and the output size is The feature map with 256 channels is output to a second GSConv module consistent with step claim 2; In the seventh step, the output size of the second GSConv module is The feature map with a channel number of 256 is output to the third aggregation module; The third aggregation module aggregates the first GSConv module and the first convolution module and the output size is The feature map with a channel number of 768 is output to the third C2f_GhostDynamicConv module consistent with the step claim 2; In the eighth step, after the third C2f_GhostDynamicConv module, the output size is The feature map with 512 channels is output to the detection head network, and the output size is The feature map with a channel number of 512 is output to a third GSConv module consistent with step claim 2; In the ninth step, the output size of the third GSConv module is The feature map with a channel number of 512 is output to the fourth aggregation module; The third aggregation module aggregates the second GSConv module and the first convolution module and the output size is The feature map with a channel number of 512 is output to the fourth C2f_GhostDynamicConv module consistent with the step of claim 2; In the tenth step, the output size of the fourth C2f_GhostDynamicConv module is The feature map with 512 channels is output to the detection head network.