Training Method, Pedestrian Detection Method and Device for Pedestrian Detection Model
By introducing attention model and classification detection model into the pedestrian detection model and adopting the intercomparison training method, the problems of pedestrian detection accuracy and speed in complex scenarios are solved, and more efficient pedestrian detection results are achieved.
Patent Information
- Application Number
- CN202210303199.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-25
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2042-03-25
AI Technical Summary
In complex scenarios, the accuracy of the SSD algorithm for pedestrian detection needs to be further improved, and there are too many pedestrian candidate areas output, which reduces the detection speed and accuracy.
By introducing attention model and classification detection model into the pedestrian detection model, combining feature extraction units and classification detection units, and using the interchange and comparison training method, the pedestrian candidate areas are screened and optimized.
It effectively improves the accuracy of pedestrian detection in complex scenarios, reduces redundant pedestrian candidate areas, and improves detection speed.
Smart Images

Figure CN114677709B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and particularly relates to a method for training a pedestrian detection model, a pedestrian detection method, and a device. Background Art
[0002] Pedestrian detection is one of the most important tasks in computer vision. Its purpose is to accurately locate pedestrians in an image or video sequence. Currently, pedestrian detection has been widely applied in intelligent vision systems, such as autonomous driving, intelligent monitoring, and road scene understanding.
[0003] SSD (single shot multibox detector) is a mainstream pedestrian detection algorithm in the industry. It uses a convolutional neural network to extract multiple features of different scales from an image and obtains the final detection result by fusing multiple feature images, with relatively high detection efficiency.
[0004] However, the accuracy of the SSD algorithm for pedestrian detection needs to be further improved in complex scenarios. Summary of the Invention
[0005] Embodiments of this application provide a method for training a pedestrian detection model, a pedestrian detection method, and a device, which can improve the accuracy of pedestrian detection.
[0006] In a first aspect, embodiments of this application provide a method for training a pedestrian detection model. The pedestrian detection model includes a feature extraction unit and a classification and detection unit. The method for training the pedestrian detection model includes:
[0007] Obtain a plurality of sample images and the standard region of each sample image;
[0008] Input the sample image into the feature extraction unit to obtain a plurality of target feature images output by the feature extraction unit; input the plurality of target feature images into the classification and detection unit to obtain the candidate regions of each target feature image output by the classification and detection unit. The candidate region is the region where the classification result of the classification and detection unit indicates the presence of a pedestrian;
[0009] Obtain the pedestrian candidate region in the sample image according to the first intersection over union of the candidate region of each target feature image and the standard region. The first intersection over union is the ratio of the intersection to the union of the candidate region of the target feature image and the standard region; train the pedestrian detection model according to the target intersection over union of the pedestrian candidate region and the standard region.
[0010] Optionally, the pedestrian detection model includes a feature extraction layer and a plurality of sampling layers. Inputting the sample image into the pedestrian detection model to obtain a plurality of target feature images output by the pedestrian detection model includes:
[0011] Input the sample image into the feature extraction layer to obtain the first feature image; for the first sampling layer, input the first feature image into the first sampling layer to obtain the second feature image; for the Nth sampling layer, input the Nth feature image output by the (N - 1)th sampling layer into the Nth sampling layer to obtain the (N + 1)th feature image, where N is an integer greater than or equal to 2; use the second to the (N + 1)th feature images as the target feature images.
[0012] Optionally, each sampling layer includes an attention model and a residual model; for each sampling layer, input the Nth feature image into the attention model and obtain the attention feature image output by the attention model; input the attention feature image into the residual model to obtain the (N + 1)th feature image.
[0013] Optionally, each attention model includes left and right branches; input the Nth feature image into the attention model and obtain the attention feature image output by the attention model, including:
[0014] Input the Nth feature image into the left branch of the attention model to obtain the left feature image output by the left branch, where the left branch is used to obtain the feature parameters of the layer dimension of the image to be processed;
[0015] Input the left feature image into the right branch of the attention model to obtain the right feature image output by the right branch, where the right branch is used to obtain the feature parameters of the width and height dimensions of the image to be processed; perform a dot product on the left feature image and the right feature image to obtain the attention feature image.
[0016] Optionally, obtain the pedestrian candidate regions in the sample image according to the first intersection over union (IoU) between the candidate regions of each target feature image and the standard region, including:
[0017] According to the first IoU and the reset IoU, screen the target candidate regions from the candidate regions through a confidence reset algorithm, where the reset IoU is the IoU used to eliminate redundant candidate regions; use the target candidate regions as the pedestrian candidate regions in the sample image.
[0018] Optionally, screen the target candidate regions from the candidate regions through a confidence reset algorithm according to the first IoU and the reset IoU, including:
[0019] Put the candidate regions with the first IoU greater than the first preset value into set A; obtain the first candidate region with the largest first IoU in set A, put it into set B, and use the remaining candidate regions in set A as the second candidate regions; screen the third candidate regions from the second candidate regions according to the reset IoU and put them into set B; use the candidate regions in set B as the target candidate regions.
[0020] Optionally, according to the reset intersection over union (IoU), filter the third candidate regions from the second candidate regions and put them into set B, including:
[0021] Obtain the reset second IoU according to the ratio of the intersection to the union of the first candidate region and the second candidate region; if the second IoU of each second candidate region is greater than the second preset value, perform weight reduction processing on the first IoU to obtain the reset third IoU; if the third IoU is greater than the first preset value, eliminate the second candidate regions with a third IoU less than the first preset value; re - sort the second candidate regions according to the target IoU, and put the candidate region with the largest target IoU among them into set B, where the target IoU includes the second IoU and the third IoU; repeat the steps of resetting the second IoU and resetting the third IoU until set A is empty.
[0022] Optionally, train the pedestrian detection model according to the target IoU between the pedestrian candidate region and the standard region, including:
[0023] Judge whether the target IoU is greater than the preset value; if not, update the parameters in the pedestrian detection model to obtain a new target IoU, and judge whether the new target IoU is greater than the preset IoU. If not, repeat the step of parameter update until the target IoU is greater than the preset IoU; if so, the training of the pedestrian detection model is completed.
[0024] In a second aspect, an embodiment of the present application further provides a pedestrian detection method, including:
[0025] Obtain the image to be processed; input the image to be processed into the pedestrian detection model to obtain the pedestrian candidate regions in the image to be processed output by the pedestrian detection model.
[0026] In a third aspect, an embodiment of the present application further provides a pedestrian detection model training device, including:
[0027] An acquisition module, configured to acquire a plurality of sample images and the standard region of each sample image;
[0028] A first training module, configured to input the sample images to be processed into the feature extraction unit to obtain a plurality of target feature images output by the feature extraction unit;
[0029] A second training module, configured to input the plurality of target feature images into the classification and detection unit to obtain the candidate regions of each target feature image output by the classification and detection unit, where the candidate region is the region where pedestrians exist in the classification result of the classification and detection unit;
[0030] A first determination module, configured to obtain pedestrian candidate regions in a sample image according to a first intersection-over-union ratio between a candidate region and a standard region of each target feature image, where the first intersection-over-union ratio is a ratio of an intersection of the candidate region and the standard region of the target feature image to a union of the candidate region and the standard region;
[0031] A second determination module, configured to train a pedestrian detection model according to a target intersection-over-union ratio between a pedestrian candidate region and a standard region.
[0032] In a fourth aspect, an embodiment of the present application further provides a pedestrian detection device, including:
[0033] An acquisition module, configured to acquire an image to be processed;
[0034] A processing module, configured to input the image to be processed into a pedestrian detection model, and obtain pedestrian candidate regions in the image to be processed output by the pedestrian detection model.
[0035] In a fifth aspect, an embodiment of the present application provides a terminal device, including: a memory and a processor;
[0036] The memory is configured to store computer instructions; the processor is configured to run the computer instructions stored in the memory to implement the method according to any one of the first aspect or the second aspect.
[0037] In a sixth aspect, an embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored, and the computer program is executed by a processor to implement the method according to any one of the first aspect or the second aspect.
[0038] In a seventh aspect, an embodiment of the present application provides a computer program product, including a computer program, and when the computer program is executed by a processor, the method according to any one of the first aspect or the second aspect is implemented. Description of the Drawings
[0039] Figure 1 It is a schematic diagram of a scenario provided by an embodiment of the present application;
[0040] Figure 2 It is a schematic diagram of a server processing an image provided by an embodiment of the present application;
[0041] Figure 3 It is a schematic flowchart of a training method of a pedestrian detection model provided by an embodiment of the present application;
[0042] Figure 4 It is a schematic diagram of the structure of a pedestrian detection model provided by an embodiment of the present application;
[0043] Figure 5 It is a schematic diagram of the structure of an attention model provided by an embodiment of the present application;
[0044] Figure 6Schematic structural diagram of the residual model provided by an embodiment of the present application;
[0045] Figure 7 Schematic process diagram of obtaining a pedestrian candidate region in a sample image provided by an embodiment of the present application;
[0046] Figure 8 Schematic process diagram of a pedestrian detection method provided by an embodiment of the present application;
[0047] Figure 9 Schematic structural diagram of a pedestrian detection model training device provided by an embodiment of the present application;
[0048] Figure 10 Schematic structural diagram of a pedestrian detection device provided by an embodiment of the present application;
[0049] Figure 11 Schematic structural diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners
[0050] For the convenience of clearly describing the technical solutions of the embodiments of the present application, the following briefly introduces some terms and technologies involved in the embodiments of the present application:
[0051] 1) Pedestrian detection is to use computer vision technology to determine whether there are pedestrians in an image or video sequence and give precise positioning. This technology can be combined with technologies such as pedestrian tracking and pedestrian re-identification, and is applied to fields such as artificial intelligence systems, vehicle assisted driving systems, intelligent robots, intelligent video surveillance, human behavior analysis, and intelligent transportation.
[0052] 2) Single Shot MultiBox Detector (SSD) is a single-stage detection model that can ensure the detection speed without sacrificing detection accuracy.
[0053] 3) Residual network is a convolutional neural network, which is characterized by being easy to optimize and being able to improve the accuracy by increasing the appropriate depth. The internal residual blocks use skip connections, alleviating the problem of gradient disappearance caused by increasing depth in deep neural networks.
[0054] 4) The attention mechanism originates from the research on human vision. In cognitive science, due to the bottleneck of information processing, humans will selectively focus on a part of all information while ignoring other visible information. The above mechanism is usually called the attention mechanism. The attention mechanism in deep learning is essentially similar to the human selective mechanism, and the core goal is also to select the information that is more critical to the current task goal from numerous information.
[0055] 5) Other terms
[0056] In the embodiments of the present application, terms such as "first" and "second" are used to distinguish identical or similar items with substantially the same functions and effects, and the order thereof is not limited. Those skilled in the art can understand that terms such as "first" and "second" do not limit the quantity and execution order, and terms such as "first" and "second" do not necessarily mean different.
[0057] It should be noted that in the embodiments of the present application, words such as "exemplary" or "for example" are used to represent examples, illustrations or explanations. Any embodiment or design solution described as "exemplary" or "for example" in the present application should not be construed as being more preferred or more advantageous than other embodiments or design solutions. Rather, the use of words such as "exemplary" or "for example" is intended to present relevant concepts in a specific manner.
[0058] The file transfer method provided by the embodiments of the present application will be introduced in detail below with reference to the accompanying drawings. It should be noted that "when... " in the embodiments of the present application can be at the instant when a certain situation occurs, or within a period of time after a certain situation occurs. The embodiments of the present application do not make specific limitations thereto.
[0059] Pedestrian detection is one of the most important tasks in computer vision. Its purpose is to accurately locate pedestrians in images or video sequences. Currently, pedestrian detection has been widely applied in intelligent vision systems, such as autonomous driving, intelligent monitoring, and road scene understanding.
[0060] SSD is a mainstream pedestrian detection algorithm in the industry. It uses a convolutional neural network to extract multiple features of different scales from an image, and obtains the final detection result by fusing multiple feature images, with relatively high detection efficiency.
[0061] However, in complex scenarios, SSD will output a relatively large number of pedestrian candidate regions, which will not only reduce the detection speed but also affect the detection accuracy.
[0062] In view of this, the present application proposes a pedestrian detection model, that is, a pedestrian detection method, which introduces an attention model and a classification detection model into the SSD network, can effectively improve the accuracy of pedestrian detection in complex scenarios and reduce redundant pedestrian candidate regions, and improve the speed of pedestrian detection.
[0063] The technical solutions of the present invention and how the technical solutions of the present application solve the above technical problems will be described in detail below with specific embodiments. These specific embodiments below can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of the present invention will be described below with reference to the accompanying drawings.
[0064] Figure 1 For the scene schematic diagram provided by the embodiments of the present application, asFigure 1 As shown, it includes a server and an image acquisition device. The image acquisition device is connected to the server through a network for data transmission. A pedestrian detection model and a classification model capable of recognizing pedestrians are set in the server.
[0065] The image acquisition device can collect information on the road in real time, including pedestrians, vehicles, buildings, etc. on the road and their locations, generate corresponding images, and send the images to the server through the network. Based on the received images, the server recognizes the pedestrians in the images and marks them in the form of a box, as Figure 2 shown.
[0066] It can be understood that the server can also use the multiple obtained images as sample images to train the pedestrian detection model to obtain the expected output results. During the training process, the pedestrians in the sample images will be marked, such as marking the area where the pedestrians are located with a box.
[0067] Exemplarily, in an autonomous driving scenario, the vehicle can collect information on the road through the mounted image acquisition device, process it into corresponding images and send them to the connected server. The server performs pedestrian recognition on the received images, marks the locations of the pedestrians and returns them to the vehicle. The vehicle can judge the locations of the pedestrians on the road based on the received information to reduce the probability of a vehicle-pedestrian collision.
[0068] The above briefly describes the application scenarios provided by the embodiments of the present application. Below, taking the server applied to Figure 1 as an example, the training method of the pedestrian detection model provided by the embodiments of the present application will be described in detail.
[0069] Figure 3 FIG. is a schematic flowchart of the training method of the pedestrian detection model provided by the embodiments of the present application, including the following steps:
[0070] S301. Obtain a plurality of sample images and the standard regions of each sample image.
[0071] The sample images are images including pedestrians in the pattern, such as Figure 2 the pictures generated from the road information collected by the image acquisition device shown. The standard region of the sample image is the region where the pedestrians are marked in the sample image.
[0072] The server can obtain a plurality of sample images and the corresponding standard regions of the sample images for training the pedestrian detection model.
[0073] S302. Input the sample images into the feature extraction unit to obtain a plurality of target feature images output by the feature extraction unit.
[0074] The target feature image is the feature image obtained by the feature extraction unit in the pedestrian detection model through sampling the sample image. Among them, this sampling can be, for example, downsampling, and the data processing volume is reduced through downsampling.
[0075] For example, the feature extraction unit in the pedestrian detection model of the embodiment of the present application can adopt the SSD network as the basic structure. For example, an attention model and a residual model can also be introduced into the SSD. The attention model is used to obtain the feature parameters of the sample image and perform weighted processing, and the residual model is used to prevent the gradient disappearance during the downsampling process. The implementation manner of the feature extraction unit in the pedestrian detection model in this embodiment is not particularly limited.
[0076] The server inputs the sample image into the feature extraction unit in the pedestrian detection model, and performs multiple sampling processes on the sample image to obtain the target feature image output after the sampling process.
[0077] S303. Input multiple target feature images into the classification and detection unit to obtain the candidate regions of each target feature image output by the classification and detection unit. The candidate region is the region where the classification result of the classification and detection unit indicates the presence of a pedestrian.
[0078] The classification and detection unit is used to label the pedestrians in the target feature image. The pedestrians identified by the classification and detection unit are marked with a rectangle. The region where each pedestrian identified by the classification and detection unit in each target feature image is located is the candidate region.
[0079] S304. According to the first intersection-over-union ratio of the candidate region of each target feature image and the standard region, obtain the pedestrian candidate region in the sample image. The first intersection-over-union ratio is the ratio of the intersection to the union of the candidate region of the target feature image and the standard region.
[0080] Calculate the ratio of the intersection to the union of each candidate region of each target feature image output by the classification and detection unit and the standard region, and determine whether it meets the first preset value. Eliminate the candidate regions whose first intersection-over-union ratio does not meet the first preset value, and use the remaining candidate regions as the pedestrian candidate regions in the sample image.
[0081] Exemplarily, for a specific pedestrian candidate region, after multiple downsamplings, multiple candidate regions of different sizes are obtained. Calculate the first intersection-over-union ratio of each candidate region and the standard region. When the candidate region coincides with the standard region, the first intersection-over-union ratio is 1. Determine whether the size of the first intersection-over-union ratio is greater than the first preset value. For example, the first preset value can be 0.5. When the first intersection-over-union ratio is less than 0.5, this candidate region will be eliminated, and the remaining candidate regions are used as the pedestrian candidate regions in the sample image. It can be understood that the size of the first preset value can be adjusted according to actual needs, and the embodiment of the present application does not limit this.
[0082] S305. Train the pedestrian detection model according to the intersection over union (IoU) between the pedestrian candidate region and the standard region.
[0083] After determining the pedestrian candidate region in the sample image, the server determines whether the target IoU between the pedestrian candidate region and the standard region meets a preset value. If not, based on the difference between the target IoU and the preset value, adjust the parameters in the pedestrian detection model, and repeat the steps shown in S301 - S304 until the first IoU of the pedestrian candidate region meets the preset value.
[0084] Exemplarily, the preset value can be 0.95. When the server determines that the first IoU of the pedestrian candidate region is less than 0.95, adjust the parameters of each layer of the pedestrian detection model. For example, the size of the convolutional kernel can be adjusted. Repeat the process of S301 to S304 multiple times until the first IoU of the pedestrian candidate region meets the preset value.
[0085] The pedestrian detection model training method provided in this embodiment inputs the sample image into the feature extraction unit in the pedestrian detection model to obtain multiple target feature images, inputs the target features into the classification and detection unit to obtain the pedestrian candidate region in the sample image, and trains the pedestrian detection model according to the IoU between the pedestrian candidate region in the sample image and the standard region. Using the trained pedestrian detection model to detect pedestrians in the image can effectively improve the accuracy of the feature images output by the pedestrian detection model and reduce redundant pedestrian candidate regions. It improves the accuracy and detection speed of pedestrian detection in complex scenarios.
[0086] Next, on the basis of the Figure 3 embodiment, taking the input of the sample image into the feature extraction unit in the pedestrian detection model as an example, the structure of the pedestrian detection model will be described in detail.
[0087] Figure 4 The structure diagram of the feature extraction unit in the pedestrian detection model provided in the embodiment of the present application includes a feature extraction layer and multiple sampling layers. Each sampling layer includes an attention model and a residual model (ResNet). In this embodiment of the present application, 4 sampling layers are taken as an example for illustration. It can be understood that the number of sampling layers can be increased or decreased according to actual needs.
[0088] Correspondingly, inputting the sample image into the feature extraction unit to obtain multiple target feature images output by the feature extraction unit includes the following steps:
[0089] S401. Input the sample image into the feature extraction layer to obtain the first feature image.
[0090] The feature extraction layer is used to extract the feature parameters of the sample image, including the feature parameters in three dimensions: the number of layers, height, and width.
[0091] Exemplarily, the first feature image output by the feature extraction layer can be represented in the form of feature parameters, denoted as C*H*W, where C is the number of layers, H is the height, and W is the width, and each feature image is the corresponding target feature image.
[0092] S402: Input the first feature image into the first sampling layer to obtain the second feature image.
[0093] The specific implementation process is as shown in S1 - S2 below.
[0094] S1: Input the first feature image into the attention model to obtain the attention feature image output by the attention model.
[0095] The attention model includes two branches, the left and the right, as Figure 5 shown.
[0096] Among them, the left branch includes: left 1 branch, left 2 branch, left 3 branch; the left 1 branch includes the first processing layer and the second processing layer, and the left 2 branch includes the third processing layer. The right branch includes: right 1 branch, right 2 branch, right 3 branch; the right 1 branch includes the fourth processing layer and the fifth processing layer, and the right 2 branch includes the sixth processing layer.
[0097] The first processing layer includes a convolutional layer and a reshape layer; the second processing layer includes a convolutional layer, a reshape layer, and a softmax classification layer; the third processing layer includes a convolutional layer, a layernorm normalization layer, and a sigmod activation layer; the fourth processing layer includes a convolutional layer and a reshape layer; the reshape layer and the activation layer; the sixth processing layer includes a convolutional layer, a global pooling layer, a reshape layer, and a classification layer. The convolution kernel size of each convolutional layer is 1*1
[0098] The specific implementation process is as follows:
[0099] S11: Input the first feature image into the left branch of the first attention model to obtain the left feature image output by the left branch.
[0100] In the specific implementation process, input the first feature image into the first processing layer to obtain the first dimensionality - reduced feature image; input the first feature image into the third processing layer to obtain the second dimensionality - reduced feature image; multiply the first dimensionality - reduced feature image and the second dimensionality - reduced feature image in matrix form, and input it into the second processing layer to obtain the third dimensionality - reduced feature image; multiply the third dimensionality - reduced feature image and the first feature image of the left 3 branch in dot - product form to obtain the left feature image.
[0101] Exemplarily, the feature parameters of the first feature map are C*H*W. It is input into the first processing layer. After being processed by the convolutional layer, the feature parameters become C / 2*H*W. Then it is input into the reconstruction layer for processing to obtain the first downsampled feature image, whose feature parameters are C / 2*HW. Among them, the reconstruction layer is used to reconstruct the shape of the feature image.
[0102] The first feature map is input into the third processing layer. After being processed by the convolutional layer, the feature parameters are 1*H*W. Then it is input into the reconstruction layer and the classification layer for processing to obtain the second downsampled feature image, whose feature parameters are HW*1*1.
[0103] The first downsampled feature image and the second downsampled feature image are subjected to matrix multiplication processing, and their feature parameters are C / 2*1*1. Then it is input into the second processing layer. After being processed by the convolutional layer, the normalization layer, and the activation layer, the third downsampled feature image is obtained, whose feature parameters are C*1*1. Among them, the normalization layer processing refers to normalizing the image features, pulling the data distribution to the non-saturated region of the activation function, having the characteristics of weight / data scaling invariance, and achieving the effects of alleviating gradient disappearance / explosion, accelerating training, and regularization.
[0104] The third downsampled feature image and the first feature image are multiplied point by point to obtain the left feature image, whose feature parameters are C*H*W.
[0105] S12: Input the left feature image into the right branch of the first attention model to obtain the right feature image output by the right branch.
[0106] In a specific implementation, the left feature image is input into the fourth processing layer to obtain the fourth downsampled feature image; the left feature image is input into the sixth processing layer to obtain the fifth downsampled feature image; the fifth downsampled feature image and the fourth downsampled feature image are multiplied matrix-wise and input into the fifth processing layer to obtain the sixth downsampled feature image; the sixth downsampled feature image and the left feature image of the right three branches are multiplied point by point to obtain the right feature image.
[0107] Exemplarily, the left feature map is input into the fourth processing layer. After being processed by the convolutional layer, the feature parameters are C / 2*H*W. Then it is input into the reconstruction layer for processing to obtain the third downsampled feature image, whose feature parameters are C / 2*HW.
[0108] The first feature map is input into the sixth processing layer. After being processed by the convolutional layer, the feature parameters are C / 2*H*W. Then it is input into the global pooling layer for processing, and the feature parameters are C / 2*1*1. Finally, it is input into the reconstruction layer and the classification layer for processing to obtain the fourth downsampled feature image, whose feature parameters are 1*C / 2.
[0109] The fourth dimensionality-reduced feature image and the fifth dimensionality-reduced feature image are subjected to matrix multiplication processing, and their feature parameters become 1*HW, and then input into the fifth processing layer. After being processed by the reconstruction layer and the activation layer, a sixth dimensionality-reduced feature image is obtained, and its feature parameters are 1*H*W.
[0110] The sixth dimensionality-reduced feature image and the left feature image are multiplied point by point to obtain a right feature image, and its feature parameters are C*H*W.
[0111] S13: Multiply the left feature image and the right feature image point by point to obtain an attention feature image.
[0112] The point multiplication of the feature images does not change their feature parameters, and the feature parameters of the attention feature image are C*H*W.
[0113] S2: Input the attention feature image into the residual model to obtain a second feature image.
[0114] The residual model includes multiple convolutional layers and activation layers. In this embodiment of the application, 3 convolutional layers are taken as an example for illustration. An activation layer is provided behind each convolutional layer, and the relu function is used as the activation function in the activation layer. The structure of the residual model is as Figure 6 shown.
[0115] The specific method for obtaining the second feature image from the attention feature image is as shown in S21-S24.
[0116] S21: Input the attention feature image into the first convolutional layer to obtain a first convolutional feature image.
[0117] The convolutional kernel size of the first convolutional layer is 1*1. After the attention feature image is subjected to convolution and activation processing, a first convolutional feature image is obtained, and its feature parameters are C*H*W / 4.
[0118] S22: Input the first convolutional feature image into the second convolutional layer to obtain a second convolutional feature image.
[0119] The convolutional kernel size of the second convolutional layer is 3*3. After the first convolutional feature image is subjected to convolution and activation processing, a second convolutional feature image is obtained, and its feature parameters are C*H*W / 4.
[0120] S23: Input the second convolutional feature image into the third convolutional layer to obtain a third convolutional feature image.
[0121] The convolutional kernel size of the third convolutional layer is 1*1. After the second convolutional feature image is subjected to convolution and activation processing, a third convolutional feature image is obtained, and its feature parameters are C*H*W.
[0122] S24: Superimpose the third convolutional feature image and the attention feature image, input it into the activation layer, and obtain a second feature image.
[0123] The superimposing process of the feature image does not change the feature parameters of the feature image. After the activation layer process, its feature parameters are C*H*W.
[0124] S403. Input the second feature image into the first sampling layer to obtain a third feature image.
[0125] S404. Input the third feature image into the first sampling layer to obtain a fourth feature image.
[0126] S405. Input the fourth feature image into the first sampling layer to obtain a fifth feature image.
[0127] In the embodiments of the present application, the specific implementation processes of S403 - S405 are similar to the process shown in S402, and will not be elaborated here.
[0128] Among them, the feature parameters of the third feature image are 4C*H / 4*W / 4, the feature parameters of the fourth feature image are 8C*H / 8*W / 8, and the feature parameters of the fifth feature image are 16C*H / 16*W / 16.
[0129] The above has described the multiple target feature images for obtaining the sample image. Next, the process of obtaining the pedestrian candidate regions in the sample image according to the target feature images will be described in detail.
[0130] Figure 7 It is a schematic flowchart of obtaining the pedestrian candidate regions in the sample image according to the target feature images provided by the embodiments of the present application, including:
[0131] S701. Input the multiple target feature images into the classification and detection unit to obtain the candidate regions of each target feature image output by the classification and detection unit. The candidate region is the region where the classification result of the classification and detection unit indicates the presence of a pedestrian.
[0132] In the embodiments of the present application, the implementation manner of S701 is similar to Figure 3 the implementation manner of S303 in
[0133] S702. Sort the candidate regions of each target feature image in descending order according to the first intersection over union ratio, and eliminate the candidate regions with the first intersection over union ratio less than the first preset value, and put them into set A.
[0134] If the first intersection over union (IoU) of the candidate regions is relatively small, it indicates that there is a large difference between the candidate regions output by the classification detection model and the ground truth regions. Setting the first preset value can screen the candidate regions, eliminating the candidate regions with a small first IoU and accelerating the detection speed of the model. In the embodiments of the present application, the first preset value is 0.5. It can be understood that the first preset value can also be changed according to actual needs, and the embodiments of the present application do not limit this.
[0135] Based on the candidate regions output by the classification detection model, calculate the first IoU of each candidate region, sort them according to the size of the first IoU, and eliminate the candidate regions with a first IoU less than 0.5.
[0136] Sets A and B can be established to store the candidate regions, and sets A and B are initially empty sets.
[0137] Put the screened candidate regions into set A, and the steps shown in S703 can be executed.
[0138] S703: Calculate the second IoU between the first candidate region and each second candidate region respectively. The second IoU is the ratio of the intersection to the union of the first candidate region and the second candidate region.
[0139] The first candidate region is the candidate region with the largest first IoU in set A, and the remaining candidate regions are the second candidate regions.
[0140] Put the first candidate region from set A into set B, and calculate the second IoU of each second candidate region.
[0141] S704: Determine whether the second IoU of each second candidate region is greater than the second preset value, and calculate the third IoU of the second candidate regions with a second IoU greater than the second preset value. Among them, the third IoU is a value obtained by de-duplicating the first IoU.
[0142] The overlap degree between the second candidate region and the first candidate region can be judged according to the second IoU. Setting the second preset value can reduce the number of overlapping regions output by the classification detection model. In the embodiments of the present application, the second preset value is 0.6. It can be understood that the second preset value can also be changed according to actual needs, and the embodiments of the present application do not limit this.
[0143] The server calculates the third IoU for each second candidate region with a second IoU greater than the second preset value. Exemplarily, the calculation method of the third IoU can be as follows:
[0144]
[0145] Where Si1 is the third IoU, Si is the first IoU, and σ is an empirical constant.
[0146] S705. Determine whether the third intersection over union (IoU) is greater than a first preset value, and remove the second candidate regions with a third IoU less than the first preset value.
[0147] The server removes the candidate regions with a third IoU less than the first preset value according to the calculated third IoU of the second candidate regions.
[0148] S706. Re - sort the second candidate regions according to the target IoU, and put the candidate region with the largest target IoU into set B, where the target IoU includes the second IoU and the third IoU.
[0149] After the processing steps shown in S702 - S705, each candidate region in set A has multiple IoUs. Sort them according to the finally calculated IoU of each candidate region. The finally calculated IoU is also called the target IoU, and put the candidate region with the largest target IoU into set B.
[0150] Exemplarily, if a certain candidate region in set A has the first, second, and third IoUs, then the third IoU is the target IoU.
[0151] S707. Repeat the steps shown in S703 - S706 until set A is an empty set, and the candidate regions in set B are the pedestrian candidate regions corresponding to the sample image output by the classification detection model.
[0152] It can be understood that during the repeated execution, the first candidate region used each time is the candidate region newly put into set B, and the second candidate region is the remaining candidate region in set A after removal. Each time it is executed, a candidate region is taken out from set A and put into set B, and the candidate regions that do not meet the requirements are removed until set A is empty.
[0153] In the above method, by calculating multiple IoUs of candidate regions and setting multiple preset values for screening, redundant and overlapping candidate regions in multiple target feature images can be removed, the detection speed of the model can be accelerated, and relatively accurate pedestrian candidate regions can be obtained.
[0154] This application embodiment also provides a pedestrian detection method, as Figure 8 shown, including the following steps:
[0155] S801. Obtain the image to be processed.
[0156] S802. Input the image to be processed into the pedestrian detection model, and obtain the pedestrian candidate regions in the image to be processed output by the pedestrian detection model.
[0157] Among them, the pedestrian detection model is obtained through Figure 3obtained by the pedestrian detection model training method shown below.
[0158] In the above method, an attention model is introduced into the feature extraction unit of the pedestrian detection model. By using the advantage that the attention model can perform weighted processing on image feature information, the feature information of the sample image is fully extracted and processed. The redundant and overlapping candidate regions are removed through the classification detection unit, improving the accuracy of the pedestrian detection model for pedestrian detection.
[0159] An embodiment of the present application further provides a pedestrian detection model training device 90, as Figure 9 shown, including: an acquisition module 901, a first training module 902, a second training module 903, a first determination module 904, and a second determination module 905.
[0160] The acquisition module 901 is configured to acquire a plurality of sample images and the standard region of each sample image.
[0161] The first training module 902 is configured to input the to-be-sample image into the feature extraction unit to obtain a plurality of target feature images output by the feature extraction unit.
[0162] The second training module 903 is configured to input a plurality of target feature images into the classification detection unit to obtain candidate regions of each target feature image output by the classification detection unit, where the candidate region is a region where a pedestrian exists in the classification result of the classification detection unit.
[0163] The first determination module 904 is configured to obtain the pedestrian candidate regions in the sample image according to the first intersection-over-union ratio of the candidate region of each target feature image and the standard region, where the first intersection-over-union ratio is the ratio of the intersection to the union of the candidate region of the target feature image and the standard region.
[0164] The second determination module 905 is configured to train the pedestrian detection model according to the intersection-over-union ratio of the pedestrian candidate region and the standard region.
[0165] The pedestrian detection model training device provided in this embodiment can execute Figure 3 the technical solutions of the method embodiments shown below. The implementation principles and technical effects are similar and will not be elaborated here.
[0166] An embodiment of the present application further provides a pedestrian detection device 10, as Figure 10 shown, including: an acquisition module 1001, a processing module 1002, and a classification module 1003.
[0167] The acquisition module 1001 is configured to acquire the to-be-processed image.
[0168] The processing module 1002 inputs the image to be processed into the pedestrian detection model, and obtains the candidate pedestrian regions in the image to be processed output by the pedestrian detection model.
[0169] The pedestrian detection device provided in this embodiment can execute Figure 8 the technical solutions of the method embodiments shown, and their implementation principles and technical effects are similar, which will not be elaborated here.
[0170] Figure 11 It is a schematic structural diagram of an electronic device provided in an embodiment of the present application. As Figure 11 shown, the device 11 provided in this embodiment may include:
[0171] A processor 1101.
[0172] A memory 1102 for storing executable instructions of the electronic device.
[0173] Among them, the processor is configured to execute the technical solutions of the above-mentioned pedestrian detection model training method or the pedestrian detection method embodiments by executing the executable instructions, and their implementation principles and technical effects are similar, which will not be elaborated here.
[0174] An embodiment of the present application also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the technical solutions of the above-mentioned pedestrian detection model training method or the pedestrian detection method embodiments, and their implementation principles and technical effects are similar, which will not be elaborated here.
[0175] In one possible implementation, the computer-readable medium may include a random access memory (RAM), a read-only memory (ROM), a compact disc read-only memory (CD-ROM), or other optical disc storage, a magnetic disk storage, or any other medium targeted to carry or store the required program code in the form of instructions or data structures and accessible by a computer. Moreover, any connection is properly termed a computer-readable medium. For example, if software is transmitted from a website, server, or other remote source using coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL), or wireless technologies such as infrared, radio, and microwave, then the coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies such as infrared, radio, and microwave are included in the definition of the medium. As used herein, disk and disc include optical disc, laser disc, optical disc, digital versatile disc (DVD), floppy disk, and Blu-ray disc, where disks typically reproduce data magnetically, while discs reproduce data optically using a laser. Combinations of the above should also be included within the scope of computer-readable media.
[0176] In an embodiment of the present application, a computer program product is further provided, including a computer program, which implements the technical solutions of the above-mentioned pedestrian detection model training method or the pedestrian detection method embodiment when executed by a processor. The implementation principle and technical effects are similar and will not be elaborated here.
[0177] In the specific implementation of the above terminal device or server, it should be understood that the processor may be a central processing unit (CPU for short), or may also be other general-purpose processors, digital signal processors (DSP for short), application-specific integrated circuits (ASIC for short), etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of the present application may be directly embodied as being executed by a hardware processor, or may be executed by a combination of hardware and software modules in the processor.
[0178] Those skilled in the art can understand that all or part of the steps of any of the above method embodiments can be completed by hardware related to program instructions. The foregoing program can be stored in a computer-readable storage medium. When the program is executed, all or part of the steps of the above method embodiments are executed.
[0179] If the technical solution of the present application is implemented in the form of software and sold or used as a product, it can be stored in a computer-readable storage medium. Based on such an understanding, all or part of the technical solution of the present application can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a computer program or several instructions. The computer software product enables a computer device (which may be a personal computer, a server, a network device, or a similar electronic device) to execute all or part of the steps of the method described in the embodiments of the present application.
[0180] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present application.
Claims
1. A training method for a pedestrian detection model, characterized in that, The pedestrian detection model includes a feature extraction unit and a classification and detection unit, and the method includes: Obtain a plurality of sample images and the standard region of each sample image; Input the sample image into the feature extraction layer of the feature extraction unit to obtain a first feature image; The feature extraction unit further includes a plurality of sampling layers; for the first sampling layer, input the first feature image into the attention model corresponding to the first sampling layer, and obtain the attention feature image output by the attention model, and input the attention feature image into the residual model corresponding to the first sampling layer to obtain a second feature image; For each of the second to the Nth sampling layers: input the feature image output by the previous sampling layer of this sampling layer into the attention model, and obtain the attention feature image output by the attention model; input the attention feature image into the residual model corresponding to this sampling layer to obtain the feature image output by this sampling layer; wherein, the ith sampling layer outputs the (i + 1)th feature image, and i is an integer greater than or equal to 2 and less than or equal to N; the multiple target feature images include the second to the (N + 1)th feature images output by the second to the (N + 1)th sampling layers respectively; Input the multiple target feature images into the classification and detection unit to obtain the candidate region of each target feature image output by the classification and detection unit, and the candidate region is the region where the classification result of the classification and detection unit indicates the presence of a pedestrian; Obtain the pedestrian candidate region in the sample image according to the first intersection-over-union ratio of the candidate region of each target feature image and the standard region, and the first intersection-over-union ratio is the ratio of the intersection to the union of the candidate region of the target feature image and the standard region; Train the pedestrian detection model according to the intersection-over-union ratio of the pedestrian candidate region and the standard region.
2. The method according to claim 1, characterized in that, Each of the attention models includes left and right branches; the step of inputting the Nth feature image into the attention model and obtaining the attention feature image output by the attention model includes: Input the Nth feature image into the left branch of the attention model to obtain the left feature image output by the left branch, and the left branch is used to obtain the feature parameters of the layer dimension of the image to be processed; Input the left feature image into the right branch of the attention model to obtain the right feature image output by the right branch, and the right branch is used to obtain the feature parameters of the width and height dimensions of the image to be processed; Perform element-wise multiplication on the left feature image and the right feature image to obtain the attention feature image.
3. The method according to claim 1, characterized in that, The step of obtaining the pedestrian candidate region in the sample image according to the first intersection-over-union ratio of the candidate region of each target feature image and the standard region includes: Filter the target candidate regions in the candidate regions through a confidence reset algorithm according to the first intersection-over-union ratio and the reset intersection-over-union ratio, and the reset intersection-over-union ratio is the intersection-over-union ratio used to eliminate redundant candidate regions; Use the target candidate regions as the pedestrian candidate regions in the sample image.
4. The method according to claim 3, characterized in that, The step of filtering the target candidate regions in the candidate regions through a confidence reset algorithm according to the first intersection-over-union ratio and the reset intersection-over-union ratio includes: Put the candidate regions with the first intersection over union (IoU) greater than the first preset value into set A; Obtain the candidate region with the largest first IoU in set A, put it into set B, and obtain the remaining candidate regions in set A; Determine the reset IoU corresponding to each remaining candidate region, put the candidate region with the largest reset IoU into set B, and eliminate the candidate regions with a reset IoU less than the first preset value, and obtain the remaining candidate regions in set A; Repeat the step of determining the reset IoU corresponding to each remaining candidate region until set A is empty; Take the candidate regions in set B as the target candidate regions.
5. The method according to claim 4, characterized in that, The step of determining the reset IoU corresponding to each remaining candidate region, putting the candidate region with the largest reset IoU into set B, and eliminating the candidate regions with a reset IoU less than the first preset value includes: Obtain the reset second IoU of each remaining candidate region according to the ratio of the intersection to the union of the new candidate region with the largest IoU in set B and each remaining candidate region; If the second intersection over union of each of the remaining candidate regions is greater than a second preset value, weight reduction processing is performed on the first intersection over union of each of the remaining candidate regions to obtain a reset third intersection over union of each of the remaining candidate regions, where the third intersection over union is the ratio of the first intersection over union and the k-th power of, where k is the ratio of the square of the first intersection over union to the empirical constant σ; If the reset third IoU of each remaining candidate region is less than the first preset value, eliminate the candidate regions with a reset third IoU less than the first preset value; Sort the remaining candidate regions according to the target IoU, and put the candidate region with the largest target IoU into set B, where the target IoU includes the reset second IoU and the reset third IoU.
6. The method according to claim 1, characterized in that, The step of training the pedestrian detection model according to the IoU between the pedestrian candidate region and the standard region includes: Judge whether the IoU is greater than the preset IoU; If not, update the parameters in the pedestrian detection model, obtain the new IoU, and judge whether the new IoU is greater than the preset IoU. If not, repeat the step of parameter update until the IoU is greater than the preset IoU; If so, the training of the pedestrian detection model is completed.
7. A pedestrian detection method, characterized in that, It includes: Obtain the image to be processed; Input the image to be processed into the pedestrian detection model, and obtain the pedestrian candidate regions in the image to be processed output by the pedestrian detection model. The pedestrian detection model is trained by the method described in any one of the above 1 to 6.
8. A pedestrian detection model training device, characterized in that, It includes: An acquisition module for acquiring a plurality of sample images and the standard region of each sample image; A first processing module for inputting the sample image into the feature extraction layer of the feature extraction unit to obtain a first feature image; A second processing module, wherein the feature extraction unit further includes a plurality of sampling layers; for the first sampling layer, input the first feature image into the attention model corresponding to the first sampling layer, and obtain the attention feature image output by the attention model, and input the attention feature image into the residual model corresponding to the first sampling layer to obtain a second feature image; The second processing module is further configured to, for each of the second to the Nth sampling layers: input the feature image output by the previous sampling layer of the sampling layer into the attention model, and obtain the attention feature image output by the attention model; input the attention feature image into the residual model corresponding to the sampling layer, and obtain the feature image output by the sampling layer; wherein, the ith sampling layer outputs the (i + 1)th feature image, and i is an integer greater than or equal to 2 and less than or equal to N; the multiple target feature images include the second to the (N + 1)th feature images respectively output by the second to the (N + 1)th sampling layers; the training module is configured to input the multiple target feature images into the classification and detection unit, and obtain the candidate regions of each target feature image output by the classification and detection unit, where the candidate region is the region where the classification result of the classification and detection unit indicates the presence of a pedestrian; The first determination module is configured to obtain the pedestrian candidate region in the sample image according to the first intersection-over-union ratio of the candidate region of each target feature image to the standard region, where the first intersection-over-union ratio is the ratio of the intersection to the union of the candidate region of the target feature image and the standard region; The second determination module is configured to train the pedestrian detection model according to the intersection-over-union ratio of the pedestrian candidate region to the standard region.
9. A pedestrian detection device, characterized in that, Comprising: An acquisition module, configured to acquire an image to be processed; A processing module, configured to input the image to be processed into the pedestrian detection model, and obtain the pedestrian candidate region in the image to be processed output by the pedestrian detection model, where the pedestrian detection model is trained by the method according to any one of the above 1-6.
10. An electronic device, characterized in that, Comprising: A memory, configured to store a computer program; A processor, configured to execute the computer program to implement the method according to any one of claims 1-6, or to implement the method according to claim 7.
11. A computer-readable storage medium, characterized in that, A computer program is stored thereon, and the computer program is executed by the processor to implement the method according to any one of claims 1-6, or to implement the method according to claim 7.
Citation Information
Patent Citations
Image processing method and device, electronic equipment and storage medium
CN109829501A
Pedestrian detection method and device, electronic device and storage medium
CN114067186A