A highway pedestrian detection method based on super-resolution

By applying super-resolution technology and feature point detection in pedestrian detection, the efficiency of traditional detection technology in low resolution and complex scenarios is solved, and higher detection accuracy and lower false alarm rate are achieved.

CN114694100BActive Publication Date: 2025-05-23ZHEJIANG UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202210375798.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-04-11
Publication Date
2025-05-23
Estimated Expiration
2042-04-11

AI Technical Summary

Technical Problem

Traditional pedestrian detection technology is difficult to achieve better detection results in low resolution and complex scenarios, resulting in high false alarm rate and low detection accuracy.

Method used

The highway pedestrian detection method based on super resolution is used to extract the fuzzy kernel through the KernelGAN network, generate a set of fuzzy kernels, and use the super resolution generation adversarial network (ESRGAN) to recover high-resolution images from low-resolution images. The feature points are extracted in combination with the feature point detection network, and the total weight is calculated to determine whether they are pedestrians.

Benefits of technology

It improves the accuracy and adaptability of pedestrian detection, reduces the false alarm rate, and can more effectively detect pedestrians on the highway and adapt to different application scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114694100B_ABST
    Figure CN114694100B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for detecting pedestrians on highways based on super-resolution. A bicubic downsampling operation is performed on a data set to obtain a corresponding high-resolution image, and blur kernels and noise are extracted from the data set to construct a low-resolution image. The obtained high-resolution image and low-resolution image are used to train a super-resolution generative adversarial network to obtain a super-resolution generator. The low-resolution image to be identified is input into the trained super-resolution generator, and then a feature point detection network is used to extract feature points, and the total weight is calculated. If the calculated total weight is greater than the set weight threshold, it is judged to be a pedestrian. The present invention uses a super-resolution algorithm to supplement detailed information for small-sized pedestrians, and uses human feature point information for further screening, thereby improving detection accuracy, reducing false alarm rate, and adapting to different application scenarios.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application belongs to the field of image recognition and computer vision technology, and specifically relates to a method for detecting pedestrians on a highway based on super-resolution. Background Art

[0002] As we all know, the presence of pedestrians on highways is very dangerous and can easily lead to serious traffic accidents such as rear-end collisions, resulting in casualties and economic losses. This special situation involving human life is also the most concerning issue in highway management. Therefore, detecting pedestrians on highways and issuing timely warnings are of great significance in avoiding major traffic accidents on highways.

[0003] Traditional pedestrian detection technology may misreport some objects on the road that have similar features to people as pedestrians, and the false alarm rate increases in some bad weather conditions. At the same time, some pedestrians at a distance usually have a smaller resolution, that is, small-sized pedestrians, which creates a greater obstacle to accurate detection. Therefore, for targets with low resolution and in some complex scenes, traditional target detectors are difficult to achieve good detection results. Summary of the invention

[0004] The purpose of this application is to provide a super-resolution-based highway pedestrian detection method to overcome the problem that traditional target detectors are difficult to achieve good detection results.

[0005] In order to achieve the above purpose, the technical solution of this application is as follows:

[0006] A method for detecting pedestrians on a highway based on super-resolution, comprising:

[0007] Obtain a training data set T, perform a bicubic downsampling operation on each pedestrian image in the data set T, and obtain a corresponding high-resolution image;

[0008] The KernelGAN network is used to extract the blur kernel from the data set T and generate a blur kernel set;

[0009] A cropped image is obtained from the image in the dataset T according to a preset step size, and the cropped image is converted into a grayscale image. If the variance of the converted grayscale image is less than the variance threshold, it is put into the noise set;

[0010] A blur kernel for degradation operation is randomly extracted from the blur kernel set, and the extracted blur kernel is used to perform blur operation on the high-resolution image, and down-sampled. After down-sampling, noise randomly selected from the noise set is injected to obtain the corresponding low-resolution image;

[0011] Using the high-resolution image and the low-resolution image as training samples to train a super-resolution generative adversarial network to obtain a super-resolution generator;

[0012] The low-resolution image to be identified is input into the trained super-resolution generator, and then the feature point detection network is used to extract the feature points and calculate the total weight. If the calculated total weight is greater than the set weight threshold, it is judged as a pedestrian.

[0013] Furthermore, the loss function of the super-resolution generative adversarial network is as follows:

[0014] L total =μ pixel L pixel +μ per L per +μ adv L adv +μ kp L kp

[0015] L pixel =E x ||X f -HR|| 1

[0016] L adv =-E x [log(1-D(HR,x f ))-E x [log(D(x f ,HR))]]

[0017]

[0018] L kp =|Ho(HR)-Ho(X f )|

[0019] Among them, L total is the total loss function, D represents the discriminator, G represents the generator, and E x represents the mean value, X f represents the image generated by the generator, represents the feature map extracted by ResNet50, W and H are the width and height of the feature map, Ho represents the hourglass feature point extraction network, μ pixel , μ per , μ adv and μ kp is the weight of the corresponding loss, L pixel Represents pixel loss, L per represents the perceptual loss, L adv represents the adversarial loss, L kpRepresents feature point loss.

[0020] Furthermore, the feature point detection network is used to extract feature points and calculate the total weight, including:

[0021] After extracting feature points using the feature point detection network, the extracted feature points are divided into two categories, namely head feature points and body feature points. Then the number of head feature points and body feature points is counted, and corresponding weights are assigned to the head feature points and body feature points. The product of the number of head feature points and body feature points and their corresponding weights is calculated respectively, and then the sum is summed to obtain the total weight.

[0022] Furthermore, the noise randomly selected from the noise set further includes:

[0023] The randomly selected noise is randomly cropped to the same size as the low-resolution image.

[0024] Furthermore, the head feature points include a left eye, a right eye, a nose, a left ear, and a right ear.

[0025] Furthermore, the body part feature points include left shoulder, right shoulder, left elbow, right elbow, left wrist, right wrist, left hip, right hip, left knee, right knee, left ankle and right ankle.

[0026] The present application proposes a method for detecting pedestrians on highways based on super-resolution. Different from traditional pedestrian detectors, it has the characteristics of low false alarm rate and strong scene adaptability, and can detect pedestrians on highways more accurately. The super-resolution algorithm is used to supplement the detailed information of small-sized pedestrians, and the human feature point information is used for further screening, which improves the detection accuracy, reduces the false alarm rate, and can adapt to different application scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] Figure 1 This is a flow chart of the highway pedestrian detection method based on super-resolution in this application;

[0028] Figure 2 This is a schematic diagram of the training process of an embodiment of the present application;

[0029] Figure 3 This is a schematic diagram of the detection process of an embodiment of the present application. DETAILED DESCRIPTION

[0030] In order to make the purpose, technical solution and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0031] In one embodiment, Figure 1As shown in the figure, a highway pedestrian detection method based on super-resolution is proposed, including:

[0032] Step S1: Obtain a training data set T, perform a bicubic downsampling operation on each pedestrian image in the data set T, and obtain a corresponding high-resolution image.

[0033] The present application first obtains a data set as a training sample. In order to train the super-resolution generative adversarial network (ESRGAN: Enhanced Super-Resolution Generative Adversarial Networks) used in the present application for pedestrian detection, it is necessary to generate corresponding high-resolution images and low-resolution images according to each image in the data set, and use the high-resolution images and low-resolution images to train the super-resolution generative adversarial network.

[0034] Specifically, a bicubic downsampling operation is performed on the dataset T to obtain a high-resolution image (HR image):

[0035] I HR =T*k bic

[0036] where k bic represents the bicubic downsampling operation. Bicubic downsampling uses the weighted sum of neighboring pixel values ​​for interpolation to achieve the cleaning effect. HR is the set of HR images generated after cleaning.

[0037] This application uses a bicubic downsampling operation, which is mainly used to remove noise in the sample set to make the image clearer. The data set directly uses higher-resolution images to obtain better high-resolution images after denoising.

[0038] Step S2: Use the KernelGAN network to extract the blur kernel from the data set T and generate a blur kernel set.

[0039] This application uses the KernelGAN network to extract blur kernels from the sample set T. The input is the image in the data set T, and the output is the corresponding blur kernel obtained by learning. The blur kernel is placed in the blur kernel set.

[0040] The size changes of the discriminator network feature graph of KernelGAN in this embodiment are shown in the following table:

[0041]

[0042] Table 1

[0043] Here, p is the padding of the convolution and s is the stride of the convolution.

[0044] The size changes of the KernelGAN generator network structure feature map are shown in the following table:

[0045]

[0046] Table 2

[0047] It should be noted that the trained KernelGAN generator can obtain the corresponding blur kernel by convolving all the convolution layers in the generator in sequence according to the images in the input dataset, which will not be described here.

[0048] Step S3: obtain a cropped image from the image in the data set T according to a preset step size, convert the cropped image into a grayscale image, and if the variance of the converted grayscale image is less than the variance threshold, put it into the noise set.

[0049] This step extracts noise from the data set T and generates a noise set. Specifically, for any image belonging to the data set T, a cropped image with a set length and width is cut with a set step size, and then the cropped image is converted into a grayscale image. The variance of the grayscale image is calculated to determine whether the variance of the grayscale image is less than the variance threshold. The formula is:

[0050] σ(L(Crop t )) <v

[0051] Where σ represents the variance of the grayscale image, L represents the grayscale image, and v represents the set variance threshold.

[0052] If the grayscale image variance of the cropped image is less than the variance threshold, the cropped image is taken as noise and put into the noise set, otherwise the cropped image is discarded.

[0053] Step S4, randomly extracting a blur kernel for degradation operation from the blur kernel set, using the extracted blur kernel to perform blur operation on the high-resolution image, and down-sampling, injecting noise randomly selected from the noise set after down-sampling, and obtaining a corresponding low-resolution image.

[0054] Specifically, the blur kernel set is K, and the blur kernel k used for the degradation operation is randomly selected from K. m to I HR The images in the data set are blurred and then downsampled to their original size. Then inject a noise crop randomly selected from the noise set l (Crop l Randomly crop to the same size as the LR image) to obtain the corresponding low-resolution (ie LR) image. The above process is expressed as:

[0055] I LR =(I HR *km )↓ s +Crop l , m∈{1,2,...,n}, l∈[1,θ]

[0056] Among them, k m represents the randomly selected blur kernel, s is the downsampling multiple, Crop l represents the extracted noise, I LR is the obtained LR dataset, and θ is the total number of noises in the noise set.

[0057] Through the above steps, we can get the paired I required for super-resolution training LR and I HR Datasets and.

[0058] Step S5: Use the high-resolution image and the low-resolution image as training samples to train a super-resolution generative adversarial network to obtain a super-resolution generator.

[0059] This step trains the ESRGAN network model (ESRGAN, Enhanced Super-Resolution Generative Adversarial Networks, super-resolution generative adversarial network). The network is trained using paired LR and HR samples. The generator obtained by training the network can super-resolve low-resolution images to high-resolution images.

[0060] In a specific embodiment, the loss function (L total ) including pixel loss L pixel , Perceptual loss L per , against the loss L adv And feature point loss L kp .

[0061] The total loss is calculated as follows:

[0062] L total =μ pixel L pixel +μ per L per +μ adv L adv +μ kp L kp

[0063] L pixel =E x ||X f -HR|| 1

[0064] L adv =-Ex [log(1-D(HR,x f ))-E x [log(D(x f ,HR))]]

[0065]

[0066] L kp =|Ho(HR)-Ho(X f )|

[0067] Among them, D represents the discriminator, G represents the generator, and E x represents the mean value, X f Represents the image generated by the generator. represents the feature map extracted by ResNet50, W and H are the width and height of the feature map, Ho represents the hourglass feature point extraction network, μ pixel , μ per , μ adv and μ kp is the weight of the corresponding loss.

[0068] This application uses the hourglass feature point extraction network to extract HR images and X f Extract feature points to calculate feature point loss, increasing the loss L kp , making the trained generator more accurate. At the same time, the perceptual loss is extracted based on ResNet50. The residual structure of ResNet50 can better integrate the features of each layer and extract the features of pedestrians more comprehensively, thereby better guiding the training of the super-resolution network, making the pedestrians after super-resolution have clearer features.

[0069] In a specific embodiment, the size change of the discriminator network feature map of the ESRGAN network model is shown in the following table, where h and w are the height and width of the discriminator input image respectively:

[0070]

[0071]

[0072] Table 3

[0073] In a specific embodiment, the number of output channels of the first convolution layer of the generator of the ESRGAN network model is 64. The result after passing through 23 RRDB modules and one convolution layer is added to the result of the first convolution layer, and then undergoes two upsampling and four convolution operations to obtain a high-resolution HR image.

[0074] In a specific embodiment, the learning rate adjustment strategy used in training is MultiStepLR, as shown in the following formula:

[0075] 1. new =lr*γ milestones[epoch]

[0076] Where lr is the learning rate before updating, lr new is the updated learning rate, epoch is the number of rounds of the current iteration, milestones is the corresponding epoch value for which the learning rate needs to be adjusted, and γ is the adjustment coefficient.

[0077] After the super-resolution generative adversarial network is trained, its generator is the super-resolution generator. It should be noted that the specific network structure of the generator and the discriminator in the super-resolution generative adversarial network ESRGAN in this embodiment has many existing good cases in the technical field, which will not be repeated here.

[0078] Step S6: input the low-resolution image to be identified into the trained super-resolution generator, then use the feature point detection network to extract feature points and calculate the total weight. If the calculated total weight is greater than the set weight threshold, it is judged to be a pedestrian.

[0079] In actual applications, according to the size and position distribution of pedestrians on the highway, considering the performance overhead, only the images in the detection frame of a certain size and position are processed as low-resolution images to be identified. The low-resolution image to be identified is input into the trained super-resolution generator, and the generated image after super-resolution processing is output. Then, the feature point detection network is used to extract feature points from the generated image.

[0080] This application uses the hourglass feature point detection network to detect feature points, and then calculates a total weight based on the number of feature points, and compares the calculated total weight with the set weight threshold. If the calculated total weight is greater than the set weight threshold, it is judged as a pedestrian. Otherwise, it is not considered a pedestrian.

[0081] In a specific application, a total weight is calculated according to the number of feature points. The total weight can be directly calculated based on the total number of feature points, or the extracted feature points can be classified for calculation.

[0082] Preferably, a feature point detection network is used to extract feature points and calculate the total weight, including:

[0083] After extracting feature points using the feature point detection network, the extracted feature points are divided into two categories, namely head feature points and body feature points. Then the number of head feature points and body feature points is counted, and corresponding weights are assigned to the head feature points and body feature points. The product of the number of head feature points and body feature points and their corresponding weights is calculated respectively, and then the sum is summed to obtain the total weight.

[0084] For example, the image in the detection frame is input into the trained super-resolution generator for super-resolution amplification, and the hourglass-based feature point detection network is further used to extract 17 feature points from the detection results, which are divided into two sets:

[0085] Set 1: {left eye, right eye, nose, left ear, right ear};

[0086] Set 2: {left shoulder, right shoulder, left elbow, right elbow, left wrist, right wrist, left hip, right hip, left knee, right knee, left ankle, right ankle}.

[0087] Then the total weight α is calculated by the following formula:

[0088]

[0089] in is the weight corresponding to the feature points in the two sets, count 1 , count 2 are the number of extracted feature points belonging to set 1 and set 2 respectively. The detection box with α greater than the weight threshold δ is judged as a pedestrian.

[0090] In one embodiment, Figure 2 As shown, the process of training the network using training samples is shown, corresponding to Figure 1 Steps S1-S5 in . For the collected data set, on the one hand, bicubic downsampling is performed for data cleaning to obtain high-resolution HR images, and on the other hand, KernelGAN is used to extract blur kernels, and noise that meets the conditions is extracted by processing the data set. Then, the HR image is degraded using randomly extracted blur kernels and noise to obtain low-resolution LR images. Finally, the HR image and LR image are input into ESRGAN to train the GAN generator. During the training process, the learning rate is adaptively adjusted and the loss is calculated, the network parameters are updated, and the training is completed. During the training process, the initial learning rate is set first, and the learning rate is reduced by 10 times every 750 rounds. When extracting qualified noise from the original data set, the step size is set to 128, the cropped image size is set to 128, and the variance threshold v is set to 20. The downsampling multiple s for downsampling the HR image in step S4 is 4. μ pixel Take 0.01, μ per Take 1, μadv Take 0.005, μ kp The value of γ is 0.5, the values ​​of milestones are [5000, 10000, 20000, 30000], and the initial learning rate is 0.0001.

[0091] In another embodiment, Figure 3 As shown, the process of detecting the image to be detected (image to be identified) is shown, corresponding to Figure 1 Step S6 in the process. First, the image to be inspected is detected using the detection network to obtain a detection frame. It is easy to understand that detecting a detection frame from an image is a relatively mature technology in this field and will not be described in detail here. Then, the image is cropped according to the detection result to obtain a cropped image, and the cropped image is input into the GAN generator (super-resolution generator) to obtain an HR image. Finally, feature point detection is performed on the HR image, and the total weight is calculated to determine whether the total weight is greater than the set weight threshold (for example, 9). The detection result is obtained by judging whether it is a pedestrian or a non-pedestrian. In this embodiment, Take 2, Take 1 and δ take 9.

[0092] The above-mentioned embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the invention patent. It should be pointed out that, for a person of ordinary skill in the art, several variations and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the attached claims.

Claims

1. A highway pedestrian detection method based on super-resolution, It is characterized in that The super-resolution-based highway pedestrian detection method comprises: Obtain a training data set T, perform a bicubic downsampling operation on each pedestrian image in the data set T, and obtain a corresponding high-resolution image; The KernelGAN network is used to extract the blur kernel from the data set T and generate a blur kernel set; A cropped image is obtained from the image in the dataset T according to a preset step size, and the cropped image is converted into a grayscale image. If the variance of the converted grayscale image is less than the variance threshold, it is put into the noise set; A blur kernel for degradation operation is randomly extracted from the blur kernel set, and the extracted blur kernel is used to perform blur operation on the high-resolution image, and down-sampled. After down-sampling, noise randomly selected from the noise set is injected to obtain the corresponding low-resolution image; Using the high-resolution image and the low-resolution image as training samples to train a super-resolution generative adversarial network to obtain a super-resolution generator; The low-resolution image to be identified is input into the trained super-resolution generator, and then the feature point detection network is used to extract the feature points and calculate the total weight. If the calculated total weight is greater than the set weight threshold, it is judged as a pedestrian. The method of extracting feature points using a feature point detection network and calculating the total weight includes: After extracting feature points using the feature point detection network, the extracted feature points are divided into two categories, namely head feature points and body feature points. Then the number of head feature points and body feature points is counted, and corresponding weights are assigned to the head feature points and body feature points. The product of the number of head feature points and body feature points and their corresponding weights is calculated respectively, and then the sum is summed to obtain the total weight.

2. The method for detecting pedestrians on highways based on super-resolution according to claim 1, It is characterized in that The loss function of the super-resolution generative adversarial network is as follows: L total =μ pixel L pixel +m per L per +m adv L adv +m kp L kp L pixel =E x ||X f -HR|| 1 L adv =-E x [log(1-D(HR,x f ))-E x [log(D(x f ,HR))]] THE kp =|I(HR)-I(X f )| Among them, L total is the total loss function, D represents the discriminator, G represents the generator, and E x represents the mean, X f represents the image generated by the generator, represents the feature map extracted by ResNet50, W and H are the width and height of the feature map, Ho represents the hourglass feature point extraction network, μ pixel , μ per , μ adv and μ kp is the weight of the corresponding loss, L pixel Represents pixel loss, L per represents the perceptual loss, L adv represents the adversarial loss, L kp Represents feature point loss.

3. The method for detecting pedestrians on highways based on super-resolution according to claim 1, It is characterized in that The noise randomly selected from the noise set also includes: The randomly selected noise is randomly cropped to the same size as the low-resolution image.

4. The method for detecting pedestrians on highways based on super-resolution according to claim 1, It is characterized in that The head feature points include a left eye, a right eye, a nose, a left ear, and a right ear.

5. The method for detecting pedestrians on highways based on super-resolution according to claim 1, It is characterized in that The body part feature points include a left shoulder, a right shoulder, a left elbow, a right elbow, a left wrist, a right wrist, a left hip, a right hip, a left knee, a right knee, a left ankle and a right ankle.