Road pedestrian detection method based on deep reinforcement learning and related equipment
Through the deep reinforcement learning method, combined with deep learning and reinforcement learning agents, the EIoU reward function and YOLOv5 model are used to improve the accuracy and efficiency of pedestrian detection, solve the problem of unstable detection accuracy in the existing technology, and enhance the obstacle avoidance ability of autonomous driving vehicles.
Patent Information
- Application Number
- CN202311588345.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-24
- Publication Date
- 2025-07-22
AI Technical Summary
The existing pedestrian detection methods based on deep learning are not stable enough in detection accuracy, which affects the safety of autonomous driving.
The method of deep reinforcement learning is adopted, through the initial positioning of deep learning and the fine positioning of deep reinforcement learning agents, pedestrian detection is used using the EIoU reward function, and combined with the YOLOv5 model and DQN algorithm, the accuracy and efficiency of pedestrian detection are improved.
It improves the accuracy and efficiency of pedestrian detection, enhances the obstacle avoidance ability of autonomous driving vehicles, reduces the error when the prediction box and the real box do not intersect, and improves the convergence speed of the model.
Smart Images

Figure CN120356174A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of autonomous driving, and in particular, to a road pedestrian detection method and related devices based on deep reinforcement learning. Background Art
[0002] Today, with the rapid development of information technology, computers are becoming more and more closely integrated with human production and life. With the rapid development of computer vision technology, the object detection level based on computer vision has made great progress and is applied to all aspects of production and life. For the unmanned vehicle autonomous driving technology, the primary guarantee is human safety, and pedestrian detection is an essential key technology.
[0003] The pedestrian detection method based on deep learning performs excellently in terms of detection speed, but the current detection method is not stable in terms of detection accuracy. Summary of the Invention
[0004] In view of the above problems, the present invention provides a road pedestrian detection method and related devices based on deep reinforcement learning, mainly aiming to solve the problem that the current pedestrian detection method based on deep learning performs excellently in terms of detection speed, but the current detection method is not stable in terms of detection accuracy.
[0005] To solve the above at least one technical problem, in a first aspect, the present invention provides a road pedestrian detection method based on deep reinforcement learning, and the method includes:
[0006] Based on the collected visible light images around the vehicle, perform an initial localization of pedestrians through deep learning to obtain the first localization box information containing pedestrians;
[0007] According to the position coordinates of pedestrians in the first localization box, use a deep reinforcement learning agent to perform a fine localization of pedestrians to obtain the second localization box information containing pedestrians, where the reward function in the deep reinforcement learning agent is EIoU;
[0008] Transmit the second localization box information to the back end of the autonomous driving vehicle through human-computer interaction, so that the autonomous driving vehicle can perform obstacle avoidance operations.
[0009] Optionally, the step of performing an initial localization of pedestrians through deep learning based on the collected visible light images around the vehicle to obtain the first localization box information containing pedestrians includes:
[0010] Based on the collected visible light images around the vehicle, perform an initial localization of pedestrians through the YOLOv5 model to obtain the first localization box information containing pedestrians.
[0011] Optionally, the neural network model adopted by the deep reinforcement learning agent is a DQN model.
[0012] Optionally, the human-computer interaction application program is developed based on the flutter cross-platform.
[0013] Optionally, the action set of the deep reinforcement learning agent includes actions of moving the box along the horizontal and vertical axes, changing the scale, modifying the aspect ratio, and terminating the search.
[0014] Optionally, it further includes:
[0015] Obtaining the obstacle avoidance evaluation score of the vehicle user after ending a driving trip;
[0016] Uploading the video data and the corresponding second positioning box information during the driving trip with the obstacle avoidance evaluation score higher than the preset score to the cloud as an updated training data set.
[0017] Optionally, it further includes:
[0018] In the case that multiple sets of second positioning box information obtained from multiple frames of visible light images indicate that the user is moving, performing face recognition in the second positioning box information to obtain the attention of pedestrians to the vehicle;
[0019] Executing a speed concession strategy when the attention is greater than the preset value, otherwise executing a stop-and-wait strategy.
[0020] In a second aspect, an embodiment of the present invention further provides a road pedestrian detection device based on deep reinforcement learning, including:
[0021] An acquisition unit, configured to perform initial positioning of pedestrians through deep learning based on the collected visible light images around the vehicle to obtain first positioning box information including pedestrians;
[0022] A first recognition unit, configured to perform fine positioning of pedestrians using a deep reinforcement learning agent according to the position coordinates of the pedestrians in the first positioning box to obtain second positioning box information including pedestrians, wherein the reward function in the deep reinforcement learning agent is EIoU;
[0023] A second recognition unit, configured to transmit the second positioning box information to the back end of the autonomous vehicle through human-computer interaction so that the autonomous vehicle performs obstacle avoidance operations.
[0024] To achieve the above object, according to a third aspect of the present invention, there is provided a computer-readable storage medium, where the computer-readable storage medium includes a stored program, and when the program is executed by a processor, the above-mentioned road pedestrian detection method based on deep reinforcement learning is implemented.
[0025] To achieve the above object, according to the fourth aspect of the present invention, there is provided an electronic device, including at least one processor and at least one memory connected to the above-mentioned processor; wherein, the above-mentioned processor is used to call program instructions in the above-mentioned memory and execute the above-mentioned road pedestrian detection method based on deep reinforcement learning.
[0026] By means of the above technical solution, the present invention provides a road pedestrian detection method and related devices based on deep reinforcement learning. The above method performs an initial positioning of pedestrians through deep learning on the visible light images collected around the vehicle to obtain the first positioning frame information containing pedestrians; according to the position coordinates of the pedestrians in the first positioning frame, a deep reinforcement learning agent is used to perform fine positioning on the pedestrians to obtain the second positioning frame information containing pedestrians, wherein the reward function in the deep reinforcement learning agent is EIoU; the second positioning frame information is transmitted to the backend of the autonomous vehicle through human-computer interaction so that the autonomous vehicle can perform obstacle avoidance operations. Thus, according to the camera feedback system of the autonomous intelligent vehicle itself, real-time road scene video frames and visible light pictures around the vehicle body are obtained. Deep learning is used to detect the visible light pictures and analyze the rough positioning position coordinates of pedestrians. According to the rough positioning pedestrian detection frame, deep reinforcement learning is used for real-time fine positioning detection. It can enable the autonomous intelligent vehicle to complete the detection and positioning of pedestrian targets, improve the detection accuracy and detection efficiency, and is beneficial to enhancing the obstacle avoidance ability of autonomous vehicles. The reward function EIoU is selected to replace the IoU reward function in the traditional model. EIoU can solve the situation where the prediction box and the real box do not intersect, reduce the error generated when the prediction box and the real box are in an inclusion state, and greatly improve the convergence speed.
[0027] Correspondingly, the road pedestrian detection device, electronic system and computer-readable storage medium provided by the embodiments of the present invention also have the above technical effects.
[0028] The above description is only an overview of the technical solution of the present invention. In order to be able to understand the technical means of the present invention more clearly, it can be implemented according to the content of the specification. And in order to make the above and other objects, features and advantages of the present invention more obvious and understandable, the following specific embodiments of the present invention are specifically given. Description of the Drawings
[0029] By reading the following detailed description of the preferred embodiments, various other advantages and benefits will become clear to those of ordinary skill in the art. The drawings are only for the purpose of showing the preferred embodiments and are not considered to be a limitation of the present invention. Moreover, throughout the drawings, the same reference numerals are used to represent the same components. In the drawings:
[0030] Figure 1Shows a schematic flow chart of a road pedestrian detection method based on deep reinforcement learning provided by an embodiment of the present invention;
[0031] Figure 2 Shows a schematic block diagram of the training process of reinforcement learning in a road pedestrian detection method based on deep reinforcement learning provided by an embodiment of the present invention;
[0032] Figure 3 Shows a schematic diagram of the actions of the agent action set in a road pedestrian detection method based on deep reinforcement learning provided by an embodiment of the present invention;
[0033] Figure 4 Shows a schematic effect diagram of the agent window positioning process in a road pedestrian detection method based on deep reinforcement learning provided by an embodiment of the present invention;
[0034] Figure 5 Shows a schematic block diagram of the composition of a road pedestrian detection device based on deep reinforcement learning provided by an embodiment of the present invention;
[0035] Figure 6 Shows a schematic block diagram of the composition of a road pedestrian detection electronic device based on deep reinforcement learning provided by an embodiment of the present invention. Detailed implementation manners
[0036] The exemplary embodiments of the present invention will be described in more detail with reference to the accompanying drawings. Although the exemplary embodiments of the present invention are shown in the drawings, it should be understood that the present invention can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided so that the present invention can be more thoroughly understood and the scope of the present invention can be completely conveyed to those skilled in the art.
[0037] To solve the problem that the current pedestrian detection method based on deep learning performs excellently in detection speed, but the current detection method is not stable in detection accuracy, an embodiment of the present invention provides a road pedestrian detection method based on deep reinforcement learning, as Figure 1 shown, the method includes:
[0038] S101, based on the collected visible light images around the vehicle, initially locate pedestrians through deep learning to obtain the first positioning frame information including pedestrians.
[0039] Exemplarily, the above visible light images can be collected by cameras set at multiple angles in the vehicle.
[0040] S102. Use the position coordinates of the pedestrian in the first positioning box and the deep reinforcement learning agent to perform fine positioning on the pedestrian to obtain the second positioning box information containing the pedestrian, where the reward function in the deep reinforcement learning agent is EIoU.
[0041] Exemplarily, a reinforcement learning environment for visible light pedestrian image data can be further built. The reinforcement learning environment can include dimensions, boundaries, the position, shape, and size of the detection target, as well as the state s, action set a, and reward function r of the agent in reinforcement learning.
[0042] Exemplarily, the state set s can be represented by the pixel values in the two-dimensional detection box b = [x1, y1, x2, y2]. The transformation action can discretely change the detection box according to the formula relative to the current scale factor: a w = a * (x1 - x2), a h = a * (y1 - y2) where a ∈ [0, 1]. Increase or decrease a w or a h to change the x and y coordinates of the diagonal points of the detection box.
[0043] Exemplarily, the action set A consists of nine actions for controlling the detection box. The reward function R can use EIoU instead of the IoU reward function in the traditional model. EIoU can solve the situation where the predicted box and the ground truth box do not intersect, reduce the error generated when the predicted box and the ground truth box are in an inclusion state, and greatly improve the convergence speed. When the agent executes an action and transfers from state s to state s', the reward function is R a (s, s').
[0044] It can be understood that in reinforcement learning, the reward function provides immediate feedback to the agent for performing a specific action in a certain state. It is the criterion for evaluating the quality of actions. IoU (Intersection over Union) is a common performance metric in object detection, used to measure the overlap between the predicted bounding box and the ground truth bounding box. EIoU is an improved version of IoU, namely Enhanced IoU. In IoU, if the predicted bounding box and the ground truth bounding box do not overlap, the IoU value is 0, which may be unfavorable for the propagation of gradients and the learning of the model. EIoU is proposed to solve this problem and can provide effective gradients even in the case of non-intersection. In some cases, the predicted bounding box may completely contain the ground truth bounding box or vice versa. In this case, IoU may give a relatively high value, but in fact, the localization is not accurate. EIoU will give a more reasonable evaluation in this case. Since EIoU provides a more refined and effective gradient signal, it can help the model learn and converge faster. This describes a reinforcement learning scenario where the agent transfers from the current state s to the new state s′ after performing an action. This formula shows that the reward is determined based on the change in the Enhanced IoU value. Where b represents the predicted bounding box in the current state, b’ represents the predicted bounding box after the action is performed, and g represents the ground truth bounding box. If the EIoU value increases after the action is performed, that is, the matching degree between the predicted bounding box and the ground truth bounding box becomes better, the agent will receive a positive reward; if the matching degree deteriorates, it will receive a negative reward. The sign function is a mathematical function that returns 1 if its argument is positive, -1 if it is negative, and 0 if it is zero. This reward setting encourages the agent to take actions that increase the EIoU value.
[0045] S103, transmit the second positioning box information to the back end of the autonomous vehicle through human-computer interaction, so that the autonomous vehicle can perform obstacle avoidance operations.
[0046] With the above technical solution, the road pedestrian detection method based on deep reinforcement learning provided by the present invention performs an initial positioning of pedestrians through deep learning based on the collected visible light images around the vehicle to obtain the first positioning frame information including pedestrians; uses a deep reinforcement learning agent to perform fine positioning on pedestrians according to the position coordinates of the pedestrians in the first positioning frame to obtain the second positioning frame information including pedestrians, wherein the reward function in the deep reinforcement learning agent is EIoU; transmits the second positioning frame information to the back end of the autonomous vehicle through human-computer interaction so that the autonomous vehicle can perform an obstacle avoidance operation. Thus, according to the camera feedback system of the autonomous intelligent vehicle itself, real-time road scene video frames and visible light pictures around the vehicle body are obtained. The deep learning is used to detect the visible light pictures and analyze the rough positioning position coordinates of pedestrians. The deep reinforcement learning is used for real-time fine positioning detection according to the rough positioning pedestrian detection frame. The autonomous intelligent vehicle can complete the detection and positioning of pedestrian targets, improve the detection accuracy and detection efficiency, and is beneficial to enhancing the obstacle avoidance ability of autonomous vehicles. The reward function EIoU is selected to replace the IoU reward function in the traditional model. EIoU can solve the situation where the predicted box and the real box do not intersect, reduce the error generated when the predicted box and the real box are in an inclusion state, and greatly improve the convergence speed.
[0047] In one embodiment, the neural network model adopted by the deep reinforcement learning agent is a DQN model.
[0048] Exemplarily, the DQN algorithm is selected due to the effect of the traditional Q-Learning algorithm. The neural network is used to calculate the Q value, and the output of the neural network is used to replace the Q value obtained by looking up the Q value table. Storage process: Update the memory bank. Each time a transition occurs, it will be stored in the memory bank, and then according to the actual situation, the action, reward, and next state are obtained. This improves the speed of the algorithm and reduces the large memory of the traditional algorithm at the same time.
[0049] It is understandable that this refers to the selection of using the Deep Q-Network (DQN) algorithm. DQN is an algorithm that combines traditional Q-Learning and deep learning. The traditional Q-Learning algorithm learns the optimal policy by updating a Q-value table, which stores the expected rewards (Q-values) for performing a certain action in a given state. In DQN, a deep neural network is used to approximate the Q-value function instead of using a table. The neural network can handle higher-dimensional input states, which is necessary for dealing with complex problems. The DQN algorithm no longer looks up a Q-value table to obtain the Q-values, but directly gives the Q-values of each possible action through the input state to the neural network. The DQN algorithm includes an experience replay mechanism that stores the agent's experiences. Each transition - that is, the combination of the current state, the action taken, the reward obtained, and the next state - will be stored in the memory bank. Every time the Agent transitions from one state to another, it will be recorded, including the current state, the action taken, the reward obtained, and the next state. The DQN algorithm randomly samples experiences from the memory bank. This is done to break the temporal correlation between experiences and more evenly cover the entire state-action space. Since the neural network can effectively handle a large number of states, DQN can learn faster than traditional Q-Learning. Since there is no need to store a large Q-value table, DQN reduces the storage requirements through the neural network.
[0050] In one embodiment, based on the collected visible light images around the vehicle, the pedestrians are initially located through deep learning to obtain the first positioning box information including pedestrians, including:
[0051] Based on the collected visible light images around the vehicle, the pedestrians are initially located through the YOLOv5 model to obtain the first positioning box information including pedestrians.
[0052] Exemplarily, the Caltech Pedestrian dataset released by the California Institute of Technology in 2009 can be used, which contains approximately 10 hours of 640*480 30Hz videos, mainly taken by cars on the driving streets. The videos total approximately 250,000 frames, containing 350,000 bounding boxes and annotations of 2,300 pedestrians, where the annotations include the corresponding relationships between the detailed labels of the bounding boxes. This dataset is used to evaluate the proposed pedestrian detection method. If the action selected by the agent results in the prediction box exceeding the boundary after execution, this action will be cancelled and the agent will be punished with a reward value of -1. The positioning result is input to the initial position of the agent in deep reinforcement learning, such as Figure 2 described.
[0053] Exemplarily, the state is represented by a vector group (o, h), where o is a feature vector representing the region where the detection box is located, and h is a vector of historical actions. The state set S includes combinations of all extended actions corresponding to any position of the detection box in the image. Therefore, it is necessary to represent the state in a generalized manner. A pre-trained CNN network is used to extract the feature vector o from the current region. Each action in the historical vector h is represented by a 9-dimensional binary vector. The value corresponding to the action to be taken is set to 1, and the others are set to 0 to record the historical actions. The action set consists of nine actions for controlling the detection box, such as Figure 3 as shown. Actions such as moving the box along the horizontal and vertical axes, changing the scale, modifying the aspect ratio, and terminating the search are set. Such a setting gives the agent four degrees of freedom to transform the detection box during the interaction with the environment.
[0054] Exemplarily, the reward function R is proportional to the improvement degree of the state after the agent selects an action. In previous reinforcement learning object detection frameworks, it is usually measured based on the difference in the intersection over union (IoU). The formula for IoU can be The calculation of the EIoU loss function is divided into: calculating the overlapping loss part of the two boxes, calculating the width and height loss part of the box, and calculating the center distance loss part of the two boxes. The formula for the loss function can be where C w and C h are the width and height of the minimum bounding rectangle of the predicted bounding box and the ground truth bounding box. ρ is the Euclidean distance between two points. w is the width of the predicted bounding box, w gt is the width of the ground truth bounding box. h is the width of the predicted bounding box, h gt is the width of the ground truth bounding box. b is the center distance of the predicted bounding box, b gt is the center distance of the ground truth bounding box. d is the diagonal distance of the minimum bounding rectangle of the predicted bounding box and the ground truth bounding box. EIoU can solve the situation where the predicted box and the ground truth box do not intersect, and can reduce the error generated when the predicted box and the ground truth box are in an inclusion state, and there is a great improvement in the convergence speed. When the agent executes an action and transfers from state s to state s′, the reward function is R a (s, s′). At this time, the reward is as follows: R a (s, s′) = sign(EIoU(b′, g) - EIoU(b, g)).
[0055] It can be understood that the visible light pedestrian image target detection task is a computer vision task aimed at detecting pedestrians in images captured by ordinary visible light cameras. And YOLOv5s is a lightweight model version in the YOLO (You Only Look Once) family for fast target detection. It is first used to make a preliminary prediction of the pedestrian position in the image. The preliminary bounding box obtained by regression is set as the initial position of the reinforcement learning agent window. Here, the "preliminary bounding box obtained by regression" refers to the rectangular box of the pedestrian position predicted by the YOLOv5s model. This box is used as the starting point of the reinforcement learning agent, that is, the area initially concerned by the agent. The detection box is the object of action, and the agent tries to locate the exact position of the pedestrian by performing actions on this box. The agent decides how to move or adjust the detection box according to this algorithm, including scale change (enlarging or reducing the size of the box), aspect ratio change (changing the width-to-height ratio of the box), and movement (changing the position of the box). Through continuous iteration and learning of the DQN algorithm, the pedestrian in the image can ultimately be accurately located (as Figure 4 shown). This method combines the accuracy of deep reinforcement learning positioning and the low false detection rate of YOLOv5s in processing multi-scale target detection, aiming to improve the performance of the overall model. Through the above combination, it is intended to make the model more intelligent, that is, able to detect pedestrians more accurately and efficiently. The ultimate goal is to enable autonomous vehicles to detect and locate pedestrians more accurately and effectively, which can improve the vehicle's obstacle avoidance ability and thus increase driving safety.
[0056] In one embodiment, the human-computer interaction application program is developed based on the flutter cross-platform.
[0057] Exemplarily, according to the fine position positioning information, it is transmitted to the backend of the autonomous vehicle through human-computer interaction, facilitating subsequent obstacle avoidance operations. The yolo format of the final pedestrian positioning detection box data is uploaded to the intelligent vehicle obstacle avoidance system, and a human-computer interaction interface is developed through the flutter cross-platform development platform, and the pedestrian positioning result is displayed through the front-end interface.
[0058] In one embodiment, it further includes:
[0059] Obtaining the obstacle avoidance evaluation score of the vehicle user after ending a driving trip;
[0060] Uploading the video data and the corresponding second positioning box information during the driving trip with the obstacle avoidance evaluation score higher than the preset score to the cloud as an updated training data set.
[0061] Exemplarily, in order to expand the dataset for model training and make the reinforcement learning model more accurate, the evaluation of pedestrian detection and automatic obstacle avoidance behavior after the actual historical driving journey of vehicle users can be carried out. Based on the evaluation scores of users, the video data and the corresponding second positioning box information during the driving journey with higher scores can be used as a video dataset including bounding boxes and pedestrian annotations to rapidly and massively expand the training set of the reinforcement learning model, improving the accuracy of the subsequent model.
[0062] In one embodiment, it further includes:
[0063] In the case where multiple sets of second positioning box information obtained from multiple frames of visible light images indicate that the user is moving, facial recognition is performed on the second positioning box information to obtain the attention of pedestrians to the vehicle;
[0064] When the attention is greater than a preset value, a speed reduction strategy is executed; otherwise, a stop-and-wait strategy is executed.
[0065] Exemplarily, in the case where multiple sets of second positioning box information obtained from multiple frames of visible light images indicate that the user is moving, facial recognition is performed on the second positioning box information, and the attention of pedestrians to the vehicle can be obtained. When the attention is greater than a preset value, it can indicate that the pedestrian is paying attention to this vehicle at this time and moving with an active avoidance behavior. In this case, it is only necessary to decelerate and avoid passing instead of making an emergency stop, avoiding the poor user experience caused by frequent emergency stops of the vehicle during driving. When the attention is less than or equal to the preset value, it can indicate that the pedestrian is moving without noticing this vehicle at this time. Then, in order to avoid active or even passive collisions with pedestrians, a stop-and-wait strategy can be executed to ensure driving safety.
[0066] Further, as an implementation of the above Figure 1 shown method, an embodiment of the present invention also provides a road pedestrian detection device based on deep reinforcement learning for implementing the above Figure 1 shown method. This device embodiment corresponds to the foregoing method embodiment. For the convenience of reading, the details in the foregoing method embodiment will not be described one by one in this device embodiment. However, it should be clear that the device in this embodiment can correspondingly implement all the contents in the foregoing method embodiment. As Figure 5 shown, the device includes: an acquisition unit 21, a first recognition unit 21, and a second recognition unit 23, where
[0067] The acquisition unit 21 is used to perform initial positioning of pedestrians through deep learning based on the collected visible light images around the vehicle to obtain first positioning box information including pedestrians;
[0068] The first recognition unit 22 is configured to use a deep reinforcement learning agent to perform fine positioning on a pedestrian based on the position coordinates of the pedestrian in the first positioning frame, so as to obtain second positioning frame information including the pedestrian, where the reward function in the deep reinforcement learning agent is EIoU;
[0069] The second recognition unit 23 is configured to transmit the second positioning frame information to the rear end of the autonomous vehicle through human-computer interaction, so that the autonomous vehicle performs an obstacle avoidance operation.
[0070] With the above technical solution, the road pedestrian detection device based on deep reinforcement learning provided by the present invention performs an initial positioning of a pedestrian through deep learning based on a collected visible light image around the vehicle to obtain first positioning frame information including the pedestrian; uses a deep reinforcement learning agent to perform fine positioning on the pedestrian based on the position coordinates of the pedestrian in the first positioning frame to obtain second positioning frame information including the pedestrian, where the reward function in the deep reinforcement learning agent is EIoU; transmits the second positioning frame information to the rear end of the autonomous vehicle through human-computer interaction, so that the autonomous vehicle performs an obstacle avoidance operation. Thus, according to the camera feedback system of the autonomous intelligent vehicle itself, a real-time road picture video frame visible light picture around the vehicle is obtained. The position coordinates of the pedestrian in the rough position are analyzed by using deep learning to detect the visible light picture. Real-time fine positioning detection is performed on the rough positioning pedestrian detection frame by using deep reinforcement learning. It can enable the autonomous intelligent vehicle to complete the detection and positioning of pedestrian targets, improve the detection accuracy and detection efficiency, and is beneficial to improving the obstacle avoidance ability of autonomous vehicles. The reward function EIoU is selected to replace the IoU reward function in the traditional model. EIoU can solve the situation where the predicted box and the real box do not intersect, reduce the error generated when the predicted box and the real box are in an inclusion state, and greatly improve the convergence speed.
[0071] The processor includes a kernel, and the kernel retrieves corresponding program units from the memory. One or more kernels can be set, and by adjusting the kernel parameters, a road pedestrian detection method based on deep reinforcement learning can be implemented, which can solve the problem of lacking a better method for coordinating device switches at the current Internet of Things level.
[0072] An embodiment of the present invention provides a computer-readable storage medium. The above computer-readable storage medium includes a stored program, and when the program is executed by a processor, the above road pedestrian detection method based on deep reinforcement learning is implemented.
[0073] An embodiment of the present invention provides a processor. The above processor is used to run a program, where the above program, when running, executes the above road pedestrian detection method based on deep reinforcement learning:
[0074] Based on the collected visible light images around the vehicle, pedestrians are initially located through deep learning to obtain the first localization box information containing pedestrians;
[0075] According to the position coordinates of pedestrians in the first localization box, a deep reinforcement learning agent is used to perform fine localization on pedestrians to obtain the second localization box information containing pedestrians, where the reward function in the deep reinforcement learning agent is EIoU;
[0076] The second localization box information is transmitted to the back end of the autonomous vehicle through human-computer interaction so that the autonomous vehicle can perform obstacle avoidance operations.
[0077] Optionally, the step of initially locating pedestrians through deep learning based on the collected visible light images around the vehicle to obtain the first localization box information containing pedestrians includes:
[0078] Based on the collected visible light images around the vehicle, pedestrians are initially located through the YOLOv5 model to obtain the first localization box information containing pedestrians.
[0079] Optionally, the neural network model adopted by the deep reinforcement learning agent is a DQN model.
[0080] Optionally, the human-computer interaction application program is developed based on the flutter cross-platform.
[0081] Optionally, the action set of the deep reinforcement learning agent includes actions such as moving the box along the horizontal and vertical axes, changing the scale, modifying the aspect ratio, and terminating the search.
[0082] Optionally, it further includes:
[0083] Obtaining the obstacle avoidance evaluation score of the vehicle user after ending a driving trip;
[0084] Uploading the video data and the corresponding second localization box information during the driving trip with the obstacle avoidance evaluation score higher than the preset score to the cloud as an updated training data set.
[0085] Optionally, it further includes:
[0086] In the case where multiple sets of second localization box information obtained from multiple frames of visible light images indicate that the user is moving, facial recognition is performed on the second localization box information to obtain the attention of the pedestrian to the vehicle;
[0087] Execute the speed concession strategy when the attention is greater than the preset value, otherwise execute the stop and wait strategy.
[0088] An embodiment of the present invention provides an electronic device, which includes at least one processor and at least one memory connected to the processor; wherein, the processor is configured to call program instructions in the memory to execute the above-mentioned road pedestrian detection method based on deep reinforcement learning:
[0089] Based on the collected visible light images around the vehicle, perform an initial localization of pedestrians through deep learning to obtain the first localization box information containing pedestrians;
[0090] According to the position coordinates of pedestrians in the first localization box, use a deep reinforcement learning agent to perform fine localization of pedestrians to obtain the second localization box information containing pedestrians, where the reward function in the deep reinforcement learning agent is EIoU;
[0091] Transmit the second localization box information to the backend of the autonomous vehicle through human-computer interaction so that the autonomous vehicle can perform obstacle avoidance operations.
[0092] Optionally, the step of performing an initial localization of pedestrians through deep learning based on the collected visible light images around the vehicle to obtain the first localization box information containing pedestrians includes:
[0093] Based on the collected visible light images around the vehicle, perform an initial localization of pedestrians through the YOLOv5 model to obtain the first localization box information containing pedestrians.
[0094] Optionally, the neural network model adopted by the deep reinforcement learning agent is a DQN model.
[0095] Optionally, the human-computer interaction application program is developed based on the flutter cross-platform.
[0096] Optionally, the action set of the deep reinforcement learning agent includes actions such as moving the box along the horizontal and vertical axes, changing the scale, modifying the aspect ratio, and terminating the search.
[0097] Optionally, it further includes:
[0098] Obtain the obstacle avoidance evaluation score of the vehicle user after ending a driving trip;
[0099] Upload the video data and the corresponding second localization box information during the driving trip with an obstacle avoidance evaluation score higher than the preset score to the cloud as an updated training data set.
[0100] Optionally, it further includes:
[0101] In the case where multiple groups of second localization box information obtained from multiple frames of visible light images indicate that the user is moving, perform face recognition in the second localization box information to obtain the attention of pedestrians to the vehicle;
[0102] Execute the speed reduction strategy when the attention level is greater than the preset value; otherwise, execute the stop-and-wait strategy.
[0103] An embodiment of the present invention provides an electronic device 30, as Figure 6 shown. The electronic device includes at least one processor 301, at least one memory 302 connected to the processor, and a bus 303. Among them, the processor 301 and the memory 302 communicate with each other through the bus 303. The processor 301 is configured to call program instructions in the memory to execute the above-mentioned road pedestrian detection method based on deep reinforcement learning.
[0104] The intelligent electronic device herein may be a PC, a PAD, a mobile phone, etc.
[0105] The present application also provides a computer program product, which is suitable for executing a program initialized with the steps of the above-mentioned road pedestrian detection method based on deep reinforcement learning when executed on a process management electronic device:
[0106] The present application is described with reference to the flowcharts and / or block diagrams of methods, electronic devices (systems), and computer program products according to embodiments of the present application. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processors of general-purpose computers, special-purpose computers, embedded processors, or other programmable process management electronic devices to generate a machine, so that the instructions executed by the processors of the computer or other programmable process management electronic devices generate means for implementing the specified functions in one Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0107] In a typical configuration, an electronic device includes one or more processors (CPUs), a memory, and a bus. The electronic device may also include an input / output interface, a network interface, etc.
[0108] The memory may include non-permanent memory in a computer-readable medium, in the form of random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash memory (flash RAM). The memory includes at least one storage chip. The memory is an example of a computer-readable medium.
[0109] A computer-readable medium includes permanent and non-permanent, removable and non-removable media that can implement information storage by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer-readable storage media for a computer include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory, or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD), or other optical storage, magnetic cassette tapes, magnetic tape magnetic disk storage, or other magnetic storage electronic devices, or any other non-transmission medium that can be used to store information that can be accessed by a computing electronic device. As defined herein, a computer-readable medium does not include transitory computer-readable media, such as modulated data signals and carrier waves.
[0110] It should also be noted that the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, commodity or electronic device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, commodity or electronic device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, commodity or electronic device comprising the element.
[0111] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable computer-readable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0112] The above are only the embodiments of the present application and are not used to limit the present application. For those skilled in the art, various changes and modifications can be made to the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application shall be included within the scope of the claims of the present application.
Claims
1. A road pedestrian detection method based on deep reinforcement learning, characterized in that, Including: Based on the collected visible light images around the vehicle, initially locate pedestrians through deep learning to obtain the first localization box information including pedestrians; According to the position coordinates of pedestrians in the first localization box, use a deep reinforcement learning agent to finely locate pedestrians to obtain the second localization box information including pedestrians, where the reward function in the deep reinforcement learning agent is EIoU; Transmit the second localization box information to the backend of the autonomous vehicle through human-computer interaction so that the autonomous vehicle can perform obstacle avoidance operations.
2. The method according to claim 1, wherein The initially locating pedestrians through deep learning based on the collected visible light images around the vehicle to obtain the first localization box information including pedestrians includes: Based on the collected visible light images around the vehicle, initially locate pedestrians through the YOLOv5 model to obtain the first localization box information including pedestrians.
3. The method according to claim 1, characterized in that, The neural network model adopted by the deep reinforcement learning agent is a DQN model.
4. The method according to claim 1, characterized in that, The human-computer interaction application program is developed based on the flutter cross-platform.
5. The method according to claim 1, characterized in that The action set of the deep reinforcement learning agent includes actions such as moving the box along the horizontal and vertical axes, changing the scale, modifying the aspect ratio, and terminating the search.
6. The method according to claim 1, characterized in that, Also including: Obtain the obstacle avoidance evaluation score of the vehicle user after ending a driving trip; Upload the video data and the corresponding second localization box information during the driving trips with the obstacle avoidance evaluation score higher than the preset score to the cloud as an updated training data set.
7. The method according to claim 1, characterized in that, Also including: In the case where multiple sets of second localization box information obtained from multiple frames of visible light images indicate that the user is moving, perform face recognition in the second localization box information to obtain the attention of pedestrians to the vehicle; Execute the speed concession strategy when the attention is greater than the preset value, otherwise execute the stop-and-wait strategy.
8. A road pedestrian detection device based on deep reinforcement learning, characterized in that, Including: An acquisition unit for initially locating pedestrians through deep learning based on the collected visible light images around the vehicle to obtain the first localization box information including pedestrians; A first recognition unit for finely locating pedestrians according to the position coordinates of pedestrians in the first localization box using a deep reinforcement learning agent to obtain the second localization box information including pedestrians, where the reward function in the deep reinforcement learning agent is EIoU; A second recognition unit for transmitting the second localization box information to the backend of the autonomous vehicle through human-computer interaction so that the autonomous vehicle can perform obstacle avoidance operations.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored program, where when the program is executed by a processor, it implements the method for road pedestrian detection based on deep reinforcement learning as described in any one of claims 1 to 7.
10. An electronic device, characterized in that, The electronic device includes at least one processor and at least one memory connected to the processor; wherein, the processor is used to call the program instructions in the memory to execute the method for road pedestrian detection based on deep reinforcement learning as described in any one of claims 1 to 7.