Behavior recognition method based on vehicle-mounted camera visual angle and deep learning attitude estimation

Through the on-board camera and deep learning attitude estimation method, combined with the bicycle speed information, pedestrian behavior is accurately identified, and the problem of insufficient pedestrian behavior recognition accuracy and real-time performance in the existing technology is solved, efficient pedestrian behavior recognition and consideration of bicycle influencing factors is achieved, and the perception ability of the autonomous driving system is improved.

CN120340122APending Publication Date: 2025-07-18YANSHAN UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510380580.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-28
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The existing pedestrian behavior recognition methods rely on image recognition technology, ignore the advantages of posture estimation and bicycle factors, and lack the accuracy of dynamic behavior recognition, resulting in low accuracy and real-time performance of behavior recognition in complex scenarios.

Method used

Image frames are collected through the on-board camera, combined with OBD sensor to obtain the speed of the bicycle, pose estimation is performed using the OpenPose model, pedestrian pose key points are extracted and bounding boxes are generated, tracking algorithms are used to assign IDs to pedestrians, input behavior recognition models for behavior classification, and use graph convolution networks and attention mechanisms to enhance model expression capabilities.

Benefits of technology

It significantly improves the accuracy and real-time nature of pedestrian behavior recognition, reduces the consumption of computing resources, and can accurately capture the dynamic changes of pedestrians in complex scenarios, considers the influencing factors of bicycles on pedestrian behavior, and ensures the system's timely response in a dynamic environment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120340122A_ABST
    Figure CN120340122A_ABST
Patent Text Reader

Abstract

The invention provides a behavior recognition method based on a visual angle of a vehicle-mounted camera and deep learning attitude estimation, and the method comprises the steps: collecting an image frame of a front road through a vehicle-mounted foresight camera, and obtaining the speed of a vehicle through an OBD sensor; performing attitude estimation on the pedestrian in the image frame through an OpenPose model to extract pedestrian attitude key point coordinate information and generate a pedestrian bounding box; providing a pedestrian ID for each pedestrian by adopting a pedestrian bounding box generated by combining a tracking algorithm with attitude estimation; and inputting the pedestrian posture key point coordinate information, the self-vehicle speed and the pedestrian ID into a trained behavior recognition model for behavior classification, and outputting a pedestrian behavior recognition result. According to the method, the accuracy and robustness of pedestrian behavior recognition can be improved from multiple dimensions, the method is particularly suitable for application scenes where dynamic behaviors of pedestrians need to be analyzed, the behaviors of the pedestrians can be quickly and accurately recognized, and the intelligent level of vehicles and the safety of the pedestrians are enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of intelligent driving technology, and particularly to a behavior recognition method based on in-vehicle camera perspective and deep learning pose estimation. Background Art

[0002] With the rapid development of autonomous driving technology, the autonomous driving system needs to have efficient and accurate perception capabilities to cope with complex traffic environments.

[0003] During the process of autonomous driving, pedestrians are one of the traffic participants, and the behavior recognition of pedestrians is the key to ensuring driving safety. However, most of the existing pedestrian behavior recognition methods rely on image recognition technology, ignoring the advantages of pose estimation and the influence of the ego-vehicle factors on the behavior of surrounding pedestrians. In addition, the accuracy of dynamic behavior recognition based on single-frame image detection is insufficient, and traditional CNNs are difficult to model the spatio-temporal dependence relationship between human key points, resulting in low accuracy and real-time performance of behavior recognition in some complex scenarios. Summary of the Invention

[0004] A behavior recognition method based on in-vehicle camera perspective and deep learning pose estimation provided by this application aims to improve the recognition ability of the vehicle perception system for pedestrian behavior in complex traffic environments. This method can obtain image data through an in-vehicle camera, and combine object detection, pose estimation, and behavior classification technologies to achieve accurate recognition of pedestrian behavior, thereby providing more accurate decision-making support for the autonomous driving system.

[0005] In a first aspect, this application provides a behavior recognition method based on in-vehicle camera perspective and deep learning pose estimation, including the following steps: collecting image frames of the front road through an in-vehicle front-view camera, and obtaining the ego-vehicle speed through an OBD sensor; performing pose estimation on pedestrians in the image frames through the OpenPose model to extract the coordinate information of pedestrian pose key points and generate pedestrian bounding boxes; using a tracking algorithm to provide a pedestrian ID for each pedestrian in combination with the pedestrian bounding boxes generated by pose estimation; inputting the pedestrian pose key point coordinate information, the ego-vehicle speed, and the pedestrian ID into a trained behavior recognition model for behavior classification, and outputting the behavior recognition result of the pedestrian.

[0006] Optionally, the pedestrian pose key point coordinate information includes the coordinate information of 18 key points obtained from the head, shoulders, elbows, hips, and knees, and the key points are used to represent the pose characteristics of pedestrians.

[0007] Optionally, the tracking algorithm is the DeepSORT algorithm.

[0008] Optionally, the behavior recognition model is trained through the following steps: Data preprocessing: Preprocess the coordinate information of pedestrian pose key points, and extract joint positions, motion speeds, and bone features; Graph convolutional network construction: Apply a graph convolutional network in skeleton-based action recognition, represent the graph convolutional operation on each frame, and use the adjacency matrix to represent the relationship between the graph convolutional operation as the input feature map and the output feature map; Normalization and convolutional operation: Normalize each input stream through the BatchNorm layer, and then perform convolutional operations through the initial block; Graph convolutional network block processing: Each input stream is processed through multiple graph convolutional network blocks, and each graph convolutional network block includes spatial graph convolution and temporal convolution; Attention mechanism application: After each graph convolutional network block, apply an attention mechanism based on the vehicle speed of the host vehicle to enhance the expression ability of the behavior recognition model; Output aggregation: The outputs of the input branches are merged and aggregated through a global average pooling layer to compress multi-dimensional features into a one-dimensional feature vector; Classification and optimization: Use a fully connected layer to map the one-dimensional feature vector to the output of the number of action categories. Finally, the fully connected layer outputs a logits vector, calculates the classification probability through the Softmax function, and uses the cross-entropy loss as the objective function for optimizing the behavior recognition model.

[0009] Optionally, the input of the joint position is formed by connecting a set of relative positions obtained by normalization, the input of the motion speed is obtained by splicing the fast motion speed and the slow motion speed, and the input of the bone feature includes the bone length and the bone angle.

[0010] Optionally, in the graph convolutional network, the graph convolutional operation on each frame t is expressed as:

[0011]

[0012] where v ti represents the i-th joint of the t-th frame, f in (·) and f out (·) are the input and output features of the corresponding joint respectively, N(v ti ) is the neighbor set of v ti , the normalization term Z ti is used to balance the contributions of different neighbors, w(·) is a weight function, and weights are assigned through the label function l ti (·);

[0013] Using the adjacency matrix, the graph convolutional operation on each frame t is expressed as:

[0014]

[0015] where D is the graph distance, the adjacent nodes are 1, f in and f outrespectively represent the input feature map and the output feature map, ⊙ represents the element-wise product, and A d represents the d-th order adjacency matrix, marking the joint pairs with a graph distance of d, and Λ d is used to normalize A d .

[0016] Optionally, the application of the attention mechanism specifically includes: first, normalizing the scalar vehicle speed to the same range as the input features, and projecting it to the same dimensions as the frame-level pooled features and the joint-level pooled features through two independent fully connected layers; performing average pooling on the output features of the graph convolutional network block at the frame level and the joint level respectively, and then splicing the vehicle speed features into the pooling results; multiplying the attention scores of the frame and the joint through the outer product of channels to obtain the final attention score.

[0017] In a second aspect, the present application provides a behavior recognition system based on the perspective of an in-vehicle camera and deep learning pose estimation, including: an image acquisition module for acquiring image frames of the road ahead through an in-vehicle front-view camera; a vehicle speed acquisition module for acquiring the vehicle speed of the host vehicle through an OBD sensor; a pose estimation module for estimating the pose of pedestrians in the image frames through an OpenPose model to extract the coordinate information of pedestrian pose key points and generate pedestrian bounding boxes; a tracking module for providing a pedestrian ID for each pedestrian by using a tracking algorithm in combination with the pedestrian bounding boxes generated by pose estimation; a behavior classification module for inputting the pedestrian pose key point coordinate information, the vehicle speed of the host vehicle, and the pedestrian ID into a trained behavior recognition model for behavior classification and outputting the behavior recognition result of the pedestrian.

[0018] In a third aspect, the present application provides an electronic device, including a processor and a memory storing a program. Among them, the program includes instructions that, when executed by the processor, cause the processor to execute the steps performed by the method in the first aspect as described above.

[0019] In a fourth aspect, the present application provides a computer storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the method in the first aspect as described above.

[0020] Compared with the prior art, the beneficial effects of the solution of the present application include:

[0021] By combining pose estimation and object detection, the present invention can more accurately capture the spatial positions of each joint point of a pedestrian, thereby effectively extracting behavioral features, significantly improving the classification accuracy of the pedestrian behavior recognition model, avoiding complex calculations and analyses of full-body or panoramic images, and thus significantly reducing the consumption of computing resources. Compared with traditional image-processing-based behavior recognition methods, the present invention can more precisely capture the dynamic changes of a pedestrian's body, taking into account the influence factors of the host vehicle on the behavior of surrounding pedestrians, and improving the recognition ability of pedestrian behavior in complex scenarios. In an in-vehicle perception system or other application scenarios that require real-time response, the present invention can complete the recognition of pedestrian behavior in a short time, ensuring the timely response of the system in a dynamic environment and having high real-time performance. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following-described drawings are only some embodiments of the present application, and those of ordinary skill in the art can obtain other drawings based on these drawings without creative efforts.

[0023] Figure 1 It is a flowchart of the steps of a behavior recognition method based on the perspective of an in-vehicle camera and deep learning pose estimation according to the present invention;

[0024] Figure 2 For Figure 1 The overall flowchart of the corresponding behavior recognition method based on the perspective of an in-vehicle camera and deep learning pose estimation;

[0025] Figure 3 It is a schematic diagram of the key points of human pose estimation;

[0026] Figure 4 It is a schematic diagram of the network architecture of the behavior recognition model;

[0027] Figure 5 It is a schematic diagram of the GCN network structure;

[0028] Figure 6 It is a schematic diagram of the TC layer structure;

[0029] Figure 7 It is a schematic diagram of the visualization effect of pedestrian behavior recognition. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0030] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are some, but not all, of the embodiments of this application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in this application without creative efforts shall fall within the scope of protection of this application.

[0031] It should be noted that the brief description of the terms in this application is only for facilitating the understanding of the subsequent described embodiments, rather than intending to limit the embodiments of this application. Unless otherwise specified, these terms should be understood in their ordinary and general meanings.

[0032] In this application, terms such as "first", "second", "third", etc. in the description, claims, and the above-mentioned drawings are used to distinguish similar or like objects or entities, and do not necessarily mean to limit a specific order or sequence, unless otherwise noted. It should be understood that such terms can be interchanged under appropriate circumstances.

[0033] The terms "comprising" and "having" and any variations thereof are intended to cover but not exclude inclusion. For example, a product or device comprising a series of components does not necessarily have to be limited to all the components clearly listed, but may include other components not clearly listed or inherent to these products or devices.

[0034] Pedestrian behavior recognition is of great significance in intelligent driving, traffic safety, and the construction of smart cities. In an autonomous driving system, accurately recognizing pedestrian behavior can significantly improve the vehicle's perception ability of complex road scenarios and reduce traffic accidents caused by improper human-vehicle interaction. In traffic management, pedestrian behavior recognition can provide a decision-making basis for signal control systems and optimize road traffic efficiency. In addition, in the field of public safety, real-time monitoring and analysis of pedestrian behavior can also effectively warn of abnormal activities and ensure the safety of urban residents. Therefore, the research and application of pedestrian behavior recognition technology not only concern technological development but also have a profound impact on social safety and convenience. It is very necessary to design a reliable, accurate, and adaptable pedestrian behavior recognition system.

[0035] See Figure 1 、 Figure 2 , this application provides a behavior recognition method based on the perspective of an in-vehicle camera and deep learning pose estimation, including the following steps:

[0036] Step S101: Collect image frames of the road ahead through an in-vehicle front-view camera, and obtain the vehicle speed of the host vehicle through an OBD sensor;

[0037] Step S102: Perform pose estimation on the pedestrians in the image frame through the OpenPose model to extract the coordinate information of the pedestrian pose key points and generate a pedestrian bounding box;

[0038] Step S103: Use a tracking algorithm combined with the pedestrian bounding box generated by the pose estimation to provide a pedestrian ID for each pedestrian;

[0039] Step S104: Input the pedestrian pose key point coordinate information, the self-vehicle speed, and the pedestrian ID into the trained behavior recognition model for behavior classification, and output the behavior recognition result of the pedestrian.

[0040] Exemplarily, obtain the image frame from the in-vehicle front-view camera. First, collect the image frames of the road ahead in real time through the in-vehicle front-view camera. The installation position of the in-vehicle front-view camera is at the front of the vehicle, which can obtain the road and traffic environment information within a certain range in front of the vehicle and ensure that the clarity and smoothness meet the requirements of real-time analysis.

[0041] Obtain the self-vehicle speed: directly obtain the self-vehicle speed through the OBD (On-Board Diagnostic) sensor.

[0042] Apply the OpenPose model to each frame of the image for pose estimation. OpenPose is a deep learning-based pose estimation technology that can extract the key point coordinates of the human body from the image and analyze the human pose. Through the OpenPose model, the coordinate information of 18 key points such as the head, shoulders, elbows, hips, and knees of each pedestrian can be obtained. The specific key point positions are shown in Figure 3 。

[0043] Use a part of the information (pedestrian bounding box, ID, image) in the PIE dataset to help train the tracking model DeepSORT to obtain the trained model one. Based on the pedestrian bounding box information provided by the pose estimation, use the trained model one to continuously track the same pedestrians in consecutive image frames, and create a unique pedestrian ID for each newly appeared pedestrian. By associating the pedestrian bounding boxes across frames, stable tracking of the pedestrian movement path is achieved, and the ID information corresponding to each pedestrian is dynamically updated.

[0044] Make a human pose dataset in the traffic scenario to train the behavior recognition model. The PIE dataset provides the image information of pedestrians in the traffic scenario, as well as the labeled position, action information, self-vehicle speed information, etc. Extract the key points of the pedestrians in the PIE dataset and correspond them to the action types labeled in the dataset. Input the pedestrian pose features extracted by OpenPose, the pedestrian ID, and the vehicle speed into the trained behavior recognition model for behavior classification, and output the behavior recognition result of the pedestrian. The behavior recognition model processes the data through a multi-layer graph neural network and outputs the behavior category, that is, the behavior recognition result: standing, walking.

[0045] Optionally, the tracking algorithm is the DeepSORT algorithm.

[0046] Optionally, the pedestrian pose key point coordinate information includes the coordinate information of 18 key points obtained from the head, shoulders, elbows, hips, and knees, and the key points are used to represent the pose characteristics of the pedestrian.

[0047] Optionally, the behavior recognition model is trained through the following steps: Data preprocessing: Preprocess the pedestrian pose key point coordinate information to extract joint positions, movement speeds, and bone features; Graph convolutional network construction: Apply a graph convolutional network in skeleton-based action recognition to represent the graph convolutional operation on each frame, and use the adjacency matrix to represent the relationship between the graph convolutional operation as the input feature map and the output feature map; Normalization and convolutional operation: Normalize each input stream through the BatchNorm layer, and then perform convolutional operations through the initial block; Graph convolutional network block processing: Each input stream is processed through multiple graph convolutional network blocks, and each graph convolutional network block includes spatial graph convolution and temporal convolution; Attention mechanism application: After each graph convolutional network block, apply an attention mechanism based on the vehicle speed of the host vehicle to enhance the expression ability of the behavior recognition model; Output aggregation: The outputs of the input branches are merged and aggregated through a global average pooling layer to compress the multi-dimensional features into a one-dimensional feature vector; Classification and optimization: Use a fully connected layer to map the one-dimensional feature vector to the output of the number of action categories. Finally, the fully connected layer outputs a logits vector, calculates the classification probability through the Softmax function, and uses the cross-entropy loss as the objective function for optimizing the behavior recognition model.

[0048] Optionally, the input of the joint position is formed by the concatenation of the relative position set obtained by normalization and the original 3D coordinate set, the input of the movement speed is obtained by splicing the fast movement speed and the slow movement speed, and the input of the bone feature includes the bone length and the bone angle.

[0049] Optionally, in the graph convolutional network, the graph convolutional operation on each frame t is expressed as:

[0050]

[0051] where, v ti represents the i-th joint of the t-th frame, f in (·) and f out (·) are the input and output features of the corresponding joints respectively, N(v ti ) is the neighbor set of v ti , the normalization term Z ti is used to balance the contributions of different neighbors, w(·) is a weight function, and through the label function lti (·) Assign weights;

[0052] Using the adjacency matrix, represent the graph convolution operation on each frame t as:

[0053]

[0054] where D is the graph distance, adjacent nodes are 1, f in and f out represent the input feature map and the output feature map respectively, ⊙ represents element-wise multiplication, A d represents the d-th order adjacency matrix, marking the joint pairs with graph distance d, Λ d is used to normalize A d .

[0055] Optionally, the application of the attention mechanism specifically includes: First, normalize the scalar vehicle speed to the same range as the input features, and project it to the same dimensions as the frame-level pooling features and joint-level pooling features through two independent fully connected layers; perform average pooling on the output features of the graph convolution network block at the frame level and joint level respectively, and then splice the vehicle speed features into the pooling results; multiply the attention scores of the frame and the joint through the outer product of channels to obtain the final attention score.

[0056] The method of the present application will be described in detail below in conjunction with embodiments.

[0057] Embodiment 1:

[0058] Step 1, use the on-vehicle front-view camera to collect the image frames of the road ahead, that is, the original video frame image I; directly obtain the vehicle speed of the vehicle through the OBD (On-Board Diagnostic system) sensor.

[0059] Step 2, transfer the collected image frames to the OpenPose model for pose estimation. OpenPose extracts the coordinates of each key point of the pedestrian body, that is, the pedestrian pose key point coordinate information, according to the image area within the detection box, that is, the pedestrian bounding box, and generates a bounding box for each pedestrian according to the body key points. The pose of each pedestrian is represented by 18 key points, namely Nose, Neck, RShoulder, RElbow, RWrist, LShoulder, LElbow, LWrist, RHip, RKnee, Rankle, LHip, LKnee, LAnkle, Reye, LEye, Rear, Lear. As Figure 3 shown.

[0060] Step 3, tracking algorithm: After the pose estimation in step 2 in consecutive video frames, input the trained DeepSORT model to identify and track the pedestrian object, and assign a unique identifier (ID) to each detected pedestrian.

[0061] Step 4, refer to Figure 4 , preprocessing of pedestrian pose key point coordinate information. Data preprocessing is very important for skeleton-based action recognition. The input features after various preprocessings are mainly divided into three categories: joint positions, motion speeds, and bone features.

[0062] (1) Joint positions: The original 3D coordinate set of an action sequence is: where C in , T in , V in represent the input coordinates, number of frames, and number of joints respectively. Then, the relative position set is obtained through normalization, that is, where r {i} = x[:, :, i] - x[:, :, c], where c represents the index of the central spine joint. Next, the input of joint positions is formed by the concatenation of and .

[0063] (2) Joint point motion speed, i.e., motion speed: Two sets of speed data are obtained by processing the time series. Define for fast motion, for slow motion, where f t = x[:, t + 2, :] - x[:, t, :], s t = x[:, t + 1, :] - x[:, t, :]. From time point t to time point t + 2 (i.e., spanning two time steps), the change in all joint positions is regarded as the fast motion speed within this time period; from time point t to time point t + 1 (i.e., adjacent time steps), the change in all joint positions is regarded as the slow motion speed within this time period. The input of motion speed is obtained by concatenating and together, forming a feature vector for each joint at each moment.

[0064] (3) Bone features: The input of bone features includes bone length and bone angle . The length and angle of each bone are calculated by the following formula: l i = x[:, :, i] - x[:, :, i adj , where i adj represents the adjacent joint of the i-th joint, and w ∈ {x, y} represents 2D coordinates.

[0065] Step 5, refer to Figure 4 , applying a graph convolutional network (GCN) in skeleton-based action recognition. The graph convolutional operation at each frame t can be expressed as:

[0066]

[0067] where v ti represents the i-th joint of the t-th frame, and f in (·) and f out (·) are the input and output features of the corresponding joint respectively. N(v ti ) is the neighbor set of v ti . The normalization term Z ti is used to balance the contributions of different neighbors. w(·) is a weight function, and weights are assigned through the label function l ti (·). The label function is used to construct multiple neighbor sets N(v ti ), assign different labels to each graph node, and select a distance-based partitioning method: l ti (v tj ) = d(v ti , v tj ). Where d(v ti , v tj ) represents the graph distance between v ti and v tj . Joints with the same distance will form a subset and share a learnable weight function w(·).

[0068] Using the adjacency matrix A, the graph convolution operation on each frame t is expressed as:

[0069]

[0070] where D is the graph distance, adjacent nodes are 1, and f in and f out represent the input and output feature maps respectively. ⊙ represents element-wise multiplication, A d represents the d-th order adjacency matrix, marking pairs of joints with graph distance d. Λ d is used to normalize A d . W d and M d are both learnable matrices.

[0071] Step 6, after data preprocessing, each input stream is normalized through the BatchNorm layer. The BatchNorm layer can make network training more stable and fast by normalizing the data distribution on the input of each layer. Next is the Initial Block, which contains a convolution operation with an output channel number of 64.

[0072] Step 7, each input stream is processed through multiple GCN blocks, i.e., multiple graph convolutional network blocks. Each GCN block contains spatial graph convolution (SGC) and temporal convolution (TC), as shown in Figure 5 . The SGC layer is implemented through the graph convolution mentioned in Step 5, and the structure of the TC layer is as shown in Figure 6 . By combining SGC and TC, the model can consider the information in both spatial and temporal dimensions when processing data.

[0073] Step 8, after each graph convolutional network block, i.e., each GCN block, an attention mechanism (Attention) considering the ego-vehicle speed is applied to enhance the model's expressive ability. The spatio-temporal joint attention module first calculates the attention weights of each joint over all frames, and then calculates the attention weights of each frame over all joints, which can capture the spatial and temporal relationships in the skeleton data more comprehensively. At the same time, the speed information is embedded into the pooled features at the frame level (temporal dimension) and joint level (spatial dimension) respectively. By dynamically adjusting the attention weights, the speed information can specifically affect the modeling of spatio-temporal features, thereby improving the model's attention to key features.

[0074] First, the scalar vehicle speed is normalized to the same range as the input features. The vehicle speed passes through two independent fully connected layers and is projected into the same dimension as the frame-level pooled features and joint-level pooled features respectively. The vehicle speed projection formula:

[0075] v t = FC t (v speed ), v v = FC v (v speed )

[0076] where is the projected vehicle speed feature vector.

[0077] Average pooling is performed on the output features of the GCN Block, i.e., the graph convolutional network block, at the frame level and joint level respectively, and then the vehicle speed feature is concatenated into the pooling result:

[0078] Frame-level pooling + vehicle speed:

[0079] pool t (f in ) = Concat(pool t (f in ), v t )

[0080] Joint-level pooling + vehicle speed:

[0081] pool v (f in ) = Concat(poolv (f in ),v v )

[0082] Internal feature f inner Calculation:

[0083]

[0084] where pool t (f in ) represents frame-level average pooling, and pool v (f in ) represents joint-level average pooling, θ is the activation function, and w is the weight matrix.

[0085] Multiply the attention scores of frames and joints through outer product of channels to obtain the final attention score:

[0086]

[0087] where f out is the output feature, σ is the activation function, W t is the weight matrix in the frame dimension, and W v is the weight matrix in the joint dimension.

[0088] Step 9, the outputs of the input branches are merged and aggregated through a global average pooling layer (Global Average Pool, GAP). The multi-dimensional features are compressed into a one-dimensional vector through the GAP layer, making the subsequent fully connected layer processing simpler and more efficient.

[0089] Step 10, use the fully connected layer to map the feature vector to the output of the number Q of action categories.

[0090] Step 11, the final fully connected (FC) layer outputs a logits vector where Q is the number of action categories. For the input sample, the classification probability can be calculated through the Softmax function, that is:

[0091]

[0092] where z i represents the i-th element of z.

[0093] The cross-entropy loss is calculated as the objective function for optimizing the behavior recognition model as follows:

[0094]

[0095] where is the one-hot vector indicating the true label.

[0096] Step 12: Use the PIE dataset extracted through key points to train the constructed behavior recognition model. Input the pedestrian pose key point coordinate information, the ego-vehicle speed, and the pedestrian ID obtained in Step 4 into the trained behavior recognition model to obtain the pedestrian behavior recognition result and visualize it on the image, as Figure 7 shown.

[0097] In summary, after the pooling layer of the spatio-temporal joint attention module in the method of the present invention, the vehicle speed information is respectively embedded into the pooling features at the frame level (temporal dimension) and the joint level (spatial dimension). By dynamically adjusting the attention weights, the vehicle speed information specifically affects the modeling of spatio-temporal features, thereby considering the impact of the ego-vehicle speed on pedestrian behavior. It has the following beneficial effects:

[0098] By combining pose estimation and object detection, it can more accurately capture the spatial positions of each joint point of the pedestrian, thereby effectively extracting behavior features, significantly improving the classification accuracy of the pedestrian behavior recognition model, avoiding complex calculations and analyses of full-body or panoramic images, and thus significantly reducing the consumption of computing resources. Compared with traditional behavior recognition methods based on image processing, the present invention can more precisely capture the dynamic changes of the pedestrian's body, consider the influencing factors of the ego-vehicle on the behavior of surrounding pedestrians, and improve the recognition ability of pedestrian behavior in complex scenarios. In an in-vehicle perception system or other application scenarios that require real-time response, the present invention can complete the recognition of pedestrian behavior in a short time, ensuring the timely response of the system in a dynamic environment and having high real-time performance.

[0099] The present application also provides a behavior recognition system based on the perspective of an in-vehicle camera and deep learning pose estimation, including:

[0100] An image acquisition module for acquiring image frames of the road ahead through an in-vehicle front-view camera;

[0101] A vehicle speed acquisition module for obtaining the ego-vehicle speed through an OBD sensor;

[0102] A pose estimation module for performing pose estimation on pedestrians in the image frame through an OpenPose model to extract pedestrian pose key point coordinate information and generate pedestrian bounding boxes;

[0103] A tracking module for providing a pedestrian ID for each pedestrian by using a tracking algorithm in combination with the pedestrian bounding boxes generated by pose estimation;

[0104] A behavior classification module for inputting the pedestrian pose key point coordinate information, the ego-vehicle speed, and the pedestrian ID into a trained behavior recognition model for behavior classification and outputting the pedestrian behavior recognition result.

[0105] It should be understood that the behavior recognition system based on the vehicle-mounted camera perspective and deep learning pose estimation in the embodiments of the present invention is used to implement the corresponding methods in the foregoing multiple method embodiments and has the beneficial effects of the corresponding method embodiments.

[0106] As another example, the embodiments of the present invention also provide an electronic device, which is an example of a hardware device that can be applied to various aspects of the present invention. The electronic device is intended to represent various forms of digital electronic computer devices, such as, laptop computers, desktop computers, workbenches, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, personal digital processors, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present invention described and / or claimed herein.

[0107] The electronic device may include: a processor, a communications interface, a memory, and a communication bus.

[0108] The processor, the communications interface, and the memory communicate with each other through the communication bus. The communications interface is used to communicate with other electronic devices or servers.

[0109] The processor is used to execute a program, specifically, it can execute the relevant steps in the foregoing method embodiments.

[0110] Specifically, the program may include program code, and the program code includes computer operation instructions.

[0111] The processor may be a processor CPU, or a specific integrated circuit ASIC (Application Specific Integrated Circuit), or one or more integrated circuits configured to implement the embodiments of the present invention. One or more processors included in the intelligent device may be of the same type of processor, such as one or more CPUs; or they may be of different types of processors, such as one or more CPUs and one or more ASICs.

[0112] The memory is used to store the program. The memory may include high-speed RAM memory, and may also include non-volatile memory, such as at least one disk memory.

[0113] When the program is executed by the processor, it is used to cause the electronic device to execute the behavior recognition method based on the vehicle-mounted camera perspective and deep learning pose estimation of the present invention.

[0114] In addition, for the specific implementation of each step in the program, reference may be made to the corresponding steps and the corresponding descriptions in the units in the foregoing method embodiments, which will not be elaborated herein. Those skilled in the art can clearly understand that, for the sake of convenience and brevity of description, for the specific working processes of the devices and modules described above, reference may be made to the corresponding process descriptions in the foregoing method embodiments, which will not be elaborated herein.

[0115] An exemplary embodiment of the present invention further provides a computer storage medium storing a computer program, wherein when the computer program is executed by a processor, the methods of the embodiments of the present invention are implemented. Reference may be made to the corresponding process descriptions in the foregoing method embodiments, which will not be elaborated herein.

[0116] The methods according to the embodiments of the present invention described above can be implemented in hardware, firmware, or be implemented as software or computer code that can be stored in a recording medium (such as a CD ROM, RAM, floppy disk, hard disk, or magneto-optical disk), or be implemented as computer code originally stored in a remote recording medium or a non-transitory machine-readable medium and to be stored in a local recording medium through network download, so that the methods described herein can be stored in such software processes on a recording medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware (such as an ASIC or FPGA). It can be understood that a computer, a processor, a microprocessor controller, or programmable hardware includes a storage component (such as a RAM, a ROM, a flash memory, etc.) that can store or receive software or computer code, and when the software or computer code is accessed and executed by the computer, the processor, or the hardware, the methods described herein are implemented. In addition, when a general-purpose computer accesses the code for implementing the methods shown herein, the execution of the code converts the general-purpose computer into a dedicated computer for executing the methods shown herein.

[0117] So far, specific embodiments of the present invention have been described. Other embodiments are within the scope of the appended claims. In some cases, the actions recited in the claims can be performed in a different order and still achieve the desired result. Additionally, the processes depicted in the figures do not necessarily require the particular order or sequential order shown to achieve the desired result. In certain embodiments, multitasking and parallel processing may be advantageous.

[0118] It should be understood that although this specification is described according to each embodiment, not every embodiment only contains an independent technical solution. This narrative manner of the specification is only for clarity. Those skilled in the art should regard the specification as a whole, and the technical solutions in each embodiment can also be appropriately combined to form other embodiments that can be understood by those skilled in the art.

[0119] Finally, it should be noted that the above embodiments are only used to illustrate the embodiments of the present invention, rather than to limit the embodiments of the present invention. Those of ordinary skill in the relevant technical field can also make various changes and modifications without departing from the spirit and scope of the embodiments of the present invention. Therefore, all equivalent technical solutions also belong to the scope of the embodiments of the present invention, and the patent protection scope of the embodiments of the present invention shall be defined by the claims.

Claims

1. A behavior recognition method based on the perspective of in-vehicle cameras and deep learning pose estimation, characterized in that, It includes the following steps: Collect image frames of the road ahead through an in-vehicle front-view camera, and obtain the vehicle's own speed through an OBD sensor; Perform pose estimation on pedestrians in the image frames through the OpenPose model to extract the coordinate information of pedestrian pose key points and generate pedestrian bounding boxes; Use a tracking algorithm combined with the pedestrian bounding boxes generated by pose estimation to provide a pedestrian ID for each pedestrian; Input the pedestrian pose key point coordinate information, the vehicle's own speed, and the pedestrian ID into a trained behavior recognition model for behavior classification, and output the behavior recognition result of the pedestrian.

2. The method according to claim 1, wherein The pedestrian pose key point coordinate information includes the coordinate information of 18 key points obtained from the head, shoulders, elbows, hips, and knees, and the key points are used to represent the pose characteristics of the pedestrian.

3. The method according to claim 1, characterized in that, The tracking algorithm is the DeepSORT algorithm.

4. The method according to claim 1, wherein The behavior recognition model is trained through the following steps: Data preprocessing: Preprocess the pedestrian pose key point coordinate information to extract joint positions, movement speeds, and skeletal features; Graph convolutional network construction: Apply a graph convolutional network in skeleton-based action recognition to represent the graph convolutional operations on each frame, and use an adjacency matrix to represent the relationship between the graph convolutional operations as the input feature map and the output feature map; Normalization and convolutional operations: Normalize each input stream through a BatchNorm layer, and then perform convolutional operations through an initial block; Graph convolutional network block processing: Each input stream is processed through multiple graph convolutional network blocks, and each graph convolutional network block includes spatial graph convolution and temporal convolution; Attention mechanism application: After each graph convolutional network block, apply an attention mechanism based on the vehicle's own speed to enhance the expression ability of the behavior recognition model; Output aggregation: The outputs of the input branches are merged and aggregated through a global average pooling layer to compress the multi-dimensional features into a one-dimensional feature vector; Classification and optimization: Use a fully connected layer to map the one-dimensional feature vector to the output of the number of action categories. Finally, the fully connected layer outputs a logits vector, calculates the classification probability through the Softmax function, and uses the cross-entropy loss as the objective function for optimizing the behavior recognition model.

5. The method according to claim 4, wherein The input of the joint position is formed by connecting a set of relative positions obtained by normalization, the input of the movement speed is obtained by splicing the fast movement speed and the slow movement speed, and the input of the skeletal feature includes the skeletal length and the skeletal angle.

6. The method according to claim 4, wherein In the graph convolutional network, the graph convolutional operation on each frame t is expressed as: Among them, vti represents the i-th joint of the t-th frame, f in (·) and fout (·) are the input and output features of the corresponding joints respectively, N(v ti ) is the neighbor set of v ti , and the normalization term Z ti is used to balance the contributions of different neighbors. w(·) is a weight function that assigns weights through the label function l ti (·); Using the adjacency matrix, the graph convolutional operation on each frame t is expressed as: Among them, D is the graph distance, the adjacent nodes are 1, f in and f out represent the input feature map and the output feature map respectively, ⊙ represents the element-wise product, A d represents the d-th order adjacency matrix, marking the joint pairs with graph distance d, Λ d is used to normalize A d .

7. The method according to claim 4, characterized in that The application of the attention mechanism specifically includes: First, normalize the scalar vehicle speed to the same range as the input features, and project it to the same dimension as the frame-level pooled features and the joint-level pooled features through two independent fully connected layers respectively; Perform average pooling on the output features of the graph convolutional network block at the frame level and the joint level respectively, and then splice the vehicle speed features into the pooling results; Multiply the attention scores of the frame and the joint through the outer product of channels to obtain the final attention score.

8. A behavior recognition system based on the perspective of in-vehicle cameras and deep learning pose estimation, characterized in that, It includes: An image acquisition module for collecting image frames of the road ahead through an in-vehicle front-view camera; A vehicle speed acquisition module, configured to acquire the vehicle speed of the host vehicle through an OBD sensor; A pose estimation module, configured to perform pose estimation on pedestrians in an image frame through an OpenPose model to extract pedestrian pose key point coordinate information and generate pedestrian bounding boxes; A tracking module, configured to provide a pedestrian ID for each pedestrian by using a tracking algorithm in combination with the pedestrian bounding boxes generated by pose estimation; A behavior classification module, configured to input the pedestrian pose key point coordinate information, the host vehicle speed, and the pedestrian ID into a trained behavior recognition model for behavior classification and output the pedestrian behavior recognition result.

9. An electronic device, characterized in that, Comprising: A processor; A memory storing a program; Wherein, the program includes instructions which, when executed by the processor, cause the processor to execute the steps performed by the method according to any one of claims 1-7.

10. A computer storage medium, characterized in that, A computer program is stored thereon, and when the program is executed by a processor, the method according to any one of claims 1-7 is implemented.

Citation Information

Cited By

  • Pedestrian behavior prediction method for autonomous vehicle

    CN121583001A