Road user behavior identification method and system based on vehicle-mounted visual angle visual information
By annotating the video data of the on-board viewing angle and building a behavior recognition model with a hybrid dual-path backbone network, neck network and multi-task head, the problem of scarce road users' behavior data in vehicle-side perspective is solved, and accurate identification of road users' behaviors in vehicle-side perspectives is achieved and behavior recognition capabilities are improved in complex traffic scenarios.
Patent Information
- Application Number
- CN202510099290.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-22
- Publication Date
- 2025-05-06
AI Technical Summary
In the prior art, road user behavior data is scarce on vehicle perspective, and there are few behavior recognition methods based on vehicle perspective, which affects the accuracy and real-time nature of behavior recognition in complex traffic scenarios.
By annotating the video data in the target dataset, label information for categories, behaviors and locations is obtained, and a road user behavior recognition model including a hybrid dual-path backbone network, a neck network and a multi-task head is constructed. The model uses slow paths and fast paths to capture static background information and dynamic action changes in the video data respectively, combines the multi-dimensional attention mechanism to integrate features, and finally uses multi-task heads to detect and classify, and obtains the category information, behavior information and location information of road users.
It realizes accurate identification of road users' behaviors at the vehicle perspective, improves behavior recognition capabilities in complex traffic scenarios, enhances stable understanding and judgment of dynamic changes in backgrounds, and makes up for the scarcity of vehicle perspective behavior data.
Smart Images

Figure CN119942501A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of autonomous driving environment perception, and specifically to a road user behavior recognition method and system based on vehicle-mounted perspective visual information. Background Art
[0002] With the development of science and technology and the continuous improvement of people's living standards, the number of civilian cars has increased rapidly. Autonomous driving is an effective solution to improve traffic efficiency and enhance driving safety. The key technologies involved in autonomous driving vehicles cover the fields of perception and cognition, decision planning and control execution. Among them, the recognition of road users' behavior is an important part of autonomous driving environmental perception and cognition, and it is also an important basis for autonomous driving decision control. The accuracy and duration of behavior recognition will directly affect the action accuracy and allowable time consumption of the lower-level decision-making and control tasks, and it occupies an important position in autonomous driving technology.
[0003] At present, the behavior recognition based on the vehicle perspective often only focuses on pedestrians, lacking data and related methods for multiple types of road users such as vehicles and bicycles. Moreover, most of the research on vehicle behavior recognition methods is based on historical trajectory information under the bird's-eye view (BEV). However, in the process of converting the target's perception data into the historical trajectory under the geodetic coordinate system, due to the influence of factors such as information loss and perspective limitation, it is easy to introduce errors that cannot be ignored, resulting in a decrease in data quality, thereby affecting the accuracy of related behavior recognition detection methods based on the BEV perspective.
[0004] Most of the current research based on vehicle-mounted perspective datasets focuses on areas such as intention recognition and trajectory prediction, and lacks research on the detection of road user behavior in traffic environments. However, in some complex traffic scenarios, behavioral changes caused by interactions between road users often have a profound impact on driving decisions. At this time, ignoring the recognition and detection of action behaviors will affect the accuracy of related algorithms. Summary of the invention
[0005] The present application provides a road user behavior recognition method based on vehicle-mounted perspective visual information to solve the problems in the prior art of scarce vehicle-mounted perspective road user behavior data and few vehicle-mounted perspective road user behavior recognition methods.
[0006] Correspondingly, the present application also provides a road user behavior recognition system based on vehicle-mounted perspective visual information, an electronic device, and a computer-readable storage medium to ensure the implementation and application of the above method.
[0007] In order to solve the above technical problems, the present application discloses a road user behavior recognition method based on vehicle-mounted visual information, the method comprising:
[0008] Label the video data in the target dataset to obtain the corresponding label information; the label information includes category, behavior and location;
[0009] Construct a road user behavior recognition model; the road user behavior recognition model includes a hybrid dual-path backbone network, a neck network based on a multi-dimensional attention mechanism, and a multi-task head; the hybrid dual-path backbone network includes a slow path and a fast path;
[0010] The video data and the corresponding label information are input into the road user behavior recognition model, and the slow path and the fast path are used to capture the static background information and the dynamic action changes in the video data respectively, so as to obtain the corresponding slow path spatiotemporal features and the fast path spatiotemporal features;
[0011] The neck network is used to fuse the slow path spatiotemporal features and the fast path spatiotemporal features to obtain fused features;
[0012] The fused features are detected and classified using a multi-task head to obtain the category information, behavior information and location information of road users; wherein the multi-task head realizes information matching through label information in the detection and classification tasks.
[0013] The present application also discloses a road user behavior recognition system based on vehicle-mounted visual information, the system comprising:
[0014] The data annotation module is used to annotate the video data in the target data set to obtain the corresponding label information; the label information includes category, behavior and location;
[0015] A model building module, used to build a road user behavior recognition model; the road user behavior recognition model includes a hybrid dual-path backbone network, a neck network based on a multi-dimensional attention mechanism, and a multi-task head; the hybrid dual-path backbone network includes a slow path and a fast path;
[0016] A feature extraction module is used to input the video data and the corresponding label information into the road user behavior recognition model, and use the slow path and the fast path to capture the static background information and dynamic action changes in the video data, respectively, to obtain the corresponding slow path spatiotemporal features and the fast path spatiotemporal features;
[0017] The feature extraction module is also used to fuse the slow path spatiotemporal features and the fast path spatiotemporal features using the neck network to obtain fused features;
[0018] The information recognition module is used to detect and classify the fused features using the multi-task head to obtain the category information, behavior information and location information of the road users; wherein the multi-task head realizes information matching through label information in the detection and classification tasks.
[0019] The present application also discloses an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, one or more methods described in the present application are implemented.
[0020] The present application also discloses a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, one or more methods described in the present application are implemented.
[0021] In this application, the video data in the target data set from the vehicle perspective are annotated with categories, behaviors and positions to obtain corresponding label information, which makes up for the scarcity of road user behavior data from the vehicle perspective and provides data support for the establishment of subsequent models. A road user behavior recognition model including a hybrid dual-path backbone network, a neck network and a multi-task head is constructed to realize the recognition of road user behavior from the vehicle perspective. Among them, the hybrid dual-path backbone network includes a slow path and a fast path, which can respectively capture dynamic action changes and static background information in the video data; then the neck network uses a multi-dimensional attention mechanism to fuse the slow path spatiotemporal features and the fast path spatiotemporal features extracted from the slow path and the fast path to obtain fused features, which can improve the model's recognition ability for complex actions, thereby enhancing the model's stable understanding and judgment of background dynamic changes. Finally, the multi-task head is used to detect and classify the fused features to obtain the category information, behavior information and location information of the road user, thereby realizing the recognition of road user behavior from the vehicle perspective.
[0022] Additional aspects and advantages of the present application will be given in the following description, which will become apparent from the following description, or will be understood through the practice of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] The above and / or additional aspects and advantages of the present application will become apparent and easily understood from the following description of the embodiments in conjunction with the accompanying drawings, in which:
[0024] Figure 1 A flow chart of a method for identifying road user behavior based on vehicle-mounted visual information provided in an embodiment of the present application;
[0025] Figure 2 A specific flow chart of the road user behavior identification method provided in an embodiment of the present application;
[0026] Figure 3 A schematic diagram of the structure of a hybrid dual-path backbone network provided in an embodiment of the present application;
[0027] Figure 4 A schematic diagram of the convolution operation of a three-dimensional convolution layer provided in an embodiment of the present application;
[0028] Figure 5 A schematic diagram of the structure of an expanded 3D convolutional neural network that integrates optical flow temporal information provided in an embodiment of the present application;
[0029] Figure 6 A schematic diagram of the structure of the channel space attention module of the neck network provided in an embodiment of the present application;
[0030] Figure 7 A schematic diagram of the structure of a channel time attention module of a neck network provided in an embodiment of the present application;
[0031] Figure 8 A schematic diagram of the structure of the spatial-temporal attention module of the neck network provided in an embodiment of the present application;
[0032] Fig. 9 A schematic diagram of the 3D-RetinaNet network structure for detection tasks provided in an embodiment of the present application;
[0033] Fig.10 A schematic diagram of behavior, position recognition loss and average loss curve during the training process of the road user behavior recognition model provided in an embodiment of the present application;
[0034] Fig.11 The target category detection curve diagram provided by the embodiment of the present application;
[0035] Fig.12 The action behavior detection curve diagram provided by the embodiment of the present application;
[0036] Fig.13 A position information detection curve diagram provided by an embodiment of the present application;
[0037] Fig.14 The mAP curve diagram of each detection item provided in the embodiment of the present application;
[0038] Fig.15 A schematic diagram of the structure of a road user behavior recognition system based on vehicle-mounted visual information provided in an embodiment of the present application;
[0039] Fig.16 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0040] The embodiments of the present application are described in detail below, and examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present application, and cannot be interpreted as limiting the present application.
[0041] It will be understood by those skilled in the art that, unless expressly stated, the singular forms "one", "said", and "the" used herein may also include plural forms. It should be further understood that the term "comprising" used in the specification of the present application refers to the presence of features, integers, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or combinations thereof. It should be understood that when we refer to an element as being "connected" or "coupled" to another element, it may be directly connected or coupled to the other element, or there may be an intermediate element. In addition, the "connection" or "coupling" used herein may include wireless connection or wireless coupling. The term "and / or" used herein includes all or any unit and all combinations of one or more associated listed items.
[0042] Those skilled in the art will appreciate that, unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as those generally understood by those skilled in the art in the field to which the present invention belongs. It should also be understood that terms such as those defined in general dictionaries should be understood to have meanings consistent with the meanings in the context of the prior art, and will not be interpreted with idealized or overly formal meanings unless specifically defined as here.
[0043] The solution provided in the embodiments of the present application can be executed by any electronic device, such as a terminal device or a server, wherein the server can be an independent physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server that provides cloud computing services. The terminal can be a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart speaker, a smart watch, etc., but is not limited to this. The terminal and the server can be directly or indirectly connected via wired or wireless communication, and the present application does not limit this. With regard to the technical problems existing in the prior art, the road user behavior recognition method and system based on vehicle-mounted perspective visual information provided in the present application is intended to solve at least one of the technical problems of the prior art.
[0044] The technical solution of the present application and how the technical solution of the present application solves the above-mentioned technical problems are described in detail below with specific embodiments. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of the present application will be described below in conjunction with the accompanying drawings.
[0045] The present application embodiment provides a possible implementation method, such as Figure 1As shown, a flowchart of a method for identifying road user behavior based on vehicle-mounted perspective visual information is provided. The solution can be executed by any electronic device, and optionally, can be executed on a server or terminal device.
[0046] like Figure 1 As shown in , the method may include the following steps:
[0047] Step 101, annotate the video data in the target data set to obtain corresponding label information; the label information includes category, behavior and position.
[0048] In the embodiment of the present application, the public Waymo data set can be selected as the target data set, and the target data set is annotated with the category, behavior, and location triplet label information to provide data support for the establishment of subsequent models, making up for the current scarcity of vehicle-mounted perspective data including multiple road users. Directly using the video data from the vehicle-mounted perspective for behavior recognition will save the conversion link of the target trajectory in the geodetic coordinate system, improve the data quality degradation during the coordinate conversion process and the real-time problem of behavior recognition.
[0049] Step 102, constructing a road user behavior recognition model; the road user behavior recognition model includes a hybrid dual-path backbone network, a neck network based on a multi-dimensional attention mechanism, and a multi-task head; the hybrid dual-path backbone network includes a slow path and a fast path.
[0050] Step 103, input the video data and the corresponding label information into the road user behavior recognition model, use the slow path and the fast path to capture the static background information and dynamic action changes in the video data respectively, and obtain the corresponding slow path spatiotemporal features and fast path spatiotemporal features.
[0051] The road user behavior recognition model adopts a hybrid dual-path backbone network including a slow path and a fast path, so that the road user behavior recognition model can simultaneously capture the dynamic action changes and static background information in the video data, improve the recognition ability of complex actions, and thus enhance the model's stable understanding and judgment ability of background dynamic changes.
[0052] Step 104: Use the neck network to fuse the slow path spatiotemporal features and the fast path spatiotemporal features to obtain fused features.
[0053] Among them, the neck network introduces a multi-dimensional attention mechanism to complete the feature fusion and splicing of the fast path and the slow path. The neck network is located after the hybrid dual-path backbone network. The application of a multi-dimensional attention mechanism in the neck network can dynamically weight the spatiotemporal features of the input, so that the entire behavior recognition model can pay more attention to the key information in the time, space and channel dimensions.
[0054] Step 105, using the multi-task head to detect and classify the fused features, and obtain the category information, behavior information and location information of the road user; wherein the multi-task head realizes information matching through label information in the detection and classification tasks.
[0055] In the embodiment of the present application, a unified behavior recognition method is studied by supplementing road user data, which is more conducive to subsequent decision-making planning and control of autonomous driving.
[0056] In the embodiment of the present application, the video data in the target data set from the vehicle perspective is annotated with categories, behaviors and positions to obtain corresponding label information, which makes up for the scarcity of road user behavior data from the vehicle perspective and provides data support for the establishment of subsequent models. A road user behavior recognition model including a hybrid dual-path backbone network, a neck network and a multi-task head is constructed to realize the recognition of road user behavior from the vehicle perspective. Among them, the hybrid dual-path backbone network includes a slow path and a fast path, which can respectively capture dynamic action changes and static background information in the video data; then the neck network uses a multi-dimensional attention mechanism to fuse the slow path spatiotemporal features and the fast path spatiotemporal features extracted from the slow path and the fast path to obtain fused features, which can improve the model's recognition ability for complex actions, thereby enhancing the model's stable understanding and judgment ability for background dynamic changes. Finally, the multi-task head is used to detect and classify the fused features to obtain the category information, behavior information and location information of the road user, thereby realizing the recognition of road user behavior from the vehicle perspective.
[0057] In an optional embodiment, the video data in the target data set is labeled to obtain corresponding label information, as follows:
[0058] The original image data and its labels of the dataset are read. In the original data, all category targets inherit the initial bounding box, and each road user is annotated according to the vehicle's onboard perspective. The annotation uses three types of labels: (1) category, such as car, cyclist or pedestrian; (2) behavior, such as moving, turning or crossing the road; (3) location, such as intersection, lane or sidewalk. Among them, all data is based on the vehicle's own viewpoint (EVV) - captured by the (front) camera installed on the vehicle. Through the above annotations, the self-built dataset is obtained. In the self-built dataset, annotations are provided for 3 category items, 10 action items and 11 location items, as shown in Table 1:
[0059] Table 1 Label annotation information of self-built dataset
[0060]
[0061] In an optional embodiment, the road user behavior recognition model further includes a frame sequence encoding network;
[0062] The video data and the corresponding label information are input into the road user behavior recognition model, and the slow path and the fast path are used to capture the static background information and dynamic action changes in the video data respectively, and the corresponding slow path spatiotemporal features and fast path spatiotemporal features are obtained, including:
[0063] Input the video data and the corresponding label information into a frame sequence coding network, and use the frame sequence coding network to process the video data into coded video frames;
[0064] The slow path is used to capture the long-term dependency and global structure information of the encoded video frames and obtain the slow path spatiotemporal features;
[0065] The fast path is used to capture the details and fast-changing behaviors in the encoded video frames and obtain the fast path spatiotemporal features.
[0066] In the present application embodiment, Figure 2 As shown in , the reconstruction of label information and video data reading are realized in the data input module. The video data processed by the self-built dataset and the reconstructed label information are input into the frame sequence encoding network. The video data is stored in the Videos folder, and the label information is named as the "train_val.json" file. Set a single frame-level annotation indicator tag to include ['annotated', 'rgb_image_id', 'width', 'height', 'annos'], where 'annos' covers the specific annotation content in a frame, including the bounding box 'box', category 'agent_ids', behavior 'action_ids' and location 'loc_ids'. In particular, 'box' is the coordinates of the bounding box coordinates normalized to (0,1) using 'xmin, ymin, xmax, ymax'.
[0067] The frame sequence coding network obtains video data input from a self-built dataset, which contains 798 video clips, and the frame rate of each video is 10s. -1 When a length of L is input (vis) The video clip is represented as a set of continuous image frame sequences, each frame is an RGB image, representing a moment in the video, and the corresponding decomposed image group is stored in the RGB folder. The input video data is processed by the encoding network into a fixed-size three-dimensional tensor (H, W, C), where H represents the height of the image, that is, 'height' in the label information, W represents the width of the image, that is, 'width' in the label information, and C represents the number of channels.
[0068] like Figure 2 As shown in , the frame sequence coding network encodes video frames from the input data set and passes them to the downstream hybrid dual-path backbone network. The coding process sets encoding parameters and extracts video frames according to the input requirements of the fast path and the slow path, and then performs fast path sampling and slow path sampling on the extracted video frames according to the corresponding parameters. The frame rate ratio between the fast path and the slow path is α=8, and the channel ratio is The time step τ = 16. The encoder network samples the video frame passed to the slow path with a time step τ slow =τ=16, number of channels C slow = C. When the number of sample frames passed to the slow path is T, the length of the video segment input to the slow path is L (vis) = Tτ; and the time step passed by the encoding network to the fast path Since the encoding process is performed on the same video, the number of sample frames passed to the fast path is αT, but the number of channels is only β times that of the slow path, that is, C fast =βC.
[0069] The hybrid dual-path backbone network in the embodiment of the present application is based on the overall architecture of slowfast. Different 3D convolutional neural network models are selected for the slow path and the fast path, and video data with different time resolutions are used to improve the efficiency and accuracy of behavior recognition. Finally, the outputs of the two paths are spliced. The structure of the backbone network is as follows: Figure 3 As shown in the figure, the slow path uses a 3D-ResNet50 network with C channels and T sampling frames. It captures long-term dependencies and global structural information in the video at a lower frame rate, and can process deeper networks due to the residual connection in the ResNet architecture. The fast path uses an inflated 3D convolutional neural network (Inflated-3D, I3D). The number of sampling frames passed to the fast path is αT and the number of channels is βC. The network captures details and rapidly changing behaviors in the video at a higher frame rate, and inherits the advantages of the 2D convolutional network, which can better process the temporal information in the video.
[0070] Among them, Figure 2 As shown in , the 3D-ResNet50 network of the slow path first performs tensor processing. Specifically, the encoded video frame is processed into a five-dimensional tensor of shape (N, T, H, W, C), where N is the batch size, T is the number of time frames, and H, W, C are the input three-dimensional tensor information. Then convolution operations and feature extraction are performed. Specifically, the tensor X∈R of the input video is N ×T×H×W×CAfter that, the convolution operation is performed in three dimensions (time dimension T, space dimension H and W) through the 3D convolution layer. A larger convolution kernel (7×7×7) slides on each dimension and calculates the weighted sum based on the weight parameter to output the spatiotemporal features. The process of the convolution operation is as follows: Figure 4 As shown. The calculation formula for the convolution operation is:
[0071]
[0072] Among them, x(t, x, y) represents the value of a position in the input tensor at time t and spatial position (x, y), ω(i, j, k) represents the weight of the convolution kernel, and y(t, x, y) represents the value of a position in the output tensor at time t and spatial position (x, y). The shape of the output tensor is (N, T', H', W', C') where T', H' and W' are the time and space dimensions after convolution, and C' is the number of output channels.
[0073] In particular, the 3DResNet-50 network inherits the residual structure of the standard ResNet, which includes multiple residual blocks, each of which contains multiple convolutional layers. After the upper layer output is used as the residual block input x, the calculation formula for the residual block output y is:
[0074] y=F(x,W)+W δ x
[0075] Among them, F(x, W) is the transformation result after multiple 3D convolutional layers, batch normalization and activation function, x is the input feature, W is the weight of the convolution operation of the convolution layer, and W δ represents a convolutional layer, W δ x is added to F(x, W) through a residual connection to form the final output.
[0076] After completing the residual layer processing, the 3DResNet-50 network has learned the deep slow path spatiotemporal features of the video data, which will be passed to the neck network and spliced and fused with the fast path output.
[0077] Similarly, if Figure 2 As shown in , the fast path also performs tensor processing, convolution operations, and feature extraction. Specifically, the I3D network of the fast path processes the encoded video frame into a five-dimensional tensor of shape (N, T, H, W, C). However, unlike the slow path, which directly uses the standard 3D convolution kernel, the I3D model is extended on the basis of the standard 2D convolution network, using the pre-trained 2D convolution network weights and converting them into 3D convolution layers through expansion. The structure of the I3D network is shown in Figure 5As shown in Figure 1. A standard 2D small convolution kernel is selected, with a size of 3×3. After expansion, it becomes 3×3×T, where T is the time dimension, which represents the increased convolution capability for processing video frames. The initial input of the network is a sequence of 1 to K frames of images and 1 to K frames of optical flow information with a time sequence, which are processed into a five-dimensional tensor. When the five-dimensional tensor X∈R is input N×T×H×W×C After that, it passes through the 3D convolution layer Conv3D, and the convolution kernel size is K T ×K H ×K W , the output feature is Y1∈R N×T×H×W×C , the operation formula of the first layer of convolution process is:
[0078] Y1 = Conv3D(X, W1)
[0079] W1 is the 3D convolution kernel. In the I3D network, after a series of 3D convolution and pooling operations, the final output is a tensor that has undergone spatiotemporal feature extraction. This multidimensional array is used to represent the fast path spatiotemporal features extracted by the model. These features will also be passed to the neck network and spliced and fused with the slow path output.
[0080] In an optional embodiment, the neck network includes a channel-time attention module, a channel-space attention module, and a space-time attention module;
[0081] The neck network is used to fuse the slow path spatiotemporal features and the fast path spatiotemporal features to obtain fused features, including:
[0082] The channel space attention module is used to weight each channel of the slow path spatiotemporal features in the spatial dimension to obtain the channel space weighted features;
[0083] The channel-time attention module is used to weight the channels of the fast-path spatiotemporal features at each time step to obtain channel-time weighted features;
[0084] The spatial-temporal attention module is used to fuse the channel spatial weighted features and the channel temporal weighted features to obtain the fused features.
[0085] In the embodiment of the present application, the core idea of applying the multi-dimensional attention mechanism is to improve the expressiveness of feature fusion and optimize feature selection by modeling the channel, space and time dimensions respectively. The multi-dimensional attention mechanism consists of three modules, which respectively model the channel-time (Channel-Temporal, CT) downstream of the fast path, the channel-space (Channel-Spatial, CS) downstream of the slow path, and the fused spatial-temporal (Spatial-Temporal, ST), so the multi-dimensional attention mechanism can also be called the channel-time-space attention (Channel-Spatial-Temporal Attention, CSTA) mechanism.
[0086] The Channel-SpatialAttention module weights each channel in the spatial dimension to highlight the key areas in space, i.e. Figure 2 The channel-space weighting shown in . The channel-temporal attention module is modeled at each time step, considering the correlation between channels and highlighting the important channel features in the time series, that is, Figure 2 After completing the weighted processing of the output features of the two paths, the spatial-temporal attention module is used to perform feature fusion operations. By modeling the relationship in the spatial and temporal dimensions, the interactive information between time and space is captured, and the interactive features in the weighted spatial and temporal dimensions are weighted to help the model capture the dynamic changes in the video, that is, Figure 2 The space-time weighting shown in .
[0087] Each module in the neck network in the embodiment of the present application is composed of a feature fusion layer that introduces an attention mechanism and a standard convolution with feature map size alignment.
[0088] In an optional embodiment, each channel of the slow path spatiotemporal feature is weighted in the spatial dimension using a channel space attention module to obtain a channel space weighted feature, including:
[0089] The slow path spatiotemporal feature map is compressed into the spatial dimension of each channel through global average pooling, and the channel-spatial attention weight of each channel in the spatial dimension is obtained using an activation function;
[0090] Based on the channel-space attention weight, each channel of the slow path spatiotemporal features is weighted in the spatial dimension to obtain the channel-space weighted features.
[0091] The structure of the channel space attention module is as follows Figure 6 As shown. For the input feature matrix X∈R N×T′×H′×W′×C′ Convolutional splicing is performed on the t1, t2, and t3 convolutional layers respectively, followed by channel-space weighting. The channel-space attention weight of each channel in the spatial dimension is obtained through global average pooling and softmax activation function. The calculation formula is as follows:
[0092] A CS =softmax(W CS GlobalAvgPool(X)
[0093] Among them, W CS ∈R C×C is a learning parameter. GlobalAvgPool(X) calculates the mean of each channel in the spatial dimension to generate a tensor of shape C×1×H×W. The channel-spatial attention weight A is obtained by the softmax activation function. CS Then it is weighted in the spatial dimension and convolved through a 1×1×1 convolution layer to obtain the channel space weighted feature X CS , the calculation formula is as follows:
[0094] X CS =X·A CS
[0095] Get X CS It is a channel-space weighted feature map.
[0096] In an optional embodiment, a channel time attention module is used to weight the channel of the fast path spatiotemporal feature at each time step to obtain a channel time weighted feature, including:
[0097] The fast path spatiotemporal features are compressed into the temporal dimension of each channel through global average pooling, and the channel-temporal attention weight of each time step is obtained using an activation function;
[0098] The fast-path spatiotemporal features are weighted on the channels at each time step based on the channel-temporal attention weights to obtain channel-time weighted features.
[0099] The structure of the channel time attention module is as follows Figure 7 As shown. For the input feature matrix X∈R N×T′×H′×W′×C′ After two dimensionality reductions, the convolution layer is used for convolution processing, and then channel-time weighting is performed along the time (T) dimension. The feature map is compressed to the time dimension of each channel through global average pooling, and then the time and channel are aligned and fused, and then input into the activation function, and the channel-time attention weight A is obtained through the softmax activation function. CT, the calculation formula is as follows:
[0100] A CT =softmax(W CT GlobalAvgPool(X)
[0101] Among them, W CT ∈R C×C is a learning parameter. GlobalAvgPool(X) calculates the mean of each channel in the time dimension, generates a tensor of shape C×T×1×1, and obtains the channel-time attention weight A by the softmax activation function. CT Then the channel of each time step is weighted and convolved through a 1×1×1 convolution layer to obtain the channel time weighted feature X CT , the calculation formula is as follows:
[0102] X CT =X·A CT
[0103] Get X CT It is a channel-time weighted feature map.
[0104] In an optional embodiment, if Figure 8 As shown, the spatial-temporal attention module is used to fuse the channel spatial weighted features and the channel temporal weighted features to obtain fused features, including:
[0105] Concatenate and pool the channel spatial weighted features and the channel temporal weighted features to obtain concatenated features;
[0106] After the concatenated feature tensor T′×H′×W′×C′ is convolved through a 1×1×1 convolutional layer, the activation function is used to capture the interaction between the concatenated features in time and space to obtain the spatial-temporal attention weight;
[0107] According to the spatial-temporal attention weight, the concatenated features (a tensor with a shape of 1×T×H×W) are convolved through a 1×1×1 convolutional layer and feature extracted to obtain the fused features.
[0108] In the embodiment of the present application, firstly, the channel space weighted feature X CS and channel time weighted features X CT The spliced features are then convolved through a 1×1×1 convolutional layer and global average pooling to obtain the spliced features. Then the space-time attention weight is obtained through the softmax activation function, and the calculation formula is as follows:
[0109]
[0110] Among them, W ST ∈R C×C are the learning parameters, The entire spatiotemporal feature map is concatenated and pooled to learn the interactive relationship between time and space.
[0111] Then the concatenated features (a tensor with a shape of 1×T×H×W) are convolved through a 1×1×1 convolutional layer to obtain Then, the features are spatially and temporally weighted based on the spatial-temporal attention weights to obtain the fused features. The calculation formula is as follows:
[0112]
[0113] Get X ST It is the spatiotemporal feature obtained through weighted fusion.
[0114] In an optional embodiment, the multi-task head adopts a 3D-RetinaNet network.
[0115] In the embodiment of the present application, the fused features are transmitted to the 3D-RetinaNet network for detection and classification tasks. The structure of the 3D-RetinaNet network is as follows: Fig. 9 As shown in the figure, the input part of the 3D-RetinaNet network is the fusion feature. The structure of the network includes the initial block on the left and the second block on the right (category + boundary subnet part). The initial block includes an extraction network that outputs a series of forward feature pyramid maps and a horizontal layer of the final feature pyramid composed of T feature maps. The second block includes a boundary subnet and a category subnet, which process T feature maps through a 1×3×3 convolution layer and a 3×3×3 convolution layer to generate C classification values for the bounding box (4 coordinates) and each anchor position (more than A possible positions). Then output the corresponding ('ids', 'agent_ids', 'action_ids', 'loc_ids') information to obtain the category, behavior and location recognition results of the road user.
[0116] The road user behavior recognition model used in the embodiments of the present application comprehensively considers several common types of road users in traffic environments, provides relevant self-built data sets as the data truth value for model training, and deeply analyzes the construction method and construction process of the behavior recognition model from a theoretical and technical level, providing support for the development of a driving assistance system based on road user motion detection.
[0117] Fig.10The figure shows the convergence of training data during the training of the road user behavior recognition model. The horizontal axis of the figure represents the number of training steps executed during model training, a total of 30 rounds of more than 300K times, and the vertical axis represents the loss value of the entire model training process, ranging from 0 to 1. After 30 epochs of training, the model obtained a relatively ideal training result, and the convergence effect met the requirements. Use Anaconda3 to build a training environment, use the self-built dataset as the training dataset, correctly place it in the Dataset folder, start the main.py script in PyCharm to train the model, and you can draw Figure 11-14 The multi-task detection curves during the 30 epoch training process are shown in Figure 1, where the target category detection curve (ClassAP-agent) is shown in Figure 2. Fig.11 As shown in Figure 2, the action behavior detection curve (ClassAP-action) is as follows: Fig.12 As shown, the position information detection curve (ClassAP-loc) is as follows Fig.13 As shown in Figure 2, the mAP curves (mAPs) of each detection item are as follows: Fig.14 As shown in Figure 2. Overall, the verification results obtained after 30 rounds of training are close to convergence, reflecting the good detection effect of the model.
[0118] Based on the same principle as the method provided in the embodiment of the present application, the embodiment of the present application also provides a road user behavior recognition system based on vehicle-mounted visual information, such as Fig.15 As shown, the system comprises:
[0119] The data labeling module 1501 is used to label the video data in the target data set to obtain corresponding label information; the label information includes target category, behavior action and location information;
[0120] The model building module 1502 is used to build a road user behavior recognition model; the road user behavior recognition model includes a hybrid dual-path backbone network, a neck network based on a multi-dimensional attention mechanism, and a multi-task head; the hybrid dual-path backbone network includes a slow path and a fast path;
[0121] The feature extraction module 1503 is used to input the video data and the corresponding label information into the road user behavior recognition model, and use the slow path and the fast path to capture the static background information and dynamic action changes in the video data, respectively, to obtain the corresponding slow path spatiotemporal features and the fast path spatiotemporal features;
[0122] The feature extraction module 1503 is also used to fuse the slow path spatiotemporal features and the fast path spatiotemporal features using the neck network to obtain fused features;
[0123] The information recognition module 1504 is used to detect and classify the fused features using the multi-task head to obtain the category information, behavior information and location information of the road user; wherein the multi-task head realizes information matching through label information in the detection and classification tasks.
[0124] In the embodiment of the present application, the video data in the target data set from the vehicle perspective is annotated with categories, behaviors and positions to obtain corresponding label information, which makes up for the scarcity of road user behavior data from the vehicle perspective and provides data support for the establishment of subsequent models. A road user behavior recognition model including a hybrid dual-path backbone network, a neck network and a multi-task head is constructed to realize the recognition of road user behavior from the vehicle perspective. Among them, the hybrid dual-path backbone network includes a slow path and a fast path, which can respectively capture dynamic action changes and static background information in the video data; then the neck network uses a multi-dimensional attention mechanism to fuse the slow path spatiotemporal features and the fast path spatiotemporal features extracted from the slow path and the fast path to obtain fused features, which can improve the model's recognition ability for complex actions, thereby enhancing the model's stable understanding and judgment ability for background dynamic changes. Finally, the multi-task head is used to detect and classify the fused features to obtain the category information, behavior information and location information of the road user, thereby realizing the recognition of road user behavior from the vehicle perspective.
[0125] The road user behavior recognition system based on vehicle-mounted visual information provided by the embodiment of the present application can achieve Figures 1 to 14 To avoid repetition, the various processes implemented in the method embodiment are not described here.
[0126] The road user behavior recognition system based on vehicle-mounted perspective visual information of the embodiment of the present application can execute the road user behavior recognition method based on vehicle-mounted perspective visual information provided by the embodiment of the present application, and the implementation principle is similar. The actions performed by each module and unit in the road user behavior recognition system based on vehicle-mounted perspective visual information in each embodiment of the present application correspond to the steps in the road user behavior recognition method based on vehicle-mounted perspective visual information in each embodiment of the present application. For the detailed functional description of each module of the road user behavior recognition system based on vehicle-mounted perspective visual information, please refer to the description of the corresponding road user behavior recognition method based on vehicle-mounted perspective visual information shown in the previous text, which will not be repeated here.
[0127] Based on the same principle as the method shown in the embodiment of the present application, the embodiment of the present application also provides an electronic device, which may include but is not limited to: a processor and a memory; a memory for storing a computer program; a processor for executing the road user behavior recognition method based on vehicle-mounted perspective visual information shown in any optional embodiment of the present application by calling a computer program. Compared with the prior art, the road user behavior recognition method based on vehicle-mounted perspective visual information provided by the present application annotates the category, behavior and position of the video data in the target data set from the vehicle perspective, obtains the corresponding label information, makes up for the scarcity of road user behavior data from the vehicle perspective, and provides data support for the establishment of subsequent models. A road user behavior recognition model including a hybrid dual-path backbone network, a neck network and a multi-task head is constructed to realize the recognition of road user behavior from the vehicle perspective. Among them, the hybrid dual-path backbone network includes a slow path and a fast path, which can capture dynamic action changes and static background information in video data respectively; then the neck network uses a multi-dimensional attention mechanism to extract the slow path and the fast path into the slow path spatiotemporal features and the fast path spatiotemporal features to obtain fusion features, which can improve the model's recognition ability for complex actions, thereby enhancing the model's stable understanding and judgment of background dynamic changes. Finally, the multi-task head is used to detect and classify the fusion features to obtain the category information, behavior information and location information of road users, thereby realizing the recognition of road user behavior from the vehicle perspective.
[0128] In an optional embodiment, an electronic device is also provided, such as Fig.16 As shown, Fig.16 The electronic device 1600 shown may be a server, including: a processor 1601 and a memory 1603. The processor 1601 and the memory 1603 are connected, such as through a bus 1602. Optionally, the electronic device 1600 may further include a transceiver 1604. It should be noted that in actual applications, the transceiver 1604 is not limited to one, and the structure of the electronic device 1600 does not constitute a limitation on the embodiments of the present application.
[0129] Processor 1601 may be a CPU (Central Processing Unit), a general-purpose processor, a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It may implement or execute various exemplary logic blocks, modules, and circuits described in conjunction with the disclosure of this application. Processor 1601 may also be a combination that implements computing functions, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, and the like.
[0130] The bus 1602 may include a path to transmit information between the above components. The bus 1602 may be a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus. The bus 1602 may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Fig.16 Only one thick line is used in the diagram, but this does not mean that there is only one bus or only one type of bus.
[0131] The memory 1603 may be a ROM (Read Only Memory) or other types of static storage devices that can store static information and instructions, a RAM (Random Access Memory) or other types of dynamic storage devices that can store information and instructions, or an EEPROM (Electrically Erasable Programmable Read Only Memory), a CD-ROM (Compact Disc Read Only Memory) or other optical disk storage, optical disk storage (including compressed optical disk, laser disk, optical disk, digital versatile disk, Blu-ray disk, etc.), a magnetic disk storage medium or other magnetic storage device, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited to these.
[0132] The memory 1603 is used to store the application code for executing the solution of the present application, and the execution is controlled by the processor 1601. The processor 1601 is used to execute the application code stored in the memory 1603 to implement the content shown in the above method embodiment.
[0133] Among them, electronic devices include but are not limited to: mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), and fixed terminals such as digital TVs and desktop computers. Fig.16 The electronic device shown is merely an example and should not bring any limitation to the functions and scope of use of the embodiments of the present application.
[0134] The server provided in this application can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. The terminal can be a smart phone, tablet computer, laptop computer, desktop computer, smart speaker, smart watch, etc., but is not limited to this. The terminal and the server can be directly or indirectly connected by wired or wireless communication, and this application does not limit this.
[0135] An embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer-readable storage medium is run on a computer, the computer can execute the corresponding content in the aforementioned method embodiment.
[0136] It should be understood that, although the steps in the flowchart of the accompanying drawings are displayed in sequence as indicated by the arrows, these steps are not necessarily executed in sequence in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least a part of the steps in the flowchart of the accompanying drawings may include multiple sub-steps or multiple stages, and these sub-steps or stages are not necessarily executed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be executed in turn or alternately with other steps or at least a part of the sub-steps or stages of other steps.
[0137] It should be noted that the computer-readable storage medium mentioned above in the present application can also be a computer-readable signal medium or a combination of a computer-readable storage medium and a computer-readable storage medium. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in combination with an instruction execution system, device or device. In the present application, a computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, which carries a computer-readable program code. This propagated data signal can take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer readable signal medium may also be any computer readable medium other than a computer readable storage medium, which may send, propagate or transmit a program for use by or in conjunction with an instruction execution system, apparatus or device. The program code contained on the computer readable medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.
[0138] The computer-readable medium may be included in the electronic device, or may exist independently without being incorporated into the electronic device.
[0139] The computer-readable medium carries one or more programs. When the one or more programs are executed by the electronic device, the electronic device executes the method shown in the above embodiment.
[0140] According to one aspect of the present application, a computer program product or a computer program is provided, the computer program product or the computer program comprising computer instructions, the computer instructions being stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the road user behavior recognition method and system based on vehicle-mounted viewpoint visual information provided in the above-mentioned various optional implementations.
[0141] Computer program code for performing the operations of the present application may be written in one or more programming languages or a combination thereof, including object-oriented programming languages, such as Java, Smalltalk, C++, and conventional procedural programming languages, such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a separate software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0142] The flow chart and block diagram in the accompanying drawings illustrate the possible architecture, function and operation of the system, method and computer program product according to various embodiments of the present application. In this regard, each square box in the flow chart or block diagram can represent a module, a program segment or a part of a code, and the module, the program segment or a part of the code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the square box can also occur in a sequence different from that marked in the accompanying drawings. For example, two square boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each square box in the block diagram and / or flow chart, and the combination of the square boxes in the block diagram and / or flow chart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0143] The modules involved in the embodiments described in this application can be implemented by software or hardware. The name of the module does not limit the module itself in some cases. For example, the data annotation module can also be described as "a data annotation module for annotating the video data in the target data set to obtain the corresponding label information".
[0144] The above description is only a preferred embodiment of the present application and an explanation of the technical principles used. Those skilled in the art should understand that the scope of disclosure involved in the present application is not limited to the technical solution formed by a specific combination of the above technical features, but should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above disclosed concept. For example, the above features are replaced with the technical features with similar functions disclosed in this application (but not limited to) by each other to form a technical solution.
Claims
1. A road user behavior recognition method based on vehicle-mounted visual information, characterized in that: The method comprises: Annotate the video data in the target data set to obtain corresponding label information; the label information includes category, behavior and location; Constructing a road user behavior recognition model; the road user behavior recognition model includes a hybrid dual-path backbone network, a neck network based on a multi-dimensional attention mechanism, and a multi-task head; the hybrid dual-path backbone network includes a slow path and a fast path; The video data and the corresponding label information are input into a road user behavior recognition model, and the slow path and the fast path are used to capture the static background information and the dynamic action changes in the video data, respectively, to obtain the corresponding slow path spatiotemporal features and the fast path spatiotemporal features; Using the neck network to fuse the slow path spatiotemporal features and the fast path spatiotemporal features to obtain fused features; The fused features are detected and classified using the multi-task head to obtain category information, behavior information and location information of road users; wherein the multi-task head realizes information matching through the label information in the detection and classification tasks.
2. The method for identifying road user behavior based on vehicle-mounted visual information according to claim 1, characterized in that: The road user behavior recognition model also includes a frame sequence encoding network; The step of inputting the video data and the corresponding label information into a road user behavior recognition model, using the slow path and the fast path to respectively capture static background information and dynamic action changes in the video data, and obtaining corresponding slow path spatiotemporal features and fast path spatiotemporal features includes: Inputting the video data and corresponding label information into the frame sequence coding network, and processing the video data into coded video frames using the frame sequence coding network; The slow path is used to capture the long-term dependency and global structure information of the encoded video frame to obtain the slow path spatiotemporal features; The fast path is used to capture details and fast-changing behaviors in the encoded video frame, and the fast path spatiotemporal features are obtained.
3. The method for identifying road user behavior based on vehicle-mounted visual information according to claim 2, characterized in that: The slow path uses the 3D-ResNet50 network; the fast path uses the expanded 3D convolutional neural network.
4. The method for identifying road user behavior based on vehicle-mounted visual information according to claim 1, characterized in that: The neck network includes a channel time attention module, a channel space attention module and a space time attention module; The using the neck network to fuse the slow path spatiotemporal features and the fast path spatiotemporal features to obtain fused features includes: Using the channel space attention module to weight each channel of the slow path spatiotemporal feature in the spatial dimension to obtain a channel space weighted feature; Using the channel time attention module to weight the channel of the fast path spatiotemporal feature at each time step to obtain a channel time weighted feature; The spatial-temporal attention module is used to fuse the channel spatial weighted features and the channel temporal weighted features to obtain the fused features.
5. The method for identifying road user behavior based on vehicle-mounted visual information according to claim 4, characterized in that: The method of using the channel space attention module to weight each channel of the slow path spatiotemporal feature in the spatial dimension to obtain a channel space weighted feature includes: The slow path spatiotemporal features are compressed into the spatial dimension of each channel by global average pooling, and the channel-spatial attention weight of each channel in the spatial dimension is obtained by using an activation function; Each channel of the slow path spatiotemporal feature is weighted in the spatial dimension based on the channel-space attention weight to obtain the channel space weighted feature.
6. The method for identifying road user behavior based on vehicle-mounted visual information according to claim 4, characterized in that: The step of using the channel time attention module to weight the channel of the fast path spatiotemporal feature at each time step to obtain a channel time weighted feature includes: The fast path spatiotemporal features are compressed into the temporal dimension of each channel by global average pooling, and the channel-temporal attention weight of each time step is obtained using an activation function; The fast path spatiotemporal feature is weighted on the channel at each time step based on the channel-time attention weight to obtain the channel-time weighted feature.
7. The method for identifying road user behavior based on vehicle-mounted visual information according to claim 4, characterized in that: The using the space-time attention module to fuse the channel space weighted features and the channel time weighted features to obtain the fused features includes: Performing splicing and pooling on the channel spatial weighted features and the channel temporal weighted features to obtain splicing features; The activation function is used to capture the interaction between the splicing features in time and space, and obtain the spatial-temporal attention weight; Feature extraction is performed on the spliced features according to the space-time attention weights to obtain the fused features.
8. A road user behavior recognition system based on vehicle-mounted visual information, characterized in that: The system comprises: A data annotation module is used to annotate the video data in the target data set to obtain corresponding label information; the label information includes category, behavior and location; A model building module, for building a road user behavior recognition model; the road user behavior recognition model includes a hybrid dual-path backbone network, a neck network based on a multi-dimensional attention mechanism, and a multi-task head; the hybrid dual-path backbone network includes a slow path and a fast path; a feature extraction module, configured to input the video data and the corresponding label information into a road user behavior recognition model, and use the slow path and the fast path to respectively capture static background information and dynamic action changes in the video data, and obtain corresponding slow path spatiotemporal features and fast path spatiotemporal features; The feature extraction module is further used to fuse the slow path spatiotemporal features and the fast path spatiotemporal features using the neck network to obtain fused features; An information recognition module is used to detect and classify the fused features using the multi-task head to obtain category information, behavior information and location information of road users; wherein the multi-task head realizes information matching through the label information in the detection and classification tasks.
9. An electronic device, characterized in that: The invention comprises a memory, a processor and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the method according to any one of claims 1 to 7 is implemented.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.