A method for classifying target behaviors, a storage medium, and a terminal

By identifying and classifying lane intrusion behavior in high-speed long-distance scenarios, and using pre-trained target behavior classification models, the problem of difficult to accurately identify and classify lane intrusion behavior in high-speed long-distance scenarios in the prior art is solved, and a fast and accurate prediction effect is achieved without navigation data and high-precision map information.

CN114612811BActive Publication Date: 2025-06-17TOYOTA JIDOSHA KK +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202011401830.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-12-04
Publication Date
2025-06-17
Estimated Expiration
2040-12-04

AI Technical Summary

Technical Problem

The prior art is difficult to accurately identify and classify lane intrusion behavior in high-speed long-distance scenarios, and has high requirements for navigation data and high-precision map information.

Method used

By obtaining multiple frames of images to be identified, identifying the moving target and lane lines, obtaining the behavior time series data of the moving target, and inputting a pre-trained target behavior classification model for prediction, including breaking into the lane from the left, breaking into the lane from the right, and not breaking into the lane.

Benefits of technology

It realizes the rapid and accurate prediction of lane breaking behavior in high-speed long-distance scenarios, without the need to use navigation data and high-precision map information, and reduces the requirements for on-board equipment and wireless network quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114612811B_ABST
    Figure CN114612811B_ABST
Patent Text Reader

Abstract

The present invention provides a method for classifying target behaviors, a storage medium, and a terminal. The method for classifying target behaviors includes: acquiring multiple frames of images to be recognized, where the multiple frames of images to be recognized reflect the lane scenes at different moments; recognizing moving targets and lane lines in the multiple frames of images to be recognized according to the multiple frames of images to be recognized; acquiring the behavioral time series data of the moving targets, where the behavioral time series data includes a plurality of relative position relationships arranged in the order of the moments corresponding to the multiple frames of images to be recognized; inputting the behavioral time series data into a pre-trained target behavior classification model to determine the behavioral prediction result of the moving targets, and the behavioral prediction result of the moving targets includes entering the lane from the left, entering the lane from the right, and not entering the lane.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing, and particularly to a method for classifying target behaviors, a storage medium, and a terminal. Background Art

[0002] The main technology for predicting lane intrusion behavior is action recognition, more precisely, action recognition in a high-speed and long-distance scenario. Video action recognition refers to recognizing the actions of targets (mainly referring to people) in a video from a continuous video (i.e., a sequence of image frames). A video contains two types of features: temporal features and spatial features. Extracting spatio-temporal features simultaneously requires calculating the relationships between pixels within a frame and between different frames.

[0003] Existing classification of lane intrusion behavior requires high-precision navigation data and map information, and has high requirements for network quality, in-vehicle device hardware, and in-vehicle device software updates.

[0004] Therefore, a new method for recognizing and classifying the actions of moving targets is needed. Summary of the Invention

[0005] The technical problem to be solved by the present invention is to propose a method for classifying target behaviors, a storage medium, and a terminal, which can realize the action recognition and classification of moving targets in a high-speed and long-distance scenario.

[0006] To solve the above problems, an embodiment of the present invention provides a method for classifying target behaviors, including: acquiring multiple frames of images to be recognized, where the multiple frames of images to be recognized reflect the lane scenes at different times; recognizing moving targets and lane lines in the multiple frames of images to be recognized according to the multiple frames of images to be recognized; acquiring the behavioral time series data of the moving targets, where the behavioral time series data includes a plurality of relative position relationships arranged in the order of the times corresponding to the multiple frames of images to be recognized; inputting the behavioral time series data into a pre-trained target behavior classification model to determine the behavioral prediction result of the moving target, where the behavioral prediction result of the moving target includes intrusion into the lane from the left, intrusion into the lane from the right, and no intrusion into the lane.

[0007] Optionally, the recognizing the moving targets and lane lines in the multiple frames of images to be recognized includes: respectively inputting each frame of the images to be recognized into a trained moving target detection model and a lane line detection model, where the moving target detection model outputs the pixel coordinates of the recognized moving targets, and the lane line detection model outputs the pixel coordinates of the recognized lane lines.

[0008] Optionally, the moving targets are represented by the pixel coordinates of the bounding boxes, and the lane lines are represented by the coordinates of their pixels in the images to be recognized.

[0009] Optionally, obtaining the behavioral time series data of the moving target includes: mapping the moving target and the current lane line recognized in each frame of the image to be recognized onto the same mapping image according to pixel coordinates, and each frame of the image to be recognized has a corresponding mapping image; determining a horizontal reference line in each mapping image, and determining the first pixel coordinate of the moving target on the horizontal reference line, as well as the second pixel coordinate and the third pixel coordinate of the two lines of the lane line on the horizontal reference line; obtaining the behavioral time series data of the moving target in the image to be recognized according to the first pixel coordinate, the second pixel coordinate, and the third pixel coordinate in each mapping image.

[0010] Optionally, the contact line between the moving target and the road surface serves as the horizontal reference line.

[0011] Optionally, obtaining the behavioral time series data of the moving target in the image to be recognized includes: determining the fourth pixel coordinate of the center point of the lane line on the horizontal reference line according to the second pixel coordinate and the third pixel coordinate in each mapping image; calculating the difference between the first pixel coordinate and the fourth pixel coordinate in each mapping image as the relative position relationship.

[0012] Optionally, obtaining the behavioral time series data of the moving target in the image to be recognized includes: determining the width of the lane line and the fourth pixel coordinate of the center point of the lane line on the horizontal reference line according to the second pixel coordinate and the third pixel coordinate in each mapping image; calculating the difference between the first pixel coordinate and the fourth pixel coordinate in each mapping image and normalizing it using the width of the lane line as the relative position relationship.

[0013] Optionally, after determining the horizontal reference line in each mapping image, and determining the first pixel coordinate of the moving target on the horizontal reference line, as well as the second pixel coordinate and the third pixel coordinate of the two lines of the lane line on the horizontal reference line, it includes: performing Kalman filtering on the normalized relative position relationship corresponding to each frame of the image to be recognized and the behavioral time series of the moving target to obtain a spatially smoothed behavioral time series.

[0014] Optionally, the target behavior classification model is a recurrent neural network model.

[0015] Optionally, the target behavior classification model is a two - layer bidirectional LSTM prediction model. The two - layer bidirectional LSTM prediction model outputs a first classification score, a second classification score, and a third classification score according to the behavior time - series data. The first classification score corresponds to the probability that the detected target breaks into the lane from the left, the second classification score corresponds to the probability that the detected target breaks into the lane from the right, and the third classification score corresponds to the probability that the detected target does not break into the lane. The maximum value among the first classification score, the second classification score, and the third classification score is the final behavior prediction result.

[0016] Optionally, the bidirectional LSTM prediction model includes: a two - layer bidirectional LSTM network and a fully - connected network. The two - layer bidirectional LSTM network serves as the input end of the bidirectional LSTM prediction model and is used to input the first 4 - order data of the behavior time - series. The data output by the last neuron of the two - layer bidirectional LSTM network serves as the input data of the fully - connected network.

[0017] Optionally, the fully - connected network serves as the output end of the bidirectional LSTM prediction model.

[0018] Optionally, the number of hidden layers of the two - layer bidirectional LSTM network is 16, its input data size is (K×1), and its output data size is (32×K); the input data size of the fully - connected network is (32×1), and the output is 3; where K is a natural number between 20 and 40.

[0019] An embodiment of the present invention also provides a storage medium, on which computer instructions are stored. When the computer instructions run, they execute the steps of the above - mentioned method.

[0020] An embodiment of the present invention also provides a terminal, including a memory and a processor. Computer instructions capable of running on the processor are stored on the memory. When the processor runs the computer instructions, it executes the steps of the above - mentioned method.

[0021] Compared with the prior art, the technical solution of the embodiment of the present invention has the following beneficial effects:

[0022] The technical solution of the present invention quickly and accurately predicts the lane - breaking behavior by obtaining the relative position relationship between the moving target and the lane line on the horizontal reference line in the mapped image, without using any navigation data and high - precision map information, and has low requirements for in - vehicle devices and wireless network quality.

[0023] Further, a recurrent neural network prediction model is used to classify the behaviors of the moving target based on the behavior time series, and a monocular camera is used for shooting. Even in a long-distance scenario, the position features of small targets can be captured, solving the problem of difficult capture of small target actions in long-distance scenarios. Moreover, the technical solution of the present invention has low cost and high accuracy.

[0024] Further, the relative position relationship is normalized by using the width of the lane line, accurately describing the motion trajectory and solving the problem of image jitter caused by high speed. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] Figure 1 is a schematic flowchart of a target behavior classification method provided by an embodiment of the present invention;

[0026] Figure 2 is a schematic flowchart of a method for obtaining the behavior time series data of the moving target provided by an embodiment of the present invention;

[0027] Figure 3 is a schematic diagram of a specific application scenario of an embodiment of the present invention;

[0028] Figure 4 is according to an embodiment of the present invention Figure 3 a schematic diagram of the pixel value curves of the lane line and the moving target obtained from the shown application scenario;

[0029] Figure 5 is Figure 2 a schematic flowchart of a specific implementation manner of S303 shown;

[0030] Figure 6 is a schematic diagram of the relative position relationship of the first pixel coordinate, the second pixel coordinate, and the second pixel coordinate changing with time in an embodiment of the present invention;

[0031] Figure 7 is a specific schematic diagram of the relative position relationship of the first pixel coordinate, the second pixel coordinate, and the second pixel coordinate changing with time after normalization in an embodiment of the present invention; and

[0032] Figure 8 is a schematic diagram of the framework structure of a bidirectional LSTM prediction model in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0033] For video action recognition, there are currently two mainstream methods: 3D convolutional neural network and two-stream method. The 3D spatio-temporal convolutional method uses a 3D convolutional kernel to capture both temporal and spatial features simultaneously. The 3D convolutional kernel adds one dimension compared to the 2D convolutional kernel. The 3D convolutional kernel can calculate the pixel relationships in the same region of several consecutive frames. Therefore, it can extract both temporal and spatial features, thereby achieving video action recognition. The two-stream method includes an optical flow network branch and an image network branch, representing temporal features with optical flow images. Optical flow is the instantaneous velocity of the pixel motion of a moving object in space on the observation imaging plane. The optical flow method is a method that uses the change of pixels in the time domain of an image sequence and the correlation between adjacent frames to find the corresponding relationship between the previous frame and the current frame, and thus calculates the motion information of the object between adjacent frames. The two-stream method extracts temporal features with a 2D convolutional kernel in the optical flow network branch, extracts spatial features in the image network branch, and then fuses the two to achieve the extraction of spatio-temporal features, and further classifies video actions.

[0034] However, 3D convolution and the two-stream method are not applicable to high-speed long-distance scenarios. A high-speed long-distance scenario means that the camera is moving at high speed and the distance between the action target and the camera is greater than a preset distance, such as 100 meters. In this case, the irregular visual changes of the moving target in the image frame and the severe jitter of the image may both cause the changes in the image frame not to truly reflect the actual motion of the target. Whether it is the 3D convolution method or the optical flow method, considering that they both extract the change features of the corresponding pixels in the image frame, it is obvious that the entire behavior cannot be recognized only based on the absolute changes of the pixels in the target area of the image. In addition, due to the long distance, the action target only occupies a small part of the pixel space in the video image frame plane. Both 3D convolution and the two-stream method will extract the pixel features of the entire image, including the background, and the proportion of the moving target features is very small. Therefore, it is difficult for these methods to capture the action features of such small targets and even more difficult to accurately predict their behavioral intentions.

[0035] Therefore, the present invention proposes a target behavior recognition method applicable to high-speed long-distance scenarios, and classifies the target behavior based on a prediction model.

[0036] The moving target in the embodiments of the present invention generally refers to a pedestrian or an animal performing a behavior of breaking into a motor vehicle lane line, or a person riding a non-motor vehicle.

[0037] The "lane line" referred to in the embodiments of the present invention may refer to a motor vehicle lane line, which usually appears in pairs. The two lines of the lane line can define the driving range of the motor vehicle, that is, the prohibited area for non-motor vehicles.

[0038] In the embodiments of the present invention, the "high-speed long-distance scenario" may refer to a scenario where the moving speed of the camera for collecting video is greater than a preset speed, such as 60 km / h; and the distance between the camera and the moving target to be photographed is greater than a preset distance, for example, 100 meters. In a non-limiting example, the camera may be disposed in a vehicle, the driving speed of the vehicle is greater than 60 km / h, and the distance between the vehicle and the moving target is greater than 100 meters.

[0039] In the embodiments of the present invention, the "target behavior" may be a lane intrusion behavior, that is, a behavior in which a moving target enters the range defined by two lane lines.

[0040] To make the above objects, features, and advantages of the present invention more obvious and understandable, the following will describe in detail the specific embodiments of the present invention with reference to the accompanying drawings.

[0041] Figure 1 It is a schematic flowchart of a target behavior classification method according to an embodiment of the present invention.

[0042] The target behavior classification method according to the embodiments of the present invention can be used on the terminal device side. For example, it can be an in-vehicle device, that is, it can be executed by the in-vehicle device. Figure 1 Each step of the method shown.

[0043] Specifically, Figure 1 The target behavior classification method shown may include:

[0044] S1: Obtain multiple frames of images to be recognized, and the multiple frames of images to be recognized reflect the lane scenarios at different times;

[0045] S2: Identify the moving target and the lane lines in the multiple frames of images according to the multiple frames of images to be recognized;

[0046] S3: Obtain the behavior time series data of the moving target, and the behavior time series data includes a plurality of relative position relationships arranged in the order of the times corresponding to the multiple frames of images to be recognized;

[0047] S4: Input the behavior time series data into a pre-trained target behavior classification model to determine the behavior prediction result of the moving target, and the behavior prediction result of the moving target includes intruding into the lane from the left, intruding into the lane from the right, and not intruding into the lane.

[0048] It should be noted that the serial numbers of the steps in this embodiment do not represent the limitation of the execution order of each step.

[0049] In a specific implementation of S1, the in-vehicle device can obtain multiple frames of images to be recognized from an in-vehicle camera. The multiple frames of images to be recognized can be the video to be recognized captured by the camera or multiple consecutive photos taken. The in-vehicle camera can be set inside the in-vehicle device or externally coupled to the in-vehicle device. Specifically, the in-vehicle camera can collect photos or videos in real time and transmit them to the in-vehicle device for target behavior recognition. In a specific embodiment, the in-vehicle camera can be a monocular camera. In other embodiments, the in-vehicle camera can be a wide-angle or telephoto camera.

[0050] In a specific implementation of S2, the in-vehicle device performs target detection on each frame of the images to be recognized respectively. That is to say, the recognition of moving targets and lane lines is carried out independently. The moving targets and lane lines in each frame of the images to be recognized are represented by their pixel coordinates in the image to be recognized. More specifically, the moving targets are represented by the pixel coordinates of the bounding box, and the lane lines are represented by the coordinates of their pixels in the image to be recognized.

[0051] In a non-limiting example, the bounding box of the moving target is defined as the coordinates in the u-v coordinate system of the image pixels, including its upper left coordinates (u l,t , v l,t ) and the lower right coordinates (u r,b , v r,b ).

[0052] In a non-limiting example, when recognizing lane lines, it can be determined through a grayscale image. The smaller the grayscale value of a pixel, the greater the probability that the pixel is a lane line. Therefore, the coordinates of the left and right lane lines in the image coordinate system can be obtained in the grayscale image.

[0053] In a non-limiting embodiment, Figure 1 S12 shown can include the following steps: input each frame of the images to be recognized into a trained moving target detection model and a lane line detection model respectively. The moving target detection model outputs the pixel coordinates of the recognized moving targets, and the lane line detection model outputs the pixel coordinates of the recognized lane lines.

[0054] In this embodiment, the moving target detection model and the lane line detection model can be constructed and pre-trained using deep learning algorithms, and can be used for the detection of moving targets and lane lines respectively. Specifically, the moving target detection model can be trained using the annotated video or multiple consecutive image data taken, and the video or multiple consecutive image data taken include the annotated moving targets; the lane line detection model can be trained using the annotated video or multiple consecutive image data taken, and the video or multiple consecutive image data taken include the annotated lane lines.

[0055] It should be noted that the specific method for constructing and training a moving target detection model and a lane line detection model using deep learning can refer to the prior art, and the embodiments of the present invention will not elaborate herein.

[0056] Figure 2 It is a schematic flowchart of a method for obtaining the behavioral time series data of the moving target provided by an embodiment of the present invention. Refer to Figure 2 , in S3, obtaining the behavioral time series data of the moving target includes:

[0057] S301: Map the moving target and the current lane line recognized in each frame of the image to be recognized onto the same mapped image according to pixel coordinates, and each frame of the image to be recognized has a corresponding mapped image;

[0058] S302: Determine a horizontal reference line in each mapped image, and determine the first pixel coordinate of the moving target on the horizontal reference line, and the second and third pixel coordinates of the two lines of the lane line on the horizontal reference line;

[0059] S303: Obtain the behavioral time series data of the moving target in the image to be recognized according to the first, second, and third pixel coordinates in each mapped image.

[0060] In the specific implementation of S301, since the moving target and the lane line are detected separately, it is necessary to map them onto the same image. Compared with the original image to be recognized, the ratio of the moving target and the lane line in the mapped image becomes larger, laying a foundation for the behavior prediction of the moving target.

[0061] In the specific implementation of S302, it can be reasonably assumed that the lower edge line of the moving target bounding box is the contact line between the moving target and the road surface, and it is used as the horizontal reference line of the current frame image.

[0062] Specifically, please refer to Figure 3 , make a straight line extension along the contact line, and this line is the horizontal reference line L that passes through the moving target T and is perpendicular to the current lane lines V1 and V2. Construct a one-dimensional coordinate system on the horizontal reference line L, and the positions of the moving target T and the lane lines V1 and V2 can be clearly represented. This one-dimensional coordinate system can take the u coordinate system in the image coordinate system. The interval [u l,t , u r,b represents the area occupied by the moving target, and the point O represents the center of the moving target, and its coordinate value is (u l,t , u r,b ) / 2, where u l,t represents the upper left coordinate and the lower right coordinate of the bounding box of the moving target T on the u coordinate system. Thus, the pixel positions of the moving target and the lane line in the unified one-dimensional coordinate system can be determined.

[0063] Further, referring to Figure 4 , Figure 4 which is a schematic diagram of the pixel value curves of the lane lines and moving targets obtained by the embodiments of the present invention according to Figure 3 the application scenario shown. Figure 4 In , the abscissa represents the pixel coordinates on the horizontal reference line, and the ordinate represents the probability p that the pixel is a lane line. In some embodiments, when determining the second pixel coordinate and the third pixel coordinate of the two lines of the lane line on the horizontal reference line, each spike represents a lane line. That is Figure 4 the peaks L1 and L2 in are the center points of the left and right lane lines respectively. Generally speaking, the distance between the two lines of the lane line, that is, the distance between point L1 and point L2 should be 2-4 times the moving target area (that is, the width occupied by the bounding box on the horizontal reference line).

[0064] Specifically, some pixel points of the lane line may not be detected due to occlusion, etc. At this time, the center points L1 and L2 of the lane line can be fitted with a quadratic polynomial to ensure that the lane line is a continuous and complete curve.

[0065] In the specific implementation of S303, the relative position relationship between the first pixel coordinate and the second pixel coordinate and the third pixel coordinate can be calculated to characterize the relative position relationship between the moving target and the lane line.

[0066] Specifically, the relative position relationship between the first pixel coordinate and the second pixel coordinate and the third pixel coordinate can be represented by the difference between the first pixel coordinate and the pixel coordinate of the midpoint of the lane line on the horizontal reference line. The midpoint position of the lane line on the horizontal reference line is the sum of the second pixel coordinate and the third pixel coordinate divided by 2.

[0067] Furthermore, according to each frame of the image to be recognized, a relative position relationship can be determined. Therefore, a behavior time series data is finally obtained. The behavior time series data includes a plurality of relative position relationships arranged in the time order of multiple frames of the image to be recognized.

[0068] In a non-limiting embodiment, Figure 2 the S303 shown may include: determining the fourth pixel coordinate of the center point of the lane line on the horizontal reference line according to the second pixel coordinate and the third pixel coordinate in each mapped image; calculating the difference between the first pixel coordinate and the fourth pixel coordinate in each mapped image as the relative position relationship.

[0069] In this embodiment, first calculate the fourth pixel coordinate L of the center point of the lane line on the horizontal reference line c =(L1 + L2) / 2. Then calculate the difference P between the first pixel coordinate and the fourth pixel coordinate r=(O - L c ), where L c represents the fourth pixel coordinate of the center point of the current lane, L1 and L2 respectively represent the pixel coordinates of the left and right lane lines of the current lane, O is the pixel coordinate of a moving object such as a pedestrian or a cyclist, and P r represents the positional relationship of the moving object relative to the center point.

[0070] In another non - restrictive embodiment, please refer to Figure 5 , Figure 5 which Figure 2 is a schematic flowchart of a specific implementation manner of S303 shown in Figure 2 The S303 shown in

[0071] S3031: Determine the width of the lane line and the fourth pixel coordinate of the center point of the lane line on the horizontal reference line according to the second pixel coordinate and the third pixel coordinate in each mapped image;

[0072] S3032: Calculate the difference between the first pixel coordinate and the fourth pixel coordinate in each mapped image, and normalize it using the width of the lane line to be used as the relative positional relationship.

[0073] Different from the foregoing embodiments, in the embodiment of the present invention, the difference between the first pixel coordinate and the fourth pixel coordinate is normalized.

[0074] In a high - speed long - distance scenario, as the camera moves at a high speed with the vehicle, there will be obvious problems such as scale change of the moving object and image jitter. At this time, the positions of the moving object and the lane line can only represent the positional relationship in the current frame, and the position time series of the moving object based on this cannot accurately reflect the movement behavior of the object. In order to accurately describe the movement of the object, this embodiment uses a normalized relative positioning method to solve the problems of scale change and image jitter.

[0075] where w = |L1 - L2|, L c =(L1 + L2) / 2, P r =(O - L c ) / w, w represents the pixel width of the current lane, L c represents the fourth pixel coordinate of the center point of the current lane, L1 and L2 respectively represent the pixel coordinates of the left and right lane lines of the current lane, O is the pixel coordinate of a moving object such as a pedestrian or a cyclist, and P r represents the positional relationship of the moving object relative to the center point.

[0076] Specifically, please refer to Figure 6 and Figure 7 , the abscissa of which represents time, Figure 6The lines l1 and l2 in the figure are obtained by fitting the second pixel coordinate and the third pixel coordinate in the behavioral time series data, respectively representing the two lines of the lane line. The line o is obtained by fitting the first pixel coordinate in the behavioral time series data, representing the moving target. From Figure 6 it can be seen that due to the camera jitter and scale change in the high-speed scenario, the width of the lane line also changes, which will lead to inaccurate judgment of the intrusion behavior of the moving target.

[0077] Figure 7 Then it shows that after the normalization process, Figure 6 the relative relationship between the line o and the lines l1 and l2 in the figure. Among them, the abscissa represents time, and the ordinate represents the normalized value of the difference between the first pixel coordinate and the fourth pixel coordinate. A negative value indicates that the moving target is on the left side of the center point of the lane line, and a positive value indicates that the moving target is on the right side of the center point of the lane line. Combining Figure 6 and Figure 7 it can be seen that after the normalization process, the width of the lane is stable, solving the problems of image scale change and image jitter in the images captured by the camera in the high-speed scenario.

[0078] In a non-limiting embodiment, Figure 2 the S302 shown may include: performing Kalman filtering on the first pixel coordinate, the second pixel coordinate, and the third pixel coordinate corresponding to each frame of the image to be recognized.

[0079] When using traditional or deep learning-based visual target detection methods to perform real-time tracking and detection of moving targets and lane lines, false detections and missed detections will inevitably occur. To reduce the adverse effects brought by false detections and missed detections, this embodiment performs filtering processing on the calculated moving targets and lane lines. For high-speed application scenarios, the number of moving targets is small, for example, within 5. At this time, the difference between adjacent two frames of images is very small, so the observed values of adjacent two frames can be compared bidirectionally to minimize their mean absolute error, so as to achieve the purpose of target alignment. After alignment, the ratio r = n / N between the number of frames n in which each moving target appears and the total number of frames N of the images is statistically calculated. If r < T (T is a preset threshold), it is considered that the target is a false detection and the target is discarded. The size of the preset threshold T can be determined with reference to the precision rate and recall rate of the target detection model. After obtaining the aligned moving target sequence, the false detections and missed detections generated by the target detection model are regarded as outliers and random noises, and the Kalman filter is used to remove the outliers and smooth the noises. Then the accurate continuous time series data of the moving targets and lane lines in the one-dimensional coordinate system can be obtained.

[0080] Refer to Figure 1, in the specific implementation of S4, the behavior time series data obtained in the foregoing steps is input into a pre-trained target behavior classification model, and the target behavior classification model classifies the behavior of the moving target. Optionally, the target behavior classification model is a recurrent neural network model. In a non-limiting embodiment, the target behavior classification model is a pre-trained bidirectional LSTM prediction model, and the bidirectional LSTM prediction model classifies the intrusion behavior of the moving target according to the relative position relationship in the behavior time series data, and outputs a first classification score, a second classification score, and a third classification score, where the first classification score corresponds to the probability that the detected target intrudes into the lane from the left, the second classification score corresponds to the probability that the detected target intrudes into the lane from the right, and the third classification score corresponds to the probability that the detected target does not intrude into the lane; the maximum value among the first classification score, the second classification score, and the third classification score is the final behavior prediction result.

[0081] Reference Figure 8 , Figure 8 is a schematic diagram of the framework structure of a bidirectional LSTM prediction model according to an embodiment of the present invention. The bidirectional LSTM prediction model includes: a two-layer bidirectional LSTM network 10 and a fully connected network 20, where the two-layer bidirectional LSTM network 10 serves as the input end of the entire prediction model and inputs the relative position relationship, that is, P ri (i = 1.....k), in the embodiment of the present invention, taking the simultaneous input of k data as an example, where k is a natural number between 20 and 40; the fully connected network 20 serves as the output end of the entire prediction model and outputs a first classification score R1, a second classification score R2, and a third classification score R3.

[0082] Continue to refer to Figure 8 , the two-layer bidirectional LSTM network 10 includes a plurality of neurons 101 ( Figure 8 shows K neurons 101), and each neuron 101 includes: an input unit S i (i = 1.....k), an output unit f i (i = 1.....k), a forward hidden layer unit F 1i (i = 1.....k) and F 2i (i = 1.....k), and a backward hidden layer unit B 1i (i = 1.....k) and B 2i (i = 1.....k). The input unit S i (i = 1.....k) is used to input the first 4-order data of the behavior time series P ri (i = 1.....k), and the input behavior time series data respectively enters the corresponding forward hidden layer unit F 1i(i = 1.....k), F 2i (i = 1.....k), and the reverse hidden layer unit B 1i (i = 1.....k), B 2i (i = 1.....k), each of the hidden layer units undertakes the main data interaction and computing work, and finally outputs the intermediate calculation results through their respective corresponding output units f i (i = 1.....k).

[0083] In a non - restrictive embodiment, the number of hidden layers of the two - layer bidirectional LSTM network 10 is 16. If the size of its input data is (K×1), then the size of its output data is (32×K). Among them, the data output by the last neuron of the two - layer bidirectional LSTM network is used as the input data of the fully - connected network 20, and its size is (32×1). The fully - connected network 20 classifies the intrusion behavior of the moving target according to the data output by the last neuron of the two - layer bidirectional LSTM network, and outputs the first classification score R1, the second classification score R2, and the third classification score R3. The maximum value among the first classification score R1, the second classification score R2, and the third classification score R3 is the behavior prediction result of the moving target. That is to say, if R1 is greater than R2 and R3, it is determined that the moving target intrudes into the lane from the left; if R2 is greater than R1 and R3, it is determined that the moving target intrudes into the lane from the right; if R3 is greater than R1 and R2, it is determined that the moving target does not intrude into the lane.

[0084] Since the bidirectional LSTM prediction model has been pre - trained, it classifies the behavior of the moving target according to the relative position relationship in the behavior time - series data, with a fast response speed and high classification result accuracy; and it does not need to use any navigation data and high - precision map information, reducing the difficulty of lane intrusion behavior classification.

[0085] In a non - restrictive embodiment, the target behavior classification method can be applied to an automotive intelligent driving solution. The in - vehicle device can give a prompt or an alarm according to the above - mentioned target behavior classification result to remind the driver or trigger the vehicle's own autonomous driving system to make an early response, improving the safety driving coefficient of the vehicle.

[0086] The embodiment of the present invention also discloses a storage medium. The storage medium is a computer - readable storage medium, on which a computer program is stored. When the computer program runs, it can execute the steps of the target behavior classification method shown in the above - mentioned method. The storage medium can include ROM, RAM, magnetic disk, or optical disc, etc. The storage medium can also include non - volatile memory or non - transitory memory, etc.

[0087] An embodiment of the present invention also discloses a terminal, which may include a memory and a processor, and a computer program that can run on the processor is stored on the memory. When the processor runs the computer program, it can execute the steps of the target behavior classification method shown in the above method.

[0088] Through the technical solution of the present invention, by obtaining the relative position relationship between the moving target and the lane line on the horizontal reference line in the mapped image, a trained prediction model is used to quickly and accurately predict the lane intrusion behavior, without using any navigation data and high-precision map information, and the requirements for in-vehicle devices and wireless network quality are low.

[0089] Furthermore, a recurrent neural network prediction model is used to classify the behavior of the moving target based on the behavior time series, and a monocular camera is used for shooting. Even in a long-distance scenario, the position features of small targets can be captured, solving the problem of difficult capture of small target actions in a long-distance scenario. Moreover, the technical solution of the present invention has low cost and high accuracy.

[0090] Furthermore, the relative position relationship is normalized by using the width of the lane line, accurately describing the movement trajectory and solving the problem of image jitter caused by high speed.

[0091] Although the present invention is disclosed as above, the present invention is not limited thereto. Any person skilled in the art can make various changes and modifications without departing from the spirit and scope of the present invention. Therefore, the protection scope of the present invention should be subject to the scope defined by the claims.

Claims

1. A method for classifying target behaviors, characterized in that, Including: Obtaining multiple frames of images to be recognized, where the multiple frames of images to be recognized reflect the lane scenes at different times; Identifying moving targets and lane lines in the multiple frames of images to be recognized according to the multiple frames of images to be recognized; Obtaining the behavioral time series data of the moving target, where the behavioral time series data includes a plurality of relative position relationships arranged in the order of the times corresponding to the multiple frames of images to be recognized; And Inputting the behavioral time series data into a pre-trained target behavior classification model to determine the behavioral prediction result of the moving target, where the behavioral prediction result of the moving target includes breaking into the lane from the left, breaking into the lane from the right, and not breaking into the lane; Wherein, the obtaining the behavioral time series data of the moving target includes: Mapping the moving target and the lane lines recognized in each frame of the image to be recognized onto the same mapped image according to pixel coordinates, and each frame of the image to be recognized has a corresponding mapped image; Determining a horizontal reference line in each mapped image, and determining the first pixel coordinate of the moving target on the horizontal reference line, and the second pixel coordinate and the third pixel coordinate of the two lines of the lane line on the horizontal reference line; and Obtaining the behavioral time series data of the moving target in the image to be recognized according to the first pixel coordinate, the second pixel coordinate, and the third pixel coordinate in each mapped image; The target behavior classification model is a two-layer bidirectional LSTM prediction model. The two-layer bidirectional LSTM prediction model outputs a first classification score, a second classification score, and a third classification score according to the behavioral time series data. Wherein, the first classification score corresponds to the probability that the detected target breaks into the lane from the left, the second classification score corresponds to the probability that the detected target breaks into the lane from the right, and the third classification score corresponds to the probability that the detected target does not break into the lane; the maximum value among the first classification score, the second classification score, and the third classification score is the final behavioral prediction result.

2. The method for classifying target behaviors according to claim 1, characterized in that, The identifying the moving targets and lane lines in the multiple frames of images to be recognized includes: Inputting each frame of the image to be recognized into a trained moving target detection model and a lane line detection model respectively. The moving target detection model outputs the pixel coordinates of the recognized moving target, and the lane line detection model outputs the pixel coordinates of the recognized lane line.

3. The method for classifying target behaviors according to claim 2, characterized in that, The moving target is represented by the pixel coordinates of the bounding box, and the lane line is represented by the coordinates of its pixels in the image to be recognized.

4. The method for classifying target behaviors according to claim 1, characterized in that, The contact line between the moving target and the road surface is used as the horizontal reference line.

5. The method for classifying target behaviors according to claim 1, characterized in that, The obtaining the behavioral time series data of the moving target in the image to be recognized includes: Determining the fourth pixel coordinate of the center point of the lane line on the horizontal reference line according to the second pixel coordinate and the third pixel coordinate in each mapped image; Calculating the difference between the first pixel coordinate and the fourth pixel coordinate in each mapped image as the relative position relationship.

6. The method for classifying target behaviors according to claim 1, characterized in that, The obtaining the behavioral time series data of the moving target in the image to be recognized includes: Determine the width of the lane line according to the second pixel coordinate and the third pixel coordinate in each mapped image, and the fourth pixel coordinate of the center point of the lane line on the horizontal reference line; Calculate the difference between the first pixel coordinate and the fourth pixel coordinate in each mapped image, and normalize it using the width of the lane line to be used as the relative position relationship.

7. The method for classifying target behaviors according to claim 6, characterized in that, The determining of the horizontal reference line in each mapped image, and the determining of the first pixel coordinate of the moving target on the horizontal reference line, and the second pixel coordinate and the third pixel coordinate of the two lines of the lane line on the horizontal reference line include: Perform Kalman filtering on the normalized relative position relationship corresponding to each frame of the image to be recognized and the behavior time series of the moving target to obtain a spatially smoothed behavior time series.

8. The method for classifying target behaviors as claimed in claim 1, characterized in that, The bidirectional LSTM prediction model includes: a two-layer bidirectional LSTM network and a fully connected network. The two-layer bidirectional LSTM network serves as the input end of the bidirectional LSTM prediction model and is used to input the data of the first 4 orders of the behavior time series. The data output by the last neuron of the two-layer bidirectional LSTM network is used as the input data of the fully connected network.

9. The method for classifying target behaviors according to claim 8, characterized in that, The fully connected network serves as the output end of the bidirectional LSTM prediction model.

10. The method for classifying target behaviors according to claim 8, characterized in that, The number of hidden layers of the two-layer bidirectional LSTM network is 16, the size of its input data is K×1, and the size of its output data is 32×K; the size of the input data of the fully connected network is 32×1, and the output data is 3; where K is a natural number between 20 and 40.

11. A storage medium, on which computer instructions are stored, characterized in that, When the computer instructions run, they execute the steps of the method according to any one of claims 1 to 10.

12. A terminal, comprising a memory and a processor, wherein computer instructions capable of running on the processor are stored on the memory, characterized in that, When the processor runs the computer instructions, it executes the steps of the method according to any one of claims 1 to 10.

Citation Information

Patent Citations

  • Vehicle control device, vehicle control method, and storage medium

    CN110271547A

  • Vehicle illegal lane change detection method and device

    CN111523464A