Front vehicle lane changing behavior recognition method and device and electronic equipment
By combining the characteristics of the running data and driving images, the identification model can accurately identify the intention of changing lanes in front of the car, solving the problem of inaccurate identification in the prior art and improving traffic safety.
Patent Information
- Application Number
- CN202510198031.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-21
- Publication Date
- 2025-06-27
AI Technical Summary
The prior art is difficult to accurately identify the intention of changing lanes in front of the car, especially in driving scenarios not included in the ideal data set, which can easily lead to misidentification and traffic accidents.
By obtaining the running data of the target front car corresponding to the vehicle and the driving image of the target front car driving, input it to the target front car lane change behavior recognition model, and outputting preliminary lane change classification results, including changing lanes to the left, changing lanes to the right and not changing lanes. The model combines running timing features and position image features to improve anti-interference ability and recognition accuracy through the fusion of attention mechanisms and features.
It significantly improves the accuracy and meticulousness of the recognition of lane change behavior in front of the car, and can stably output preliminary lane change classification results in complex environments, reduce the probability of traffic accidents, and ensure driving safety.
Smart Images

Figure CN120217073A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of intelligent driving, and particularly to a method, device and electronic device for identifying the lane-changing behavior of a preceding vehicle. Background Art
[0002] Lane-changing behavior is generally considered to be one of the most difficult driving tasks to identify, and it is also an important cause of traffic congestion and collisions. In real driving scenarios, the ambiguity of lane-changing intention presents multi-dimensional characteristics. In the time-sequence dimension, there is a phenomenon of delayed triggering of the turn signal by the driver, and even a reverse operation mode where the turning behavior is performed first and then the turn signal is prompted. In the behavior dimension, the driver may have non-intentional mis-touching, and the turn signal is turned on by mis-touching without the intention of changing lanes. Therefore, early and accurate identification of lane-changing intention can better judge the lane-changing behavior of the preceding vehicle, which is of great significance for reducing traffic accidents.
[0003] However, the data used in existing lane-changing intention recognition research usually comes from a real vehicle driving dataset in a certain period. Such a dataset is an ideal dataset. If there is a driving scenario in the real world that is not included in the ideal dataset, the existing research will be difficult to handle, and even incorrect recognition may occur, leading to traffic accidents.
[0004] Therefore, how to accurately determine the lane-changing behavior of the preceding vehicle has become an urgent problem to be solved. Summary of the Invention
[0005] In view of this, the present invention provides a method, device and electronic device for identifying the lane-changing behavior of a preceding vehicle to solve the problem of how to accurately determine the lane-changing behavior of the preceding vehicle.
[0006] In a first aspect, the present invention provides a method for identifying the lane-changing behavior of a preceding vehicle, the method comprising:
[0007] Obtaining the running data of a target preceding vehicle corresponding to the host vehicle; the running data includes at least one of the relative speed of the target preceding vehicle relative to the host vehicle, the relative distance of the target preceding vehicle relative to the host vehicle, the lateral displacement of the target preceding vehicle, and the lateral acceleration of the target preceding vehicle;
[0008] Obtaining multiple frames of target preceding vehicle driving images corresponding to the target preceding vehicle; the target preceding vehicle driving images include the target preceding vehicle;
[0009] Inputting the running data and each frame of the preceding vehicle driving image into a target preceding vehicle lane-changing behavior recognition model, and outputting a preliminary lane-changing classification result corresponding to the target preceding vehicle, the preliminary lane-changing classification result including changing lanes to the left, changing lanes to the right, and not changing lanes.
[0010] The method for identifying the lane-changing behavior of the leading vehicle provided by the embodiment of the present application obtains the operation data of the target leading vehicle corresponding to the host vehicle, and obtains multiple frames of target leading vehicle driving images corresponding to the target leading vehicle, thereby providing rich and comprehensive information for the target leading vehicle lane-changing behavior recognition model, and greatly improving the accuracy of target lane-changing behavior recognition. Then, the operation data and each frame of leading vehicle driving image are input into the target leading vehicle lane-changing behavior recognition model, and a preliminary lane-changing classification result corresponding to the target leading vehicle is output. It is ensured that the output preliminary lane-changing classification result is accurate and the classification is detailed, covering three situations: changing lanes to the left, changing lanes to the right, and not changing lanes. This enables the driver of the host vehicle or the automatic driving system to know the lane-changing intention of the leading vehicle in advance and have enough time to make corresponding responses, such as adjusting the vehicle speed, maintaining a safe distance, or preparing to avoid, etc., effectively reducing the probability of traffic accidents such as rear-end collisions and scratches, and ensuring driving safety.
[0011] In an alternative embodiment, the target leading vehicle lane-changing behavior recognition model includes a first target feature extraction network and a second target feature extraction network; inputting the operation data and each frame of leading vehicle driving image into the target leading vehicle lane-changing behavior recognition model and outputting a preliminary lane-changing classification result corresponding to the target leading vehicle includes:
[0012] Input the operation data into the first target feature extraction network in the target leading vehicle lane-changing behavior recognition model, and output the operation timing feature corresponding to the current operation state of the target leading vehicle;
[0013] Input each frame of target leading vehicle driving image into the second target feature extraction network in the target leading vehicle lane-changing behavior recognition model, and output the position image feature corresponding to the target leading vehicle;
[0014] Based on the operation timing feature and the position image feature, output a preliminary lane-changing classification result corresponding to the target leading vehicle.
[0015] The method for identifying the lane-changing behavior of the leading vehicle provided by the embodiment of the present application inputs the operation data into the first target feature extraction network in the target leading vehicle lane-changing behavior recognition model, and outputs the operation time-series features corresponding to the current operation state of the target leading vehicle, ensuring the accuracy of the output operation time-series features. Input each frame of the driving image of the target leading vehicle into the second target feature extraction network in the target leading vehicle lane-changing behavior recognition model, and output the position image features corresponding to the target leading vehicle, ensuring the accuracy of the output position image features corresponding to the target leading vehicle. It provides richer and more detailed information for lane-changing behavior recognition. Then, based on the operation time-series features and the position image features, output the preliminary lane-changing classification result corresponding to the target leading vehicle, ensuring the accuracy of the output preliminary lane-changing classification result corresponding to the target leading vehicle. In addition, during actual driving, there will be various interference factors, such as changes in lighting and weather effects, which may cause the information of the driving image of the target leading vehicle to be distorted or the operation data to have noise. Relying solely on a certain type of data for lane-changing behavior recognition, the anti-interference ability of the target leading vehicle lane-changing behavior recognition model is weak. However, by separately extracting the operation time-series features and the position image features and making a comprehensive judgment based on both, when the image is affected by lighting, the operation data can still provide stable vehicle motion information; conversely, if the operation data is interfered by sensor failures, etc., the driving image of the target leading vehicle can assist in making a judgment. This multi-source feature fusion method effectively improves the anti-interference ability of the model, ensuring that the preliminary lane-changing classification result can be stably and accurately output in a complex environment.
[0016] In an optional implementation manner, the first target feature extraction network includes a first sub-feature extraction network and a second sub-feature extraction network. Inputting the operation data into the first target feature extraction network in the target leading vehicle lane-changing behavior recognition model and outputting the operation time-series features corresponding to the current operation state of the target leading vehicle includes:
[0017] Input the operation data into the first sub-feature extraction network in the first target feature extraction network, and output the global time-series features corresponding to the operation data;
[0018] Input the operation data into the second sub-feature extraction network in the first target feature extraction network, and output the local time-series features corresponding to the operation data;
[0019] Fuse the global time-series features and the local time-series features, and output the operation time-series features corresponding to the target leading vehicle.
[0020] The method for identifying the lane-changing behavior of the leading vehicle provided by the embodiment of the present application inputs the operation data into the first sub-feature extraction network in the first target feature extraction network, and outputs the global time-series feature corresponding to the operation data. Thus, the time-series information in the operation data can be effectively captured. For example, the change trend of the relative speed of the target leading vehicle over a period of time, whether it is continuously accelerating, decelerating or maintaining a constant speed, is crucial for judging the lane-changing intention because the vehicle usually adjusts its speed before changing lanes. Then, the operation data is input into the second sub-feature extraction network in the first target feature extraction network, and the local time-series feature corresponding to the operation data is output, ensuring the accuracy of the output local time-series feature. Then, the global time-series feature and the local time-series feature are fused to output the operation time-series feature corresponding to the target leading vehicle. Thus, based on the global time-series feature and the local time-series feature, they work together to comprehensively and accurately extract the key information in the operation data, making the subsequent lane-changing judgment based on these features more accurate.
[0021] In an alternative embodiment, the second target feature includes a third sub-feature extraction network and a fourth sub-feature extraction network; inputting each frame of the driving image of the target leading vehicle into the second target feature extraction network in the target leading vehicle lane-changing behavior recognition model, and outputting the position image feature corresponding to the target leading vehicle, including:
[0022] Inputting each frame of the driving image of the target leading vehicle into the third sub-feature extraction network in the second target feature extraction network, and outputting the global image feature corresponding to the target leading vehicle;
[0023] Inputting each frame of the driving image of the target leading vehicle into the fourth sub-feature extraction network in the second target feature extraction network, and outputting the local image feature corresponding to the target leading vehicle;
[0024] Performing a fusion process on the global image feature and the local image feature to generate the position image feature corresponding to the target leading vehicle.
[0025] The method for identifying the lane-changing behavior of the leading vehicle provided by the embodiment of the present application inputs each frame of the driving image of the target leading vehicle into the third sub-feature extraction network in the second target feature extraction network, and outputs the global image features corresponding to the target leading vehicle, ensuring the accuracy of the output global image features, so as to identify the position of the target leading vehicle in the entire scene, the relative relationship with the surrounding lane lines and vehicles, etc., which is of great significance for judging the lane-changing direction and possibility. Then, each frame of the driving image of the target leading vehicle is input into the fourth sub-feature extraction network in the second target feature extraction network, and the local image features corresponding to the target leading vehicle are output, ensuring the accuracy of the output local image features. Thus, local features such as whether the turn signal of the target leading vehicle is on and whether the body posture changes to one side can be identified. Then, the global image features and the local image features are fused to generate the position image features corresponding to the target leading vehicle, providing richer and more detailed information for lane-changing behavior recognition, and greatly improving the accuracy of the preliminary lane-changing classification result.
[0026] In an alternative embodiment, based on the running time series features and the position image features, the preliminary lane-changing classification result corresponding to the target leading vehicle is output, including:
[0027] Use the first linear transformation matrix to map the running time series features to the first query space, the first key space, and the first value space;
[0028] Use the second linear transformation matrix to map the position image features to the second query space, the second key space, and the second value space;
[0029] Based on the first query space and the second key space, calculate the first attention score;
[0030] Based on the second query space and the first key space, calculate the second attention score;
[0031] Normalize the first attention score and the second attention score respectively to obtain the first normalized attention score and the second normalized attention score;
[0032] Multiply the first normalized attention score by the second value space to get the first product, and add the second product obtained by multiplying the second normalized attention score by the first value space to obtain the position feature encoding corresponding to the target leading vehicle;
[0033] Based on the position feature encoding, output the preliminary lane-changing classification result corresponding to the target leading vehicle.
[0034] The method for identifying the lane-changing behavior of the leading vehicle provided by the embodiment of the present application maps the running time-series features to the first query space, the first key space, and the first value space by using the first linear transformation matrix, ensuring the accuracy of the obtained first query space, first key space, and first value space. The position image features are mapped to the second query space, the second key space, and the second value space by using the second linear transformation matrix, ensuring the accuracy of the obtained second query space, second key space, and second value space. Then, based on the first query space and the second key space, the first attention score is calculated; based on the second query space and the first key space, the second attention score is calculated, ensuring the accuracy of the obtained first attention score and second attention score. Then, the first attention score and the second attention score are respectively normalized to obtain the first normalized attention score and the second normalized attention score; the first product obtained by multiplying the first normalized attention score by the second value space is added to the second product obtained by multiplying the second normalized attention score by the first value space to obtain the position feature encoding corresponding to the target leading vehicle, ensuring the accuracy of the obtained position feature encoding, enabling the position feature encoding to more comprehensively describe the state of the target leading vehicle, and providing a richer feature representation for subsequent lane-changing classification. Based on the position feature encoding, the preliminary lane-changing classification result corresponding to the target leading vehicle is output, ensuring the accuracy of the obtained preliminary lane-changing classification result. In the above method, by separately calculating the attention scores of the running time-series features and the position image features and normalizing them, the lane-changing behavior recognition model of the target leading vehicle can dynamically assign weights to them according to the importance of different features in the lane-changing behavior judgment, ensuring that the lane-changing behavior recognition model of the target leading vehicle can fully utilize the key information in different types of features and improve the adaptability to complex traffic scenarios.
[0035] In an alternative embodiment, after inputting the running data and the driving images of the leading vehicle in each frame into the lane-changing behavior recognition model of the target leading vehicle and outputting the preliminary lane-changing classification result corresponding to the target leading vehicle, the method further includes:
[0036] Obtain the driving environment information corresponding to the host vehicle, where the driving environment information includes the running information of the surrounding vehicles corresponding to the host vehicle and the obstacle information;
[0037] Convert the driving environment information into target text information;
[0038] Input the target text information into the target large language model to output the candidate lane-changing classification result corresponding to the target leading vehicle;
[0039] Combine the candidate lane-changing classification result with the preliminary lane-changing classification result to output the target lane-changing classification result corresponding to the target leading vehicle.
[0040] The method for identifying the lane-changing behavior of the leading vehicle provided by the embodiment of the present application obtains the driving environment information corresponding to the vehicle itself, converts the driving environment information into target text information, so that the target large language model can perform human-like reasoning based on the traffic domain common sense, driving rules, and semantic logic it has learned. Then, the target text information is input into the target large language model, and the candidate lane-changing classification results corresponding to the target leading vehicle are output, ensuring the accuracy of the output candidate lane-changing classification results. Then, the candidate lane-changing classification results are combined with the preliminary lane-changing classification results to output the target lane-changing classification results corresponding to the target leading vehicle, ensuring the accuracy of the output target lane-changing classification results. The above method considers both the running data and video data corresponding to the target leading vehicle, and also introduces the expert opinions of the large language model, ensuring the accuracy of the output target lane-changing classification results. Furthermore, it enables the driver of the vehicle itself or the autonomous driving system to know in advance the lane-changing intention of the leading vehicle. In addition, when facing driving scenarios that do not appear in the ideal dataset, the large language model can dynamically update its knowledge base and reasoning ability by continuously receiving new driving environment information and feedback data, and will also give certain expert opinions, enabling the vehicle itself to have enough time to make corresponding responses, such as adjusting the vehicle speed, maintaining a safe distance, or preparing to avoid, effectively reducing the probability of traffic accidents such as rear-end collisions and scratches, and ensuring driving safety.
[0041] In an alternative embodiment, inputting the target text information into the target large language model and outputting the candidate lane-changing classification results corresponding to the target leading vehicle includes:
[0042] Input the target text information into the target large language model;
[0043] The target large language model extracts features from the target text information to obtain key elements corresponding to the target text information; the key elements include at least one of the number of surrounding vehicles, the speed of each surrounding vehicle, the driving direction of each surrounding vehicle, the distance between each surrounding vehicle and the vehicle itself, and the type, position, and size of obstacles;
[0044] Generate the first prompt information based on the key elements; among them, traffic rules and driving behavior patterns are introduced into the target large language model;
[0045] Input the first prompt information into the target large language model to obtain the first-round output result;
[0046] Analyze the first-round output result, extract the doubtful or unclear parts in the first-round output result, and reconstruct the second prompt information again;
[0047] Input the second prompt information into the target large language model to obtain the second-round output result; the second-round output result includes sub-candidate lane-changing classification results;
[0048] Evaluate the sub-candidate lane change classification results, and generate the candidate lane change classification results based on the evaluation results.
[0049] In the method for identifying the leading vehicle's lane change behavior provided by the embodiments of the present application, the target text information is input into the target large language model; the target large language model extracts features from the target text information to obtain the key elements corresponding to the target text information, ensuring the accuracy of the key elements corresponding to the obtained target text information. In addition, traffic rules and driving behavior patterns are introduced into the target large language model to make the process of extracting key elements more intelligent and accurate. Then, based on the key elements, the first prompt information is generated, ensuring the accuracy of the generated first prompt information. The first prompt information is input into the target large language model to obtain the first round of output results; the first round of output results is analyzed, and the doubtful or unclear parts in the first round of output results are extracted to construct the second prompt information again; the second prompt information is input into the target large language model to obtain the second round of output results; the second round of output results includes sub-candidate lane change classification results. Evaluate the sub-candidate lane change classification results, and generate the candidate lane change classification results based on the evaluation results. Through the above method, the sub-candidate lane change classification results can be refined through multiple rounds of reasoning processes, from simple lane change or no lane change classification to considering conditions, timing, risks, etc. of lane change. This step-by-step refinement process helps to comprehensively consider various factors, such as the speed difference between different vehicles, the distance from obstacles, etc., to conduct a more detailed evaluation of different lane change situations, making the final candidate lane change classification results more comprehensive and accurate, and providing richer references for the final lane change decision. In addition, evaluating the sub-candidate lane change classification results can screen and optimize the final candidate lane change classification results. Through evaluation, contradictions and unreasonable parts in the sub-candidate results can be found, or weight assignment can be performed on different results. For example, for the contradictory parts in multiple sub-candidate results, selection can be made according to certain rules (such as being more in line with traffic rules or more in line with most situations) through evaluation, avoiding the one-sidedness of single judgment, and improving the reliability of the final candidate lane change classification results. This helps to avoid result deviations caused by the uncertainty of the large language model itself and improve the stability and credibility of the output results.
[0050] In an alternative embodiment, the candidate lane change classification result is combined with the preliminary lane change classification result to output the target lane change classification result corresponding to the target leading vehicle, including:
[0051] Quantify the candidate lane change classification result and the preliminary lane change classification result respectively to obtain the candidate quantified lane change classification result and the preliminary quantified lane change classification result;
[0052] Obtain the weight information corresponding to the candidate quantified lane change classification result and the preliminary quantified lane change classification result;
[0053] Multiply the candidate quantized lane change classification result and the preliminary quantized lane change classification result by their corresponding weight information respectively to obtain the target quantized lane change classification result;
[0054] Based on the target quantized lane change classification result, determine the target lane change classification result corresponding to the target leading vehicle.
[0055] In the method for identifying the lane change behavior of the leading vehicle provided by the embodiments of the present application, the candidate lane change classification result and the preliminary lane change classification result are respectively quantized to obtain the candidate quantized lane change classification result and the preliminary quantized lane change classification result, so that the classification results that may originally be presented in different forms (such as text form or different category representations) are unified into quantifiable values, which is convenient for subsequent calculations and fusions. This quantization method helps to comprehensively consider information from different sources and avoid the limitations of a single classification result. Then, obtain the weight information corresponding to the candidate quantized lane change classification result and the preliminary quantized lane change classification result, so that the importance of the two in the final decision can be flexibly adjusted according to different scenarios and requirements. Multiply the candidate quantized lane change classification result and the preliminary quantized lane change classification result by their corresponding weight information respectively to obtain the target quantized lane change classification result, ensuring the accuracy of the obtained target quantized lane change classification result. Then, based on the target quantized lane change classification result, determine the target lane change classification result corresponding to the target leading vehicle, ensuring the accuracy of the determined target lane change classification result.
[0056] In a second aspect, the present invention provides a device for identifying the lane change behavior of a leading vehicle, the device comprising:
[0057] A first acquisition module, configured to acquire the running data of the target leading vehicle corresponding to the host vehicle; the running data includes at least one of the relative speed of the target leading vehicle relative to the host vehicle, the relative distance of the target leading vehicle relative to the host vehicle, the lateral displacement of the target leading vehicle, and the lateral acceleration of the target leading vehicle;
[0058] A second acquisition module, configured to acquire multiple frames of target leading vehicle driving images corresponding to the target leading vehicle; the target leading vehicle driving images include the target leading vehicle;
[0059] An identification module, configured to input the running data and each frame of leading vehicle driving image into the target leading vehicle lane change behavior recognition model, and output a preliminary lane change classification result corresponding to the target leading vehicle, the preliminary lane change classification result including changing lanes to the left, changing lanes to the right, and not changing lanes.
[0060] The front vehicle lane change behavior recognition device provided by the embodiment of the present application obtains the operation data of the target front vehicle corresponding to the own vehicle, and obtains multiple frames of target front vehicle driving images corresponding to the target front vehicle, thereby providing rich and comprehensive information for the target front vehicle lane change behavior recognition model, and greatly improving the accuracy of target lane change behavior recognition. Then, the operation data and each frame of front vehicle driving image are input into the target front vehicle lane change behavior recognition model, and a preliminary lane change classification result corresponding to the target front vehicle is output. It is ensured that the output preliminary lane change classification result is accurate and the classification is detailed, covering three situations: changing lanes to the left, changing lanes to the right, and not changing lanes. This enables the driver of the own vehicle or the automatic driving system to know the lane change intention of the front vehicle in advance and have enough time to make corresponding responses, such as adjusting the vehicle speed, maintaining a safe distance, or preparing to avoid, etc., effectively reducing the occurrence probability of traffic accidents such as rear-end collisions and scratches, and ensuring driving safety.
[0061] In a third aspect, the present invention provides an electronic device, including: a memory and a processor, which are communicatively connected to each other. The memory stores computer instructions, and the processor executes the computer instructions to execute the front vehicle lane change behavior recognition method according to the first aspect or any corresponding embodiment thereof. BRIEF DESCRIPTION OF THE DRAWINGS
[0062] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for the description of the specific embodiments or the prior art. Obviously, the following drawings are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0063] Figure 1 It is a flowchart of the front vehicle lane change behavior recognition method according to the embodiment of the present invention;
[0064] Figure 2 It is a flowchart of another front vehicle lane change behavior recognition method according to the embodiment of the present invention;
[0065] Figure 3 It is a flowchart of yet another front vehicle lane change behavior recognition method according to the embodiment of the present invention;
[0066] Figure 4 It is a structural block diagram of the front vehicle lane change behavior recognition device according to the embodiment of the present invention;
[0067] Figure 5 It is a structural block diagram of another front vehicle lane change behavior recognition device according to the embodiment of the present invention;
[0068] Figure 6 It is a schematic hardware structure diagram of the electronic device according to the embodiment of the present invention. Detailed implementation manners
[0069] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0070] It should be noted that for the method for identifying the lane-changing behavior of the leading vehicle provided in the embodiments of the present application, the execution subject may be a device for identifying the lane-changing behavior of the leading vehicle. The device for identifying the lane-changing behavior of the leading vehicle may be implemented as part or all of an electronic device through software, hardware, or a combination of software and hardware. Among them, the electronic device may be a control device in the vehicle of the host vehicle. In the following method embodiments, the execution subject is taken as an electronic device as an example for description.
[0071] According to an embodiment of the present invention, an embodiment of a method for identifying the lane-changing behavior of a leading vehicle is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings may be executed in a computer system such as a set of computer-executable instructions. And although the logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than here.
[0072] In this embodiment, a method for identifying the lane-changing behavior of a leading vehicle is provided, which can be used for the above-mentioned electronic device. Figure 1 is a flowchart of the method for identifying the lane-changing behavior of a leading vehicle according to an embodiment of the present invention, as Figure 1 shown, the process includes the following steps:
[0073] Step S101, obtain the running data of the target leading vehicle corresponding to the host vehicle.
[0074] Among them, the running data includes at least one of the relative speed of the target leading vehicle relative to the host vehicle, the relative distance of the target leading vehicle relative to the host vehicle, the lateral displacement of the target leading vehicle, and the lateral acceleration of the target leading vehicle.
[0075] Specifically, the host vehicle may be equipped with various sensors, such as millimeter-wave radar sensors, lidar sensors, cameras, speed sensors, acceleration sensors, gyroscopes, GPS positioning systems, etc. The electronic device may collect the running data of the target leading vehicle corresponding to the host vehicle based on the various sensors equipped on the host vehicle.
[0076] Exemplarily, the electronic device may obtain the relative speed of the target vehicle ahead relative to the host vehicle based on a millimeter-wave radar sensor. The electronic device may obtain the relative distance of the target vehicle ahead relative to the host vehicle based on lidar. The electronic device may also calculate the lateral displacement and lateral acceleration of the target vehicle ahead according to the obtained relative speed of the target vehicle ahead relative to the host vehicle, the relative distance of the target vehicle ahead relative to the host vehicle, as well as the speed and displacement corresponding to the host vehicle.
[0077] Step S102, obtain multiple frames of driving images of the target vehicle ahead.
[0078] Among them, the driving images of the target vehicle ahead include the target vehicle ahead.
[0079] Specifically, the electronic device may also collect multiple frames of driving images of the target vehicle ahead based on a camera equipped on the host vehicle.
[0080] Step S103, input the operation data and each frame of driving image of the vehicle ahead into the lane change behavior recognition model of the target vehicle ahead, and output a preliminary lane change classification result corresponding to the target vehicle ahead.
[0081] Among them, the preliminary lane change classification result includes changing lanes to the left, changing lanes to the right, and not changing lanes.
[0082] Specifically, the electronic device may input the operation data and each frame of driving image of the vehicle ahead into the lane change behavior recognition model of the target vehicle ahead. The lane change behavior recognition model of the target vehicle ahead respectively recognizes the operation data and each frame of driving image of the vehicle ahead, and outputs a preliminary lane change classification result corresponding to the target vehicle ahead.
[0083] This step will be introduced in detail below.
[0084] The lane change behavior recognition method for the vehicle ahead provided in the embodiments of the present application obtains the operation data of the target vehicle ahead corresponding to the host vehicle, and obtains multiple frames of driving images of the target vehicle ahead corresponding to the target vehicle ahead, thereby providing rich and comprehensive information for the lane change behavior recognition model of the target vehicle ahead, and greatly improving the accuracy of target lane change behavior recognition. Then, the operation data and each frame of driving image of the vehicle ahead are input into the lane change behavior recognition model of the target vehicle ahead, and a preliminary lane change classification result corresponding to the target vehicle ahead is output. It ensures that the output preliminary lane change classification result is accurate and the classification is detailed, covering three situations: changing lanes to the left, changing lanes to the right, and not changing lanes. This enables the driver of the host vehicle or the automatic driving system to know the lane change intention of the vehicle ahead in advance, and has enough time to make corresponding responses, such as adjusting the vehicle speed, maintaining a safe distance, or preparing to avoid, etc., effectively reducing the probability of traffic accidents such as rear-end collisions and scratches, and ensuring driving safety.
[0085] In this embodiment, a lane change behavior recognition method for the vehicle ahead is provided, which can be used for the above-mentioned electronic device. Figure 2is a flowchart of a method for identifying the lane - changing behavior of a leading vehicle according to an embodiment of the present invention. As Figure 2 shown, the process includes the following steps:
[0086] Step S201, obtain the running data of the target leading vehicle corresponding to the host vehicle.
[0087] Among them, the running data includes at least one of the relative speed of the target leading vehicle relative to the host vehicle, the relative distance of the target leading vehicle relative to the host vehicle, the lateral displacement of the target leading vehicle, and the lateral acceleration of the target leading vehicle.
[0088] For this step, please refer to the above introduction to step S101, and details will not be repeated here.
[0089] Step S202, obtain multiple frames of target leading - vehicle driving images corresponding to the target leading vehicle.
[0090] Among them, the target leading - vehicle driving image includes the target leading vehicle.
[0091] For this step, please refer to the above introduction to step S102, and details will not be repeated here.
[0092] Step S203, input the running data and each frame of leading - vehicle driving images into the target leading - vehicle lane - changing behavior recognition model, and output a preliminary lane - changing classification result corresponding to the target leading vehicle.
[0093] Among them, the preliminary lane - changing classification result includes changing lanes to the left, changing lanes to the right, and not changing lanes.
[0094] Specifically, the target leading - vehicle lane - changing behavior recognition model includes a first target feature extraction network and a second target feature extraction network; the above step S203 may include the following steps:
[0095] Step S2031, input the running data into the first target feature extraction network in the target leading - vehicle lane - changing behavior recognition model, and output the running - time series features corresponding to the current running state of the target leading vehicle.
[0096] Specifically, the first target feature extraction network includes a first sub - feature extraction network and a second sub - feature extraction network. The above step S2031 may include the following steps:
[0097] Step a1, input the running data into the first sub - feature extraction network in the first target feature extraction network, and output the global time - series features corresponding to the running data.
[0098] Specifically, before the electronic device inputs the operation data into the first sub-feature extraction network in the first target feature extraction network, the electronic device can perform normalization processing on the operation data. Thus, operation data of different dimensions and magnitudes (such as relative speed, relative distance, lateral displacement, lateral acceleration, etc.) can be mapped to a unified scale range, such as the interval [0, 1] or [-1, 1]. Taking relative speed and lateral acceleration as examples, their numerical ranges and units vary greatly. Through normalization processing, the model training deviation caused by data magnitude differences can be avoided, enabling the network to treat each feature more fairly, and improving the training efficiency and the accuracy of feature extraction.
[0099] Then, the electronic device inputs the operation data after normalization processing into the first sub-feature extraction network in the first target feature extraction network, and outputs the global temporal features corresponding to the operation data.
[0100] Exemplarily, the first sub-feature extraction network can be an LSTM neural network. The specific process is as follows:
[0101] The electronic device inputs the operation data after normalization processing into the LSTM neural network in chronological order. At each time step, the LSTM neural network can calculate the attention weights between the current operation data and the hidden states of all previous time steps. The attention weights reflect the degree of association between the current operation data and the operation data at each historical moment. For example, when the target leading vehicle is about to change lanes, the changes in its lateral acceleration and relative speed are crucial for lane change judgment in a short period of time. The attention mechanism will automatically increase the attention to the data at these key time steps, making the hidden state more focused on these important features.
[0102] The input gate in the LSTM neural network determines the target operation data that needs to be added to the cell state from the current operation data according to the attention weights between the current operation data and the hidden states of all previous time steps. For example, when it is detected that the relative speed of the target leading vehicle suddenly changes, the input gate will increase the intake of the relative speed operation data of the target leading vehicle and integrate it into the cell state so that the subsequent network can remember this important speed change feature.
[0103] The forget gate in the LSTM neural network determines the operation data to be discarded that needs to be discarded from the cell state from the current operation data according to the attention weights between the current operation data and the hidden states of all previous time steps. For example, in a traffic scenario, for some early operation data that no longer has an important impact on the current lane change judgment (such as the relative distance information a few minutes ago, while the vehicle position has changed significantly in the current scenario), the forget gate will appropriately reduce its weight in the cell state to avoid the interference of these outdated information on the current feature extraction.
[0104] The output gate in the LSTM neural network generates the hidden state of the current time step based on the cell state and the current input. This hidden state contains the comprehensive information of the running data at the current moment and the previous historical moments. As the time step progresses, the hidden state gradually accumulates the key information of the entire time series. Based on each hidden state, the global time series features corresponding to the running data are formed.
[0105] Step a2: input the operation data into the second sub-feature extraction network in the first target feature extraction network, and output the local temporal features corresponding to the operation data.
[0106] Specifically, the electronic device can arrange the operation data in chronological order to generate time series data, which includes the relative speed, relative distance, lateral displacement and lateral acceleration of the target front vehicle relative to the vehicle. In order to extract local time series features, the electronic device can use the sliding window technology to divide the time series data into multiple local time series segments. For example, the operation data is sampled by seconds, and we can set a sliding window of length n seconds (such as n = 5) and slide across the entire time series with a certain step size (such as 1 second). In this way, the electronic device can divide the time series data into multiple shorter local time series segments, each of which contains the operation data of the target front vehicle within n seconds. This helps to focus on short-term, local changes in vehicle behavior.
[0107] Then, the electronic device regards the local time series fragments in each sliding window as a one-dimensional sequence. In order to make full use of the different dimensions of the operating data, the electronic device can input the relative speed, relative distance, lateral displacement and lateral acceleration in the operating data as the corresponding local time series fragments as different channels into the second sub-feature extraction network, similar to the RGB channels of an image. In this way, the convolution kernel can process data of different dimensions at the same time, learn the relationship between them, and thus extract richer local time series features. For example, a convolution kernel can simultaneously capture the local changes in relative speed and the corresponding local changes in lateral displacement, dig out the intrinsic connection between them, and provide more valuable information for subsequent lane change behavior judgment.
[0108] In addition, for the one-dimensional second sub-feature extraction network, the electronic device can slide multiple convolution kernels of different sizes on local time series segments, and each convolution kernel is responsible for extracting local features of different scales. Convolution kernels of different sizes can capture local patterns of different lengths. For example, small-size convolution kernels (such as 3 or 5 time steps) can extract short-term local features, which can be sensitive to small changes in the speed or acceleration of the target leading vehicle within a few seconds and are suitable for capturing rapid operation behaviors; large-size convolution kernels (such as 7 or 9 time steps) can capture local features within a relatively long time and focus more on detecting some local trends or patterns, such as the trend of the target leading vehicle gradually changing the relative distance within a short period of time.
[0109] After the convolution operation, the second sub-feature extraction network can introduce non-linearity using an activation function (such as ReLU) to enhance the network's expressive ability. Subsequently, the second sub-feature extraction network can use a pooling layer (such as max pooling or average pooling) to downsample the convolved feature map and reduce the feature dimension. Max pooling can extract the most prominent local features, while average pooling can reflect the average level of local features.
[0110] After being processed by multiple convolutional layers, activation functions, and pooling layers, the finally output feature map contains local temporal features of the running data.
[0111] In an alternative embodiment, the second sub-feature extraction network can add or concatenate the feature maps of the previous layer and the feature maps of the subsequent layer, enabling the network to learn richer local features. At the same time, feature fusion techniques can be used to fuse the features extracted by different convolution kernels and different channels to form a comprehensive local temporal feature vector. For example, fuse the local features extracted by small convolution kernels with the local features extracted by large convolution kernels, or fuse the local features of different data dimension channels, so that the final local temporal features include both subtle short-term behavior changes and relatively long local trend information. To make the output local temporal features more general, they can be normalized, mapping the elements of the feature vector to a certain range (such as [0,1] or [-1,1]) for fusing with other features or inputting into subsequent networks. Feature encoding techniques such as principal component analysis (PCA) or autoencoders can also be used to reduce the dimension or reconstruct the local temporal features to reduce redundant information and highlight the most representative local features.
[0112] Step a3: Fuse the global temporal features and the local temporal features, and output the running temporal features corresponding to the target leading vehicle.
[0113] Specifically, the electronic device can use a fully connected layer or a convolutional layer to expand features with lower dimensions or reduce the dimensions of features with higher dimensions, so that the global temporal features and the local temporal features have matching dimensions. Then, the electronic device performs normalization processing on the adjusted global temporal features and local temporal features respectively to ensure their comparability in the numerical range and avoid a certain feature being ignored or overemphasized due to magnitude differences during the fusion process.
[0114] Next, the normalized global temporal features and local temporal features are concatenated together along the feature dimension to output the running temporal features corresponding to the target leading vehicle.
[0115] Step S2032: Input each frame of the driving image of the target leading vehicle into the second target feature extraction network in the target leading vehicle lane change behavior recognition model to output the position image features corresponding to the target leading vehicle.
[0116] Specifically, the second target feature includes a third sub-feature extraction network and a fourth sub-feature extraction network. The above step S2032 may include the following steps:
[0117] Step b1: Input each frame of the driving image of the target leading vehicle into the third sub-feature extraction network in the second target feature extraction network to output the global image features corresponding to the target leading vehicle.
[0118] Specifically, before inputting each frame of the driving image of the target leading vehicle into the third sub-feature extraction network in the second target feature extraction network, the electronic device can perform normalization processing and cropping processing on each frame of the driving image of the target leading vehicle to obtain the processed driving images of the target leading vehicle for each frame.
[0119] Then, the electronic device can divide the processed driving images of the target leading vehicle for each frame into multiple sub-image blocks of a fixed size. For example, a 224×224 pixel driving image of the target leading vehicle is divided into 14×14 sub-image blocks of 16×16 pixels. Each sub-image block is linearly mapped into a sub-vector, and these sub-vectors serve as the input to the third sub-feature extraction network. This way converts the image into sequence data for easy processing.
[0120] Multiple parallel attention heads in the third sub-feature extraction network capture the relationships between the sub-image blocks from different subspaces. Specifically, the third sub-feature extraction network can generate a dynamic mask according to the position and size of the target leading vehicle in the processed driving images of the target leading vehicle for each frame. For regions far from the target leading vehicle, the mask reduces its attention weight and decreases the attention to irrelevant background information.
[0121] Specifically, the third sub-feature extraction network can first determine the bounding box of the target leading vehicle in the image through an object detection algorithm, and then generate a mask matrix based on the bounding box. Then, for each parallel attention head, query (Query), key (Key), and value (Value) matrices are calculated, and then the attention weight matrix is calculated based on the similarity between the query and the key. Then, the mask matrix is multiplied by the attention weight matrix to obtain the target attention weight, and the target attention weight is multiplied by the corresponding value matrix to obtain the sub-vector corresponding to the sub-image block.
[0122] Then, the third sub-feature extraction network adds position encoding to each sub-image block. Then, the position encoding vector is added to the sub-vector corresponding to the sub-image block to obtain the target vector corresponding to each sub-image block.
[0123] Finally, the third sub-feature extraction network fuses the target vectors corresponding to each sub-image block to generate the global image feature corresponding to the target leading vehicle.
[0124] Step b2, input each frame of the target leading vehicle driving image into the fourth sub-feature extraction network in the second target feature extraction network, and output the local image feature corresponding to the target leading vehicle.
[0125] Specifically, before inputting each frame of the target leading vehicle driving image into the fourth sub-feature extraction network, the electronic device can perform enhancement operations on each frame of the target leading vehicle driving image to generate the enhanced target leading vehicle driving image, so as to improve the adaptability of the network to different lighting, weather, and scenarios. Common enhancement techniques include random brightness adjustment, contrast adjustment, adding noise, etc. For example, to randomly adjust the brightness of each frame of the target leading vehicle driving image, the following formula can be used:
[0126] I enhanced = I + α × rand(-1, 1)
[0127] where I is the target leading vehicle driving image, α is an adjustable brightness adjustment parameter, and rand(-1, 1) generates a random number between -1 and 1. This can simulate images under different lighting conditions and enable the network to learn more robust features.
[0128] Then, the electronic device can adjust the size of the enhanced target leading vehicle driving image according to the input requirements of the fourth sub-feature extraction network. If the network requires an input size of 256×256 pixels and the size of the enhanced target leading vehicle driving image is different, methods such as bilinear interpolation can be used to scale the image to this size. At the same time, ensure that during the size adjustment process, the proportion and position in the enhancement are relatively reasonable to avoid information distortion caused by excessive stretching or compression.
[0129] Next, the electronic device inputs the target front vehicle driving image after size adjustment into the fourth sub-feature extraction network. Among them, the fourth sub-feature extraction network can be a CNN with multiple convolutional layers, pooling layers, and activation functions. Taking the classic VGG architecture as an example, the network contains multiple convolutional layers, and each convolutional layer uses convolutional kernels of different sizes, such as 3x3 or 5x5 convolutional kernels. The convolutional kernel slides on the image to extract local features.
[0130] For the target front vehicle driving image after size adjustment, the first convolutional layer can extract low-level features of the target front vehicle driving image after size adjustment, such as edges, textures, etc. For example, performing a convolution operation using a 3x3 convolutional kernel, the parameters of the convolutional kernel are learned through network training, and it will automatically capture edge information in different directions in the image, and extract the contour of the target front vehicle and the edge features of the lane lines.
[0131] After the convolution operation, the ReLU activation function is used to introduce non-linearity and enhance the expression ability of the network. The formula is: f(x) = max(0, x), where x is the output of the convolutional layer, setting negative values to zero and retaining positive values to simulate the activation state of neurons.
[0132] Then, a max-pooling or average-pooling layer is used to reduce the size of the feature map while retaining the main features. Max-pooling takes the maximum value of each region in the feature map, which helps to highlight the most significant features in the local region. For example, in a 2x2 max-pooling layer, the maximum value of the pixel values within the 2x2 region is taken as the output, reducing the amount of data while retaining the most significant local information. This helps the network to focus on the most representative features in the local region, and plays an important role in extracting local details of the target front vehicle, such as local features of vehicle signs, headlights, windows, etc., because these local features may be key information for lane-changing intentions.
[0133] Finally, after being processed by multiple convolutional layers, pooling layers, and possibly attention modules, the finally obtained feature map contains the local image features of the target front vehicle. These features reflect the local details of the target front vehicle, such as local features of vehicle components and local relationships with the surrounding environment. To obtain a fixed-length feature vector, a fully connected layer can be used to flatten the final feature map and map it to the required feature dimension. For example, the size of the last feature map is, and it is mapped to a vector with a length of 1024 as the local image feature through a fully connected layer.
[0134] Step b3, fuse the global image feature and the local image feature to generate the position image feature corresponding to the target front vehicle.
[0135] Specifically, the electronic device can splice the global image feature and the local image feature to generate the position image feature corresponding to the target front vehicle.
[0136] Step S2033: Output a preliminary lane-changing classification result corresponding to the target leading vehicle based on the running timing features and the position image features.
[0137] Specifically, the above-mentioned step S2033 may include the following steps:
[0138] Step c1: Map the running timing features to a first query space, a first key space, and a first value space by using a first linear transformation matrix.
[0139] Specifically, the electronic device may map the running timing features to a first query space, a first key space, and a first value space by using a first linear transformation matrix.
[0140] Exemplarily, is the first linear transformation matrix for the linear transformation of the running timing features, and the specific transformation formula is as follows:
[0141]
[0142] where X r is the running timing feature, Q r is the first query space, K r is the first key space, and V r is the first value space.
[0143] Step c2: Map the position image features to a second query space, a second key space, and a second value space by using a second linear transformation matrix.
[0144] Specifically, the electronic device may map the position image features to a second query space, a second key space, and a second value space by using a second linear transformation matrix.
[0145] Exemplarily, is the second linear transformation matrix for the linear transformation of the position image features, and the specific transformation formula is as follows:
[0146]
[0147] where X i is the position image feature, Q i is the second query space, K i is the second key space, and V i is the second value space.
[0148] Step c3: Calculate a first attention score based on the first query space and the second key space.
[0149] Specifically, the electronic device calculates a first attention score based on the first query space and the second key space.
[0150] Exemplarily, the electronic device may calculate a first attention score based on the following formula:
[0151]
[0152] where Score ri is the first attention score, Q r is the first query space, K i is the second key space, d k is the dimension of the second key space.
[0153] Step c4: Calculate a second attention score based on the second query space and the first key space.
[0154] Specifically, the electronic device may calculate a second attention score based on the second query space and the first key space.
[0155] Exemplarily, the electronic device may calculate the second attention score based on the following formula:
[0156]
[0157] where Score ir is the second attention score, Q i is the second query space, K r is the first key space, d k is the dimension of the first key space.
[0158] Step c5: Normalize the first attention score and the second attention score respectively to obtain a first normalized attention score and a second normalized attention score.
[0159] Specifically, the electronic device normalizes the first attention score and the second attention score respectively to obtain a first normalized attention score and a second normalized attention score.
[0160] Exemplarily, the electronic device may apply the Softmax function to normalize the first attention score and the second attention score. The specific formula is as follows:
[0161] α ri = Softmax(Score ri );
[0162] α ir = Softmax(Score ir );
[0163] where α ri is the first normalized attention score, and α ir is the second normalized attention score.
[0164] Step c6: Add the first product obtained by multiplying the first normalized attention score by the second value space to the second product obtained by multiplying the second normalized attention score by the first value space to obtain the position feature encoding corresponding to the target leading vehicle.
[0165] Specifically, the electronic device adds the first product obtained by multiplying the first normalized attention score by the second value space to the second product obtained by multiplying the second normalized attention score by the first value space to obtain the position feature encoding corresponding to the target leading vehicle.
[0166] Exemplarily, the formula is as follows:
[0167] F ri = α ri V i ;
[0168] F ir = α ir V r ;
[0169] F = F ri + F ir ;
[0170] Wherein, V r is the first value space, and V i is the second value space.
[0171] Step c7: Based on the position feature encoding, output the preliminary lane change classification result corresponding to the target leading vehicle.
[0172] Specifically, the electronic device may input the position feature encoding into a preset classifier, and the preset classifier identifies the position feature encoding and outputs the preliminary lane change classification result corresponding to the target leading vehicle.
[0173] Wherein, the preset classifier may be a fully connected neural network (FCN) classifier, or a support vector machine (SVM), or a decision tree classifier. The embodiments of the present application do not make specific limitations on the preset classifier.
[0174] The method for identifying the lane-changing behavior of the leading vehicle provided by the embodiment of the present application inputs the running data into the first sub-feature extraction network in the first target feature extraction network, and outputs the global temporal features corresponding to the running data. Thus, the time series information in the running data can be effectively captured, such as the change trend of the relative speed of the target leading vehicle over a period of time, whether it is continuously accelerating, decelerating or moving at a constant speed, which is crucial for judging the lane-changing intention because the vehicle usually adjusts its speed before changing lanes. Then, the running data is input into the second sub-feature extraction network in the first target feature extraction network, and the local temporal features corresponding to the running data are output, ensuring the accuracy of the output local temporal features. Then, the global temporal features and the local temporal features are fused to output the running temporal features corresponding to the target leading vehicle. Thus, based on the global temporal features and the local temporal features, they work together to comprehensively and accurately extract the key information in the running data, making the subsequent lane-changing judgment based on these features more accurate.
[0175] Each frame of the driving image of the target leading vehicle is input into the third sub-feature extraction network in the second target feature extraction network, and the global image features corresponding to the target leading vehicle are output, ensuring the accuracy of the output global image features. Thus, the position of the target leading vehicle in the entire scene, the relative relationship with the surrounding lane lines and vehicles, etc. can be identified, which is of great significance for judging the lane-changing direction and possibility. Then, each frame of the driving image of the target leading vehicle is input into the fourth sub-feature extraction network in the second target feature extraction network, and the local image features corresponding to the target leading vehicle are output, ensuring the accuracy of the output local image features. Thus, local features such as whether the turn signal of the target leading vehicle is on and whether there is a change in the body posture towards one side can be identified. Then, the global image features and the local image features are fused to generate the position image features corresponding to the target leading vehicle, providing richer and more detailed information for lane-changing behavior recognition and greatly improving the accuracy of the preliminary lane-changing classification result.
[0176] Then, the running time series features are mapped to the first query space, the first key space, and the first value space by using the first linear transformation matrix, ensuring the accuracy of the obtained first query space, first key space, and first value space. The position image features are mapped to the second query space, the second key space, and the second value space by using the second linear transformation matrix, ensuring the accuracy of the obtained second query space, second key space, and second value space. Then, based on the first query space and the second key space, the first attention score is calculated; based on the second query space and the first key space, the second attention score is calculated, ensuring the accuracy of the obtained first attention score and second attention score. Then, the first attention score and the second attention score are respectively normalized to obtain the first normalized attention score and the second normalized attention score; the first product obtained by multiplying the first normalized attention score by the second value space is added to the second product obtained by multiplying the second normalized attention score by the first value space to obtain the position feature encoding corresponding to the target leading vehicle, ensuring the accuracy of the obtained position feature encoding, enabling the position feature encoding to more comprehensively describe the state of the target leading vehicle, and providing a richer feature representation for subsequent lane change classification. Based on the position feature encoding, the preliminary lane change classification result corresponding to the target leading vehicle is output, ensuring the accuracy of the obtained preliminary lane change classification result.
[0177] In the above method, by separately calculating the attention scores of the running time series features and the position image features and normalizing them, the target leading vehicle lane change behavior recognition model can dynamically assign weights to them according to the importance of different features in the lane change behavior judgment, ensuring that the target leading vehicle lane change behavior recognition model can make full use of the key information in different types of features and improve the adaptability to complex traffic scenarios. In addition, during actual driving, there will be various interference factors, such as changes in lighting and weather effects, which may cause the distortion of the driving image information of the target leading vehicle or noise in the running data. Relying solely on a certain type of data for lane change behavior recognition, the anti-interference ability of the target leading vehicle lane change behavior recognition model is weak. However, by separately extracting the running time series features and the position image features and comprehensively judging the two, when the image is affected by lighting, the running data can still provide stable vehicle motion information; conversely, if the running data is interfered by sensor failures, etc., the driving image of the target leading vehicle can assist in the judgment. This multi-source feature fusion method effectively improves the anti-interference ability of the model, ensuring that the preliminary lane change classification result can be stably and accurately output in a complex environment.
[0178] In this embodiment, a method for recognizing the lane change behavior of a leading vehicle is provided, which can be used for the above-mentioned electronic device. Figure 3 It is a flowchart of the method for recognizing the lane change behavior of a leading vehicle according to an embodiment of the present invention, as Figure 3 shown, and this process includes the following steps:
[0179] Step S301: Obtain the running data of the target leading vehicle corresponding to the host vehicle.
[0180] Among them, the running data includes at least one of the relative speed of the target leading vehicle relative to the host vehicle, the relative distance of the target leading vehicle relative to the host vehicle, the lateral displacement of the target leading vehicle, and the lateral acceleration of the target leading vehicle.
[0181] For this step, please refer to the introduction of step S201 above and will not be elaborated here.
[0182] Step S302: Obtain multiple frames of target leading vehicle driving images corresponding to the target leading vehicle.
[0183] Among them, the target leading vehicle driving image includes the target leading vehicle.
[0184] For this step, please refer to the introduction of step S202 above and will not be elaborated here.
[0185] Step S303: Input the running data and each frame of leading vehicle driving image into the target leading vehicle lane change behavior recognition model, and output the preliminary lane change classification result corresponding to the target leading vehicle.
[0186] Among them, the preliminary lane change classification result includes changing lanes to the left, changing lanes to the right, and not changing lanes.
[0187] For this step, please refer to the introduction of step S203 above and will not be elaborated here.
[0188] Step S304: Obtain the driving environment information corresponding to the host vehicle.
[0189] Among them, the driving environment information includes the surrounding vehicle running information and obstacle information corresponding to the host vehicle.
[0190] Specifically, the host vehicle can be equipped with various sensors, such as millimeter wave radar sensors, lidar sensors, cameras, speed sensors, acceleration sensors, gyroscopes, GPS positioning systems, etc. The electronic device can obtain the driving environment information corresponding to the host vehicle based on the various sensors equipped on the host vehicle.
[0191] Step S305: Convert the driving environment information into target text information.
[0192] Specifically, the electronic device can use image recognition algorithms to identify target objects such as vehicles, pedestrians, traffic signs, and lane lines in the driving environment information. For example, through deep learning models, such as object detection algorithms based on convolutional neural networks (CNNs) (such as YOLO, Faster R-CNN), the category, position, and size information of different objects can be detected from camera images. Suppose a car is detected ahead, and its position coordinates in the image are (x1, y1, x2, y2), where (x1, y1) and (x2, y2) are the upper left and lower right coordinates of the target object's bounding box, respectively.
[0193] In addition, the electronic device can also analyze the radar echo signals in the driving environment information to obtain the distance, speed, and angle information of the target object. For example, the radar data shows that there is a target object 30 meters away from the vehicle, with a relative speed of 10 m / s and an angle of 5 degrees to the left of the due front.
[0194] The electronic device can also analyze the point cloud data generated by lidar in the driving environment information to identify the three-dimensional shape and position of the object. Through point cloud segmentation algorithms, different objects in the point cloud data are segmented, and their positions and poses in three-dimensional space are determined. For example, the lidar data shows that there is a cuboid object with a height of 1.5 meters and a length of 4 meters, 20 meters in the front left of the vehicle, which is identified as a minibus.
[0195] Next, the electronic device can organize the extracted various types of information according to a certain logical structure to facilitate subsequent conversion into text. The electronic device can create a dictionary or data structure containing different information categories. During the structuring process, the electronic device can clarify the semantic associations between different pieces of information. For example, determine the relative position relationship between a certain target object and the vehicle itself, as well as its relationship with surrounding road features (such as lane lines, traffic signs). For instance, the detected car mentioned above is located directly in front of the vehicle and within the vehicle's driving lane; the traffic sign "Speed Limit 60" is located 100 meters ahead on the right side of the road. The electronic device designs corresponding text templates according to the categories and characteristics of the structured information. For example: For vehicle information: "The current speed of the vehicle is {self_vehicle.speed} km / h, the acceleration is {self_vehicle.acceleration} m / s 2 , and the steering angle is {self_vehicle.steering_angle} degrees."
[0196] For information about surrounding objects: "At a direction of {abs(surrounding_objects[0].angle)} degrees {surrounding_objects[0].angle>0?'to the right in front':'to the left in front'} of the vehicle, there is a {surrounding_objects[0].type} at a distance of {surrounding_objects[0].distance} meters, and its relative speed is {surrounding_objects[0].relative_speed} m / s".
[0197] For information about road features: "There are {road_features.lane_lines} visible on the road, and there is a traffic sign of '{road_features.traffic_signs[0].content}' on the right roadside {road_features.traffic_signs[0].distance} meters ahead".
[0198] Finally, the electronic device fills the specific numerical values and descriptions in the structured information into the corresponding text templates to generate the complete target text information. For example: "The current speed of the vehicle is 60 km / h, the acceleration is 1 m / s 2 , and the steering angle is 5 degrees. At a direction of 5 degrees to the left in front of the vehicle, there is a car at a distance of 30 meters, and its relative speed is 10 m / s. There are double-lane lane lines visible on the road, and there is a traffic sign of'speed limit 60' on the right roadside 100 meters ahead".
[0199] Step S306: Input the target text information into the target large language model to output the candidate lane change classification result corresponding to the target leading vehicle.
[0200] Specifically, the above step S306 may include the following steps:
[0201] Step S3061: Input the target text information into the target large language model.
[0202] Specifically, the electronic device may input the target text information into the target large language model.
[0203] Step S3062: The target large language model extracts features from the target text information to obtain the key elements corresponding to the target text information.
[0204] Among them, the key elements include at least one of the number of surrounding vehicles, the speed of each surrounding vehicle, the driving direction of each surrounding vehicle, the distance between each surrounding vehicle and the vehicle, and the type, position, and size of the obstacle.
[0205] Specifically, the target large language model can segment the target text information by word or sub-word. For example, for the target text information "There is a car 30 meters away in the direction 5 degrees to the left in front of this vehicle, and its relative speed is 10 m / s", natural language processing tools (such as NLTK, spaCy, etc.) can be used to tokenize it into ["In", "this vehicle", "front", "to the left", "5", "degrees", "direction", ",", "distance", "30", "meters", "away", ",", "there is", "a", "car", ",", "its", "relative speed", "is", "10", "m / s"].
[0206] Then, the target large language model tags the part of speech for each word, such as noun, verb, adjective, numeral, etc. For example, "car" is tagged as a noun, "10" is tagged as a numeral, and "m / s" is tagged as a unit noun. This helps to identify key elements subsequently, because words with different parts of speech play different roles in expressing key elements. Next, the target large language model removes common words that are not substantially helpful for extracting key elements, such as "In", "there is", "its", etc. After removing the stop words, the above text becomes ["this vehicle", "front", "to the left", "5", "degrees", "direction", "distance", "30", "meters", "away", "car", "relative speed", "10", "m / s"], making the text more concise and highlighting the key information.
[0207] Then, the electronic device extracts key elements from the text information after removing the stop words.
[0208] Exemplarily, look for words representing the number of vehicles in the text, such as "a", "two", "multiple", etc. For expressions containing specific numbers, the numbers can be directly extracted. For example, for the text "There are three cars in front", the number "3" can be directly extracted as the number of surrounding vehicles. Look for combinations of numerical values and units representing speed, such as "10 m / s", "60 km / h", etc. First, identify the numeral, and then check whether the subsequent word is a speed unit. In addition, pay attention to words related to speed, such as "speed", "rate", "fast or slow", etc., and use these words as clues to look for nearby speed numerical values and units. For example, for "The driving speed of the car is 60 km / h", through the word "speed", the subsequent speed numerical value "60" and unit "km / h" are located. Look for words representing directions, such as "front", "rear", "left", "right", "east", "south", "west", "north", etc., and some words representing relative directions, such as "to the left in front", "to the right in the rear", etc. For example, for the text "There is a car in the direction 5 degrees to the left in front of this vehicle", by matching "to the left in front", the driving direction of the vehicle is determined as "to the left in front", and the specific angle is determined in combination with "5 degrees".
[0209] Step S3063, generate the first prompt information based on the key elements.
[0210] Among them, knowledge bases such as traffic rules and driving behavior patterns are introduced into the target large language model.
[0211] Specifically, the target large language model can construct a basic prompt framework according to the task requirements and the characteristics of key elements. This framework will serve as the basis for generating specific prompt information to ensure the integrity and logic of the information.
[0212] Exemplarily, if the ultimate goal is to analyze the lane-changing behavior of the target vehicle in front, the prompt framework can be constructed around the impact of surrounding vehicles and obstacles on lane-changing. For example: "Analyze the possibility, direction, and timing of the target vehicle in front changing lanes based on the [quantity, speed, driving direction, distance from the vehicle] of surrounding vehicles and the [type, location, size] of obstacles. The information of surrounding vehicles is as follows: [detailed information of surrounding vehicles]. The information of obstacles is as follows: [detailed information of obstacles]." For a more general traffic scenario analysis, the prompt framework can be: "Combine the [list of key elements] of surrounding vehicles and the [list of key elements] of obstacles to analyze the safety, potential risks, and reasonable driving strategies of the current traffic scenario. Surrounding vehicles: [specific key elements of surrounding vehicles]. Obstacles: [specific key elements of obstacles]".
[0213] Then, the target large language model can fill in the previously extracted key elements according to the requirements of the prompt framework, so that the generated first prompt information has actual content.
[0214] Step S3064, input the first prompt information into the target large language model to obtain the first-round output result.
[0215] Specifically, after receiving the first prompt information, the target large language model encodes it and converts the text into a vector form that the target large language model can process. Then, through its internal Transformer architecture, it uses the multi-head attention mechanism to capture the relationships between elements in the text and understand the semantics of the first prompt information. After being processed by multiple layers of Transformer blocks, in-depth analysis and reasoning are performed on the first prompt information.
[0216] The target large language model generates the first-round output result based on its understanding and analysis of the first prompt information. This result may be a preliminary judgment on the possibility of the target vehicle in front changing lanes and a description of potential risks, such as "The target vehicle in front has a certain possibility of changing lanes, but attention should be paid to the construction obstacle in the left front. When changing lanes, the safety distance from the following vehicle may be affected due to avoiding the obstacle."
[0217] Step S3065, analyze the first-round output result, extract the doubtful or unclear parts in the first-round output result, and construct the second prompt information again.
[0218] Specifically, the target large language model analyzes the first round of output results to identify the parts of the first round of output results that are questionable or unclear. For example, the output mentions "there is a certain possibility of changing lanes", but the expression "certain" is vague and does not clarify the degree of possibility of changing lanes; or it mentions paying attention to construction obstacles, but does not specify how to pay attention and the specific forms of risks that may arise.
[0219] Then, the target large language model constructs a second prompt for these uncertain or ambiguous contents. For example, "Regarding the previously mentioned possibility that the target vehicle in front has a certain lane change, is this possibility high, medium or low? In addition, regarding the construction obstacle in front of the left, please explain in detail the specific risks that may be caused and how the target vehicle in front should avoid these risks." This second prompt is more specific and targeted, and is designed to guide the model to give a clearer and more accurate answer.
[0220] Step S3066: input the second prompt information into the target large language model to obtain a second round of output results.
[0221] Among them, the second round of output results includes the sub-candidate lane change classification results.
[0222] Specifically, after receiving the second prompt information, the target large language model repeats a similar processing flow to conduct a deeper understanding and analysis of the new second prompt information. Since the second prompt information is more targeted, the model can answer the previous questions more focusedly.
[0223] The second round of output results includes the classification results of the lane change sub-candidates, such as "The possibility of the target front vehicle changing lanes is medium. The specific risk is that when changing lanes, the vehicle may lose control due to being close to the construction obstacle, and the rear vehicle may rear-end due to lack of time to react. The target front vehicle should observe the speed and distance of the rear vehicle in advance, and change lanes slowly and maintain a safe distance from the obstacle while ensuring safety." These results provide a clearer classification of the possibility of lane changes and elaborate on the risks and countermeasures.
[0224] Step S3067: Evaluate the sub-candidate lane change classification results, and generate candidate lane change classification results based on the evaluation results.
[0225] Specifically, the electronic device can evaluate the sub-candidate lane-changing classification results output in the second round according to the evaluation criteria. If the sub-candidate results meet the evaluation criteria in all aspects, then it can be directly used as the candidate lane-changing classification result; if there are some situations that do not meet the criteria, it may need to be adjusted or further optimized, and finally an accurate and reliable candidate lane-changing classification result is generated, such as "The possibility of the target leading vehicle changing lanes is medium. It is necessary to carefully consider the lane-changing operation and closely monitor the construction obstacles and the dynamics of the vehicles behind." This result will provide an important basis for the subsequent combination with the preliminary lane-changing classification result and the final lane-changing decision.
[0226] Step S307: Combine the candidate lane-changing classification result with the preliminary lane-changing classification result, and output the target lane-changing classification result corresponding to the target leading vehicle.
[0227] Specifically, the above step S307 may include the following steps:
[0228] Step S3071: Quantify the candidate lane-changing classification result and the preliminary lane-changing classification result respectively to obtain the candidate quantified lane-changing classification result and the preliminary quantified lane-changing classification result.
[0229] Specifically, the electronic device can quantify the candidate lane-changing classification result and the preliminary lane-changing classification result respectively to obtain the candidate quantified lane-changing classification result and the preliminary quantified lane-changing classification result.
[0230] Step S3072: Obtain the weight information corresponding to the candidate quantified lane-changing classification result and the preliminary quantified lane-changing classification result.
[0231] Specifically, the electronic device can receive the weight information corresponding to the candidate quantified lane-changing classification result and the preliminary quantified lane-changing classification result input by the user. The electronic device can also receive the weight information corresponding to the candidate quantified lane-changing classification result and the preliminary quantified lane-changing classification result sent by other devices.
[0232] Step S3073: Multiply the candidate quantified lane-changing classification result and the preliminary quantified lane-changing classification result by the corresponding weight information respectively to obtain the target quantified lane-changing classification result.
[0233] Specifically, the electronic device can multiply the candidate quantified lane-changing classification result and the preliminary quantified lane-changing classification result by the corresponding weight information respectively to obtain the target quantified lane-changing classification result.
[0234] Step S3074: Based on the target quantified lane-changing classification result, determine the target lane-changing classification result corresponding to the target leading vehicle.
[0235] Specifically, the electronic device can determine the target lane-changing classification result corresponding to the target leading vehicle according to the corresponding relationship between the target quantified lane-changing classification result and the target lane-changing classification result.
[0236] The method for identifying the lane-changing behavior of the leading vehicle provided by the embodiment of the present application obtains the driving environment information corresponding to the host vehicle, converts the driving environment information into target text information, so that the target large language model can perform reasoning based on the traffic domain common sense, driving rules, and semantic logic it has learned. Then, the target text information is input into the target large language model; the target large language model extracts features from the target text information to obtain the key elements corresponding to the target text information, ensuring the accuracy of the key elements corresponding to the obtained target text information. In addition, knowledge bases such as traffic rules and driving behavior patterns are introduced into the target large language model to make the process of extracting key elements more intelligent and accurate. Then, based on the key elements, a first prompt message is generated, ensuring the accuracy of the generated first prompt message. The first prompt message is input into the target large language model to obtain the first-round output result; the first-round output result is analyzed, and the doubtful or unclear parts in the first-round output result are extracted to construct a second prompt message again; the second prompt message is input into the target large language model to obtain the second-round output result; the second-round output result includes sub-candidate lane-changing classification results. The sub-candidate lane-changing classification results are evaluated, and based on the evaluation results, candidate lane-changing classification results are generated. The above method can refine the sub-candidate lane-changing classification results through multiple rounds of reasoning processes, from a simple lane-changing or non-lane-changing classification to considering conditions, timing, risks, etc. of lane-changing. This gradual refinement process helps to more comprehensively consider various factors, such as the speed difference between different vehicles, the distance from obstacles, etc., and conduct a more detailed evaluation of different lane-changing situations, making the final candidate lane-changing classification results more comprehensive and accurate, and providing richer references for the final lane-changing decision. In addition, evaluating the sub-candidate lane-changing classification results can screen and optimize the final candidate lane-changing classification results. Through evaluation, contradictions and unreasonable parts in the sub-candidate results can be found, or weight allocation can be performed on different results. For example, for the contradictory parts in multiple sub-candidate results, selection can be made according to certain rules (such as being more in line with traffic rules or more in line with the majority of situations) through evaluation, avoiding the one-sidedness of single judgment and improving the reliability of the final candidate lane-changing classification results. This helps to avoid result deviations caused by the uncertainty of the large language model itself and improve the stability and credibility of the output results.
[0237] Then, the candidate lane change classification results and the preliminary lane change classification results are respectively quantified to obtain the candidate quantified lane change classification results and the preliminary quantified lane change classification results, so that the classification results that may originally be presented in different forms (such as text form or different category representations) are unified into quantifiable numerical values, facilitating subsequent calculations and fusions. This quantification method helps to comprehensively consider information from different sources and avoid the limitations of a single classification result. Then, the weight information corresponding to the candidate quantified lane change classification results and the preliminary quantified lane change classification results is obtained, so that the importance of the two in the final decision can be flexibly adjusted according to different scenarios and requirements. The candidate quantified lane change classification results and the preliminary quantified lane change classification results are respectively multiplied by the corresponding weight information to obtain the target quantified lane change classification results, ensuring the accuracy of the obtained target quantified lane change classification results. Then, based on the target quantified lane change classification results, the target lane change classification results corresponding to the target leading vehicle are determined, ensuring the accuracy of the determined target lane change classification results.
[0238] The above method considers both the operation data and video data corresponding to the target leading vehicle, and also introduces the expert opinions of the large language model, ensuring the accuracy of the output target lane change classification results. At the same time, when facing driving scenarios that do not appear in the ideal dataset, the large language model can continuously receive new driving environment information and feedback data, dynamically update its knowledge base and reasoning ability, and also give certain expert opinions. As a result, the driver of the vehicle or the autonomous driving system can learn in advance the lane change intention of the leading vehicle and have enough time to make corresponding responses, such as adjusting the vehicle speed, maintaining a safe distance, or preparing to avoid, effectively reducing the occurrence probability of traffic accidents such as rear-end collisions and scratches, and ensuring driving safety.
[0239] In this embodiment, a device for identifying the lane change behavior of a leading vehicle is also provided. This device is used to implement the above embodiments and preferred implementation manners, and those that have been described will not be repeated. As used below, the term "module" can be a combination of software and / or hardware that realizes a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation in hardware, or a combination of software and hardware is also possible and contemplated.
[0240] This embodiment provides a device for identifying the lane change behavior of a leading vehicle, as Figure 4 shown, including:
[0241] A first acquisition module 401, configured to acquire the operation data of the target leading vehicle corresponding to the vehicle; the operation data includes at least one of the relative speed of the target leading vehicle relative to the vehicle, the relative distance of the target leading vehicle relative to the vehicle, the lateral displacement of the target leading vehicle, and the lateral acceleration of the target leading vehicle;
[0242] A second acquisition module 402, configured to acquire multiple frames of target leading vehicle driving images corresponding to the target leading vehicle; the target leading vehicle driving images include the target leading vehicle;
[0243] An identification module 403 is configured to input the running data and the driving images of the preceding vehicle in each frame into a target preceding-vehicle lane-changing behavior identification model, and output a preliminary lane-changing classification result corresponding to the target preceding vehicle. The preliminary lane-changing classification result includes changing lanes to the left, changing lanes to the right, and not changing lanes.
[0244] In some alternative embodiments, the target preceding-vehicle lane-changing behavior identification model includes a first target feature extraction network and a second target feature extraction network. The identification module 403 is specifically configured to input the running data into the first target feature extraction network in the target preceding-vehicle lane-changing behavior identification model to output the running timing features corresponding to the current running state of the target preceding vehicle; input the driving images of the target preceding vehicle in each frame into the second target feature extraction network in the target preceding-vehicle lane-changing behavior identification model to output the position image features corresponding to the target preceding vehicle; and output the preliminary lane-changing classification result corresponding to the target preceding vehicle based on the running timing features and the position image features.
[0245] In some alternative embodiments, the first target feature extraction network includes a first sub-feature extraction network and a second sub-feature extraction network. The identification module 403 is specifically configured to input the running data into the first sub-feature extraction network in the first target feature extraction network to output the global timing features corresponding to the running data; input the running data into the second sub-feature extraction network in the first target feature extraction network to output the local timing features corresponding to the running data; and fuse the global timing features and the local timing features to output the running timing features corresponding to the target preceding vehicle.
[0246] In some alternative embodiments, the second target feature extraction network includes a third sub-feature extraction network and a fourth sub-feature extraction network. The identification module 403 is specifically configured to input the driving images of the target preceding vehicle in each frame into the third sub-feature extraction network in the second target feature extraction network to output the global image features corresponding to the target preceding vehicle; input the driving images of the target preceding vehicle in each frame into the fourth sub-feature extraction network in the second target feature extraction network to output the local image features corresponding to the target preceding vehicle; and perform a fusion process on the global image features and the local image features to generate the position image features corresponding to the target preceding vehicle.
[0247] In some alternative embodiments, the recognition module 403 is specifically configured to map the running timing features to a first query space, a first key space, and a first value space by using a first linear transformation matrix; map the position image features to a second query space, a second key space, and a second value space by using a second linear transformation matrix; calculate a first attention score based on the first query space and the second key space; calculate a second attention score based on the second query space and the first key space; perform normalization processing on the first attention score and the second attention score respectively to obtain a first normalized attention score and a second normalized attention score; obtain a position feature encoding corresponding to the target leading vehicle by adding a first product obtained by multiplying the first normalized attention score by the second value space and a second product obtained by multiplying the second normalized attention score by the first value space; and output a preliminary lane change classification result corresponding to the target leading vehicle based on the position feature encoding.
[0248] In an alternative embodiment of the present application, as Figure 5 shown, the above-mentioned leading vehicle lane change behavior recognition device further includes:
[0249] A third acquisition module 404, configured to acquire driving environment information corresponding to the vehicle itself, where the driving environment information includes surrounding vehicle running information and obstacle information corresponding to the vehicle itself;
[0250] A conversion module 405, configured to convert the driving environment information into target text information;
[0251] An input module 406, configured to input the target text information into a target large language model and output a candidate lane change classification result corresponding to the target leading vehicle;
[0252] An output module 407, configured to combine the candidate lane change classification result with the preliminary lane change classification result and output a target lane change classification result corresponding to the target leading vehicle.
[0253] In an alternative embodiment of the present application, the above input module 406 is specifically configured to input target text information into a target large language model; the target large language model extracts features from the target text information to obtain key elements corresponding to the target text information; the key elements include at least one of the number of surrounding vehicles, the speed of each surrounding vehicle, the driving direction of each surrounding vehicle, the distance between each surrounding vehicle and the vehicle itself, and the type, position, and size of obstacles; based on the key elements, a first prompt message is generated; wherein, traffic rules and driving behavior patterns are introduced into the target large language model; the first prompt message is input into the target large language model to obtain a first-round output result; the first-round output result is analyzed to extract the doubtful or unclear parts in the first-round output result, and a second prompt message is constructed again; the second prompt message is input into the target large language model to obtain a second-round output result; the second-round output result includes a sub-candidate lane-changing classification result; the sub-candidate lane-changing classification result is evaluated, and based on the evaluation result, a candidate lane-changing classification result is generated.
[0254] In an alternative embodiment of the present application, the above output module 407 is specifically configured to perform quantization processing on the candidate lane-changing classification result and the preliminary lane-changing classification result respectively to obtain a candidate quantized lane-changing classification result and a preliminary quantized lane-changing classification result; obtain the weight information corresponding to the candidate quantized lane-changing classification result and the preliminary quantized lane-changing classification result; multiply the candidate quantized lane-changing classification result and the preliminary quantized lane-changing classification result by the corresponding weight information respectively to obtain a target quantized lane-changing classification result; based on the target quantized lane-changing classification result, determine the target lane-changing classification result corresponding to the target leading vehicle.
[0255] The further functional descriptions of the above various modules and units are the same as those in the corresponding above embodiments, and will not be repeated here.
[0256] The leading vehicle lane-changing behavior recognition device in this embodiment is presented in the form of functional units. Here, the unit refers to an ASIC (Application Specific Integrated Circuit) circuit, a processor and a memory that execute one or more software or fixed programs, and / or other devices that can provide the above functions.
[0257] The embodiment of the present invention also provides an electronic device having the above Figure 4 and / or Figure 5 shown leading vehicle lane-changing behavior recognition device.
[0258] Please refer to Figure 6 , Figure 6 which is a schematic structural diagram of an electronic device provided by an alternative embodiment of the present invention. As shown in Figure 6As shown, the electronic device includes: one or more processors 10, a memory 20, and interfaces for connecting the components, including a high-speed interface and a low-speed interface. Each component communicates with each other using different buses and can be installed on a common motherboard or installed in other ways as needed. The processor can process instructions executed within the electronic device, including instructions stored in the memory or on the memory to display graphical information of the GUI on an external input / output device (such as a display device coupled to the interface). In some alternative embodiments, if needed, multiple processors and / or multiple buses can be used with multiple memories and multiple memories. Similarly, multiple electronic devices can be connected, and each device provides some necessary operations (such as an array of servers, a set of blade servers, or a multi-processor system). Figure 6 Here, a processor 10 is taken as an example.
[0259] The processor 10 can be a central processing unit, a network processor, or a combination thereof. Among them, the processor 10 can further include a hardware chip. The above hardware chip can be an application-specific integrated circuit, a programmable logic device, or a combination thereof. The above programmable logic device can be a complex programmable logic device, a field programmable gate array, a generic array logic, or any combination thereof.
[0260] Among them, the memory 20 stores instructions executable by at least one processor 10, so that at least one processor 10 executes the method shown in the above embodiments.
[0261] The memory 20 can include a program storage area and a data storage area. Among them, the program storage area can store an operating system and application programs required for at least one function; the data storage area can store data created according to the use of the electronic device, etc. In addition, the memory 20 can include a high-speed random access memory, and can also include a non-transitory memory, such as at least one disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some alternative embodiments, the memory 20 can optionally include a memory remotely set relative to the processor 10, and these remote memories can be connected to the electronic device through a network. Examples of the above network include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.
[0262] The memory 20 can include a volatile memory, such as a random access memory; the memory can also include a non-volatile memory, such as a flash memory, a hard disk, or a solid-state hard disk; the memory 20 can also include a combination of the above types of memories.
[0263] The electronic device further includes an input device 30 and an output device 40. The processor 10, the memory 20, the input device 30, and the output device 40 can be connected through a bus or other means.Figure 6 Take the bus connection as an example.
[0264] The input device 30 can receive input digital or character information, and generate key signal inputs related to the user settings and function controls of the electronic device, such as a touch screen, a keypad, a mouse, a trackpad, a touchpad, a pointing stick, one or more mouse buttons, a trackball, a joystick, etc. The output device 40 may include a display device, an auxiliary lighting device (e.g., an LED), and a haptic feedback device (e.g., a vibration motor), etc. The above display device includes but is not limited to a liquid crystal display, a light emitting diode, a display, and a plasma display. In some alternative embodiments, the display device may be a touch screen.
[0265] The embodiments of the present invention also provide a computer-readable storage medium. The method according to the embodiments of the present invention can be implemented in hardware, firmware, or be implemented as computer code that can be recorded on a storage medium, or be implemented by downloading via a network the original computer code stored in a remote storage medium or a non-transitory machine-readable storage medium and to be stored in a local storage medium, so that the method described herein can be stored as such software processing on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only memory, a random access memory, a flash memory, a hard disk, or a solid-state drive, etc.; further, the storage medium can also include a combination of the above types of memories. It can be understood that a computer, a processor, a microprocessor controller, or programmable hardware includes a storage component that can store or receive software or computer code, and when the software or computer code is accessed and executed by the computer, the processor, or the hardware, the method shown in the above embodiments is implemented.
[0266] A part of the present invention can be applied as a computer program product, such as computer program instructions, which when executed by a computer, can call or provide the method and / or technical solution according to the present invention through the operation of the computer. Those skilled in the art should be able to understand that the forms of existence of computer program instructions in a computer-readable medium include but are not limited to source files, executable files, installation package files, etc. Correspondingly, the ways in which computer program instructions are executed by a computer include but are not limited to: the computer directly executes the instruction, or the computer compiles the instruction and then executes the corresponding compiled program, or the computer reads and executes the instruction, or the computer reads and installs the instruction and then executes the corresponding installed program. Herein, the computer-readable medium can be any available computer-readable storage medium or communication medium accessible by the computer.
[0267] Although embodiments of the present invention have been described in conjunction with the accompanying drawings, those skilled in the art can make various modifications and variations without departing from the spirit and scope of the present invention, and such modifications and variations fall within the scope defined by the appended claims.
Claims
1. A method for identifying lane-changing behavior of a preceding vehicle, characterized in that: The method comprises: Acquire operation data of a target preceding vehicle corresponding to the vehicle; the operation data includes at least one of a relative speed of the target preceding vehicle relative to the vehicle, a relative distance of the target preceding vehicle relative to the vehicle, a lateral displacement of the target preceding vehicle, and a lateral acceleration of the target preceding vehicle; Acquire a plurality of frames of target front vehicle driving images corresponding to the target front vehicle; the target front vehicle driving images include the target front vehicle; The operation data and each frame of the preceding vehicle driving image are input into a target preceding vehicle lane changing behavior recognition model, and a preliminary lane changing classification result corresponding to the target preceding vehicle is output, wherein the preliminary lane changing classification result includes changing lanes to the left, changing lanes to the right, and not changing lanes.
2. The method according to claim 1, characterized in that: The target front vehicle lane-changing behavior recognition model includes a first target feature extraction network and a second target feature extraction network; The step of inputting the operation data and each frame of the preceding vehicle driving image into a target preceding vehicle lane-changing behavior recognition model and outputting a preliminary lane-changing classification result corresponding to the target preceding vehicle includes: Inputting the operation data into the first target feature extraction network in the lane-changing behavior recognition model of the target preceding vehicle, and outputting the operation timing features corresponding to the current operation state of the target preceding vehicle; Inputting each frame of the target front vehicle driving image into the second target feature extraction network in the target front vehicle lane change behavior recognition model, and outputting the position image features corresponding to the target front vehicle; Based on the running timing features and the position image features, a preliminary lane change classification result corresponding to the target front vehicle is output.
3. The method according to claim 2, characterized in that The first target feature extraction network includes a first sub-feature extraction network and a second sub-feature extraction network. The first target feature extraction network inputs the operation data into the target front vehicle lane change behavior recognition model, and outputs the operation timing features corresponding to the current operation state of the target front vehicle, including: Inputting the operation data into the first sub-feature extraction network in the first target feature extraction network, and outputting the global temporal features corresponding to the operation data; Inputting the operation data into the second sub-feature extraction network in the first target feature extraction network, and outputting the local temporal features corresponding to the operation data; The global timing features and the local timing features are fused to output the running timing features corresponding to the target preceding vehicle.
4. The method according to claim 2, characterized in that: The second target feature includes a third sub-feature extraction network and a fourth sub-feature extraction network; the second target feature extraction network inputs each frame of the target front vehicle driving image into the target front vehicle lane change behavior recognition model, and outputs the position image feature corresponding to the target front vehicle, including: Inputting each frame of the target front vehicle driving image into the third sub-feature extraction network in the second target feature extraction network, and outputting the global image features corresponding to the target front vehicle; Inputting each frame of the target front vehicle driving image into the fourth sub-feature extraction network in the second target feature extraction network, and outputting the local image features corresponding to the target front vehicle; The global image features and the local image features are fused to generate the position image features corresponding to the target front vehicle.
5. The method according to claim 2, characterized in that: The outputting a preliminary lane change classification result corresponding to the target front vehicle based on the running time sequence feature and the position image feature includes: Mapping the runtime features to a first query space, a first key space, and a first value space using a first linear transformation matrix; Mapping the position image features to a second query space, a second key space, and a second value space using a second linear transformation matrix; Calculating a first attention score based on the first query space and the second key space; Calculating a second attention score based on the second query space and the first key space; Normalizing the first attention score and the second attention score respectively to obtain a first normalized attention score and a second normalized attention score; A first product obtained by multiplying the first normalized attention score by the second value space is added to a second product obtained by multiplying the second normalized attention score by the first value space to obtain a position feature code corresponding to the target preceding vehicle; Based on the position feature coding, the preliminary lane change classification result corresponding to the target front vehicle is output.
6. The method according to claim 1, characterized in that After inputting the operation data and each frame of the preceding vehicle driving image into a target preceding vehicle lane-changing behavior recognition model and outputting a preliminary lane-changing classification result corresponding to the target preceding vehicle, the method further includes: Acquiring driving environment information corresponding to the vehicle, wherein the driving environment information includes surrounding vehicle operation information and obstacle information corresponding to the vehicle; Converting the driving environment information into target text information; Inputting the target text information into a target large language model, and outputting a candidate lane change classification result corresponding to the target front vehicle; The candidate lane-changing classification result is combined with the preliminary lane-changing classification result, and a target lane-changing classification result corresponding to the target preceding vehicle is output.
7. The method according to claim 6, characterized in that The step of inputting the target text information into a target large language model and outputting a candidate lane change classification result corresponding to the target front vehicle includes: Inputting the target text information into a target large language model; The target large language model performs feature extraction on the target text information to obtain key elements corresponding to the target text information; the key elements include the number of surrounding vehicles, the speed of each of the surrounding vehicles, the driving direction of each of the surrounding vehicles, the distance between each of the surrounding vehicles and the vehicle, and at least one of the type, position and size of obstacles; Based on the key elements, generating first prompt information; wherein the target large language model introduces traffic rules and driving behavior patterns; Inputting the first prompt information into the target large language model to obtain a first round output result; Analyze the first round of output results, extract the doubtful or unclear parts of the first round of output results, and re-construct the second prompt information; Inputting the second prompt information into the target large language model to obtain a second round of output results; the second round of output results includes a sub-candidate lane change classification result; The sub-candidate lane change classification results are evaluated, and based on the evaluation results, the candidate lane change classification results are generated.
8. The method according to claim 6, characterized in that The combining the candidate lane change classification result with the preliminary lane change classification result to output the target lane change classification result corresponding to the target front vehicle includes: Quantizing the candidate lane-changing classification result and the preliminary lane-changing classification result respectively to obtain a candidate quantized lane-changing classification result and a preliminary quantized lane-changing classification result; Obtaining weight information corresponding to the candidate quantized lane-changing classification result and the preliminary quantized lane-changing classification result; Multiplying the candidate quantized lane-changing classification result and the preliminary quantized lane-changing classification result by the corresponding weight information respectively to obtain a target quantized lane-changing classification result; Based on the target quantified lane change classification result, the target lane change classification result corresponding to the target front vehicle is determined.
9. A device for identifying lane-changing behavior of a preceding vehicle, characterized in that: The device comprises: A first acquisition module is used to acquire the operation data of the target preceding vehicle corresponding to the vehicle; the operation data includes at least one of the relative speed of the target preceding vehicle relative to the vehicle, the relative distance of the target preceding vehicle relative to the vehicle, the lateral displacement of the target preceding vehicle, and the lateral acceleration of the target preceding vehicle; A second acquisition module is used to acquire a plurality of frames of target front vehicle driving images corresponding to the target front vehicle; the target front vehicle driving images include the target front vehicle; The recognition module is used to input the operating data and each frame of the preceding vehicle driving image into a target preceding vehicle lane changing behavior recognition model, and output a preliminary lane changing classification result corresponding to the target preceding vehicle, wherein the preliminary lane changing classification result includes changing lanes to the left, changing lanes to the right, and not changing lanes.
10. An electronic device, characterized in that: include: A memory and a processor, wherein the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the preceding vehicle lane changing behavior identification method according to any one of claims 1 to 8 by executing the computer instructions.
Citation Information
Patent Citations
Car lane change early warning method based on continuous image constraint pose estimation
CN110745140A
Preceding vehicle lane changing intention prediction method and prediction system
CN111746559A
Vehicle lane change prediction method and device and computer storage medium
CN111950394A
Autonomous lane changing decision planning method and system adaptive to different driving styles and road environments
CN118238847A
Vehicle control method, apparatus, vehicle, electronic device and storage medium
US20210206378A1
Cited By
Air suspension decision-making and switching method
CN120552545A