Path planning method and device, electronic equipment, storage medium and vehicle

By using static and dynamic object detection models combined with point cloud data for road condition recognition in autonomous driving, the problem of unclear distinction between static and dynamic objects in path planning is solved, thus improving the accuracy and efficiency of path planning.

CN120840618APending Publication Date: 2025-10-28BEIJING CO WHEELS TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410509483.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-04-25
Publication Date
2025-10-28

AI Technical Summary

Technical Problem

Existing autonomous driving technologies cannot effectively distinguish between static and dynamic objects, resulting in inaccurate reference information for path planning and affecting the accuracy of path planning.

Method used

Static and dynamic object detection models are used to identify static and dynamic objects in video data, respectively. Road condition recognition is performed by combining point cloud data to generate the target driving path.

Benefits of technology

It improves the accuracy of identifying static and dynamic objects, and enhances the accuracy and efficiency of path planning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120840618A_ABST
    Figure CN120840618A_ABST
Patent Text Reader

Abstract

The invention relates to a path planning method and device, electronic equipment, a storage medium and a vehicle, and the method comprises the steps: obtaining video data and point cloud data around the vehicle; inputting each video frame in the video data into a static object detection model, and outputting a static object recognition result by the static object detection model; inputting the video data and the point cloud data into a dynamic object detection model, and outputting a dynamic object recognition result by the dynamic object detection model; performing road condition recognition based on the static object recognition result and the dynamic object recognition result to obtain a road condition prediction result; and path planning is performed according to the road condition prediction result, and the target driving path is generated, so that the static object and the dynamic object can be identified through the static object detection model and the dynamic object detection model respectively, the accuracy of the obtained static identification result and the dynamic object identification result is improved, and the accuracy of path planning is further improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of autonomous driving technology, and more particularly to a path planning method, apparatus, electronic device, storage medium, and vehicle. Background Technology

[0002] Object detection is crucial in autonomous driving. In existing autonomous driving processes, cameras installed on the vehicle capture images of the vehicle's surroundings. These images are then input into an object detection model for object detection, resulting in recognition results. Path planning is then performed based on these results.

[0003] In actual vehicle operation, there are not only static objects but also dynamic objects around the vehicle. However, the target detection methods mentioned above cannot distinguish between static and dynamic objects, which leads to inaccurate reference information required for path planning and thus affects the accuracy of path planning. Summary of the Invention

[0004] To address the aforementioned technical problems, this disclosure provides a path planning method, apparatus, electronic device, storage medium, and vehicle.

[0005] A first aspect of this disclosure provides a path planning method, the method comprising:

[0006] Acquire video and point cloud data around the vehicle;

[0007] Each video frame in the video data is input into the static object detection model, which outputs the static object recognition result. The static object detection model is used to identify static objects in the video data.

[0008] Video data and point cloud data are input into the dynamic object detection model, which outputs the dynamic object recognition result. The dynamic object detection model is used to identify dynamic objects in video data by combining point cloud data.

[0009] Traffic condition identification is performed based on static object recognition results and dynamic object recognition results to obtain traffic condition prediction results;

[0010] Based on the road condition prediction results, route planning is performed to generate the target driving route.

[0011] A second aspect of this disclosure provides a path planning apparatus, the apparatus comprising:

[0012] The data acquisition module is used to acquire video data and point cloud data around the vehicle;

[0013] The first recognition module is used to input each video frame in the video data into the static object detection model, and the static object detection model outputs the static object recognition result.

[0014] The second recognition module is used to input video data and point cloud data into the dynamic object detection model, and the dynamic object detection model outputs the dynamic object recognition result.

[0015] The third identification module is used to identify road conditions based on static object identification results and dynamic object identification results, and to obtain road condition prediction results.

[0016] The route planning module is used to plan routes based on road condition prediction results and generate target driving routes.

[0017] A third aspect of this disclosure provides an electronic device, the device comprising:

[0018] Memory;

[0019] Processor; and

[0020] A computer program, wherein the computer program is stored in memory and configured to be executed by a processor to implement the path planning method as described in the first aspect above.

[0021] A fourth aspect of this disclosure provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the path planning method of the first aspect described above.

[0022] A fifth aspect of this disclosure provides a vehicle that includes the path planning device of the second aspect or the electronic device of the third aspect described above.

[0023] The technical solution provided in this disclosure has the following advantages compared with the prior art:

[0024] The path planning method, apparatus, electronic device, storage medium, and vehicle provided in this disclosure can acquire video data and point cloud data around the vehicle. Each video frame in the video data is input into a static object detection model, which outputs a static object recognition result. The video data and point cloud data are input into a dynamic object detection model, which outputs a dynamic object recognition result. The static object detection model is used to identify static objects in the video data, and the dynamic object detection model is used to identify dynamic objects in the video data in combination with the point cloud data. After obtaining the static and dynamic object recognition results, road condition recognition is performed based on the static and dynamic object recognition results to obtain a road condition prediction result. Path planning is then performed based on the road condition prediction result to generate a target driving path. Thus, static and dynamic objects can be identified by the static and dynamic object detection models respectively, improving the accuracy of the obtained static and dynamic object recognition results, thereby improving the accuracy of path planning. Attached Figure Description

[0025] The accompanying drawings, which are incorporated in and form a part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure.

[0026] To more clearly illustrate the technical solutions in the embodiments of this disclosure or the prior art, the accompanying drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, those skilled in the art can obtain other drawings based on these drawings without creative effort.

[0027] Figure 1 This is a flowchart of a path planning method provided in an embodiment of this disclosure;

[0028] Figure 2 This is a flowchart of another path planning method provided in this embodiment of the disclosure;

[0029] Figure 3 This is a schematic diagram of the structure of a path planning device provided in an embodiment of this disclosure;

[0030] Figure 4 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this disclosure. Detailed Implementation

[0031] To better understand the above-mentioned objectives, features, and advantages of this disclosure, the solutions disclosed herein will be further described below. It should be noted that, unless otherwise specified, the embodiments and features described herein can be combined with each other.

[0032] Numerous specific details are set forth in the following description in order to provide a full understanding of this disclosure, but this disclosure may also be implemented in other ways different from those described herein; obviously, the embodiments in the specification are only some, and not all, of the embodiments of this disclosure.

[0033] It should be understood that the steps described in the method embodiments of this disclosure may be performed in different orders and / or in parallel. Furthermore, the method embodiments may include additional steps and / or omit the steps shown. The scope of this disclosure is not limited in this respect.

[0034] It should be noted that, in this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0035] It should be noted that the terms "a" and "a plurality of" used in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".

[0036] In general, existing autonomous driving processes suffer from inaccurate dynamic object detection results, leading to inaccurate path planning. To address this issue, this disclosure provides a path planning method, which is described below with reference to specific embodiments.

[0037] Figure 1 This is a flowchart of a route planning method provided in an embodiment of the present disclosure. The method can be executed by a route planning device, which can be implemented in software and / or hardware. The route planning device can be configured in an electronic device, such as a server, terminal, or server cluster. Specifically, the terminal can include a computer or tablet computer, a vehicle terminal, or any device capable of processing the route planning method.

[0038] like Figure 1 As shown in the embodiments of this disclosure, the path planning method includes the following steps.

[0039] S110: Acquire video data and point cloud data around the vehicle.

[0040] In this embodiment of the disclosure, the electronic device acquires video data and point cloud data around the vehicle in response to vehicle startup or vehicle entering autonomous driving mode or semi-autonomous driving mode.

[0041] Specifically, when the vehicle starts or enters an autonomous or semi-autonomous driving mode, the electronic device controls the activation of image acquisition devices and point cloud data acquisition devices such as cameras and lidar installed on the vehicle. There can be multiple image acquisition devices and point cloud data acquisition devices, which are located at different positions on the vehicle and are used to collect environmental data around the vehicle. The image acquisition devices upload the collected video data to the electronic device, and the point cloud data acquisition devices upload the collected point cloud data to the electronic device, thereby obtaining video data and point cloud data around the vehicle.

[0042] The acquisition parameters of the image acquisition device and the point cloud data acquisition device can be preset according to the application scenario, etc.

[0043] For example, the point cloud data acquisition device is a lidar, which has an acquisition accuracy of 0.1 degrees angular resolution and a detection range of 200 meters; the image acquisition device is a camera, which has a resolution of 1920×1080 and an acquisition frequency of 30 frames per second; the acquisition frequencies of the lidar and the camera can be the same.

[0044] S120. Each video frame in the video data is input into the static object detection model, and the static object detection model outputs the static object recognition result. The static object detection model is used to identify static objects in the video data.

[0045] In this embodiment of the disclosure, after acquiring video data, the electronic device inputs each video frame in the video data into the static object detection model, and the static object detection model outputs the static object recognition result.

[0046] In this embodiment of the disclosure, the static object detection model can be any machine learning model capable of detecting and recognizing static objects, such as convolutional neural networks, YOLO models, etc., without limitation. Specifically, the static object detection model is used to recognize static objects in video data.

[0047] Optionally, static objects are objects that remain stationary, such as lanes, traffic signs, trees, buildings, etc.

[0048] Static object recognition results can include the category of the static object, the probability that the static object belongs to that category, and the detection bounding box information of the static object.

[0049] Specifically, after acquiring video data, the electronic device inputs each video frame into the static object detection model according to the chronological order of the timestamps of each video frame in the video data. The static object detection model extracts the features of the video frames to obtain the feature maps corresponding to the video frames, and then detects and identifies each object in the video frames based on the feature maps corresponding to the video frames, and determines the static object, the category of the static object, and the probability of the static object belonging to the category, and uses this as the static object recognition result, and outputs the static object recognition result through the output layer.

[0050] For example, the categories of static objects can include multiple major categories such as lanes, traffic signs, buildings, and trees. Each major category can also include multiple subcategories. For example, lanes can include double yellow lines, single white lines, etc., and traffic signs can include no entry, slow down, etc.

[0051] S130, which is parallel to S120, inputs video data and point cloud data into the dynamic object detection model, and outputs the dynamic object recognition result. The dynamic object detection model is used to identify dynamic objects in video data by combining point cloud data.

[0052] In this embodiment of the disclosure, after acquiring video data and point cloud data around the vehicle, the electronic device inputs the video data and point cloud data into the dynamic object detection model, and the dynamic object detection model outputs the dynamic object recognition result.

[0053] In this embodiment of the disclosure, the dynamic object detection model can be any machine learning model capable of detecting and recognizing dynamic objects, such as recurrent neural networks, backpropagation neural networks, etc., without limitation. Specifically, the dynamic object detection model is used to recognize dynamic objects in video data by combining point cloud data.

[0054] Optionally, a dynamic object is an object that can move, that is, an object whose position can change, such as pedestrians, vehicles around a vehicle, etc.

[0055] The results of dynamic object recognition can include the category of the dynamic object and the predicted trajectory of the dynamic object in the next moment.

[0056] Specifically, after acquiring video data and point cloud data around the vehicle, the electronic device inputs the video data and point cloud data into the dynamic object detection model. The dynamic object detection model detects objects in the video data based on the video data, determines target objects belonging to the dynamic object category, and / or determines the target object based on the positional changes of multiple objects in the current and previous video data and point cloud data. Based on the target object's current and previous position information, it predicts the target object's trajectory for the next moment and outputs the target object's trajectory for the next moment based on the output layer of the dynamic object detection model, thereby obtaining the dynamic object recognition result.

[0057] It should be noted that steps S120 and S130 can be performed simultaneously.

[0058] S140. Based on the static object recognition results and dynamic object recognition results, road condition recognition is performed to obtain road condition prediction results.

[0059] In this embodiment of the disclosure, after obtaining the static object recognition result and the dynamic object recognition result, the electronic device performs road condition recognition based on the static object recognition result and the dynamic object recognition result to obtain the road condition prediction result.

[0060] Specifically, after acquiring the static object recognition results and the dynamic object recognition results, the electronic device uses an ensemble learning model, such as random forest, to comprehensively analyze the static object recognition results and the dynamic object recognition results, thereby obtaining the road condition prediction results.

[0061] The specific implementation method of using ensemble learning models, such as random forests, to comprehensively analyze the static and dynamic object recognition results is similar to the real-time method of data comprehensive analysis using existing ensemble learning models, and will not be elaborated here.

[0062] For example, the static object recognition result is "Slow down" traffic sign ahead, and the dynamic object recognition result is "a bicycle, located 10 meters away from this vehicle, with a speed of 20 kilometers per hour, and the next travel trajectory is to cross the vehicle travel route". At this time, the road condition prediction result is: continuing to travel along the current route will result in a collision with the bicycle.

[0063] S150. Based on the road condition prediction results, perform route planning and generate the target driving route.

[0064] In this embodiment of the disclosure, after obtaining the traffic prediction result, the electronic device performs path planning based on the traffic prediction result and generates a target driving path.

[0065] Specifically, after obtaining the traffic prediction results, the electronic device uses a path planning algorithm such as the A* algorithm or a variant of the A* algorithm to analyze the traffic prediction results and perform path planning, generating a new path, i.e., the target driving path.

[0066] The specific implementation method of using a path planning algorithm based on road condition prediction results is similar to the existing implementation method of path planning based on path planning algorithms, and will not be described in detail here.

[0067] In this embodiment, video data and point cloud data around the vehicle can be acquired. Each video frame in the video data is input into a static object detection model, which outputs a static object recognition result. The video data and point cloud data are input into a dynamic object detection model, which outputs a dynamic object recognition result. The static object detection model is used to identify static objects in the video data, while the dynamic object detection model is used to identify dynamic objects in the video data in combination with the point cloud data. After obtaining the static and dynamic object recognition results, road condition recognition is performed based on the static and dynamic object recognition results to obtain a road condition prediction result. Path planning is then performed based on the road condition prediction result to generate a target driving path. Thus, static and dynamic objects can be identified by the static and dynamic object detection models respectively, improving the accuracy of the obtained static and dynamic object recognition results, and consequently improving the accuracy of path planning.

[0068] In this embodiment of the disclosure, two object detection models can be used to identify static objects and dynamic objects respectively, which improves the accuracy of the obtained static object identification results and dynamic object identification results, as well as the efficiency of obtaining dynamic object identification results.

[0069] Based on the above embodiments of this disclosure, each video frame in the video data is input into a static object detection model, and the static object detection model outputs a static object recognition result. Specifically, this may include: inputting each video frame in the video data into the static object detection model; for each video frame, performing a convolution operation on the video frame based on the convolutional layer in the static object detection model to obtain a first feature map corresponding to the video frame; performing a nonlinear operation on the first feature map based on a preset activation function to obtain a second feature map; performing pooling processing on the second feature map based on the pooling layer in the static object detection model to obtain a third feature map; inputting the third feature map into a fully connected layer in the static object detection model, and having the fully connected layer recognize the third feature map to obtain a static object recognition result, wherein the static object recognition result includes the category corresponding to the static object and the probability of belonging to the category.

[0070] In this embodiment, the static object detection model can be a CNN model. The static object detection model can be trained on training data containing millions of labeled images, including lanes, traffic signs, etc., under different climate and lighting conditions. During the training process, the weights of each network in the model can be updated using a backpropagation algorithm combined with an optimization algorithm to minimize the loss value of the loss function, resulting in a trained detection model. The loss function can be a cross-entropy loss function, with the following formula: L=-\sum_{i}^{C}t_i\log(p_i), where C is the number of categories, t_i is the one-hot encoding of the true label, and p_i is the probability predicted by the model.

[0071] Specifically, the electronic device inputs each video frame from the video data into the static object detection model, and performs convolution operations on the video frames based on the convolutional layers in the static object detection model. The convolutional layers consist of multiple convolutional kernels, and each convolutional kernel performs convolution operations on the input video frames to generate the first feature map.

[0072] In this embodiment of the disclosure, the expression for the convolution operation is: F(i,j)=(K*I)(i,j)=\sum_m\sum_nK(m,n)\cdot I(im,jn).

[0073] Where, F(i,j): represents the pixel value of the feature map at position ((i),(j)) after the convolution operation; K: represents the convolution kernel or filter, which is a matrix of predefined size used to extract specific features of the image; I: represents the input image; *: represents the convolution operation; K(m,n): represents the element value at position ((m),(n)) in the convolution kernel; I(im,jn)): represents the element value at the position corresponding to the convolution kernel in the input image; \sum_m\sum_n: represents the summation operation on all elements in the convolution kernel; cdot: represents the dot product, also known as the scalar product.

[0074] In this embodiment, the specific steps of the convolution operation are as follows: First, the convolution kernel slides, that is, the convolution kernel (K) slides from left to right and from top to bottom on the input image, i.e., the input video frame. The stride of each slide can be set as needed, and the stride determines the size of the feature map; second, element-wise product summation is performed, that is, at each position, the corresponding convolution kernel (K) and the local region elements on the input image (i.e., the input video frame) are multiplied element by element, and then all products are summed. This sum is the pixel value of the feature map (F) at the current position ((i), (j)); third, boundary processing is performed, that is, according to the requirements, a certain number of zeros are padded around the input image to control the size of the feature map, and finally the first feature map of the input image, i.e., the video frame, is obtained.

[0075] After obtaining the first feature map, a nonlinear operation is performed on the first feature map based on a preset activation function to obtain the second feature map. The preset activation function can be the ReLU activation function, or it can be an activation function such as sigmoid or tanh, etc. There are no restrictions here.

[0076] Taking the ReLU activation function as an example, the first feature map is subjected to a nonlinear operation based on the ReLU activation function, and nonlinearity is introduced into the first feature map to obtain the second feature map. The specific formula is: \text{ReLU}(x)=\max(0,x).

[0077] Where, \text{ReLU}(x): represents the standard notation of the ReLU function, indicating the output after processing by the ReLU function. Here, (x) usually refers to the output value of a neural network layer (such as a convolutional layer), that is, a pixel value in the feature map or the output value of a neuron in a fully connected layer; (x): represents the input value. In the scenario of applying ReLU, (x) can be any real value, positive or negative, representing the original output of a certain layer of the network; \max(0,x): This part is the core of the ReLU function. It means taking the maximum value between (0) and (x). The result leads to two possibilities for the output of the ReLU function: if (x) is positive, that is, (x>0), then \text{ReLU}(x)=x, indicating that the input (x) is directly passed while keeping its value unchanged; if (x) is negative, that is, (x\leq 0), then \text{ReLU}(x)=0, indicating that the negative value of the input (x) is "corrected" to (0), that is, ignored.

[0078] In this embodiment of the disclosure, a nonlinear operation is performed on the first feature map based on a preset activation function to introduce nonlinearity into the first feature map and avoid linear combinations in the neural network.

[0079] In this embodiment of the disclosure, after obtaining the second feature map, the second feature map is pooled based on the pooling layer in the static object detection model to obtain the third feature map. The pooling process includes max pooling and average pooling. Performing the pooling process reduces the size of the second feature map, which helps to reduce the amount of computation while ensuring that the information is kept as inconvenient as possible.

[0080] Furthermore, after obtaining the third feature map, the third feature map is input into the fully connected layer in the static object detection model, and the fully connected layer recognizes the third feature map to obtain the static object recognition result.

[0081] The specific processing formula for recognizing the third feature map through the fully connected layer is: P(y_i)=\frac{e^{z_i}}{\sum_{k=1}^Ke^{z_k}}.

[0082] Where, P(y_i): represents the probability of predicting category i; e^{z_i}: where e is the base of the natural logarithm, and z_i is the raw score (logits) output by the fully connected layer for category i; K: represents the total number of categories; e^{z_i} is performed once for each category; \sum_{k=1}^Ke^{z_k}: represents summing e^{z_k} for all categories, as the normalized denominator.

[0083] Each P(y_i) represents the probability that the network predicts the input data belongs to category i. The category with the highest probability is selected as the category corresponding to the static object.

[0084] In this embodiment of the disclosure, before inputting video data and point cloud data into the dynamic object detection model and before the dynamic object detection model outputs the dynamic object recognition result, the path planning method may further include: performing data fusion processing on the video frame and point cloud data at each moment to obtain the fusion result at the current moment.

[0085] Specifically, after obtaining video data and point cloud data, the electronic device aligns the video data and point cloud data in a time series, and acquires video frames and point cloud data at each moment. The video frames and point cloud data are then fused based on a preset data fusion algorithm to obtain the fusion result at the current moment. The fusion result includes the visual information of the video frames and the spatial data of the point cloud data, which means it contains a balanced and rich data input, thereby further improving the processing speed and accuracy of the dynamic object detection model.

[0086] Furthermore, video data and point cloud data are input into the dynamic object detection model, and the dynamic object detection model outputs the dynamic object recognition result. Specifically, this may include: inputting the fusion result at the current moment into the sequence prediction layer of the dynamic object detection model, and the sequence prediction layer predicting the fusion result at the current moment and the position information of the target object at the previous moment to obtain the movement trajectory of the target object at the next moment, and determining the movement trajectory of the target object at the next moment as the dynamic object recognition result, and the output layer of the dynamic object detection model outputs the dynamic object recognition result.

[0087] In this embodiment of the disclosure, the dynamic object detection model can be a recurrent neural network (RNN). The dynamic object detection model can be a model trained using time-series data collected from actual traffic flow. The time-series data collected from actual traffic flow can include information such as the speed, direction, and position of a vehicle.

[0088] During the training of a dynamic object detection model, backpropagation can be used to update the model's parameters over time. Specifically, this involves calculating the gradient at each time step in reverse chronological order and updating the weights accordingly.

[0089] Taking vehicle trajectory prediction as an example, the loss function can be:

[0090] L=\frac{1}{N}\sum_{t=1}^{N}(y_t-\hat{y}_t)^2.

[0091] Where y_t represents the actual next position; \hat{y}_t represents the predicted position; N represents the length of the sequence; (y_t-\hat{y}_t)^2 represents the square of the difference between the model's predicted value and the actual value at time t. The squaring operation ensures that the loss value is always positive, while amplifying the impact of larger errors.

[0092] In this embodiment of the disclosure, the processing formula for the sequence prediction layer at each time point is as follows:

[0093] h_t=\text{tanh}(W_{xh}x_t+W_{hh}h_{t-1}+b)

[0094] Where, h_t: represents the hidden state at time step t. The hidden state is the memory part of the sequence prediction layer, containing information from the previous time step and used for calculation at the current time step. The hidden state can also be passed to the next time step or used for output calculation; x_t: the input at time t. In sequence data, x_t is usually a vector representing the data at that time step; W_{xh}: represents the weight matrix from the input to the hidden layer. This matrix is ​​one of the learning parameters used to adjust the influence of the input data on the hidden state; W_{hh}: represents the weight matrix from the hidden layer to the hidden layer. The weight matrix itself is the core of the recurrent structure because it processes the hidden state of the previous time step and affects the current state; h_{t-1}: represents the hidden state of the previous time step t-1. In the recursive structure, this value runs through the entire sequence and is updated at each step; b: represents the bias term, which is a learnable parameter, usually a vector, that adds extra degrees of freedom to the calculation of the hidden state; tanh: represents the hyperbolic tangent activation function, which compresses the linearly transformed value to the range (-1,1) and introduces nonlinearity.

[0095] Specifically, for each time step, the sequence prediction layer first performs a linear transformation, that is, it calculates the product of the current input x_t and the weight W_{xh}, and adds it to the product of the previous hidden state h_{t-1} and the weight W_{hh} to obtain the first result. Next, a bias is added, and the result calculated in the previous step is added to the bias vector b to obtain the second result. Then, an activation function is applied, such as using the biplane tangent function to process the second result, mapping the value of the second result to the range (-1,1), while adding necessary nonlinearity. Furthermore, the hidden state is updated to obtain a new hidden state h_t, which is passed to the next time step to calculate the final output result. Finally, the output layer maps the hidden state to the required output form for output.

[0096] For example, taking the final output as the position of the next moment as an example, the processing formula of this output layer is: \hat{y}t=W{hy}h_t+b_y).

[0097] Where, \hat{y}t: represents the output of the predicted next time step; W{hy}: represents the weight matrix from the sequence prediction layer (i.e., the hidden layer) to the output layer; b_y: represents the bias term of the output layer.

[0098] In this embodiment of the disclosure, a fusion result is obtained by fusing video frames and point cloud data in video data at the same time. The fusion result is then input into a dynamic object detection model to obtain a dynamic object recognition result, thereby improving the accuracy and efficiency of dynamic object recognition.

[0099] In this embodiment of the disclosure, road condition recognition is performed based on static object recognition results and dynamic object recognition results to obtain road condition prediction results. Specifically, this may include: performing a comprehensive analysis of static object recognition results and dynamic object recognition results based on a preset integration strategy to obtain road condition prediction results.

[0100] Optionally, the preset integration strategy may include one or more of the following: averaging (such as simple averaging, weighted averaging, etc.), voting, learning (such as stacking).

[0101] Specifically, after acquiring the static object recognition results and the dynamic object recognition results, the electronic device performs a comprehensive analysis of the static object recognition results and the dynamic object recognition results through a preset integration strategy to obtain the road condition prediction results.

[0102] In this embodiment of the disclosure, before performing road condition identification based on static object identification results and dynamic object identification results to obtain road condition prediction results, the path planning method may further include: acquiring first data of target vehicles adjacent to the vehicle within the target area and second data of roadside units.

[0103] In this embodiment of the disclosure, the target area may be a pre-defined range area for acquiring target vehicles and road test units around the vehicle.

[0104] In this embodiment of the disclosure, the target vehicle adjacent to the vehicle may include other vehicles around the vehicle, or vehicles that are within a preset distance from the vehicle but indirectly affect the vehicle's movement; the first data may include information such as the target vehicle's position, driving direction, speed, and distance from the vehicle; the second data may include basic information of the road testing unit and information such as the road testing unit's position and distance from the vehicle.

[0105] In some embodiments of this disclosure, electronic devices can use vehicle-to-everything (V2X) technology to obtain first data of target vehicles adjacent to the vehicle within the target area and second data of roadside units from the V2X platform.

[0106] In other embodiments of this disclosure, when the target area is within the acquisition range of the lidar, the first data and the second data can be calculated from the point cloud data acquired by the lidar.

[0107] Furthermore, traffic condition recognition is performed based on the static object recognition results and dynamic object recognition results to obtain traffic condition prediction results. Specifically, this may include: performing traffic condition recognition based on the static object recognition results, dynamic object recognition results, first data, and second data to obtain traffic condition prediction results.

[0108] Specifically, the implementation method for obtaining traffic prediction results by performing traffic condition identification based on static object identification results, dynamic object identification results, first data, and second data is similar to the implementation method for obtaining traffic prediction results by performing traffic condition identification based on static object identification results and dynamic object identification results described above, except that more traffic data is added on this basis, which will not be elaborated here.

[0109] In this embodiment of the disclosure, by acquiring the first data of the target vehicle and the second data of the road test unit, road condition identification can be performed based on the static object identification result, the dynamic object identification result, the first data, and the second data to obtain the road condition prediction result, thereby improving the accuracy of the road condition prediction result and thus improving the accuracy of the route planning, which can avoid traffic congestion and other situations.

[0110] In this embodiment of the disclosure, after generating a target driving path by performing path planning based on the road condition prediction results, the path planning method may further include: generating a control command based on the target driving path and a preset command generation frequency, and sending the control command to the controller corresponding to the control command so that the controller executes the control command.

[0111] In this embodiment of the disclosure, the control commands may include vehicle speed control commands, vehicle driving direction control commands, vehicle braking commands, etc.

[0112] In this embodiment, the electronic device is pre-set with a command generation frequency for generating control commands. The preset command generation frequency is used to control the frequency of generating the control commands so that the vehicle's stability is greater than a preset stability threshold. Thus, control commands can be generated based on the preset command generation frequency and the target driving path to ensure the vehicle's driving stability and vehicle safety during autonomous driving.

[0113] In this embodiment of the disclosure, control commands can be generated based on the target driving path, and the control commands can be sent to the controller corresponding to the control commands so that the controller can execute the control commands, thereby achieving the effect of controlling the vehicle to drive safely.

[0114] Figure 2 This is a flowchart of another path planning method provided in this disclosure embodiment, such as... Figure 2 As shown, the path planning method may include the following steps:

[0115] S210: Acquire video data and point cloud data around the vehicle.

[0116] S220. Input each video frame in the video data into the static object detection model, and output the static object recognition result from the static object detection model.

[0117] S230, which is parallel to S220, inputs video data and point cloud data into the dynamic object detection model, and the dynamic object detection model outputs the dynamic object recognition result.

[0118] S240, Obtain first data of target vehicles adjacent to the vehicle within the target area and second data of roadside units.

[0119] S250, based on the static object recognition results, dynamic object recognition results, first data and second data, performs road condition recognition to obtain road condition prediction results.

[0120] S260. Based on the road condition prediction results, perform route planning and generate the target driving route.

[0121] S270. Generate control commands based on the target driving path and send the control commands to the controller corresponding to the control commands so that the controller executes the control commands.

[0122] It should be noted that the specific implementation of steps S210-S270 is similar to the implementation of the relevant steps in the above embodiments of this disclosure, and will not be repeated here.

[0123] In this embodiment of the disclosure, static objects and dynamic objects in video data can be detected by two models, a static object detection model and a dynamic object detection model, respectively, which improves the accuracy and speed of object detection. At the same time, road condition recognition can be performed by combining the first data of target vehicles adjacent to the vehicle in the target area and the second data of the roadside unit, which improves the accuracy of the obtained road condition prediction results, and further improves the accuracy of the obtained target driving path.

[0124] Figure 3 This is a schematic diagram of the structure of a path planning device provided in an embodiment of this disclosure. The path planning device in this embodiment can be installed in an electronic device, which can be a server, a terminal, or a server cluster. Specifically, the terminal can include a computer or tablet computer, a vehicle terminal, or any device capable of processing path planning methods, etc., without limitation.

[0125] like Figure 3 As shown, the path planning device 300 may include a data acquisition module 310, a first identification module 320, a second identification module 330, a third identification module 340, and a path planning module 350.

[0126] The data acquisition module 310 can be used to acquire video data and point cloud data around the vehicle.

[0127] The first recognition module 320 can be used to input each video frame in the video data into the static object detection model, and the static object detection model outputs the static object recognition result. The static object detection model is used to recognize static objects in the video data.

[0128] The second recognition module 330 can be used to input video data and point cloud data into the dynamic object detection model, and the dynamic object detection model outputs the dynamic object recognition result. The dynamic object detection model is used to recognize dynamic objects in video data by combining point cloud data.

[0129] The third recognition module 340 can be used to identify road conditions based on static object recognition results and dynamic object recognition results, and obtain road condition prediction results.

[0130] The route planning module 350 can be used to plan routes based on road condition prediction results and generate target driving routes.

[0131] In this embodiment, video data and point cloud data around the vehicle can be acquired. Each video frame in the video data is input into a static object detection model, which outputs a static object recognition result. The video data and point cloud data are input into a dynamic object detection model, which outputs a dynamic object recognition result. The static object detection model is used to identify static objects in the video data, while the dynamic object detection model is used to identify dynamic objects in the video data in combination with the point cloud data. After obtaining the static and dynamic object recognition results, road condition recognition is performed based on the static and dynamic object recognition results to obtain a road condition prediction result. Path planning is then performed based on the road condition prediction result to generate a target driving path. Thus, static and dynamic objects can be identified by the static and dynamic object detection models respectively, improving the accuracy of the obtained static and dynamic object recognition results, and consequently improving the accuracy of path planning.

[0132] In some embodiments of this disclosure, the first recognition module 320 may be specifically used to input each video frame in the video data into a static object detection model. For each video frame, a convolution operation is performed on the video frame based on the convolutional layer in the static object detection model to obtain a first feature map corresponding to the video frame; a nonlinear operation is performed on the first feature map based on a preset activation function to obtain a second feature map; the second feature map is pooled based on the pooling layer in the static object detection model to obtain a third feature map; the third feature map is input into the fully connected layer in the static object detection model, and the fully connected layer recognizes the third feature map to obtain a static object recognition result. The static object recognition result includes the category corresponding to the static object and the probability of belonging to the category.

[0133] In some embodiments of this disclosure, the path planning device 300 may further include a data processing module.

[0134] The data processing module can be used to perform data fusion processing on the video frames and point cloud data at each moment before the dynamic object detection model outputs the dynamic object recognition result, after inputting the video data and point cloud data into the dynamic object detection model.

[0135] In some embodiments of this disclosure, the second identification module 330 may be specifically used to input the fusion result at the current moment into the sequence prediction layer of the dynamic object detection model, and the sequence prediction layer predicts the fusion result at the current moment and the position information of the target object at the previous moment to obtain the movement trajectory of the target object at the next moment, and determines the movement trajectory of the target object at the next moment as the dynamic object identification result, and the output layer of the dynamic object detection model outputs the dynamic object identification result.

[0136] In some embodiments of this disclosure, the path planning device 300 may further include an information acquisition module.

[0137] The information acquisition module can be used to acquire first data of target vehicles adjacent to the vehicle in the target area and second data of roadside units before obtaining road condition prediction results based on static object recognition results and dynamic object recognition results.

[0138] In some embodiments of this disclosure, the third identification module 340 may be specifically used to identify road conditions based on static object identification results, dynamic object identification results, first data, and second data, and obtain road condition prediction results.

[0139] In some embodiments of this disclosure, the third identification module 340 can be specifically used to perform a comprehensive analysis of the static object identification results and the dynamic object identification results based on a preset integration strategy to obtain the road condition prediction results.

[0140] In some embodiments of this disclosure, the path planning device 300 may further include a control module.

[0141] The control module can be used to perform path planning based on road condition prediction results, generate a target driving path, generate control commands based on the target driving path and a preset command generation frequency, and send the control commands to the controller corresponding to the control commands so that the controller executes the control commands. The preset command generation frequency is the frequency used to generate control commands so that the vehicle's stability is greater than a preset stability threshold.

[0142] It should be noted that, Figure 3The path planning device 300 shown can execute the various steps in the above method embodiments and realize the various processes and effects in the above method embodiments, which will not be elaborated here.

[0143] Figure 4 A schematic diagram of the structure of an electronic device provided in an embodiment of this disclosure is shown.

[0144] In this embodiment of the disclosure, Figure 4 The electronic device shown can be a server, a terminal, or a server cluster. Specifically, the terminal can include a computer or tablet computer, a vehicle terminal, or any device that can be used for path planning methods, etc., without any limitation.

[0145] like Figure 4 As shown, the electronic device may include a processor 410 and a memory 420 storing computer program instructions.

[0146] Specifically, the processor 410 may include a central processing unit (CPU), an application-specific integrated circuit (ASIC), or one or more integrated circuits that can be configured to implement the embodiments of this application.

[0147] Memory 420 may include mass storage for information or instructions. For example, and not limitingly, memory 420 may include a hard disk drive (HDD), floppy disk drive, flash memory, optical disk, magneto-optical disk, magnetic tape, or Universal Serial Bus (USB) drive, or a combination of two or more of these. Where appropriate, memory 420 may include removable or non-removable (or fixed) media. Where appropriate, memory 420 may be internal or external to the integrated gateway device. In a particular embodiment, memory 420 is non-volatile solid-state memory. In a particular embodiment, memory 420 includes read-only memory (ROM). Where appropriate, the ROM may be a mask-programmed ROM, a programmable ROM (PROM), an erasable PROM (Electrically Programmable ROM, EPROM), an electrically erasable programmable PROM (EEPROM), an electrically alterable ROM (EAROM), or flash memory, or a combination of two or more of these.

[0148] The processor 410 reads and executes computer program instructions stored in the memory 420 to perform the steps of the path planning method provided in the embodiments of this disclosure.

[0149] In one example, the electronic device may also include a transceiver 430 and a bus 440. Wherein, as... Figure 4 As shown, the processor 410, memory 420 and transceiver 430 are connected via bus 440 and communicate with each other.

[0150] Bus 440 may include hardware, software, or both. For example, and not limitingly, the bus may include an Accelerated Graphics Port (AGP) or other graphics bus, an Extended Industry Standard Architecture (EISA) bus, a Front Side Bus (FSB), a Hyper Transport (HT) interconnect, an Industrial Standard Architecture (ISA) bus, an Infinite Bandwidth Interconnect, a Low Pin Count (LPC) bus, a memory bus, a MicroChannel Architecture (MCA) bus, a Peripheral Component Interconnect (PCI) bus, a PCI-Express (PCI-X) bus, a Serial Advanced Technology Attachment (SATA) bus, a Video Electronics Standards Association Local Bus (VLB) bus, or other suitable buses, or a combination of two or more of these. Where appropriate, bus 440 may include one or more buses.

[0151] This disclosure also provides a computer-readable storage medium that can store a computer program that, when executed by a processor, enables the processor to implement the path planning method provided in this disclosure.

[0152] The aforementioned storage medium may, for example, include a memory 420 containing computer program instructions, which can be executed by a processor 410 of an electronic device to complete the path planning method provided in the embodiments of this disclosure. Optionally, the storage medium may be a non-transitory computer-readable storage medium, such as a ROM, random access memory (RAM), compact disc-only memory (CD-ROM), magnetic tape, floppy disk, and optical data storage device.

[0153] This disclosure also provides a vehicle that includes the path planning device or electronic device described in the above embodiments of this disclosure, which can realize the various processes and effects described in the above embodiments of this disclosure, and will not be elaborated here.

[0154] It should be noted that, in this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the term "comprising" is intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus.

[0155] The above description is merely a specific embodiment of this disclosure, enabling those skilled in the art to understand or implement it. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this disclosure. Therefore, this disclosure is not to be limited to the embodiments described herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A path planning method, characterized in that, The method includes: Acquire video and point cloud data around the vehicle; Each video frame in the video data is input into the static object detection model, and the static object detection model outputs the static object recognition result. The static object detection model is used to identify static objects in the video data. The video data and the point cloud data are input into the dynamic object detection model, and the dynamic object detection model outputs the dynamic object recognition result. The dynamic object detection model is used to identify dynamic objects in the video data by combining the point cloud data. Based on the static object recognition results and the dynamic object recognition results, road condition recognition is performed to obtain road condition prediction results; Based on the road condition prediction results, route planning is performed to generate the target driving route.

2. The method according to claim 1, characterized in that, The step of inputting each video frame from the video data into the static object detection model, and having the static object detection model output the static object recognition result, includes: Each video frame in the video data is input into the static object detection model. For each video frame, a convolution operation is performed on the video frame based on the convolutional layer in the static object detection model to obtain the first feature map corresponding to the video frame. A second feature map is obtained by performing a nonlinear operation on the first feature map based on a preset activation function; The second feature map is pooled based on the pooling layer in the static object detection model to obtain the third feature map. The third feature map is input into the fully connected layer of the static object detection model, and the fully connected layer identifies the third feature map to obtain the static object identification result. The static object identification result includes the category corresponding to the static object and the probability of belonging to the category.

3. The method according to claim 1, characterized in that, Before inputting the video data and the point cloud data into the dynamic object detection model, and before the dynamic object detection model outputs the dynamic object recognition result, the method further includes: For each moment's video frame and point cloud data, the video frame and point cloud data are fused together to obtain the fusion result for the current moment; The step of inputting the video data and the point cloud data into the dynamic object detection model, and having the dynamic object detection model output the dynamic object recognition result, includes: The fusion result at the current moment is input into the sequence prediction layer of the dynamic object detection model. The sequence prediction layer predicts the fusion result at the current moment and the position information of the target object at the previous moment to obtain the movement trajectory of the target object at the next moment. The movement trajectory of the target object at the next moment is determined as the dynamic object recognition result. The output layer of the dynamic object detection model outputs the dynamic object recognition result.

4. The method according to claim 1, characterized in that, Before performing road condition identification based on the static object identification result and the dynamic object identification result to obtain the road condition prediction result, the method further includes: Acquire first data of target vehicles adjacent to the vehicle within the target area and second data of roadside units; The process of identifying road conditions based on the static object recognition results and the dynamic object recognition results to obtain road condition prediction results includes: Based on the static object recognition result, the dynamic object recognition result, the first data, and the second data, road condition recognition is performed to obtain the road condition prediction result.

5. The method according to claim 1, characterized in that, The process of identifying road conditions based on the static object recognition results and the dynamic object recognition results to obtain road condition prediction results includes: The static object recognition results and the dynamic object recognition results are comprehensively analyzed based on a preset integration strategy to obtain the road condition prediction results. The preset integration strategy includes one or more of the following: averaging method, voting method, and learning method.

6. The method according to claim 1, characterized in that, After generating the target driving route based on the road condition prediction results, the method further includes: Based on the target driving path and a preset command generation frequency, a control command is generated and sent to the controller corresponding to the control command so that the controller executes the control command. The preset command generation frequency is the frequency used to control the generation of the control command so that the vehicle's stability is greater than a preset stability threshold.

7. A path planning device, characterized in that, include: The data acquisition module is used to acquire video data and point cloud data around the vehicle; The first recognition module is used to input each video frame in the video data into the static object detection model, and the static object detection model outputs the static object recognition result. The static object detection model is used to recognize static objects in the video data. The second recognition module is used to input the video data and the point cloud data into the dynamic object detection model, and the dynamic object detection model outputs the dynamic object recognition result. The dynamic object detection model is used to recognize dynamic objects in the video data by combining the point cloud data. The third identification module is used to identify road conditions based on the static object identification results and the dynamic object identification results, and to obtain road condition prediction results. The route planning module is used to plan routes based on the road condition prediction results and generate the target driving route.

8. An electronic device, characterized in that, include: Memory; processor; as well as Computer programs; The computer program is stored in the memory and configured to be executed by the processor to implement the method as described in any one of claims 1-6.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method as described in any one of claims 1-6.

10. A vehicle, characterized in that, Includes the path planning device as described in claim 7 or the electronic device as described in claim 8.