Vision-based intelligent connected vehicle speed prediction method and device

By acquiring historical motion data of intelligent connected vehicles and road environment image data, and using a dual-branch neural network model to fuse environmental information, the problem of ignoring the environment in traditional vehicle speed prediction methods is solved, and more accurate vehicle speed prediction is achieved.

CN121459599BActive Publication Date: 2026-07-21HIGER
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
HIGER
Filing Date
2025-09-11
Publication Date
2026-07-21

AI Technical Summary

Technical Problem

Traditional vehicle speed prediction methods ignore the influence of the vehicle's surrounding environment, making it difficult to accurately predict vehicle speed.

Method used

A vision-based speed prediction method for intelligent connected vehicles is adopted. By acquiring historical motion data and road environment image data of intelligent connected vehicles, a two-branch speed prediction neural network model is used to integrate road environment image data captured by the vehicle's front camera into the speed prediction. This includes a speed feature extraction network, an environmental perception feature extraction network, and a speed prediction network to perform speed prediction.

Benefits of technology

It improves the accuracy of vehicle speed prediction and can make full use of environmental information around the vehicle to improve prediction accuracy in various scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121459599B_ABST
    Figure CN121459599B_ABST
Patent Text Reader

Abstract

The application discloses a visual-based intelligent networked vehicle speed prediction method and device, and relates to the technical field of intelligent network connection. The method comprises the following steps: acquiring historical motion data and road environment image data of an intelligent networked vehicle; determining a historical speed sequence and a historical image sequence of the intelligent networked vehicle at each time point according to the historical motion data and the road environment image data; and inputting the historical speed sequence and the historical image sequence into a preset vehicle speed prediction model to perform speed prediction, so as to obtain a predicted speed sequence of the intelligent networked vehicle. The preset vehicle speed prediction model comprises a speed feature extraction network, an environment perception feature extraction network and a speed prediction network. When performing speed prediction, the historical speed sequence and the historical image sequence are input into the speed prediction network after the features of the historical speed sequence and the historical image sequence are extracted by the speed feature extraction network and the environment perception feature extraction network. The application can improve the prediction accuracy of vehicle speed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of intelligent connected vehicle technology, and in particular to a vision-based method and apparatus for predicting the speed of intelligent connected vehicles. Background Technology

[0002] In recent years, speed prediction has played an important role in areas such as hybrid vehicle energy management systems, adaptive cruise control, and intelligent connected vehicle path planning, and is a key technology for improving energy economy and road safety.

[0003] Currently, traditional vehicle speed prediction methods typically use the vehicle's historical speed as input and output a predicted speed. However, this approach ignores the influence of the vehicle's surrounding environment on its speed, making it difficult to accurately predict vehicle speed. Summary of the Invention

[0004] In view of this, this application provides a vision-based method and device for predicting the speed of intelligent connected vehicles, which mainly improves the accuracy of predicting the future speed of vehicles.

[0005] According to a first aspect of this application, a vision-based method for predicting the speed of intelligent connected vehicles is provided, the method comprising:

[0006] Acquire historical motion data and road environment image data of intelligent connected vehicles;

[0007] Based on the historical motion data and the road environment image data, the historical speed sequence and historical image sequence of the intelligent connected vehicle at each moment are determined, wherein the historical speed sequence includes the historical lateral speed sequence and the historical longitudinal speed sequence of the intelligent connected vehicle.

[0008] The historical speed sequence and the historical image sequence are input into a preset vehicle speed prediction model for speed prediction to obtain the predicted speed sequence of the intelligent connected vehicle. The preset vehicle speed prediction model includes a speed feature extraction network, an environmental perception feature extraction network, and a speed prediction network. When performing speed prediction, the historical speed sequence and the historical image sequence are processed by the speed feature extraction network and the environmental perception feature extraction network, respectively, to extract features before being input into the speed prediction network for speed prediction.

[0009] According to a second aspect of this application, a vision-based intelligent connected vehicle speed prediction device is provided, the device comprising:

[0010] The acquisition unit is used to acquire historical motion data and road environment image data of intelligent connected vehicles;

[0011] The determining unit is configured to determine the historical speed sequence and historical image sequence of the intelligent connected vehicle at each moment based on the historical motion data and the road environment image data, wherein the historical speed sequence includes the historical lateral speed sequence and the historical longitudinal speed sequence of the intelligent connected vehicle.

[0012] The prediction unit is used to input the historical speed sequence and the historical image sequence into a preset vehicle speed prediction model to predict the speed and obtain the predicted speed sequence of the intelligent connected vehicle. The preset vehicle speed prediction model includes a speed feature extraction network, an environmental perception feature extraction network, and a speed prediction network. When performing speed prediction, the historical speed sequence and the historical image sequence are processed by the speed feature extraction network and the environmental perception feature extraction network to extract features, respectively, and then input into the speed prediction network for speed prediction.

[0013] According to a third aspect of this application, a storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the above-described vision-based intelligent connected vehicle speed prediction method.

[0014] According to a fourth aspect of this application, an electronic device is provided, including a storage medium, a processor, and a computer program stored on the storage medium and executable on the processor, wherein the processor executes the program to implement the above-described vision-based intelligent connected vehicle speed prediction method.

[0015] By means of the above technical solution, this application provides a vision-based intelligent connected vehicle speed prediction method and device. Compared with traditional vehicle speed prediction methods, by adopting a dual-branch speed prediction neural network model, it can integrate road environment image data captured by the vehicle's front camera into the speed prediction technology, thereby making full use of the environmental information around the vehicle and improving the prediction accuracy of vehicle speed in various scenarios.

[0016] The above description is only an overview of the technical solution of this application. In order to better understand the technical means of this application and to implement it in accordance with the contents of the specification, and to make the above and other objects, features and advantages of this application more obvious and understandable, the following are specific embodiments of this application. Attached Figure Description

[0017] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments of this application and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings:

[0018] Figure 1 A flowchart illustrating a vision-based intelligent connected vehicle speed prediction method provided in an embodiment of this application is shown.

[0019] Figure 2 This paper shows a schematic diagram of the structure of the preset vehicle speed prediction model provided in an embodiment of this application;

[0020] Figure 3 This paper shows a schematic diagram of the structure of the environmental perception feature extraction network provided in an embodiment of this application;

[0021] Figure 4 A flowchart illustrating the training method for a preset vehicle speed prediction model provided in an embodiment of this application is shown.

[0022] Figure 5 A schematic diagram of the structure of a vision-based intelligent connected vehicle speed prediction device provided in an embodiment of this application is shown. Detailed Implementation

[0023] The present application will be described in detail below with reference to the accompanying drawings and embodiments. It should be noted that, unless otherwise specified, the embodiments and features described in the embodiments of the present application can be combined with each other.

[0024] Traditional vehicle speed prediction methods ignore the influence of the vehicle's surrounding environment on its speed, making it difficult to accurately predict vehicle speed.

[0025] To address the aforementioned problems, embodiments of the present invention provide a vision-based method for predicting the speed of intelligent connected vehicles, such as... Figure 1 As shown, the method includes:

[0026] Step 101: Obtain historical motion data and road environment image data of intelligent connected vehicles.

[0027] The historical motion data includes the lateral speed, longitudinal speed, and heading angle of the intelligent connected vehicle, while the road environment image data consists of video image data that includes road environment information.

[0028] The embodiments of the present invention are mainly applicable to the speed prediction scenario of intelligent connected vehicles. The execution subject of the embodiments of the present invention is a device or equipment capable of predicting vehicle speed, such as an edge computing unit.

[0029] In this embodiment of the invention, during the driving process of an intelligent connected vehicle, continuous motion data of the vehicle in the world coordinate system can be obtained via satellite, specifically including the lateral velocity of the intelligent connected vehicle. Longitudinal velocity and heading angle At the same time, video image data containing road environment information, i.e., road environment image data, can be collected through the front-facing camera of intelligent connected vehicles.

[0030] Step 102: Based on the historical motion data and the road environment image data, determine the historical speed sequence and historical image sequence of the intelligent connected vehicle at each moment.

[0031] The historical speed sequence includes the historical lateral speed sequence and the historical longitudinal speed sequence of the intelligent connected vehicle.

[0032] In this embodiment of the invention, after collecting historical motion data and road environment image data of intelligent connected vehicles, it is necessary to determine the historical speed sequence and historical image sequence of the intelligent connected vehicles. Specifically, step 102 includes: converting the continuous lateral and longitudinal velocities of the intelligent connected vehicles in the world coordinate system from the historical motion data into continuous lateral and longitudinal velocities in the vehicle coordinate system; sampling the continuous lateral and longitudinal velocities in the vehicle coordinate system to obtain the historical lateral speed sequence and the historical longitudinal speed sequence of the intelligent connected vehicles; extracting target area image data related to road information from the road environment image data; and sequentially cropping and sampling the target area image data to obtain the historical image sequence of the intelligent connected vehicles.

[0033] Specifically, suppose that at a certain time, the lateral velocity of the intelligent connected vehicle in the world coordinate system is... Longitudinal velocity is heading angle is Convert it into lateral velocity in the vehicle coordinate system and longitudinal velocity The specific formula is as follows:

[0034]

[0035] Therefore, the continuous lateral and longitudinal velocities of the intelligent connected vehicle in the vehicle coordinate system can be obtained according to the above formula. Then, these velocities are sampled. Assuming that the historical time length of the lateral and longitudinal velocities at each moment is 4 seconds and the sampling frequency is 2Hz, the number of data points obtained is 8. These 8 lateral velocities and 8 longitudinal velocities in the vehicle coordinate system are respectively used as the historical lateral velocity sequence and the historical longitudinal velocity sequence.

[0036] At the same time, target area image data related to road information is extracted from road environment image data, and the image size is cropped, such as the cropped image size is 160×160. Then, at each moment, sampling is performed at the same sampling frequency as the historical motion data, and the sampled historical image data is aligned with the historical motion data to finally obtain the historical image sequence of intelligent connected vehicles.

[0037] Step 103: Input the historical speed sequence and the historical image sequence into a preset vehicle speed prediction model to predict the speed and obtain the predicted speed sequence of the intelligent connected vehicle.

[0038] The preset vehicle speed prediction model includes a speed feature extraction network, an environmental perception feature extraction network, and a speed prediction network. When performing speed prediction, the historical speed sequence and the historical image sequence are processed by the speed feature extraction network and the environmental perception feature extraction network to extract features, respectively, and then input into the speed prediction network for speed prediction.

[0039] In this embodiment of the invention, when performing speed prediction, the historical speed sequence is input into the speed feature extraction network for feature extraction to obtain the speed feature vector corresponding to the intelligent connected vehicle; the historical image sequence is input into the environmental perception feature extraction network for feature extraction to obtain the environmental perception feature vector corresponding to the intelligent connected vehicle; the speed feature vector and the environmental perception feature vector are input into the speed prediction network for speed prediction to obtain the predicted speed sequence of the intelligent connected vehicle.

[0040] As can be seen from the above speed prediction process, in order to improve the prediction accuracy of vehicle speed, the embodiments of the present invention employ a dual-branch preset vehicle speed prediction model, such as... Figure 2 As shown, the preset vehicle speed prediction model can be a neural network model, with two branches: a horizontal and vertical speed branch and an environmental perception branch. The horizontal and vertical speed branches correspond to a speed feature extraction network, and the environmental perception branch corresponds to an environmental perception feature extraction network. The speed feature extraction network can be an LSTM (Long Short-Term Memory) network, or an RNN (Recursive Neural Network), RBFNN (Radial Basis Function Neural Network), BPNN (Back Propagation Neural Network), etc. This embodiment of the invention does not specifically limit the type of speed feature extraction network. The environmental perception feature extraction network can be an Inception network, or a CNN (Convolutional Neural Network) or other image processing networks, but the convolutional kernels in their convolutional layers must be expanded from 2D to 3D to process temporal information.

[0041] Furthermore, the speed feature extraction network in this embodiment of the invention includes a lateral speed feature extraction network and a longitudinal speed feature extraction network. During speed feature extraction, the historical lateral speed sequence is input into the lateral speed feature extraction network to extract lateral speed features, obtaining a lateral speed feature vector corresponding to the intelligent connected vehicle; the historical longitudinal speed sequence is input into the longitudinal speed feature extraction network to extract longitudinal speed features, obtaining a longitudinal speed feature vector corresponding to the intelligent connected vehicle; and the speed feature vector corresponding to the intelligent connected vehicle is determined based on the lateral speed feature vector and the longitudinal speed feature vector.

[0042] Specifically, to further improve the prediction accuracy of vehicle speed, this embodiment of the invention employs two speed feature extraction networks to extract features from historical lateral speed sequences and historical longitudinal speed sequences, respectively. This encodes the historical lateral and longitudinal speed sequences of the intelligent connected vehicle, generating a feature vector of vehicle motion. The lateral and longitudinal speed feature extraction networks can be LSTM networks, or deep neural networks such as RNNs, RBFNNs, and BPNNs; this embodiment of the invention does not impose specific limitations on these.

[0043] For example, embodiments of the present invention employ two two-layer LSTM networks as the horizontal velocity feature extraction network and the vertical velocity feature extraction network, respectively, such as... Figure 2 As shown. Since the two two-layer LSTM structures are identical, taking the LSTM network that processes historical longitudinal velocity sequences as an example, for historical 4s longitudinal velocity data, assuming the sampling frequency is 2Hz, the historical longitudinal velocity sequence input to the two-layer LSTM network has 8 data points. After processing by the two-layer LSTM network, the historical longitudinal velocity sequence yields a 128-dimensional longitudinal velocity feature vector.

[0044] For environment-aware feature extraction networks, RGB images are stacked to form a historical image sequence, which is then input into the network to extract environment-aware feature vectors. These networks can employ deep neural networks such as CNNs and Inception to process images. Since CNNs and Inception can only extract features from a single image, when processing image sequences with an additional temporal dimension, the convolutional kernels of the deep neural network need to be dilated from 2D to 3D to process 3D image sequence information. Taking the Inception network as an example... Figure 2As shown, the kernel size of the convolutional layers and the pooling size of the pooling layers in the original Inception network are expanded from 2D to 3D to obtain the 3D Inception network. This allows it to process image sequences composed of stacked RGB images, ultimately generating a high-dimensional vector of 1024 dimensions. This embodiment of the invention adds a temporal dimension to the 3D Inception network, enabling it to process image information from spatiotemporal data while maintaining good 2D image information extraction capabilities.

[0045] The specific processing steps of the 3D Inception network are as follows: Figure 3 As shown, the input data structure of the 3D Inception network is (T, W, H, C), where C is the number of channels (typically 3 channels for RGB images); W and H are the width and height of the image, respectively (e.g., both width and height are 160); T is the length of the historical image sequence, aligned with the historical vertical and horizontal velocity sequences, also 4 seconds, with 8 data points. During feature extraction, the historical image sequence is first input into a 3D convolutional layer and a 3D max-pooling layer to downsample the image data spatially and temporally. Then, nine Inception modules are used to further extract spatiotemporal information from the image data. Finally, the vector output from the last Inception module is input into a 3D average pooling layer for processing, resulting in an environment-aware feature vector containing spatiotemporal information.

[0046] Therefore, in the environmental perception branch of this invention, the 2D convolutional kernels of the Inception network are expanded to 3D convolutional kernels, enabling the Inception network to process image sequence information and ultimately generate high-dimensional feature vectors of the surrounding environment. Compared to 3D CNN networks, using a 3D Inception network can effectively reduce training parameters, lower memory and computational costs, and is more suitable for training high-dimensional image information with time series data.

[0047] Furthermore, the speed prediction network includes feature fusion, a first fully connected layer network, and a second fully connected layer network. During speed prediction, the speed feature vector and the environmental perception feature vector are horizontally concatenated to obtain a fused feature vector. The fused feature vector is then input into the first and second fully connected layer networks respectively for speed prediction, resulting in a predicted lateral speed sequence and a predicted longitudinal speed sequence. Based on the predicted lateral speed sequence and the predicted longitudinal speed sequence, the predicted speed sequence of the intelligent connected vehicle is determined.

[0048] Specifically, the three feature vectors from the horizontal and vertical velocity branches and the environment perception branch are concatenated horizontally to obtain a 1280-dimensional vector. Finally, the predicted horizontal velocity sequence and the predicted vertical velocity sequence are output through two double-layer fully connected layers, such as the predicted horizontal velocity sequence and the predicted vertical velocity sequence for the next 8 seconds.

[0049] Furthermore, embodiments of the present invention also provide a training method for a preset vehicle speed prediction model, such as... Figure 4 As shown, it includes:

[0050] Step 104: Obtain the sample historical speed sequence, sample historical image sequence, and real future speed sequence of the intelligent connected vehicle.

[0051] In this embodiment of the invention, the lateral velocity of the intelligent connected vehicle in the world coordinate system is first collected. and sample longitudinal velocity Then, the lateral velocity of the sample in the world coordinate system is calculated according to formulas (1) and (2). and sample longitudinal velocity Transformed into sample lateral velocity in the vehicle coordinate system and sample longitudinal velocity Next, the continuous horizontal and vertical velocity data are sliced, that is, the M-second continuous velocity data is divided into M-N+1 N-second slices, and the historical time range and prediction time range of each slice are set. For example, each slice is 12 seconds long, with the historical data time range of 4 seconds and the prediction data time range of 8 seconds. Then, each slice is sampled at a certain sampling frequency (e.g., 2Hz) to obtain the sample historical horizontal velocity sequence and the sample historical vertical velocity sequence.

[0052] Simultaneously, target area image data related to road information is extracted from the sample road environment image data, and the image size is re-cropped, for example, to 160×160. Then, the sample image data is sliced, matching the sample lateral and longitudinal velocities, and aligned with the historical lateral and longitudinal velocity sequences. Next, the image slice data is sampled to obtain the historical image sequence, ensuring that the sampling frequency of the sample image data is the same as the sampling frequency of the sample velocity data.

[0053] Step 105: Construct an initial velocity prediction network, and input the sample historical velocity sequence and the sample historical image sequence into the initial velocity prediction network to perform velocity prediction, thereby obtaining the sample predicted velocity sequence.

[0054] The sample prediction velocity sequence includes the sample horizontal prediction velocity sequence and the sample vertical prediction velocity sequence. The initial velocity prediction network includes the initial velocity feature extraction network, the initial environment perception feature extraction network, and the initial velocity prediction network. The initial velocity feature extraction network can be a deep neural network such as LSTM, RNN, RBF, or BP. The initial environment perception feature extraction network can be a 3D deep neural network for image processing such as CNN or Inception. The initial velocity prediction network includes feature fusion, a first fully connected layer network, and a second fully connected layer network.

[0055] Let the historical lateral velocity sequence of the sample be... The historical longitudinal velocity sequence of the sample is The historical image sequence of the sample is Where Δt is the sampling interval, k is an integer, and k≥1. The intelligent connected vehicle in the future time t... f (t f The sample lateral prediction velocity sequence corresponding to >t0) is And the longitudinal prediction velocity sequence of the sample is f p From the sample history lateral velocity sequence Sample historical longitudinal velocity sequence and sample historical image sequences To the sample lateral prediction velocity sequence and sample longitudinal prediction velocity sequence The mapping is determined by the parameter θ, and the specific formula is expressed as:

[0056]

[0057] Step 106: Calculate the mean absolute error between the predicted velocity sequence of the sample and the actual future velocity sequence, and construct a loss function based on the mean absolute error.

[0058] In the embodiments of the present invention, after obtaining the sample lateral predicted velocity sequence and sample longitudinal prediction velocity sequence Subsequently, in order to find the optimal solution for the parameters θ of the deep neural network based on the two branches, this embodiment of the invention employs the gradient descent algorithm to minimize the MAE (Mean absolute error) between the predicted velocity sequence of the samples and the true future velocity sequence. The specific expression of the corresponding loss function is as follows:

[0059]

[0060] in, and These are the actual future horizontal velocity sequence and the vertical velocity sequence, respectively. and These are the predicted future horizontal velocity sequence and vertical velocity sequence, respectively, and N is the total number of sample slices.

[0061] Step 107: Based on the loss function, iteratively train the initial speed prediction network to construct the preset vehicle speed prediction model.

[0062] In this embodiment of the invention, after constructing the loss function, the parameter θ is updated iteratively using the gradient descent algorithm based on the loss function, as shown in the following formula:

[0063]

[0064] Where, θ m θ represents the parameter values ​​of the deep neural network in the current iteration. m+1 η represents the parameter values ​​of the deep neural network for the next iteration; η is the learning rate.

[0065] The networks designed in this embodiment of the invention are all trained using the Adam optimizer, with an initial learning efficiency of 0.0003. If the training loss of the validation formula (4) does not decrease within 10 epochs (one complete training cycle), the learning rate of formula (5) is increased by 0.5 times the previous learning rate. The batch size is set to 32, and the training epochs are set to 500. To avoid overfitting, training is terminated if the training loss of the validation formula (4) does not decrease further within 30 epochs. The model with the minimum training loss on the validation set is saved and used for further testing.

[0066] This invention provides a vision-based intelligent connected vehicle speed prediction method. Compared with traditional vehicle speed prediction methods, this method uses a dual-branch speed prediction neural network model to integrate road environment image data captured by the vehicle's front-facing camera into the speed prediction technology. This allows for full utilization of the environmental information around the vehicle, improving the accuracy of vehicle speed prediction in various scenarios.

[0067] Furthermore, as Figure 1 The specific implementation of the method shown in this embodiment provides a vision-based intelligent connected vehicle speed prediction device, such as... Figure 5 As shown, the device includes: an acquisition unit 31, a determination unit 32, and a prediction unit 33.

[0068] The acquisition unit 31 can be used to acquire historical motion data and road environment image data of intelligent connected vehicles.

[0069] The determining unit 32 can be used to determine the historical speed sequence and historical image sequence of the intelligent connected vehicle at each moment based on the historical motion data and the road environment image data, respectively. The historical speed sequence includes the historical lateral speed sequence and the historical longitudinal speed sequence of the intelligent connected vehicle.

[0070] The prediction unit 33 can be used to input the historical speed sequence and the historical image sequence into a preset vehicle speed prediction model to predict the speed and obtain the predicted speed sequence of the intelligent connected vehicle. The preset vehicle speed prediction model includes a speed feature extraction network, an environmental perception feature extraction network, and a speed prediction network. When performing speed prediction, the historical speed sequence and the historical image sequence are processed by the speed feature extraction network and the environmental perception feature extraction network to extract features, respectively, and then input into the speed prediction network for speed prediction.

[0071] In some embodiments, the determining unit 32 may be specifically configured to convert the continuous lateral and longitudinal velocities of the intelligent connected vehicle in the world coordinate system in the historical motion data into continuous lateral and longitudinal velocities in the vehicle coordinate system; sample the continuous lateral and longitudinal velocities in the vehicle coordinate system to obtain the historical lateral velocity sequence and the historical longitudinal velocity sequence of the intelligent connected vehicle; extract target area image data related to road information from the road environment image data; and sequentially crop and sample the target area image data to obtain the historical image sequence of the intelligent connected vehicle.

[0072] In some embodiments, the prediction unit 33 includes: a first extraction module, a second extraction module, and a prediction module.

[0073] The first extraction module can be used to input the historical speed sequence into the speed feature extraction network for feature extraction to obtain the speed feature vector corresponding to the intelligent connected vehicle.

[0074] The second extraction module can be used to input the historical image sequence into the environmental perception feature extraction network for feature extraction, and obtain the environmental perception feature vector corresponding to the intelligent connected vehicle.

[0075] The prediction module can be used to input the speed feature vector and the environmental perception feature vector into the speed prediction network to predict the speed and obtain the predicted speed sequence of the intelligent connected vehicle.

[0076] In some embodiments, the speed feature extraction network includes a lateral speed feature extraction network and a longitudinal speed feature extraction network. The first extraction module can be specifically used to input the historical lateral speed sequence into the lateral speed feature extraction network to extract lateral speed features, thereby obtaining a lateral speed feature vector corresponding to the intelligent connected vehicle; input the historical longitudinal speed sequence into the longitudinal speed feature extraction network to extract longitudinal speed features, thereby obtaining a longitudinal speed feature vector corresponding to the intelligent connected vehicle; and determine the speed feature vector corresponding to the intelligent connected vehicle based on the lateral speed feature vector and the longitudinal speed feature vector.

[0077] In some embodiments, the environment-aware feature extraction network includes a 3D Inception network, which is obtained by expanding the convolution kernel size of the convolutional layer and the pooling size of the pooling layer in the original Inception network from 2D to 3D.

[0078] In some embodiments, the speed prediction network includes feature fusion, a first fully connected layer network, and a second fully connected layer network. The prediction module can be specifically used to laterally concatenate the speed feature vector and the environmental perception feature vector to obtain a fused feature vector; input the fused feature vector into the first fully connected layer network and the second fully connected layer network respectively for speed prediction to obtain a predicted lateral speed sequence and a predicted longitudinal speed sequence; and determine the predicted speed sequence of the intelligent connected vehicle based on the predicted lateral speed sequence and the predicted longitudinal speed sequence.

[0079] In some embodiments, the apparatus further includes a construction unit.

[0080] The construction unit can be used to acquire the sample historical speed sequence, sample historical image sequence, and real future speed sequence of the intelligent connected vehicle; construct an initial speed prediction network, and input the sample historical speed sequence and the sample historical image sequence into the initial speed prediction network to perform speed prediction, thereby obtaining a sample predicted speed sequence; calculate the mean absolute error between the sample predicted speed sequence and the real future speed sequence, and construct a loss function based on the mean absolute error; and iteratively train the initial speed prediction network based on the loss function to construct the preset vehicle speed prediction model.

[0081] It should be noted that other corresponding descriptions of the functional units involved in the vision-based intelligent connected vehicle speed prediction device provided in this embodiment can be found in [reference needed]. Figure 1 The corresponding description in [the document] will not be repeated here.

[0082] Based on the above, Figure 1Accordingly, this embodiment also provides a storage medium storing a computer program that, when executed by a processor, implements the above-described method. Figure 1 The method shown is a vision-based method for predicting the speed of intelligent connected vehicles.

[0083] Based on this understanding, the technical solution of this application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as CD-ROM, USB flash drive, mobile hard drive, etc.) and includes several instructions to cause an electronic device (such as a personal computer, server, or network device, etc.) to execute the methods of various implementation scenarios of this application.

[0084] Based on the above, Figure 1 The method shown, and Figure 5 To achieve the above objectives, the present application also provides an electronic device, specifically a personal computer, tablet computer, server, or other network device, as shown in the virtual device embodiment. This device includes a storage medium and a processor; the storage medium stores a computer program; the processor executes the computer program to achieve the above-described objectives. Figure 1 The method shown is a vision-based method for predicting the speed of intelligent connected vehicles.

[0085] Optionally, the aforementioned physical devices may also include a user interface, a network interface, a camera, radio frequency (RF) circuitry, sensors, audio circuitry, a Wi-Fi module, etc. The user interface may include a display screen, input units such as a keyboard, etc., and optional user interfaces may also include USB interfaces, card reader interfaces, etc. The network interface may optionally include standard wired interfaces, wireless interfaces (such as Wi-Fi interfaces), etc.

[0086] Those skilled in the art will understand that the physical device structure provided in this embodiment does not constitute a limitation on the physical device, and may include more or fewer components, or combine certain components, or have different component arrangements.

[0087] The storage medium may also include an operating system and a network communication module. The operating system is a program that manages the hardware and software resources of the aforementioned physical device, supporting the operation of information processing programs and other software and / or programs. The network communication module is used to enable communication between the various components within the storage medium, as well as communication with other hardware and software in the information processing physical device.

[0088] Through the above description of the embodiments, those skilled in the art can clearly understand that this application can be implemented by means of software plus necessary general-purpose hardware platform, or it can be implemented by hardware.

[0089] This invention employs a dual-branch speed prediction neural network model, which integrates road environment image data captured by the vehicle's front-facing camera into the speed prediction technology. This allows for full utilization of the environmental information surrounding the vehicle, improving the accuracy of vehicle speed prediction in various scenarios.

[0090] Those skilled in the art will understand that the accompanying drawings are merely schematic diagrams of a preferred embodiment, and the modules or processes shown in the drawings are not necessarily essential for implementing this application. Those skilled in the art will understand that the modules in the apparatus of the embodiment can be distributed within the apparatus of the embodiment as described, or can be modified to be located in one or more apparatuses different from this embodiment. The modules of the above-described embodiment can be combined into one module, or further divided into multiple sub-modules.

[0091] The serial numbers in this application are for descriptive purposes only and do not represent the superiority or inferiority of any particular implementation scenario. The above disclosures are merely a few specific implementation scenarios of this application; however, this application is not limited thereto, and any variations conceived by those skilled in the art should fall within the protection scope of this application.

Claims

1. A vision-based method for predicting the speed of intelligent connected vehicles, characterized in that, include: Acquire historical motion data and road environment image data of intelligent connected vehicles; Based on the historical motion data and the road environment image data, the historical speed sequence and historical image sequence of the intelligent connected vehicle at each moment are determined, wherein the historical speed sequence includes the historical lateral speed sequence and the historical longitudinal speed sequence of the intelligent connected vehicle. The historical speed sequence and the historical image sequence are input into a preset vehicle speed prediction model to predict the speed, thereby obtaining the predicted speed sequence of the intelligent connected vehicle. The preset vehicle speed prediction model includes a speed feature extraction network, an environmental perception feature extraction network, and a speed prediction network. The step of inputting the historical speed sequence and the historical image sequence into a preset vehicle speed prediction model to predict the speed and obtain the predicted speed sequence of the intelligent connected vehicle includes: The historical speed sequence is input into the speed feature extraction network for feature extraction to obtain the speed feature vector corresponding to the intelligent connected vehicle. The historical image sequence is input into the environmental perception feature extraction network for feature extraction to obtain the environmental perception feature vector corresponding to the intelligent connected vehicle. The environmental perception feature extraction network includes a 3D Inception network. The convolution kernel size of the convolutional layer and the pooling size of the pooling layer in the original Inception network are expanded from 2D to 3D to obtain the 3D Inception network. The speed feature vector and the environmental perception feature vector are input into the speed prediction network to predict the speed, thereby obtaining the predicted speed sequence of the intelligent connected vehicle. The speed prediction network includes feature fusion, a first fully connected layer network, and a second fully connected layer network. The step of inputting the speed feature vector and the environmental perception feature vector into the speed prediction network to perform speed prediction and obtain the predicted speed sequence of the intelligent connected vehicle includes: The velocity feature vector and the environment perception feature vector are horizontally concatenated to obtain a fused feature vector. The fused feature vectors are input into the first fully connected layer network and the second fully connected layer network respectively to perform velocity prediction, resulting in a predicted horizontal velocity sequence and a predicted vertical velocity sequence. The predicted speed sequence of the intelligent connected vehicle is determined based on the predicted lateral speed sequence and the predicted longitudinal speed sequence.

2. The method according to claim 1, characterized in that, The step of determining the historical speed sequence and historical image sequence of the intelligent connected vehicle at each moment based on the historical motion data and the road environment image data includes: The continuous lateral and longitudinal velocities of the intelligent connected vehicle in the world coordinate system in the historical motion data are converted into continuous lateral and longitudinal velocities in the vehicle coordinate system. The continuous lateral and longitudinal velocities in the vehicle coordinate system are sampled to obtain the historical lateral velocity sequence and the historical longitudinal velocity sequence of the intelligent connected vehicle. Extract target area image data related to road information from the road environment image data; The target area image data is cropped and sampled sequentially to obtain the historical image sequence of the intelligent connected vehicle.

3. The method according to claim 1, characterized in that, The speed feature extraction network includes a horizontal speed feature extraction network and a vertical speed feature extraction network. The historical speed sequence is input into the speed feature extraction network for feature extraction to obtain the speed feature vector corresponding to the intelligent connected vehicle, including: The historical lateral velocity sequence is input into the lateral velocity feature extraction network to extract lateral velocity features, thereby obtaining the lateral velocity feature vector corresponding to the intelligent connected vehicle. The historical longitudinal velocity sequence is input into the longitudinal velocity feature extraction network to extract longitudinal velocity features, thereby obtaining the longitudinal velocity feature vector corresponding to the intelligent connected vehicle. Based on the lateral velocity feature vector and the longitudinal velocity feature vector, the velocity feature vector corresponding to the intelligent connected vehicle is determined.

4. The method according to any one of claims 1-3, characterized in that, Before inputting the historical speed sequence and the historical image sequence into a preset vehicle speed prediction model to predict the speed and obtain the predicted speed sequence of the intelligent connected vehicle, the method further includes: Obtain the sample historical speed sequence, sample historical image sequence, and real future speed sequence of the intelligent connected vehicle; An initial velocity prediction network is constructed, and the sample historical velocity sequence and the sample historical image sequence are input into the initial velocity prediction network to perform velocity prediction, thereby obtaining the sample predicted velocity sequence. Calculate the mean absolute error between the predicted velocity sequence of the sample and the actual future velocity sequence, and construct a loss function based on the mean absolute error; Based on the loss function, the initial speed prediction network is iteratively trained to construct the preset vehicle speed prediction model.

5. A vision-based intelligent connected vehicle speed prediction device, characterized in that, include: The acquisition unit is used to acquire historical motion data and road environment image data of intelligent connected vehicles; The determining unit is configured to determine the historical speed sequence and historical image sequence of the intelligent connected vehicle at each moment based on the historical motion data and the road environment image data, wherein the historical speed sequence includes the historical lateral speed sequence and the historical longitudinal speed sequence of the intelligent connected vehicle. The prediction unit is used to input the historical speed sequence and the historical image sequence into a preset vehicle speed prediction model to predict the speed and obtain the predicted speed sequence of the intelligent connected vehicle. The preset vehicle speed prediction model includes a speed feature extraction network, an environmental perception feature extraction network, and a speed prediction network. The prediction unit includes: a first extraction module, a second extraction module, and a prediction module. The first extraction module is used to input the historical speed sequence into the speed feature extraction network for feature extraction to obtain the speed feature vector corresponding to the intelligent connected vehicle; The second extraction module is used to input the historical image sequence into the environmental perception feature extraction network for feature extraction to obtain the environmental perception feature vector corresponding to the intelligent connected vehicle. The environmental perception feature extraction network includes a 3D Inception network, which expands the convolution kernel size of the convolutional layer and the pooling size of the pooling layer in the original Inception network from 2D to 3D to obtain the 3D Inception network. The prediction module is used to input the speed feature vector and the environmental perception feature vector into the speed prediction network to predict the speed and obtain the predicted speed sequence of the intelligent connected vehicle. The speed prediction network includes feature fusion, a first fully connected layer network and a second fully connected layer network. The prediction module is specifically used to horizontally concatenate the speed feature vector and the environmental perception feature vector to obtain a fused feature vector; input the fused feature vector into the first fully connected layer network and the second fully connected layer network respectively for speed prediction to obtain a predicted horizontal speed sequence and a predicted vertical speed sequence; and determine the predicted speed sequence of the intelligent connected vehicle based on the predicted horizontal speed sequence and the predicted vertical speed sequence.

6. A storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method of any one of claims 1 to 4.

7. An electronic device comprising a storage medium, a processor, and a computer program stored on the storage medium and executable on the processor, characterized in that, When the processor executes the computer program, it implements the method of any one of claims 1 to 4.

Citation Information

Patent Citations

  • Short-term vehicle speed working condition real-time prediction method based on interaction between preceding vehicle and self vehicle

    CN111009134A