Driving data generation method and device, electronic equipment and readable storage medium
By encoding and predicting the driving data in time and space, generating feature marks at the second moment and decoding them into driving data, the problem of difficulty in realizing the joint generation of multimodal driving data in the prior art is solved, and the accuracy of driving data and the safety and reliability of the system are improved.
Patent Information
- Application Number
- CN202510323382.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-18
- Publication Date
- 2025-06-10
AI Technical Summary
The prior art is difficult to realize the joint generation of multimodal driving data, and cannot meet the safety and reliability requirements of intelligent driving systems in complex driving environments.
By acquiring the first driving data of at least one first moment, encoding it to obtain the feature identifier, and generating the feature identifier of the second moment based on the time and space prediction, finally decoding to obtain the driving data of the second moment, and achieving joint generation of multimodal data.
Through spatial and temporal prediction, the driving sub-data of different modes are effectively fused, and the joint generation of driving data of multiple modes is realized, which improves the accuracy of the driving sub-data at the second moment.
Smart Images

Figure CN120126237A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of autonomous driving technology, and in particular, to a method, an apparatus, an electronic device, and a readable storage medium for generating driving data. Background Art
[0002] Intelligent driving systems rely on large-scale and diverse driving data (such as ego-vehicle decision-making data, traffic participant data, map data, ego-vehicle perception data, etc.) for training and testing to ensure their safety and reliability in complex driving environments. Therefore, how to achieve the joint generation of large-scale and diverse multi-modal driving data has become a technical problem that urgently needs to be solved. Summary of the Invention
[0003] To solve the above technical problems, the present disclosure provides a method, an apparatus, an electronic device, and a readable storage medium for generating driving data to achieve the joint generation of multi-modal driving data.
[0004] In a first aspect of the present disclosure, a method for generating driving data is provided, including: obtaining first driving data at at least one first moment, where the first driving data includes at least one type of first driving sub-data in at least one modality; encoding the first driving sub-data to obtain a first feature identifier corresponding to the first driving sub-data; performing spatio-temporal domain prediction on the first feature identifier to obtain a second feature identifier corresponding to second driving sub-data at a second moment, where the first moment is before the second moment; decoding the second feature identifier to obtain the second driving sub-data; and determining second driving data at the second moment based on the second driving sub-data.
[0005] In a second aspect of the present disclosure, a device for generating driving data is provided, including: a driving data acquisition module for obtaining first driving data at at least one first moment, where the first driving data includes at least one type of first driving sub-data in at least one modality; a driving data encoding module for encoding the first driving sub-data to obtain a first feature identifier corresponding to the first driving sub-data; a feature identifier prediction module for performing spatio-temporal domain prediction on the first feature identifier to obtain a second feature identifier corresponding to second driving sub-data at a second moment, where the first moment is before the second moment; a feature identifier decoding module for decoding the second feature identifier to obtain the second driving sub-data; and a driving data generation module for determining second driving data at the second moment based on the second driving sub-data.
[0006] In a third aspect of the present disclosure, a computer-readable storage medium is provided, where the storage medium stores a computer program for executing the method for generating driving data provided in the first aspect above.
[0007] In a fourth aspect of the present disclosure, an electronic device is provided, which includes: a processor; a memory for storing executable instructions executable by the processor; the processor is configured to read the executable instructions from the memory and execute the instructions to implement the driving data generation method provided in the first aspect above.
[0008] In an embodiment of a fifth aspect of the present disclosure, a computer program product is proposed. When the instruction processor in the computer program product executes, it executes the driving data generation method provided in the first aspect of the present disclosure.
[0009] In the embodiments of the present disclosure, through spatio-temporal domain prediction, different modalities of driving sub-data can be effectively fused, realizing the joint generation of multiple modalities of driving data, and improving the accuracy of the predicted driving sub-data at the second moment. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] Figure 1 It is a schematic flowchart of a driving data generation method provided by an exemplary embodiment of the present disclosure.
[0011] Figure 2 It is a schematic flowchart of a driving data generation method provided by another exemplary embodiment of the present disclosure.
[0012] Figure 3 It is a schematic flowchart of a driving data generation method provided by another exemplary embodiment of the present disclosure.
[0013] Figure 4 It is a schematic diagram of spatio-temporal feature prediction provided by another exemplary embodiment of the present disclosure.
[0014] Figure 5 It is a schematic flowchart of a driving data generation method provided by another exemplary embodiment of the present disclosure.
[0015] Figure 6 It is a schematic flowchart of a driving data generation method provided by another exemplary embodiment of the present disclosure.
[0016] Figure 7 It is a schematic diagram of map data provided by another exemplary embodiment of the present disclosure.
[0017] Figure 8 It is a schematic flowchart of a driving data generation method provided by another exemplary embodiment of the present disclosure.
[0018] Figure 9 It is a schematic structural diagram of a driving data generation device provided by an exemplary embodiment of the present disclosure.
[0019] Figure 10It is a schematic structural diagram of a driving data generation device provided by another exemplary embodiment of the present disclosure.
[0020] Figure 11 It is a schematic structural diagram of a driving data generation device provided by another exemplary embodiment of the present disclosure.
[0021] Figure 12 It is a schematic structural diagram of a driving data generation device provided by another exemplary embodiment of the present disclosure.
[0022] Figure 13 It is a schematic structural diagram of a driving data generation device provided by another exemplary embodiment of the present disclosure.
[0023] Figure 14 It is a structural diagram of an electronic device provided by an exemplary embodiment of the present disclosure. Detailed implementation manners
[0024] To explain the present disclosure, exemplary embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present disclosure, rather than all embodiments. It should be understood that the present disclosure is not limited by the exemplary embodiments.
[0025] It should be noted that: Unless otherwise specifically stated, the relative arrangements of components and steps, numerical expressions and values set forth in these embodiments do not limit the scope of the present disclosure.
[0026] Application overview
[0027] The intelligent driving described in the present disclosure can cover multiple fields such as autonomous driving, assisted driving, and robot systems. Autonomous driving technology aims to achieve the fully autonomous driving of vehicles in various complex road conditions without human intervention, and is also known as driverless driving, which is the advanced form of intelligent driving. Assisted driving provides real-time road condition information, warnings, and partial driving operation support for drivers through a series of sensors and algorithms, such as automatic parking, adaptive cruise control, etc., aiming to improve the safety and convenience of driving. The robot system further extends the intelligent driving technology to fields such as service robots and industrial robots, enabling robots to autonomously navigate, avoid obstacles, and complete specific tasks in complex environments, such as logistics distribution, warehouse management, etc., demonstrating the broad application potential of intelligent driving technology in different scenarios.
[0028] The autonomous driving system relies on large-scale and diverse driving data (such as ego vehicle decision-making data, traffic participant data, map data, ego vehicle perception data, etc.) for training and testing to ensure its safety and reliability in complex driving environments.
[0029] For different modalities of driving data required by an autonomous driving system, different deep learning models are usually used for generation currently. For example, a denoising model is used to generate image data in the ego-vehicle perception data, and a generation model based on Transformer (a deep learning model architecture based on the attention mechanism) is used to generate traffic participant data. Therefore, the current driving data generation technology can only realize the generation of single-modal driving data and cannot realize the joint generation of multiple-modal driving data.
[0030] To solve the above technical problems, the embodiments of the present disclosure provide a driving data generation method, which obtains first driving data at at least one first moment, encodes the first driving sub-data to obtain a first feature identifier corresponding to the first driving sub-data, performs spatio-temporal domain prediction on the first feature identifier to obtain a second feature identifier corresponding to the second driving sub-data at a second moment, decodes the second feature identifier to obtain the second driving sub-data, and determines the second driving data at the second moment based on the second driving sub-data. In this way, through spatio-temporal domain prediction, different-modal driving sub-data can be effectively fused, and further the joint generation of multiple-modal driving data can be realized.
[0031] Exemplary method
[0032] Figure 1 is a schematic flowchart of a driving data generation method provided by an exemplary embodiment of the present disclosure. This embodiment can be applied to an electronic device, such as Figure 1 shown, and includes the following steps:
[0033] Step 101, obtain first driving data at at least one first moment. The first driving data includes first driving sub-data of at least one modality.
[0034] Exemplarily, the electronic device can obtain perception data collected by a vehicle through sensors (such as an image sensor, a radar sensor, an inertial sensor, etc.), and obtain the first driving data at the first moment based on the perception data. The electronic device can also obtain publicly available and authorized-downloadable driving data on the network and obtain the first driving data at the first moment therefrom.
[0035] Further, the electronic device can classify the first driving data into first driving sub-data of different modalities according to classification methods such as data type, data source, data subject, and data function. For example, according to data function, the electronic device can classify the first driving data into four modalities of first driving sub-data: ego-vehicle decision-making data, traffic participant data, map data, and ego-vehicle perception data. For another example, according to data subject, the electronic device can classify the first driving data into three modalities of first driving sub-data: ego-vehicle subject data, road user subject data, and environmental subject data.
[0036] Of course, those skilled in the art can also divide the first driving data into first driving sub-data of other modalities based on any classification method, and the embodiments of the present disclosure do not make limitations. Taking the division of the first driving data into first driving sub-data of four modalities: ego vehicle decision-making data, traffic participant data, map data, and ego vehicle perception data as an example, the other situations are similar, and the embodiments of the present disclosure will not be elaborated further.
[0037] Step 102: Encode the first driving sub-data to obtain a first feature identifier corresponding to the first driving sub-data.
[0038] Exemplarily, the data types and complexities of the first driving sub-data of different modalities are also different. Unifying the feature representations of the first driving sub-data of different modalities is the basis for realizing the joint generation of driving data of multiple modalities. Therefore, for each first driving sub-data, the electronic device can encode the first driving sub-data to obtain a first feature identifier corresponding to the first driving sub-data). Among them, the first feature identifiers of the first driving sub-data of different modalities are unified in the form of feature representation.
[0039] In one embodiment, tokens can be obtained by tokenizing the first driving sub-data (Tokenization), and the tokens are used as the first feature identifiers corresponding to the first driving sub-data.
[0040] Step 103: Perform spatio-temporal domain prediction on the first feature identifier to obtain a second feature identifier corresponding to the second driving sub-data at the second moment. Wherein, the first moment is before the second moment.
[0041] Exemplarily, there is a causal correlation relationship in the time dimension for the driving sub-data of one modality, and there is a causal correlation relationship in the space dimension between the driving sub-data of different modalities. Therefore, after the electronic device obtains the first feature identifier corresponding to the first driving sub-data, it can further perform spatio-temporal domain prediction on the first feature identifier to obtain a second feature identifier corresponding to the second driving sub-data at the second moment. Wherein, the first moment is before the second moment.
[0042] Step 104: Decode the second feature identifier to obtain the second driving sub-data.
[0043] Exemplarily, the data types and complexities of the first driving sub-data of different modalities are also different. Unifying the feature representations of the first driving sub-data of different modalities is the basis for realizing the joint generation of driving data of multiple modalities. Therefore, after the electronic device obtains the second feature identifier corresponding to the second driving sub-data at the second moment, for each second feature identifier corresponding to the second driving sub-data, the electronic device can decode the second feature identifier of the second driving sub-data to obtain the second driving sub-data.
[0044] In one embodiment, the second feature identifier may be represented by a token, and the second driving sub-data may be obtained by performing detokenization on the second feature identifier (token).
[0045] Step 105: Determine the second driving data at the second moment based on the second driving sub-data.
[0046] Exemplarily, after the electronic device obtains the second driving sub-data at the second moment, it may further determine the second driving data at the second moment based on the second driving sub-data.
[0047] In the embodiments of the present disclosure, through spatio-temporal domain prediction, driving sub-data of different modalities can be effectively fused, joint generation of driving data of multiple modalities can be achieved, and the accuracy of the predicted driving sub-data at the second moment can be improved.
[0048] As Figure 2 shown, based on the above Figure 1 shown embodiment, step 102 may include the following steps:
[0049] Step 201: Determine the data type of the first driving sub-data.
[0050] Exemplarily, the data types of the first driving sub-data of different modalities are usually different, and the encoding methods adopted by the electronic device for the first driving sub-data of different data types are also different. Therefore, for each first driving sub-data, the electronic device can determine the data type of the first driving sub-data. Among them, based on the organization form, data format, and parsability of the data, the data type may include a structured type and an unstructured type. The structured type refers to a data type with a fixed format and organization form, usually stored in a relational database, a table, or a file with a fixed format, and each field of these data has a clear meaning and type; the unstructured type refers to a data type without a fixed format and organization form, usually including text, images, audio, video, etc. The content of these data is rich, but the format and structure are not fixed, and the processing difficulty is relatively large. For example, the data types of the ego vehicle decision-making data and the traffic participant data are structured types. The ego vehicle decision-making data mainly includes lateral speed, longitudinal speed, and angular velocity. The traffic participant data mainly includes length, width, height, orientation, lateral position, longitudinal position, vertical position, lateral speed, longitudinal speed, vertical speed, and traffic participant category (such as motor vehicles, non-motor vehicles, vulnerable traffic participants, etc.). The map data and the ego vehicle perception data are unstructured types. The map data is mainly image data, and the ego vehicle perception data is mainly image data, radar data, point cloud data, etc.
[0051] Of course, those skilled in the art can also define data types based on other data characteristics between driving sub-data of different modalities, and the embodiments of the present disclosure do not make limitations. Taking the data types including structured types and unstructured types as an example, the embodiments of the present disclosure are introduced. The same applies to other situations, and the embodiments of the present disclosure will not be elaborated herein.
[0052] Step 202: Encode the first driving sub-data based on the data type of the first driving sub-data to obtain a first feature identifier corresponding to the first driving sub-data.
[0053] Exemplarily, for each first driving sub-data, based on the data type of the first driving sub-data, the electronic device can use the encoding algorithm corresponding to the first driving sub-data to encode the first driving sub-data to obtain a first feature identifier corresponding to the first driving sub-data.
[0054] In the above embodiments of the present disclosure, by encoding the first driving sub-data of different modalities based on the data types of the first driving sub-data of different modalities using different encoding algorithms, the feature representations of the first driving sub-data of different modalities can be unified, and rich information of the driving sub-data of different modalities can be retained to a large extent, thereby providing a basis for realizing the joint generation of driving data of multiple modalities.
[0055] In some examples, when the data type of the first driving sub-data is a structured type, on the basis of the above Figure 2 shown embodiments, step 202 may include the following steps:
[0056] Step 1: In response to the data type of the first driving sub-data being a structured type, determine the parameter interval in which the value of each driving data parameter in the first driving sub-data is located based on the value of each driving data parameter in the first driving sub-data.
[0057] Exemplarily, when the data type of the first driving sub-data is a structured type, the first driving sub-data usually includes at least one driving data parameter. For example, the ego vehicle decision data includes 3 driving data parameters: lateral speed, longitudinal speed, and angular velocity. The traffic participant data includes 11 driving data parameters: length, width, height, orientation, lateral position, longitudinal position, vertical position, lateral speed, longitudinal speed, vertical speed, and traffic participant category (such as motor vehicle, non-motor vehicle, vulnerable traffic participant, etc.).
[0058] Each parameter interval corresponding to each driving data parameter can be pre-stored in the electronic device. For each driving data parameter, first, the value range of the driving data parameter can be counted, and then, according to the preset number of parameter intervals, the value range of the driving data parameter can be evenly or unevenly divided into multiple parameter intervals. For example, taking the longitudinal speed in traffic participant data as an example, the value range of the longitudinal speed is from 0 km / h to 250 km / h, and the preset number of parameter intervals is 250, then the parameter intervals corresponding to the longitudinal speed are [0, 1), [1, 2),..., [248, 249), [249, 250]. Another example is the traffic participant category in traffic participant data. Usually, motor vehicles are 0, non-motor vehicles are 1, and vulnerable traffic participants are 2. The preset number of parameter intervals is 3, then the parameter intervals corresponding to the traffic participant category are [0, 1), [1, 2), [2, 3).
[0059] It should be noted that the number of parameter intervals corresponding to different driving data parameters can be the same or different, which is not limited in the embodiments of the present disclosure.
[0060] In response to the data type of the first driving sub-data being a structured type, the electronic device can query the target parameter interval in which the value of each driving data parameter in the first driving sub-data is located among the parameter intervals corresponding to the pre-stored driving data parameters. For example, taking the longitudinal speed in traffic participant data as an example, the parameter intervals corresponding to the longitudinal speed are [0, 1), [1, 2),..., [248, 249), [249, 250]. If the value of the longitudinal speed is 30.5 km / h, then the target parameter interval corresponding to the longitudinal speed is [30, 31).
[0061] Step 2: Determine the first feature identifier corresponding to the first driving sub-data based on the interval number of the parameter interval.
[0062] Exemplarily, for each driving data parameter, after the electronic device queries the target parameter interval corresponding to the driving data parameter, it can determine the interval number of the target parameter interval corresponding to the driving data parameter as the first feature identifier corresponding to the driving data parameter. For example, taking the longitudinal speed in traffic participant data as an example, the parameter intervals corresponding to the longitudinal speed are [0, 1), [1, 2),..., [248, 249), [249, 250]. If the value of the longitudinal speed is 30.5 km / h, then the target parameter interval corresponding to the longitudinal speed is [30, 31), and the interval number is 30. Correspondingly, the first feature identifier corresponding to the longitudinal speed is 30.
[0063] Further, for each first driving sub-data, after the electronic device determines the first feature identifiers corresponding to the driving data parameters included in the first driving sub-data, the first feature identifiers corresponding to the driving data parameters can be determined as the first feature identifier corresponding to the first driving sub-data. For example, the ego vehicle decision data includes 3 driving data parameters. Correspondingly, the ego vehicle decision data corresponds to 3 first feature identifiers. Similarly, the traffic participant data includes 11 driving data parameters. Correspondingly, the traffic participant data corresponds to 11 first feature identifiers.
[0064] For example, it is assumed that the driving data parameter can be expressed as D i , 1 ≤ i ≤ N, where N is the number of driving data parameters, and the first moment can be expressed as t m , 1 ≤ m ≤ M, M is the number of first moments, and the driving data parameter D i at the first moment t m The corresponding first feature identifier can be expressed as T i,m . The first feature identifiers corresponding to the driving data parameter D i at different first moments belong to the same type of feature identifiers, which can be expressed as M i , that is, one driving data parameter corresponds to one type of feature identifier. For example, T i,1 , T i,2 …T i,m belong to the same type of feature identifiers.
[0065] In the above embodiments of the present disclosure, for the structured type of the first driving sub-data, through a targeted coding algorithm, high-dimensional and complex data can be converted into low-dimensional and discrete feature representations (i.e., the first feature identifiers), thereby reducing the computational complexity of the joint generation of multiple modalities of driving data and improving the speed of the joint generation of multiple modalities of driving data.
[0066] In some examples, when the data type of the first driving sub-data is unstructured, on the basis of the above Figure 2 shown embodiments, step 202 may include the following steps:
[0067] Step 1, in response to the data type of the first driving sub-data being unstructured, compress the first driving sub-data to obtain the first feature identifier corresponding to the first driving sub-data.
[0068] Exemplarily, when the data type of the first driving sub-data is unstructured, for the unstructured type of the first driving sub-data (such as map data and ego vehicle perception data), the electronic device can compress the first driving sub-data to obtain the first feature identifier corresponding to the first driving sub-data.
[0069] The compression of the first driving sub-data by the above-mentioned electronic device can be achieved through a pre-trained VQGAN (Vector Quantized Generative Adversarial Network) model. Based on the VQGAN model, the high-dimensional image data input by the electronic device passes through an encoder of a Convolutional Neural Network (CNN) to map the high-dimensional image data into a continuous latent space representation. The continuous latent space representation output by the encoder is mapped to a discrete codebook and replaced with the discrete vector in the codebook that is closest to it, thereby discretizing the continuous latent space representation. Among them, the codebook is a predefined set of discrete vectors, and each vector represents a "prototype" of a continuous latent space representation. In this way, the original high-dimensional image data is represented as a series of discrete latent encodings (i.e., the first feature identification sequence), and these encodings are the indices of the vectors in the codebook. This discrete representation form greatly compresses the image data. Therefore, the electronic device can input the first driving sub-data into the VQGAN model, and the first driving sub-data is mapped to the discretized vectors in the VQ space to obtain a feature identification sequence containing multiple first feature identifications, that is, the first feature identifications corresponding to the first driving sub-data are obtained. For example, both the map data and the ego-vehicle perception data correspond to 512 first feature identifications.
[0070] Further, for example, taking the first driving sub-data as image data and the image data containing multiple image blocks as an example, assume that the image block can be represented as B j , 1 ≤ j ≤ K, where K is the number of image blocks, and the first moment can be represented as t m , 1 ≤ m ≤ M, M is the number of the first moments, and the image block B j at the first moment t m the corresponding first feature identification can be represented as T j,m . The first feature identifications corresponding to the image block B j at different first moments belong to the same class of feature identifications, which can be represented as M j , that is, an image block corresponds to a class of feature identifications. For example, T j,1 , T j,2 …T j,m belong to the same class of feature identifications.
[0071] It should be noted that the feature identifications corresponding to the above-mentioned driving data parameters and the feature identifications corresponding to the image blocks constitute all the feature identifications.
[0072] Of course, those skilled in the art can also use other deep learning models to achieve the compression of the first driving sub-data, and the embodiments of the present disclosure are not limited thereto.
[0073] In the above embodiments of the present disclosure, for the unstructured first driving sub-data, through a targeted compression algorithm, high-dimensional and complex data can be converted into low-dimensional and discrete feature representations (i.e., the second feature identifier), thereby reducing the computational complexity of the joint generation of multi-modal driving data and improving the speed of the joint generation of multi-modal driving data.
[0074] As Figure 3 shown, based on the above Figure 1 shown embodiments, step 103 may include the following steps:
[0075] Step 301, determine the high-dimensional features corresponding to each first feature identifier.
[0076] Exemplarily, a feature codebook may be pre-stored in the electronic device. The feature codebook contains the correspondence between the feature identifier and the high-dimensional feature. In one embodiment, for each feature identifier, first, a set of frequency components corresponding to the feature identifier may be generated based on the feature identifier, and then, a Fourier transform is performed on the set of frequency components corresponding to the feature identifier to obtain the high-dimensional feature corresponding to the feature identifier. Finally, the correspondence between the feature identifier and the high-dimensional feature is stored as the feature codebook. Among them, one feature identifier uniquely corresponds to one high-dimensional feature.
[0077] For each first feature identifier included in each type of feature identifier, the electronic device may query the high-dimensional features corresponding to each first feature identifier in the pre-stored correspondence between the feature identifier and the high-dimensional feature. For example, as Figure 4 shown, taking 3 first moments as an example for illustration, there are multiple first moments (i.e., Figure 4 the moments 1 to 3 in Figure 4 ), and the high-dimensional features corresponding to the first feature identifiers (T 1 , T 1,1 , T 1,2 ) included in the first type of feature identifier (i.e., the feature identifier M 1,3 ) in 1,1 are (G 1,2 , G 1,3 ); the high-dimensional features corresponding to the first feature identifiers (T Figure 4 ), T 2 ), T 2,1 ), T 2,2 ), T 2,3 ) included in the second type of feature identifier (i.e., the feature identifier M 2,1 ) in 2,2 are (G 2,3 ); the high-dimensional features corresponding to the first feature identifiers (T Figure 4 ), T i ) included in the i-th type of feature identifier (i.e., the feature identifier M i,1 ) ini,2 , T i,3 ) The corresponding high-dimensional feature is (G i,1 , G i,2 , G i,3 ); The feature identifier of the nth class (i.e., Figure 4 the feature identifier M in n ) contains the first feature identifier (T n,1 , T n,2 , T n,3 ) The corresponding high-dimensional feature is (G n,1 , G n,2 , G n,3 ). Among them, the feature identifier M of the first class 1 to the feature identifier M of the nth class n The total first feature identifiers included are the first feature identifiers corresponding to the first driving sub-data of all modalities.
[0078] Step 302, based on the high-dimensional feature, determine the spatio-temporal feature corresponding to the first feature identifier.
[0079] Exemplarily, there is a causal correlation relationship in the time dimension for the driving sub-data of one modality, and there is a causal correlation relationship in the space dimension between the driving sub-data of different modalities. Therefore, after the electronic device obtains the high-dimensional feature corresponding to the first feature identifier, it can further determine the spatio-temporal feature corresponding to the first feature identifier based on the high-dimensional feature corresponding to the first feature identifier. Based on the method of separating and extracting the temporal sequence feature and the spatial feature, the present disclosure first performs feature aggregation in the time dimension based on the high-dimensional feature of the first feature identifier to obtain the temporal sequence feature corresponding to the first feature identifier. Then, based on the temporal sequence feature corresponding to the first feature identifier, feature aggregation is performed in the space dimension to obtain the spatio-temporal feature corresponding to the first feature identifier.
[0080] Step 303, based on the spatio-temporal features corresponding to each first feature identifier, determine the second feature identifier.
[0081] Exemplarily, after the electronic device obtains the spatio-temporal features corresponding to each first feature identifier, it can further determine the second feature identifier corresponding to the second driving sub-data at the second moment based on the spatio-temporal features corresponding to each first feature identifier.
[0082] In some examples, after the electronic device obtains the spatio-temporal features corresponding to each first feature identifier, for each first feature identifier, the spatio-temporal feature corresponding to the first feature identifier can be input into a Multilayer Perceptron (MLP) to output the second feature identifier corresponding to the second driving sub-data at the second moment. The multilayer perceptron is a neural network composed of multiple fully connected layers, which is used to learn the non-linear features of the input data and output the prediction result.
[0083] In the above embodiments of the present disclosure, by extracting temporal features and spatial features, driving sub-data of different modalities can be effectively fused, and the accuracy of the predicted driving sub-data at the second moment can be improved.
[0084] As Figure 5 shown, based on the above Figure 3 shown embodiments, step 302 may include the following steps:
[0085] Step 501: For each first feature identifier, splice the high-dimensional features corresponding to the first feature identifier at each first moment to obtain a first spliced feature corresponding to the first feature identifier.
[0086] Exemplarily, as Figure 4 shown, after the electronic device obtains the high-dimensional features corresponding to the first feature identifier included in each type of feature identifier, for each type of feature identifier, the high-dimensional features corresponding to the first feature identifier of this type of feature identifier at each first moment can be spliced to obtain a first spliced feature corresponding to this type of feature identifier. For example, as Figure 4 shown, for the feature identifier M 1 , splice the high-dimensional features (G 1 , G 1,1 , G 1,2 ) corresponding to the first feature identifier (T 1,3 ) of the feature identifier M at moments 1 to 3 to obtain a first spliced feature corresponding to the feature identifier M 1,1 ; for the feature identifier M 1,2 , splice the high-dimensional features (G 1,3 , G 1 ) corresponding to the first feature identifier (T 2 ) of the feature identifier M at moments 1 to 3 to obtain a first spliced feature corresponding to the feature identifier M 2 , and so on. 2,1 , T 2,2 , T 2,3 ) corresponding to the first feature identifier (T 2,1 , G 2,2 , G 2,3 ) corresponding to the first feature identifier (T 2 ) of the feature identifier M at moments 1 to 3 to obtain a first spliced feature corresponding to the feature identifier M
[0087] Step 502: Perform causal-based temporal feature aggregation on the first spliced feature corresponding to the first feature identifier to obtain a temporal feature corresponding to the first feature identifier.
[0088] Exemplarily, after the electronic device obtains the first spliced feature corresponding to each type of feature identifier, for each type of feature identifier, causal-based temporal feature aggregation can be performed on the first spliced feature corresponding to this type of feature identifier to obtain a temporal feature corresponding to this type of feature identifier. For example, asFigure 4 As shown, the feature identifier M 1 The corresponding timing feature is A 1 , the feature identifier M 2 The corresponding timing feature is A 2 , and so on.
[0089] It can be understood that the causal-based timing feature aggregation refers to the influence among the first feature identifiers based on the timing relationship. For example, the first feature identifier at an earlier time will affect the first feature identifier at a later time, while the first feature identifier at a later time will not affect the first feature identifier at an earlier time. Further, the first feature identifier at a later time can be affected only by the first feature identifier at its previous time. For example, the first feature identifier at time 10 is affected only by the first feature identifier at time 9. Or, the first feature identifier at a later time can be affected by the first feature identifiers at the previous preset number of earlier times. For example, the first feature identifier at time 10 is affected by the first feature identifiers at times 5 to 9. Or, the first feature identifier at a later time is affected by all the first feature identifiers at earlier times. For example, the first feature identifier at time 10 is affected by the first feature identifiers at times 1 to 9. It should be noted that for the case where the first feature identifier at a later time is affected by multiple first feature identifiers at earlier times, the influence weights of the multiple first feature identifiers at earlier times on the first feature identifier at a later time can be the same or different. For the case where the influence weights of the multiple first feature identifiers at earlier times are different, in some examples, the influence weights can be set in the order of time.
[0090] In some examples, the timing feature extraction can adopt a Transformer model, which includes multiple layers of causal self-attention mechanisms and a feed-forward network. For each type of feature identifier, the electronic device concatenates the high-dimensional features corresponding to the first feature identifiers of this type of feature identifier at each first time in chronological order to obtain the first concatenated feature corresponding to this type of feature identifier. Then, the electronic device inputs the first concatenated feature corresponding to this type of feature identifier into the Transformer model for feature aggregation in the time dimension to obtain the timing feature corresponding to this type of feature identifier.
[0091] Step 503, add the timing feature corresponding to the first feature identifier and the target high-dimensional feature to obtain the added feature corresponding to the first feature identifier.
[0092] Exemplarily, after the electronic device obtains the temporal features corresponding to the feature identifiers of each category, for each category of feature identifiers, the electronic device may add the temporal features corresponding to the category of feature identifiers to the target high-dimensional features to obtain the added features corresponding to the category of feature identifiers. The embodiments of the present disclosure adopt a misaligned addition method. Therefore, for the first category of feature identifiers, the corresponding target high-dimensional feature is a preset high-dimensional feature. For other categories of feature identifiers, the corresponding target high-dimensional feature is the high-dimensional feature corresponding to the second feature identifier of this category of feature identifiers at the second moment. For example, as Figure 4 shown, for the first category of feature identifier M 1 , the corresponding target high-dimensional feature is the preset high-dimensional feature G 0,0 , the first category of feature identifier M 1 corresponding temporal feature A 1 and the high-dimensional feature G 0,0 are added to obtain the added feature B 1 . For the second category of feature identifier M 2 , the corresponding target high-dimensional feature is the high-dimensional feature (G 1,4 ) corresponding to the second feature identifier (T 1,4 ) of this category of feature identifiers at the second moment (moment 4). The second category of feature identifier M 2 corresponding temporal feature A 2 and the high-dimensional feature G 1,4 are added to obtain the added feature B 2 , and so on.
[0093] Step 504: Perform causal-based spatial feature aggregation according to the added feature corresponding to the first feature identifier to obtain the spatio-temporal feature corresponding to the first feature identifier.
[0094] Exemplarily, after the electronic device obtains the added features corresponding to the feature identifiers of each category, for each category of feature identifiers, it may perform causal-based spatial feature aggregation according to the added features corresponding to all the previous feature identifiers of this category to obtain the spatio-temporal feature corresponding to this category of feature identifiers. For example, as Figure 4 shown, for the first category of feature identifier M 1 , the electronic device performs causal-based spatial feature aggregation on the added feature B 1 corresponding to the feature identifier M 1 to obtain the spatio-temporal feature C 1 corresponding to the feature identifier M 1 ; for the second category of feature identifier M 2 , the electronic device combines the added feature B 2 corresponding to the feature identifier M 2 with the added feature B 1 corresponding to the feature identifier M 1Perform causal-based spatial feature aggregation together to obtain feature identifier M 2 The corresponding spatio-temporal feature C 2 , and so on.
[0095] In some examples, the spatial feature extraction may adopt a Transformer model, which includes multiple layers of causal self-attention mechanisms and feed-forward networks. For each type of feature identifier, the electronic device may input the sum feature corresponding to all the previous feature identifiers of this type into the Transformer model for feature aggregation in the spatial dimension.
[0096] It can be understood that the causal-based spatial feature aggregation means that there is an impact among the first feature identifiers based on spatial changes. For example, the first feature identifiers corresponding to the ego vehicle decision data respectively affect the first feature identifiers corresponding to the traffic participant data, the map data, and the ego vehicle perception data; the first feature identifiers corresponding to the traffic participant data respectively affect the first feature identifiers corresponding to the map data and the ego vehicle perception data; the first feature identifiers corresponding to the map data affect the first feature identifiers corresponding to the ego vehicle perception data. The space described in this disclosure refers to one or more of a physical space, a topological space, a semantic space, and an interaction space. Among them, the physical space mainly focuses on geometric attributes such as the position, speed, direction, and distance of objects in the real world; the topological space mainly focuses on the connectivity of the space structure, road feasibility, and driving area relevance; the semantic space mainly focuses on the high-level abstract meaning defined by rules, labels, and risks, and this high-level abstract meaning is used to represent the functions and intentions in the environment; the interaction space mainly focuses on the spatial relationships formed by the positions, behaviors, and interactions of entities (such as vehicles and pedestrians).
[0097] In the embodiments of the present disclosure, it is assumed that the number of the first feature identifiers at a first moment is N, and the dimension of the high-dimensional feature dimension of each first feature identifier is C. Then, the dimension of all the first feature identifiers at T first moments is T * N * C. The traditional non-separated method for extracting temporal features and spatial features is performed in the T + N dimension. Therefore, the complexity corresponding to the traditional non-separated method is O((T + N) 2 ). Compared with the traditional non-separated method, the separated method adopted in the embodiments of the present disclosure for extracting temporal features is performed in the T dimension, and the complexity corresponding to the temporal feature extraction is O(T 2 ). The spatial feature extraction is performed in the N dimension, and the complexity corresponding to the spatial feature extraction is O(N 2 ). Therefore, the complexity corresponding to the separated method adopted in the embodiments of the present disclosure is O(T 2 + N 2) This significantly reduces the complexity of extracting temporal and spatial features. At the same time, the traditional non-separable method cannot achieve the spatial fusion of various modalities of driving sub-data. The separable method adopted in the embodiments of the present disclosure can effectively fuse different modalities of driving sub-data, which can improve the accuracy of the predicted driving sub-data at the second moment.
[0098] In some examples, the map data is generated from road elements within a preset distance centered on the host vehicle. As the host vehicle moves, the road elements in the map will move relative to the host vehicle, and the road elements originally located outside the preset distance of the host vehicle will move within the preset distance of the host vehicle. For example, as the host vehicle moves, the lane lines gradually extend in the driving direction of the host vehicle, and the lane segments originally located outside the preset distance of the host vehicle will gradually move within the preset distance of the host vehicle. As the host vehicle moves, some of the road elements within the preset distance of the host vehicle still remain within the preset distance of the host vehicle, only their positions have changed. Another part of the road elements within the preset distance of the host vehicle moves outside the preset distance of the host vehicle. At the same time, due to the continuity of the road elements. Therefore, based on the road elements originally within the preset distance, the road elements that gradually move from outside the preset distance of the host vehicle to within the preset distance of the host vehicle can be generated. Based on this, after the electronic device predicts the host vehicle decision data at the second moment, before predicting the second feature identifier of the map data at the second moment, the second high-dimensional feature of the second feature identifier corresponding to the map data at the second moment can be further generated based on the host vehicle decision data at the second moment and the first high-dimensional feature of the first feature identifier corresponding to the map data at the first moment, and the obtained second high-dimensional feature of the second feature identifier is used to predict the second feature identifier of the map data at the second moment. As Figure 6 shown, the above process may include the following steps:
[0099] Step 601, determine the second position of the second high-dimensional feature in the map feature matrix based on the host vehicle decision data at the second moment and the first position of the first high-dimensional feature in the map feature matrix. Wherein, the first high-dimensional feature is the high-dimensional feature of the first feature identifier corresponding to the map data at the first moment, and the second high-dimensional feature is the high-dimensional feature of the second feature identifier corresponding to the map data at the second moment.
[0100] Exemplarily, when encoding the map data at the first moment, the electronic device encodes each map block in the map data in the order from left to right and from top to bottom to obtain the first feature identifier corresponding to each map block. After the electronic device queries the second high-dimensional feature corresponding to each first feature identifier in the pre-stored correspondence between the feature identifier and the high-dimensional feature, based on the positional relationship between the map blocks, the second high-dimensional feature corresponding to the first feature identifier can be generated into a map feature matrix. For example, in the case where the map data includes 256 map blocks, in the map feature matrix, the second high-dimensional features corresponding to each map block are respectively m 1,1, m 1,2 …m 1,16 …m 16,16 。
[0101] After the electronic device obtains the ego-vehicle decision data at the second moment, it can determine the movement vector (lateral speed, longitudinal speed, and angular velocity) of the ego-vehicle based on the ego-vehicle decision data. Then, for each first high-dimensional feature in the map feature matrix, the electronic device can perform an inverse affine transformation on the first position of the first high-dimensional feature and the movement vector, so as to obtain the second position of the second high-dimensional feature in the map feature matrix. For example, as Figure 7 shown, the first high-dimensional features in the map feature matrix are respectively M 1,1 , M 1,2 …M 1,16 …M 16,16 ; taking the first high-dimensional feature M 1,1 as an example, if the movement vector of the ego-vehicle is (0, 1, 0), indicating that the ego-vehicle moves straight forward by one map block, after the inverse affine transformation, the second high-dimensional feature corresponding to the first high-dimensional feature M 1,1 is m -1,1 .
[0102] Step 602: Determine the interpolated high-dimensional feature based on the second position of the second high-dimensional feature, and determine the first high-dimensional feature based on the interpolated high-dimensional feature.
[0103] Exemplarily, after the electronic device determines the second position of the second high-dimensional feature, it can select other second high-dimensional features in the map feature matrix that satisfy the preset position proximity condition with the second position as the center, as the interpolated high-dimensional feature. For example, the second high-dimensional feature is m -1,1 , if the position proximity condition is within a range of 1 map block horizontally and vertically, the interpolated high-dimensional features are M 1,1 , M 1,2 .
[0104] Furthermore, after the electronic device determines the interpolated high-dimensional feature, it can perform an interpolation operation on the interpolated high-dimensional feature to obtain the first high-dimensional feature.
[0105] Step 603: Perform causal-based spatial feature aggregation on the sum feature of the first feature identifier corresponding to the map data and the second high-dimensional feature, to obtain the spatio-temporal feature corresponding to the first feature identifier.
[0106] Exemplarily, before predicting the map data at the second moment, the electronic device may add the sum feature of the first feature identifiers corresponding to the map data and the second high-dimensional feature, and further perform causal spatial feature aggregation to obtain the spatio-temporal feature corresponding to the first feature identifier. Among them, the process of the electronic device performing causal-based spatial feature aggregation according to the sum feature of the first feature identifiers corresponding to the map data and the second high-dimensional feature to obtain the spatio-temporal feature corresponding to the first feature identifier is similar to the process of step 504 above, and will not be elaborated here.
[0107] In the embodiments of the present disclosure, before predicting the map data at the second moment, based on the ego vehicle decision data at the second moment and the second high-dimensional feature of the first feature identifier corresponding to the map data at the first moment, the electronic device generates the first high-dimensional feature of the second feature identifier corresponding to the map data at the second moment, and adds the generated first high-dimensional feature to the corresponding sum feature, thereby improving the prediction accuracy of the map data at the second moment.
[0108] As Figure 8 shown, on the basis of the above Figure 1 shown embodiment, step 104 may include the following steps:
[0109] Step 801, determine the data type of the second driving sub-data corresponding to the second feature identifier.
[0110] Exemplarily, the data types of different modalities of driving sub-data are usually different. For driving sub-data of different data types, the encoding methods adopted by the electronic device are also different. Correspondingly, for the feature identifiers corresponding to driving sub-data of different data types, the decoding methods adopted by the electronic device are also different. Therefore, for each second driving sub-data, the electronic device can determine the data type of the second driving sub-data. Among them, based on the organization form, data format, and parsability of the data, the data type may include a structured type and an unstructured type. For example, the data types of ego vehicle decision data and traffic participant data are structured types, while map data and ego vehicle perception data are unstructured types.
[0111] Of course, those skilled in the art can also define the data type based on other characteristics of the data, which is not limited in the embodiments of the present disclosure. The embodiments of the present disclosure are introduced by taking the data type including structured type and unstructured type as an example. Other situations are similar, and the embodiments of the present disclosure will not be elaborated.
[0112] Step 802, based on the data type of the second driving sub-data corresponding to the second feature identifier, decode the second feature identifier to obtain the second driving sub-data.
[0113] Exemplarily, for each second driving sub-data, in response to the data type of the second driving sub-data, the electronic device may decode the second feature identifier corresponding to the second driving sub-data based on the decoding algorithm corresponding to the second driving sub-data to obtain the second driving sub-data.
[0114] In the above embodiments of the present disclosure, based on the data types of the second driving sub-data of different modalities, different decoding algorithms are used to decode the second feature identifiers of the second driving sub-data, so as to realize the joint generation of driving data of multiple modalities.
[0115] In some examples, when the data type of the second driving sub-data corresponding to the second feature identifier is a structured type, on the basis of the above Figure 8 shown embodiments, step 802 may include the following steps:
[0116] Step 1, in response to the data type of the second driving sub-data corresponding to the second feature identifier being a structured type, for each driving data parameter in the second driving sub-data, based on the target parameter interval with the second feature identifier corresponding to the driving data parameter as the interval number, determine the driving data parameter.
[0117] Exemplarily, similar to the structured type of the first driving sub-data, when the data type of the second driving sub-data is a structured type, the second driving sub-data usually includes at least one driving data parameter. For example, the ego vehicle decision data includes 3 driving data parameters: lateral speed, longitudinal speed, and angular velocity. The traffic participant data includes 11 driving data parameters: length, width, height, orientation, lateral position, longitudinal position, vertical position, lateral speed, longitudinal speed, vertical speed, and traffic participant category (such as motor vehicle, non-motor vehicle, vulnerable traffic participant, etc.).
[0118] Each parameter interval corresponding to each driving data parameter may be pre-stored in the electronic device. For each driving data parameter, first, the numerical range of the driving data parameter may be counted, and then, according to the preset number of parameter intervals, the numerical range of the driving data parameter is evenly / non-uniformly divided into multiple parameter intervals. For example, taking the longitudinal speed in the traffic participant data as an example, the numerical range of the longitudinal speed is 0 km / h to 250 km / h, and the preset number of parameter intervals is 250, then the parameter intervals corresponding to the longitudinal speed are [0, 1), [1, 2) …… [248, 249), [249, 250].
[0119] Based on the introduction in the foregoing step 202, each driving data parameter included in the first driving sub-data corresponds to a first feature identifier, and the first feature identifier corresponding to the driving data parameter is the interval number of the target parameter interval corresponding to the driving data parameter. Therefore, in response to the data type of the second driving sub-data being an unstructured type, for each driving data parameter in the second driving sub-data, the electronic device may query, in the respective parameter intervals corresponding to the pre-stored driving data parameters, the target parameter interval with the second feature identifier corresponding to the driving data parameter as the interval number. For example, taking the longitudinal speed in the traffic participant data as an example, the respective parameter intervals corresponding to the longitudinal speed are [0, 1), [1, 2) …… [248, 249), [249, 250]. If the second feature identifier corresponding to the longitudinal speed is 35, that is, the interval number of the target parameter interval corresponding to the longitudinal data is 35, then the target parameter interval corresponding to the longitudinal speed is [34, 35).
[0120] For each driving data parameter, after the electronic device determines the target parameter interval corresponding to the driving data parameter, it may determine the value of the driving data parameter based on the upper limit value and the lower limit value of the target parameter interval. Among them, the electronic device may determine the average value of the upper limit value and the lower limit value as the value of the driving data parameter; the electronic device may also determine the weighted average value of the upper limit value and the lower limit value as the value of the driving data parameter. For example, taking the longitudinal speed in the traffic participant data as an example, if the target parameter interval corresponding to the longitudinal speed is [34, 35), then the value corresponding to the longitudinal speed is (34 + 35) / 2 = 34.5 km / h. Another example, taking the longitudinal speed in the traffic participant data as an example, if the target parameter interval corresponding to the longitudinal speed is [34, 35), the weight corresponding to the upper limit value is 0.7, and the weight corresponding to the lower limit value is 0.3, then the value corresponding to the longitudinal speed is 34*0.3 + 35*0.7 = 34.7 km / h.
[0121] Of course, those skilled in the art may also determine the value of the driving data parameter based on the target parameter interval and a preset other parameter determination algorithm, and the embodiments of the present disclosure do not make limitations.
[0122] Step three, determine the second driving sub-data based on each driving data parameter.
[0123] Exemplarily, after the electronic device determines each driving data parameter in the second driving sub-data, it may further determine the second driving sub-data based on each driving data parameter.
[0124] In the above embodiments of the present disclosure, for the second driving sub-data of the structured type, through a targeted decoding algorithm, the low-dimensional and discrete feature representation (i.e., the second feature identifier) can be converted into high-dimensional and complex data, thereby obtaining the second driving sub-data.
[0125] In some examples, when the data type of the second driving sub-data corresponding to the second feature identifier is of the unstructured type, based on the above Figure 8 shown embodiments, step 802 may include the following steps:
[0126] Step 1, in response to the data type of the second driving sub-data corresponding to the second feature identifier being of the unstructured type, decompress the second feature identifier corresponding to the second driving sub-data to obtain the second driving sub-data.
[0127] Exemplarily, similar to the first driving sub-data of the unstructured type, when the data type of the second driving sub-data is of the unstructured type, for the unstructured type of the second driving sub-data (such as map data and ego-vehicle perception data), the electronic device can decompress the second feature identifier of the second driving sub-data to obtain the second driving sub-data.
[0128] The decompression of the second feature identifier by the above electronic device can be implemented by a pre-trained VQGAN model. Based on the VQGAN model, the electronic device looks up the vector corresponding to the latent code (i.e., the second feature identifier sequence) in the codebook, and then uses a convolutional neural network to map the vector to a continuous latent space to obtain high-dimensional data. That is, the electronic device inputs the second feature identifier corresponding to the second driving sub-data into the VQGAN model, and the decoder generates the second driving sub-data based on the second feature identifier. For example, taking the image data in the ego-vehicle perception data as an example, after the electronic device obtains the second feature identifier corresponding to the image data at the second moment, it can input the second feature identifier into the decoder of the VQGAN model. The decoder looks up the vector corresponding to the second feature identifier in the codebook based on the second feature identifier and maps the vector corresponding to the second feature identifier back to the continuous latent space, thereby obtaining the image data at the second moment.
[0129] Of course, those skilled in the art can also use other deep learning models to implement the decompression of the second feature identifier, which is not limited in the embodiments of the present disclosure.
[0130] In the above embodiments of the present disclosure, for the second driving sub-data of the unstructured type, through a targeted decompression algorithm, the low-dimensional and discrete feature representation (i.e., the second feature identifier) can be converted into high-dimensional and complex data, thereby obtaining the second driving sub-data.
[0131] Exemplary device
[0132] Figure 9 is a schematic structural diagram of a driving data generation device provided by an exemplary embodiment of the present disclosure. As Figure 9 shown, the driving data generation device 900 includes a driving data acquisition module 901, a driving data encoding module 902, a feature identifier prediction module 903, a feature identifier decoding module 904, and a driving data generation module 905.
[0133] The driving data acquisition module 901 is configured to acquire first driving data at at least one first moment, and the first driving data includes first driving sub-data of at least one modality;
[0134] The driving data encoding module 902 is configured to encode the first driving sub-data to obtain a first feature identifier corresponding to the first driving sub-data;
[0135] The feature identifier prediction module 903 is configured to perform spatio-temporal domain prediction on the first feature identifier to obtain a second feature identifier corresponding to second driving sub-data at a second moment; the first moment is before the second moment;
[0136] The feature identifier decoding module 904 is configured to decode the second feature identifier to obtain the second driving sub-data;
[0137] The driving data generation module 905 is configured to determine second driving data at the second moment based on the second driving sub-data.
[0138] In some embodiments, as Figure 10 shown, the driving data encoding module 902 includes: a data type determination unit 9021 and a driving data encoding unit 9022.
[0139] The data type determination unit 9021 is configured to determine the data type of the first driving sub-data;
[0140] The driving data encoding unit 9022 is configured to encode the first driving sub-data based on the data type of the first driving sub-data to obtain a first feature identifier corresponding to the first driving sub-data.
[0141] In some embodiments, the driving data encoding unit 9022 is specifically configured to:
[0142] In response to the data type of the first driving sub-data being a structured type, determine a parameter interval in which the value of each driving data parameter in the first driving sub-data is located based on the value of each driving data parameter in the first driving sub-data;
[0143] Determine a first feature identifier corresponding to the first driving sub-data based on the interval number of the parameter interval;
[0144] Or,
[0145] In response to the data type of the first driving sub-data being an unstructured type, compress the first driving sub-data to obtain a first feature identifier corresponding to the first driving sub-data.
[0146] In some embodiments, as Figure 11 shown, the feature identifier prediction module 903 includes: a high-dimensional feature determination unit 9031, a spatio-temporal feature determination unit 9032, and a feature identifier determination unit 9033.
[0147] The high-dimensional feature determination unit 9031 is configured to determine the high-dimensional feature corresponding to each first feature identifier based on the correspondence between the pre-stored feature identifier and the high-dimensional feature;
[0148] The spatio-temporal feature determination unit 9032 is configured to determine the spatio-temporal feature corresponding to the first feature identifier based on the high-dimensional feature;
[0149] The feature identifier determination unit 9033 is configured to determine the second feature identifier based on the spatio-temporal features corresponding to each first feature identifier.
[0150] In some embodiments, the spatio-temporal feature determination unit 9032 is specifically configured to:
[0151] For each first feature identifier, splice the high-dimensional features corresponding to the first feature identifier at each first moment to obtain a first spliced feature corresponding to the first feature identifier;
[0152] Perform causal-based temporal feature aggregation on the first spliced feature corresponding to the first feature identifier to obtain a temporal feature corresponding to the first feature identifier;
[0153] Add the temporal feature corresponding to the first feature identifier and the target high-dimensional feature to obtain an added feature corresponding to the first feature identifier;
[0154] Perform causal-based spatial feature aggregation according to the added feature corresponding to the first feature identifier to obtain a spatio-temporal feature corresponding to the first feature identifier.
[0155] In some embodiments, as Figure 12 shown, it further includes: a position determination module 906 and a high-dimensional feature determination module 907.
[0156] The position determination module 906 is configured to determine the second position of the first high-dimensional feature in the map feature matrix based on the ego vehicle decision data at the second moment and the first position of the second high-dimensional feature in the map feature matrix; the first high-dimensional feature is the high-dimensional feature of the first feature identifier corresponding to the map data at the first moment, and the second high-dimensional feature is the high-dimensional feature of the second feature identifier corresponding to the map data at the second moment;
[0157] The high-dimensional feature determination module 907 is configured to determine an interpolated high-dimensional feature based on the second position of the second high-dimensional feature, and determine the second high-dimensional feature based on the interpolated high-dimensional feature;
[0158] Performing causal spatial feature aggregation on the sum feature corresponding to the first feature identifier to obtain the spatio-temporal feature corresponding to the first feature identifier, including:
[0159] Performing causal spatial feature aggregation on the sum feature of the first feature identifier corresponding to the map data and the second high-dimensional feature to obtain the spatio-temporal feature corresponding to the first feature identifier.
[0160] In some embodiments, as Figure 13 shown, the feature identifier decoding module 904 includes: a data type determination unit 9041 and a feature identifier decoding unit 9042.
[0161] The data type determination unit 9041 is configured to determine the data type of the second driving sub-data corresponding to the second feature identifier;
[0162] The feature identifier decoding unit 9042 is configured to decode the second feature identifier based on the data type of the second driving sub-data corresponding to the second feature identifier to obtain the second driving sub-data.
[0163] In some embodiments, the feature identifier decoding unit 9042 is specifically configured to:
[0164] In response to the data type of the second driving sub-data corresponding to the second feature identifier being a structured type, for each driving data parameter in the second driving sub-data, based on the target parameter interval numbered by the second feature identifier corresponding to the driving data parameter, determine the driving data parameter;
[0165] Based on each driving data parameter, determine the second driving sub-data;
[0166] Or,
[0167] In response to the data type of the second driving sub-data corresponding to the second feature identifier being a structured type, decompress the second feature identifier to obtain the second driving sub-data.
[0168] For the beneficial technical effects corresponding to the exemplary embodiments of the present device, reference may be made to the corresponding beneficial technical effects in the above-mentioned exemplary method section, which will not be elaborated here.
[0169] Exemplary electronic device
[0170] Figure 14 It is a structural diagram of an electronic device provided by an embodiment of the present disclosure, including at least one processor 11 and a memory 12.
[0171] The processor 11 may be a central processing unit (CPU) or other form of processing unit with data processing capabilities and / or instruction execution capabilities, and may control other components in the electronic device 10 to perform desired functions.
[0172] The memory 12 may include one or more computer program products, and the computer program products may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. Volatile memory may include, for example, random access memory (RAM) and / or cache memory, etc. Non-volatile memory may include, for example, read-only memory (ROM), hard disk, flash memory, etc. One or more computer program instructions may be stored on the computer-readable storage media, and the processor 11 may run one or more computer program instructions to implement the driving data generation methods of various embodiments of the present disclosure described above and / or other desired functions.
[0173] In one example, the electronic device 10 may further include: an input device 13 and an output device 14, and these components are interconnected through a bus system and / or other forms of connection mechanisms (not shown).
[0174] The input device 13 may further include, for example, a keyboard, a mouse, and the like.
[0175] The output device 14 may output various information to the outside, which may include, for example, a display, a speaker, a printer, and a communication network and its connected remote output devices, and the like.
[0176] Of course, for simplicity, Figure 14 only some of the components in the electronic device 10 related to the present disclosure are shown, and components such as a bus, an input / output interface, etc. are omitted. In addition, according to specific application scenarios, the electronic device 10 may further include any other appropriate components.
[0177] Exemplary Computer Program Product and Computer Readable Storage Medium
[0178] In addition to the above methods and devices, embodiments of the present disclosure may further provide a computer program product, including computer program instructions, and when the computer program instructions are run by a processor, the processor is caused to execute the steps in the driving data generation methods of various embodiments of the present disclosure described in the above "Exemplary Method" section.
[0179] A computer program product may write program code for performing the operations of the embodiments of the present disclosure in any combination of one or more programming languages, including object-oriented programming languages such as Java, C++, etc., and also including conventional procedural programming languages such as the "C" language or similar programming languages. The program code may be executed entirely on the user's computing device, partially on the user's device, executed as a stand-alone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server.
[0180] In addition, an embodiment of the present disclosure may also be a computer-readable storage medium having computer program instructions stored thereon, and when the computer program instructions are run by a processor, the processor is caused to execute the steps in the driving data generation method of various embodiments of the present disclosure described in the above "Exemplary Method" section.
[0181] The computer-readable storage medium may adopt any combination of one or more readable media. The readable media may be a readable signal medium or a readable storage medium. The readable storage medium, for example but not limited to, includes systems, devices or components of electricity, magnetism, light, electromagnetic, infrared, or semiconductor, or any combination of the above. More specific examples (non-exhaustive list) of the readable storage medium include: an electrical connection having one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.
[0182] The basic principles of the present disclosure have been described above in conjunction with specific embodiments. However, the advantages, benefits, effects, etc. mentioned in the present disclosure are only examples and not limitations, and it cannot be considered that they are essential for each embodiment of the present disclosure. In addition, the above-disclosed specific details are only for the purposes of illustration and easy understanding, and not for limitation. The above details do not limit the present disclosure to necessarily adopt the above specific details for implementation.
[0183] Those skilled in the art can make various changes and modifications to the present disclosure without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present disclosure and their equivalent technologies, the present disclosure also intends to include these changes and modifications.
Claims
1. A driving data generating method, comprising: Acquire at least one first driving data at a first moment, wherein the first driving data includes at least one first driving sub-data of a modality; Encoding the first driving sub-data to obtain a first feature identifier corresponding to the first driving sub-data; Performing spatiotemporal prediction on the first feature identifier to obtain a second feature identifier corresponding to the second driving sub-data at a second moment; The first moment is before the second moment; Decoding the second feature identifier to obtain the second driving sub-data; Based on the second driving sub-data, second driving data at the second moment is determined.
2. The method according to claim 1, wherein: The encoding of the first driving sub-data to obtain a first feature identifier corresponding to the first driving sub-data includes: Determine a data type of the first driving sub-data; Based on the data type of the first driving sub-data, the first driving sub-data is encoded to obtain the first feature identifier corresponding to the first driving sub-data.
3. The method according to claim 2, wherein: The encoding of the first driving sub-data based on the data type of the first driving sub-data to obtain the first feature identifier corresponding to the first driving sub-data includes: In response to the data type of the first driving sub-data being a structured type, determining, based on the value of each driving data parameter in the first driving sub-data, a parameter interval in which the value of the driving data parameter lies; Determining a first feature identifier corresponding to the first driving sub-data based on the interval number of the parameter interval; or, In response to the data type of the first driving sub-data being an unstructured type, the first driving sub-data is compressed to obtain a first feature identifier corresponding to the first driving sub-data.
4. The method according to claim 1, wherein: The performing spatiotemporal prediction on the first feature identifier to obtain a second feature identifier corresponding to the second driving sub-data at the second moment includes: Determine a high-dimensional feature corresponding to each of the first feature identifiers; Based on the high-dimensional feature, determining the spatiotemporal feature corresponding to the first feature identifier; The second feature identifier is determined based on the spatiotemporal features corresponding to each of the first feature identifiers.
5. The method according to claim 4, wherein: The determining, based on the high-dimensional feature, the spatiotemporal feature corresponding to the first feature identifier includes: For each first feature identifier, concatenate the high-dimensional features corresponding to the first feature identifier at each first moment to obtain a first concatenated feature corresponding to the first feature identifier; Performing causal-based temporal feature aggregation on the first concatenated feature corresponding to the first feature identifier to obtain the temporal feature corresponding to the first feature identifier; Adding the time series feature and the target high-dimensional feature corresponding to the first feature identifier to obtain a summed feature corresponding to the first feature identifier; Perform causal-based spatial feature aggregation according to the summed feature corresponding to the first feature identifier to obtain the spatiotemporal feature corresponding to the first feature identifier.
6. The method according to claim 5, wherein: Also includes: Determine the second position of the first high-dimensional feature in the map feature matrix based on the vehicle decision data at the second moment and the first position of the second high-dimensional feature in the map feature matrix; the first high-dimensional feature is a high-dimensional feature of the first feature identifier corresponding to the map data at the first moment, and the second high-dimensional feature is a high-dimensional feature of the second feature identifier corresponding to the map data at the second moment; Determine an interpolated high-dimensional feature based on the second position of the second high-dimensional feature, and determine the second high-dimensional feature based on the interpolated high-dimensional feature; The performing causal-based spatial feature aggregation according to the summed feature corresponding to the first feature identifier to obtain the spatiotemporal feature corresponding to the first feature identifier includes: A causal-based spatial feature aggregation is performed according to the sum feature of the first feature identifier corresponding to the map data and the second high-dimensional feature to obtain a spatiotemporal feature corresponding to the first feature identifier.
7. The method according to claim 1, wherein: The decoding of the second feature identifier to obtain the second driving sub-data includes: Determine a data type of the second driving sub-data corresponding to the second feature identifier; Based on the data type of the second driving sub-data corresponding to the second feature identifier, the second feature identifier is decoded to obtain the second driving sub-data.
8. The method according to claim 7, wherein: The decoding of the second feature identifier based on the data type of the second driving sub-data corresponding to the second feature identifier to obtain the second driving sub-data includes: In response to the data type of the second driving sub-data corresponding to the second feature identifier being a structured type, for each driving data parameter in the second driving sub-data, determining the driving data parameter based on a target parameter interval with the second feature identifier corresponding to the driving data parameter as an interval number; Determining the second driving sub-data based on each of the driving data parameters; or, In response to the data type of the second driving sub-data corresponding to the second feature identifier being an unstructured type, the second feature identifier is decompressed to obtain the second driving sub-data.
9. A driving data generating device, comprising: A driving data acquisition module, configured to acquire first driving data at least at a first moment, wherein the first driving data includes first driving sub-data of at least one mode; a driving data encoding module, configured to encode the first driving sub-data and obtain a first feature identifier corresponding to the first driving sub-data; a feature identifier prediction module, configured to perform spatiotemporal prediction on the first feature identifier to obtain a second feature identifier corresponding to second driving sub-data at a second moment; the first moment being before the second moment; a feature identifier decoding module, used for decoding the second feature identifier to obtain the second driving sub-data; A driving data generating module is used to determine second driving data at the second moment based on the second driving sub-data.
10. A computer-readable storage medium, wherein the storage medium stores a computer program, wherein the computer program is used to execute the driving data generating method described in any one of claims 1 to 8.
11. An electronic device, comprising: processor; a memory for storing instructions executable by the processor; The processor is used to read the executable instructions from the memory and execute the instructions to implement the driving data generation method described in any one of claims 1-8 above.