Vehicle control method and device, vehicle, storage medium, program product and chip
By encoding and denoising diffusion processing of vehicle multimodal perception data, and optimizing the sampling direction of trajectory information using probabilistic information, the problem of insufficient trajectory information generation accuracy in intelligent driving of vehicles is solved, thereby improving generation accuracy and driving safety.
Patent Information
- Application Number
- CN202511377514.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-24
- Publication Date
- 2026-01-23
- Estimated Expiration
- 2045-09-24
AI Technical Summary
In existing technologies, the accuracy of vehicle intelligent driving trajectory information generation is insufficient, making it difficult to adaptively optimize the sampling direction of trajectory information, resulting in inaccurate generated target trajectory information.
By acquiring multimodal perception data of the vehicle and encoding it, the probability information of the previous denoising step is used to guide the feature sampling of the current denoising step during the denoising diffusion process, optimizing the sampling direction of trajectory information and generating target trajectory information with better reward scores.
It improves the accuracy of target trajectory information generation, enhances vehicle driving safety and user experience, and better meets users' actual needs.
Smart Images

Figure CN121375835A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates to the technical field of intelligent driving, and particularly relates to a vehicle control method and device, a vehicle, a storage medium, a program product and a chip. BACKGROUND
[0002] Intelligent driving of a vehicle refers to the ability of the vehicle to autonomously drive without human intervention through artificial intelligence, sensors and other technologies. Intelligent driving of the vehicle can rely on the cooperation of artificial intelligence, visual computing, radar, monitoring devices and global positioning systems to enable the terminal to automatically and safely operate the motor vehicle without any human initiative. SUMMARY
[0003] To overcome the problems in the related art, the present disclosure provides a vehicle control method and device, a vehicle, a non-transitory computer-readable storage medium, a chip and a computer program product, which can adaptively optimize the sampling direction of trajectory information, so as to generate target trajectory information with a more optimal reward score, greatly improving the generation accuracy of the target trajectory information.
[0004] According to a first aspect of an embodiment of the present disclosure, a vehicle control method is provided, including: obtaining multi-modal perception data of a vehicle, and encoding the multi-modal perception data to obtain encoded feature data; in any non-first current denoising step in a denoising diffusion process, performing feature sampling according to the encoded feature data and probability information of a previous denoising step to obtain sampling feature data of the current denoising step; wherein the probability information is used to represent probability differences in different reward score trajectory information generated based on the sampling feature data of the previous denoising step; generating target trajectory information according to the sampling feature data corresponding to multiple denoising steps in the denoising diffusion process; and controlling the vehicle to travel according to the target trajectory information.
[0005] In the above embodiment, by obtaining the multi-modal perception data of the vehicle, and encoding the multi-modal perception data to obtain the encoded feature data, and according to the encoded feature data and the probability information of the last denoising step, performing feature sampling to obtain the sampled feature data of the current denoising step at any non-first current denoising step in the denoising diffusion process, wherein the probability information is used to represent the probability difference of the trajectory information based on the sampled feature data of the last denoising step generating different reward scores, and according to the sampled feature data corresponding to the plurality of denoising steps in the denoising diffusion process, generating the target trajectory information, and controlling the vehicle to travel according to the target trajectory information. Therefore, since the "probability information is used to represent the probability difference of the trajectory information based on the sampled feature data of the last denoising step generating different reward scores", when the probability information of the last denoising step is used to guide the sampling direction of the feature sampling of the current denoising step, the sampling direction of the trajectory information can be adaptively optimized, so that the target trajectory information with better reward score is generated, and the generation accuracy of the target trajectory information is greatly improved.
[0006] Optionally, in some embodiments of the present disclosure, according to the encoded feature data and the probability information of the last denoising step, performing feature sampling includes:
[0007] According to the probability information of the last denoising step, determining the feature sampling parameter of the current denoising step;
[0008] According to the feature sampling parameter of the current denoising step, performing feature sampling on the encoded feature data.
[0009] In the above embodiment, in the process of performing feature sampling according to the encoded feature data and the probability information of the last denoising step, the feature sampling parameter of the current denoising step can be determined according to the probability information of the last denoising step, and the feature sampling is performed on the encoded feature data according to the feature sampling parameter of the current denoising step. Therefore, the accuracy of feature sampling can be improved, and the feature sampling of each denoising step can be dynamically adjusted to improve the accuracy of feature sampling.
[0010] Optionally, in some embodiments of the present disclosure, the method further includes:
[0011] Obtaining the positive probability information and the negative probability information based on the last denoising step, wherein the positive probability information is information of a conditional probability of generating first trajectory information based on the sampled feature data of the last denoising step, the negative probability information is information of a conditional probability of generating second trajectory information based on the sampled feature data of the last denoising step, and the reward score of the first trajectory information is higher than the reward score of the second trajectory information;
[0012] Determining the positive probability information and the negative probability information as the probability information of the last denoising step.
[0013] In the above embodiments, the feature sampling of each denoising step can be gradually optimized, and the sampling direction can be guided by the positive probability information and the negative probability information, so as to improve the sampling accuracy and the sampling effect.
[0014] Optionally, in some embodiments of the present disclosure, the feature sampling parameter of the current denoising step is determined according to the probability information of the last denoising step, including:
[0015] The feature sampling parameter of the last denoising step is determined.
[0016] The feature sampling parameter of the last denoising step is adjusted according to the positive probability information and the negative probability information, wherein the feature sampling parameter is used to make the target trajectory information tend to the first trajectory information.
[0017] The adjusted feature sampling parameter is determined as the feature sampling parameter of the current denoising step.
[0018] In the above embodiments, the feature sampling of each denoising step can be gradually optimized to guide the sampling direction, so as to ensure that the generated target trajectory information can tend to the first trajectory information with a higher reward score.
[0019] Optionally, in some embodiments of the present disclosure, the feature sampling parameter of the last denoising step is adjusted according to the positive probability information and the negative probability information, including:
[0020] The weight information is determined, wherein the weight information is used to control the intensity of the adjustment of the feature sampling parameter.
[0021] The parameter adjustment information is determined according to the weight information, the positive probability information and the negative probability information.
[0022] The feature sampling parameter of the last denoising step is adjusted according to the parameter adjustment information.
[0023] In the above embodiments, in the process of adjusting the feature sampling parameter of the last denoising step according to the positive probability information and the negative probability information, the weight information can be determined, wherein the weight information is used to control the intensity of the adjustment of the feature sampling parameter, the parameter adjustment information is determined according to the weight information, the positive probability information and the negative probability information, and the feature sampling parameter of the last denoising step is adjusted according to the parameter adjustment information. Therefore, the feature sampling parameter of the last denoising step can be adaptively adjusted, which can greatly improve the adjustment accuracy and improve the practicability.
[0024] Optionally, in some embodiments of the present disclosure, any denoising step of the denoising diffusion process is executed by a prediction model, wherein the prediction model includes a decoder, a first sub-model connected to the decoder, and a second sub-model.
[0025] The forward probability information and the negative probability information obtained based on the previous denoising step are obtained, including:
[0026] In the previous denoising step, the encoded feature data is sampled by the decoder to obtain the sampling feature data of the previous denoising step;
[0027] The sampling feature data of the previous denoising step is processed by the first sub-model to obtain the forward probability information;
[0028] The sampling feature data of the previous denoising step is processed by the second sub-model to obtain the negative probability information;
[0029] The first sub-model has modeled and learned the mapping relationship between the sampling feature data of the previous denoising step and the forward probability information, and the second sub-model has modeled and learned the mapping relationship between the sampling feature data of the previous denoising step and the negative probability information.
[0030] In the above embodiment, any denoising step of the denoising diffusion process is completed by a pre-trained prediction model, which can include an encoder, a decoder connected to the encoder, a first sub-model and a second sub-model connected to the decoder. The first sub-model has modeled and learned the mapping relationship between the sampling feature data of the previous denoising step and the forward probability information, and the second sub-model has modeled and learned the mapping relationship between the sampling feature data of the previous denoising step and the negative probability information. The prediction accuracy and efficiency of the target trajectory information can be greatly improved, which can be effectively applied to real-time vehicle driving control, and greatly ensures the driving safety.
[0031] Optionally, in some embodiments of the present disclosure, the prediction model further includes an encoder connected to the decoder; wherein the multi-modal perception data is encoded to obtain the encoded feature data, including:
[0032] The multi-modal perception data is encoded by the encoder to obtain the encoded feature data.
[0033] In the above embodiment, the efficiency and accuracy of the encoding process can be greatly improved to ensure that more accurate encoded feature data is extracted, thereby further improving the prediction accuracy and efficiency of the target trajectory information.
[0034] Optionally, in some embodiments of the present disclosure, the multi-modal perception data of the vehicle is obtained, including:
[0035] Collecting environment images of the driving environment of the vehicle and spatial point cloud data of the driving environment;
[0036] Determine the state information and navigation information of the vehicle;
[0037] The environmental image, the spatial point cloud data, the state information of the vehicle, and the navigation information are determined as the multi-modal perception data.
[0038] In the above embodiment, by collecting an environmental image of a driving environment of a vehicle, and spatial point cloud data of the driving environment, and determining state information of the vehicle and navigation information, the environmental image, the spatial point cloud data, the state information of the vehicle, and the navigation information are determined as the multi-modal perception data. In this way, more detailed multi-modal perception data can be collected, and more accurate trajectory information can be predicted.
[0039] According to a second aspect of the embodiments of the present disclosure, a vehicle control apparatus is provided, including: an acquisition unit configured to acquire multi-modal perception data of a vehicle, and encode the multi-modal perception data to obtain encoded feature data; a processing unit configured to, at any non-first current denoising step in a denoising diffusion process, perform feature sampling according to the encoded feature data and probability information of a previous denoising step to obtain sampled feature data of the current denoising step; wherein the probability information is used to represent a probability difference in a case that trajectory information of different reward scores is generated based on the sampled feature data of the previous denoising step; a generation unit configured to generate target trajectory information according to the sampled feature data corresponding to multiple denoising steps in the denoising diffusion process; and a control unit configured to control the vehicle to travel according to the target trajectory information.
[0040] Optionally, in some embodiments of the present disclosure, the processing unit is further configured to:
[0041] determine the feature sampling parameter of the current denoising step according to the probability information of the previous denoising step;
[0042] perform feature sampling on the encoded feature data according to the feature sampling parameter of the current denoising step.
[0043] Optionally, in some embodiments of the present disclosure, the processing unit is further configured to:
[0044] acquire positive probability information and negative probability information based on the previous denoising step, wherein the positive probability information is information of a conditional probability that first trajectory information is generated based on the sampled feature data of the previous denoising step, the negative probability information is information of a conditional probability that second trajectory information is generated based on the sampled feature data of the previous denoising step, and a reward score of the first trajectory information is higher than a reward score of the second trajectory information;
[0045] determine the positive probability information and the negative probability information as the probability information of the previous denoising step.
[0046] Optionally, in some embodiments of the present disclosure, the processing unit is further configured to:
[0047] determine the feature sampling parameter of the previous denoising step;
[0048] According to the positive probability information and the negative probability information, the feature sampling parameter of the last denoising step is adjusted, wherein the feature sampling parameter is used to make the target trajectory information tend to the first trajectory information.
[0049] The adjusted feature sampling parameter is determined as the feature sampling parameter of the current denoising step.
[0050] According to a third aspect of the embodiments of the present disclosure, a vehicle is provided, comprising a processor, a memory for storing processor-executable instructions, wherein the processor is configured to implement the steps of the vehicle control method provided by the first aspect of the present disclosure.
[0051] According to a fourth aspect of the embodiments of the present disclosure, a non-transitory computer-readable storage medium is provided, when the instructions in the storage medium are executed by the processor of the mobile terminal, the mobile terminal can execute a vehicle control method, the method comprising: obtaining multi-modal perception data of the vehicle, and encoding the multi-modal perception data to obtain encoded feature data; at any non-first current denoising step in the denoising diffusion process, performing feature sampling according to the encoded feature data and the probability information of the last denoising step to obtain the sampling feature data of the current denoising step; wherein the probability information is used to represent the probability difference under the condition that the trajectory information based on the sampling feature data of the last denoising step generates different reward scores; generating target trajectory information according to the sampling feature data corresponding to the plurality of denoising steps in the denoising diffusion process; and controlling the vehicle to travel according to the target trajectory information.
[0052] According to a fifth aspect of the embodiments of the present disclosure, a chip is provided, comprising a processing circuit and an interface circuit; wherein the interface circuit is used to read instructions, and the interface circuit sends the instructions to the processing circuit to make the processing circuit execute the vehicle control method as proposed in the first aspect of the present disclosure.
[0053] According to a sixth aspect of the embodiments of the present disclosure, a computer program product is provided, comprising a computer program which is executed by a processor to implement the vehicle control method as proposed in the first aspect of the present disclosure.
[0054] The technical solutions provided by the embodiments of the present disclosure can include the following beneficial effects:
[0055] By acquiring the multi-modal perception data of the vehicle, and encoding the multi-modal perception data to obtain encoded feature data, and at any non-first current denoising step in the denoising diffusion process, according to the encoded feature data and the probability information of the last denoising step, feature sampling is performed to obtain the sampling feature data of the current denoising step, wherein the probability information is used to represent the probability difference under the condition that the trajectory information of different reward scores is generated based on the sampling feature data of the last denoising step, and the target trajectory information is generated according to the sampling feature data corresponding to the plurality of denoising steps in the denoising diffusion process, and the vehicle is controlled to travel according to the target trajectory information. Therefore, since the "probability information is used to represent the probability difference under the condition that the trajectory information of different reward scores is generated based on the sampling feature data of the last denoising step", when the sampling direction of the feature sampling of the current denoising step is guided based on the probability information of the last denoising step, the sampling direction of the trajectory information can be adaptively optimized, so that the target trajectory information with better reward score is generated, the generation accuracy of the target trajectory information is greatly improved, and the actual needs of users can be better met and the vehicle experience is improved.
[0056] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. BRIEF DESCRIPTION OF DRAWINGS
[0057] The accompanying drawings, which are incorporated into and form part of the specification, illustrate embodiments consistent with the present disclosure and, together with the specification, serve to explain the principles of the present disclosure.
[0058] Figure 1 is a flowchart of a vehicle control method according to some embodiments of the present disclosure;
[0059] Figure 2 is a flowchart of another vehicle control method according to some embodiments of the present disclosure;
[0060] Figure 3 is a flowchart of yet another vehicle control method according to some embodiments of the present disclosure;
[0061] Figure 4 is a structural diagram of a vehicle control device according to some embodiments of the present disclosure;
[0062] Figure 5 is a functional block diagram of a vehicle according to an exemplary embodiment;
[0063] Figure 6 is a structural diagram of a chip according to an embodiment of the present disclosure;
[0064] Figure 7 is a structural diagram of another chip according to an embodiment of the present disclosure. DETAILED DESCRIPTION
[0065] Some embodiments of the present disclosure will be described in detail herein with reference to the drawings, in which some embodiments of the present disclosure are shown by way of illustration. The description of the methods, devices and / or systems described herein can be relevant to the drawings when the following description is made. Unless otherwise indicated, the same numbers in different drawings indicate the same or similar elements. Various changes, modifications and equivalents of the methods, devices and / or systems described herein will become apparent to those skilled in the art after understanding the present disclosure. For example, the order of the operations described herein is merely an example, and is not limited to those set forth herein, but can be changed as apparent after understanding the present disclosure, except for the operations that must be performed in a specific order. In addition, the description of features known in the art can be omitted for the sake of clarity and brevity.
[0066] The implementations described in some embodiments of the present disclosure below do not represent all implementations consistent with the present disclosure. Rather, they are merely examples of apparatuses and methods consistent with some aspects of the present disclosure as detailed in the appended claims.
[0067] It should be noted that all actions of acquiring signals, information or data in the present disclosure are carried out in accordance with the corresponding data protection regulations and policies of the country where the device is located, and with the authorization given by the owner of the corresponding device.
[0068] Figure 1 is a flowchart of a vehicle control method according to some embodiments of the present disclosure, as shown in Figure 1 The vehicle control method can be used in electronic devices such as mobile terminals, vehicle-mounted devices, and can be applied in intelligent driving scenarios of vehicles, or can also be used in assisted driving scenarios, without limitation. The vehicle control method comprises the following steps:
[0069] Step S101: acquiring multi-modal perception data of a vehicle, and encoding the multi-modal perception data to obtain encoded feature data.
[0070] Optionally, in some embodiments, the multi-modal perception data can be various modal perception data collected during driving of the vehicle. For example, image data of the driving environment collected by a vehicle-mounted camera, vehicle state data collected by sensors (such as vehicle speed, direction, control parameters, etc.), and data of obstacles in the driving environment, sound data, etc.
[0071] Optionally, in some embodiments, the acquiring the multi-modal perception data of the vehicle comprises: collecting an environmental image of a driving environment of the vehicle and spatial point cloud data of the driving environment; determining state information and navigation information of the vehicle; and determining the environmental image, the spatial point cloud data, the state information and the navigation information of the vehicle as the multi-modal perception data. In this way, more detailed multi-modal perception data can be collected, which can assist in predicting more accurate trajectory information.
[0072] Optionally, in some embodiments, in the process of collecting the environmental image of the driving environment of the vehicle, the environmental image of the driving environment can be captured in real time by a vehicle-mounted visual sensor (such as a camera or a laser radar), or the environmental image of the driving environment can be captured based on a certain period. In the process of collecting the spatial point cloud data of the driving environment, a laser pulse can be emitted by a vehicle-mounted laser radar, and the reflection time can be measured to generate three-dimensional point cloud data, which can be used as spatial point cloud data to describe the spatial features of the driving environment of the vehicle in a three-dimensional manner. In the process of determining the state information of the vehicle, the state can refer to the driving state of the vehicle, such as speed, direction, acceleration, fuel consumption, etc. In the process of determining the navigation information of the vehicle, the navigation information can be, for example, navigation data in the vehicle-mounted device, such as global navigation satellite system data, without limitation.
[0073] Of course, any possible driving environment data and vehicle information can be collected during vehicle driving to predict trajectory information, without limitation.
[0074] Optionally, in some embodiments, after the multi-modal perception data of the vehicle is collected, the multi-modal perception data can be encoded to obtain encoded feature data. The encoded feature data is used for uniform feature representation of the multi-modal perception data, and the encoded feature data is, for example, a feature vector obtained by encoding. For example, the multi-modal perception data can be feature extracted, and some encoding techniques can be combined to convert the extracted features to uniform feature representation, i.e., to obtain the encoded feature data, without limitation.
[0075] Step S102: In any non-first current denoising step in the denoising diffusion process, feature sampling is performed according to the encoded feature data and the probability information of the last denoising step to obtain sampling feature data of the current denoising step, wherein the probability information is used to represent the probability difference of different reward score trajectory information based on the sampling feature data of the last denoising step.
[0076] The denoising diffusion process is a data processing process in artificial intelligence, and the goal is to gradually recover the original structured data from noisy data. The noisy data may be, for example, the original input data, such as the encoded feature data in the embodiments of the present disclosure, and the output may be the recovered data, such as the target trajectory information in the embodiments of the present disclosure. The target trajectory information may be information about the predicted future trajectory of the vehicle, and the target trajectory information is used to assist intelligent driving or autonomous driving of the vehicle. The denoising diffusion process can be applied to diffusion models in artificial intelligence. The process can be implemented by a reverse Markov chain, and is symmetric to the forward noise adding process.
[0077] Optionally, the denoising diffusion process can include multiple denoising steps, each denoising step can be understood as an iterative step, and in each denoising step, the input data is processed once.
[0078] Optionally, in the embodiments of the present disclosure, the encoded feature data can be processed based on the denoising diffusion process, and each denoising step in the denoising diffusion process can be performed to gradually denoise the encoded feature data and obtain optimal target trajectory information.
[0079] Optionally, in some embodiments, at any current denoising step other than the first denoising step in the denoising diffusion process, feature sampling can be performed based on the encoded feature data and the probability information of the last denoising step. The sampled feature can be referred to as the sampled feature data of the current denoising step, that is, each denoising step outputs a sampled feature data, and the multiple sampled feature data output by the multiple denoising steps can be used to predict the target trajectory information.
[0080] The step of "sampling features according to the encoded feature data and the probability information of the previous de-noising step" is described as follows: each de-noising step can output a probability information, which is used to represent the probability difference of generating trajectory information with different reward scores based on the sampled feature data of the corresponding de-noising step. For example, the trajectory information can include at least two types: first trajectory information and second trajectory information, wherein the reward score of the first trajectory information can be higher than that of the second trajectory information, and the first trajectory information is, for example, trajectory information that tends to be safe and reliable, and the second trajectory information is, for example, relatively dangerous and unreliable trajectory information. In the process of performing the current de-noising step, the probability information of the previous de-noising step can be obtained, which is used to represent the probability difference of generating trajectory information with different reward scores based on the sampled feature data of the previous de-noising step. The current de-noising step has corresponding parameters for feature sampling, which can be guided and adjusted based on the probability information of the previous de-noising step, that is, the parameters for feature sampling are adjusted, and then the feature sampling of the current de-noising step is performed based on the adjusted parameters. Moreover, the "probability difference" can guide the generation of target trajectory information that tends to the first trajectory information, so as to improve the reliability of the target trajectory information and improve the driving safety.
[0081] Optionally, in some embodiments, each de-noising step can be executed one by one for multiple de-noising steps of the de-noising diffusion process. In the first de-noising step, the encoded feature data can be directly sampled to obtain the sampled feature data of the current de-noising step, and at the same time, a probability information can be output at this de-noising step. Then, the next de-noising step is executed, which can be regarded as the "current de-noising step". In the "current de-noising step", the parameters required for feature sampling of the current de-noising step are determined according to the "output probability information", and then the encoded feature data is sampled by the "parameters required for feature sampling" to obtain the sampled feature data of the current de-noising step. Moreover, in the current de-noising step, a probability information can also be output based on the sampled feature data of the current de-noising step, which is used for feature sampling of the encoded feature data in the next de-noising step. That is, due to the successive execution of each de-noising step, the "next de-noising step" can be regarded as the next "current de-noising step" for iterative execution of the de-noising diffusion process.
[0082] Optionally, in the process of implementing the feature sampling according to the encoded feature data and the probability information of the previous de-noising step, the feature sampling parameters of the current de-noising step can be determined according to the probability information of the previous de-noising step, and the encoded feature data can be sampled according to the feature sampling parameters of the current de-noising step. In this way, the precision of feature sampling can be improved, the dynamic adjustment of feature sampling of each de-noising step can be realized, and the accuracy of feature sampling can be improved.
[0083] For example, the probability information of the last de-noising step is used to represent the probability difference in the case that the trajectory information with different reward scores is generated based on the sampling feature data of the last de-noising step. The probability difference can be a quantitative value. For example, based on the sampling feature data of the last de-noising step, trajectory information A and trajectory information B can be generated, the reward score of the trajectory information A is higher than that of the trajectory information B, the probability information A (for example, the probability gradient A or other arbitrary probability quantitative value) of generating the trajectory information A based on the sampling feature data of the last de-noising step can be determined, and the probability information B (for example, the probability gradient B or other arbitrary probability quantitative value) of generating the trajectory information B based on the sampling feature data of the last de-noising step can be determined. Then, the probability difference between the probability information A and the probability information B is determined, for example, the probability information A and the probability information B are subtracted to obtain the probability difference as the probability information of the last de-noising step. Then, the sampling of the current de-noising step can be guided based on the probability difference between the probability information A and the probability information B. For example, the sampling parameter P of the current de-noising step (that is, the sampling parameter P of each de-noising step can be determined based on the probability information of the last de-noising step, and the sampling parameter P is, for example, the sampling type, the sampling granularity, etc., which is not limited) can be determined based on the probability difference between the probability information A and the probability information B. Then, the feature sampling parameter P of the current de-noising step can be used to sample the encoded feature data, and the sampled feature data can be referred to as the sampling feature data of the current de-noising step. The sampling feature data of the corresponding de-noising step can be obtained through the above process at any current de-noising step in the de-noising diffusion process.
[0084] Optionally, in the above process, since the probability information is used to represent the probability difference in the case that the trajectory information with different reward scores is generated based on the sampling feature data of the last de-noising step, the sampling direction of the feature sampling of the current de-noising step can be guided based on the probability information of the last de-noising step, the sampling direction of the trajectory information can be adaptively optimized, the target trajectory information with a better reward score can be generated, and the generation accuracy of the target trajectory information is greatly improved.
[0085] Step S103: generating the target trajectory information according to the sampling feature data corresponding to the multiple de-noising steps in the de-noising diffusion process.
[0086] Optionally, the sampling feature data corresponding to each de-noising step in the de-noising diffusion process can be obtained, and then the target trajectory information can be generated according to the sampling feature data corresponding to the multiple de-noising steps in the de-noising diffusion process. For example, the target trajectory information can be obtained by trajectory prediction based on the multiple sampling feature data.
[0087] Step S104: controlling the vehicle to travel according to the target trajectory information.
[0088] Optionally, after the target trajectory information is generated, the vehicle can be controlled to travel based on the target trajectory information. For example, a travel control instruction is generated based on the target trajectory information, and the vehicle is controlled to travel through the travel control instruction, which is not limited.
[0089] In this embodiment, the multi-modal perception data of the vehicle is obtained, and the multi-modal perception data is encoded to obtain encoded feature data. At any non-first current denoising step in the denoising diffusion process, the encoded feature data and the probability information of the last denoising step are used for feature sampling to obtain the sampling feature data of the current denoising step, wherein the probability information is used to represent the probability difference of the trajectory information under different reward scores based on the sampling feature data of the last denoising step, and the target trajectory information is generated based on the sampling feature data corresponding to the plurality of denoising steps in the denoising diffusion process, and the vehicle is controlled to travel according to the target trajectory information. Therefore, since the "probability information is used to represent the probability difference of the trajectory information under different reward scores based on the sampling feature data of the last denoising step", the sampling direction of the feature sampling of the current denoising step is guided based on the probability information of the last denoising step, which can adaptively optimize the sampling direction of the trajectory information, so as to generate the target trajectory information with better reward scores, greatly improve the generation accuracy of the target trajectory information, and better meet the actual needs of users and improve the vehicle experience.
[0090] Figure 2 is a flowchart of another vehicle control method according to some embodiments of the present disclosure, as shown in Figure 2 The vehicle control method can be used in an electronic device, such as a mobile terminal or a vehicle-mounted device. The vehicle control method can be applied in a smart driving scenario of a vehicle, or can be used in an assisted driving scenario, which is not limited. The vehicle control method includes the following steps:
[0091] Step S201: obtaining multi-modal perception data of a vehicle, and encoding the multi-modal perception data to obtain encoded feature data.
[0092] The description of S201 can be specifically referred to the above embodiments, which will not be repeated here.
[0093] Step S202: performing a denoising diffusion process to obtain positive probability information and negative probability information based on the sampling feature data obtained in the last denoising step, wherein the positive probability information is information of a conditional probability of generating first trajectory information based on the sampling feature data in the last denoising step, and the negative probability information is information of a conditional probability of generating second trajectory information based on the sampling feature data in the last denoising step, and the reward score of the first trajectory information is higher than that of the second trajectory information.
[0094] The explanation of the denoising diffusion process can be specifically referred to the above embodiments.
[0095] Optionally, the trajectory information with different reward scores can include the first trajectory information and the second trajectory information, and the reward score of the first trajectory information is higher than that of the second trajectory information. Optionally, the higher the reward score is, the safer and more reliable the corresponding trajectory information is, and the lower the reward score is, the less safe and reliable the corresponding trajectory information is. For example, the first trajectory information can be the future trajectory information of the vehicle under the driving behavior conforming to safety and comfort, and the second trajectory information can be the future trajectory information of the vehicle under the risky driving behavior. In addition, the positive probability information is information of a conditional probability of generating the first trajectory information based on the sampling feature data in the last denoising step, and the negative probability information is information of a conditional probability of generating the second trajectory information based on the sampling feature data in the last denoising step.
[0096] Step S203: determining the positive probability information and the negative probability information as the probability information of the last denoising step.
[0097] Optionally, the positive probability information and the negative probability information based on the last denoising step can be determined as the probability information of the last denoising step.
[0098] Optionally, the positive probability information and the negative probability information obtained in the last denoising step can be used to guide the sampling direction of the feature sampling in the current denoising step, so as to gradually optimize the feature sampling process and ensure that more accurate target trajectory information is finally generated.
[0099] Step S204: determining the feature sampling parameter of the last denoising step.
[0100] Optionally, in each denoising step of the denoising diffusion process, the feature sampling parameter can be used to sample the encoded feature data to obtain the sampling feature data of the corresponding denoising step.
[0101] Optionally, in the current denoising step, the feature sampling parameter used in the current denoising step can be determined based on the feature sampling parameter used in the last denoising step.
[0102] Step S205: adjusting the feature sampling parameter of the last de-noising step according to the positive probability information and the negative probability information, and determining the adjusted feature sampling parameter as the feature sampling parameter of the current de-noising step, wherein the feature sampling parameter is used to make the target trajectory information tend to the first trajectory information.
[0103] Optionally, in the embodiments of the present disclosure, after the feature sampling parameter used in the last de-noising step is determined, the feature sampling parameter of the last de-noising step can be adjusted based on the positive probability information and the negative probability information of the last de-noising step, and then the adjusted feature sampling parameter is determined as the feature sampling parameter of the current de-noising step. As can be seen, in each de-noising step, the feature sampling parameter of the current de-noising step can be optimized using this method, thereby the feature sampling of each de-noising step can be gradually optimized to guide the sampling direction, and ensure that the generated target trajectory information can tend to the first trajectory information with a higher reward score.
[0104] Optionally, in some embodiments, in the process of adjusting the feature sampling parameter of the last de-noising step according to the positive probability information and the negative probability information, weight information can be determined, wherein the weight information is used to control the intensity of the adjustment of the feature sampling parameter, and the parameter adjustment information is determined according to the weight information, the positive probability information and the negative probability information, and the feature sampling parameter of the last de-noising step is adjusted according to the parameter adjustment information. Thus, the feature sampling parameter of the last de-noising step can be adaptively adjusted, which can greatly improve the accuracy of the adjustment and improve the practicability.
[0105] For example, the probability difference between the positive probability information and the negative probability information can be determined, and the probability difference is weighted by the weight information, and the feature sampling parameter of the last de-noising step is adjusted based on the weighted result to obtain the feature sampling parameter of the current de-noising step. In addition, the "feature sampling parameter of the current de-noising step", and the "positive probability information and the negative probability information output by the current de-noising step" can also be used to determine the "feature sampling parameter of the next de-noising step", which will not be described here.
[0106] Optionally, in the process of determining the parameter adjustment information according to the weight information, the positive probability information and the negative probability information, the first weight parameter corresponding to the positive probability information and the second weight parameter corresponding to the negative probability information can be determined according to the weight information, and the positive probability information is weighted according to the first weight parameter, and the negative probability information is weighted according to the second weight parameter, and the difference information between the weighted positive probability information and the weighted negative probability information is determined, and the difference information is determined as the parameter adjustment information. Thus, more accurate parameter adjustment information can be determined to obtain better feature sampling parameters.
[0107] In an example, the first weight information and the positive probability information can be multiplied to obtain a multiplication result, the second weight information and the negative probability information can be multiplied to obtain another multiplication result, and the multiplication result and the another multiplication result can be subtracted to obtain a difference information, which is used as the parameter adjustment information to adjust the feature sampling parameter of the previous denoising step. The adjusted feature sampling parameter is used for feature sampling of the current denoising step.
[0108] Step S206: According to the feature sampling parameter of the current denoising step, the encoded feature data is sampled to obtain the sampling feature data of the current denoising step.
[0109] Optionally, after the feature sampling parameter of the current denoising step is obtained by optimization, the encoded feature data can be sampled based on the feature sampling parameter of the current denoising step to obtain the sampling feature data of the current denoising step. In an example, if the encoded feature data is sampled by the decoder, the encoder can be configured using the feature sampling parameter of the current denoising step, and the configured encoder can be used to sample the encoded feature data to obtain the sampling feature data of the current denoising step, which is not limited.
[0110] Step S207: According to the sampling feature data corresponding to the plurality of denoising steps in the denoising diffusion process, the target trajectory information is generated.
[0111] Step S208: According to the target trajectory information, the vehicle is controlled to drive.
[0112] The description of S207-S208 can be specifically referred to the above-mentioned embodiments, which will not be repeated here.
[0113] In this embodiment, since the probability information is used to represent the probability difference of the trajectory information generated based on the sampling feature data of the last denoising step under different reward score conditions, the sampling direction of the feature sampling of the current denoising step can be guided based on the probability information of the last denoising step, which can adaptively optimize the sampling direction of the trajectory information, so as to generate target trajectory information with better reward score, and greatly improve the generation accuracy of the target trajectory information. By obtaining the positive probability information and the negative probability information based on the last denoising step, wherein the positive probability information is information of a conditional probability of generating first trajectory information based on the sampling feature data of the last denoising step, the negative probability information is information of a conditional probability of generating second trajectory information based on the sampling feature data of the last denoising step, the reward score of the first trajectory information is higher than that of the second trajectory information, and the positive probability information and the negative probability information are determined as the probability information of the last denoising step. And determine the feature sampling parameter of the last denoising step, and adjust the feature sampling parameter of the last denoising step according to the positive probability information and the negative probability information, wherein the feature sampling parameter is used to make the target trajectory information tend to the first trajectory information, and the adjusted feature sampling parameter is determined as the feature sampling parameter of the current denoising step. The feature sampling parameter of the current denoising step can be used to perform feature sampling on the encoded feature data to obtain the sampling feature data of the current denoising step. Thus, the feature sampling of each denoising step can be gradually optimized to guide the sampling direction, so as to ensure that the generated target trajectory information can tend to the first trajectory information with higher reward score, and can also better meet the actual needs of the user and improve the vehicle experience. It also enables adaptive adjustment of the feature sampling parameter of the last denoising step, which can greatly improve the accuracy of the adjustment and improve the practicality. Wherein, the higher reward score means that the closer the parameter (data) is to the user's setting requirement, the higher the reward score is.
[0114] Figure 3 is a flowchart of still another vehicle control method according to some embodiments of the present disclosure, as Figure 3 shown, the vehicle control method can be used in an electronic device, such as a mobile terminal, a vehicle-mounted device, the vehicle control method can be applied in the intelligent driving scene of the vehicle, or also can be used in the assisted driving scene, which is not limited. The vehicle control method comprises the following steps:
[0115] Step S301: obtaining multi-modal perception data of the vehicle.
[0116] The description of S301 can be specifically referred to the above embodiments, which will not be repeated here.
[0117] Step S302: performing any denoising step of the denoising diffusion process by a prediction model, wherein the prediction model comprises a decoder, an encoder connected to the decoder, a first sub-model connected to the decoder, and a second sub-model connected to the decoder.
[0118] Optionally, in the embodiment, any denoising step of the denoising diffusion process can be implemented by an artificial intelligence (AI) model. The AI model is used to predict trajectory information, and the AI model can be referred to as a prediction model. The prediction model can be a trained model.
[0119] Optionally, the prediction model can include an encoder, a decoder connected to the encoder, a first sub-model connected to the decoder, and a second sub-model connected to the decoder. The first sub-model has modeled and learned a mapping relationship between the sampling feature data of the previous denoising step and the positive probability information, and the second sub-model has modeled and learned a mapping relationship between the sampling feature data of the previous denoising step and the negative probability information. The encoder can be used for feature extraction and encoding of the multi-modal perception data to obtain encoded feature data, which can be input into the decoder. The decoder can be used for feature sampling of the encoded feature data. The backbone network of the decoder can also be connected to the first sub-model and the second sub-model. The first sub-model can be obtained by fine-tuning in a positive driving behavior, which can strengthen the generation of trajectory information of the vehicle in the future in a driving behavior conforming to safety and comfort. The second sub-model can be obtained by fine-tuning in a negative driving behavior, which can strengthen the generation of trajectory information of the vehicle in the future in a risky driving behavior.
[0120] Optionally, during the training of the prediction model, the encoder can be fixed, i.e., the encoder is not trained, while the decoder and the first sub-model and the second sub-model connected to the decoder are trained. The prediction model can be trained on a large amount of human driving data by imitation learning, and the output is trajectory information of the vehicle in the future for a period of time, which is not limited.
[0121] Optionally, the first sub-model described above can be implemented using a fine-tuning algorithm Low-Rank Adaptation (LoRA). Alternatively, the second sub-model described above can be implemented using the fine-tuning algorithm LoRA. Alternatively, the first sub-model and the second sub-model described above can be implemented using the fine-tuning algorithm LoRA, which is not limited.
[0122] Step S303: encoding the multi-modal perception data by the encoder to obtain encoded feature data.
[0123] Optionally, in any denoising step of the denoising diffusion process performed by the prediction model, the multi-modal perception data can be encoded by an encoder to obtain encoded feature data. The encoded feature data can be input into a decoder to perform each denoising step in the denoising diffusion process.
[0124] Step S304: In the execution of the previous denoising step, the encoded feature data is feature-sampled by a decoder to obtain the sampled feature data of the previous denoising step.
[0125] Optionally, in the execution of the previous denoising step, the encoded feature data can be feature-sampled by a decoder in the prediction model to obtain the sampled feature data of the previous denoising step. That is, the decoder is called in each denoising step to sample the encoded feature data output by the encoder to obtain the sampled feature data under the corresponding denoising step. And the "sampled feature data of the previous denoising step" can be used to guide the sampling direction of the feature sampling process of the current denoising step.
[0126] Step S305: The sampled feature data of the previous denoising step is processed by a first sub-model to obtain forward probability information.
[0127] Step S306: The sampled feature data of the previous denoising step is processed by a second sub-model to obtain negative probability information, and the forward probability information and the negative probability information are determined as the probability information of the previous denoising step.
[0128] Optionally, after the execution of the previous denoising step and the obtaining of the sampled feature data of the previous denoising step, the sampled feature data can be further input into the first sub-model and the second sub-model, the sampled feature data of the previous denoising step is processed by the first sub-model to obtain the forward probability information, the forward probability information is information of a conditional probability of generating first trajectory information based on the sampled feature data of the previous denoising step, the sampled feature data of the previous denoising step is processed by the second sub-model to obtain the negative probability information, the negative probability information is information of a conditional probability of generating second trajectory information based on the sampled feature data of the previous denoising step, the reward score of the first trajectory information is higher than that of the second trajectory information, and then the forward probability information and the negative probability information are determined as the probability information of the previous denoising step.
[0129] The construction and training of the above model are explained as follows:
[0130] For example, multi-modal perception data can be collected, which can be represented as s (state) and used to describe the state of the driving environment in an intelligent driving scenario (e.g., by collecting images through an in-vehicle camera or constructing spatial point cloud data), and also used to describe the state of the vehicle (e.g., speed, acceleration, driving direction, etc.), navigation information of the vehicle, and the like. These data can be used as input of the model, and the output of the model can be represented as a (action) and used to represent a trajectory corresponding to N points in a fixed time T in the future, for example, 60 coordinate points (6s (i.e., the corresponding time T is 6s), 10 Hz (Hz, indicating a sampling frequency of 10 Hz)), that is, the output of the model is the future trajectory information of the vehicle (which can be an optional example of the target trajectory information described above). In the model training process, the policy function can be represented as π (a|s), indicating the probability of taking action a when the input state is s, and the reward function used in the model training can be represented as R (s, a) and used to estimate the reward score. The dimensions of the score can include: collision, out-of-bound, comfort, driving habit, and the like. In the embodiments of the present disclosure, the state s can be used to represent the multi-modal perception data described above, and the action a can be used to represent the predicted trajectory information described above.
[0131] For example, multi-modal perception data can be collected and encoded to obtain encoded feature data, and then the encoded feature data can be input into a model (an optional example of the prediction model described above). The decoder in the model can perform feature sampling based on the encoded feature data in the current denoising step to obtain sampled feature data, which can be used to predict trajectory information. The model can further predict trajectory information based on the sampled feature data, and then evaluate the predicted trajectory information based on various dimensions of the scores described above to obtain a reward score. Different trajectory information can be predicted, and each trajectory information corresponds to a reward score. The model can also evaluate the probability difference of the trajectory information with different reward scores to guide the feature sampling process in the next denoising step, until the final target trajectory information is output. In this process, the reward score of the trajectory information is calculated based on various dimensions such as collision, out-of-bound, comfort, and driving habit. Therefore, the final output trajectory information can have good performance in the various dimensions of the scores (e.g., no collision, no out-of-bound, high comfort, or compliance with the user's driving habit, and the like), thereby improving the prediction accuracy and effect of the trajectory information while making the output trajectory information more in line with the actual needs of the user and improving the user's driving experience.
[0132] For example, assuming that it is the kth round of training, the model (an optional example of the prediction model described above) optimized in the (k-1) th round can be represented as π (k-1)(a|s), the model obtained in the kth round of optimization (an optional example of the prediction model described above) can be represented as π (k) (a|s). The model can be a model trained based on imitation learning (encoder + decoder of diffusion model). The model can include an encoder, a decoder connected to the encoder, the decoder being a decoder based on a diffusion model, and can be used to perform a denoising diffusion process, wherein the encoder can not be trained, and the decoder based on the diffusion model can be divided into two branches, one of which can be represented as (an optional example of the first sub-model described above), and the other of which can be represented as (an optional example of the second sub-model described above). Wherein the input of the encoder can be s, the output of the encoder can be some representation of s (an optional example of the encoding feature data described above), and the decoder based on the diffusion model can receive the representation of the encoder output and output a (an optional example of the predicted trajectory information described above).
[0133] For example, the loss function of the above can be represented as:
[0134]
[0135] The loss function of the above can be represented as:
[0136]
[0137] Wherein, the state s can be used to represent the "multimodal perception data" described above, the action a can be used to represent the predicted trajectory information described above, p(·) is the state distribution of s, π k-1 (·|s) is the policy of the last iteration, and the reward function used in model training can be represented as R(s, a), σ(x) represents the sigmoid function, t represents the noise adding time step of the diffusion model, T represents the total noise adding time step of the diffusion model, z represents the Gaussian noise, N(0, I) represents the standard Gaussian distribution, I represents the unit covariance, u represents that the noise adding time step is uniformly distributed, and E() represents the expectation.
[0138] In addition, in the process of model training, the loss value output by the loss function of (an optional example of the positive probability information described above) and the loss value output by the loss function of (an optional example of the negative probability information described above) can be weighted to obtain an operation result which can guide the training process of the entire model. The formula of the weighted operation is as follows:
[0139]
[0140] wherein, represents a loss value output by a loss function of represents a loss value output by a loss function of (1+ω) can be an optional example of the first weight parameter, and ω can be an optional example of the second weight parameter.
[0141] Step S307: In the current denoising step, feature sampling is performed according to the encoded feature data and the probability information of the previous denoising step, to obtain the sampled feature data of the current denoising step.
[0142] Step S308: Target trajectory information is generated according to the sampled feature data corresponding to the plurality of denoising steps in the denoising diffusion process.
[0143] Step S309: The vehicle is controlled to travel according to the target trajectory information.
[0144] The description of S307-S309 can be specifically referred to the above embodiments, which will not be repeated here.
[0145] In this embodiment, since the probability information is used to represent the probability difference in the case that the trajectory information based on the sampled feature data of the previous denoising step generates different reward scores, the sampling direction of the feature sampling of the current denoising step can be guided based on the probability information of the previous denoising step, which can adaptively optimize the sampling direction of the trajectory information, so as to generate target trajectory information with better reward scores, greatly improving the generation accuracy of the target trajectory information. Moreover, since any denoising step of the denoising diffusion process is completed by the pre-trained prediction model, the prediction model can include an encoder, a decoder connected to the encoder, a first sub-model and a second sub-model connected to the decoder. The first sub-model has modeled and learned the mapping relationship between the sampled feature data of the previous denoising step and the positive probability information, and the second sub-model has modeled and learned the mapping relationship between the sampled feature data of the previous denoising step and the negative probability information. The encoder is used to encode the multi-modal perception data to obtain the encoded feature data, which can greatly improve the prediction accuracy and efficiency of the target trajectory information, and can be effectively applied to real-time vehicle driving control, greatly ensuring the driving safety.
[0146] Figure 4 is a structural diagram of a vehicle control device according to some embodiments of the present disclosure.
[0147] As shown in Figure 4 , the vehicle control device 40 includes:
[0148] The acquisition unit 401 is configured to acquire multi-modal perception data of a vehicle, and encode the multi-modal perception data to obtain encoded feature data.
[0149] The processing unit 402 is configured to, in any non-first current denoising step in a denoising diffusion process, perform feature sampling on the encoded feature data and probability information of a previous denoising step to obtain sampled feature data of the current denoising step, where the probability information is used to represent a probability difference in a case where trajectory information based on the sampled feature data of the previous denoising step generates different reward scores.
[0150] The generation unit 403 is configured to generate target trajectory information according to the sampled feature data corresponding to a plurality of denoising steps in the denoising diffusion process.
[0151] The control unit 404 is configured to control the vehicle to travel according to the target trajectory information.
[0152] Optionally, in some embodiments of the present disclosure, the processing unit 402 is configured to:
[0153] determine the feature sampling parameter of the current denoising step according to the probability information of the previous denoising step;
[0154] perform feature sampling on the encoded feature data according to the feature sampling parameter of the current denoising step.
[0155] Optionally, in some embodiments of the present disclosure, the processing unit 402 is configured to:
[0156] obtain positive probability information and negative probability information based on the previous denoising step, where the positive probability information is information of a conditional probability of generating first trajectory information based on the sampled feature data of the previous denoising step, the negative probability information is information of a conditional probability of generating second trajectory information based on the sampled feature data of the previous denoising step, and a reward score of the first trajectory information is higher than a reward score of the second trajectory information;
[0157] determine the positive probability information and the negative probability information as the probability information of the previous denoising step.
[0158] Optionally, in some embodiments of the present disclosure, the processing unit 402 is configured to:
[0159] determine the feature sampling parameter of the previous denoising step;
[0160] adjust the feature sampling parameter of the previous denoising step according to the positive probability information and the negative probability information, where the feature sampling parameter is used to make the target trajectory information tend to the first trajectory information;
[0161] determine the adjusted feature sampling parameter as the feature sampling parameter of the current denoising step.
[0162] Optionally, in some embodiments of the present disclosure, the processing unit 402 is configured to:
[0163] determine weight information, wherein the weight information is used to control the intensity of the adjustment of the feature sampling parameters;
[0164] determine parameter adjustment information according to the weight information, the positive probability information and the negative probability information;
[0165] adjust the feature sampling parameters of the previous denoising step according to the parameter adjustment information.
[0166] Optionally, in some embodiments of the present disclosure, any denoising step of the denoising diffusion process is performed by a prediction model, wherein the prediction model comprises a decoder, a first sub-model and a second sub-model connected to the decoder.
[0167] wherein the processing unit 402 is configured to:
[0168] in the execution of the previous denoising step, perform feature sampling on the encoded feature data by the decoder to obtain the sampled feature data of the previous denoising step;
[0169] process the sampled feature data of the previous denoising step by the first sub-model to obtain the positive probability information;
[0170] process the sampled feature data of the previous denoising step by the second sub-model to obtain the negative probability information;
[0171] wherein the first sub-model has modeled and learned the mapping relationship between the sampled feature data of the previous denoising step and the positive probability information, and the second sub-model has modeled and learned the mapping relationship between the sampled feature data of the previous denoising step and the negative probability information.
[0172] Optionally, in some embodiments of the present disclosure, the prediction model further comprises an encoder connected to the decoder; wherein the acquisition unit 401 is configured to:
[0173] encode the multi-modal perception data by the encoder to obtain the encoded feature data.
[0174] Optionally, in some embodiments of the present disclosure, the acquisition unit 401 is configured to:
[0175] acquire an environmental image of a driving environment of the vehicle and spatial point cloud data of the driving environment;
[0176] determine state information and navigation information of the vehicle;
[0177] determine the environmental image, the spatial point cloud data, the state information and the navigation information of the vehicle as the multi-modal perception data.
[0178] With regard to the apparatus in the above-described embodiments, in which the specific manner in which each unit performs an operation has been described in detail in the embodiments relating to the method, no detailed elaboration will be made here.
[0179] In this embodiment, by acquiring the multi-modal perception data of the vehicle, and encoding the multi-modal perception data to obtain the encoded feature data, and under any non-first current denoising step in the denoising diffusion process, according to the encoded feature data and the probability information of the last denoising step, the feature sampling is performed to obtain the sampling feature data of the current denoising step, wherein the probability information is used to represent the probability difference under the condition that the trajectory information of different reward scores is generated based on the sampling feature data of the last denoising step, and the target trajectory information is generated according to the sampling feature data corresponding to the plurality of denoising steps in the denoising diffusion process, and the vehicle is controlled to travel according to the target trajectory information. Therefore, since the "probability information is used to represent the probability difference under the condition that the trajectory information of different reward scores is generated based on the sampling feature data of the last denoising step", when the sampling direction of the feature sampling of the current denoising step is guided based on the probability information of the last denoising step, the sampling direction of the trajectory information can be adaptively optimized, so that the target trajectory information with better reward score is generated, and the generation accuracy of the target trajectory information is greatly improved.
[0180] Figure 5 FIG. 1 is a functional block diagram of a vehicle 500, according to an example embodiment. For example, the vehicle 500 can be a hybrid vehicle, a non-hybrid vehicle, an electric vehicle, a fuel cell vehicle, or another type of vehicle. The vehicle 500 can be a self-driving vehicle, a semi-self-driving vehicle, or a non-self-driving vehicle.
[0181] Referring to Figure 5 The vehicle 500 can include various subsystems, such as an infotainment system 510, a perception system 520, a decision control system 530, a drive system 540, and a computing platform 550. The vehicle 500 can include more or fewer subsystems, and each subsystem can include multiple components. In addition, each subsystem of the vehicle 500 and each component can be interconnected by wired or wireless means.
[0182] In some embodiments, the infotainment system 510 can include a communication system, an entertainment system, a navigation system, and the like.
[0183] The perception system 520 can include several sensors for sensing information of the environment around the vehicle 500. For example, the perception system 520 can include a global positioning system (which can be a GPS system, a Beidou system, or another positioning system), an inertial measurement unit, a laser radar, a millimeter wave radar, an ultrasonic radar, and a camera device.
[0184] The decision control system 530 can include a computing system, a vehicle controller, a steering system, a throttle, and a braking system.
[0185] The drive system 540 can include components that provide motive power for the vehicle 500. In one embodiment, the drive system 540 can include an engine, an energy source, a transmission system, and wheels. The engine can be one or a combination of an internal combustion engine, an electric motor, an air compression engine. The engine is capable of converting energy provided by the energy source into mechanical energy.
[0186] Some or all functions of the vehicle 500 are controlled by the computing platform 550. The computing platform 550 can include at least one processor 551 and a memory 552, the processor 551 can execute instructions 553 stored in the memory 552.
[0187] The processor 551 can be any conventional processor, such as commercially available CPUs. The processor can also include a Graphics Processing Unit (GPU), a Field Programmable Gate Array (FPGA), a System on Chip (SOC), an Application Specific Integrated Circuit, or a combination thereof.
[0188] The memory 552 can be implemented by any type of volatile or nonvolatile memory or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read only memory (EEPROM), erasable programmable read only memory (EPROM), programmable read only memory (PROM), read only memory (ROM), magnetic memory, flash memory, magnetic disk, or optical disk.
[0189] In addition to the instructions 553, the memory 552 can also store data, such as road maps, route information, the location, direction, speed, and the like of the vehicle. The data stored in the memory 552 can be used by the computing platform 550.
[0190] In the embodiments of the present disclosure, the processor 551 can execute the instructions 553 to complete all or part of the steps of the vehicle control method described above.
[0191] The present disclosure also provides a computer readable storage medium having computer program instructions stored thereon, the program instructions being executed by a processor to implement the steps of the vehicle control method provided by the present disclosure.
[0192] To implement the above-mentioned embodiments, the present disclosure also provides a chip, comprising: the chip comprises a processing circuit, the processing circuit is configured to execute the method provided in the above-mentioned embodiments.
[0193] Figure 6 is a structural schematic diagram of a chip according to an embodiment of the present disclosure. Referring to FIG. 6, Figure 6 a structural schematic diagram of the chip 600 is shown, but is not limited thereto.
[0194] The chip 600 includes a processing circuit 601 and an interface circuit 602. The interface circuit 602 is configured to read instructions and send the instructions to the processing circuit 601, so that the processing circuit 601 performs the method in the above embodiments.
[0195] Optionally, as shown in FIG. 6, Figure 7 Figure 7 is another structural schematic diagram of a chip according to an embodiment of the present disclosure. The chip 600 can further include a memory 603 configured to store instructions. The interface circuit 602 can be configured to read the instructions stored in the memory 603.
[0196] Optionally, the interface circuit 602 is connected with the memory 603. The interface circuit 602 can be configured to receive signals from the memory 603 or other devices, and the interface circuit 602 can be configured to send signals to the memory 603 or other devices. For example, the interface circuit 602 can read the instructions stored in the memory 603 and send the instructions to the processing circuit 601.
[0197] Optionally, the number of the memory 603 can be one or more. The number of the interface circuit 602 can also be one or more.
[0198] In some embodiments, the interface circuit 602 performs at least one of the communication steps such as sending and / or receiving in the above method, and the processing circuit 601 performs other steps.
[0199] In some embodiments, the terms such as interface circuit, interface, transceiver pin, and transceiver can be replaced with each other.
[0200] Optionally, all or part of the memory 603 can also be outside the chip 600.
[0201] In order to implement the above embodiments, the present disclosure further provides a computer program product. When instructions in the computer program product are executed by a processor, the method according to the above embodiments of the present disclosure is performed.
[0202] Furthermore, the word "exemplary" is used herein to mean serving as an example, instance, or illustration. Any aspect or design described herein as "exemplary" is not necessarily to be construed as advantageous over other aspects or designs. Rather, use of the word exemplary is intended to present concepts in a concrete manner. As used in this application, the term "or" is intended to mean an inclusive "or" rather than an exclusive "or". That is, unless specified otherwise, or clear from context, "X employs A or B" is intended to mean any of the natural inclusive permutations. That is, if X employs A; X employs B; or X employs both A and B, then "X employs A or B" is satisfied under any of the foregoing instances. In addition, the articles "a" and "an" as used in this application and the appended claims should generally be construed to mean "one or more" unless specified otherwise or clear from context to be directed to a singular form. Thus, use of the articles in this application and the following claims is not limiting.
[0203] Also, although the disclosure has been described with respect to one or more implementations, those skilled in the art will readily appreciate that other alternatives can be used. It is contemplated that the disclosure can be carried out in alternate embodiments that do not depart from the spirit and scope of the present disclosure. Accordingly, other implementations are within the scope of the following claims.
[0204] Other embodiments of the disclosure will be apparent to those skilled in the art from consideration of the specification and practice of the features of the disclosure disclosed herein. It is intended that the specification and examples be considered as exemplary only, with the true scope and spirit of the disclosure being indicated by the following claims.
[0205] It is to be understood that the disclosure is not limited to the precise construction described and as shown in the accompanying drawings, which can be varied as desired. The scope of the disclosure is only limited by the claims appended hereto.
[0206] It should be understood that the features of various some embodiments of the present disclosure described herein can be combined with each other, unless specifically noted otherwise. As used herein, the term "and / or" includes any permissible combination of the associated listed terms and any combination of two or more of the associated listed terms; similarly, "at least one of" includes any permissible combination of the associated listed terms and any combination of two or more of the associated listed terms. In addition, the terms "first", "second", etc. are used only for descriptive purposes and are not to be construed as indicating or implying relative importance or an indicated number of technical features. Thus, features defined with "first", "second" can explicitly or implicitly include at least one of the features. In the description herein, the meaning of "a plurality of" is at least two, e.g., two, three, etc., unless specifically noted otherwise.
Claims
1. A vehicle control method characterized by, The method comprises: obtaining multi-modal perception data of a vehicle and encoding the multi-modal perception data to obtain encoded feature data; in any non-first current denoising step in a denoising diffusion process, performing feature sampling according to the encoded feature data and probability information of a previous denoising step to obtain sampled feature data of the current denoising step; wherein the probability information is used to represent probability differences in trajectory information based on the sampled feature data of the previous denoising step to generate different reward scores; generating target trajectory information according to the sampled feature data corresponding to multiple denoising steps in the denoising diffusion process; controlling the vehicle to travel according to the target trajectory information.
2. The method of claim 1, wherein, The feature sampling according to the encoded feature data and the probability information of the previous denoising step comprises: determining feature sampling parameters of the current denoising step according to the probability information of the previous denoising step; performing feature sampling on the encoded feature data according to the feature sampling parameters of the current denoising step.
3. The method of claim 2, wherein, The method further comprises: obtaining positive probability information and negative probability information based on the previous denoising step, wherein the positive probability information is information of a conditional probability of generating first trajectory information based on the sampled feature data of the previous denoising step, and the negative probability information is information of a conditional probability of generating second trajectory information based on the sampled feature data of the previous denoising step, and the reward score of the first trajectory information is higher than that of the second trajectory information; determining the positive probability information and the negative probability information as the probability information of the previous denoising step.
4. The method of claim 3, wherein, The determination of the feature sampling parameters of the current denoising step according to the probability information of the previous denoising step comprises: determining the feature sampling parameters of the previous denoising step; adjusting the feature sampling parameters of the previous denoising step according to the positive probability information and the negative probability information, wherein the feature sampling parameters are used to make the target trajectory information tend to the first trajectory information; determining the adjusted feature sampling parameters as the feature sampling parameters of the current denoising step.
5. The method of claim 4, wherein, The adjustment of the feature sampling parameters of the previous denoising step according to the positive probability information and the negative probability information comprises: determining weight information, wherein the weight information is used to control the intensity of the adjustment of the feature sampling parameters; determining parameter adjustment information according to the weight information, the positive probability information, and the negative probability information; adjusting the feature sampling parameters of the previous denoising step according to the parameter adjustment information.
6. The method of claim 3, wherein, Any denoising step of the denoising diffusion process is executed through a prediction model, wherein the prediction model comprises a decoder, a first sub-model connected to the decoder, and a second sub-model connected to the decoder. The obtaining of the positive probability information and the negative probability information based on the previous denoising step comprises: performing feature sampling on the encoded feature data through the decoder to obtain the sampled feature data of the previous denoising step in the execution of the previous denoising step; processing the sampled feature data of the previous denoising step through the first sub-model to obtain the positive probability information; processing the sampled feature data of the previous denoising step through the second sub-model to obtain the negative probability information. obtaining the negative probability information by processing the sampling feature data of the last denoising step by the second sub-model; wherein the first sub-model has modeled and learned a mapping relationship between the sampling feature data of the last denoising step and the positive probability information, and the second sub-model has modeled and learned a mapping relationship between the sampling feature data of the last denoising step and the negative probability information.
7. The method of claim 6, wherein, The prediction model further comprises an encoder connected to the decoder, wherein the encoding of the multi-modal perception data to obtain the encoded feature data comprises: encoding the multi-modal perception data by the encoder to obtain the encoded feature data.
8. The method according to any one of claims 1 to 7, characterized in that, The obtaining of the multi-modal perception data of the vehicle comprises: collecting environment images of a driving environment of the vehicle and spatial point cloud data of the driving environment; determining state information and navigation information of the vehicle; determining the environment images, the spatial point cloud data, the state information and the navigation information of the vehicle as the multi-modal perception data.
9. A vehicle control device characterized by comprising: comprise: an obtaining unit, configured to obtain multi-modal perception data of a vehicle, and encode the multi-modal perception data to obtain encoded feature data; a processing unit, configured to, at any non-first current denoising step in a denoising diffusion process, perform feature sampling according to the encoded feature data and probability information of a last denoising step to obtain sampling feature data of the current denoising step, wherein the probability information is used to represent probability differences in different reward score trajectory information generated based on the sampling feature data of the last denoising step; a generating unit, configured to generate target trajectory information according to the sampling feature data corresponding to multiple denoising steps in the denoising diffusion process; a control unit, configured to control driving of the vehicle according to the target trajectory information.
10. The apparatus of claim 9, wherein, The processing unit is further configured to: determine feature sampling parameters of the current denoising step according to the probability information of the last denoising step; perform feature sampling on the encoded feature data according to the feature sampling parameters of the current denoising step.
11. The apparatus of claim 10, wherein, The processing unit is further configured to: obtain positive probability information and negative probability information based on the last denoising step, wherein the positive probability information is information of a conditional probability of generating first trajectory information based on the sampling feature data of the last denoising step, and the negative probability information is information of a conditional probability of generating second trajectory information based on the sampling feature data of the last denoising step, and a reward score of the first trajectory information is higher than a reward score of the second trajectory information; determine the positive probability information and the negative probability information as the probability information of the last denoising step.
12. The apparatus of claim 11, wherein, The processing unit is further configured to: determine feature sampling parameters of the last denoising step; adjust the feature sampling parameters of the last denoising step according to the positive probability information and the negative probability information, wherein the feature sampling parameters are used to make the target trajectory information tend to the first trajectory information; determine the adjusted feature sampling parameters as the feature sampling parameters of the current denoising step.
13. A vehicle characterized by comprising: comprise: a processor; a memory for storing processor-executable instructions; wherein the processor is configured to: implement the steps of the method of any one of claims 1-8. 14.A non-transitory computer-readable storage medium, when instructions in the storage medium are executed by a processor of a mobile terminal, enabling the mobile terminal to perform a method, the method comprising: obtaining multi-modal perception data of a vehicle, and encoding the multi-modal perception data to obtain encoded feature data; at any non-first current denoising step in a denoising diffusion process, performing feature sampling according to the encoded feature data and probability information of a previous denoising step to obtain sampled feature data of the current denoising step, wherein the probability information is used to represent probability differences in trajectory information based on the sampled feature data of the previous denoising step to generate different reward scores; generating target trajectory information according to the sampled feature data corresponding to multiple denoising steps in the denoising diffusion process; controlling the vehicle to travel according to the target trajectory information.
15. A computer program product, characterised in that, a computer program that, when executed by a processor, implements the method of any one of claims 1-8.
16. A chip comprising processing circuitry, interface circuitry; wherein, an interface circuit for reading instructions, the interface circuit sending the instructions to the processing circuit to cause the processing circuit to perform the method of any one of claims 1-8.
Citation Information
Patent Citations
Vehicle track recovery method and device based on unified imaging data modeling
CN119887844A
Visual motion prediction acceleration method and system directly based on pre-training model
CN120029162A
Driving track planning method, and driving track planning model training method and device
CN120609378A
Personalized customized video object replacement method capable of maintaining layout and motion modes
CN120676201A
Automatic driving-oriented kinematics priori guided vehicle trajectory generation method
CN120686825A