Method for acquiring planning and control information, and related device

By using a method of serially generating control information and incorporating machine learning models to consider the impact of the previous moment, the problem of discontinuous vehicle control information is solved, thereby improving the continuity and safety of vehicle control.

WO2026051566A1PCT designated stage Publication Date: 2026-03-12YINWANG INTELLIGENT TECHNOLOGIES CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-07-01
Publication Date
2026-03-12

AI Technical Summary

Technical Problem

The vehicle control information generated by machine learning models has poor consistency across different times, resulting in poor consistency in the vehicle control process.

Method used

The control information is generated serially. The control information for the current time and subsequent time is generated sequentially through a machine learning model. The influence of the control information of the previous time on the generation of the current time is considered. The control information for multiple consecutive time moments is generated serially using a machine learning model.

Benefits of technology

It improves the continuity of the vehicle control process and the coherence between control information, thereby enhancing the safety, reliability, and user experience of the vehicle during intelligent driving.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025106452_12032026_PF_FP_ABST
    Figure CN2025106452_12032026_PF_FP_ABST
Patent Text Reader

Abstract

Embodiments of the present application provide a method for acquiring planning and control information, and a related device, applied to the field of intelligent driving. The method comprises: acquiring environmental information at a current moment; and on the basis of the environmental information at the current moment, obtaining planning and control information of a vehicle corresponding to each moment among the current moment and at least one subsequent moment, wherein the planning and control information corresponding to each moment is obtained by means of a machine learning model, the planning and control information corresponding to the current moment is obtained by the machine learning model on the basis of the environmental information at the current moment, and the planning and control information at a first moment (any one of the at least one moment subsequent to the current moment) is obtained by the machine learning model on the basis of the planning and control information corresponding to a moment preceding the first moment. The planning and control information corresponding to the current moment and the subsequent moments is sequentially generated in a serial manner, and the influence of the planning and control information at the previous moment on generating the planning and control information at the first moment is considered, thereby improving the coherence among the planning and control information at a plurality of moments.
Need to check novelty before this filing date? Find Prior Art

Description

Method for acquiring regulatory information and related device

[0001] The present application claims priority to the Chinese patent application No. 202411254293.6, filed on September 6, 2024, and entitled "Method for acquiring regulatory information and related device", the content of which is incorporated herein by reference in its entirety. TECHNICAL FIELD

[0002] The present application relates to intelligent driving technology, and in particular, to a method for acquiring regulatory information and related device. BACKGROUND

[0003] Artificial intelligence (AI) is the use of digital computers or digital computer-controlled machines to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results. In other words, artificial intelligence is a branch of computer science that aims to understand the essence of intelligence and produce a new intelligent machine that can react in a similar way to human intelligence. Artificial intelligence is the design principle and implementation method of various intelligent machines, enabling machines to have perception, reasoning and decision-making functions.

[0004] Among them, the field of intelligent driving is a common application field of artificial intelligence technology. For example, the current environmental information can be obtained through a sensor, and the current environmental information is input into a machine learning model to obtain the regulatory information of the vehicle corresponding to the current time generated by the machine learning model.

[0005] However, since the machine learning model only generates the regulatory information of the vehicle corresponding to the current time each time, the continuity between the regulatory information of the vehicle corresponding to different times may be poor, thereby leading to poor continuity of the vehicle control process. SUMMARY

[0006] The present application provides a method for acquiring regulatory information and related device, which generates the regulatory information corresponding to the current time and the time thereafter in a serial manner, and considers the influence of the regulatory information of the previous time on the regulatory information of the first time, which is beneficial to improve the continuity between the regulatory information of the vehicle corresponding to multiple times, and thereby beneficial to improve the continuity of the vehicle control process.

[0007] The present application provides the following technical solutions:

[0008] In a first aspect, the present application provides a method for obtaining regulation information, which can be applied in the field of intelligent driving. In the method, a first device obtains environment information at a current time, and then obtains regulation information of a vehicle corresponding to each of at least two times based on the environment information at the current time. The at least two times include the current time and at least one time after the current time. The regulation information corresponding to each of the at least two times is obtained by the same machine learning model. The regulation information corresponding to the current time is obtained by the machine learning model based on the environment information at the current time. The first time is any one of the at least one time after the current time. The regulation information of the first time is obtained by the machine learning model based on the regulation information corresponding to the last time of the first time. That is, the machine learning model generates the regulation information corresponding to each of the current time and the at least one time after the current time in a serial manner.

[0009] For example, the at least two times can be at least two continuous times. The time interval between two adjacent times of the at least two times can be a preset time step. For example, the time step can be 0.1 seconds, 0.2 seconds, 0.3 seconds, or other time lengths.

[0010] In one case, the first device obtains the environment information at the current time by collecting the environment information at the current time through a sensor deployed in the vehicle. The first device obtains the regulation information of the vehicle corresponding to each of the at least two times based on the environment information at the current time by generating the regulation information of the vehicle corresponding to each of the at least two times through the machine learning model based on the environment information at the current time.

[0011] In another case, the first device obtains the environment information at the current time by collecting the environment information at the current time through a sensor deployed in the vehicle. The first device obtains the regulation information of the vehicle corresponding to each of the at least two times based on the environment information at the current time by sending the environment information at the current time to an execution device, and receiving the regulation information of the vehicle corresponding to each of the at least two times sent by the execution device.

[0012] In another case, the first device obtains the environment information at the current time by receiving the environment information at the current time sent by an intelligent driving system in the vehicle. The first device obtains the regulation information of the vehicle corresponding to each of the at least two times based on the environment information at the current time by generating the regulation information of the vehicle corresponding to each of the at least two times through the machine learning model based on the environment information at the current time.

[0013] In the present implementation, based on the obtained environment information at the current time, the machine learning model obtains the regulation and control information of the vehicle corresponding to the current time and at least one time after the current time, that is, after obtaining the environment information at the current time, the machine learning model generates the regulation and control information corresponding to a plurality of continuous times, so that the machine learning model not only considers the situation at the current time when generating the regulation and control information, but also considers the influence on the future time, which is beneficial to improve the continuity of the vehicle control process; and the regulation and control information corresponding to the current time is obtained by the machine learning model based on the environment information at the current time, the first time is any one of at least one time after the current time, and the regulation and control information at the first time is obtained by the machine learning model based on the regulation and control information corresponding to the last time of the first time, that is, the machine learning model generates the regulation and control information corresponding to the current time and the time after the current time in a serial manner, and considers the influence of the regulation and control information at the last time on the generation of the regulation and control information at the first time, which is beneficial to further improve the continuity between the regulation and control information of the vehicle corresponding to a plurality of times, and further beneficial to improve the continuity of the vehicle control process.

[0014] In a possible implementation, the regulation and control information at the first time is obtained by the machine learning model based on the predicted feature information of the environment at the first time; wherein the predicted feature information of the environment at the first time is obtained by the machine learning model based on the regulation and control information corresponding to the last time of the first time and the feature information of the environment at the last time of the first time.

[0015] For example, if the last time of the first time is the current time, the feature information of the environment at the last time of the first time can be the feature information obtained based on the environment information at the current time (i.e. the real environment information); if the last time of the first time is a time after the current time, the feature information of the environment at the last time of the first time can be the predicted feature information of the environment at the last time of the first time.

[0016] In the implementation, when the machine learning model generates the regulation and control information corresponding to the first time point (i.e., any time point after the current time point), the predicted feature information of the environment at the first time point is generated based on the feature information of the environment at the last time point before the first time point and the regulation and control information corresponding to the last time point before the first time point, i.e., the machine learning model can imagine the environment at the future time point after the current time point in the hidden layer space. Since the machine learning model in the application has the ability to imagine the non-existent environment, the traffic environment seen by the machine learning model is greatly expanded, which is conducive to reducing the probability of encountering a completely unknown traffic environment by the machine learning model and improving the response capability of the machine learning model to the unknown traffic environment, i.e., the above covariate problem can be solved, in other words, the performance of the regulation and control information of the vehicle generated by the machine learning model is improved, so as to improve the safety and reliability of the intelligent driving process.

[0017] In a possible implementation, the predicted feature information of the environment at the first time point (i.e., any time point after the current time point) is generated by a first module in the machine learning model, and the machine learning model further includes a second module. The input of the second module includes the environment information at the current time point, and the feature information corresponding to the environment information at the current time point is obtained through the second module. Optionally, the input of the second module further includes the state information of the vehicle at the current time point. The first module is configured to obtain the regulation and control information corresponding to each of the at least two time points based on the feature information generated by the second module. Optionally, the first module and the second module can both be neural network modules, i.e., the first module can include a plurality of neural network layers, and the second module can also include a plurality of neural network layers.

[0018] For example, the feature information generated by the second module in the machine learning model can be understood as information obtained by processing the environment information at the current time point (optionally, the state information of the vehicle at the current time point) based on the neural network layers in the second module.

[0019] In the implementation, the architecture of the machine learning model is refined, and the implementation difficulty of the scheme is reduced. The entire machine learning model is divided into the first module and the second module, the input of the second module includes the environment information at the current time point, and the feature information corresponding to the environment information at the current time point is obtained through the second module. The first module is configured to generate the regulation and control information at the current time point, the predicted feature information of the environment at each of the at least one time point after the current time point, and the regulation and control information. The entire machine learning model is managed more finely, i.e., the process of generating the regulation and control information of the vehicle corresponding to each of the at least two time points by using the machine learning model is managed more finely, which is conducive to improving the quality of the finally obtained regulation and control information, and further improving the user experience of the vehicle in the intelligent driving process.

[0020] In a possible implementation, the first module includes a first submodule corresponding to the current time and at least one second submodule corresponding to at least one time after the current time, and the at least one second submodule can be in a serial structure. The first submodule corresponding to the current time is configured to generate the regulation and control information of the vehicle corresponding to the current time based on the feature information generated by the second module in the machine learning model. The second submodule corresponding to the first time is configured to generate the regulation and control information of the first time based on the regulation and control information of the vehicle corresponding to the last time of the first time and the predicted feature information of the environment of the first time. For example, the second submodule corresponding to the first time is configured to generate the predicted feature information of the environment of the first time based on the regulation and control information of the vehicle corresponding to the last time of the first time, and then generate the regulation and control information of the vehicle corresponding to the first time based on the predicted feature information of the environment of the first time.

[0021] In the implementation, the first module includes a first submodule corresponding to the current time and at least one second submodule corresponding to at least one time after the current time, and each second submodule is specially configured to generate the regulation and control information of one time after the current time, thereby further improving the refinement degree of the machine learning model, that is, the process of generating the regulation and control information of the vehicle corresponding to each of the at least two times by using the machine learning model is managed more finely, so as to further improve the quality of the obtained regulation and control information.

[0022] In a possible implementation, the training phase of the first module is reinforcement learning. Optionally, the training phase of the first module is obtained after reinforcement learning in a simulation environment by using a simulator. In the implementation, the predicted feature information of the environment of the first time (that is, any time after the current time) is generated by the first module, the first module is further configured to generate the regulation and control information corresponding to the current time and at least one time after the current time, and the training phase of the first module adopts the reinforcement learning manner. Therefore, in the training phase of the first module, the first module can sufficiently explore the feature information of various traffic environments, obtain the corresponding regulation and control information based on the various traffic environments, and then evaluate the foregoing regulation and control information by using the reinforcement learning manner, that is, the reinforcement learning manner can feed back the regulation and control information generated by the machine learning model for various traffic environments, so that the first module can learn the response capability for various traffic environments in the training phase, greatly expand the traffic scenarios that can be responded to by the machine learning model, and be beneficial to improving the quality of the regulation and control information generated by the machine learning model in various application environments, thereby being beneficial to improving the safety and reliability of the intelligent driving process of the vehicle and improving the user experience of the intelligent driving process.

[0023] In a possible implementation, the feature information generated by the second module in the machine learning model includes a semantic image in a top view perspective. The top view perspective in the present application can also be understood as a bird's eye view (BEV). The semantic image in the top view perspective can also be referred to as a BEV perspective. The semantic image indicates a semantic category of at least one region in an image in the BEV perspective corresponding to the environment information at the current moment. For example, the semantic category can be a lane, an intersection, a ramp, a traffic sign, a static obstacle, a social vehicle, a bus, a pedestrian, or another semantic category.

[0024] In the present implementation, the second module in the machine learning model generates a semantic image in the BEV perspective based on the environment information at the current moment. The semantic image in the BEV perspective can not only reflect the environment in the BEV perspective, but also carry the semantic category of at least one region in the image in the BEV perspective. That is, the second module in the machine learning model obtains effective information in the environment at the current moment based on the environment information at the current moment. Then, the first module in the machine learning model generates the regulation and control information of the vehicle corresponding to each of the at least two moments based on the semantic image in the BEV perspective, which is beneficial to obtain more accurate regulation and control information. In addition, if the training stage of the first module in the machine learning model is reinforcement learning, the reinforcement learning is generally trained in a simulation environment with the aid of a simulator. The environment information obtained from the simulation environment is quite different from the environment information obtained from the real traffic environment. If the environment information is directly input into the first module, the environment information input into the first module in the training stage and the application stage may be quite different, which may result in poor quality of the regulation and control information generated by the first module in the application stage. Therefore, the second module is independently set to generate the semantic image in the BEV perspective based on the environment information at the current moment. The input of the first module includes the semantic image in the BEV perspective, which can make the data input into the first module in the training stage and the application stage as similar as possible, so as to improve the quality of the regulation and control information generated by the first module, and further improve the user experience of the intelligent driving process of the vehicle.

[0025] In a possible implementation, the feature information generated by the second module in the machine learning model further includes first feature information obtained by feature extraction based on the environment information at the current moment. For example, after the first device inputs the environment information at the current moment (optionally, further including the state information of the vehicle at the current moment) into the machine learning model, the second module in the machine learning model performs feature extraction on the input information to obtain the first feature information. The second module in the machine learning model performs feature processing on the first feature information to obtain the semantic image in the BEV perspective.

[0026] In the implementation, the feature information generated by the second module includes not only the image in the BEV perspective, but also the first feature information obtained based on the environment information at the current time. Therefore, in the process of generating the regulation and control information, the first module of the machine learning model can not only use the effective information (i.e., the image in the BEV perspective) extracted by the second module based on the environment information, but also obtain more comprehensive first feature information corresponding to the environment information, that is, not only a clearer understanding of the environment at the current time, but also a more comprehensive understanding of the environment at the current time, which is conducive to further improving the quality of the obtained regulation and control information.

[0027] In a possible implementation, the feature information generated by the second module in the machine learning model further includes second feature information corresponding to the time before the first time, and the second feature information is obtained by the second module in the machine learning model based on the first feature information and the regulation and control information of the vehicle corresponding to the time before the first time. Then, in the process of generating the regulation and control information of the vehicle corresponding to each of the at least two times, the first module in the machine learning model also uses the second feature information.

[0028] In the implementation, the second feature information corresponding to the time before the first time can also be obtained by the second module of the machine learning model based on the first feature information and the regulation and control information corresponding to the time before the first time, that is, in the process of generating the regulation and control information by the first module in the machine learning model, the second feature information obtained based on the first feature information and the regulation and control information corresponding to the time before the first time can also be used, and more comprehensive feature information is used to generate the regulation and control information, which is conducive to obtaining regulation and control information with higher quality.

[0029] In a possible implementation, the predicted feature information of the environment at the first time is obtained by the first module based on the feature information of the environment at the time before the first time, the regulation and control information corresponding to the time before the first time, and the second feature information. For example, in the process of generating the regulation and control information at the first time by the first module in the machine learning model, the predicted feature information of the environment at the first time can be obtained by the first module in the machine learning model based on the feature information of the environment at the time before the first time, the regulation and control information corresponding to the time before the first time, and the second feature information; and then, the regulation and control information at the first time is generated by the first module in the machine learning model based on the predicted feature information of the environment at the first time.

[0030] In the implementation, the second feature information corresponding to the previous moment of the first moment is used to generate the predicted feature information of the environment of the first moment by the first module, and the predicted feature information of the environment of the first moment is determined in combination with the feature information of the environment of the previous moment of the first moment, the regulation and control information corresponding to the previous moment of the first moment, and the second feature information, which is conducive to reducing the difficulty of predicting the feature information of the environment of the first moment, thereby assisting the first module to better predict the feature information of the environment of the first moment, and on the basis of obtaining the predicted feature information of the environment of the first moment with better quality, the regulation and control information with better quality is generated.

[0031] In a possible implementation, the first module in the machine learning model is configured to generate predicted feature information of an environment of each of at least one moment after the current moment. In the training phase of the first module, the predicted feature information of the environment of each of at least one moment after the current moment is input into a decoder to obtain predicted environment information corresponding to each of the at least one moment generated by the decoder after feature processing of the predicted feature information of the environment of each of the at least one moment after the current moment. The correct environment information corresponding to the predicted environment information is obtained by the simulator, and the loss function of the training phase of the first module indicates the similarity between the predicted environment information and the correct environment information. The target of training the first module by using the first loss function includes improving the similarity between the predicted environment information and the correct environment information.

[0032] In the implementation, in the training phase of the first module, since the training phase of the first module is performed in the simulator, the correct environment information of at least one moment after the current moment can be obtained by the simulator, the predicted environment information is obtained by performing reverse feature processing on the predicted feature information of the environment generated by the first module, and the similarity between the predicted environment information and the correct environment information is narrowed by using the loss function of the first module. In this way, the first module is trained to generate more accurate predicted feature information of the environment, that is, the accuracy of the first module in predicting the feature information of the environment is improved.

[0033] In a possible implementation, the first module is configured to generate predicted feature information of an environment of each of at least one moment after the current moment. The loss function of the training phase of the first module indicates the similarity between the predicted feature information generated by the first module and expected feature information, and the expected feature information is obtained based on correct environment information, for example, the expected feature information is obtained by performing feature extraction on the correct environment information, and the correct environment information is obtained by the simulator.

[0034] In the training stage of the first module, the correct environment information at each of the at least one time moment after the current time moment can also be subjected to feature extraction to obtain expected feature information of the environment at each of the at least one time moment, and the similarity between the predicted feature information and the expected feature information of the environment at each of the at least one time moment can be narrowed down by using the loss function of the first module, so that the first module can be trained to generate more accurate predicted feature information of the environment, and the accuracy of the first module in predicting the feature information of the environment can also be improved.

[0035] In a possible implementation, the regulation and control information of the vehicle corresponding to each of the at least two time moments can include at least one of the following: control information of the vehicle corresponding to each of the at least two time moments, a planned position of the vehicle corresponding to each of the at least two time moments, a driving strategy of the vehicle corresponding to each of the at least two time moments, or other types of regulation and control information, and the like. And / or, the output of the machine learning model further includes perception information corresponding to the environment information of the current time moment, wherein the perception information includes road information and / or obstacle information in the environment of the current time moment.

[0036] Optionally, the control information of the vehicle corresponding to each of the time moments can include a control signal of at least one component in the vehicle corresponding to each of the time moments, and / or state information of the planned vehicle corresponding to each of the time moments. For example, the at least one component can include a steering wheel, an engine, a brake, a clutch, a turn signal, or a component used during vehicle driving, and the like; and the state information of the planned vehicle corresponding to each of the at least one time moment after the current time moment can be understood as state information that the vehicle needs to reach at each of the at least one time moment.

[0037] The planned position of the vehicle corresponding to each of the at least two time moments can also be understood as a planned trajectory point of the vehicle corresponding to each of the at least two time moments, and the at least two planned trajectory points of the vehicle corresponding to the at least two time moments can also be understood as a planned trajectory of the vehicle at the at least two time moments.

[0038] In the present implementation, the machine learning model can not only obtain the regulation and control information of the current time moment and the at least one time moment after the current time moment, but also obtain the road information and / or obstacle information in the environment information of the current time moment, so that the user can obtain more information, and the user experience of the present solution can be further improved.

[0039] In a second aspect, the present application provides a device for obtaining regulation information, which can be used in the field of artificial intelligence. The device comprises: an obtaining module, configured to obtain environment information at a current time; and a processing module, configured to obtain regulation information of a vehicle corresponding to each of at least two times based on the environment information at the current time, the at least two times comprising the current time and at least one time after the current time; wherein the regulation information corresponding to each of the at least two times is obtained by a machine learning model, the regulation information corresponding to the current time is obtained by the machine learning model based on the environment information at the current time, the first time is any one of the at least one time after the current time, and the regulation information of the first time is obtained by the machine learning model based on the regulation information corresponding to the last time before the first time.

[0040] In the second aspect of the present application, the device for obtaining regulation information is also configured to perform the steps performed by the first device in the first aspect and various possible implementation manners of the first aspect. The specific implementation manners of the steps, the meanings of the terms and the beneficial effects brought by the second aspect can be referred to the first aspect, and will not be described here.

[0041] In a third aspect, the present application provides a device comprising a processor and a memory, the processor being coupled to the memory, the memory being configured to store a program, and the processor being configured to execute the program in the memory so that the device performs the method of the first aspect.

[0042] In a fourth aspect, the present application provides a computer-readable storage medium, which stores a computer program, and when the computer program is executed on a computer, the computer is caused to perform the method of the first aspect.

[0043] In a fifth aspect, the present application provides a computer program product, which comprises a program, and when the program is executed on a computer, the computer is caused to perform the method of the first aspect.

[0044] In a sixth aspect, the present application provides a chip system, which comprises a processor configured to support the functions involved in the above aspects, such as sending or processing the data and / or information involved in the above method. In a possible design, the chip system further comprises a memory configured to save necessary program instructions and data of the terminal device or the communication device. The chip system can be composed of a chip, or can comprise a chip and other discrete devices.

[0045] The third aspect to the sixth aspect of the present application correspond to the first aspect or the various possible manners of the first aspect, and have corresponding beneficial effects. BRIEF DESCRIPTION OF DRAWINGS

[0046] FIG. 1 is a structural schematic diagram of an artificial intelligence subject framework provided by the present application.

[0047] Fig. 2a is a system architecture diagram of a system for acquiring regulation information according to an embodiment of the present application;

[0048] Fig. 2b is another system architecture diagram of a system for acquiring regulation information according to an embodiment of the present application;

[0049] Fig. 3 is a flow diagram of a method for acquiring regulation information according to an embodiment of the present application;

[0050] Fig. 4 is a diagram of a machine learning model according to an embodiment of the present application;

[0051] Fig. 5 is another diagram of a machine learning model according to an embodiment of the present application;

[0052] Fig. 6 is a diagram of a first module of a machine learning model according to an embodiment of the present application;

[0053] Fig. 7 is another diagram of a machine learning model according to an embodiment of the present application;

[0054] Fig. 8 is a flow diagram of a method for training a model according to an embodiment of the present application;

[0055] Fig. 9 is another flow diagram of a method for training a model according to an embodiment of the present application;

[0056] Fig. 10 is a structural diagram of an apparatus for acquiring regulation information according to an embodiment of the present application;

[0057] Fig. 11 is a structural diagram of an apparatus according to an embodiment of the present application;

[0058] Fig. 12 is a structural diagram of a vehicle according to an embodiment of the present application. DETAILED DESCRIPTION

[0059] The embodiments of the present application will be described in detail with reference to the drawings. Obviously, the described embodiments are only a part of the embodiments of the present application, but not all the embodiments. Those skilled in the art can know that with the emergence of new application scenarios, the technical solutions provided by the embodiments of the present application are also applicable to similar technical problems.

[0060] The terms "first", "second", etc. in the specification and claims of the present application and the above drawings are used to distinguish similar objects, and do not necessarily have to be used to describe a specific order or sequence. It should be understood that the terms used in this way can be interchanged under appropriate circumstances, and this is merely a way of distinguishing the objects of the same attribute used in the description of the embodiments of the present application. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, so that a process, method, system, product or apparatus including a series of units does not have to be limited to those units, but can include other units not clearly listed or inherent to the process, method, product or apparatus.

[0061] In the embodiments of the present application, "indication" can include direct indication and indirect indication, and can also include explicit indication and implicit indication. The information indicated by certain information (indication information as described below) is referred to as to-be-indicated information, and there are many ways to indicate the to-be-indicated information in the specific implementation process, for example, but not limited to, the to-be-indicated information can be directly indicated, such as the to-be-indicated information itself or an index of the to-be-indicated information. The to-be-indicated information can also be indirectly indicated by indicating other information, where the other information and the to-be-indicated information have an association relationship; the to-be-indicated information can also be indicated only by a part, and the other part of the to-be-indicated information is known or agreed in advance, for example, the indication of a specific information can be achieved by means of the arrangement order of each information agreed in advance (for example, protocol predefined), thereby reducing the indication overhead to a certain extent. The specific manner of indication is not limited in the present application. It can be understood that the indication information can be used to indicate the to-be-indicated information for the sender of the indication information, and the indication information can be used to determine the to-be-indicated information for the receiver of the indication information.

[0062] First, the overall workflow of the artificial intelligence system is described, please refer to FIG. 1, which is a structural schematic diagram of an artificial intelligence main framework provided by the present application. The above artificial intelligence theme framework is described below from two dimensions of "intelligent information chain" (horizontal axis) and "IT value chain" (vertical axis). The "intelligent information chain" reflects a series of processes from data acquisition to processing. For example, it can be a general process of intelligent information perception, intelligent information representation and formation, intelligent reasoning, intelligent decision-making, intelligent execution and output. In this process, data undergoes a condensation process of "data-information-knowledge-wisdom". The "IT value chain" reflects the value brought by artificial intelligence to the information technology industry from the underlying infrastructure of human intelligence, information (provision and processing technology implementation) to the industrial ecological process of the system.

[0063] (1) Infrastructure

[0064] The infrastructure provides computing power support for the artificial intelligence system, realizes communication with the external world, and realizes support through the underlying platform. Communication with the outside world is realized through sensors; computing power is provided by an intelligent chip, which can specifically adopt a central processing unit (CPU), a neural-network processing unit (NPU), a graphics processing unit (GPU), an application specific integrated circuit (ASIC), or a field programmable gate array (FPGA) hardware acceleration chip; the underlying platform includes a distributed computing framework and related platform guarantees and supports such as a network, which can include cloud storage and computing, an interconnection network, etc. For example, the sensor and external communication obtain data, which are provided to the intelligent chip in the distributed computing system provided by the underlying platform for calculation.

[0065] (2) Data

[0066] The data of the upper layer of the infrastructure is used to represent the data source in the field of artificial intelligence. The data relates to graphics, images, voice, text, and also relates to the Internet of Things data of traditional devices, including the business data of existing systems and sensing data such as force, displacement, liquid level, temperature, and humidity.

[0067] (3) Data processing

[0068] Data processing usually includes data training, machine learning, deep learning, search, reasoning, decision-making, etc.

[0069] Among them, machine learning and deep learning can model, extract, preprocess, train, etc. symbolic and formalized intelligent information of data.

[0070] Reasoning refers to the process of simulating human intelligent reasoning methods in a computer or intelligent system, using formalized information to perform machine thinking and solve problems according to reasoning control strategies, and the typical function is search and matching.

[0071] Decision-making refers to the process of decision-making after intelligent information is reasoned, which usually provides functions such as classification, sorting, and prediction.

[0072] (4) General capabilities

[0073] After the data is processed as mentioned above, some general capabilities can be formed based on the results of the data processing, such as algorithms or a general system, for example, translation, text analysis, computer vision processing, speech recognition, image recognition, etc.

[0074] (5) Intelligent product and industry application

[0075] Intelligent product and industry application refers to the product and application of artificial intelligence system in various fields, which is the packaging of the overall solution of artificial intelligence, and realizes the application of intelligent information decision product. The application fields mainly include intelligent terminal, intelligent manufacturing, intelligent transportation, smart home, intelligent medical treatment, intelligent security, intelligent driving, smart city, etc.

[0076] The method provided in the present application can be applied in the field of intelligent driving. For example, the method provided in the present application can be used in the application scenario of path planning of an intelligent vehicle, and / or in the application scenario of generating control information of the intelligent vehicle, etc. The vehicle can be a car, a truck, a motorcycle, a bus, a ship, a mower, an entertainment vehicle, an amusement park vehicle, a construction equipment, an electric car, a golf cart, a train, an airplane, a helicopter, etc., which is not particularly limited in the embodiments of the present application.

[0077] In the related art, the current environmental information can be obtained through a sensor, and the current environmental information is input into a machine learning model to obtain the regulation and control information of the vehicle corresponding to the current time generated by the machine learning model. However, since the machine learning model only generates the regulation and control information of the vehicle corresponding to the current time each time, the continuity between the regulation and control information of the vehicle corresponding to different times can be poor, which leads to poor continuity of the vehicle control process.

[0078] To solve the above problems, the application discloses: obtaining environment information at the current time, based on the obtained environment information at the current time, the machine learning model will obtain the vehicle control information corresponding to the current time and at least one time after the current time, that is, after obtaining the environment information at the current time, the machine learning model will generate control information corresponding to multiple continuous times, so that the machine learning model will not only consider the current situation when generating control information, but also consider the influence on future time, which is conducive to improving the continuity of the vehicle control process; and the control information corresponding to the current time is obtained by the machine learning model based on the environment information at the current time, the first time is any one of the at least one time after the current time, and the control information of the first time is obtained by the machine learning model based on the control information corresponding to the last time of the first time, that is, the machine learning model generates the control information corresponding to the current time and the time after the current time in a serial manner, and considers the influence of the control information of the last time on the generation of the control information of the first time, which is conducive to further improving the continuity between the control information of the vehicle corresponding to multiple times, and further improving the continuity of the vehicle control process.

[0079] Optionally, the method provided by the application can also solve the problem of covariate shift. At present, the training phase of the machine learning model used to generate the control information of the vehicle often adopts the idea of behavior cloning. The idea of behavior cloning is to pre-acquire a training data set, which includes multiple training samples and the expected control information of the vehicle corresponding to each training sample. Each training sample can include environment information, and the expected control information corresponding to each training sample can also be referred to as the true value corresponding to each training sample. The expected control information corresponding to each training sample can be understood as the correct control information of the vehicle demonstrated by an expert. The machine learning model is trained in a supervised learning manner to learn the correct control information of the vehicle demonstrated by the expert, so as to imitate the behavior of the expert. However, due to the fact that the environment information in some traffic environments cannot be collected, when the machine learning model executed after the training operation processes the environment information of the traffic environment not in the training data set, the quality of the control information of the vehicle generated by the machine learning model is poor; and although the machine learning model tries to imitate the behavior of the expert, the machine learning model often cannot imitate one hundred percent, and the accumulation of such errors may lead to the need for the machine learning model to process the traffic environment not in the training data set, which is called the covariate shift problem.

[0080] Before the method provided in the present application is described in detail, the architecture of the rule control information acquisition system provided in the present application is described. Please refer to FIG. 2a, which is a system architecture diagram of the rule control information acquisition system provided in an embodiment of the present application. In FIG. 2a, the rule control information acquisition system 200 includes a training device 210, a database 220, an execution device 230, and a data storage system 240. The execution device 230 includes a computing module 231.

[0081] The database 220 stores a training data set. In the training phase of the machine learning model 201, the training device 210 uses the training data set to iteratively train the machine learning model 201, obtaining a trained machine learning model 201. The machine learning model 201 can be a neural network or a non-neural network model. Optionally, the machine learning model in the present application can use a deep learning model, such as a neural network based on an attention mechanism, a residual neural network, a fully connected neural network, a convolutional neural network, or other types of neural networks.

[0082] The trained machine learning model 201 described above can be deployed in the computing module 231 of the execution device 230. Optionally, as shown in FIG. 2a, the execution device 230 can be integrated into a vehicle, and the user can directly interact with the vehicle in which the execution device 230 is deployed. For example, the execution device 230 can be a module in the host CPU of the vehicle that uses the machine learning model to process data. The execution device 230 can also be a graphics processing unit (GPU), a neural network processing unit (NPU), or a tensor processing unit (TPU) in the vehicle, etc. The aforementioned GPU, NPU, or TPU is mounted as a co-processor to the host CPU, and the host CPU allocates tasks, etc.

[0083] In the application phase of the machine learning model 201, for example, after the intelligent driving system in the vehicle acquires the environmental information at the current time, it can use the trained machine learning model 201 deployed in the computing module 231 to obtain the rule control information corresponding to the current time and at least one time after the current time based on the environmental information at the current time.

[0084] The execution device 230 can call data, code, etc. in the data storage system 240, or store data, instructions, etc. in the data storage system 240. The data storage system 240 can be placed in the execution device 230, or it can be an external storage relative to the execution device 230.

[0085] Please continue to refer to FIG. 2b, which is another system architecture diagram of the rule control information acquisition system provided by the embodiments of the present application. The system architecture shown in FIG. 2b can be understood in combination with the description of FIG. 2a. The difference is that the execution device 230 and the vehicle 250 in the rule control information acquisition system 200 shown in FIG. 2b are independent devices respectively. The execution device 230 is configured with an input / output (I / O) interface. The execution device 230 can interact with the vehicle 250 through the I / O interface. For example, in the application stage of the machine learning model 201, the intelligent driving system in the vehicle 250 can send the environment information at the current time to the execution device 230 through the I / O interface. After the execution device 230 obtains the rule control information corresponding to the current time and each time after the current time through the machine learning model 201, the execution device 230 can send the rule control information to the intelligent driving system in the vehicle 250 through the I / O interface.

[0086] It should be noted that FIG. 2a and FIG. 2b are only the architecture diagrams of the rule control information acquisition system provided by the embodiments of the present application. The positional relationship between the devices, components, modules and the like shown in the diagrams does not constitute any limitation. For example, in some other embodiments of the present application, the training device 210 and the execution device 230 can also be integrated into the same device. The specific architecture of the rule control information acquisition system can be determined according to the actual application scenario. The application stage of the machine learning model and the specific implementation process of the application stage will be described respectively as follows.

[0087] I. Application stage

[0088] Please refer to FIG. 3, which is a flowchart of the rule control information acquisition method provided by the embodiments of the present application. The rule control information acquisition method provided by the embodiments of the present application can include:

[0089] 301. Obtain the environment information at the current time.

[0090] Exemplarily, the environment information of the current moment can include at least one of the following: one or more images of the environment within a first range around the vehicle at the current moment, one or more sets of point cloud data of the environment within a second range around the vehicle at the current moment, or other types of environment information, such as map information corresponding to the vehicle position at the current moment, and the like. The specific information used can be determined in combination with the actual application scenario. Wherein, all images of the environment within the first range around the vehicle can be collected by a first sensor deployed in the vehicle, for example, the first sensor can be an optical sensor, for example, the optical sensor can be a camera or an event camera, and the first range can be the field of view range of the first sensor in the vehicle. All point cloud data of the environment within the second range around the vehicle can be collected by a second sensor deployed in the vehicle, for example, the second sensor can be an ultrasonic sensor, a laser radar sensor, a millimeter wave radar sensor or other sensors capable of measuring point cloud data, and the second range can be the field of view range of the second sensor in the vehicle.

[0091] Optionally, in step 301, the first device not only obtains the environment information of the current moment, but also obtains the state information of the vehicle at the current moment. Exemplarily, the state information of the vehicle can include at least one of the following information of the vehicle: speed, acceleration, orientation, angular acceleration or other information, etc. For example, the state information of the vehicle can be obtained by an inertial measurement unit (IMU) in the vehicle, for example, the acceleration of the vehicle can be measured by an accelerometer in the IMU, the angular acceleration of the vehicle can be measured by a gyroscope in the IMU, and the like. The examples are provided herein only for the convenience of understanding the scheme, and the specific state information of the vehicle and the acquisition method thereof can be determined in combination with the actual application scenario.

[0092] Optionally, after obtaining the environment information of the current moment, the first device can also pre-process the environment information of the current moment, and the foregoing pre-processing includes but is not limited to: enlarging or reducing the size of the image, voxelizing the point cloud data or other pre-processing operations, etc.

[0093] 302, based on the environment information of the current moment, obtain the regulation and control information of the vehicle corresponding to each of at least two moments, the at least two moments including the current moment and at least one moment after the current moment, wherein the regulation and control information corresponding to each of the at least two moments is obtained by a machine learning model, the regulation and control information corresponding to the current moment is obtained by the machine learning model based on the environment information of the current moment, the first moment is any one of the at least one moment after the current moment, and the regulation and control information of the first moment is obtained by the machine learning model based on the regulation and control information corresponding to the last moment of the first moment.

[0094] For example, the "at least two time instants" in the present application can be at least two continuous time instants, and a time interval between two adjacent time instants in the at least two time instants can be a preset time step, for example, the time step can be 0.1 seconds, 0.2 seconds, 0.3 seconds, or other lengths of time, etc., and the specific length of the time step can be set in combination with an actual application scenario. For example, the at least two time instants can include a T0 time instant, a T1 time instant, a T2 time instant, a T3 time instant, a T4 time instant,..., a T19 time instant, and a T20 time instant, where the T0 time instant represents a current time instant, and any two adjacent time instants in the T0 time instant, the T1 time instant, the T2 time instant, the T3 time instant, the T4 time instant,..., the T19 time instant, and the T20 time instant can be separated by 0.1 seconds, the T1 time instant, the T3 time instant, the T4 time instant,..., the T19 time instant, and the T20 time instant represent 20 time instants within 2 seconds after the current time instant, and it should be understood that the examples herein are only for the convenience of understanding the present solution and do not limit the present solution.

[0095] For example, the control information of the vehicle corresponding to each of the at least two time instants can include at least one of the following: control information of the vehicle corresponding to each of the at least two time instants, a planned position of the vehicle corresponding to each of the at least two time instants, a driving strategy of the vehicle corresponding to each of the at least two time instants, or other types of control information, etc., which are not exhaustively listed herein.

[0096] For example, the control information of the vehicle corresponding to each of the at least two time instants can include at least one of the following: control information of the vehicle corresponding to each of the at least two time instants, a planned position of the vehicle corresponding to each of the at least two time instants, a driving strategy of the vehicle corresponding to each of the at least two time instants, or other types of control information, etc., which are not exhaustively listed herein.

[0097] For example, the meaning of the state information of the vehicle can refer to the description above, which is not repeated herein, and the state information of the planned vehicle corresponding to each of the at least one time instant after the current time instant can be understood as the state information that the vehicle needs to reach at each of the at least one time instant after the current time instant, for example, the planned state information corresponding to the T1 time instant includes a speed 1, an acceleration 1, and a direction 1 corresponding to the T1 time instant, which means that the speed of the vehicle needs to reach the speed 1, the acceleration of the vehicle needs to reach the acceleration 1, and the direction of the vehicle needs to reach the direction 1 at the T1 time instant, and it should be understood that the examples herein are only for the convenience of understanding the present solution and do not limit the present solution.

[0098] For example, the planned position of the vehicle corresponding to each of the at least two time points can also be understood as a planned trajectory point of the vehicle corresponding to each of the at least two time points, and the at least two planned trajectory points of the vehicle corresponding to the at least two time points can also be understood as a planned trajectory of the vehicle at the at least two time points.

[0099] For example, the driving strategy of the vehicle corresponding to each time point can include a lateral driving strategy and / or a longitudinal driving strategy corresponding to each time point, where the lateral driving strategy corresponding to a time point can be left turn, straight driving or right turn, the longitudinal driving strategy corresponding to a time point can be acceleration, constant speed or deceleration, and the specific implementation of the driving strategy can be determined in combination with the actual application scenario.

[0100] Optionally, in step 302, the first device can also obtain perception information corresponding to the environment information of the current time, that is, the output of the machine learning model can also include perception information corresponding to the environment information of the current time, where the perception information corresponding to the environment information of the current time includes road information and / or obstacle information in the environment of the current time. In the embodiment of the present application, not only can the regulation and control information of the current time and at least one time point after the current time be obtained through the machine learning model, but also the road information and / or obstacle information in the environment information of the current time can be obtained through the machine learning model, so that the user can obtain more information, which is beneficial to further improve the user experience of the present scheme.

[0101] For example, the road information in the environment can indicate at least one of the following: which lane lines exist in the environment, the position of the lane line, the type of the lane line, which lane the vehicle is located on, the position of the intersection, the sign information or other information in the road, etc. Here, no exhaustive enumeration is made, and the type of the lane can be single solid line, double solid line, dashed line, solid line + dashed line, etc.

[0102] For example, the obstacle information in the environment can include static obstacle information and dynamic obstacle information in the environment. The static obstacle information can indicate the position, type, footprint, shape or other information of the static obstacle in the environment, for example, the type of the static obstacle can be a fire hydrant, a cone barrel, a warning column, a crash barrel, a construction board or other static obstacles, etc. Here, no exhaustive enumeration is made. The dynamic obstacle information can indicate the type, position, speed, orientation or other information of the dynamic obstacle in the environment, for example, the type of the dynamic obstacle can be a social vehicle, a bus, a pedestrian, an electric vehicle or other types, etc. It should be understood that the examples herein are only for the convenience of understanding the present scheme.

[0103] Exemplarily, in one case, the first device is the execution device in the regulation information acquisition system shown in FIG. 2a, and the execution device is integrated in a vehicle in which an intelligent driving system is deployed. Step 301 can include: the first device collecting environment information at the current time (optionally, also including state information of the vehicle at the current time) through sensors deployed in the vehicle. Step 302 can include: the first device generating, based on the environment information at the current time (optionally, also including state information of the vehicle at the current time), regulation information of the vehicle corresponding to each of at least two time points (optionally, also including perception information corresponding to the environment information at the current time) through a machine learning model.

[0104] In another case, the first device is the vehicle in which an intelligent driving system is deployed in the regulation information acquisition system shown in FIG. 2b. Step 301 can include: the first device collecting environment information at the current time (optionally, also including state information of the vehicle at the current time) through sensors deployed in the vehicle. Step 302 can include: the first device sending the aforementioned environment information at the current time (optionally, also including state information of the vehicle at the current time) to the execution device, and receiving regulation information of the vehicle corresponding to each of at least two time points (optionally, also including perception information corresponding to the environment information at the current time) sent by the execution device.

[0105] In another case, the first device is the execution device in the regulation information acquisition system shown in FIG. 2b. Step 301 can include: the first device receiving environment information at the current time (optionally, also including state information of the vehicle at the current time) sent by the intelligent driving system in the vehicle. Step 302 can include: the first device generating, based on the environment information at the current time (optionally, also including state information of the vehicle at the current time), regulation information of the vehicle corresponding to each of at least two time points (optionally, also including perception information corresponding to the environment information at the current time) through a machine learning model.

[0106] The specific implementation process of obtaining regulation information of the vehicle corresponding to each of at least two time points through a machine learning model is described as follows. For convenience of description, in this application, any one of at least one time point after the current time is referred to as the first time point, and the regulation information corresponding to the current time is obtained by the machine learning model based on the environment information at the current time, and the regulation information of the first time point is obtained by the machine learning model based on the regulation information corresponding to the last time point of the first time point, in other words, the machine learning model can generate the regulation information corresponding to each of the aforementioned at least two time points in a serial manner.

[0107] Exemplarily, the execution device inputs the environment information of the current moment (optionally, also including the state information of the vehicle at the current moment) into the machine learning model, generates the regulation and control information of the vehicle corresponding to the current moment through the machine learning model; further generates the regulation and control information of the vehicle corresponding to the next moment of the current moment through the machine learning model based on the regulation and control information of the vehicle corresponding to the current moment; further generates the regulation and control information of the vehicle corresponding to the moment after the next moment of the current moment through the machine learning model based on the regulation and control information of the vehicle corresponding to the next moment of the current moment; and so on, until the regulation and control information of the vehicle corresponding to the last moment of the at least two moments is generated through the machine learning model, so as to obtain the regulation and control information of the vehicle corresponding to each moment of the at least two moments output by the machine learning model.

[0108] For example, the machine learning model can include a first submodule corresponding to the current moment and at least one second submodule corresponding to at least one moment after the current moment one by one, and the at least one second submodule is in a serial relationship. For example, the first submodule can be understood as a neural network module including at least one neural network layer, and each second submodule can also be understood as a neural network module including at least one neural network layer. The first submodule is used to generate the regulation and control information of the vehicle corresponding to the current moment, and the second submodule corresponding to the first moment in the at least one second submodule is used to obtain the regulation and control information of the vehicle at the first moment based on the regulation and control information of the vehicle at the moment before the first moment.

[0109] In order to more intuitively understand the scheme, at least two moments including T0 moment, T1 moment, T2 moment, T3 moment and T4 moment are taken as an example for illustration, T0 moment represents the current moment, T1 moment, T2 moment, T3 moment and T4 moment represent four moments after the current moment. When the environment information of the current moment (optionally, also including the state information of the vehicle at the current moment) is input into the machine learning model, the regulation and control information 0 of the vehicle corresponding to T0 moment is generated through the machine learning model; further, the regulation and control information 1 of the vehicle corresponding to T1 moment is generated through the machine learning model based on the regulation and control information 0; further, the regulation and control information 2 of the vehicle corresponding to T2 moment is generated through the machine learning model based on the regulation and control information 1; further, the regulation and control information 3 of the vehicle corresponding to T3 moment is generated through the machine learning model based on the regulation and control information 2; further, the regulation and control information 4 of the vehicle corresponding to T4 moment is generated through the machine learning model based on the regulation and control information 3, so as to obtain the regulation and control information 0, regulation and control information 1, regulation and control information 2, regulation and control information 3 and regulation and control information 4 corresponding to T0 moment, T1 moment, T2 moment, T3 moment and T4 moment one by one output by the machine learning model. It should be understood that the example herein is only for the convenience of understanding the scheme and is not used to limit the scheme.

[0110] Optionally, the regulation information of the vehicle corresponding to the first time (i.e., any time after the current time) is obtained by the machine learning model based on the predicted feature information of the environment at the first time; wherein the predicted feature information of the environment at the first time is obtained by the machine learning model based on the regulation information corresponding to the last time before the first time; optionally, the predicted feature information of the environment at the first time is obtained by the machine learning model based on the regulation information corresponding to the last time before the first time and the feature information of the environment at the last time before the first time. In the embodiments of the present application, since the machine learning model generates the regulation information corresponding to the first time (i.e., any time after the current time) based on the feature information of the environment at the last time before the first time and the regulation information corresponding to the last time before the first time, i.e., the machine learning model imagines the environment at the future time after the current time in the hidden layer space, and since the machine learning model in the present application has the ability to imagine the non-existent environment, the traffic environment seen by the machine learning model is greatly expanded, which is conducive to reducing the probability of the machine learning model encountering a completely unseen traffic environment, and is also conducive to improving the response capability of the machine learning model to the unseen traffic environment, i.e., the above covariate problem can be solved, in other words, it is conducive to improving the performance of the regulation information of the vehicle generated by the machine learning model, so as to improve the safety and reliability of the intelligent driving process.

[0111] For example, if the last time before the first time is the current time, the feature information of the environment at the last time before the first time can be the feature information obtained based on the environment information at the current time (i.e., the real environment information); if the last time before the first time is a time after the current time, the feature information of the environment at the last time before the first time can be the predicted feature information of the environment at the last time before the first time.

[0112] For a more intuitive understanding of the scheme, at least two time points are taken as examples here, including T0 time, T1 time, T2 time, T3 time and T4 time, T0 time represents the current time, T1 time, T2 time, T3 time and T4 time represent four time points after the current time, in the machine learning model, the current time environment information (optionally, also including the current time vehicle state information) is input into the machine learning model, and the feature information 0 corresponding to the T0 time environment information and the vehicle control information 0 corresponding to the T0 time are generated through the machine learning model; Then, through the machine learning model, the predicted feature information 1 of the T1 time environment is generated based on the feature information 0 and the control information 0, and the vehicle control information 1 corresponding to the T1 time is generated based on the predicted feature information 1 of the T1 time environment; Then, through the machine learning model, the predicted feature information 2 of the T2 time is generated based on the predicted feature information 1 and the control information 1, and the vehicle control information 2 corresponding to the T2 time is generated based on the predicted feature information 2 of the T2 time environment; Then, through the machine learning model, the predicted feature information 3 of the T3 time is generated based on the predicted feature information 2 and the control information 2, and the vehicle control information 3 corresponding to the T3 time is generated based on the predicted feature information 3 of the T3 time environment; Then, through the machine learning model, the predicted feature information 4 of the T4 time is generated based on the predicted feature information 3 and the control information 3, and the vehicle control information 4 corresponding to the T4 time is generated based on the predicted feature information 4 of the T4 time environment, the vehicle control information 0, vehicle control information 1, vehicle control information 2, vehicle control information 3 and vehicle control information 4 corresponding to T0 time, T1 time, T2 time, T3 time and T4 time are obtained, which are output by the machine learning model, It should be understood that the examples here are only for easy understanding of the scheme and do not limit the scheme.

[0113] Optionally, the predicted feature information of the environment at the first time point (i.e., any one of the at least one time point after the current time point) is generated by a first module in the machine learning model, and the machine learning model further includes a second module. The input of the second module includes the environment information at the current time point, and the feature information corresponding to the environment information at the current time point is obtained through the second module. Optionally, the input of the second module further includes the state information of the vehicle at the current time point, the map information corresponding to the vehicle position, or other information, etc. The first module is configured to obtain the regulation and control information corresponding to each of the at least two time points based on the feature information generated by the second module. In the embodiments of the present application, the architecture of the machine learning model is refined, and the implementation difficulty of the present solution is reduced. The entire machine learning model is divided into the first module and the second module, the input of the second module includes the environment information at the current time point, and the feature information corresponding to the environment information at the current time point is obtained through the second module. The first module is used to generate the regulation and control information at the current time point, the predicted feature information of the environment at each of the at least one time point after the current time point, and the regulation and control information. The entire machine learning model is more finely managed, i.e., the process of generating the regulation and control information of the vehicle corresponding to each of the at least two time points by using the machine learning model is more finely managed, which is beneficial to improving the quality of the finally obtained regulation and control information, and further beneficial to improving the user experience of the vehicle in the intelligent driving process.

[0114] Optionally, the first module and the second module can both be neural network modules, i.e., the first module can include a plurality of neural network layers, and the second module can also include a plurality of neural network layers. For example, the first module can be a convolutional neural network, a fully connected neural network, a residual neural network, a neural network based on an attention mechanism, etc., and the second module can also be a convolutional neural network, a fully connected neural network, a residual neural network, a neural network based on an attention mechanism, etc. The specific implementation can be set according to the actual application scenario.

[0115] Exemplarily, after the execution device deploying the machine learning model inputs the environment information at the current time point (optionally, further including the state information of the vehicle at the current time point) into the machine learning model, the feature information corresponding to the environment information at the current time point can be generated through the second module in the machine learning model, and the regulation and control information corresponding to each of the at least two time points can be obtained through the first module in the machine learning model based on the feature information generated by the second module. It should be noted that if the first device is the execution device deploying the machine learning model, the foregoing steps can also be understood as being executed by the first device.

[0116] For a more intuitive understanding of the scheme, please refer to FIG. 4, which is a schematic diagram of a machine learning model provided by an embodiment of the present application. As shown in FIG. 4, the machine learning model can include a first module and a second module. After the environment information at the current time is input into the machine learning model, the second module in the machine learning model can generate feature information corresponding to the environment at the current time. Then, the first module of the machine learning model generates regulation and control information corresponding to each of the at least two time points based on the feature information corresponding to the environment at the current time. It should be understood that the example in FIG. 4 is only for the convenience of understanding the scheme and does not limit the scheme.

[0117] Optionally, the training phase of the first module is reinforcement learning. Optionally, the training phase of the first module is obtained after reinforcement learning in a simulation environment using a simulator. The specific implementation process of the training phase of the first module will be described later when the training phase of the machine learning model is described. Here, no further description is given. In an embodiment of the present application, since the predicted feature information of the environment at the first time (i.e., any time after the current time) is generated by the first module, and the first module is also used to generate regulation and control information corresponding to the current time and at least one time point after the current time, the training phase of the first module adopts the reinforcement learning mode. Therefore, in the training phase of the first module, the first module can fully explore the feature information of various traffic environments, obtain corresponding regulation and control information based on various traffic environments, and then evaluate the foregoing regulation and control information using the reinforcement learning mode. That is, the reinforcement learning mode can feed back the regulation and control information generated by the machine learning model for various traffic environments, so that the first module can learn the response capability for various traffic environments in the training phase, greatly expanding the traffic scenarios that the machine learning model can respond to. This is conducive to improving the quality of the regulation and control information generated by the machine learning model in various application environments, thereby improving the safety and reliability of the intelligent driving process of the vehicle and improving the user experience of the intelligent driving process.

[0118] For example, the feature information generated by the second module in the machine learning model can be understood as information obtained by processing the environment information at the current time (optionally, also including the state information of the vehicle at the current time) based on the neural network layer in the second module in the machine learning model. Optionally, the feature information generated by the second module includes a semantic image from a top-down perspective (for the convenience of description, hereinafter referred to as "first semantic image"). The top-down perspective in the present application can also be understood as a bird's eye view (BEV). Therefore, the top-down perspective in the present application can also be referred to as a BEV perspective. The foregoing first semantic image indicates the semantic category of at least one region in the image from the BEV perspective corresponding to the environment information at the current time.

[0119] For example, the image in the BEV perspective can be an image after semantic segmentation, and optionally, the image after semantic segmentation includes at least one region, different regions in the at least one region can adopt different colors, image regions of the same semantic category adopt the same color, and image regions of different semantic categories adopt different colors. The first semantic image can be understood as each type of object in the image in the segmented BEV perspective plus a semantic category. The semantic category in this application can also be referred to as a semantic label. For example, the semantic category can be a lane, an intersection, a ramp, a traffic sign, a static obstacle, a social vehicle, a bus, a pedestrian, or another semantic category. The examples are only for easy understanding of the scheme, and the specific application scenarios can be determined according to actual application scenarios.

[0120] In the embodiments of the application, the second module in the machine learning model generates a semantic image in the BEV perspective based on the environment information at the current moment. Since the semantic image in the BEV perspective can not only reflect the environment in the BEV perspective, but also carry the semantic category of at least one region in the image in the BEV perspective, that is, the second module in the machine learning model obtains effective information in the environment at the current moment based on the environment information at the current moment, and then the first module in the machine learning model generates the regulation and control information of the vehicle corresponding to each moment in the at least two moments based on the semantic image in the BEV perspective, which is beneficial to obtain more accurate regulation and control information. In addition, if the training stage of the first module in the machine learning model is reinforcement learning, since reinforcement learning is generally trained in a simulation environment with the help of a simulator, and the environment information obtained from the simulation environment is very different from the environment information obtained from the real traffic environment. If the environment information is directly input into the first module, the environment information input into the first module in the training stage and the application stage may be very different, which may result in poor quality of the regulation and control information generated by the first module in the application stage. Therefore, the second module is independently set to generate a semantic image in the BEV perspective based on the environment information at the current moment, and the input of the first module includes the semantic image in the BEV perspective, which can make the data input into the first module in the training stage and the application stage as similar as possible, so as to improve the quality of the regulation and control information generated by the first module, and further improve the user experience of the intelligent driving process of the vehicle.

[0121] Optionally, the feature information generated by the second module further includes first feature information obtained by feature extraction based on the environment information at the current time. Illustratively, the execution device inputs the environment information at the current time (optionally, also including the state information of the vehicle at the current time) into the machine learning model, and obtains the first feature information by feature extraction of the input information by the second module in the machine learning model, and obtains the first semantic image in the BEV perspective based on the first feature information by feature processing by the second module in the machine learning model. Further, based on the first semantic image in the BEV perspective and the aforementioned first feature information, the regulation and control information of the vehicle corresponding to each of the at least two time points is generated by the first module in the machine learning model; it should be noted that if the first device is an execution device deploying the machine learning model, the aforementioned steps can also be understood as being executed by the first device.

[0122] Optionally, the feature information generated by the second module can also include feature information in the BEV perspective; illustratively, the execution device can also obtain the feature information in the BEV perspective based on the first feature information by feature processing by the second module in the machine learning model; the feature information in the BEV perspective can be understood as the feature information of the environment at the current time in the BEV perspective, that is, based on the environment information at the current time, the feature information of the environment at the current time in the BEV perspective is determined by the second module. The execution device can also generate the regulation and control information of the vehicle corresponding to each of the at least two time points by the first module in the machine learning model based on the first semantic image in the BEV perspective, the feature information in the BEV perspective, and the aforementioned first feature information.

[0123] Illustratively, the feature information in the BEV perspective can be generated by the first neural network layer in the second module included in the machine learning model, and the first semantic image in the BEV perspective can be generated by the second neural network layer in the second module included in the machine learning model; optionally, the first neural network layer is located before the second neural network layer, for example, the execution device inputs the environment information at the current time (optionally, also including the state information of the vehicle at the current time) into the machine learning model, and obtains the first feature information by feature extraction of the input information by the second module in the machine learning model, and in the process of feature processing by the second module in the machine learning model based on the first feature information, the feature information in the BEV perspective is first generated by the first neural network layer in the second module, and then the first semantic image in the BEV perspective is generated by the second neural network layer in the second module; it should be noted that if the first device is an execution device deploying the machine learning model, the aforementioned steps can also be understood as being executed by the first device.

[0124] For a more intuitive understanding of the scheme, please refer to FIG. 5, which is another schematic diagram of the machine learning model provided by the embodiments of the present application. As shown in FIG. 5, the machine learning model can include a first module and a second module. The environmental information and the state information of the vehicle at the current time are input into the machine learning model, the first feature information is obtained by feature extraction through the second module in the machine learning model, the feature processing is performed based on the first feature information through the second module in the machine learning model, and the first semantic image under the BEV perspective and the feature information under the BEV perspective are obtained. The rule control information corresponding to each time in at least two time periods is generated based on the first semantic image under the BEV perspective, the feature information under the BEV perspective, and the first feature information through the first module in the machine learning model. It should be understood that the example in FIG. 5 is only for the convenience of understanding the scheme and is not used to limit the scheme.

[0125] In the embodiments of the present application, the feature information generated by the second module not only includes the image under the BEV perspective, but also includes the first feature information obtained by feature extraction based on the environmental information at the current time. Therefore, in the process of generating the rule control information, the first module of the machine learning model can not only utilize the effective information (i.e., the image under the BEV perspective) extracted by the second module based on the environmental information, but also obtain more comprehensive first feature information corresponding to the environmental information, i.e., not only a clearer understanding of the environment at the current time, but also a more comprehensive understanding of the environment at the current time, which is conducive to further improving the quality of the obtained rule control information.

[0126] Alternatively, the feature information generated by the second module can also include the second semantic image under the surround view perspective corresponding to the environment at the current time; or the feature information generated by the second module can also include the feature information under the surround view perspective corresponding to the environment at the current time, etc. What kind of feature information is generated by the second module can be determined in combination with the actual application scenario, which is not limited here.

[0127] Optionally, the first module in the machine learning model includes a first sub-module corresponding to the current time and at least one second sub-module corresponding to at least one time after the current time in a one-to-one manner. The at least one second sub-module can be in a serial structure.

[0128] The first submodule corresponding to the current time is configured to generate the regulation and control information of the vehicle corresponding to the current time based on the feature information generated by the second module in the machine learning model. The second submodule corresponding to the first time (i.e., any one of the at least one time after the current time) is configured to generate the regulation and control information of the vehicle corresponding to the first time based on the regulation and control information of the vehicle corresponding to the time before the first time and the predicted feature information of the environment of the first time. For example, the second submodule corresponding to the first time is configured to generate the predicted feature information of the environment of the first time based on the regulation and control information of the vehicle corresponding to the time before the first time, and then generate the regulation and control information of the vehicle corresponding to the first time based on the predicted feature information of the environment of the first time.

[0129] For example, at least two times are taken as T0, T1, T2, T3 and T4, i.e., the machine learning model is taken as an example to generate the regulation and control information of the vehicle corresponding to the current time and four times after the current time. T0 represents the current time, T1, T2, T3 and T4 represent four times after the current time. The second module in the machine learning model includes one first submodule and four second submodules, and the four second submodules include second submodule 1, second submodule 2, second submodule 3 and second submodule 4. The first submodule is configured to generate the regulation and control information 0 of the vehicle corresponding to T0. The second submodule 1 is configured to generate the regulation and control information 1 of the vehicle corresponding to T1 based on the regulation and control information 0 and the predicted feature information 1 of the environment of T1. The second submodule 2 is configured to generate the regulation and control information 2 of the vehicle corresponding to T2 based on the regulation and control information 1 and the predicted feature information 2 of the environment of T2. The second submodule 3 is configured to generate the regulation and control information 3 of the vehicle corresponding to T3 based on the regulation and control information 2 and the predicted feature information 3 of the environment of T3. The second submodule 4 is configured to generate the regulation and control information 4 of the vehicle corresponding to T4 based on the regulation and control information 3 and the predicted feature information 4 of the environment of T4. It should be understood that the examples are only for the convenience of understanding the present scheme.

[0130] For a more intuitive understanding of the scheme, please refer to FIG. 6, which is a schematic diagram of a first module of a machine learning model provided by an embodiment of the present application. Here, the machine learning model is used to generate control information at T0, T1, T2, T3 and T4, where T0 represents the current time, and T1, T2, T3 and T4 represent four time points after the current time. As shown in FIG. 6, the first module in the machine learning model includes one first submodule and four second submodules. The one first submodule and the four second submodules are in a serial relationship. The four second submodules include second submodule 1, second submodule 2, second submodule 3 and second submodule 4. The first submodule is used to generate control information at T0. The second submodule 1 is used to generate control information at T1. The second submodule 2 is used to generate control information at T2. The second submodule 3 is used to generate control information at T3. The second submodule 4 is used to generate control information at T4. It should be understood that the example in FIG. 6 is only for the convenience of understanding the scheme and does not limit the scheme.

[0131] In the embodiment of the present application, the first module includes a first submodule corresponding to the current time and at least one second submodule corresponding to at least one time point after the current time. Each second submodule is specially used to generate control information at a time point after the current time, which further improves the refinement degree of the machine learning model, i.e., the process of using the machine learning model to generate vehicle control information corresponding to each time point in at least two time points is more finely managed to further improve the quality of the obtained control information.

[0132] Alternatively, the first module in the machine learning model comprises a first submodule corresponding to the current time and a second submodule corresponding to at least one time after the current time, and the control information of the vehicle corresponding to at least one time after the current time can be generated by calling the second submodule at least once. For example, at least two time points are taken as T0, T1, T2, T3 and T4, T0 represents the current time, T1, T2, T3 and T4 represent four time points after the current time, the second module in the machine learning model comprises one first submodule and one second submodule; wherein the first submodule is used to generate the control information 0 of the vehicle corresponding to T0; the second submodule is based on the control information 0 and the predicted feature information 1 of the environment of T1 to obtain the control information 1 of the vehicle corresponding to T1, and then the second submodule is based on the control information 1 and the predicted feature information 2 of the environment of T2 to obtain the control information 2 of the vehicle corresponding to T2, and then the second submodule is based on the control information 2 and the predicted feature information 3 of the environment of T3 to obtain the control information 3 of the vehicle corresponding to T3, and then the second submodule 4 is based on the control information 3 and the predicted feature information 4 of the environment of T4 to obtain the control information 4 of the vehicle corresponding to T4, it should be understood that the example is only for the convenience of understanding the scheme.

[0133] Alternatively, the second module in the machine learning model generates feature information corresponding to the last time of the first time, and the second feature information is obtained by the second module in the machine learning model based on the first feature information and the control information of the vehicle corresponding to the last time of the first time. Then the first module in the machine learning model also uses the second feature information in the process of generating the control information of the vehicle corresponding to each time in the at least two time points. For example, the first module in the machine learning model obtains the control information of the vehicle corresponding to each time in the at least two time points based on the first semantic image in the BEV perspective, the first feature information and the second feature information; or the first module in the machine learning model obtains the control information of the vehicle corresponding to each time in the at least two time points based on the first semantic image in the BEV perspective and the second feature information; or the first module in the machine learning model obtains the control information of the vehicle corresponding to each time in the at least two time points based on the first semantic image in the BEV perspective, the feature information in the BEV perspective, the first feature information and the second feature information.

[0134] In the embodiments of the present application, the second module of the machine learning model can also obtain second feature information corresponding to the previous moment of the first moment based on the first feature information and the regulation and control information corresponding to the previous moment of the first moment, that is, in the process of generating the regulation and control information by the first module in the machine learning model, the second feature information obtained based on the first feature information and the regulation and control information corresponding to the previous moment of the first moment can also be used, and more comprehensive feature information is used to generate the regulation and control information, which is beneficial to obtaining regulation and control information with higher quality.

[0135] Optionally, the predicted feature information of the environment at the first moment is obtained by the first module based on the feature information of the environment at the previous moment of the first moment, the regulation and control information corresponding to the previous moment of the first moment, and the second feature information. For example, in the process of generating the regulation and control information at the first moment by the first module in the machine learning model, the predicted feature information of the environment at the first moment can be obtained by the first module in the machine learning model based on the feature information of the environment at the previous moment of the first moment, the regulation and control information corresponding to the previous moment of the first moment, and the second feature information; and then the regulation and control information at the first moment is generated by the first module in the machine learning model based on the predicted feature information of the environment at the first moment. In the embodiments of the present application, the second feature information corresponding to the previous moment of the first moment is used to generate the predicted feature information of the environment at the first moment by the first module, and the predicted feature information of the environment at the first moment is determined based on the feature information of the environment at the previous moment of the first moment, the regulation and control information corresponding to the previous moment of the first moment, and the second feature information, which is beneficial to reducing the difficulty of predicting the feature information of the environment at the first moment, thereby assisting the first module to better predict the feature information of the environment at the first moment, and obtaining better predicted feature information of the environment at the first moment, which is beneficial to generating better regulation and control information.

[0136] In order to more intuitively understand the present scheme, please refer to FIG. 7, which is another schematic diagram of the machine learning model provided by the embodiments of the present application. As shown in FIG. 7, the environment information at the current moment and the state information of the vehicle are input into the machine learning model, the first feature information is obtained by the second module of the machine learning model through feature extraction on the environment information at the current moment and the state information of the vehicle, and the first semantic image under the BEV perspective and the feature information under the BEV perspective are obtained by the second module of the machine learning model through feature processing based on the first feature information. Based on the first semantic image under the BEV perspective, the feature information under the BEV perspective, and the first feature information, the feature information of the environment at T0 is generated by the first submodule in the first module of the machine learning model, and the regulation and control information 0 at T0 is generated by the first submodule based on the feature information of the environment at T0.

[0137] Further, based on the first feature information and the regulation and control information 0 at the T0 moment, the feature is updated through the second module in the machine learning model to obtain the second feature information 0 generated by the second module; through the second submodule 1 in the first module of the machine learning model, based on the regulation and control information 0 at the T0 moment, the feature information of the environment at the T0 moment and the second feature information 0, the predicted feature information 1 of the environment at the T1 moment is generated, and based on the predicted feature information 1 of the environment at the T1 moment, the regulation and control information 1 at the T1 moment is generated through the second submodule 1.

[0138] Further, based on the first feature information and the regulation and control information 1 at the T1 moment, the feature is updated through the second module in the machine learning model to obtain the second feature information 1 generated by the second module, and based on the regulation and control information 1 at the T1 moment and the second feature information 1, the predicted feature information 2 of the environment at the T2 moment is generated through the second submodule 2 in the first module of the machine learning model, and based on the predicted feature information 2 of the environment at the T2 moment, the regulation and control information 2 at the T2 moment is generated through the second submodule 2. Similarly, the regulation and control information 3 at the T3 moment and the regulation and control information 4 at the T4 moment are obtained. It should be understood that the example in FIG. 7 is only for the convenience of understanding the scheme and does not limit the scheme.

[0139] Optionally, after obtaining the regulation and control information corresponding to each of the at least two moments, the first device or the execution device of the machine learning model can further perform smoothing processing on all the regulation and control information corresponding to the at least two moments; for example, if the regulation and control information corresponding to each of the at least two moments includes the control signals of the current moment and the 10 moments after the current moment, since the prediction operation can be performed once every 1 second through the machine learning model, the control signals of the 10 moments after the current moment in the 2 historical rounds before the current round can be obtained, and the overlapping part is weighted and averaged to obtain the control signal after the smoothing processing.

[0140] For example, the control signals obtained in the current round include the control signals of T0, T1, …, T10, T0 is the current moment, the control signals obtained in the last historical round of the current round include the control signals of T-1, T0, T1, …, T9, the control signals obtained in the second last historical round of the current round include the control signals of T-2, T-1, T0, T1, …, T8, then the control signals of T0 moment in the current round, the last historical round of the current round and the second last historical round of the current round can be weighted and averaged to obtain the rule control information after smoothing processing of the T0 moment in the current round; the control signals of T1 moment in the current round, the last historical round of the current round and the second last historical round of the current round are weighted and averaged to obtain the rule control information after smoothing processing of the T1 moment in the current round; and so on, until the control signals of T0, T1, …, T10 after smoothing processing in the current round are obtained. It should be understood that the examples herein are only for the convenience of understanding the scheme.

[0141] Optionally, the machine learning model further includes a third module for generating perception information corresponding to the environment information of the current moment, wherein the perception information includes road information and / or obstacle information in the environment of the current moment. For further understanding of the meanings of the road information and the obstacle information, reference can be made to the above description, which will not be repeated here.

[0142] Optionally, the third module in the machine learning model obtains the perception information corresponding to the environment information of the current moment based on the feature information generated by the second module; for example, the third module in the machine learning model generates the perception information corresponding to the environment information of the current moment based on the first semantic image in the BEV perspective; or the third module in the machine learning model generates the perception information corresponding to the environment information of the current moment based on the first semantic image in the BEV perspective and the feature information in the BEV perspective; or the third module in the machine learning model generates the perception information corresponding to the environment information of the current moment based on the first semantic image in the BEV perspective and the first feature information; or the third module in the machine learning model generates the perception information corresponding to the environment information of the current moment based on the first semantic image in the BEV perspective, the feature information in the BEV perspective and the first feature information, etc. The specific implementation mode can be determined in combination with the actual application scenario.

[0143] In the embodiments of the present application, after obtaining the environment information of the current time each time, the regulation and control information corresponding to a plurality of continuous time points is generated, so that the machine learning model not only considers the current situation when generating the regulation and control information, but also considers the influence on the future time, which is beneficial to improve the continuity of the vehicle control process; and the regulation and control information corresponding to the current time is obtained by the machine learning model based on the environment information of the current time, and the regulation and control information of the first time is obtained by the machine learning model based on the regulation and control information corresponding to the last time of the first time, that is, the machine learning model generates the regulation and control information corresponding to the current time and the time after the current time in a serial manner, and considers the influence of the regulation and control information of the last time on the generation of the regulation and control information of the first time, which is beneficial to further improve the continuity between the regulation and control information of the vehicle corresponding to a plurality of time points, and thus is beneficial to improve the continuity of the vehicle control process.

[0144] In order to further understand the present scheme, the method provided by the present application is described in combination with the specific traffic scene of automatic parking, and here the execution device of the machine learning model is integrated in the vehicle in which the intelligent driving system is deployed. For example, after the intelligent driving system in the vehicle determines that the user starts the automatic parking function, the target position that the vehicle needs to reach can be generated according to the target parking space selected by the user. Then, the intelligent driving system in the vehicle can collect the environment information around the vehicle and the state information of the vehicle at the current time through the sensors deployed in the vehicle. The environment information around the vehicle can be the environment information within the visual range of the sensors deployed in the vehicle, and the environment information can include image and point cloud data.

[0145] The intelligent driving system in the vehicle can input the environment information around the vehicle and the state information of the vehicle at the current time into the machine learning model, generate the control signal corresponding to each of the at least two time points and the planning position through the machine learning model, and generate the perception information corresponding to the current environment through the machine learning model; wherein the at least two time points include the current time and at least one time point after the current time, the control signal of the current time is obtained by the machine learning model based on the environment information of the current time, the planning position of the current time can be understood as the actual position of the vehicle at the current time, the first time is any one of the at least one time point after the current time, and the control signal and the planning position of the first time are obtained based on the control signal and the planning position of the last time of the first time. The foregoing control signal corresponding to each of the at least two time points can be understood as the control signal for controlling the vehicle to automatically park, and the foregoing planning position corresponding to each of the at least two time points can form the moving track planned by the vehicle during the automatic parking process.

[0146] The intelligent driving system in the vehicle outputs the perception information corresponding to the environment at the current moment and the planning position corresponding to each of the at least two moments to a display system of the vehicle, and then displays the planning position on a center console; and based on the control signals corresponding to each of the at least two moments, the intelligent driving system sends the control signals to the components in the vehicle to control the vehicle to automatically park. It should be understood that the example is only for facilitating understanding of the present scheme and is not used to limit the present scheme.

[0147] II. Training phase

[0148] Optionally, in the case that the machine learning model comprises the first module and the second module, the training phase of the machine learning model can be divided into a training phase of the first module and a training phase of the second module. Optionally, the training phase of the first module can be reinforcement learning. The training phase of the first module and the training phase of the second module are described as follows, respectively.

[0149] (I) Training phase of the first module

[0150] Please refer to FIG. 8, which is a flowchart of the training method of the model provided by the embodiments of the present application. The training method of the model provided by the embodiments of the present application can comprise:

[0151] 801. Input the first training sample into the first module to obtain the predicted feature information of the environment at each of the at least one moment after the current moment and the predicted regulation and control information of the vehicle corresponding to each of the at least two moments generated by the first module.

[0152] For example, in the training phase of the first module, the training device can obtain the first training sample. Optionally, the first training sample can be a semantic image in the BEV perspective, or the first training sample can also be a semantic image in the perspective, etc. The specific information contained in the first training sample can be determined in combination with the information input into the first module described in the above-mentioned embodiments corresponding to FIG. 3.

[0153] After inputting the first training sample into the first module, the training device can generate, through the first module, the predicted regulation and control information of the vehicle corresponding to the current moment, the predicted feature information of the environment at each of the at least one moment after the current moment, and the predicted regulation and control information of the vehicle corresponding to each of the at least one moment after the current moment.

[0154] 802. Train the first module.

[0155] Exemplarily, the training device can generate a function value of a loss function (hereinafter referred to as a "first loss function" for convenience of distinction) of the first module based on the predicted regulation and control information of the vehicle corresponding to each of the at least two time points, update the weight parameters of the first module by using a back propagation algorithm based on the function value of the first loss function, so as to realize one training of the first module; the training device repeatedly executes steps 801 and 802 for multiple times, so as to realize iterative training of the first module until a convergence condition is met. The convergence condition can include that the number of times of training of the first module reaches a preset number of times, and / or the convergence condition of the first loss function is met.

[0156] Optionally, the training device can train the first module by using a reinforcement learning manner; exemplarily, the first loss function can indicate a benefit of the predicted regulation and control information of the vehicle corresponding to each of the at least two time points, and exemplarily, the benefit of the predicted regulation and control information of the vehicle corresponding to each of the time points can include a short-term benefit and a long-term benefit. The target of training the first module by using the first loss function includes improving the benefit of the predicted regulation and control information of the vehicle corresponding to each of the at least two time points. Alternatively, if the training stage of the first module adopts a supervised learning manner, the first loss function of the first module can indicate a similarity between the predicted regulation and control information and the expected regulation and control information of the vehicle corresponding to each of the time points. The target of training the first module by using the first loss function includes improving the similarity between the predicted regulation and control information and the expected regulation and control information of the vehicle corresponding to each of the time points.

[0157] Optionally, in the training stage of the first module, the training device can further generate the predicted environment information corresponding to each of the at least one time point after the current time point by performing feature processing on the predicted feature information of the environment of each of the at least one time point after the current time point. For example, the predicted feature information of the environment of each of the at least one time point after the current time point is input into a decoder to obtain the predicted environment information corresponding to each of the at least one time point after the current time point generated by the decoder. The training device can also obtain the correct environment information corresponding to the predicted environment information by using the simulator, that is, obtain the correct environment information corresponding to each of the at least one time point after the current time point by using the simulator.

[0158] Optionally, the predicted environment information and the correct environment information can specifically be images in a BEV perspective reflecting the predicted environment; optionally, the image in the BEV perspective can be a semantic image in the BEV perspective. For example, taking the first module to generate the predicted feature information of the environment at T1, T2, T3 and T4 as an example, the training device controls the vehicle to execute the control information 0 corresponding to T0 in the simulation environment simulated by the simulator, and the training device can obtain the real environment at T1 in the simulation environment, so as to obtain the correct environment information corresponding to T1. The training device controls the vehicle to continue to execute the control information 1 corresponding to T1 in the simulation environment simulated by the simulator, and the training device can obtain the real environment at T2 in the simulation environment, so as to obtain the correct environment information corresponding to T2. In this way, the training device can obtain the correct environment information corresponding to each time after the current time through the simulator. It should be understood that the above examples are only for the convenience of understanding the present application, and are not used to limit the present application.

[0159] The first loss function of the training stage of the first module also indicates the similarity between the predicted environment information and the correct environment information corresponding to each time, and the target of training the first module by using the first loss function also includes improving the similarity between the predicted environment information and the correct environment information corresponding to each time.

[0160] In the training stage of the first module, since the training stage of the first module is executed in the simulator, the correct environment information of at least one time after the current time can be obtained through the simulator, the predicted environment information is obtained by performing reverse feature processing on the predicted feature information of the environment generated by the first module, and the similarity between the predicted environment information and the correct environment information is narrowed by using the loss function of the first module. In this way, the first module is trained to generate more accurate predicted feature information of the environment, that is, the accuracy of the first module in predicting the feature information of the environment is improved.

[0161] Optionally, after obtaining the correct environment information of each of the at least one time point after the current time point, the training device can perform feature extraction on the correct environment information of each of the at least one time point after the current time point to obtain expected feature information of the environment of each of the at least one time point. The first loss function of the training stage of the first module can also indicate the similarity between the predicted feature information of the environment of each of the at least one time point after the current time point generated by the first module and the expected feature information, and the target of training the first module by using the first loss function also includes improving the similarity between the predicted feature information of the environment of each of the at least one time point and the expected feature information, the expected feature information of the environment of each of the at least one time point after the current time point is obtained based on the correct environment information of each of the at least one time point, and the correct environment information of each of the at least one time point is obtained by using the simulator.

[0162] In the training stage of the first module, the correct environment information of each of the at least one time point after the current time point can also be subjected to feature extraction to obtain expected feature information of the environment of each of the at least one time point, and the loss function of the first module can be used to narrow the similarity between the predicted feature information of the environment of each of the at least one time point and the expected feature information, so as to train the first module to generate more accurate predicted feature information of the environment, and also to improve the accuracy of the first module in predicting the feature information of the environment.

[0163] (II) Training stage of the second module

[0164] Referring to FIG. 9, FIG. 9 is another flowchart of the training method of the model provided in the embodiments of the present application. The training method of the model provided in the embodiments of the present application can include:

[0165] 901, inputting the second training sample into the machine learning model.

[0166] For example, after performing the training operation of the first module in the machine learning model, the training device can also train the second module in the machine learning model. In the training stage of the second module, the training device can obtain a second training sample, which can include environment information, and optionally, the second training sample can include state information of the vehicle, or can also include map information corresponding to the position of the vehicle, etc. The training device inputs the second training sample into the machine learning model to generate predicted regulation and control information of the vehicle corresponding to each of the at least two time points through the machine learning model.

[0167] Exemplarily, the training device inputs the second training sample into the machine learning model, and the feature information can be generated by the second module in the machine learning model. Based on the feature information generated by the second module, the first module in the machine learning model generates the predicted regulation and control information of the vehicle corresponding to each time in the at least two times. The training device can also obtain the predicted third feature information generated by the first module in the machine learning model in the process of data processing based on the feature information generated by the second module.

[0168] Optionally, in the training phase of the second module, the training device can also input the first training sample corresponding to the second training sample into the first module of the machine learning model, and obtain the expected third feature information generated by the first module in the process of data processing based on the first training sample. The predicted third feature information and the expected third feature information can be generated by the same neural network layer in the first module. Optionally, the first training sample corresponding to the second training sample can be a correct semantic image in the BEV perspective corresponding to the environmental information.

[0169] 902, keep the weight parameters of the first module in the machine learning model unchanged, and train the second module in the machine learning model.

[0170] Exemplarily, the training device can generate the function value of the loss function (for convenience, referred to as "second loss function" hereinafter) of the second module based on the predicted regulation and control information of the vehicle corresponding to each time in the at least two times, keep the weight parameters of the first module in the machine learning model unchanged, update the weight parameters of the second module based on the function value of the second loss function by using the back propagation algorithm, so as to realize one training of the second module in the machine learning model. The training device repeatedly executes steps 801 and 802 for multiple times to realize the iterative training of the first module until the convergence condition is met.

[0171] Exemplarily, the training phase of the second module can adopt a supervised learning manner. The second loss function can indicate the similarity between the predicted regulation and control information of the vehicle corresponding to each time and the expected regulation and control information. The target of training the second module by using the second loss function includes improving the similarity between the predicted regulation and control information of the vehicle corresponding to each time and the expected regulation and control information.

[0172] Optionally, the second loss function can also indicate the similarity between the predicted third feature information and the expected third feature information, and the training target of training the second module by using the second loss function includes improving the similarity between the predicted third feature information and the expected third feature information. Since the predicted third feature information is obtained by the second module based on the feature information generated by the first module, and the expected third feature information is obtained by the second module based on the correct semantic image in the BEV perspective, by adding the training target of improving the similarity between the predicted third feature information and the expected third feature information, the similarity between the semantic image in the BEV perspective included in the feature information generated by the second module and the correct semantic image in the BEV perspective can be assisted, so as to further improve the quality of the feature information generated by the second module.

[0173] On the basis of the embodiments corresponding to FIG. 1 to FIG. 9, in order to better implement the above-mentioned scheme of the embodiments of the present application, the related device for implementing the above-mentioned scheme is further provided below. Referring to FIG. 10, FIG. 10 is a structural schematic diagram of a rule control information acquisition device provided by the embodiments of the present application, the rule control information acquisition device 1000 comprises: an acquisition module 1001, configured to acquire environment information at a current time; a processing module 1002, configured to obtain rule control information of a vehicle corresponding to each of at least two time points based on the environment information at the current time, the at least two time points comprising the current time and at least one time point after the current time; wherein the rule control information corresponding to each of the at least two time points is obtained by a machine learning model, the rule control information corresponding to the current time is obtained by the machine learning model based on the environment information at the current time, the first time point is any one of the at least one time point after the current time, and the rule control information of the first time point is obtained by the machine learning model based on the rule control information corresponding to the last time point of the first time point.

[0174] Optionally, the rule control information of the first time point is obtained by the machine learning model based on predicted feature information of the environment at the first time point, and the predicted feature information of the environment at the first time point is obtained by the machine learning model based on the feature information of the environment at the last time point of the first time point and the rule control information corresponding to the last time point of the first time point.

[0175] Optionally, the predicted feature information of the environment at the first time point is generated by a first module in the machine learning model, the machine learning model further comprises a second module, an input of the second module comprises the environment information at the current time, feature information corresponding to the environment information at the current time is obtained by the second module, and the first module is configured to obtain the rule control information corresponding to each of the at least two time points based on the feature information.

[0176] Optionally, the first module comprises a first submodule corresponding to the current time and at least one second submodule corresponding to at least one time point after the current time one by one.

[0177] The first sub-module is configured to obtain the regulation and control information of the current time based on the feature information; and the second sub-module corresponding to the first time is configured to obtain the regulation and control information of the first time based on the regulation and control information of the previous time of the first time and the predicted feature information of the environment of the first time.

[0178] Optionally, the training phase of the first module is reinforcement learning.

[0179] Optionally, the feature information comprises a semantic image in a top-down BEV perspective, and the semantic image indicates a semantic category of at least one region in an image in a BEV perspective corresponding to the environment information of the current time.

[0180] Optionally, the feature information further comprises first feature information obtained by feature extraction based on the environment information of the current time.

[0181] Optionally, the feature information further comprises second feature information corresponding to the previous time of the first time, and the second feature information is obtained by the second module based on the first feature information and the regulation and control information corresponding to the previous time of the first time.

[0182] Optionally, the predicted feature information of the environment of the first time is obtained by the first module based on the feature information of the environment of the previous time of the first time, the regulation and control information corresponding to the previous time of the first time, and the second feature information.

[0183] Optionally, the first module is configured to generate predicted feature information of the environment of each of at least one time after the current time, and in the training phase of the first module, the predicted feature information of the environment generated by the first module is processed by feature processing to obtain predicted environment information, correct environment information corresponding to the predicted environment information is obtained by a simulator, and a loss function of the training phase of the first module indicates a similarity between the predicted environment information and the correct environment information.

[0184] Optionally, the first module is configured to generate predicted feature information of the environment of each of at least one time after the current time, and a loss function of the training phase of the first module indicates a similarity between the predicted feature information generated by the first module and expected feature information, the expected feature information is obtained based on correct environment information, and the correct environment information is obtained by a simulator.

[0185] Optionally, the regulation and control information of the vehicle corresponding to each of the at least two times comprises control information of the vehicle corresponding to each of the at least two times and / or a planned position of the vehicle corresponding to each of the at least two times; and / or, the output of the machine learning model further comprises perception information corresponding to the environment information of the current time, wherein the perception information comprises road information and / or obstacle information in the environment of the current time.

[0186] It should be noted that the information interaction between the modules / units in the regulation information acquisition device 1000, the execution process, and the like are based on the same concept as the various method embodiments corresponding to FIGS. 1 to 9 of the present application, and the specific content can be referred to the description in the foregoing method embodiments of the present application, which will not be described here.

[0187] Next, a device provided in an embodiment of the present application is introduced. Referring to FIG. 11, FIG. 11 is a structural schematic diagram of a device provided in an embodiment of the present application. Optionally, the device 1100 performs the functions of the first device and / or the execution device in the various method embodiments corresponding to FIGS. 1 to 9.

[0188] The device 1100 includes a memory 1102 and at least one processor 1101. Optionally, the processor 1101 implements the method in the foregoing embodiments by reading the instructions saved in the memory 1102, or the processor 1101 can also implement the method in the foregoing embodiments by reading the instructions saved in the internal memory. In the case where the processor 1101 implements the method in the foregoing embodiments by reading the instructions saved in the memory 1102, the memory 1102 saves the instructions for implementing the method provided in the foregoing embodiments of the present application.

[0189] Optionally, the at least one processor 1101 is one or more CPUs, or a single core CPU, or a multi-core CPU. The memory 1102 includes, but is not limited to, a random access memory (RAM), a read-only memory (ROM), a flash memory, or an optical memory, etc. The memory 1102 saves the instructions of an operating system. After the program instructions stored in the memory 1102 are read by the at least one processor 1101, the device 1100 performs the corresponding operations in the foregoing embodiments.

[0190] Optionally, the device 1100 further includes a network interface 1103. The network interface 1103 can be a wired interface or a wireless interface. The network interface 1103 is configured to perform the transceiving of data in the various method embodiments corresponding to FIGS. 1 to 7.

[0191] It should be understood that the network interface 1103 has the functions of receiving data and sending data. The function of “receiving data” and the function of “sending data” can be integrated in the same transceiving interface, or the function of “receiving data” and the function of “sending data” can be implemented in different interfaces, which is not limited here. In other words, the network interface 1103 can include one or more interfaces for implementing the function of “receiving data” and the function of “sending data”.

[0192] The device 1100 can perform other functions described in the foregoing method embodiments after the processor 1101 reads the program instructions in the memory 1102.

[0193] Optionally, the device 1100 further includes a bus 1104, and the processor 1101 and the memory 1102 are usually connected with each other through the bus 1104, and can be connected with each other in other manners.

[0194] The device 1100 provided by the embodiments of the present application is configured to execute the method of the first device and / or the method executed by the execution device in each of the method embodiments, and achieve the corresponding beneficial effects. The specific implementation modes of the device 1100 shown in FIG. 11 can all refer to the descriptions in the foregoing method embodiments, which will not be repeated here.

[0195] The embodiments of the present application also provide a vehicle. Please refer to FIG. 12, which is a structural schematic diagram of a vehicle provided by the embodiments of the present application. The vehicle 100 is configured to be in a fully or partially autonomous driving mode. For example, the vehicle 100 can control itself while being in the autonomous driving mode, and can determine the current state of the vehicle and its surrounding environment, determine the possible behavior of at least one other vehicle in the surrounding environment, and determine the confidence level corresponding to the possibility of the other vehicle executing the possible behavior through human operation, and control the vehicle 100 based on the determined information. When the vehicle 100 is in the autonomous driving mode, the vehicle 100 can also be configured to operate without human interaction.

[0196] The vehicle 100 can include various subsystems, such as a travel system 102, a sensor system 104, a control system 106, one or more peripheral devices 108, a power source 110, a computer system 112, and a user interface 116. Optionally, the vehicle 100 can include more or fewer subsystems, and each subsystem can include multiple components. In addition, each subsystem and component of the vehicle 100 can be interconnected by wire or wirelessly.

[0197] The travel system 102 can include components that provide powered movement for the vehicle 100. In one embodiment, the travel system 102 can include an engine 118, an energy source 119, a transmission 120, and wheels / tires 121.

[0198] The engine 118 can be a combustion engine, an electric motor, an air compression engine, or other types of engine combinations, such as a hybrid engine composed of a gasoline engine and an electric motor, a hybrid engine composed of a combustion engine and an air compression engine. The engine 118 converts an energy source 119 into mechanical energy. Examples of the energy source 119 include gasoline, diesel, other petroleum-based fuels, propane, other compressed gas-based fuels, ethanol, solar panels, batteries, and other sources of electricity. The energy source 119 can also provide energy for other systems of the vehicle 100. The transmission 120 can transmit mechanical power from the engine 118 to the wheels 121. The transmission 120 can include a gearbox, a differential, and a drive shaft. In one embodiment, the transmission 120 can also include other devices, such as a clutch. The drive shaft can include one or more shafts that can be coupled to one or more wheels 121.

[0199] The sensor system 104 can include several sensors that sense information about the environment surrounding the vehicle 100. For example, the sensor system 104 can include a positioning system 122 (which can be a global positioning GPS system, a Beidou system, or other positioning systems), an inertial measurement unit (IMU) 124, a radar 126, a laser rangefinder 128, and a camera 130. The sensor system 104 can also include sensors that monitor internal systems of the vehicle 100 (e.g., an in-vehicle air quality monitor, a fuel gauge, an oil temperature gauge, etc.). Sensing data from one or more of these sensors can be used to detect objects and their respective characteristics (location, shape, direction, speed, etc.). Such detection and identification are key functions for the safe operation of the autonomous vehicle 100.

[0200] The positioning system 122 can be used to estimate the geographic location of the vehicle 100. The IMU 124 is used to perceive changes in the position and orientation of the vehicle 100 based on inertial acceleration. In one embodiment, the IMU 124 can be a combination of an accelerometer and a gyroscope. The radar 126 can utilize radio signals to perceive objects within the surrounding environment of the vehicle 100, which can be manifested as a millimeter wave radar or a laser radar. In some embodiments, in addition to perceiving objects, the radar 126 can also be used to perceive the speed and / or direction of advance of the objects. The laser rangefinder 128 can utilize laser light to perceive objects in the environment in which the vehicle 100 is located. In some embodiments, the laser rangefinder 128 can include one or more laser sources, a laser scanner, and one or more detectors, as well as other system components. The camera 130 can be used to capture multiple images of the surrounding environment of the vehicle 100. The camera 130 can be a still camera or a video camera.

[0201] The control system 106 controls the operation of the vehicle 100 and its components. The control system 106 can include various components, including a steering system 132, a throttle 134, a braking unit 136, a computer vision system 140, a route control system 142, and an obstacle avoidance system 144.

[0202] The steering system 132 is operable to adjust the heading of the vehicle 100. In one embodiment, the steering system 132 can be a steering wheel system. The throttle 134 is used to control the speed of the engine 118 and, in turn, the speed of the vehicle 100. The braking unit 136 is used to control the deceleration of the vehicle 100. The braking unit 136 can use friction to slow the wheels 121. In other embodiments, the braking unit 136 can convert the kinetic energy of the wheels 121 into electrical current. The braking unit 136 can also take other forms to slow the wheels 121 and, in turn, control the speed of the vehicle 100. The computer vision system 140 is operable to process and analyze images captured by the camera 130 to identify objects and / or features in the environment surrounding the vehicle 100. The objects and / or features can include traffic signals, road boundaries, and obstacles. The computer vision system 140 can use object recognition algorithms, Structure from Motion (SFM) algorithms, video tracking, and other computer vision techniques. In some embodiments, the computer vision system 140 can be used to map the environment, track objects, estimate the speed of objects, and the like. The route control system 142 is used to determine the route and speed of travel of the vehicle 100. In some embodiments, the route control system 142 can include a lateral planning module 1421 and a longitudinal planning module 1422 that are used to determine the route and speed of travel of the vehicle 100 in conjunction with data from the obstacle avoidance system 144, the GPS 122, and one or more predetermined maps, respectively. The obstacle avoidance system 144 is used to identify, evaluate, and avoid or otherwise navigate around obstacles in the environment of the vehicle 100, which can be actual obstacles and virtual moving obstacles that can collide with the vehicle 100. In one instance, the control system 106 can include additional components in addition to those shown and described, or some of the components shown can be reduced.

[0203] The vehicle 100 interacts with external sensors, other vehicles, other computer systems, or users through the peripherals 108. The peripherals 108 can include a wireless communication system 146, an on-board computer 148, a microphone 150, and / or a speaker 152. In some embodiments, the peripherals 108 provide a means for a user of the vehicle 100 to interact with the user interface 116. For example, the on-board computer 148 can provide information to a user of the vehicle 100. The user interface 116 can also operate the on-board computer 148 to receive input from the user. The on-board computer 148 can be operated through a touch screen. In other cases, the peripherals 108 can provide a means for the vehicle 100 to communicate with other devices located within the vehicle. For example, the microphone 150 can receive audio (e.g., voice commands or other audio input) from a user of the vehicle 100. Similarly, the speaker 152 can output audio to a user of the vehicle 100. The wireless communication system 146 can wirelessly communicate with one or more devices, either directly or via a communication network. For example, the wireless communication system 146 can use 3G cellular communication, such as CDMA, EVDO, GSM / GPRS, or 4G cellular communication, such as LTE. Or 5G cellular communication. The wireless communication system 146 can utilize wireless local area network (WLAN) communication. In some embodiments, the wireless communication system 146 can utilize an infrared link, Bluetooth, or ZigBee to communicate directly with devices. Other wireless protocols, such as various vehicle communication systems, for example, the wireless communication system 146 can include one or more dedicated short range communications (DSRC) devices, which can include public and / or private data communication between vehicles and / or roadside stations.

[0204] The power supply 110 can provide power to various components of the vehicle 100. In one embodiment, the power supply 110 can be a rechargeable lithium-ion or lead-acid battery. One or more battery packs of such a battery can be configured as the power supply to provide power to various components of the vehicle 100. In some embodiments, the power supply 110 and the energy source 119 can be implemented together, such as in some all-electric vehicles.

[0205] Some or all of the functionality of vehicle 100 is controlled by computer system 112. Computer system 112 can include at least one processor 113 that executes instructions 115 stored in a non-transitory computer readable medium such as memory 114. Computer system 112 can also be a plurality of computing devices that control individual components or subsystems of vehicle 100 in a distributed manner. Processor 113 can be any conventional processor such as a commercially available central processing unit (CPU). Alternatively, processor 113 can be a dedicated device such as an application specific integrated circuit (ASIC) or other hardware-based processor. Although FIG. 12 functionally illustrates the processor, memory, and other components of computer system 112 within the same block, it will be understood by those skilled in the art that the processor, or memory, can actually comprise a plurality of processors, or memories, that can or can not be stored within the same physical housing. For example, memory 114 can be a hard drive or other storage media located in a housing different from that of computer system 112. Accordingly, references to processor 113 or memory 114 will be understood to include references to a collection of processors or memories that can or can not operate in parallel. Rather than using a single processor to perform the steps described herein, some components such as the steering assembly and the deceleration assembly can each have their own processor that only performs calculations related to the functionality specific to that component.

[0206] In various aspects described herein, processor 113 can be located remotely from vehicle 100 and in wireless communication with vehicle 100. In other aspects, some of the processes described herein are performed on processor 113 disposed within vehicle 100 while others are performed by a remote processor 113, including taking the necessary steps to perform a single maneuver.

[0207] In some embodiments, the memory 114 can include instructions 115 (e.g., program logic) that can be executed by the processor 113 to perform various functions of the vehicle 100, including those described above. The memory 114 can also include additional instructions, including instructions to send data to, receive data from, interact with, and / or control one or more of the travel system 102, the sensor system 104, the control system 106, and the peripherals 108. In addition to the instructions 115, the memory 114 can also store data, such as road maps, route information, the vehicle's location, direction, speed, and other such vehicle data, and other information. Such information can be used by the vehicle 100 and the computer system 112 during operation of the vehicle 100 in autonomous, semi-autonomous, and / or manual modes. The user interface 116 is used to provide information to or receive information from a user of the vehicle 100. Optionally, the user interface 116 can include one or more input / output devices within the set of peripherals 108, such as the wireless communication system 146, the on-board computer 148, the microphone 150, and the speaker 152.

[0208] The computer system 112 can control the functions of the vehicle 100 based on inputs received from various subsystems (e.g., the travel system 102, the sensor system 104, and the control system 106), as well as from the user interface 116. For example, the computer system 112 can utilize inputs from the control system 106 in order to control the steering system 132 to avoid obstacles detected by the sensor system 104 and the obstacle avoidance system 144. In some embodiments, the computer system 112 can be operable to provide control over many aspects of the vehicle 100 and its subsystems.

[0209] Optionally, one or more of the above-described components can be installed separately from or associated with the vehicle 100. For example, the memory 114 can exist partially or entirely separately from the vehicle 100. The above-described components can be communicatively coupled together in a wired and / or wireless manner.

[0210] Optionally, the above-described components are just one example, and in actual applications, components in each of the above-described modules can be added or deleted according to actual needs, and FIG. 12 should not be understood as a limitation on the embodiments of the present application. A vehicle that travels on a road, such as the vehicle 100 above, can identify objects within its surrounding environment to determine an adjustment to a current speed. The objects can be other vehicles, traffic control devices, or other types of objects. In some examples, each identified object can be considered independently, and based on respective characteristics of the object, such as its current speed, acceleration, spacing from the vehicle, etc., can be used to determine a speed at which the vehicle is to adjust.

[0211] Optionally, the vehicle 100 or a computing device associated with the vehicle 100, such as the computer system 112, the computer vision system 140, the memory 114 of FIG. 12, can predict the behavior of the identified object based on the characteristics of the identified object and the state of the surrounding environment (e.g., traffic, rain, ice on the road, etc.). Optionally, each identified object is dependent on the behavior of the other identified objects, so all of the identified objects can also be considered together to predict the behavior of a single identified object. The vehicle 100 can adjust its speed based on the predicted behavior of the identified object. In other words, the vehicle 100 can determine what steady state the vehicle will need to adjust to (e.g., accelerate, decelerate, or stop) based on the predicted behavior of the object. Other factors can also be considered in determining the speed of the vehicle 100 during this process, such as the lateral position of the vehicle 100 in the road, the curvature of the road, the proximity of static and dynamic objects, etc. In addition to providing instructions to adjust the speed of the vehicle, the computing device can also provide instructions to modify the steering angle of the vehicle 100 to cause the vehicle 100 to follow a given trajectory and / or maintain a safe lateral and longitudinal distance from objects in the vicinity of the vehicle 100 (e.g., a car in the adjacent lane on the road).

[0212] In the embodiments of the present application, the processor 113 in the vehicle 100 is configured to execute the first device and / or the method executed by the execution device in the embodiments corresponding to FIG. 1 to FIG. 9. It should be noted that the specific manner in which the processor 113 executes the foregoing steps is based on the same concept as the method embodiments corresponding to FIG. 1 to FIG. 9 in the present application, and the resulting technical effects are the same as the method embodiments corresponding to FIG. 1 to FIG. 9 in the present application. For specific content, please refer to the description of the method embodiments in the foregoing description of the present application, which will not be described here.

[0213] In the embodiments of the present application, a computer readable storage medium is also provided, which stores a program, and when the program runs on a computer, the computer executes the steps performed by the first device in the method described in the foregoing embodiments of FIG. 1 to FIG. 9.

[0214] In the embodiments of the present application, a computer program product is also provided, which includes a program, and when the program runs on a computer, the computer executes the steps performed by the first device in the method described in the foregoing embodiments of FIG. 1 to FIG. 9.

[0215] In the embodiments of the present application, a circuit system is also provided, which includes a processing circuit configured to execute the steps performed by the first device in the method described in the foregoing embodiments of FIG. 1 to FIG. 9.

[0216] The first device, the execution device or the acquisition apparatus of the regulation information provided in the embodiments of the present application can specifically be a chip. The chip includes a processing unit, for example, a processor. Optionally, the chip further includes a communication unit, for example, an input / output interface, a pin or a circuit and the like. The processing unit can execute computer execution instructions stored in a storage unit, so that the chip executes the method described in the embodiments shown in FIG. 1 to FIG. 9. Optionally, the storage unit is a storage unit in the chip, such as a register, a cache and the like. The storage unit can also be a storage unit outside the chip in the wireless access device, such as a read-only memory (ROM) or other types of static storage devices that can store static information and instructions, a random access memory (RAM) and the like.

[0217] The processor mentioned in any of the above can be a general central processing unit, a microprocessor, an ASIC or one or more integrated circuits for controlling execution of the programs of the method of the first aspect.

[0218] In addition, it should be noted that the apparatus embodiments described above are merely illustrative, and the units described as separate components can or can not be physically separated, and the components displayed as units can or can not be physical units, that is, they can be located in one place, or can be distributed on multiple network units. According to actual needs, part or all of the modules can be selected to achieve the purpose of the embodiments. In addition, the connection relationship between the modules in the apparatus embodiments provided in the present application indicates that there is a communication connection between them, which can be implemented as one or more communication buses or signal lines.

[0219] Those skilled in the art can clearly understand that the application can be implemented by means of software plus necessary universal hardware, and of course can also be implemented by means of dedicated hardware including special integrated circuit, special CLU, special memory, special component, etc. Generally, any function completed by computer program can be easily implemented by corresponding hardware, and the specific hardware structure for implementing the same function can also be various, such as analog circuit, digital circuit or special circuit, etc. However, for the application, the software program implementation is a better embodiment. Based on such understanding, the technical solution of the application or the part of the application which makes contribution to the prior art can be embodied in the form of software product, which is stored in readable storage medium, such as floppy disk, U disk, mobile hard disk, ROM, RAM, magnetic disk or optical disk, etc., and includes a plurality of instructions for making a computer device (which can be personal computer, server or network device, etc.) execute the method described in various embodiments of the application.

[0220] In the above embodiments, the implementation can be achieved by software, hardware, firmware or any combination thereof, entirely or partially. When implemented by software, the implementation can be in the form of computer program product entirely or partially.

[0221] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on the computer, the flow or function described in the embodiments of the application is entirely or partially generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transferred from one computer-readable storage medium to another, for example, the computer instructions can be transferred from one website, computer, server or data center to another website, computer, server or data center through wired (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.) mode. The computer-readable storage medium can be any available medium that can be stored by the computer or a data storage device such as server, data center, etc. integrated with one or more available media. The available medium can be magnetic medium (such as floppy disk, hard disk, magnetic tape), optical medium (such as DVD) or semiconductor medium (such as solid state disk (SSD)) etc.

Claims

1. A method of acquiring regulatory information, characterized by, The method comprises: obtaining environment information at a current time; based on the environment information at the current time, obtaining the regulation and control information of the vehicle corresponding to each of at least two time points, the at least two time points including the current time and at least one time point after the current time; wherein the regulation and control information corresponding to each of the at least two time points is obtained by a machine learning model, the regulation and control information corresponding to the current time is obtained by the machine learning model based on the environment information at the current time, the first time point is any one of the at least one time point after the current time, and the regulation and control information of the first time point is obtained by the machine learning model based on the regulation and control information corresponding to the last time point of the first time point.

2. The method of claim 1, wherein, The regulation and control information of the first time point is obtained by the machine learning model based on the predicted feature information of the environment of the first time point, and the predicted feature information of the environment of the first time point is obtained by the machine learning model based on the feature information of the environment of the last time point of the first time point and the regulation and control information corresponding to the last time point of the first time point.

3. The method of claim 2, wherein, The predicted feature information of the environment of the first time point is generated by a first module in the machine learning model, the machine learning model further comprises a second module, the input of the second module comprises the environment information at the current time, the feature information corresponding to the environment information at the current time is obtained through the second module, and the first module is used to obtain the regulation and control information corresponding to each of the at least two time points based on the feature information.

4. The method of claim 3, wherein, The first module comprises a first submodule corresponding to the current time and at least one second submodule corresponding one-to-one to at least one time point after the current time; wherein the first submodule is used to obtain the regulation and control information of the current time based on the feature information, and the second submodule corresponding to the first time point is used to obtain the regulation and control information of the first time point based on the regulation and control information of the last time point of the first time point and the predicted feature information of the environment of the first time point.

5. The method according to claim 3 or 4, characterized in that, The training stage of the first module is reinforcement learning.

6. The method according to claim 3 or 4, characterized in that, The feature information comprises a semantic image in a top-down BEV perspective, and the semantic image indicates a semantic category of at least one region in an image in a BEV perspective corresponding to the environment information at the current time.

7. The method of claim 6, wherein, The feature information further comprises first feature information obtained by feature extraction based on the environment information at the current time.

8. The method of claim 6, wherein, The feature information further comprises second feature information corresponding to the last time point of the first time point, and the second feature information is obtained by the second module based on the first feature information and the regulation and control information corresponding to the last time point of the first time point.

9. The method of claim 8, wherein, The predicted feature information of the environment of the first time point is obtained by the first module based on the feature information of the environment of the last time point of the first time point, the regulation and control information corresponding to the last time point of the first time point, and the second feature information.

10. The method of claim 3 or 4, wherein, The first module is configured to generate predicted feature information of the environment at each of the at least one time instant after the current time instant, and in a training phase of the first module, the predicted feature information of the environment generated by the first module is processed to obtain predicted environment information, correct environment information corresponding to the predicted environment information is obtained by using a simulator, and a loss function of the training phase of the first module indicates a similarity between the predicted environment information and the correct environment information.

11. The method of claim 3 or 4, wherein, The first module is configured to generate predicted feature information of the environment at each of the at least one time instant after the current time instant, and a loss function of a training phase of the first module indicates a similarity between the predicted feature information generated by the first module and expected feature information, the expected feature information being obtained based on correct environment information, and the correct environment information being obtained by using a simulator.

12. The method according to any one of claims 1 to 6, characterized in that, The control information of the vehicle corresponding to each of the at least two time instants includes control information of the vehicle corresponding to each of the at least two time instants and / or a planned position of the vehicle corresponding to each of the at least two time instants; and / or the output of the machine learning model further includes perception information corresponding to the environment information of the current time instant, wherein the perception information includes road information and / or obstacle information in the environment of the current time instant.

13. A device for acquiring information for regulation, characterized by The apparatus comprises: an acquisition module configured to acquire environment information of a current time instant; a processing module configured to obtain control information of a vehicle corresponding to each of at least two time instants based on the environment information of the current time instant, the at least two time instants including the current time instant and at least one time instant after the current time instant; wherein the control information corresponding to each of the at least two time instants is obtained by using a machine learning model, the control information corresponding to the current time instant is obtained by the machine learning model based on the environment information of the current time instant, a first time instant is any one of the at least one time instant after the current time instant, and the control information of the first time instant is obtained by the machine learning model based on control information corresponding to a previous time instant of the first time instant.

14. The apparatus of claim 13, wherein, The control information of the first time instant is obtained by the machine learning model based on predicted feature information of the environment of the first time instant, and the predicted feature information of the environment of the first time instant is obtained by the machine learning model based on feature information of the environment of a previous time instant of the first time instant and the control information corresponding to the previous time instant of the first time instant.

15. The apparatus of claim 14, wherein, The predicted feature information of the environment of the first time instant is generated by a first module in the machine learning model, the machine learning model further comprises a second module, an input of the second module includes the environment information of the current time instant, feature information corresponding to the environment information of the current time instant is obtained by the second module, and the first module is configured to obtain the control information corresponding to each of the at least two time instants based on the feature information.

16. The apparatus of claim 15, wherein, The first module includes a first submodule corresponding to the current time instant and at least one second submodule corresponding to each of the at least one time instant after the current time instant. The first sub-module is configured to obtain the regulation and control information of the current time based on the feature information; and the second sub-module corresponding to the first time is configured to obtain the regulation and control information of the first time based on the regulation and control information of a time preceding the first time and the predicted feature information of the environment of the first time.

17. The apparatus of claim 15 or 16, wherein, The training phase of the first module is reinforcement learning.

18. The apparatus of claim 15 or 16, wherein, The feature information includes a semantic image in a top-down BEV perspective, and the semantic image indicates a semantic category of at least one region in an image in a BEV perspective corresponding to the environment information of the current time.

19. The apparatus of claim 18, wherein, The feature information further includes first feature information obtained by feature extraction based on the environment information of the current time.

20. The apparatus of claim 18, wherein, The feature information further includes second feature information corresponding to a time preceding the first time, which is obtained by the second module based on the first feature information and the regulation and control information corresponding to the time preceding the first time.

21. The apparatus of claim 20, wherein, The predicted feature information of the environment of the first time is obtained by the first module based on the feature information of the environment of the time preceding the first time, the regulation and control information corresponding to the time preceding the first time, and the second feature information.

22. The apparatus of claim 15 or 16, wherein, The first module is configured to generate predicted feature information of the environment of each of at least one time succeeding the current time, and the loss function of the training phase of the first module indicates a similarity between the predicted feature information generated by the first module and correct environment information obtained by a simulator after feature processing of the predicted feature information generated by the first module.

23. The apparatus of claim 15 or 16, wherein, The first module is configured to generate predicted feature information of the environment of each of at least one time succeeding the current time, and the loss function of the training phase of the first module indicates a similarity between the predicted feature information generated by the first module and expected feature information obtained based on correct environment information obtained by a simulator.

24. The apparatus of claim 15 or 16, wherein, The regulation and control information of the vehicle corresponding to each of the at least two times includes control information of the vehicle corresponding to each of the at least two times and / or a planned position of the vehicle corresponding to each of the at least two times; and / or, the output of the machine learning model further includes perception information corresponding to the environment information of the current time, wherein the perception information includes road information and / or obstacle information in the environment of the current time.

25. An apparatus comprising: The processor and the memory are coupled, and the memory stores program instructions, which, when executed by the processor, implement the method of any one of claims 1 to 12.

26. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a program, which, when executed on a computer, causes the computer to perform the method of any one of claims 1 to 12.

27. A computer program product, characterised in that, The computer program product comprises a program which, when run on a computer, causes the computer to perform the method of any one of claims 1 to 12.

28. A chip, characterized by The chip comprises a processor for performing the steps of the method of any one of claims 1 to 12.

Citation Information

Patent Citations

  • Autonomous vehicle control method and system based on deep learning

    CN112721950A

  • Automatic driving model and method based on time sequence recursive autoregression reasoning and vehicle

    CN117539260A

  • Reinforcement learning method, device and equipment

    CN117709481A

  • System and method for distributed aware target prediction for modular autonomous vehicle control

    CN117789502A

  • Machine learning model training method and related device

    US20230237333A1