Method for generating planning and control information and related device

By combining a two-stage machine learning model and a large language model, the problems of latency and resource consumption in the generation of regulatory control information in intelligent driving are solved, achieving more efficient and higher-quality generation of regulatory control information and improving the performance of intelligent driving.

WO2026091775A1PCT designated stage Publication Date: 2026-05-07HUAWEI TECH CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
HUAWEI TECH CO LTD
Filing Date
2025-08-13
Publication Date
2026-05-07

AI Technical Summary

Technical Problem

Existing technologies suffer from high latency and high computational resource consumption when generating vehicle control information, making it difficult to meet the demands for improved intelligent driving performance.

Method used

A two-stage machine learning model is adopted. First, a first machine learning model with fewer parameters is used to extract environmental features. Then, a second machine learning model with more parameters is used to further process the features to generate regulatory information. A large language model (LLM) is combined to supplement the three-dimensional understanding ability and mine more information.

Benefits of technology

This reduces the latency and computational resource consumption in the process of generating regulatory information, while improving the quality and performance of the information, ensuring the stability and flexibility of the intelligent driving system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025114483_07052026_PF_FP_ABST
    Figure CN2025114483_07052026_PF_FP_ABST
Patent Text Reader

Abstract

Provided in the embodiments of the present application are a method for generating planning and control information and a related device, which are applied to the field of intelligent driving. The method comprises: performing feature extraction on environment information comprising an image and / or point cloud data of a traffic environment by means of a first machine learning model, so as to obtain a first feature; obtaining first information on the basis of the first feature by means of a second machine learning model, wherein since the second machine learning model has a larger number of parameters, the first information may carry more effective information mined from the environment information; and further generating planning and control information of a vehicle on the basis of the first feature and the first information by means of the first machine learning model. In this way, the first information carrying more effective information is additionally introduced, and combining the first feature and the first information helps to obtain planning and control information with better performance. Compared with directly inputting the environment information into the second machine learning model, the present application reuses the feature extraction capability of the first machine learning mode, which helps to reduce consumption of computer resources.
Need to check novelty before this filing date? Find Prior Art

Description

A method for generating regulatory information and related equipment

[0001] This application claims priority to Chinese Patent Application No. 202411549817.4, filed with the State Intellectual Property Office of China on October 31, 2024, entitled "A Method for Generating Regulatory Information and Related Equipment", the entire contents of which are incorporated herein by reference. Technical Field

[0002] This application relates to intelligent driving technology, and more particularly to a method for generating regulatory information and related equipment. Background Technology

[0003] Artificial intelligence (AI) is the theory, methods, technology, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. In other words, AI is a branch of computer science that attempts to understand the essence of intelligence and produce a new kind of intelligent machine that can react in a way similar to human intelligence. AI studies the design principles and implementation methods of various intelligent machines, enabling them to possess the functions of perception, reasoning, and decision-making.

[0004] Intelligent driving is a common application area of ​​artificial intelligence technology. For example, environmental information can be acquired through sensors and input into a machine learning model to obtain vehicle control information generated by the model. As the performance requirements for intelligent driving become increasingly demanding, a more efficient solution for generating control information is urgently needed. Summary of the Invention

[0005] This application provides a method and related equipment for generating regulatory information, which can obtain higher-performance regulatory information while minimizing the latency of the entire process of generating vehicle regulatory information.

[0006] This application provides the following technical solution:

[0007] Firstly, this application provides a method for generating traffic control information, which can be applied to the field of intelligent driving. In this method, an execution device can input environmental information into a first machine learning model, and extract features from the aforementioned environmental information through the first machine learning model to obtain a first feature. The environmental information includes at least one image and / or point cloud data corresponding to the traffic environment, and the first feature may include features of at least one image and / or features of point cloud data corresponding to the traffic environment. Based on the first feature, the execution device obtains first information through a second machine learning model, wherein the number of parameters of the second machine learning model is greater than the number of parameters of the first machine learning model. Then, the execution device can generate vehicle traffic control information based on the first feature and the first information through the first machine learning model.

[0008] In this implementation, not only is a first machine learning model used to generate traffic control information used, but a second machine learning model is also introduced. The second machine learning model has more parameters than the first machine learning model. In the process of generating vehicle traffic control information, after the first machine learning model extracts features from the environmental information to obtain the first feature, the second machine learning model with a larger number of parameters obtains the first information based on the first feature. Since the second machine learning model has a larger number of parameters, the first information can carry more effective information mined from the environmental information. Then, the first machine learning model generates vehicle traffic control information based on the first feature and the first information. Compared with generating vehicle traffic control information based solely on the first feature, the introduction of the first information carrying more effective information in the process of generating vehicle traffic control information is beneficial to obtaining better-performing traffic control information and improving the driving experience. In addition, compared with directly inputting environmental information into the second machine learning model, this application uses the second machine learning model to process based on the first feature, that is, it reuses the feature extraction capability of the first machine learning model, which helps to reduce the computer resources consumed in the entire process of generating vehicle traffic control information and also helps to reduce the latency of the entire process of generating vehicle traffic control information.

[0009] In one possible implementation, the second machine learning model is based on a large language model. The large language model (LLM) can be understood as a large model with language expression capabilities. In this application, the large model can be understood as a machine learning model with a huge number of parameters. For example, the large model can be a machine learning model with a parameter scale of more than 100 million parameters. Optionally, the number of parameters of the large model can be greater than or equal to one billion. "Having language expression capabilities" can be understood as the LLM being able to output information indicating the content of the text. The first information includes the features of each first token in at least one first token generated by the LLM and / or at least one first token generated by the LLM. The features of each first token are generated by a neural network layer in the LLM (hereinafter referred to as the "target neural network layer" for convenience). The features of each first token can be understood as follows: in the process of generating the first token through the second machine learning model, the features of each first token are first generated by the second machine learning model, and then each first token is obtained based on the features of each first token by the second machine learning model. Then, the execution device can obtain the features of each first token in the process of generating at least one first token through the second machine learning model. For example, the target neural network layer can include the Mth neural network layer from the end of the second machine learning model, where M is an integer greater than 1. Alternatively, a second machine learning model is used to update the first feature, and the first information includes the updated first feature.

[0010] This implementation provides several possible ways to implement the second machine learning model and the first information, improving the flexibility of the solution and facilitating the expansion of the application scenarios it can be adapted to. When the second machine learning model is based on LLM, since LLM often cannot extract 3D features from the original environmental information, the first features obtained by feature extraction from traffic environment images and / or point cloud data using the first machine learning model carry 3D information. Processing the first features using the second machine learning model helps to supplement the 3D understanding capability of LLM, thereby improving the quality of the first information obtained by the second machine learning model and further improving the quality of the final regulatory information. When the second machine learning model is used to update the first features, it is beneficial to use the second machine learning model with a larger number of parameters to mine richer and more effective information in the environmental information, thus obtaining better-performing regulatory information based on the first features and the updated first features.

[0011] In one possible implementation, the second machine learning model is an LLM, and at least one first token includes the first N tokens generated by the second machine learning model using an autoregressive approach, where N is an integer greater than or equal to 1, and the value of N is preset; for example, the value of N can be 2, 3, 4, 5, 6, 7, 8 or other values, etc.

[0012] For example, if the first information includes features of each of at least one first token, then the execution device obtains the first information based on the first features using a second machine learning model. This can include: the execution device generating N first tokens stepwise using an autoregressive approach with the second machine learning model based on the first features, where N is a preset value; and acquiring features of each first token during the generation of the N first tokens. If the first information includes at least one first token, then the execution device obtains the first information based on the first features using a second machine learning model. This can include: the execution device generating N first tokens stepwise using an autoregressive approach with the second machine learning model based on the first features.

[0013] In this implementation, since the second machine learning model generates at least one first token using an autoregressive approach, meaning that the second machine learning model generates one of the first tokens in each inference operation, the number of inference operations performed by the second machine learning model increases as the number of tokens included in the at least one first token increases. This leads to increased computer resource consumption and higher latency. In this application, the number of at least one first tokens is pre-set to N, thereby limiting the number of inference operations performed by the second machine learning model to N. This improves the controllability of the computer resources consumed and the latency caused by the process of generating at least one first token through the second machine learning model, and also improves the controllability of the computer resources consumed and the latency caused by the process of obtaining first information through the second machine learning model. This helps to reduce the computer resources consumed and the latency caused by the process of obtaining first information through the second machine learning model.

[0014] In one possible implementation, the first information includes features of at least one first token. The execution device generates vehicle control information based on the first features and the first information using a first machine learning model. This can include: the execution device processing the features of at least one first token through a first module to obtain second features; and generating vehicle control information based on the first and second features using the first machine learning model. For example, processing the features of at least one first token through the first module can be understood as extracting more information from the features of at least one first token to obtain the second features; in other words, the second features can carry richer information compared to the features of at least one first token.

[0015] In this implementation, after obtaining the features of at least one first token, the features of at least one first token can be processed by the first module to obtain the second features. That is, the first module is added between the first machine learning model and the second machine learning model. The process of processing the features of at least one first token by the first module is conducive to obtaining richer information from at least one first token, which in turn helps the first machine learning model to generate better regulatory information based on the first and second features.

[0016] In one possible implementation, the first module includes at least one sub-module corresponding one-to-one with at least one task. Optionally, each sub-module processes a portion of the features of the at least one first token to obtain a sub-feature, and the second feature may include all the sub-features generated by the at least one sub-module. For example, any one of the at least one tasks is referred to as the first task, and the sub-module corresponding to the first task is referred to as the first sub-module. The sub-feature obtained by updating a portion of the features of the at least one first token through the first sub-module can be understood as a feature used to execute the first task. In other words, the purpose of updating a portion of the features of the at least one first token through the first sub-module includes obtaining the information needed to execute the first task based on the first feature.

[0017] At least one of the following tasks includes: target detection of dynamic and / or static obstacles in a traffic environment; 3D modeling of the traffic environment; text recognition in the traffic environment; determination of spatial relationships between different objects in the traffic environment; intention judgment and / or behavior prediction of obstacles in the traffic environment; determination of strategies to be used when interacting with obstacles in the traffic environment; identification of risky objects in the traffic environment; determination of vehicle driving strategies; determination of the speed limit of vehicles when driving in the traffic environment; path planning for vehicles; or determination of vehicle control information, etc.

[0018] Optionally, during the joint training phase of the first and second machine learning models, each submodule can be connected to a task head. Each task head is used to generate prediction information corresponding to one of the at least one tasks mentioned above. Each task head can be understood as a feature processing module. For example, each task head can be a fully connected neural network layer, a multilayer perceptron, a support vector machine, a convolutional neural network layer, or other representations. A third loss function can be used during the joint training phase of the first and second machine learning models. The third loss function indicates the similarity between the prediction information generated by each task head and the expected information. The goal of training using the third loss function includes improving the similarity between the prediction information generated by each task head and the expected information.

[0019] In this implementation, the first module includes at least one sub-module corresponding to at least one task. Each sub-module can obtain information corresponding to one of the at least one tasks from the features of at least one first token. That is, each sub-module can obtain the information required to execute a certain task. This is beneficial to obtain the rich information required to generate vehicle control information in a targeted manner through at least one sub-module, which in turn helps the first machine learning model to generate better control information.

[0020] In one possible implementation, the first N tokens do not correspond to the text; in other words, the first N tokens cannot directly indicate text with semantic meaning. These first N tokens can be special tokens that do not correspond to the text. During the training phase of the second machine learning model—for example, in the joint training phase of the first and second machine learning models, or in the fine-tuning phase of the independent training phase of the second machine learning model—after the training device gradually generates the first N tokens using an autoregressive approach through the second machine learning model, it can further generate at least one second token corresponding to the predicted text using the same autoregressive approach. In other words, at least one second token can indicate the predicted text with semantic meaning. The training phase of the second machine learning model uses a loss function (hereinafter referred to as the "first loss function" for ease of distinction). The first loss function indicates the similarity between at least one second token and at least one third token, where the at least one third token includes tokens corresponding to the expected text.

[0021] Optionally, the first loss function also indicates the similarity between the top N tokens and the N special tokens, and the goal of training using the first loss function also includes improving the similarity between the top N tokens and the N special tokens.

[0022] In this implementation, since the first N tokens do not correspond to text, each of the first N tokens does not need to uniquely indicate a specific word. This allows the first N tokens to carry richer information. Furthermore, during the training phase of the second machine learning model, after generating the first N special tokens that do not correspond to text, a second token corresponding to the predicted text is also generated. This not only preserves the language expressive power of the LLM, but also, since the first N tokens and the second token are generated by the same LLM when performing a certain task, the predicted text pointed to by the second token can be used to interpret the meaning of the first N tokens. During the training phase of the second machine learning model, the loss function can be used to guide the second machine learning model to generate second tokens that can indicate the expected text, thereby guiding the first N tokens to carry the meaning of the expected text, which is beneficial to further improve the quality of the first N tokens.

[0023] In one possible implementation, the execution device obtains first information based on a first feature using a second machine learning model, including: in a first case, the execution device obtains first information based on the first feature using a second machine learning model. The method further includes: in a second case, generating vehicle control information based on the first feature and preset information using the first machine learning model; exemplarily, the preset information can be pre-stored information, for example, the preset information includes values ​​all of 0; or, the preset information can be information obtained based on preset rules, for example, the preset rules can be sequential filling with 0s and 1s, etc.

[0024] For example, the first case and the second case can be different. The first case can be understood as the case where the second machine learning model obtains the first information and works, while the second case can be understood as the case where the second machine learning model does not obtain the first information and does not work. For example, the second case can be the case where the second machine learning model does not work. In other words, the first case can be understood as the case where the second machine learning model is called, while the second case can be understood as the case where the second machine learning model is not called. Alternatively, the second case can also be the case where, although the second machine learning model is working, the first information has not yet been generated because the second machine learning model generates the first information slowly, or other cases where the first information has not been obtained through the second machine learning model.

[0025] In this implementation, in the first case, the first information can be generated by the second machine learning model, and then the vehicle's control information can be generated by the first machine learning model based on the first feature and the first information; in the second case, the vehicle's control information can be generated by the first machine learning model based on the first feature and preset information. This decoupling between the first machine learning model and the second machine learning model is achieved, which helps to ensure that the vehicle's control information can be generated smoothly under various conditions, thereby improving the stability of the intelligent driving system.

[0026] In one possible implementation, the frequency of generating vehicle control information through the first machine learning model is higher than the frequency of obtaining the first information through the second machine learning model. For example, the frequency at which the execution device generates vehicle control information through the first machine learning model is Q times the frequency at which the first information is obtained through the second machine learning model, where Q is greater than 1, for example, Q can be 4, 5, 6, 7 or other values.

[0027] In this implementation, since the number of parameters in the second machine learning model is greater than that in the first machine learning model, the computer resources consumed when generating the first information through the second machine learning model are often greater, and the latency of generating the first information through the second machine learning model is also higher. Since the vehicle's intelligent driving system generates vehicle control information multiple times at a high frequency during vehicle operation, setting the frequency of generating vehicle control information through the first machine learning model to be higher than the frequency of obtaining the first information through the second machine learning model is beneficial to reduce the computer resources consumed in the overall process of generating vehicle control information multiple times while improving the performance of generating vehicle control information multiple times. It is also beneficial to reduce the latency of the overall process of generating vehicle control information multiple times.

[0028] In one possible implementation, the execution device obtains first information based on a first feature through a second machine learning model. This includes: the execution device inputting the first feature into a second module, which compresses the first feature to obtain a third feature, where the data size of the third feature is smaller than that of the first feature; then, the execution device inputs the third feature into the second machine learning model to obtain the first information. In this implementation, by compressing the first feature before inputting the compressed third feature into the second machine learning model, the amount of data that the second machine learning model needs to process is reduced. This helps reduce the computer resources consumed in generating the first information through the second machine learning model and also reduces the latency associated with this process.

[0029] Optionally, the function of the second module includes converting the first feature in feature form into the third feature in token form. For example, the feature form can be understood as a matrix form, and the token form can be understood as a vector form.

[0030] Optionally, the function of the second module includes mapping the feature values ​​in the first feature to obtain the feature values ​​in the third feature. For example, the first feature can be continuous feature information; in other words, the first value space of the feature values ​​in the first feature can be continuous. The first value space refers to the set of all possible values ​​that the feature values ​​in the first feature can take. The number of values ​​in the first value space is not finite; in other words, the number of values ​​in the first value space can be infinite. For example, the second value space can be any real value within a certain value range, or the second value space can be any real value. The third feature can be discrete feature information; in other words, the second value space of the feature values ​​in the first feature can be discrete. The second value space refers to the set of all possible values ​​that the feature values ​​in the third feature can take. The number of values ​​in the second value space is finite, and any feature value in the first feature is contained within a finite set of values ​​(i.e., the first value interval).

[0031] In one possible implementation, the vehicle's control information includes at least one of the following: the vehicle's planned location, the vehicle's driving strategy, or the vehicle's control information. This implementation provides multiple possible scenarios for the vehicle's control information, which improves the flexibility of the solution and expands the application scenarios that this application can adapt to.

[0032] Secondly, this application provides a device for generating traffic control information, which can be used in the field of artificial intelligence. The device includes: an input module for inputting environmental information into a first machine learning model, and extracting features from the environmental information through the first machine learning model to obtain a first feature, wherein the environmental information includes images and / or point cloud data corresponding to the traffic environment; a processing module for obtaining first information based on the first feature through a second machine learning model, wherein the number of parameters of the second machine learning model is greater than the number of parameters of the first machine learning model; and a generation module for generating traffic control information for vehicles based on the first feature and the first information through the first machine learning model.

[0033] In the second aspect of this application, the control information generation device is also used to perform the steps of the execution device in the first aspect and various possible implementations of the first aspect. The specific implementation methods, the meanings of the terms, and the beneficial effects of the steps in the second aspect can all be found in the first aspect, and will not be repeated here.

[0034] Thirdly, this application provides an apparatus including a processor and a memory, the processor being coupled to the memory, the memory storing program instructions, and the method described in the first aspect being implemented when the program instructions stored in the memory are executed by the processor.

[0035] Fourthly, this application provides an apparatus including a processor and a memory, the processor being coupled to the memory, the memory storing program instructions, and the method described in the first aspect being implemented when the program instructions stored in the memory are executed by the processor.

[0036] Fifthly, this application provides a computer-readable storage medium storing a computer program that, when run on a computer, causes the computer to perform the method described in the first aspect.

[0037] Sixthly, this application provides a computer program product comprising a program that, when run on a computer, causes the computer to perform the method described in the first aspect.

[0038] Seventhly, this application provides a chip system including a processor for supporting the implementation of the functions involved in the foregoing aspects, such as transmitting or processing data and / or information involved in the foregoing methods. In one possible design, the chip system further includes a memory for storing program instructions and data necessary for the terminal device or communication device. This chip system may be composed of chips or may include chips and other discrete devices.

[0039] The third to seventh aspects of this application correspond to the first aspect or multiple possible ways of the first aspect, and have corresponding beneficial effects. Attached Figure Description

[0040] Figure 1 is a structural diagram of an artificial intelligence main framework provided in this application;

[0041] Figure 2 is a system architecture diagram of a regulatory information acquisition system provided in an embodiment of this application;

[0042] Figure 3 is a flowchart illustrating a method for generating regulatory information according to an embodiment of this application.

[0043] Figure 4 is a schematic diagram of a process for obtaining first information based on a first feature through a second machine learning model, according to an embodiment of this application.

[0044] Figure 5 is a schematic diagram of a second machine learning model provided in an embodiment of this application;

[0045] Figure 6 is a schematic diagram of another method for generating regulatory information provided in this application;

[0046] Figure 7 is a schematic diagram of another method for generating regulatory information provided in this application;

[0047] Figure 8 is a schematic diagram of processing the features of at least one first token using the first module provided in this application;

[0048] Figure 9 is a schematic diagram of another method for generating regulatory information provided in an embodiment of this application;

[0049] Figure 10 is a schematic diagram of a regulatory information generation device provided in an embodiment of this application;

[0050] Figure 11 is a schematic diagram of a device provided in an embodiment of this application;

[0051] Figure 12 is a structural schematic diagram of a vehicle provided in an embodiment of this application. Detailed Implementation

[0052] The embodiments of this application will now be described with reference to the accompanying drawings. Obviously, the described embodiments are merely some, and not all, of the embodiments of this application. Those skilled in the art will recognize that, with the emergence of new application scenarios, the technical solutions provided by the embodiments of this application are also applicable to similar technical problems.

[0053] The terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such terms are interchangeable where appropriate; this is merely a way of distinguishing objects with the same attributes in the description of embodiments of this application. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion, so that a process, method, system, product, or apparatus that comprises a series of units is not necessarily limited to those units, but may include other units not explicitly listed or inherent to those processes, methods, products, or apparatuses.

[0054] In the embodiments of this application, "instruction" can include direct and indirect instructions, as well as explicit and implicit instructions. The information indicated by a certain piece of information (hereinafter referred to as instruction information) is called the information to be instructed. In specific implementation, there are many ways to indicate the information to be instructed, such as, but not limited to, directly indicating the information to be instructed, such as the information to be instructed itself or its index. It can also indirectly indicate the information to be instructed by indicating other information, where there is an association between the other information and the information to be instructed; or it can indicate only a part of the information to be instructed, while the other parts are known or pre-agreed upon. For example, the instruction can be implemented by using a pre-agreed (e.g., protocol predefined) arrangement of various information, thereby reducing the instruction overhead to a certain extent. This application does not limit the specific method of instruction. It is understood that for the sender of the instruction information, the instruction information can be used to indicate the information to be instructed; for the receiver of the instruction information, the instruction information can be used to determine the information to be instructed.

[0055] First, the overall workflow of the artificial intelligence system is described, as shown in Figure 1. Figure 1 is a structural diagram of one aspect of the artificial intelligence framework provided in this application. The framework is then elaborated on from two dimensions: the "intelligent information chain" (horizontal axis) and the "IT value chain" (vertical axis). The "intelligent information chain" reflects a series of processes from data acquisition to processing. For example, it could be the general process of intelligent information perception, intelligent information representation and formation, intelligent reasoning, intelligent decision-making, and intelligent execution and output. In this process, data undergoes a condensation process of "data—information—knowledge—wisdom." The "IT value chain" reflects the value that artificial intelligence brings to the information technology industry, from the underlying infrastructure of human intelligence and information (provided and processed by technology) to the industrial ecosystem of the system.

[0056] (1) Infrastructure

[0057] The infrastructure provides computing power to support artificial intelligence systems, enabling communication with the external world and providing support through a basic platform. Communication with the outside world is achieved through sensors; computing power is provided by intelligent chips, which can specifically employ hardware acceleration chips such as central processing units (CPUs), embedded neural network processing units (NPUs), graphics processing units (GPUs), tensor processing units (TPUs), application-specific integrated circuits (ASICs), or field-programmable gate arrays (FPGAs). The basic platform includes distributed computing frameworks and related platform guarantees and support, which may include cloud storage and computing, interconnected networks, etc. For example, sensors communicate with the outside world to acquire data, and this data is provided to intelligent chips in the distributed computing system provided by the basic platform for computation.

[0058] (2) Data

[0059] The data at the next layer of infrastructure is used to represent the data sources in the field of artificial intelligence. The data involves graphics, images, voice, text, and IoT data from traditional devices, including business data from existing systems and sensor data such as force, displacement, liquid level, temperature, and humidity.

[0060] (3) Data processing

[0061] Data processing typically includes methods such as data training, machine learning, deep learning, search, reasoning, and decision-making.

[0062] Among them, machine learning and deep learning can perform intelligent information modeling, extraction, preprocessing, and training on data, including symbolization and formalization.

[0063] Reasoning refers to the process in which, in a computer or intelligent system, the machine thinks and solves problems by simulating human intelligent reasoning, based on reasoning control strategies and using formalized information. Typical functions include search and matching.

[0064] Decision-making refers to the process of making decisions based on intelligent information after reasoning, and it typically provides functions such as classification, sorting, and prediction.

[0065] (4) General ability

[0066] After the data processing mentioned above, the results of the data processing can be used to form some general capabilities, such as algorithms or a general system, for example, translation, text analysis, computer vision processing, speech recognition, image recognition, etc.

[0067] (5) Smart Products and Industry Applications

[0068] Intelligent products and industry applications refer to products and applications of artificial intelligence systems in various fields. They encapsulate overall artificial intelligence solutions, productize intelligent information decision-making, and realize practical applications. Their application areas mainly include: intelligent terminals, intelligent manufacturing, intelligent transportation, smart homes, intelligent healthcare, intelligent security, intelligent driving, and smart cities.

[0069] The method provided in this application can be applied to the field of intelligent driving. For example, the method provided in this application can be used in application scenarios for generating vehicle control information. The vehicle can be a car, truck, motorcycle, bus, boat, lawnmower, recreational vehicle, amusement park vehicle, construction equipment, tram, golf cart, train, airplane and helicopter, etc. The embodiments of this application do not impose any special limitations.

[0070] For example, the vehicle control information in this application may include at least one of the following: the planned position of the vehicle corresponding to each of at least one time moments, the driving strategy of the vehicle corresponding to each of at least one time moments, the control information of the vehicle corresponding to each of at least one time moments, or other types of control information, which may be determined in combination with the actual application scenario; the aforementioned at least one time moment may be at least one time moment after the current time moment.

[0071] For example, the planned position of a vehicle corresponding to each time moment can also be understood as the planned trajectory point of the vehicle corresponding to each time moment. If at least one time moment is specifically at least two time moments, then the at least two planned trajectory points of the vehicle corresponding to at least two time moments can also be understood as the planned trajectory of the vehicle at the aforementioned at least two time moments.

[0072] The driving strategy for a vehicle at each time point can include the lateral driving strategy and / or longitudinal driving strategy corresponding to each time point. The lateral driving strategy corresponding to a certain time point can be left turn, straight, right turn, lane change or lateral avoidance, etc., and the longitudinal driving strategy corresponding to a certain time point can be acceleration, constant speed or deceleration, etc. The specific manifestation of the driving strategy can be determined in combination with the actual application scenario.

[0073] The vehicle control information corresponding to each moment may include: control signals of at least one component in the vehicle at each moment, and / or, the planned vehicle state information at each moment. For example, at least one component may include a steering wheel, engine, brakes, clutch, turn signals, or other components used during vehicle operation; the planned vehicle state information corresponding to each moment after the current moment can be understood as the state information that the vehicle needs to reach at each of the aforementioned moments.

[0074] In this application embodiment, multiple possible scenarios of vehicle control information are provided, which helps to improve the implementation flexibility of this solution and expand the application scenarios that this application can adapt to.

[0075] In related technologies, environmental information is often acquired through sensors and input into a machine learning model to obtain vehicle control information generated by the model. This machine learning model can be understood as an end-to-end machine learning model. However, as the performance requirements for intelligent driving increase, the performance requirements for control information also increase. To provide a better control information generation solution, this application discloses that: the execution device not only inputs environmental information into a first machine learning model and extracts features from the environmental information using the first machine learning model to obtain a first feature (including images and / or point cloud data corresponding to the traffic environment), but also obtains first information based on the first feature through a second machine learning model, wherein the second machine learning model has more parameters than the first machine learning model; and based on the first feature and the first information, generates vehicle control information through the first machine learning model.

[0076] Before providing a detailed description of the method provided in this application, the architecture of the regulatory information acquisition system provided in this application will be described first. Please refer to Figure 2. Figure 2 is a system architecture diagram of the regulatory information acquisition system provided in an embodiment of this application. In Figure 2, the regulatory information acquisition system 200 includes a training device 210, a database 220, an execution device 230, and a data storage system 240. The execution device 230 includes a computing module 231.

[0077] The database 220 stores a training dataset. During the training phase, the training device 210 can use the training dataset to perform training operations to obtain a trained first machine learning model 201 and a trained second machine learning model 202. For example, the aforementioned training operations may include at least one of the following: jointly training the first machine learning model 201 and the second machine learning model 202 using the training dataset; independently training the first machine learning model 201; or independently training the second machine learning model 202.

[0078] It should be noted that if the training device 210 uses the training dataset to jointly train the first machine learning model 201 and the second machine learning model 202, the execution subject of the independent training of the first machine learning model 201 can be the training device 210 or other training devices. The execution subject of the independent training of the second machine learning model 201 can be the training device 210 or other training devices, which can be determined according to the actual application scenario.

[0079] The trained first machine learning model 201 and trained second machine learning model 202 can be deployed in the computing module 231 of the execution device 230. Optionally, as shown in Figure 2, the execution device 230 can be integrated into the vehicle, allowing users to directly interact with the vehicle where the execution device 230 is deployed. For example, the execution device 230 can be a module in the vehicle's host CPU that uses machine learning models for data processing. The execution device 230 can also be a graphics processing unit (GPU), neural network processing unit (NPU), or tensor processing unit (TPU) in the vehicle, etc. The aforementioned GPU, NPU, or TPU is mounted as a coprocessor on the vehicle's host processor, with tasks assigned by the host processor. In the application phase, after acquiring environmental information, the intelligent driving system in the vehicle can obtain vehicle control information through the trained first machine learning model 201 and trained second machine learning model 202 deployed in the computing module 231.

[0080] The execution device 230 can access data, code, etc., in the data storage system 240, and can also store data, instructions, etc., in the data storage system 240. The data storage system 240 can be located within the execution device 230, or it can be an external memory relative to the execution device 230.

[0081] It should be noted that Figure 2 is merely a schematic diagram of one architecture of the regulatory information acquisition system provided in this application embodiment, and the positional relationships between the devices, components, modules, etc. shown in the figure do not constitute any limitation. For example, in some other embodiments of this application, the execution device 230 and the vehicle can be separate and independent devices. The execution device 230 is configured with an input / output (I / O) interface, through which the execution device 230 can interact with the vehicle. For example, in the application phase, the intelligent driving system in the vehicle can send environmental information to the execution device 230 through the I / O interface. After the execution device 230 obtains the vehicle's regulatory information through the trained first machine learning model 201 and trained second machine learning model 202 deployed in the computing module 231, it can send the aforementioned regulatory information to the intelligent driving system in the vehicle through the I / O interface.

[0082] For example, in some other embodiments of this application, the trained first machine learning model 201 and the trained second machine learning model 202 can be deployed in different devices. The trained first machine learning model 201 is deployed in a vehicle, and the trained second machine learning model 202 is deployed in an execution device 230. The execution device 230 is configured with an input / output (I / O) interface, through which the execution device 230 can interact with the vehicle for data. For example, in the application phase, after the intelligent driving system in the vehicle obtains environmental information, it obtains a first feature through the trained first machine learning model 201 and sends the first feature to the execution device 230 through the I / O interface. After the execution device 230 obtains the first information through the trained second machine learning model 202, it can send the aforementioned first information to the intelligent driving system in the vehicle through the I / O interface.

[0083] For example, in some other embodiments of this application, the training device 210 and the execution device 230 may also be integrated into the same device, and the specific architecture of the regulatory information acquisition system can be determined according to the actual application scenario. The following describes the specific implementation process of the regulatory information generation method provided in this application by taking the deployment of the first trained machine learning model 201 and the second trained machine learning model 202 in the same device as an example.

[0084] Please refer to Figure 3, which is a flowchart illustrating a method for generating regulatory information according to an embodiment of this application. The method for generating regulatory information according to an embodiment of this application may include:

[0085] 301. Input the environmental information into the first machine learning model, and extract the first feature from the environmental information through the first machine learning model. The environmental information includes images and / or point cloud data corresponding to the traffic environment.

[0086] For example, if the executing device is a vehicle, the intelligent driving system in the executing device can collect environmental information through sensors before executing step 301; if the executing device is a cloud server connected to the vehicle, the executing device can receive environmental information sent by the vehicle before executing step 301. After obtaining the environmental information, the executing device can input the environmental information into the first machine learning model, and extract the first feature from the environmental information through the first machine learning model.

[0087] The environmental information includes at least one image and / or point cloud data corresponding to the traffic environment. For example, the vehicle can acquire at least one image of the traffic environment around the vehicle through a first sensor, and / or acquire point cloud data of the traffic environment around the vehicle through a second sensor. For example, the acquired image can be an image under a perspective view (PV). The "traffic environment around the vehicle" can be understood as the environment within the field of view of the sensors deployed on the vehicle (such as the aforementioned first or second sensor).

[0088] For example, the first sensor can be a photoelectric sensor, such as a camera or an event camera; the second sensor can be an ultrasonic sensor, a lidar sensor, a millimeter-wave radar sensor, or other sensors capable of measuring and obtaining point cloud data, etc., and this application does not exhaustively list them in the embodiments.

[0089] The first feature may include features of at least one image corresponding to the traffic environment and / or features of point cloud data. Optionally, the first feature may be in the form of a feature map, and the first feature may be continuous feature information. In other words, the first value space of the feature values ​​in the first feature may be continuous. The first value space refers to the set of all possible values ​​that the feature values ​​in the first feature can take. The number of values ​​in the first value space is not finite; in other words, the number of values ​​in the first value space may be infinite. For example, the second value space may be any real value within a certain value range, or the second value space may be any real value.

[0090] 302. Based on the first feature, the first information is obtained through the second machine learning model, wherein the number of parameters of the second machine learning model is greater than the number of parameters of the first machine learning model.

[0091] For example, after obtaining the first feature, in one implementation, the execution device can directly input the first feature into a second machine learning model, and process the first feature through the second machine learning model to obtain the first information.

[0092] Alternatively, in another implementation, after obtaining the first feature, the execution device can input the first feature into the second module, and the second module processes the first feature to obtain the third feature; then the execution device can input the third feature into the second machine learning model, and the second machine learning model processes the third feature to obtain the first information.

[0093] Optionally, the function of the second module includes compressing the first feature. The third feature is obtained by compressing the first feature using the second module, and the data size of the third feature is smaller than that of the first feature. In this embodiment, after compressing the first feature, the compressed third feature is input into the second machine learning model, thereby reducing the amount of data that the second machine learning model needs to process. This helps to reduce the computer resources consumed in the process of generating the first information through the second machine learning model, and also helps to reduce the latency caused by the process of generating the first information through the second machine learning model.

[0094] Optionally, the function of the second module includes converting the first feature in feature form into the third feature in token form. For example, the feature form can be understood as a matrix form, and the token form can be understood as a vector form.

[0095] Optionally, the function of the second module includes mapping the feature values ​​in the first feature to obtain the feature values ​​in the third feature; for example, the third feature can be discrete feature information, in other words, the second value space of the feature values ​​in the first feature can be discrete, the second value space refers to the set of all numerical values ​​that the feature values ​​in the third feature can take, the number of numerical values ​​in the second value space is finite, and any feature value in the first feature is contained in a finite set of numerical values ​​(i.e., the first value interval).

[0096] For example, the second module may include at least one of the following: a fully connected neural network layer, an attention-based neural network layer, a convolutional neural network layer, or a multilayer perceptron (MLP), etc. For example, the second module may be an encoder, and the specific form of the second module may be combined with the actual application scenario.

[0097] To understand this solution more intuitively, please refer to Figure 4. Figure 4 is a flowchart illustrating how a first feature is used to obtain first information through a second machine learning model, as provided in an embodiment of this application. As shown in Figure 4, after obtaining the first feature, the execution device inputs the first feature into a second module, which processes the first feature to obtain a third feature. Then, the execution device can input the third feature into the second machine learning model, which processes the third feature to obtain the first information. It should be understood that the example in Figure 4 is only for the convenience of understanding this solution and is not intended to limit this solution.

[0098] For example, in one scenario, the second machine learning model is derived from a large language model (LLM), which can be understood as a large model with language expression capabilities. In this application, a large model can be understood as a machine learning model with a huge number of parameters; for example, a large model can be a machine learning model with more than one hundred million parameters, or optionally, the number of parameters can be greater than or equal to one billion. "Having language expression capabilities" can be understood as the LLM being able to output information indicating the content of the text. Optionally, the LLM can be a machine learning model based on an attention mechanism, that is, the LLM can include neural network layers based on an attention mechanism. For example, the LLM can adopt the Pangu model or other types of large language models, etc., and the specific choice can be determined based on the actual application scenario.

[0099] Further, in one implementation, the first information may include features of each first token among at least one first tokens generated by the LLM, and the second machine learning model may be an LLM. Step 302 may include: the execution device inputting the first feature (or third feature) into the second machine learning model to obtain features of each first token among at least one first tokens generated by the second machine learning model. The features of each first token are generated by a neural network layer (hereinafter referred to as the "target neural network layer" for convenience) in the second machine learning model. The features of each first token can be understood as follows: during the generation of the first token through the second machine learning model, the features of each first token are first generated by the second machine learning model, and then each first token is obtained based on the features of each first token. Thus, the execution device can obtain the features of each first token during the generation of at least one first token through the second machine learning model. For example, the target neural network layer may include the Mth-last neural network layer in the second machine learning model, where M is an integer greater than 1.

[0100] Optionally, when the second machine learning model is an LLM, the input to the second machine learning model may further include a prompt, such as an image and / or first text. For example, the aforementioned image may include at least one image of a traffic environment; optionally, the execution device may input the aforementioned image into a visual encoder, generate initial features in the form of tokens for the aforementioned image through the visual encoder, and then input the initial features of the aforementioned image into the second machine learning model.

[0101] The first text can be used to prompt the second machine learning model what task to perform; for example, the task type prompted by the first text to the second machine learning model may include at least one of the following: target detection of static and / or dynamic obstacles in a traffic environment, 3D modeling of a traffic environment, text recognition in a traffic environment, determining the spatial relationship between different objects in a traffic environment, intention judgment and / or behavior prediction of dynamic obstacles in a traffic environment, determining the strategy adopted when interacting with dynamic obstacles in a traffic environment, identifying risky objects in a traffic environment, determining the driving strategy of a vehicle, determining the speed limit of a vehicle when driving in the traffic environment, performing path planning for a vehicle, or determining the control information of a vehicle, etc. The specific task can be determined in combination with the actual application scenario, and is not exhaustive in the embodiments of this application.

[0102] For example, a 3D model of the traffic environment is performed to obtain 3D information about the traffic environment. This 3D information may include the 3D position information of obstacles in the traffic environment. For example, determining the spatial relationships between different objects in the traffic environment may include: determining the spatial relationship between at least one object in the traffic environment and lanes in the traffic environment, thereby allowing the location of objects in the traffic environment to be understood using the lanes. For example, the aforementioned relationship may include: object A and object B are located in lane 1, object C is located in lane 2, and object D is located outside all lanes, etc. It should be understood that this example is only for the convenience of understanding the solution and is not intended to limit the solution. Alternatively, determining the spatial relationships between different objects in the traffic environment may include: determining the spatial relationship between objects in the traffic environment and the vehicle. For example, the aforementioned relationship may include: the distance and / or orientation of at least one object in the traffic environment relative to the vehicle. The specific information included in the aforementioned relationship can be flexibly determined based on the actual application scenario.

[0103] For example, the intention of a dynamic obstacle can be: to cut in, to give way, to stop moving, to cross the road, to cut into the vehicle's lane, or other intentions. The predicted behavior of a dynamic obstacle can be: to cross the road, to cut into the vehicle's lane, to merge with the vehicle's lane, to drive in the opposite direction to the vehicle, to drive parallel to the vehicle in the same direction, to stop or to start.

[0104] For example, strategies employed when interacting with dynamic obstacles in a traffic environment may include: cutting in, yielding, swerving to the left, swerving to the right, or other strategies. Risky objects in a traffic environment may include other vehicles, electric vehicles, pedestrians, animals, or other types of objects. The meaning of vehicle driving strategies and vehicle control information can be found in the above description and will not be repeated here.

[0105] For example, the second machine learning model can generate at least one first token using an autoregressive approach. Specifically, the second machine learning model generates at least one first token incrementally. After inputting the first feature (or third feature) into the second machine learning model, the model can generate the first first token among at least one first tokens. This process can also be referred to as full inference using the second machine learning model. After inputting the first feature (or third feature) plus the first first token into the second machine learning model, the model can generate the second first token among at least one first tokens. This process can also be referred to as incremental inference using the second machine learning model. After repeating the aforementioned steps at least once, at least one first token is generated using the second machine learning model. Optionally, at least one first token includes the first N tokens generated by the second machine learning model using an autoregressive approach, where N is an integer greater than or equal to 1. The value of N is preset; for example, N can be 2, 3, 4, 5, 6, 7, 8, or other values, which can be determined based on the actual application scenario. Step 302 may include: the execution device inputs the first feature (or the third feature) into the second machine learning model, and the second machine learning model generates N first tokens step by step in an autoregressive manner, where the value of N is preset, and in the process of generating each first token among the N first tokens, the features of each first token are obtained, thus obtaining the features of each first token among the N first tokens.

[0106] Since the second machine learning model generates at least one first token using an autoregressive approach, meaning that the second machine learning model generates one of the first tokens in each inference operation, the number of inference operations performed by the second machine learning model increases as the number of tokens included in the at least one first token increases. This leads to increased computer resource consumption and latency. In this application, the number of at least one first token is pre-set to N, thereby limiting the number of inference operations performed by the second machine learning model to N. This improves the controllability of computer resources consumed and latency caused by the process of generating at least one first token through the second machine learning model, and also improves the controllability of computer resources consumed and latency caused by the process of obtaining first information through the second machine learning model. This helps to reduce the computer resources consumed and latency caused by the process of obtaining first information through the second machine learning model.

[0107] Alternatively, when the execution device inputs the first feature (or the third feature) into the second machine learning model and generates at least one first token through the second machine learning model using an autoregressive approach, the number of at least one first token is not limited to N. Instead, the number of first tokens included in the at least one first token generated is determined by the second machine learning model.

[0108] Optionally, in one case, the aforementioned first token includes the first N tokens generated by the second machine learning model in an autoregressive manner. The first N tokens do not correspond to text; in other words, the first N tokens cannot directly indicate text with semantic meaning. The first N tokens can be special tokens that do not correspond to text.

[0109] Optionally, the training phase of the second machine learning model may include, for example, a joint training phase of the first and second machine learning models, or a fine-tuning phase within the independent training phase of the second machine learning model. For example, in the joint training phase of the first and second machine learning models, the input to the second machine learning model includes a first feature (or a third feature); in the fine-tuning phase of the independent training phase of the second machine learning model, the input to the second machine learning model may include the aforementioned image of the traffic environment and / or the first text. Optionally, before the fine-tuning stage in the independent training stage of the second machine learning model, a pre-training stage of the second machine learning model may be included. The pre-training stage of the second machine learning model can utilize a large number of training samples to perform numerous understanding tasks. For example, the training samples may include images, speech, text, or video. The understanding tasks performed by the second machine learning model may include: generating text descriptions corresponding to images, generating text descriptions corresponding to speech, generating text descriptions corresponding to text, generating text descriptions corresponding to videos, etc. It should be noted that the aforementioned images, speech, text, or video can be data from various fields. These various fields include not only the field of intelligent driving but also other fields, such as the field of smart terminals, smart homes, smart healthcare, or other fields, etc. This application does not exhaustively list them. Performing understanding tasks on data from various fields can greatly improve the understanding ability of the second machine learning model.

[0110] In the joint training phase of the first and second machine learning models, or in the fine-tuning phase of the independent training phase of the second machine learning model, after the training device gradually generates the first N tokens using an autoregressive approach through the second machine learning model, it can continue to generate at least one second token corresponding to the predicted text using the second machine learning model using an autoregressive approach. In other words, at least one second token can indicate the predicted text with semantic meaning. The training phase of the second machine learning model employs a loss function (hereinafter referred to as the "first loss function" for ease of description). The first loss function indicates the similarity between at least one second token and at least one third token. The at least one third token includes a token corresponding to the expected text (hereinafter referred to as the "first expected text" for ease of description). The purpose of training using the first loss function includes improving the similarity between at least one second token and at least one third token; that is, the purpose of training using the first loss function includes improving the similarity between the predicted text generated by the second machine learning model and the first expected text. For example, the predicted text corresponding to at least one second token can be understood as the semantic meaning carried by the first N tokens. The predicted text corresponding to at least one second token can help understand the semantic meaning of the first N tokens. Optionally, the first loss function also indicates the similarity between the top N tokens and the N special tokens, and the goal of training using the first loss function also includes improving the similarity between the top N tokens and the N special tokens.

[0111] For example, the first expected text may include at least one of the following: the detection result obtained by target detection of static and / or dynamic obstacles in the traffic environment, the three-dimensional information obtained by three-dimensional modeling of the traffic environment, the recognition result obtained by recognizing text in the traffic environment, the spatial relationship between different objects in the traffic environment, the predicted intent and / or predicted behavior of dynamic obstacles in the traffic environment, the strategy adopted when interacting with dynamic obstacles in the traffic environment, the recognition result of risky objects in the traffic environment, the driving strategy of the vehicle, the speed limit of the vehicle when driving in the traffic environment, the planned path of the vehicle, or the control information of the vehicle, etc. Since the goal of training using the first loss function includes improving the similarity between the predicted text and the first expected text, the meaning of the predicted text is similar to the meaning of the first expected text, and will not be repeated here.

[0112] Optionally, the first loss function also indicates the similarity between the first N tokens and the N special tokens. The purpose of training with the first loss function also includes improving the similarity between the first N tokens and the N special tokens, thereby ensuring that the first N tokens generated by the second machine learning model are N feature tokens that do not correspond to the text.

[0113] For example, after obtaining at least one second token, the training device can generate the function value of the first loss function based on at least one second token and at least one third token (optionally, also including the first N tokens and N special tokens). Based on the function value of the first loss function, the weight parameters of the second machine learning model are updated using the backpropagation algorithm to achieve one training of the second machine learning model. The training device can repeat the above steps to achieve iterative training of the second machine learning model until the convergence condition is met. For example, the convergence condition may include: satisfying the convergence condition of the first loss function and / or the number of iterative training reaches a preset number.

[0114] To better understand this solution, please refer to Figure 5. Figure 5 is a schematic diagram of a second machine learning model provided in an embodiment of this application. As shown in Figure 5, during the joint training phase of the first and second machine learning models, the input of the second machine learning model includes a third feature and a prompt. After the training device generates the first N special tokens through the second machine learning model using an autoregressive approach, it will also generate at least one second token corresponding to the predicted text. The training device can generate the function value of the first loss function based on at least one second token and at least one third token (optionally, also including the first N tokens and N special tokens). Based on the function value of the first loss function, the weight parameters of the first and second machine learning models are updated using the backpropagation algorithm to achieve one training of the first and second machine learning models. It should be understood that the example in Figure 5 is only for the convenience of understanding this solution and is not intended to limit this solution.

[0115] In this embodiment, since the first N tokens do not correspond to text, each of the first N tokens does not need to uniquely indicate a specific word. This allows the first N tokens to carry richer information. Furthermore, during the training phase of the second machine learning model, after generating the first N special tokens that do not correspond to text, a second token corresponding to the predicted text is also generated. This not only preserves the language expression capabilities of the LLM, but also, since the first N tokens and the second token are generated by the same LLM when performing a certain task, the predicted text pointed to by the second token can be used to explain the meaning of the first N tokens. During the training phase of the second machine learning model, a loss function can be used to guide the second machine learning model to generate second tokens that can indicate the expected text, thereby guiding the first N tokens to carry the meaning of the expected text, which is beneficial to further improve the quality of the first N tokens.

[0116] Optionally, in the application stage of the second machine learning model, in addition to generating the first N tokens through the second machine learning model to obtain the first information, at least one second token corresponding to the predicted text can be generated through the second machine learning model. The predicted text indicated by at least one second token can be used to assist in understanding the behavior of the vehicle's intelligent driving system, thereby improving the interpretability of the intelligent driving system.

[0117] In another scenario, the aforementioned first token includes the first N tokens generated by the second machine learning model using an autoregressive approach. The first N tokens correspond to the text; in other words, the first N tokens can directly indicate the second text with semantic meaning. For example, the information included in the second text can be found in the above explanation of the information included in the first expected text, which will not be repeated here.

[0118] In another case, the number of tokens included in at least one first token is determined by a second machine learning model. The aforementioned at least one first token corresponds to text; in other words, the aforementioned at least one first token can directly indicate a third text with semantic meaning. For example, the information included in the third text can be found in the above explanation of the information included in the first expected text, which will not be repeated here.

[0119] Optionally, during the training phase of the second machine learning model, such as in the joint training phase of the first and second machine learning models, or in the fine-tuning phase of the independent training phase of the second machine learning model, the training device can train the second machine learning model using multiple training data sets. Each training data set may include at least one fourth token, where the at least one fourth token indicates a second expected text corresponding to the third text. The number of tokens included in the at least one fourth token may be less than or equal to P, where P is an integer greater than or equal to 1. The value of P is preset, for example, P can be 2, 3, 4, 5, 6, 7, 8, or other values, which can be determined based on the actual application scenario. For example, during the joint training phase of the first and second machine learning models, or in the fine-tuning phase of the independent training phase of the second machine learning model, the training device can generate at least one first token through the second machine learning model and train the second machine learning model using a second loss function. The second loss function indicates the similarity between at least one first token and at least one fourth token; in other words, the second loss function indicates the similarity between the third text and the second expected text. The purpose of training using the second loss function includes improving the similarity between at least one first token and at least one fourth token. Since the training process of the second machine learning model guides the generation of at least one first token to mimic at least one fourth token, by controlling the number of at least one fourth token, the generation of at least one first token by the second machine learning model can contain fewer tokens. This also helps to reduce the computer resources consumed by the second machine learning model in the process of generating at least one first token and the resulting latency.

[0120] For example, after obtaining at least one first token, the training device can generate the function value of the second loss function based on at least one first token and at least one fourth token. Based on the function value of the second loss function, the weight parameters of the first machine learning model are updated using the backpropagation algorithm to achieve one training of the first machine learning model. The training device can repeat the above steps to achieve iterative training of the first machine learning model until the convergence condition is met. For example, the convergence condition may include: satisfying the convergence condition of the second loss function and / or the number of iterative training reaches a preset number.

[0121] In another implementation, the first information may include at least one first token generated by the LLM, and the second machine learning model may be an LLM. Step 302 may include: the execution device inputting the first feature (or the third feature) into the second machine learning model to obtain at least one first token generated by the second machine learning model.

[0122] Optionally, when the second machine learning model is an LLM, the input to the second machine learning model may also include a prompt, such as an image and / or first text. For example, the aforementioned image may include at least one image of a traffic environment; the first text may be used to prompt the second machine learning model what task it should perform.

[0123] For example, the second machine learning model can generate at least one first token by means of autoregression; optionally, the at least one first token includes the first N tokens generated by the second machine learning model by means of autoregression, where the value of N is preset. Step 302 may include: the execution device inputs the first feature (or the third feature) into the second machine learning model, and the second machine learning model generates N first tokens step by step by means of autoregression, wherein the first information includes the N first tokens.

[0124] Optionally, during the training phase of the second machine learning model, after the training device generates the first N tokens step by step using the second machine learning model in an autoregressive manner, it can continue to generate at least one second token corresponding to the predicted text using the second machine learning model in an autoregressive manner. The training phase of the second machine learning model uses a first loss function, which can indicate the similarity between at least one second token and at least one third token. At least one third token includes the token corresponding to the first expected text.

[0125] Alternatively, when the execution device inputs the first feature (or the third feature) into the second machine learning model and generates at least one first token through the second machine learning model using an autoregressive approach, the number of at least one first token is not limited to N. Instead, the number of first tokens included in the at least one first token generated is determined by the second machine learning model.

[0126] For an understanding of the meaning of at least one first token, the first text, the autoregressive method, the value of N, the training stage of the second machine learning model, the first loss function, the predicted text, and the first expected text, please refer to the above description, which will not be repeated here.

[0127] In another implementation, the first information includes the features of each first token in at least one first token generated by the LLM and the features of each first token in at least one first token generated by the LLM. Step 302 may include: the execution device inputting the first feature (or the third feature) into the second machine learning model to obtain at least one first token generated by the second machine learning model and the features of each first token in at least one first token generated by the second machine learning model. It should be noted that the specific meanings of the terms and the specific implementation methods of the steps in this implementation can be referred to the above description, and will not be repeated here.

[0128] In another scenario, a second machine learning model is used to update the first feature. The first information includes the updated first feature. That is, the first feature is updated using a second machine learning model with a larger number of parameters to fully extract the effective information carried in the environmental information. In this case, the updated first feature may include richer and more effective information obtained from the environmental information.

[0129] For example, the second machine learning model can be a convolutional neural network, a fully connected neural network, a residual neural network, or a neural network based on an attention mechanism, etc., which can be determined based on the actual application scenario.

[0130] This application provides several possible implementations of the second machine learning model and the first information, improving the implementation flexibility of this solution and facilitating the expansion of the application scenarios to which this application is adapted. When the second machine learning model is used to update the first feature, it is beneficial to use the second machine learning model with a larger number of parameters to mine richer and more effective information in the environmental information, thereby obtaining better-performing regulatory information based on the first feature and the updated first feature. When the second machine learning model is based on LLM, since LLM often cannot extract three-dimensional features from the original environmental information, the first feature obtained by feature extraction from traffic environment images and / or point cloud data by the first machine learning model carries three-dimensional information. Processing the first feature by the second machine learning model helps to supplement the three-dimensional understanding capability of LLM, thereby improving the quality of the first information obtained by the second machine learning model and further improving the quality of the final regulatory information.

[0131] 303. Based on the first feature and the first information, generate vehicle control information through the first machine learning model.

[0132] After acquiring the first information, the execution device can generate vehicle control information based on the first feature and the first information using a first machine learning model. The meaning of "vehicle control information" can be found in the above description and will not be repeated here. For example, the first machine learning model may include a feature extraction module and a feature processing module. The feature extraction module is used to extract features from the environmental information to obtain the first feature. For example, both the feature extraction module and the feature processing module may include at least one neural network layer.

[0133] For example, in one implementation, step 303 may include: the execution device inputs the first feature and the first information into the feature processing module of the first machine learning model, and processes the first feature and the first information through the feature processing module of the first machine learning model to obtain the vehicle control information generated by the first machine learning model.

[0134] To understand this solution more intuitively, please refer to Figure 6. Figure 6 is another flowchart of the method for generating regulatory information provided in this application. As shown in Figure 6, the first machine learning model and the second machine learning model can be understood as being in parallel. The execution device inputs environmental information into the first machine learning model, and extracts features from the environmental information through the feature extraction module in the first machine learning model to obtain the first feature. The first feature is then input into the second machine learning model, and the second machine learning model generates the first information. The first feature and the first information are then input into the feature processing module of the first machine learning model, and the first feature and the first information are processed by the feature processing module of the first machine learning model to obtain the vehicle regulatory information generated by the first machine learning model. It should be understood that the example in Figure 6 is only for the convenience of understanding this solution and is not intended to limit this solution.

[0135] In another implementation, step 303 may include: the execution device fusing the first feature and the first information to obtain fused information; the execution device inputting the fused information into the feature processing module of the first machine learning model; and the feature processing module of the first machine learning model processing the fused information to obtain the vehicle control information generated by the first machine learning model.

[0136] To better understand this solution, please refer to Figure 7. Figure 7 is another flowchart illustrating the method for generating regulatory information provided in this application. As shown in Figure 7, after the execution device generates a first feature through the feature extraction module of the first machine learning model, it can process the first feature through the first module to obtain a third feature. The third feature is then input into the second machine learning model to generate first information. The first feature and the first information are then fused to obtain fused information. The fused information is then input into the feature processing module of the first machine learning model to process the fused information, thereby obtaining the vehicle regulatory information generated by the first machine learning model. It should be understood that the example in Figure 7 is only for the convenience of understanding this solution and is not intended to limit this solution.

[0137] In another implementation, if the first information includes features of at least one first token, step 303 may include: the execution device processing the features of at least one first token through a first module to obtain second features; and generating vehicle control information based on the first and second features through a first machine learning model. For example, processing the features of at least one first token through the first module can be understood as: extracting more information from the features of at least one first token to obtain the second features; in other words, the second features can carry richer information compared to the features of at least one first token.

[0138] Optionally, the execution device can fuse the first feature and the second feature to obtain the fused feature, input the fused feature into the feature processing module of the first machine learning model, and process the fused feature through the feature processing module of the first machine learning model to obtain the vehicle control information generated by the first machine learning model.

[0139] Optionally, the first module may include at least one sub-module, and each sub-module may be a neural network module based on an attention mechanism.

[0140] Optionally, the first module may include at least one sub-module corresponding to at least one task; optionally, each sub-module processes a portion of the features of the first token to obtain a sub-feature, and the second feature may include all the sub-features generated by the at least one sub-module. For example, the features of at least one first token include six features of the first token, namely the features of first token1, first token2, first token3, first token4, first token5, and first token6. The first module includes six sub-modules, namely sub-module1, sub-module2, sub-module3, sub-module4, sub-module5, and sub-module6. The features of first token1 are input into sub-module1 to obtain sub-feature1 generated by sub-module1; the features of first token2 are input into sub-module2 to obtain... Sub-feature 2 is generated by submodule 2; the feature of the first token 3 is input into submodule 3 to obtain sub-feature 3 generated by submodule 3; the feature of the first token 4 is input into submodule 4 to obtain sub-feature 4 generated by submodule 4; the feature of the first token 5 is input into submodule 5 to obtain sub-feature 5 generated by submodule 5; the feature of the first token 6 is input into submodule 6 to obtain sub-feature 6 generated by submodule 6. The second feature includes sub-feature 1, sub-feature 2, sub-feature 3, sub-feature 4, sub-feature 5 and sub-feature 6. It should be understood that the example here is only for the convenience of understanding this scheme and is not intended to limit this scheme.

[0141] For example, any one of the at least one tasks is referred to as the first task, and the sub-module corresponding to the first task is referred to as the first sub-module. The sub-features obtained by updating some of the features of the first token through the first sub-module can be understood as features used to perform the first task. In other words, the purpose of updating some of the features of the first token through the first sub-module includes obtaining the information needed to perform the first task based on the first features.

[0142] For example, the above-mentioned at least one task includes at least one of the following: target detection of dynamic and / or static obstacles in a traffic environment, 3D modeling of the traffic environment, text recognition in the traffic environment, determination of spatial relationships between different objects in the traffic environment, intention judgment and / or behavior prediction of obstacles in the traffic environment, determination of the strategy adopted when interacting with obstacles in the traffic environment, identification of objects with risks in the traffic environment, determination of vehicle driving strategy, determination of the speed limit of the vehicle when driving in the traffic environment, path planning of the vehicle, or determination of vehicle control information, etc. The specific task can be determined in combination with the actual application scenario. It should be noted that the understanding of the above-mentioned at least one task can be referred to the above description, which will not be repeated here.

[0143] For example, if each submodule is a neural network module based on an attention mechanism, the execution device can also input query information into each submodule. The execution device multiplies a portion of the features of at least one first token with a first matrix to obtain key information, and multiplies a portion of the features of at least one first token with a second matrix to obtain value information. Based on the query information, key information, and value information, a sub-feature is generated through each submodule.

[0144] To better understand this solution, please refer to Figure 8. Figure 8 is a schematic diagram of processing the features of at least one first token using a first module, as provided in this application. Figure 8 uses at least one task, including Task 1 and Task 2, as an example. Task 1 is to determine the spatial relationships between different objects in a traffic environment, and Task 2 is to perform route planning for vehicles. The first module includes submodule 1 corresponding to Task 1 and submodule 2 corresponding to Task 2. The features of at least one first token are divided into two features of first token 1 and two features of first token 2. The execution device obtains key information 1 and value information 1 based on the features of the two first token 1. Based on query information 1, key information 1, and value information 1, submodule 1 processes the data to obtain subfeature 1. The execution device identifies subfeature 1 as query information 2. Based on the features of the two first token 2, it obtains key information 2 and value information 2. Based on query information 2, key information 2, and value information 2, submodule 2 processes the data to obtain subfeature 2. The execution device fuses sub-feature 1, sub-feature 2 and the first feature to obtain the fused feature. The feature processing module of the first machine learning model processes the fused feature to obtain the vehicle control information. It should be understood that the example in Figure 8 is only for the convenience of understanding this scheme and is not intended to limit this scheme.

[0145] Optionally, during the joint training phase of the first and second machine learning models, each submodule can be connected to a task head. Each task head is used to generate prediction information corresponding to one of the at least one tasks mentioned above. Each task head can be understood as a feature processing module. For example, each task head can be a fully connected neural network layer, a multilayer perceptron, a support vector machine, a convolutional neural network layer, or other representations. The meaning of the prediction information corresponding to each task can be found in the above description and will not be repeated here. A third loss function can be used during the joint training phase of the first and second machine learning models. The third loss function indicates the similarity between the prediction information generated by each task head and the expected information. The goal of training using the third loss function includes improving the similarity between the prediction information generated by each task head and the expected information. It should be noted that "first expected text," "second expected text," and "expected information" in this application can all be understood as ground truth.

[0146] For example, in the joint training phase of the first machine learning model and the second machine learning model, the training device inputs environmental information into the first machine learning model, extracts features from the environmental information through the feature extraction module of the first machine learning model to obtain a first feature; based on the first feature, the second machine learning model obtains features of at least one first token; the features of at least one first token are processed by at least one sub-module included in the first module to obtain at least one sub-feature (i.e., a second feature) corresponding to at least one sub-module; a sub-feature generated by each sub-module is input into a task head to obtain prediction information generated by each task head; the training device also fuses the first feature and the second feature to obtain a fused feature, and processes the fused feature through the feature processing module of the first machine learning model to obtain vehicle planning information generated by the first machine learning model.

[0147] The training device generates a third loss function value based on the prediction information generated by each task head and the expected information corresponding to each task head. Based on the third loss function value, the backpropagation algorithm is used to update the weight parameters of the first machine learning model, the second machine learning model, and the first module (optionally, also including the second module). Optionally, the training device also generates a fourth loss function value based on the vehicle planning information and expected planning information generated by the first machine learning model. Based on the fourth loss function value, the backpropagation algorithm is used to update the weight parameters of the first machine learning model and the second machine learning model (optionally, also including the first module and / or the second module) to achieve one training cycle.

[0148] The training device can repeat the above steps at least once to achieve iterative training of the first machine learning model and the second machine learning model until the convergence condition is met. The convergence condition may include: meeting the convergence conditions of the third loss function and the fourth loss function, and / or, the iterative training of the first machine learning model and the second machine learning model has reached a preset number of times.

[0149] In this embodiment of the application, after obtaining the features of at least one first token, the features of at least one first token can be processed by the first module to obtain the second features. That is, the first module is added between the first machine learning model and the second machine learning model. The process of processing the features of at least one first token by the first module is conducive to obtaining richer information from at least one first token, which in turn is conducive to the first machine learning model generating better regulatory information based on the first features and the second features.

[0150] Optionally, the first module includes at least one sub-module corresponding to at least one task. Each sub-module can obtain information corresponding to one of the at least one tasks from the features of at least one first token. That is, each sub-module can obtain the information required to execute a certain task. This is beneficial to obtain the rich information required to generate vehicle control information in a targeted manner through at least one sub-module, which in turn is beneficial to the first machine learning model to generate better control information.

[0151] Optionally, based on the embodiment corresponding to Figure 3 above, please refer to Figure 9. Figure 9 is another flowchart illustrating the method for generating regulatory information provided in this application embodiment. The method for generating regulatory information provided in this application embodiment may include:

[0152] 901. Input the environmental information into the first machine learning model, and extract the first feature from the environmental information through the first machine learning model. The environmental information includes images and / or point cloud data corresponding to the traffic environment.

[0153] 902. In the first case, based on the first feature, the first information is obtained through the second machine learning model.

[0154] 903. Based on the first feature and the first information, generate vehicle control information through the first machine learning model.

[0155] 904. In the second case, based on the first feature and preset information, vehicle control information is generated through the first machine learning model.

[0156] For example, step 904 is an optional step. The first case and the second case can be different. The first case can be understood as the case where the first information is obtained through the second machine learning model, and the second case can be understood as the case where the first information is not obtained through the second machine learning model. For example, the second case can be the case where the second machine learning model is not working. In other words, the second case can be the case where the second machine learning model is not invoked. Alternatively, the second case can also be the case where, although the second machine learning model is working, the first information has not yet been generated because the second machine learning model generates the first information slowly, or other cases where the first information is not obtained through the second machine learning model. The specific case can be determined based on the actual situation, and will not be exhaustively listed here.

[0157] For example, the preset information can be pre-stored information, such as the preset information including all values ​​of 0; or, the preset information can be information obtained based on preset rules, such as the preset rules can be to fill in 0 and 1 in sequence, etc. The specific situation can be set according to the actual application scenario.

[0158] Optionally, the frequency of generating vehicle control information through the first machine learning model is higher than the frequency of obtaining the first information through the second machine learning model. For example, the frequency at which the execution device generates vehicle control information through the first machine learning model is Q times the frequency at which the execution device obtains the first information through the second machine learning model, where Q is greater than 1, for example, Q can be 4, 5, 6, 7 or other values.

[0159] It should be noted that this application does not limit the relationship between the number of times steps 902 and 903 are executed and step 904. The number of times steps 902 and 903 are executed is greater than the number of times step 904 is executed, or the number of times step 904 is executed is greater than the number of times steps 902 and 903 are executed. The specific number of times can be determined in combination with the actual application scenario.

[0160] In this embodiment, in the first case, the first information can be generated by the second machine learning model, and then the vehicle's control information can be generated by the first machine learning model based on the first feature and the first information; in the second case, the vehicle's control information can be generated by the first machine learning model based on the first feature and preset information, thereby achieving decoupling between the first machine learning model and the second machine learning model, which is beneficial to ensuring that the vehicle's control information can be generated smoothly under various conditions, so as to improve the stability of the intelligent driving system.

[0161] Since the number of parameters in the second machine learning model is greater than that in the first machine learning model, the computer resources consumed when generating the first information through the second machine learning model are often greater, and the latency of generating the first information through the second machine learning model is also higher. Since the vehicle's intelligent driving system generates vehicle control information multiple times at a high frequency during vehicle operation, setting the frequency of generating vehicle control information through the first machine learning model to be higher than the frequency of obtaining the first information through the second machine learning model is beneficial to reduce the computer resources consumed in the overall process of generating vehicle control information multiple times while improving the performance of generating vehicle control information multiple times. It is also beneficial to reduce the latency of the overall process of generating vehicle control information multiple times.

[0162] To provide a more intuitive understanding of the beneficial effects of this application, the following explanation, based on experimental data, further illustrates these effects. When using only an end-to-end machine learning model to generate vehicle control information, the collision rate in closed-loop testing is 11.64%. When using the method provided in this application to generate vehicle control information, where the first information is at least one first token generated by LLM, and at least one first token represents the vehicle's planned trajectory, the collision rate in closed-loop testing is 11.11%, reducing the probability of a collision. Further optionally, by adding a first module between the first and second machine learning models, the collision rate in closed-loop testing is 10.94%, further reducing the probability of a collision.

[0163] Optionally, when using the method provided in this application to generate vehicle control information, tests were conducted on two cases: one where at least one first token included in the first information corresponds to the text, and the other where the first N tokens included in the first information are special tokens that do not correspond to the text. In the case where at least one first token included in the first information corresponds to the text, the accuracy rate was 89.1%, and the inference process took 5.55 seconds. In the case where the first N tokens included in the first information are special tokens that do not correspond to the text, the accuracy rate was 90.3%, and the inference process took 0.20 seconds. This not only improved the accuracy rate but also significantly reduced the inference process time.

[0164] Based on the embodiments corresponding to Figures 1 to 9, in order to better implement the above-mentioned solutions of the embodiments of this application, related equipment for implementing the above-mentioned solutions is also provided below. Specifically, referring to Figure 10, Figure 10 is a structural schematic diagram of a traffic control information generation device provided in the embodiments of this application. The traffic control information generation device 1000 includes: an input module 1001, used to input environmental information into a first machine learning model, and extract features from the environmental information through the first machine learning model to obtain a first feature, wherein the environmental information includes images and / or point cloud data corresponding to the traffic environment; a processing module 1002, used to obtain first information based on the first feature through a second machine learning model, wherein the number of parameters of the second machine learning model is greater than the number of parameters of the first machine learning model; and a generation module 1003, used to generate vehicle traffic control information based on the first feature and the first information through the first machine learning model.

[0165] Optionally, the second machine learning model is obtained based on a large language model (LLM), and the first information includes features of at least one first token generated by the LLM and / or at least one first token generated by the LLM, wherein the features of the first token are generated by a neural network layer in the LLM; or, the second machine learning model is used to update the first features, and the first information includes the updated first features.

[0166] Optionally, the second machine learning model is an LLM, and at least one first token includes the first N tokens generated by the second machine learning model in an autoregressive manner, where N is an integer greater than or equal to 10, and the value of N is preset.

[0167] Optionally, the first information includes features of at least one first token. The generation module 1003 is specifically used to: process the features of at least one first token through the first module to obtain second features; and generate vehicle control information through a first machine learning model based on the first and second features.

[0168] Optionally, the first module includes at least one sub-module corresponding to at least one task, wherein the at least one task includes at least one of the following: object detection, 3D modeling of the traffic environment, text recognition in the traffic environment, determining the spatial relationship between different objects in the traffic environment, intention judgment and / or behavior prediction of obstacles in the traffic environment, determining the strategy adopted when interacting with obstacles in the traffic environment, identifying objects with risks in the traffic environment, determining the vehicle's driving strategy, determining the speed limit of the vehicle when driving in the traffic environment, performing path planning for the vehicle, or determining the vehicle's control information.

[0169] Optionally, the first N tokens do not correspond to the text. During the training phase of the second machine learning model, the second machine learning model is also used to generate a second token corresponding to the predicted text. The second machine learning model is trained using a loss function that indicates the similarity between the second token and the third token, which includes the token corresponding to the expected text.

[0170] Optionally, the processing module 1002 is specifically used to obtain first information based on the first feature through a second machine learning model in the first case; the generation module 1003 is also used to generate vehicle control information based on the first feature and preset information through a first machine learning model in the second case.

[0171] Optionally, the frequency of generating vehicle control information through the first machine learning model is higher than the frequency of obtaining the first information through the second machine learning model.

[0172] Optionally, the processing module 1002 is specifically used for: inputting the first feature into the second module, compressing the first feature through the second module to obtain the third feature; inputting the third feature into the second machine learning model, and obtaining the first information through the second machine learning model.

[0173] Optionally, the vehicle's planning and control information includes at least one of the following: the vehicle's planned location, the vehicle's driving strategy, or the vehicle's control information.

[0174] It should be noted that the information interaction and execution process between the modules / units in the regulatory information generation device 1000 are based on the same concept as the various method embodiments corresponding to Figures 1 to 9 in this application. For details, please refer to the description in the method embodiments shown above in this application, which will not be repeated here.

[0175] The following describes a device provided by an embodiment of this application. Please refer to Figure 11. Figure 11 is a schematic diagram of the structure of a device provided by an embodiment of this application. Optionally, the device 1100 performs the functions of the device in the various method embodiments corresponding to Figures 1 to 9.

[0176] Device 1100 includes a memory 1102 and at least one processor 1101. Optionally, device 1100 further includes at least one accelerator 1103. Optionally, processor 1101 implements the methods in the above embodiments by reading program instructions stored in memory 1102; or, processor 1101 reads program instructions stored in memory 1102 and implements the steps executed by the first machine learning model and / or the second machine learning model in the methods in the above embodiments through accelerator 1103; or, processor 1101 may also implement the methods in the above embodiments by reading program instructions stored internally; or, processor 1101 may also read program instructions stored internally and implement the steps executed by the first machine learning model and / or the second machine learning model in the methods in the above embodiments through accelerator 1103.

[0177] When the processor 1101 reads the program instructions stored in the memory 1102 to implement the method in the above embodiments, the memory 1102 stores the program instructions that implement the method provided in the above embodiments of this application.

[0178] Optionally, at least one processor 1101 is one or more CPUs, either a single-core CPU or a multi-core CPU. For example, memory 1102 includes, but is not limited to, random access memory (RAM), read-only memory (ROM), flash memory, or optical memory. Memory 1102 stores program instructions for the operating system. For example, at least one accelerator 1103 may include at least one of the following: GPU, NPU, TPU, ASIC, FPGA, or other types of accelerators. After the program instructions stored in memory 1102 are read by the at least one processor 1101, device 1100 executes the corresponding operations in the foregoing embodiments.

[0179] Optionally, the device 1100 also includes a network interface 1104, which can be a wired interface or a wireless interface. The network interface 1104 is used to perform data transmission and reception in the various method embodiments corresponding to Figures 1 to 7.

[0180] It should be understood that network interface 1104 has the functions of receiving and sending data. The functions of "receiving data" and "sending data" can be integrated into the same transceiver interface, or the functions of "receiving data" and "sending data" can be implemented in different interfaces, which is not limited here. In other words, network interface 1104 may include one or more interfaces for implementing the functions of "receiving data" and "sending data".

[0181] After the processor 1101 reads the program instructions from the memory 1102, other functions that the device 1100 can perform are described in the preceding method embodiments.

[0182] Optionally, the device 1100 also includes a bus 1105, through which the processor 1101 and memory 1102 are typically interconnected, or in other ways.

[0183] The device 1100 provided in this application embodiment is used to execute the execution device and / or the method executed by the execution device in the above-described method embodiments, and to achieve the corresponding beneficial effects. The specific implementation of the device 1100 shown in FIG11 can be referred to the description in the foregoing method embodiments, and will not be repeated here.

[0184] This application also provides a vehicle, as shown in Figure 12. Figure 12 is a structural schematic diagram of a vehicle provided in this application embodiment. The vehicle 100 is configured for fully or partially automated driving mode. For example, the vehicle 100 can control itself while in automated driving mode, and can determine the current state of the vehicle and its surrounding environment through human operation, determine the possible behavior of at least one other vehicle in the surrounding environment, and determine the confidence level corresponding to the probability of other vehicles performing possible behaviors, and control the vehicle 100 based on the determined information. When the vehicle 100 is in automated driving mode, the vehicle 100 can also be set to operate without human interaction.

[0185] Vehicle 100 may include various subsystems, such as a mobility system 102, a sensor system 104, a control system 106, one or more peripheral devices 108, a power supply 110, a computer system 112, and a user interface 116. Optionally, vehicle 100 may include more or fewer subsystems, and each subsystem may include multiple components. Furthermore, each subsystem and component of vehicle 100 may be interconnected via wired or wireless means.

[0186] The mobility system 102 may include components that provide powered motion to the vehicle 100. In one embodiment, the mobility system 102 may include an engine 118, an energy source 119, a transmission 120, and wheels / tires 121.

[0187] Engine 118 can be an internal combustion engine, an electric motor, an air-compressed engine, or other combinations of engines, such as a hybrid engine consisting of a gasoline engine and an electric motor, or a hybrid engine consisting of an internal combustion engine and an air-compressed engine. Engine 118 converts energy source 119 into mechanical energy. Examples of energy source 119 include gasoline, diesel, other petroleum-based fuels, propane, other compressed gas-based fuels, ethanol, solar panels, batteries, and other sources of electricity. Energy source 119 can also provide energy to other systems of vehicle 100. Transmission 120 transmits mechanical power from engine 118 to wheels 121. Transmission 120 may include a gearbox, a differential, and a drive shaft. In one embodiment, transmission 120 may also include other components, such as a clutch. The drive shaft may include one or more axles that can be coupled to one or more wheels 121.

[0188] Sensor system 104 may include several sensors for sensing information about the environment surrounding vehicle 100. For example, sensor system 104 may include a positioning system 122 (which may be a GPS system, a BeiDou system, or another positioning system), an inertial measurement unit (IMU) 124, a radar 126, a laser rangefinder 128, and a camera 130. Sensor system 104 may also include sensors for the internal systems of the monitored vehicle 100 (e.g., an in-vehicle air quality monitor, fuel gauge, oil temperature gauge, etc.). Sensing data from one or more of these sensors can be used to detect objects and their corresponding characteristics (position, shape, orientation, speed, etc.). This detection and identification is a key function for the safe operation of the autonomous vehicle 100.

[0189] The positioning system 122 can be used to estimate the geographical location of the vehicle 100. An IMU 124 is used to sense changes in the position and orientation of the vehicle 100 based on inertial acceleration. In one embodiment, the IMU 124 can be a combination of an accelerometer and a gyroscope. A radar 126 can use radio signals to sense objects in the surrounding environment of the vehicle 100, specifically millimeter-wave radar or lidar. In some embodiments, in addition to sensing objects, the radar 126 can also be used to sense the speed and / or direction of travel of objects. A laser rangefinder 128 can use lasers to sense objects in the environment in which the vehicle 100 is located. In some embodiments, the laser rangefinder 128 may include one or more laser sources, a laser scanner, and one or more detectors, as well as other system components. A camera 130 can be used to capture multiple images of the surrounding environment of the vehicle 100. The camera 130 can be a still camera or a video camera.

[0190] The control system 106 controls the operation of the vehicle 100 and its components. The control system 106 may include various components, including a steering system 132, a throttle 134, a braking unit 136, a computer vision system 140, a trajectory control system 142, and an obstacle avoidance system 144.

[0191] The steering system 132 is operable to adjust the forward direction of the vehicle 100. For example, in one embodiment, it may be a steering wheel system. The throttle 134 controls the operating speed of the engine 118 and thus the speed of the vehicle 100. The braking unit 136 controls the deceleration of the vehicle 100. The braking unit 136 may use friction to slow down the wheels 121. In other embodiments, the braking unit 136 may convert the kinetic energy of the wheels 121 into electrical current. The braking unit 136 may also take other forms to slow down the rotational speed of the wheels 121 to control the speed of the vehicle 100. The computer vision system 140 is operable to process and analyze images captured by the camera 130 to identify objects and / or features in the environment surrounding the vehicle 100. The objects and / or features may include traffic signals, road boundaries, and obstacles. The computer vision system 140 may use object recognition algorithms, Structure from Motion (SFM) algorithms, video tracking, and other computer vision techniques. In some embodiments, the computer vision system 140 may be used to map the environment, track objects, estimate the speed of objects, etc. The route control system 142 is used to determine the driving route and speed of the vehicle 100. In some embodiments, the route control system 142 may include a lateral planning module 1421 and a longitudinal planning module 1422, which are respectively used to combine data from the obstacle avoidance system 144, GPS 122, and one or more predetermined maps to determine the driving route and speed for the vehicle 100. The obstacle avoidance system 144 is used to identify, evaluate, and avoid or otherwise traverse obstacles in the environment of the vehicle 100, which may specifically be physical obstacles and virtual moving bodies that may collide with the vehicle 100. In one example, the control system 106 may add or alternatively include components other than those shown and described. Alternatively, some of the components shown above may be reduced.

[0192] Vehicle 100 interacts with external sensors, other vehicles, other computer systems, or users via peripheral device 108. Peripheral device 108 may include wireless communication system 146, on-board computer 148, microphone 150, and / or speaker 152. In some embodiments, peripheral device 108 provides a means for a user of vehicle 100 to interact with user interface 116. For example, on-board computer 148 may provide information to a user of vehicle 100. User interface 116 may also operate on-board computer 148 to receive user input. On-board computer 148 may be operated via a touchscreen. In other cases, peripheral device 108 may provide a means for vehicle 100 to communicate with other devices located within the vehicle. For example, microphone 150 may receive audio (e.g., voice commands or other audio input) from a user of vehicle 100. Similarly, speaker 152 may output audio to a user of vehicle 100. Wireless communication system 146 may communicate wirelessly with one or more devices, either directly or via a communication network. For example, the wireless communication system 146 may use 3G cellular communication, such as CDMA, EVDO, GSM / GPRS, or 4G cellular communication, such as LTE, or 5G cellular communication. The wireless communication system 146 may utilize a wireless local area network (WLAN) for communication. In some embodiments, the wireless communication system 146 may utilize an infrared link, Bluetooth, or ZigBee to communicate directly with the device. Other wireless protocols, such as various vehicle communication systems, may also be used. For example, the wireless communication system 146 may include one or more dedicated short-range communications (DSRC) devices that can enable public and / or private data communication between the vehicle and / or a roadside station.

[0193] Power source 110 can provide power to various components of vehicle 100. In one embodiment, power source 110 can be a rechargeable lithium-ion or lead-acid battery. One or more such battery packs can be configured to provide power to various components of vehicle 100. In some embodiments, power source 110 and energy source 119 can be implemented together, as is the case in some fully electric vehicles.

[0194] Some or all of the functions of vehicle 100 are controlled by computer system 112. Computer system 112 may include at least one processor 113, which executes program instructions 115 stored in a non-transitory computer-readable medium such as memory 114. Computer system 112 may also be multiple computing devices that control individual components or subsystems of vehicle 100 in a distributed manner. Processor 113 may include any conventional processor, such as a commercially available central processing unit (CPU). Alternatively, processor 113 may also include a dedicated device such as a GPU, NPU, TPU, ASIC, FPGA, or other hardware-based processor. Although FIG12 functionally illustrates the processor, memory, and other components of computer system 112 in the same block, those skilled in the art will understand that the processor or memory may actually include multiple processors or memories not stored in the same physical housing. For example, memory 114 may be a hard disk drive or other storage medium located in a housing different from that of computer system 112. Therefore, references to processor 113 or memory 114 will be understood to include references to a collection of processors or memories that may or may not operate in parallel. Unlike using a single processor to perform the steps described herein, some components, such as steering and deceleration components, can each have their own processor that performs only calculations related to the component's specific function.

[0195] In all the aspects described herein, processor 113 may be located remotely from vehicle 100 and may communicate wirelessly with vehicle 100. In other aspects, some of the processes described herein are executed on processor 113 located within vehicle 100, while others are executed by remote processor 113, including taking the necessary steps to perform a single operation.

[0196] In some embodiments, memory 114 may contain instructions 115 (e.g., program logic) that can be executed by processor 113 to perform various functions of vehicle 100, including those described above. Memory 114 may also contain additional program instructions, including instructions for sending data to, receiving data from, interacting with, and / or controlling one or more of the mobility system 102, sensor system 104, control system 106, and peripheral devices 108. In addition to instructions 115, memory 114 may also store data such as road maps, route information, vehicle position, direction, speed, and other such vehicle data, as well as other information. This information may be used by vehicle 100 and computer system 112 during operation of vehicle 100 in autonomous, semi-autonomous, and / or manual modes. A user interface 116 is provided to or receives information from a user of vehicle 100. Optionally, user interface 116 may include one or more input / output devices within the set of peripheral devices 108, such as wireless communication system 146, on-board computer 148, microphone 150, and speaker 152.

[0197] Computer system 112 can control the functions of vehicle 100 based on input received from various subsystems (e.g., driving system 102, sensor system 104, and control system 106) and from user interface 116. For example, computer system 112 can utilize input from control system 106 to control steering system 132 to avoid obstacles detected by sensor system 104 and obstacle avoidance system 144. In some embodiments, computer system 112 is operable to provide control over many aspects of vehicle 100 and its subsystems.

[0198] Alternatively, one or more of these components may be installed separately from or associated with vehicle 100. For example, memory 114 may exist partially or completely separately from vehicle 100. The components may be communicatively coupled together in a wired and / or wireless manner.

[0199] Optionally, the above components are merely examples. In practical applications, components in each of the above modules may be added or removed according to actual needs. Figure 12 should not be construed as a limitation on the embodiments of this application. A vehicle traveling on a road, such as vehicle 100 above, can identify objects in its surrounding environment to determine adjustments to its current speed. These objects can be other vehicles, traffic control equipment, or other types of objects. In some examples, each identified object can be considered independently, and based on the object's individual characteristics, such as its current speed, acceleration, and distance from the vehicle, the speed adjustment to be made by the vehicle can be determined.

[0200] Optionally, vehicle 100 or computing devices associated with vehicle 100, such as computer system 112, computer vision system 140, and memory 114 as shown in Figure 12, can predict the behavior of the identified objects based on the characteristics of the identified objects and the state of the surrounding environment (e.g., traffic, rain, ice on the road, etc.). Optionally, each identified object depends on the behavior of each other, so all identified objects can also be considered together to predict the behavior of a single identified object. Vehicle 100 can adjust its speed based on the predicted behavior of the identified objects. In other words, vehicle 100 can determine what steady state the vehicle will need to adjust to (e.g., accelerate, decelerate, or stop) based on the predicted behavior of the objects. In this process, other factors can also be considered in determining the speed of vehicle 100, such as the lateral position of vehicle 100 in the road, the curvature of the road, the proximity of static and dynamic objects, etc. In addition to providing program instructions to adjust the speed of the vehicle, the computing device may also provide program instructions to modify the steering angle of the vehicle 100 so that the vehicle 100 follows a given trajectory and / or maintains a safe lateral and longitudinal distance from objects near the vehicle 100 (e.g., cars in adjacent lanes on the road).

[0201] In this embodiment, the processor 113 in the vehicle 100 is used to execute the execution device and / or the method executed by the execution device in the embodiments corresponding to FIG1 to FIG9. It should be noted that the specific manner in which the processor 113 executes the aforementioned steps is based on the same concept as the various method embodiments corresponding to FIG1 to FIG9 in this application, and the resulting technical effects are the same as those in the various method embodiments corresponding to FIG1 to FIG9 in this application. For details, please refer to the description in the method embodiments shown above in this application, which will not be repeated here.

[0202] This application also provides a computer-readable storage medium storing a program that, when run on a computer, causes the computer to perform the steps executed by the execution device in the methods described in the embodiments shown in Figures 1 to 9.

[0203] This application also provides a computer program product, which includes a program that, when run on a computer, causes the computer to perform the steps performed by the execution device in the methods described in the embodiments shown in Figures 1 to 9.

[0204] This application also provides a circuit system including a processing circuit configured to perform the steps executed by the execution device in the methods described in the embodiments shown in Figures 1 to 9 above.

[0205] The execution device or control information generation apparatus provided in this application embodiment can specifically be a chip. The chip includes a processing unit, such as a processor. Optionally, the chip also includes a communication unit, such as an input / output interface, pins, or circuits. The processing unit can execute computer execution instructions stored in the storage unit to cause the chip to execute the methods described in the embodiments shown in Figures 1 to 9. Optionally, the storage unit is a storage unit within the chip, such as a register or cache. The storage unit can also be a storage unit located outside the chip within the wireless access device, such as a read-only memory (ROM) or other types of static storage devices capable of storing static information and instructions, such as random access memory (RAM).

[0206] The processor mentioned above can be a general-purpose central processing unit, microprocessor, GPU, NPU, TPU, ASIC, FPGA, or one or more integrated circuits used to control the execution of the program in the first aspect of the above method.

[0207] It should also be noted that the device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. In addition, in the device embodiment drawings provided in this application, the connection relationship between modules indicates that they have a communication connection, which can be implemented as one or more communication buses or signal lines.

[0208] Through the above description of the embodiments, those skilled in the art can clearly understand that this application can be implemented by means of software plus necessary general-purpose hardware, or it can be implemented by special-purpose hardware including application-specific integrated circuits, special-purpose CLUs, special-purpose memory, special-purpose components, etc. Generally, any function performed by a computer program can be easily implemented by corresponding hardware, and the specific hardware structure used to implement the same function can also be diverse, such as analog circuits, digital circuits, or special-purpose circuits. However, for this application, software program implementation is more often the preferred implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a readable storage medium, such as a computer floppy disk, USB flash drive, mobile hard disk, ROM, RAM, magnetic disk, or optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments of this application.

[0209] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, as a computer program product.

[0210] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions may be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium may be any available medium that a computer can store or a data storage device such as a server or data center that integrates one or more available media. The available medium may be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state disk (SSD)).

Claims

1. A method for generating regulatory information, characterized in that, The method includes: Environmental information is input into a first machine learning model, and the first machine learning model extracts features from the environmental information to obtain a first feature. The environmental information includes images and / or point cloud data corresponding to the traffic environment. Based on the first feature, first information is obtained through the second machine learning model, wherein the number of parameters of the second machine learning model is greater than the number of parameters of the first machine learning model. Based on the first feature and the first information, vehicle control information is generated through the first machine learning model.

2. The method according to claim 1, characterized in that, The second machine learning model is based on a Large Language Model (LLM), where the first information includes features of at least one first token generated by the LLM and / or the at least one first token generated by the LLM, the features of which are generated by a neural network layer in the LLM; or, The second machine learning model is used to update the first feature, and the first information includes the updated first feature.

3. The method according to claim 2, characterized in that, The second machine learning model is the LLM, and the at least one first token includes the first N tokens generated by the second machine learning model in an autoregressive manner, where N is an integer greater than or equal to 1, and the value of N is preset.

4. The method according to claim 2, characterized in that, The first information includes features of the at least one first token, and the generation of vehicle control information based on the first features and the first information using the first machine learning model includes: The first module processes the features of the at least one first token to obtain the second feature; Based on the first feature and the second feature, vehicle control information is generated through the first machine learning model.

5. The method according to claim 4, characterized in that, The first module includes at least one sub-module corresponding to at least one task, wherein the at least one task includes at least one of the following: object detection, 3D modeling of the traffic environment, text recognition in the traffic environment, determining the spatial relationship between different objects in the traffic environment, intention judgment and / or behavior prediction of obstacles in the traffic environment, determining the strategy adopted when interacting with obstacles in the traffic environment, identifying objects with risks in the traffic environment, determining the vehicle's driving strategy, determining the speed limit of the vehicle when driving in the traffic environment, performing path planning for the vehicle, or determining the vehicle's control information.

6. The method according to claim 3, characterized in that, The first N tokens do not correspond to the text. During the training phase of the second machine learning model, the second machine learning model is also used to generate a second token corresponding to the predicted text. The second machine learning model is trained using a loss function, which indicates the similarity between the second token and the third token. The third token includes tokens corresponding to the expected text.

7. The method according to any one of claims 1 to 6, characterized in that, The step of obtaining the first information based on the first feature through the second machine learning model includes: in a first case, obtaining the first information based on the first feature through the second machine learning model; The method further includes: in the second case, generating vehicle control information through the first machine learning model based on the first feature and preset information.

8. The method according to claim 7, characterized in that, The frequency of generating vehicle control information through the first machine learning model is higher than the frequency of obtaining the first information through the second machine learning model.

9. The method according to any one of claims 1 to 6, characterized in that, The step of obtaining the first information based on the first feature through the second machine learning model includes: The first feature is input into the second module, and the third feature is obtained by compressing the first feature through the second module; The third feature is input into the second machine learning model, and the first information is obtained through the second machine learning model.

10. The method according to any one of claims 1 to 6, characterized in that, The vehicle's planning and control information includes at least one of the following: the vehicle's planned location, the vehicle's driving strategy, or the vehicle's control information.

11. A device for generating regulatory information, characterized in that, The device includes: An input module is used to input environmental information into a first machine learning model, and to extract features from the environmental information through the first machine learning model to obtain a first feature. The environmental information includes images and / or point cloud data corresponding to the traffic environment. The processing module is used to obtain first information based on the first feature through the second machine learning model, wherein the number of parameters of the second machine learning model is greater than the number of parameters of the first machine learning model. The generation module is used to generate vehicle control information based on the first feature and the first information through the first machine learning model.

12. The apparatus according to claim 11, characterized in that, The second machine learning model is based on a Large Language Model (LLM), where the first information includes features of at least one first token generated by the LLM and / or the at least one first token generated by the LLM, the features of which are generated by a neural network layer in the LLM; or, The second machine learning model is used to update the first feature, and the first information includes the updated first feature.

13. The apparatus according to claim 12, characterized in that, The second machine learning model is the LLM, and the at least one first token includes the first N tokens generated by the second machine learning model using an autoregressive approach, where N is an integer greater than or equal to 11, and the value of N is preset.

14. The apparatus according to claim 12, characterized in that, The first information includes the characteristics of the at least one first token, and the generation module is specifically used for: The first module processes the features of the at least one first token to obtain the second feature; Based on the first feature and the second feature, vehicle control information is generated through the first machine learning model.

15. The apparatus according to claim 14, characterized in that, The first module includes at least one sub-module corresponding to at least one task, wherein the at least one task includes at least one of the following: object detection, 3D modeling of the traffic environment, text recognition in the traffic environment, determining the spatial relationship between different objects in the traffic environment, intention judgment and / or behavior prediction of obstacles in the traffic environment, determining the strategy adopted when interacting with obstacles in the traffic environment, identifying objects with risks in the traffic environment, determining the vehicle's driving strategy, determining the speed limit of the vehicle when driving in the traffic environment, performing path planning for the vehicle, or determining the vehicle's control information.

16. The apparatus according to claim 13, characterized in that, The first N tokens do not correspond to the text. During the training phase of the second machine learning model, the second machine learning model is also used to generate a second token corresponding to the predicted text. The second machine learning model is trained using a loss function, which indicates the similarity between the second token and the third token. The third token includes tokens corresponding to the expected text.

17. The apparatus according to any one of claims 11 to 16, characterized in that, The processing module is specifically used, in the first case, to obtain first information based on the first feature through the second machine learning model; The generation module is further configured to, in the second case, generate vehicle control information based on the first feature and preset information using the first machine learning model.

18. The apparatus according to claim 17, characterized in that, The frequency of generating vehicle control information through the first machine learning model is higher than the frequency of obtaining the first information through the second machine learning model.

19. The apparatus according to any one of claims 11 to 16, characterized in that, The processing module is specifically used for: The first feature is input into the second module, and the third feature is obtained by compressing the first feature through the second module; The third feature is input into the second machine learning model, and the first information is obtained through the second machine learning model.

20. The apparatus according to any one of claims 11 to 16, characterized in that, The vehicle's planning and control information includes at least one of the following: the vehicle's planned location, the vehicle's driving strategy, or the vehicle's control information.

21. A device, characterized in that, The method includes a processor coupled to a memory storing program instructions that, when executed by the processor, implement the method of any one of claims 1 to 10.

22. A vehicle, characterized in that, The method includes a processor coupled to a memory storing program instructions that, when executed by the processor, implement the method of any one of claims 1 to 10.

23. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a program that, when run on a computer, causes the computer to perform the method as described in any one of claims 1 to 10.

24. A computer program product, characterized in that, The computer program product includes a program that, when run on a computer, causes the computer to perform the method as described in any one of claims 1 to 10.

25. A chip, characterized in that, The chip includes a processor for performing the steps of the method according to any one of claims 1 to 10.

Citation Information

Patent Citations

  • Automatic driving control method and device, electronic equipment, vehicle and computer program product

    CN117962928A

  • Automatic driving multi-mode perception decision-making method and device based on large language model

    CN118115969A

  • Generative large language model training method for automatic driving and storage medium

    CN118227761A

  • Vehicle control method and device, vehicle and computer readable storage medium

    CN118419067A

  • Domain-customizable models for conversational ai systems and applications

    US20240193445A1