Information processing method and related device
By introducing a deep learning model that processes text information in autonomous driving vehicles, the problem of poor interpretability in the model operation process is solved, and a more transparent and intuitive understanding of bicycle behavior is achieved.
Patent Information
- Application Number
- PCT/CN2024/124108
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-11-24
- Filing Date
- 2024-10-11
- Publication Date
- 2025-05-30
AI Technical Summary
The operation process of deep learning models in autonomous vehicles is poorly interpretable, which makes it difficult for users to understand the decisions and behaviors of the vehicles.
By introducing a deep learning model that can process text information, text information related to traffic scenes around the bicycle is input to the model, and prediction information about bicycle behavior decisions, trajectory planning or control is generated.
It improves the interpretability of the operation process of the deep learning model, makes the decisions and behaviors of autonomous vehicles more transparent, and users can understand the behaviors of vehicles more intuitively.
Smart Images

Figure CN2024124108_30052025_PF_FP_ABST
Abstract
Description
Information processing method and related equipment
[0001] This application claims priority to the Chinese patent application filed with the State Intellectual Property Office on November 24, 2023, with application number 202311594214.1 and invention name “An information processing method and related equipment”, the entire contents of which are incorporated by reference into this application. Technical Field
[0002] The present application relates to the field of artificial intelligence, and in particular to an information processing method and related equipment. Background Art
[0003] Artificial Intelligence (AI) is the theory, methods, techniques, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, to perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. In other words, AI is a branch of computer science that seeks to understand the essence of intelligence and develop new intelligent machines that can respond in a manner similar to human intelligence. AI also studies the design principles and implementation methods of various intelligent machines, enabling them to possess the capabilities of perception, reasoning, and decision-making. Research in the field of AI includes robotics, natural language processing, computer vision, decision-making and reasoning, human-computer interaction, recommendation and search, and basic AI theory.
[0004] The field of autonomous driving is an application field of one scenario in the field of artificial intelligence. For example, it can obtain environmental information around the vehicle, generate prediction information through one or more neural networks, and then make decisions on the vehicle's behavior and plan the trajectory of the vehicle based on the aforementioned prediction information; for example, when there is a car illegally parked in the road ahead, decide whether to change lanes, and plan when and how to change lanes.
[0005] However, since what is obtained is the environmental information around the ego vehicle, the prediction information obtained through the neural network is directly used to determine the behavior of the ego vehicle, and the entire operation process of the neural network has poor interpretability.
[0006] Summary of the Invention
[0007] The present application provides an information processing method and related equipment, which improves the interpretability of the operation process of the deep learning model, that is, makes the decision-making, trajectory planning or control process of the autonomous driving vehicle more transparent, so that users can understand the behavior of the autonomous driving vehicle more intuitively.
[0008] This application provides the following technical solutions:
[0009] In a first aspect, the present application provides an information processing method that can be used in the field of autonomous driving within the field of artificial intelligence. In this method, a first device can input first information corresponding to a traffic scene surrounding a vehicle into a deep learning model that has been trained, thereby obtaining second information corresponding to the first information. The first information includes first text information, and the second information is obtained based on the deep learning model that has been trained, and the second information corresponds to any of the following tasks: making decisions about the vehicle's behavior, planning the vehicle's trajectory, or controlling the vehicle.
[0010] The first device may be specifically a vehicle or a cloud server. For example, in one scenario, the first device is a vehicle. If the deep learning model is deployed in the vehicle, the first device inputs first information corresponding to the traffic scene around the vehicle to the deep learning model that has been trained. This may include: the first device inputs the first information corresponding to the traffic scene around the vehicle to the locally deployed deep learning model. The first device obtains second information corresponding to the first information. This may include: the first device determines second information corresponding to the first information based on first prediction information generated by the deep learning model.
[0011] In another scenario, the first device is a vehicle, and if the trained deep learning model is deployed on a cloud server, the first device inputting first information corresponding to the traffic scene around the vehicle into the trained deep learning model may include: the first device sending the first information corresponding to the traffic scene around the vehicle to the cloud server where the deep learning model is deployed. The first device obtaining second information corresponding to the first information may include: the first device receiving first prediction information sent by the cloud server, and then determining second information corresponding to the first information based on the first prediction information.
[0012] In another scenario, where the first device is a cloud server that deploys a deep learning model, the first device inputting first information corresponding to the traffic scene around the vehicle into the trained deep learning model may include: after receiving the first information sent by the vehicle, the first device inputting the first information corresponding to the traffic scene around the vehicle into the locally deployed deep learning model. The first device obtaining second information corresponding to the first information may include: the first device sending first prediction information to the vehicle, the first prediction information being used by the vehicle to determine the second information.
[0013] Among them, the first information includes first text information, and the first text information can correspond to a task performed by the deep learning model (for the convenience of description, hereinafter referred to as the "first task").
[0014] For example, "making decisions on the behavior of the vehicle" can be understood as the second information being used to indicate what behavior the vehicle should perform next; "planning the trajectory of the vehicle" can be understood as the second information may include the next trajectory information of the vehicle; and "controlling the vehicle" can be understood as the second information may include the vehicle control instructions of the vehicle.
[0015] In this implementation, a deep learning model capable of processing text information is introduced into the field of autonomous driving. Based on the second information generated by the deep learning model that processes text information, decisions are made on the behavior of the vehicle, the trajectory of the vehicle is planned, or the vehicle is controlled. Since the input information of the deep learning model provided by this application contains text information that is easy for users to understand, the interpretability of the operation process of the deep learning model is improved, and the decision-making, trajectory planning or control process of the autonomous driving vehicle is made more transparent, so that users can understand the behavior of the autonomous driving vehicle more intuitively.
[0016] In one possible implementation, the first text information includes information about preset text corresponding to the first task executed by the deep learning model. For example, the vehicle may be equipped with preset text corresponding to each of the at least one second task. The vehicle retrieves the preset text corresponding to the first task from the preset text corresponding to the at least one second task, thereby obtaining the first text information. Alternatively, the first text information includes text information input by the user.
[0017] Exemplarily, the first text information may include initial feature information of the preset text (or text entered by the user) obtained after feature extraction of the preset text (or text entered by the user). The aforementioned "feature extraction" can also be understood as "initial encoding", or the aforementioned "feature extraction" can also be understood as vectorization (embedding); exemplarily, the first text information can be specifically expressed in the form of a token.
[0018] In this implementation, two situations in which the first text information may include information are listed, which improves the implementation flexibility of this solution; when the first text information includes preset text information corresponding to the first task performed by the deep learning model, it is beneficial to quickly determine the first text information after determining the first task, so as to improve the efficiency of obtaining the second information; when the first text information includes text information input by the user, that is, questions can be answered based on the user's questions, which is beneficial to improve the user stickiness of this solution and makes it more convenient for users to understand the behavior of the vehicle during the automatic driving process.
[0019] In one possible implementation, in this method, the first device may further input a first question corresponding to a traffic scenario into the deep learning model. The first question is used to obtain a first answer corresponding to the first question, and the first answer is obtained by the deep learning model. The first information is the second question, the second information is the second answer corresponding to the second question, and the first text information contains information in the first answer.
[0020] In this implementation, a first answer to the first question is first obtained, and then a second question is generated based on the first answer, and then the answer to the second question is obtained. The first question and the second question correspond to the same traffic scenario, that is, the questions to the deep learning model are asked in a progressive manner, which is conducive to reducing the difficulty of the deep learning model in answering the second question, and is also conducive to improving the ability of the deep learning model in dealing with problems corresponding to complex traffic scenarios; in addition, the use of a progressive questioning method is also conducive to incorporating the logical thinking of step-by-step reasoning into the deep learning model used in the field of autonomous driving, thereby improving the human-likeness of the second answer finally obtained.
[0021] In one possible implementation, the first information further includes first feature information, where the first feature information includes feature information of the environment surrounding the vehicle. The environment information includes physical attribute information of objects surrounding the vehicle. For example, the objects surrounding the vehicle may include dynamic obstacles surrounding the vehicle, static obstacles surrounding the vehicle, traffic markings surrounding the vehicle, traffic signs surrounding the vehicle, or other types of objects.
[0022] Optionally, the first characteristic information also includes characteristic information of the vehicle's driving behavior information and / or characteristic information of navigation information.
[0023] In this implementation, the first information includes not only the first text information, but also the first feature information. The first feature information includes the feature information of the environmental information around the vehicle, so that the deep learning model can fully combine the feature information of the environmental information around the vehicle to determine the second information, which is conducive to improving the rationality of the second information output by the deep learning model.
[0024] In one possible implementation, the first feature information is obtained based on a feature extraction network (hereinafter referred to as the "first feature extraction network" for the convenience of description), wherein the first feature extraction network belongs to a first neural network, and the first neural network is used to perform at least two of the following tasks: trajectory prediction of objects around the vehicle, decision-making on the vehicle's behavior, trajectory planning for the vehicle, prediction of the vehicle's speed range, or control of the vehicle.
[0025] In this implementation, since the first neural network is used to perform at least two of the multiple tasks of predicting the trajectory of objects around the vehicle, making decisions on the vehicle's behavior, planning the trajectory of the vehicle, predicting the speed range of the vehicle, or controlling the vehicle, the first feature information needs to cover richer information, that is, the first information containing the first feature information will cover more information, which is conducive to the deep learning model obtaining more information, and thus is conducive to improving the accuracy of the information output by the deep learning model.
[0026] In one possible implementation, the vehicle can input the environmental information around the vehicle into the first feature extraction network. Optionally, the vehicle's driving behavior information and / or navigation information can also be input into the first feature extraction network to obtain the third feature information generated by the first feature extraction network, and then obtain the first feature information.
[0027] For example, in one case, "third feature information" and "first feature information" have the same meaning, that is, the vehicle can directly use the third feature information generated by the first feature extraction network as the first feature information.
[0028] In another case, after obtaining the third feature information, the vehicle can also input the third feature information into the second feature extraction network, and update the features of the third feature information through the second feature extraction network to obtain the first feature information. Exemplarily, the third feature information can be specifically expressed in the form of a feature map, and the first feature information can be specifically expressed in the form of a token, that is, the third feature information in the form of a feature map can be converted into the first feature information in the form of a token through the second feature extraction network. Since the deep learning model is a machine learning model that processes text information, after obtaining the first feature information in the form of a token and then inputting it into the deep learning model, the deep learning model can more easily understand the first feature information, reducing the difficulty of the deep learning model in understanding the first feature information, which is conducive to improving the accuracy of the information output by the deep learning model.
[0029] In one possible implementation, the first information is a second question, and the second information is a second answer corresponding to the second question, wherein the second answer is included in a preset answer set corresponding to the second question, and the second answer is obtained from at least one alternative answer generated from a deep learning model based on the preset answer set. Exemplarily, a preset answer set may be deployed in the execution device, and the preset answer set includes one or more answers. After the execution device obtains the first prediction information, it may determine the intersection between at least one alternative answer indicated by the first prediction information and the preset answer set, wherein the aforementioned intersection includes at least one first alternative answer; the execution device may determine the second answer from at least one first alternative answer based on the first probability value corresponding to each first alternative answer in the at least one first alternative answer, and the first probability value corresponding to the second answer is the largest among the at least one first alternative answer.
[0030] In this implementation, since the deep learning model may sometimes give irrelevant answers or speak nonsense, that is, at least one alternative answer determined by the deep learning model may contain an answer that is completely irrelevant to the question, in order to avoid the occurrence of the aforementioned problem, a second preset answer set corresponding to the second question can be deployed, and the second preset answer set corresponding to the second question includes at least one answer that is related to the second question. The second answer corresponding to the second question is finally determined based on the aforementioned second preset answer set and at least one alternative answer determined based on the deep learning model, which is conducive to avoiding the problem that the final second answer is irrelevant to the second question, thereby improving the accuracy of the second answer obtained by the deep learning model.
[0031] In one possible implementation, the first information is a second question, the second information is a second answer corresponding to the second question, the second answer is obtained through a classification network, the second feature information includes feature information of a start flag [CLS] bit in the feature information of the second answer, and the classification network is used to determine the second answer corresponding to the second feature information from at least one preset answer. Exemplarily, a trained classification network may be deployed in the execution device. After obtaining first prediction information corresponding to the second question generated by a deep learning model, the execution device may determine feature information of a second alternative answer from the first prediction information, wherein the first probability value of the second alternative answer is the highest among the at least one alternative answer indicated by the first prediction information. The execution device may obtain second feature information from the feature information of the second alternative answer, input the second feature information into the classification network, and obtain the second answer generated by the classification network.
[0032] In this implementation, since the characteristic information of the starting flag [CLS] bit (that is, the second characteristic information) can represent the overall meaning of the alternative answers, it is reasonable to use the characteristic information of the starting flag [CLS] bit to replace the characteristic information of the alternative answers. The classification network is used to determine a second answer belonging to the second characteristic information from at least one preset answer, and the second answer is limited to at least one preset answer, which is conducive to avoiding the problem that the final second answer is irrelevant to the second question, thereby improving the accuracy of the second answer obtained by the deep learning model.
[0033] In one possible implementation, the first information is the second question, and the second information is the second answer corresponding to the second question. The first device inputs the first information corresponding to the traffic scene around the vehicle into the deep learning model, which may include: the first device inputs multiple second questions into the deep learning model; illustratively, the multiple second questions correspond to the same task, that is, the multiple second questions have similar semantics, and different second questions in the multiple second questions carry different prompts. The first device obtains the second information corresponding to the first information, which may include: the first device obtains multiple reference answers corresponding one-to-one to the multiple second questions, and then determines the second answer based on the multiple reference answers, wherein the multiple reference answers are all obtained through the deep learning model.
[0034] In this implementation, multiple second questions are input into the deep learning model, and different second questions carry different prompts. Multiple reference answers are obtained by using multiple different prompts, and then the final second answer can be obtained from the multiple reference answers, which improves the rigor of obtaining the final second answer and is conducive to improving the accuracy of the final second questions.
[0035] In one possible implementation, the first device determines the second answer based on multiple reference answers, which may include: the first device adopts a majority voting strategy to determine the second answer based on the multiple reference answers, that is, the final second answer is included in the multiple reference answers.
[0036] In a second aspect, the present application provides an information processing method that can be used in the field of autonomous driving in the field of artificial intelligence. In this method, a first device inputs a first question into a deep learning model to obtain a first answer corresponding to the first question; a second question is determined based on the first answer, and the second question carries information from the first answer. The first device inputs a second question into the deep learning model to obtain a second answer corresponding to the second question, wherein the first question and the second question correspond to the same traffic scenario.
[0037] Exemplarily, the first question may include a prompt. Optionally, the first question may also include first characteristic information, where the first characteristic information includes characteristic information of the environment surrounding the vehicle, where the environment information includes physical attribute information of objects surrounding the vehicle. Optionally, the first characteristic information also includes characteristic information of the vehicle's driving behavior and / or characteristic information of navigation information.
[0038] Optionally, the first and second questions can be used to obtain different levels of information from the target traffic scene. For example, the first question can be used to obtain basic attribute information and semantic information in the traffic scene; for another example, the first question can be used to obtain the behavior of objects in the target traffic scene or the relationship between different objects.
[0039] In one possible implementation, the second answer corresponds to any of the following tasks: identifying risky obstacles around the vehicle, identifying the behavior of objects around the vehicle, predicting the behavior of objects around the vehicle, predicting the trajectory of objects around the vehicle, making decisions about the vehicle's behavior, planning the trajectory of the vehicle, or controlling the vehicle. In this embodiment, multiple tasks are provided that can be performed using a deep learning model, increasing the implementation flexibility of this solution.
[0040] In the second aspect, the first device is also used to execute the steps performed by the first device in the first aspect and various possible implementation methods of the first aspect. For the meaning of the nouns in the second aspect of this application and the various possible implementation methods of the second aspect, the specific implementation methods of the steps, and the beneficial effects brought about by each possible implementation method, please refer to the description of the various possible implementation methods in the first aspect, and will not be repeated here one by one.
[0041] On the third aspect, the present application provides an information processing method that can be used in the field of autonomous driving in the field of artificial intelligence. The method is used to train a deep learning model, and the deep learning model includes at least two training stages, and the at least two training stages include a first training stage and a second training stage. In this method, in the first training stage, the training device inputs a third question to the deep learning model to obtain a predicted answer corresponding to the third question; in the second training stage, the training device inputs a fourth question to the deep learning model to obtain a predicted answer corresponding to the fourth question. Wherein, a first loss function is used when training the deep learning model. In the first training stage, the first loss function indicates the similarity between the predicted answer corresponding to the third question and the expected answer. In the second training stage, the first loss function indicates the similarity between the predicted answer corresponding to the fourth question and the expected answer. The third question and the fourth question are both related to traffic scenes. The third question and the fourth question are used to obtain information at different levels.
[0042] In this implementation, at least two training stages are used to train the deep learning model. The aforementioned at least two training stages respectively use the third question and the fourth question. The third question and the fourth question are used to obtain information at different levels from the traffic scene. That is, the training process of the deep learning model adopts a layer-by-layer progressive training method, which is conducive to enabling the deep learning model to learn logical thinking of step-by-step reasoning, thereby improving the human-likeness of the deep learning model, and also helping the trained deep learning model to be more reasonable when performing autonomous driving related tasks.
[0043] In a possible implementation, the third question includes first feature information, where the first feature information includes feature information of environmental information around the vehicle, and the environmental information includes physical attribute information of objects around the vehicle.
[0044] In one possible implementation, before the training device inputs the third question into the deep learning model, the method also includes: the training device inputs the environmental information into the feature extraction network to obtain second feature information generated by the feature extraction network, and the second feature information is used to obtain the first feature information; the first feature information is input into the feature processing network to obtain prediction information generated by the feature processing network, the feature extraction network and the feature processing network belong to the same neural network, and the neural network is used to perform at least the following multiple tasks: trajectory prediction of objects around the vehicle, decision-making on the behavior of the vehicle, trajectory planning of the vehicle, prediction of the speed range of the vehicle, or control of the vehicle; wherein, the training stage of the deep learning model includes using a first loss function and a second loss function to train the deep learning model and the neural network, and the second loss function indicates the similarity between the prediction information corresponding to the environmental information and the second expected information.
[0045] In the third aspect, the training device is also used to execute the steps performed by the first device in the first aspect and various possible implementations of the first aspect. For the meaning of the nouns in the third aspect of this application and various possible implementations of the third aspect, the specific implementation methods of the steps, and the beneficial effects brought about by each possible implementation method, please refer to the description of the various possible implementation methods in the first aspect, and will not be repeated here.
[0046] Fourthly, the present application provides an information processing device that can be used in the field of autonomous driving in the field of artificial intelligence. The information processing device includes: an input module for inputting first information corresponding to the traffic scene around the vehicle into the deep learning model, where the first information includes first text information; an acquisition module for acquiring second information corresponding to the first information, where the second information is obtained based on the deep learning model, and the second information corresponds to any of the following tasks: making decisions on the behavior of the vehicle, planning the trajectory of the vehicle, or controlling the vehicle.
[0047] In the fourth aspect, the information processing device is also used to execute the steps performed by the first device in the first aspect and various possible implementation methods of the first aspect. For the meaning of the nouns in the fourth aspect and various possible implementation methods of the fourth aspect of this application, the specific implementation methods of the steps, and the beneficial effects brought about by each possible implementation method, please refer to the description of the various possible implementation methods in the first aspect, and will not be repeated here one by one.
[0048] In a fifth aspect, the present application provides an information processing device that can be used in the field of autonomous driving in the field of artificial intelligence. The information processing device includes: an input module for inputting a first question into a deep learning model to obtain a first answer corresponding to the first question; a determination module for determining a second question based on the first answer, wherein the second question carries information in the first answer; the input module is also used to input a second question into the deep learning model to obtain a second answer corresponding to the second question, wherein the first question and the second question correspond to the same traffic scenario.
[0049] In the fifth aspect, the information processing device is also used to execute the steps performed by the first device in the first aspect and various possible implementation methods of the first aspect. For the meaning of the nouns in the fifth aspect and various possible implementation methods of the fifth aspect of this application, the specific implementation methods of the steps, and the beneficial effects brought about by each possible implementation method, you can refer to the description of the various possible implementation methods in the first aspect, and will not repeat them one by one here.
[0050] In a sixth aspect, the present application provides an information processing device that can be used in the field of autonomous driving in the field of artificial intelligence. The device is used to train a deep learning model. The deep learning model includes at least two training stages, and the at least two training stages include a first training stage and a second training stage. The device includes: an input module, which is used to input a third question into the deep learning model in the first training stage to obtain a predicted answer corresponding to the third question; the input module is also used to input a sixth question into the deep learning model in the second training stage to obtain a predicted answer corresponding to the sixth question; wherein, a first loss function is used when training the deep learning model. In the first training stage, the first loss function indicates the similarity between the predicted answer corresponding to the third question and the expected answer. In the second training stage, the first loss function indicates the similarity between the predicted answer corresponding to the sixth question and the expected answer. The third question and the sixth question are both related to traffic scenes, and the third question and the sixth question are used to obtain information at different levels.
[0051] In the sixth aspect, the information processing device is also used to execute the steps performed by the training device in the third aspect and various possible implementations of the third aspect. For the meaning of the nouns in the sixth aspect of this application and various possible implementations of the sixth aspect, the specific implementation methods of the steps, and the beneficial effects brought about by each possible implementation method, please refer to the description of the various possible implementation methods in the third aspect, and will not be repeated here one by one.
[0052] In the seventh aspect, an embodiment of the present application provides a device including a processor and a memory, where the processor is coupled to the memory, the memory is used to store programs, and the processor is used to execute the programs in the memory, so that the device executes the method described in the first aspect, the second aspect or the third aspect above.
[0053] In an eighth aspect, an embodiment of the present application provides a vehicle comprising a processor and a memory, wherein the processor is coupled to the memory, the memory being used to store programs; and the processor being used to execute the programs in the memory so that the vehicle executes the method described in the first or second aspect above.
[0054] In a ninth aspect, the present application provides a computer-readable storage medium, in which a computer program is stored. When the computer-readable storage medium is run on a computer, the computer executes the method described in the first, second or third aspect above.
[0055] In a tenth aspect, the present application provides a circuit system, which includes a processing circuit, and the processing circuit is configured to execute the method described in the first aspect, the second aspect or the third aspect above.
[0056] In an eleventh aspect, the present application provides a computer program product, which includes a computer program. When the computer program product is run on a computer, it enables the computer to execute the method described in the first aspect, the second aspect or the third aspect above.
[0057] In a twelfth aspect, the present application provides a chip system, which includes a processor for supporting the implementation of the functions involved in the above aspects, such as sending or processing the data and / or information involved in the above methods. In one possible design, the chip system also includes a memory, which is used to store program instructions and data necessary for the server or communication device. The chip system can be composed of a chip or can include a chip and other discrete devices. BRIEF DESCRIPTION OF THE DRAWINGS
[0058] FIG1 is a schematic diagram of a structure of an artificial intelligence main framework provided in an embodiment of the present application;
[0059] FIG2 is a system architecture diagram of a data processing system provided in an embodiment of the present application;
[0060] FIG3 is a flow chart of an information processing method provided in an embodiment of the present application;
[0061] FIG4 is a schematic diagram of a deep learning model provided in an embodiment of the present application;
[0062] FIG5a is a schematic diagram of first information provided in an embodiment of the present application;
[0063] FIG5 b is another schematic diagram of the first information provided in an embodiment of the present application;
[0064] FIG6 is a schematic diagram of a vehicle obtaining first information according to an embodiment of the present application;
[0065] FIG7 is another flow chart of an information processing method according to an embodiment of the present application;
[0066] FIG8 is a schematic diagram of an alternative answer provided in an embodiment of the present application;
[0067] FIG9 is a schematic diagram of a traffic scenario provided in an embodiment of the present application;
[0068] FIG10 is another flow chart of an information processing method according to an embodiment of the present application;
[0069] FIG11 is a schematic diagram of multiple second questions provided by an embodiment of the present application;
[0070] FIG12 is a schematic diagram of the structure of an information processing device provided in an embodiment of the present application;
[0071] FIG13 is another schematic diagram of the structure of an information processing device provided in an embodiment of the present application;
[0072] FIG14 is another schematic diagram of the structure of an information processing device provided in an embodiment of the present application;
[0073] FIG15 is a schematic diagram of a structure of a device provided in an embodiment of the present application;
[0074] FIG16 is a schematic structural diagram of a vehicle provided in an embodiment of the present application;
[0075] FIG17 is a schematic diagram of the structure of a chip provided in an embodiment of the present application. DETAILED DESCRIPTION
[0076] The embodiments of the present application are described below in conjunction with the accompanying drawings. Those skilled in the art will appreciate that, with the development of technology and the emergence of new scenarios, the technical solutions provided in the embodiments of the present application are also applicable to similar technical problems.
[0077] The terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequential order. It should be understood that the terms used in this way can be interchangeable under appropriate circumstances, and this is merely a way of distinguishing the objects of the same attributes when describing them in the embodiments of the present application. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, so that the process, method, system, product or equipment comprising a series of units need not be limited to those units, but may include other units that are not clearly listed or inherent to these processes, methods, products or equipment.
[0078] In the embodiments of the present application, "sending" and "receiving" indicate the direction of signal transmission. For example, "sending information to XX device" can be understood as the destination of the information being XX device, which can include direct sending through the air interface, as well as indirect sending through the air interface by other units or modules. "Receiving information from YY device" can be understood as the source of the information being YY device, which can include direct receiving from YY device through the air interface, as well as indirect receiving from YY device through the air interface from other units or modules. "Sending" can also be understood as the "output" of the chip interface, and "receiving" can also be understood as the "input" of the chip interface. In other words, sending and receiving can be performed between devices or within a device, for example, between components, modules, chips, software modules or hardware modules within the device through a bus, trace or interface. It is understandable that information may undergo necessary processing, such as encoding, modulation, etc., between the source and destination of the information, but the destination can understand the valid information from the source. Similar expressions in this application can be understood similarly and will not be repeated.
[0079] In the embodiments of the present application, "indication" may include direct indication and indirect indication, and may also include explicit indication and implicit indication. The information indicated by a certain information (such as the indication information described below) is called information to be indicated. In the specific implementation process, there are many ways to indicate the information to be indicated, such as but not limited to, the information to be indicated can be directly indicated, such as the information to be indicated itself or the index of the information to be indicated. The information to be indicated can also be indirectly indicated by indicating other information, wherein there is an association between the other information and the information to be indicated; it is also possible to indicate only a part of the information to be indicated, while the other parts of the information to be indicated are known or agreed in advance, for example, the indication of specific information can be achieved with the help of the arrangement order of each information agreed in advance (such as predefined by the protocol), thereby reducing the indication overhead to a certain extent. The present application does not limit the specific method of indication. It is understandable that, for the sender of the indication information, the indication information can be used to indicate the information to be indicated, and for the receiver of the indication information, the indication information can be used to determine the information to be indicated.
[0080] First, let's describe the overall workflow of an AI system. See Figure 1, which shows a schematic diagram of the AI framework. This framework will be explained from two perspectives: the "intelligent information chain" (horizontal axis) and the "IT value chain" (vertical axis). The "intelligent information chain" reflects the entire process from data acquisition to processing. For example, it could be the general process of intelligent information perception, intelligent information representation and formation, intelligent reasoning, intelligent decision-making, and intelligent execution and output. Throughout this process, data undergoes a condensed journey from "data-information-knowledge-wisdom." The "IT value chain," spanning the underlying infrastructure of human intelligence, information (provided and processed by technology), and the system's industrial ecosystem, reflects the value that AI brings to the information technology industry.
[0081] (1) Infrastructure
[0082] The infrastructure provides computing power for AI systems, enabling communication with the outside world and providing support through a basic platform. Communication with the outside world is achieved through sensors; computing power is provided by smart chips, which can specifically adopt hardware acceleration chips such as central processing units (CPUs), embedded neural network processing units (NPUs), graphics processing units (GPUs), application-specific integrated circuits (ASICs), or field programmable gate arrays (FPGAs); the basic platform includes related platform guarantees and support such as distributed computing frameworks and networks, and can include cloud storage and computing, and interconnected networks. For example, sensors communicate with the outside world to obtain data, which is then provided to the smart chips in the distributed computing system provided by the basic platform for calculation.
[0083] (2) Data
[0084] Data above the infrastructure layer represents data sources for AI. This data includes graphics, images, voice, and text, as well as IoT data from traditional devices. This includes business data from existing systems and sensor data such as force, displacement, liquid level, temperature, and humidity.
[0085] (3) Data processing
[0086] Data processing generally includes data training, machine learning, deep learning, search, reasoning, decision-making, etc.
[0087] Among them, machine learning and deep learning can symbolize and formalize data for intelligent information modeling, extraction, preprocessing, and training.
[0088] Reasoning refers to the process of simulating human intelligent reasoning in computers or intelligent systems, using formalized information to perform machine thinking and solve problems based on reasoning control strategies. Typical functions are search and matching.
[0089] Decision-making refers to the process of making decisions after intelligent information is reasoned, and usually provides functions such as classification, sorting, and prediction.
[0090] (4) General ability
[0091] After the data has undergone the data processing mentioned above, some general capabilities can be further formed based on the results of the data processing, such as algorithms or a general system, for example, translation, text analysis, computer vision processing, speech recognition, image recognition, etc.
[0092] (5) Smart products and industry applications
[0093] Smart products and industry applications refer to the products and applications of artificial intelligence systems in various fields. They are the encapsulation of the overall artificial intelligence solution, which productizes intelligent information decision-making and realizes practical application. Its application areas mainly include: smart terminals, smart manufacturing, smart transportation, smart homes, smart medical care, smart security, autonomous driving, smart cities, etc.
[0094] The method provided in this application can be applied to the field of autonomous driving. For example, in the method provided in this application, a deep learning model capable of processing text information is applied to the field of autonomous driving. Before describing the method provided in this application in detail, please refer to Figure 2, which is a system architecture diagram of a data processing system provided in an embodiment of this application. In Figure 2, the data processing system 200 includes a training device 210, a database 220, an execution device 230, a data storage system 240, and a client device 250. The execution device 230 includes a computing module 231.
[0095] Among them, the database 220 stores a training data set. During the training phase of the deep learning model 201, the training device 210 generates the deep learning model 201 and iteratively trains the deep learning model 201 using the training data set to obtain the trained deep learning model 201. The deep learning model 201 can be specifically represented by a neural network or a non-neural network model. In the embodiments of the present application, only the deep learning model 201 represented by a neural network is used as an example for description.
[0096] The deep learning model 201 obtained by the training device 210 after the training operation can be deployed to the computing module 231 of the execution device 230. The execution device 230 can call data, code, etc. in the data storage system 240, or store data, instructions, etc. in the data storage system 240. The data storage system 240 can be placed in the execution device 230, or the data storage system 240 can be an external memory relative to the execution device 230.
[0097] During the application stage of the deep learning model 201 of the U-shaped over-training operation, after the execution device 230 inputs the first information into the deep learning model 201 in the calculation module 231, the second information output by the deep learning model 201 can be obtained, wherein the first information includes the first text information, and the specific information carried by the second information is related to the type of task performed by the deep learning model 201 (for the convenience of description, it will be referred to as the "first task" below).
[0098] In some embodiments of the present application, please refer to Figure 2, the execution device 230 and the client device 250 can be independent devices. The execution device 230 is configured with an input / output (I / O) interface to interact with the client device 250 for data. After determining the first information, the client device sends the first information to the execution device 230 through the I / O interface. After the execution device 230 generates second information corresponding to the first information through the deep learning model 201 in the computing module 231, the execution device 230 can return the aforementioned second information to the client device through the I / O interface.
[0099] It is worth noting that Figure 2 is only an architectural diagram of the data processing system provided by an embodiment of the present invention, and the positional relationship between the devices, components, modules, etc. shown in the figure does not constitute any limitation. For example, in other embodiments of the present application, the execution device 230 and the client device can be integrated into the same device, and the user can interact directly with the execution device 230. Exemplarily, the execution device 230 can be a module in the host processor (Host CPU) of the client device that uses a deep learning model to process data. The execution device 230 can also be a graphics processing unit (GPU) or a neural network processor (NPU) in the client device. The GPU or NPU is mounted on the host processor as a coprocessor, and the host processor assigns tasks.
[0100] In combination with the above description, the present application provides an information processing method. Specifically, please refer to FIG3 , which is a flow chart of the information processing method provided by the embodiment of the present application.
[0101] 301. Input first information corresponding to a traffic scene around the vehicle into a deep learning model, where the first information includes first text information.
[0102] In an embodiment of the present application, before the first device inputs the first information corresponding to the traffic scene around the vehicle into the deep learning model, the vehicle must first obtain the first information. In one embodiment, the first device is a vehicle, and if the deep learning model that has performed the training operation is deployed in the vehicle (i.e., the vehicle is also the execution device of the deep learning model that has performed the training operation), step 301 includes: the vehicle inputs the first information corresponding to the traffic scene around the vehicle into the locally deployed deep learning model.
[0103] In another case, the first device is a vehicle. If the deep learning model that has performed the training operation is deployed on a cloud server (that is, the device that performs the deep learning model that has performed the training operation is a cloud server), step 301 may include: the first device sends first information corresponding to the traffic scene around the vehicle to the cloud server where the deep learning model is deployed.
[0104] In another case, the first device is a cloud server that deploys a deep learning model, and step 301 may include: after receiving the first information sent by the vehicle corresponding to the traffic scene around the vehicle, the first device inputs the first information into the deep learning model.
[0105] Exemplarily, the above-mentioned deep learning model can specifically adopt a machine learning model based on the attention mechanism, or the above-mentioned deep learning model can also adopt a recurrent neural network, a convolutional neural network, a fully connected neural network or other types of machine learning models, etc., which are not limited here.
[0106] Exemplarily, when the deep learning model specifically adopts a machine learning model based on an attention mechanism, the deep learning model may include N first neural network modules, where N is an integer greater than or equal to 1, and each first neural network module is a neural network module based on an attention mechanism. The deep learning model may also include other neural network layers, for example, other neural network layers may include linear neural network layers, neural network layers for normalization, or other neural network layers, etc. It should be noted that the specific composition of the deep learning model can be flexibly determined in combination with the actual application scenario. The examples here are only for the convenience of understanding this solution and are not used to limit this solution.
[0107] In order to understand this solution more intuitively, please refer to Figure 4, which is a schematic diagram of a deep learning model provided in an embodiment of the present application. As shown in Figure 4, the deep learning model may include N first neural network modules. The deep learning model may also include a normalized exponential activation function (Softmax) layer, a linear fully connected (Linear) layer and other neural network layers.
[0108] Among them, each first neural network module is a neural network module based on the attention mechanism. As shown in Figure 4, each first neural network module may include 3 different Linear layers, a neural network layer based on the attention mechanism (Attention), a residual link and normalization (Add&Norm) layer, a feedforward (Feed Forward) neural network layer and a Softmax layer. Figure 4 also shows the order of processing information in each first neural network module. It should be understood that the example in Figure 4 is only for the convenience of understanding this solution and is not used to limit this solution.
[0109] The first information includes first text information, which may correspond to a task performed by the deep learning model (hereinafter referred to as the "first task" for ease of description), and may be used to prompt the deep learning model to generate prediction information corresponding to the first task. For example, the first information may be specifically represented by a second question, and the first text information may be specifically represented by the prompt information in the aforementioned second question. It should be understood that the term "first question" will appear in subsequent steps and will not be introduced here.
[0110] Optionally, the first information further includes first feature information, the first feature information including feature information of environmental information surrounding the vehicle, the environmental information surrounding the vehicle including physical attribute information of objects surrounding the vehicle. Exemplarily, the environmental information surrounding the vehicle may include physical attribute information of objects surrounding the vehicle in each frame of one or more frames of image (or point cloud data).
[0111] Exemplarily, the objects around the vehicle may include dynamic obstacles around the vehicle, static obstacles around the vehicle, traffic markings around the vehicle, traffic signs around the vehicle, or other types of objects; for example, the aforementioned dynamic obstacles may be other vehicles, pedestrians, electric vehicles, or other dynamic obstacles around the vehicle; the aforementioned static obstacles may be houses, fences, grass, or other static obstacles around the vehicle; the aforementioned traffic markings may be stop lines, crosswalks, or other traffic sign lines; the aforementioned traffic signs may be traffic lights, speed limit signs, or other traffic signs; and the specific details may be determined in combination with the actual application scenario and are not limited here.
[0112] For example, the physical property information of dynamic obstacles may include the category, position, speed, direction, height, color, shape or other types of physical property information of the dynamic obstacles; the physical property information of static obstacles may include the category, position, direction, height, material, shape, color or other physical property information of the static obstacles; the physical property information of traffic markings may include the category, position, shape, color or other physical property information of traffic markings; the physical properties of traffic signs may include the category, position, shape, color, content of traffic signs or other physical property information. It should be noted that the examples of various physical property information here are only for the convenience of understanding this solution and are not used to limit this solution.
[0113] Optionally, the first characteristic information also includes characteristic information of the vehicle's driving behavior information and / or characteristic information of navigation information; illustratively, the vehicle's driving behavior information may indicate information about the vehicle's driving behavior in each frame of one or more frames of image (or point cloud data), and the navigation information also includes navigation information in each frame of one or more frames of image (or point cloud data). Exemplarily, the vehicle's driving behavior information may include the vehicle's speed, acceleration, orientation, position, or other information, and the navigation information may include the position of a navigation line, the direction indicated by the navigation line, the shape of the navigation line, or other information. The specific information may be determined based on actual circumstances and is not limited here.
[0114] The process of the vehicle acquiring the first information may include: the vehicle acquiring the first feature information. With respect to the specific implementation process of the vehicle acquiring the first feature information, illustratively, the first feature information can be obtained based on a feature extraction network (hereinafter referred to as the "first feature extraction network" for the convenience of description). For example, the vehicle can input the environmental information around the vehicle into the first feature extraction network, and optionally, can also input the driving behavior information and / or navigation information of the vehicle into the first feature extraction network to obtain the third feature information generated by the first feature extraction network, and then obtain the first feature information. "Third feature information" can also be understood as the hidden feature information of the environmental information around the vehicle (optionally, it can also include the driving behavior information and / or navigation information of the vehicle).
[0115] Furthermore, in one case, "third feature information" and "first feature information" have the same meaning, that is, the vehicle can directly use the third feature information generated by the first feature extraction network as the first feature information.
[0116] In another case, after obtaining the third feature information, the vehicle can also input the third feature information into the second feature extraction network, and update the features of the third feature information through the second feature extraction network to obtain the first feature information. Exemplarily, the third feature information can be specifically expressed in the form of a feature map, and the first feature information can be specifically expressed in the form of a token, that is, the third feature information in the form of a feature map can be converted into the first feature information in the form of a token through the second feature extraction network. Since the deep learning model is a machine learning model that processes text information, after obtaining the first feature information in the form of a token and then inputting it into the deep learning model, the deep learning model can more easily understand the first feature information, reducing the difficulty of the deep learning model in understanding the first feature information, which is conducive to improving the accuracy of the information output by the deep learning model.
[0117] Exemplarily, the second feature extraction network can be specifically expressed as a neural network based on the attention mechanism, a convolutional neural network, a fully connected neural network or other types of neural networks, etc. The specific form of expression of the second feature extraction network can be determined according to the actual application scenario and is not limited in the embodiments of this application.
[0118] Among them, the first feature extraction network belongs to the first neural network, and the first neural network may also include a first feature processing network. Exemplarily, the first neural network can be specifically manifested as a neural network based on an attention mechanism, a convolutional neural network, a fully connected neural network, or other types of neural networks, etc., which can be specifically determined in combination with actual application scenarios and are not limited here. Exemplarily, the first feature extraction network can also be called an encoder, and the first feature processing network can also be called a decoder.
[0119] Exemplarily, the first neural network is used to perform at least one of the following tasks: predicting the trajectory of objects around the vehicle, making decisions on the vehicle's behavior, planning the trajectory of the vehicle, predicting the speed range of the vehicle, controlling the vehicle, or other tasks; optionally, the first neural network is used to perform at least two of the aforementioned multiple tasks.
[0120] In an embodiment of the present application, the first information not only includes the first text information, but also includes the first feature information. The first feature information includes the feature information of the environmental information around the vehicle, so that the deep learning model can fully combine the feature information of the environmental information around the vehicle to determine the second information, which is conducive to improving the rationality of the second information output by the deep learning model.
[0121] In addition, if the first neural network is used to perform at least two of the following tasks: trajectory prediction of objects around the vehicle, decision-making on the vehicle's behavior, trajectory planning of the vehicle, prediction of the vehicle's speed range, or control of the vehicle, then the first feature information needs to cover richer information, that is, the first information containing the first feature information will cover more information, which is conducive to the deep learning model obtaining more information, and thus is conducive to improving the accuracy of the information output by the deep learning model.
[0122] The process of the vehicle obtaining the first information includes: the vehicle obtaining the first text information.
[0123] Regarding the specific implementation process of the vehicle obtaining the first text information, in one case, the first text information includes information about a preset text of the first task performed by the deep learning model.
[0124] For example, the vehicle may be equipped with preset text corresponding to each of the at least one second task. Optionally, when the first information specifically represents a second question, each preset text may also be understood as a preset prompt. The at least one second task may include any one or more of the following: making decisions about vehicle behavior, trajectory planning, vehicle control, or other tasks, the list of which is not exhaustive.
[0125] In order to further understand the relationship between the words "decision-making", "trajectory planning" and "control", the aforementioned nouns are further explained here. For example, "making decisions on the vehicle's behavior" refers to making decisions at the vehicle's behavior level; for example, "making decisions on the vehicle's behavior" can be specifically manifested as turning left at the intersection ahead, and for another example, "making decisions on the vehicle's behavior" can be specifically manifested as overtaking the vehicle ahead or other behaviors, etc., and the examples are not exhaustive here.
[0126] "Trajectory planning for the vehicle" can also be called "path planning for the vehicle". "Trajectory planning for the vehicle" refers to determining the path for implementing the decision based on the decision on the vehicle's behavior (it can also be understood as "determining the trajectory for implementing the behavior"). For example, after the decision on the vehicle's behavior is to turn left at the intersection ahead, trajectory planning for the vehicle can be to determine what path the vehicle should take to implement the behavior of "turning left at the intersection ahead"; for another example, if the decision on the vehicle's behavior is to overtake the vehicle ahead, trajectory planning for the vehicle can be to determine what path the vehicle should take to implement the behavior of "overtaking the vehicle ahead", etc.
[0127] "Controlling the vehicle" refers to control instructions for the vehicle's components to achieve the planned trajectory of the vehicle. The aforementioned "control instructions for the vehicle's components" can also be called "vehicle control instructions." For example, "controlling the vehicle" can be specifically manifested as controlling the angle of the steering wheel, controlling the amount of accelerator pedal applied, controlling the amount of brake pedal applied, or control instructions for other components in the vehicle, etc., and these are not exhaustive here.
[0128] For example, the preset text corresponding to the task of "making decisions about the vehicle" can be: "What interactive behavior will the vehicle perform next, and what is the reason for performing the aforementioned behavior", "For what reason, the vehicle needs to perform what behavior", "What behavior is the vehicle about to perform" or other preset text corresponding to the task of "making decisions about the vehicle", etc.
[0129] The preset text corresponding to the task of "planning a path for the vehicle" can be: "What is the path planned for the vehicle?", "How should the vehicle go next?", "What is the next trajectory of the vehicle?" or other preset text corresponding to the task of "planning a path for the vehicle".
[0130] The preset text corresponding to the task of "controlling the vehicle" can be: "Please confirm the vehicle control instructions", "How to control the vehicle's movement", "How to control the various components of the vehicle", "Can you give a set of vehicle control instructions" or other preset texts corresponding to the task of "controlling the vehicle", etc. It should be noted that the examples of various preset texts here are only for the convenience of understanding this solution and are not used to limit this solution.
[0131] When the vehicle needs to perform a first task among the at least one second task, the vehicle is triggered to automatically obtain first information, that is, the vehicle is triggered to automatically obtain first text information. The vehicle obtaining the first text information may include: the vehicle obtains one or more preset texts corresponding to the first task from the preset texts corresponding to the at least one second task, thereby obtaining one or more first text messages; each first text message includes information about a preset text.
[0132] Exemplarily, in this scenario, the first text information may include the initial feature information of the preset text obtained after feature extraction of the preset text. The aforementioned "feature extraction" can also be understood as "initial encoding", or the aforementioned "feature extraction" can also be understood as vectorization (embedding); exemplarily, the initial feature information of the preset text can be specifically expressed in the form of a token.
[0133] In another case, the first text information may also include text input by the user (hereinafter referred to as "first text" for ease of distinction), that is, the first text input by the user may trigger the vehicle to obtain the first information. For example, the vehicle may receive any text input by the user (hereinafter referred to as "second text" for ease of distinction), and upon determining that the second text input by the user corresponds to the at least one second task, the second text may be determined to be the first text, triggering the vehicle to obtain the first information, and then executing step 301.
[0134] Exemplarily, in this scenario, the first text information may be the initial feature information of the first text obtained after feature extraction of the first text input by the user. The aforementioned "feature extraction" may also be understood as "initial encoding"; exemplarily, the initial feature information of the first text may be specifically expressed in the form of a token.
[0135] Regarding the specific implementation method of the vehicle obtaining the first text input by the user, for example, in one implementation method, the vehicle can provide the user with a receiving icon corresponding to the first function (which can also be replaced by a "button"), and the first function is used for the user to understand the decision-making situation, trajectory planning situation, or vehicle control situation during the automatic driving process. Then, when the user clicks the aforementioned receiving icon (or "button") corresponding to the first function, the vehicle can receive the first text input by the user, and determine that the first text input by the user corresponds to the at least one second task mentioned above, thereby triggering the vehicle to obtain the first information.
[0136] In another implementation, a first machine learning model that has performed a training operation may be pre-deployed in the vehicle, and the first machine learning model is used to determine whether the text input by the user corresponds to one or more tasks in at least one second task; illustratively, the first machine learning model may be specifically manifested as a machine learning model for performing classification functions in N categories, and the aforementioned N categories may include each second task and irrelevant ones. The vehicle can receive any text input by the user (i.e., the second text), input the aforementioned second text into the first machine learning model, and obtain prediction information output by the first machine learning model, which indicates to which category of the N categories the second text input by the user belongs; if it is determined that the second text belongs to any one of the at least one second task based on the prediction information output by the first machine learning model, then the aforementioned second text is determined to be the first text, thereby triggering the vehicle to obtain the first information.
[0137] It should be noted that other methods can also be used to trigger the vehicle to obtain the first information based on the first text input by the user. The example here is only to prove the feasibility of this solution and is not used to limit this solution.
[0138] After obtaining the first text information and the first feature information, the vehicle can combine the first text information and the first feature information to obtain combined information; illustratively, the aforementioned "combination" can be splicing, in which case there is a clear separation between the first text information and the first feature information; alternatively, the first text information in token form and the first feature information in token form can be combined in a mixed manner, that is, there is no clear separation between the first text information and the first feature information; or the first text information and the first feature information can be combined in other ways, which is not limited in the embodiments of the present application.
[0139] In one case, the aforementioned "combined information" can be directly used as the "first information." In another case, the vehicle can perform type encoding on the aforementioned combined information to obtain the first information; the purpose of performing type encoding is to indicate that the first text information and the first feature information are of different types (or also called "different modalities").
[0140] For a more intuitive understanding of the present solution, please refer to Figures 5a, 5b and 6. Figure 5a is a schematic diagram of the first information provided in an embodiment of the present application, Figure 5b is another schematic diagram of the first information provided in an embodiment of the present application, and Figure 6 is a schematic diagram of the vehicle obtaining the first information provided in an embodiment of the present application. First refer to Figure 5a. As shown in 5a, the first text information in token form and the first feature information in token form are completely separated. In Figure 5a, the first text information in token form is in front and the first feature information in token form is in the back, thereby realizing the splicing of the first text information in token form and the first feature information in token form to obtain the combined information; and then the combined information is type encoded to obtain the first information. It should be understood that the example in Figure 5a is only for the convenience of understanding the present solution and is not used to limit the present solution.
[0141] Continuing to refer to Figure 5b, Figure 5b takes the first text information in token form and the first feature information in token form as an example of being combined together in a mixed manner. The method adopted in Figure 5b is to mix the first text information and the first feature information together according to the logic of the language. The first text information is in token form, "Based on the given..., please tell me how to cross this intersection." Then the first information after mixing the first text information and the first feature information can be understood as: Based on the given <first feature information>, please tell me how to cross this intersection. It should be noted that the method of text description here is to facilitate understanding of how the first text information in token form and the first feature information in token form are mixed, and is not used to limit this solution.
[0142] Continuing with FIG6 , as shown in FIG6 , a training sample including environmental information surrounding the vehicle, driving behavior information, and navigation information is input into the first feature extraction network to obtain third feature information generated by the first feature extraction network. The first feature extraction network and the first feature processing network both belong to the first neural network. The first neural network is used to perform tasks including: trajectory prediction of objects around the vehicle, decision-making on the vehicle's behavior, and prediction of the vehicle's speed range. The third feature information is input into the second feature extraction network to obtain first feature information generated by the second feature extraction network. A preset text corresponding to the first task is obtained to obtain first text information. Based on the first text information and the first feature information, a type encoding operation is performed to obtain first information, which is then input into the deep learning model. It should be understood that the example in FIG6 is only for the convenience of understanding this solution and is not intended to limit this solution.
[0143] 302. Obtain second information corresponding to the first information, wherein the second information is obtained based on a deep learning model, and the second information corresponds to any of the following tasks: making decisions on the behavior of the vehicle, planning the trajectory of the vehicle, or controlling the vehicle.
[0144] In an embodiment of the present application, after the first device inputs first information corresponding to the traffic scene around the vehicle into the deep learning model, it can obtain first prediction information generated by the deep learning model, and obtain second information corresponding to the first information based on the aforementioned first prediction information. The second information corresponds to any of the following tasks: making decisions on the behavior of the vehicle, planning the trajectory of the vehicle, controlling the vehicle, or other tasks; that is, the second information is used to indicate what behavior the vehicle will perform next, or the second information may include the next trajectory information of the vehicle, or the second information may include the vehicle control instructions of the vehicle, etc. The specific information included in the second information can be determined according to the specific application scenario of the second information and is not limited here.
[0145] In one scenario, the first device is a vehicle, and if the deep learning model is deployed in the vehicle, step 302 may include: the vehicle obtaining first prediction information output by the locally deployed deep learning model, and then determining the second information based on the first prediction information. In another scenario, the first device is a vehicle, and if the deep learning model is deployed on a cloud server, step 302 may include: the vehicle receiving the first prediction information sent by the server deploying the deep learning model, and then determining the second information based on the first prediction information.
[0146] In another case, the first device is a server that deploys a deep learning model, then step 302 may include: the first device obtains first prediction information output by the locally deployed deep learning model, determines second information based on the first prediction information, and sends the second information to the vehicle.
[0147] In an embodiment of the present application, a deep learning model capable of processing text information is introduced into the field of autonomous driving. Based on the second information generated by the deep learning model that processes text information, decisions are made on the behavior of the vehicle, the trajectory of the vehicle is planned, or the vehicle is controlled. Since the input information of the deep learning model provided by the present application contains text information that is easy for users to understand, the interpretability of the operation process of the deep learning model is improved, and the decision-making, trajectory planning or control process of the autonomous driving vehicle is made more transparent, so that users can understand the behavior of the autonomous driving vehicle more intuitively.
[0148] Based on the embodiment corresponding to Figure 3 above, the following describes the specific implementation process of the training phase and application phase of the deep learning model provided in the embodiment of the present application.
[0149] 1. Training Phase
[0150] In the embodiments of the present application, as can be seen from the above description, after completing the training operation of the deep learning model, the first task performed by the deep learning model can be making decisions about the vehicle's behavior, planning the vehicle's trajectory, or controlling the vehicle. Optionally, the entire training process of the deep learning model can include at least two training phases, wherein the aforementioned at least two training phases are used to train the machine learning model to learn information at different levels. This application will disclose the four different training phases of the deep learning model in subsequent steps (i.e., the first training phase, the second training phase, the third training phase, and the fourth training phase described below).
[0151] For example, in the first training stage, the purpose is to allow the deep learning model to learn basic attribute information and semantic information, such as allowing the deep learning model to learn basic attribute information such as speed, position or orientation. For example, the deep learning model can learn obstacle labels, the concept of "frame", the concept of "second" or other basic semantic information.
[0152] In the second training phase, the deep learning model learns the semantics of the directional behavior and / or motion behavior of the vehicle or obstacles in multiple frames of images (or point cloud data). For example, the aforementioned "directional behavior" can refer to lateral behavior, longitudinal behavior, or other behaviors, and the aforementioned "motion behavior" can refer to acceleration behavior, steering behavior, lane changing behavior, or other behaviors. And / or, the deep learning model learns the semantics of the relative spatial relationship between different objects in a single frame of image (or point cloud data), for example, learning that obstacle 1 and obstacle 2 are the front and rear vehicles, and that obstacle 1 is within road topology 1.
[0153] The third training stage is to allow the deep learning model to learn the semantics of the interaction behaviors between different objects. The aforementioned "interaction behavior" can also be called "game behavior", "interaction relationship" or other names with equivalent meanings. For example, the aforementioned "interaction behavior" can be obstacle 1 pressing on obstacle 2, and obstacle 2 avoiding it by moving sideways to the left. For another example, the aforementioned "interaction behavior" can be the interaction behavior of the vehicle changing lanes and overtaking with obstacle 3.
[0154] The fourth training phase allows the deep learning model to learn the cause-and-effect and logical relationships between interactions between different objects. For example, because obstacle 3 in front of the ego vehicle is moving slowly and the lane to the left of the ego vehicle is free, the ego vehicle changes lanes to overtake to improve driving efficiency. It should be noted that these examples are only for facilitating understanding of this solution and are not intended to limit it.
[0155] Specifically, please refer to FIG. 7 , which is another flowchart of an information processing method provided in an embodiment of the present application. The information processing method provided in the present application may include:
[0156] 701. In the first training phase, the training device obtains the first-level questions corresponding to the traffic scenario.
[0157] In the embodiment of the present application, exemplarily, the "first-level question" includes at least the initial feature information of the first prompt, and the "first-level question" may also include the first feature information, and the first feature information includes the feature information of the environmental information around the vehicle, and the environmental information includes the physical attribute information of the objects around the vehicle. The meaning of the "first feature information" can refer to the description in the embodiment corresponding to Figure 3 above, and will not be repeated here. Exemplarily, the first prompt may include at least one mask (MASK), and the content of the MASK is predicted by a deep learning model, that is, the content of the MASK in the first prompt is completed by a deep learning model. Since the first training stage is to allow the deep learning model to learn basic attribute information and semantic information, the content of the MASK predicted each time by the deep learning model may include at least one of the following data frames, vehicle or obstacle labels, attribute names or attribute values.
[0158] For example, the "first prompt" might be expressed as "In frame 0, the vehicle's speed is ()." The portion enclosed by "()" is a mask, representing the portion that the deep learning model needs to predict. In this example, the deep learning model predicts the value of the "speed" attribute. Another example might be expressed as "In frame 0, the vehicle's x-coordinate position is ()." In this example, the deep learning model predicts the value of the "x-coordinate position" attribute.
[0159] For another example, the "first prompt" can be specifically expressed as "In the -1 frame, the speed of () is 10m / s". In the above example, the deep learning model predicts the label of the vehicle or obstacle. For another example, the "first prompt" can be specifically expressed as "In the () frame, the y-coordinate position of obstacle 1 is 2.9m". In the above example, the deep learning model predicts the frame of the data. For another example, the "first prompt" can be specifically expressed as "In the -1 frame, the speed of the vehicle () is 10m / s". In the above example, the deep learning model predicts the attribute name, etc. It should be noted that the various examples here are only for the convenience of understanding the concept of "first prompt" and are not used to limit this solution.
[0160] Step 701 may include: the training device obtains a first prompt, performs feature extraction on the first prompt to obtain initial feature information of the first prompt, and obtains a first-level question corresponding to the traffic scene based on the initial feature information of the first prompt and the first feature information. The aforementioned "feature extraction" can also be understood as "initial encoding"; illustratively, the initial feature information of the first prompt can be specifically expressed in the form of a token. The specific implementation method of the training device "obtaining a first-level question corresponding to the traffic scene based on the initial feature information and the first feature information of the first prompt" is similar to the specific implementation method of "generating the first information based on the first text information and the first feature information" in the embodiment corresponding to Figure 3, with the difference being that the "first text information" in the embodiment corresponding to Figure 3 is replaced by the "initial feature information of the first prompt" in step 701, and the "first information" in the embodiment corresponding to Figure 3 is replaced by the "first-level question" in step 701. The specific implementation methods of the aforementioned steps will not be repeated here.
[0161] Regarding the specific implementation method for obtaining the first feature information by the training device, the training device may be deployed with a training data set comprising multiple training samples; each training sample may include environmental information surrounding the vehicle, and optionally, each training sample may also include driving behavior information and / or navigation information of the vehicle. The training device inputs the training sample into the first feature extraction network to obtain third feature information generated by the first feature extraction network, and the third feature information is used to obtain the first feature information.
[0162] It should be noted that the meanings of "environmental information around the vehicle", "driving behavior information of the vehicle", "navigation information of the vehicle", "third characteristic information" and "first characteristic information" can all be found in the description in the corresponding embodiment of Figure 3. The specific implementation method of "obtaining the first characteristic information based on the third characteristic information" can also be found in the description in the corresponding embodiment of Figure 3, and will not be repeated here.
[0163] Optionally, the training device can also input the first feature information into the first feature processing network to obtain second prediction information generated by the first feature processing network corresponding to the training sample (that is, the second prediction information corresponding to the environmental information around the vehicle). The first feature extraction network and the first feature processing network belong to the same first neural network, and "the second prediction information generated by the first feature processing network" can also be understood as "the second prediction information generated by the first neural network."
[0164] The first neural network is used to perform at least one of the following tasks: predicting the trajectory of objects around the vehicle, making decisions on the vehicle's behavior, planning the trajectory of the vehicle, predicting the speed range of the vehicle, controlling the vehicle, or other tasks; optionally, the first neural network is used to perform at least two of the aforementioned tasks.
[0165] A specific implementation method for obtaining the first prompt by the training device. In one implementation method, a second neural network can be pre-deployed in the training device. The training device inputs the training sample into the second neural network to obtain the first prompt corresponding to the first level generated by the second neural network. Exemplarily, the second neural network can be understood as a generator. The second neural network can specifically adopt an attention-based neural network, a convolutional neural network, a recurrent neural network, or other types of neural networks, etc., which are not limited here.
[0166] Optionally, the second neural network can be used to generate prompts in the first training phase, the second training phase, the third training phase, and the fourth training phase. In the first training phase, the training device may input the training sample and the first parameter into the second neural network to obtain the first prompt generated by the second neural network, where the first parameter is used to instruct the second neural network to generate the aforementioned first prompt corresponding to the first level. In the subsequent second training phase, the training device may input the training sample and the second parameter into the second neural network to obtain the second prompt generated by the second neural network, where the second parameter is used to instruct the second neural network to generate the aforementioned second prompt corresponding to the second level. In the subsequent third training phase, the training device may input the training sample and the third parameter into the second neural network to obtain the third prompt generated by the second neural network, where the third parameter is used to instruct the second neural network to generate the aforementioned third prompt corresponding to the third level. In the subsequent fourth training phase, the training device may input the training sample and the fourth parameter into the second neural network to obtain the fourth prompt generated by the second neural network, where the fourth parameter is used to instruct the second neural network to generate the aforementioned fourth prompt corresponding to the fourth level.
[0167] Exemplarily, the first parameter, the second parameter, the third parameter, and the fourth parameter are all different. For example, the first parameter may be 1111, the second parameter may be 2222, the third parameter may be 3333, and the fourth parameter may be 4444; for another example, the first parameter may be the first level, the second parameter may be the second level, the third parameter may be the third level, and the fourth parameter may be the fourth level, etc. It should be noted that the examples here are only for the convenience of understanding the differences between the first parameter, the second parameter, the third parameter, and the fourth parameter, and are not used to limit this solution.
[0168] It should be noted that in other implementations, different neural networks can also be used to generate prompts at different levels in different training stages of the deep learning model. In this case, it is only necessary to input training samples into the neural network used to generate prompts, and there is no need to input additional parameters (such as the first parameter mentioned above) into the neural network used to generate prompts.
[0169] In an embodiment of the present application, during the training phase of the deep learning model, a second neural network is used to generate questions that need to be answered by the deep learning model, and a rich variety of questions can be generated. In order to be able to answer the aforementioned rich variety of questions, it is beneficial for the deep learning model to fully understand various information and to improve the deep learning model's ability to understand traffic scenes. In addition, a training sample containing environmental information around the vehicle is input into the second neural network to obtain the question output by the second neural network, that is, the traffic scene corresponding to the question generated by the second neural network is the same traffic scene as the traffic scene that the deep learning model needs to understand, so that the question output by the second neural network can be closer to the traffic scene that the deep learning model needs to understand, thereby helping to improve the deep learning model's ability to understand traffic scenes.
[0170] In one implementation, multiple prompts corresponding to the first level may also be pre-deployed in the training device. Then, each time the training device obtains a question of the first level, it may obtain a first prompt corresponding to the first level from the multiple prompts corresponding to the first level.
[0171] 702. The training device inputs a first-level question into the deep learning model to obtain a first predicted answer corresponding to the first-level question.
[0172] In an embodiment of the present application, after inputting a first-level question into a deep learning model, the training device may obtain third prediction information generated by the deep learning model corresponding to the first-level question; and based on the third prediction information generated by the deep learning model, obtain a first predicted answer corresponding to the first-level question. For example, the third prediction information may indicate at least one alternative answer corresponding to the first-level question and a first probability value corresponding to each alternative answer.
[0173] Exemplarily, each alternative answer may be specifically expressed in the form of a token. The alternative answer in token form may also be understood as feature information of the alternative answer. The alternative answer in token form is used to indicate one or more words.
[0174] Optionally, a start marker [CLS] bit may be inserted at the beginning of each token-based alternative answer. The start marker [CLS] bit in the token-based alternative answer not only identifies the beginning of the alternative answer but also represents the overall meaning of the alternative answer. The characteristic information of each alternative answer also includes characteristic information of the [CLS] bit. Optionally, an end marker [SEP] bit may also be inserted at the end of each token-based alternative answer. The characteristic information of each alternative answer also includes characteristic information of the [SEP] bit.
[0175] For a more intuitive understanding of this solution, please refer to Figure 8, which is a schematic diagram of the alternative answers provided in an embodiment of the present application. As shown in Figure 8, there is a start flag [CLS] bit at the head of the alternative answer in token form, and there is an end flag [SEP] bit at the tail of the alternative answer in token form. It should be understood that the example in Figure 8 is only for the convenience of understanding this solution and is not used to limit this solution.
[0176] In one implementation, after obtaining the third prediction information, the training device can determine at least one alternative answer indicated by the third prediction information and a first probability value corresponding to each alternative answer, and then can determine a first predicted answer from the at least one alternative answer, and the first probability value corresponding to the first predicted answer is the highest among the at least one alternative answer.
[0177] In another implementation, the first predicted answer is included in a first preset answer set corresponding to the first-level question, and the first predicted answer is obtained from at least one alternative answer indicated by the third prediction information based on the first preset answer set.
[0178] Exemplarily, a first preset answer set may be deployed in the training device, and the first preset answer set includes one or more answers. After the training device obtains the third prediction information, the intersection between at least one alternative answer indicated by the third prediction information and the first preset answer set may be determined, and the aforementioned intersection includes at least one first alternative answer; the training device may determine the first predicted answer from at least one first alternative answer based on the first probability value corresponding to each first alternative answer in the at least one first alternative answer, and the first probability value corresponding to the first predicted answer is the largest among the at least one first alternative answer.
[0179] To further understand the concept of "the intersection between at least one alternative answer and the first set of preset answers," exemplarily, the first-level question is specifically expressed as follows: given the first feature information, at frame -1, the speed of () is 10 m / s. The content in () is the content that needs to be predicted by the deep learning model. The first set of preset answers may include: the ego vehicle, obstacle x, other vehicles, other vehicles, electric vehicles, and pedestrians. Obstacle x means obstacle + label. For example, obstacle 1, obstacle 2, obstacle 3, or obstacle 4 all meet the obstacle + label requirement, that is, they all meet the requirement of obstacle x. After obtaining the third prediction information output by the deep learning model, the training device determines at least one alternative answer indicated by the third prediction information. For example, the at least one alternative answer may include obstacle 1, streetlight, ego vehicle, Chinese textbook, and obstacle 2. Then, the intersection between the at least one alternative answer and the first set of preset answers may include obstacle 1, ego vehicle, and obstacle 2. It should be understood that this example is only for the convenience of understanding this solution and is not intended to limit this solution.
[0180] In another implementation, the first predicted answer can be obtained through a classification network, which is used to determine the first predicted answer corresponding to the second feature information from at least one preset answer. The first predicted answer can be one of the at least one preset answer. The second feature information may include a second alternative answer in token form (also referred to as feature information of the second alternative answer) including feature information of the start flag [CLS] bit, and the first probability value corresponding to the second alternative answer is the largest among at least one alternative answer.
[0181] For example, a trained classification network may be deployed in the training device. After obtaining the third prediction information, the training device may determine feature information of the second candidate answer from the third prediction information, obtain second feature information from the feature information of the second candidate answer, input the second feature information into the classification network, and obtain the first predicted answer generated by the classification network. The classification network is configured to determine a predicted category to which the second feature information belongs from M categories, where the M categories may be at least one of the preset answers.
[0182] For example, the above-mentioned "classification network" can also be called a "classifier". The classification network can be specifically expressed as a recurrent neural network, a fully connected neural network, a convolutional neural network or other types of neural networks, etc., which is not limited in the embodiments of this application.
[0183] 703. The training device trains the deep learning model according to the first loss function. In the first training stage, the first loss function indicates the similarity between the first predicted answer and the first expected answer corresponding to the first level question.
[0184] In an embodiment of the present application, exemplarily, in the first training stage, the training device can obtain the first expected answer corresponding to the above-mentioned first-level question. The aforementioned "first expected answer corresponding to the first-level question" can also be referred to as "the first true value corresponding to the first-level question", "the first annotation information corresponding to the first-level question", "the first label corresponding to the first-level question" or other names with equivalent meanings. The training device can generate a function value of a first loss function based on the first predicted answer and the first expected answer. The goal of training using the first loss function includes improving the similarity between the first predicted answer and the first expected answer corresponding to the first-level question. Based on the function value of the first loss function, the training device uses a backpropagation algorithm to update the weight parameters of the deep learning model to achieve one-time training of the deep learning model.
[0185] Optionally, step 703 may include: the training device trains the deep learning model and the first neural network according to the first loss function and the second loss function, the second loss function indicates the similarity between the second prediction information corresponding to the training sample and the second expected information, and the goal of training using the second loss function includes improving the similarity between the second prediction information corresponding to the training sample and the second expected information.
[0186] Exemplarily, the specific content carried by the second expected information is related to the type of task performed by the first neural network.
[0187] For example, the training device can generate a function value of a first loss function based on the first predicted answer and the first expected answer; and generate a function value of a second loss function based on second predicted information and second expected information corresponding to the training sample. Based on the function values of the first loss function and the second loss function, a total loss function value can be determined; and based on the total loss function value, a backpropagation algorithm is used to update weight parameters of the deep learning model and the first neural network to achieve single-stage training of the deep learning model and the first neural network.
[0188] The training device can repeat steps 701 to 703 multiple times until the first convergence condition is met, so as to train the deep learning model multiple times in the first training stage; wherein, the first convergence condition can be that steps 701 to 703 reach the first number, or the first convergence condition can be that the convergence condition of the first loss function (optionally, also including the second loss function) is met.
[0189] It should be noted that steps 701 to 703 are optional steps. If steps 701 to 703 are not performed, step 704 may be performed directly.
[0190] 704. In the second training phase, the training device obtains second-level questions corresponding to the traffic scenario.
[0191] In the embodiment of the present application, for example, the "second-level question" includes at least the initial feature information of the second prompt, and the "second-level question" may also include the first feature information. The meaning of the "first feature information" can be found in the description of the embodiment corresponding to FIG3 above, and will not be repeated here. For example, the second prompt may include at least one mask (MASK), and the content of the MASK is predicted by a deep learning model, that is, the content of the MASK in the second prompt is completed by the deep learning model.
[0192] Since the second training phase is designed to allow the deep learning model to learn the semantics of the directional behavior of the ego vehicle or obstacles, the semantics of the movement of the ego vehicle or obstacles, or the relative spatial relationships between different objects, the content of the mask predicted by the deep learning model can include at least one of the following: directional behavior, movement behavior, or the relative spatial relationships between different objects. Optionally, in the second training phase, the content of the mask predicted by the deep learning model can also include the data frame, the ego vehicle or obstacle label, or other content.
[0193] For example, the "second prompt" can be specifically expressed as "In frame 0, the vehicle has made a lateral movement ()", and the part in "()" is a MASK, which is the part that needs to be predicted by the deep learning model. In the above example, the deep learning model predicts the action behavior, for example, the content in "()" can be a left lane change. For another example, the "second prompt" can be specifically expressed as "In frame 0, the vehicle has made a (left lane change) on ()," and the deep learning model predicts the direction behavior, for example, the content in "()" can be a lateral movement, etc. It should be noted that the various examples here are only for the convenience of understanding the concept of the "second prompt" and are not used to limit this solution.
[0194] The specific implementation method of step 704 is similar to that of step 701, except that the "first prompt" in step 701 is replaced by the "second prompt" in step 704, and the "first-level question" in step 701 is replaced by the "second-level question" in step 704. The specific implementation method of step 704 will not be repeated here.
[0195] 705. The training device inputs the second-level question into the deep learning model to obtain a second predicted answer corresponding to the second-level question.
[0196] 706. The training device trains the deep learning model according to the first loss function. In the second training phase, the first loss function indicates the similarity between the second predicted answer and the second expected answer corresponding to the second-level question.
[0197] In the embodiment of the present application, the meanings of the nouns in steps 705 and 706 and the specific implementation methods of the steps can be referred to the descriptions in steps 702 and 703. The difference is that the "first-level question" in steps 702 and 703 is replaced by "second-level question", the "first predicted answer" in steps 702 and 703 is replaced by "second predicted answer", and the "first expected answer" in steps 702 and 703 is replaced by "second expected answer". The specific implementation methods of steps 705 and 706 are not repeated here.
[0198] It should be noted that steps 704 to 706 are optional steps. If steps 704 to 706 are not executed, step 707 can be executed directly after executing step 703; if steps 701 to 707 are not executed, step 707 can be executed directly; if steps 701 to 707 are executed, steps 701 to 703 can also be executed crosswise during the process of the training device executing steps 704 to 706 multiple times.
[0199] 707. In the third training phase, the training device obtains third-level questions corresponding to the traffic scenario.
[0200] In the embodiment of the present application, illustratively, the "third-level question" includes at least the initial feature information of the third prompt, and the "third-level question" may also include the first feature information. The meaning of the "first feature information" can be found in the description of the embodiment corresponding to FIG3 above, and will not be repeated here. Exemplarily, the third prompt may include at least one mask (MASK), and the content of the MASK is predicted by a deep learning model, that is, the content of the MASK in the third prompt is completed by a deep learning model.
[0201] Since the third training stage is to allow the deep learning model to learn the semantics of the interaction behavior between different objects, the content of the MASK predicted by the deep learning model may include interaction behavior; optionally, in the third training stage, the content of the MASK predicted by the deep learning model may also include direction behavior and / or action behavior.
[0202] For example, the "third prompt" can be specifically expressed as "the self-car and social car 3 have an interactive game behavior, the category is ()", and the part in "()" is MASK, which is the part that needs to be predicted by the deep learning model. The above example predicts the interactive behavior through the deep learning model. For example, the content in "()" can be overtaking, etc. It should be noted that the various examples here are only for the convenience of understanding the concept of the "third prompt" and are not used to limit this solution.
[0203] The specific implementation method of step 707 is similar to that of step 701, except that the "first prompt" in step 701 is replaced by the "third prompt" in step 707, and the "first-level question" in step 701 is replaced by the "third-level question" in step 707. The specific implementation method of step 707 will not be repeated here.
[0204] 708. The training device inputs the third-level question into the deep learning model to obtain a third predicted answer corresponding to the third-level question.
[0205] 709. The training device trains the deep learning model according to the first loss function. In the third training stage, the first loss function indicates the similarity between the third predicted answer and the third expected answer corresponding to the third level question.
[0206] In the embodiment of the present application, the meanings of the nouns in steps 708 and 709 and the specific implementation methods of the steps can be referred to the descriptions in steps 702 and 703. The difference is that the "first-level question" in steps 702 and 703 is replaced by "third-level question", the "first predicted answer" in steps 702 and 703 is replaced by "third predicted answer", and the "first expected answer" in steps 702 and 703 is replaced by "third expected answer". The specific implementation methods of steps 705 and 709 are not repeated here.
[0207] It should be noted that steps 707 to 709 are optional steps. If steps 707 to 709 are not executed, step 710 can be executed directly after executing step 707; if steps 701 to 709 are not executed, step 710 can be executed directly; if steps 701 to 709 are executed, then in the process of the training device executing steps 707 to 709 multiple times, steps 701 to 703 can also be executed crosswise, and / or steps 704 to 706 can be executed crosswise.
[0208] 710. In the fourth training stage, the training device obtains the fourth level of questions corresponding to the traffic scenario.
[0209] In the embodiment of the present application, illustratively, the "fourth-level question" includes at least the initial feature information of the fourth prompt, and the "fourth-level question" may also include the first feature information. The meaning of the "first feature information" can be found in the description of the embodiment corresponding to FIG3 above, and will not be repeated here. Exemplarily, the fourth prompt may include at least one mask (MASK), and the content of the MASK is predicted by a deep learning model, that is, the content of the MASK in the fourth prompt is completed by a deep learning model.
[0210] Since the fourth training stage is to allow the deep learning model to learn the causes and consequences and logical relationships of interactive behaviors between different objects, the content of the MASK predicted by the deep learning model may include the reasons for the interactive behaviors; optionally, in the fourth training stage, the content of the MASK predicted by the deep learning model may also include interactive behaviors.
[0211] For example, the "fourth prompt" can be specifically expressed as "Social car 1 and social car 3 have an interactive game behavior, the category is giving way, because ()", the part in "()" is MASK, that is, the part that needs to be predicted by the deep learning model. The above example predicts the reason for the interactive behavior of giving way through the deep learning model. It should be noted that the various examples here are only for the convenience of understanding the concept of the "fourth prompt" and are not used to limit this solution.
[0212] The specific implementation method of step 710 is similar to that of step 701, except that the "first prompt" in step 701 is replaced by the "fourth prompt" in step 710, and the "first-level question" in step 701 is replaced by the "fourth-level question" in step 710. The specific implementation method of step 710 will not be repeated here.
[0213] 711. The training device inputs a fourth-level question into the deep learning model to obtain a fourth predicted answer corresponding to the fourth-level question.
[0214] 712. The training device trains the deep learning model according to the first loss function. In the fourth training stage, the first loss function indicates the similarity between the fourth predicted answer and the fourth expected answer corresponding to the fourth level question.
[0215] In the embodiment of the present application, the meanings of the nouns in steps 711 and 712 and the specific implementation methods of the steps can be referred to the descriptions in steps 702 and 703. The difference is that the "first-level question" in steps 702 and 703 is replaced by "fourth-level question", the "first predicted answer" in steps 702 and 703 is replaced by "fourth predicted answer", and the "first expected answer" in steps 702 and 703 is replaced by "fourth expected answer". The specific implementation methods of steps 705 and 712 are not repeated here.
[0216] It should be noted that steps 710 to 712 are optional steps. If steps 710 to 712 are not executed, the retraining of the deep learning model can be stopped directly after step 710 is executed; if steps 701 to 712 are all executed, then in the process of the training device executing steps 710 to 712 multiple times, steps 701 to 703 can be executed crosswise, and / or steps 704 to 706 can be executed crosswise, and / or steps 707 to 709 can be executed crosswise.
[0217] In order to more intuitively understand the "first prompt", "second prompt", "third prompt" and "fourth prompt", the following is an example of a traffic scene, please refer to Figure 9, which is a schematic diagram of a traffic scene provided by an embodiment of the present application. As shown in Figure 9, Figure 9 shows a traffic scene of "the vehicle changing lanes and overtaking". For example, in this traffic scene, the first prompt used in the first training stage may include: in the (0) frame, the (driving speed) of (the vehicle) is (15m / s), and for another example, the first prompt may include: in the (-2) frame, the (x coordinate position) of (obstacle 1) is (16.8m). It should be noted that the content in "()" can be set to [MASK]. The content of [MASK] is the content that needs to be predicted by the deep learning model. In each example of the first prompt mentioned above, there are multiple "()". In each actual training process, the content in one of the multiple "()" can be set to [MASK].
[0218] For example, the second prompt used in the second training stage may include: from the (-10)th frame to the (0)th frame, (the vehicle) has (changed lanes to the left) in (lateral behavior); for another example, from the (-10)th frame to the (0)th frame, (the vehicle) has (slightly accelerated) in (longitudinal behavior). The content of one of the multiple "()"s included in each of the aforementioned second prompts can be set to [MASK], and the deep learning model is used to predict the content of [MASK]. From the aforementioned examples, it can be seen that the second training stage is to allow the machine learning model to learn the behavior of the object and the relationship between different objects.
[0219] For example, the third prompt used in the third training stage may include: from the (-9)th frame to the (-2)th frame, (the self-car) and (the social car 1) had an interactive game behavior, the category is (lane change and overtaking), specifically: (the self-car changes lanes to the left and overtakes the social car 1), the content of one of the multiple "()"s included in each of the aforementioned second prompts can be set to [MASK], and the deep learning model is used to predict the content of [MASK]. From the above examples, it can be seen that the third training stage is to allow the machine learning model to learn the interactive behaviors between different objects.
[0220] For example, the fourth prompt used in the fourth training stage may include: in the past 3 seconds, the self-vehicle changed lanes to overtake because (the private vehicle 1 in front of the self-vehicle was driving slowly and the left lane was idle. In order to improve driving efficiency, the self-vehicle changed lanes to overtake). The content in "()" can be set to [MASK]. From the above examples, it can be seen that the fourth training stage is to allow the machine learning model to learn the causal relationship of interactive behavior. It should be understood that the example in Figure 9 is only for the convenience of understanding this solution and is not used to limit this solution.
[0221] In an embodiment of the present application, the training process of the deep learning model may include at least two training stages, and the aforementioned at least two training stages may include at least two of the first training stage, the second training stage, the third training stage, and the fourth training stage. The specific training stages to be adopted for the deep learning model can be flexibly determined in combination with the actual application scenario, and are not limited in the embodiment of the present application. For example, the "first training stage" in this application may refer to any one of the above-mentioned first training stage, the second training stage, the third training stage, and the fourth training stage, and the "second training stage" and the "first training stage" in this application are different training stages.
[0222] The third question represents the question used in the first training phase. After determining which training phase the "first training phase" is, it can be determined whether the third question represents one of the first-level, second-level, third-level, or fourth-level questions. For example, if the first training phase is the first training phase, the third question represents a first-level question; for another example, if the first training phase is the second training phase, the third question represents a second-level question; for another example, if the first training phase is the third training phase, the third question represents a third-level question; for another example, if the first training phase is the fourth training phase, the third question represents a fourth-level question.
[0223] Correspondingly, the fourth question represents the question used in the second training stage. After determining which training stage the "second training stage" is, it can be determined which of the above-mentioned first-level questions, second-level questions, third-level questions and fourth-level questions the third question represents.
[0224] Exemplarily, at least two training stages may include a first training stage and a second training stage; or, at least two training stages may include a second training stage and a third training stage; or, at least two training stages may include a third training stage and a fourth training stage; or, at least two training stages may include a first training stage, a second training stage, and a third training stage; or, at least two training stages may include a first training stage, a second training stage, and a fourth training stage; or, at least two training stages may include a second training stage, a third training stage, and a fourth training stage; or, at least two training stages may include a first training stage, a second training stage, a third training stage, and a fourth training stage, etc. The specific training stages used in the training process of the deep learning model can be flexibly determined in combination with the actual application scenario, and are not limited in the embodiments of the present application.
[0225] In an embodiment of the present application, at least two training stages are used to train the deep learning model. The aforementioned at least two training stages respectively use the third question and the fourth question. The third question and the fourth question are used to obtain information at different levels from the traffic scene. That is, the training process of the deep learning model adopts a layer-by-layer progressive training method, which is conducive to enabling the deep learning model to learn logical thinking of step-by-step reasoning, thereby improving the human-likeness of the deep learning model, and also helping the trained deep learning model to be more reasonable when performing autonomous driving related tasks.
[0226] 2. Application Phase
[0227] In the embodiment of the present application, specifically, please refer to FIG10 , which is another flowchart of the information processing method provided in the embodiment of the present application. The information processing method provided in the present application may include:
[0228] 1001. The execution device inputs a first question into the deep learning model to obtain a first answer corresponding to the first question.
[0229] In the embodiment of the present application, step 1001 is an optional step. For example, the first information in the present application can be represented by a first question, and the second information in the present application can be a second answer corresponding to the second question. Optionally, before inputting the first information into the deep learning model, the training device can also input the first question into the deep learning model to obtain a first answer corresponding to the first question.
[0230] Exemplarily, the first question may include a prompt. Optionally, the first question may also include first characteristic information; the first characteristic information includes characteristic information of the environmental information around the vehicle. For a detailed introduction to the meaning of "first characteristic information" and "vehicle generating first characteristic information", please refer to the description in the corresponding embodiment of Figure 3, which will not be repeated here.
[0231] The first question and the second question may correspond to the same traffic scene (hereinafter referred to as the "target traffic scene" for the convenience of description). Optionally, the first question and the second question may be used to obtain information of different levels from the target traffic scene. For the concept of "different levels", please refer to the description in the embodiment corresponding to Figure 7 above. For example, the first question is used to obtain basic attribute information and semantic information in the target traffic scene. The meaning of the aforementioned "basic attribute information and semantic information" may refer to the description in the embodiment corresponding to Figure 7 above. For another example, the first question is used to obtain the behavior of the object in the target traffic scene or the relationship between different objects, etc. The specific manifestation of the "first question" can be determined in combination with the actual application scenario and is not limited here.
[0232] Exemplarily, the execution device refers to a device deployed with a machine learning model that has performed trained operations. In one case, the execution device can be specifically manifested as a vehicle, that is, a trained deep learning model is deployed in the vehicle; in another case, the execution device can be a device other than a vehicle, in which case the vehicle can send the first question to the execution device, and the execution device inputs the first question into the deep learning model; in another case, the execution device can be a device other than a vehicle, or after the vehicle sends the first information to the execution device, the execution device generates the first question and inputs the first question into the deep learning model. Exemplarily, after inputting the first question into the deep learning model, the execution device can obtain the prediction information generated by the deep learning model, and determine the first answer corresponding to the first question based on the prediction information generated by the deep learning model.
[0233] 1002. The execution device determines a second question based on the first answer, where the second question contains information in the first answer.
[0234] In the embodiment of the present application, step 1002 is an optional step. After determining the first answer corresponding to the first question, the execution device can generate a second question based on the first answer. For example, after determining the first answer, the execution device can update the second question generated for the vehicle (which can also be understood as the "second question before the update") based on the first answer to obtain an updated second question. The "updated second question" can also be understood as the "first information."
[0235] Exemplarily, the execution device may perform feature extraction on all or part of the information included in the first answer, and after obtaining the initial feature information of all or part of the information included in the first answer, add the aforementioned initial feature information to the prompt of the second question before the update to obtain the updated second question.
[0236] For example, the second answer to the second question can correspond to any of the following tasks: identifying risky obstacles around the ego vehicle, identifying the behavior of objects around the ego vehicle, predicting the behavior of objects around the ego vehicle, predicting the trajectory of objects around the ego vehicle, making decisions about the ego vehicle's behavior, planning the trajectory of the ego vehicle, or controlling the ego vehicle. In other words, the second question is used to obtain information corresponding to the aforementioned tasks, and thus the second question can correspond to any of the aforementioned tasks.
[0237] For example, the second answer corresponds to the task of "identifying the behavior of objects around the vehicle." The semantics of the prompt for the second question before the update generated by the vehicle may include: From frame () to frame (), the obstacle in an interactive game relationship with the vehicle is (), and the specific behavior of the obstacle is (). The information carried in the first answer may include: From frame -26 to frame -5, the obstacle in an interactive game relationship with the vehicle is obstacle 3. The semantics of the prompt in the updated second question may include: From frame -26 to frame -5, the obstacle in an interactive game relationship with the vehicle is obstacle 3, and the specific behavior of the obstacle is (); the content of "()" in the updated second question is the content that needs to be predicted by the deep learning model.
[0238] For another example, the second answer corresponds to the task of "predicting the behavior of objects around the vehicle." The semantics of the prompt for the second question before the update generated by the vehicle may include: In frame 0, the topology of obstacle 1 is (); Based on the topological relationship of obstacle 1 and the historical frame behavior, the navigation behavior of obstacle 1 is (), because (). The information carried in the first answer may include: The topology of obstacle 1 is topology 1. The semantics of the prompt in the updated second question may include: In frame 0, the topology of obstacle 1 is topology 1; Based on the topological relationship of obstacle 1 and the historical frame behavior, the navigation behavior of obstacle 1 is (), because (); The content of "()" in the updated second question is the content that needs to be predicted by the deep learning model.
[0239] For another example, the second answer corresponds to the task of "predicting the behavior of objects around the vehicle". The semantics of the prompt for the second question before the update generated by the vehicle may include: in frame 0, the object that has an interaction relationship with obstacle 1 is (). According to the topological relationship between obstacle 1 and () and the historical frame behavior, the interaction behavior that obstacle 1 will have is (). The information carried in the first answer may include: the object that has an interaction relationship with obstacle 1 is obstacle 2. The semantics of the prompt in the updated second question may include: in frame 0, the object that has an interaction relationship with obstacle 1 is obstacle 2. According to the topological relationship between obstacle 1 and obstacle 2 and the historical frame behavior, the interaction behavior that obstacle 1 will have is (); the content of "()" in the updated second question is the content that needs to be predicted by the deep learning model.
[0240] For another example, the second answer corresponds to the task of "making decisions about the vehicle's behavior." The semantics of the prompt for the second question before the update generated by the vehicle may include: In frame 0, the vehicle's topology is (), and the obstacles with which it interacts are (). Based on the topological relationship between the vehicle and the obstacle with which it interacts () and the historical frame behavior, the vehicle's next behavior is (), and the reason for deciding to execute the aforementioned interactive behavior is (). The information carried in the first answer may include: In frame 0, the vehicle's topology is topology 2, and the obstacles with which it interacts are obstacle 4. The semantics of the prompt in the updated second question may include: In frame 0, the vehicle's topology is topology 2, and the obstacles with which it interacts are obstacle 4. Based on the topological relationship between the vehicle and the obstacle with which it interacts (obstacle 4) and the historical frame behavior, the vehicle's next behavior is (), and the reason for deciding to execute the aforementioned interactive behavior is (); the content of "()" in the updated second question requires prediction by the deep learning model.
[0241] It should be noted that the above examples of "the first text in the first message (which can also be understood as the prompt in the second question)", "the information carried in the first answer" and "the updated second question" are only for the convenience of understanding the relationship between "the first text information", "the information in the first answer" and "the updated second question", and are not used to limit this solution.
[0242] In an embodiment of the present application, a first answer to the first question is first obtained, and then a second question is generated based on the first answer, and then the answer to the second question is obtained. The first question and the second question correspond to the same traffic scene, that is, the questions to the deep learning model are asked in a step-by-step manner, which is conducive to reducing the difficulty of the deep learning model in answering the second question, and is also conducive to improving the ability of the deep learning model in dealing with problems corresponding to complex traffic scenes; in addition, the use of a step-by-step questioning method is also conducive to incorporating the logical thinking of step-by-step reasoning into the deep learning model used in the field of autonomous driving, thereby improving the human-likeness of the second answer finally obtained.
[0243] 1003. The execution device inputs first information corresponding to the traffic scene around the vehicle into the deep learning model, and obtains second information corresponding to the first information, wherein the first information includes first text information.
[0244] In an embodiment of the present application, since steps 1001 and 1002 are optional steps, if steps 1001 and 1002 are not performed, in step 1003, the first text information in the "first information" is generated by the vehicle; if steps 1001 and 1002 are performed, in step 1003, the "first information" can be understood as the "updated second question", and the first text information in the first information is determined based on the second question and the first answer before the update generated by the vehicle.
[0245] For example, if the first text information includes initial feature information of a preset text corresponding to the first task performed by the deep learning model, the first text information may include the initial feature information of the preset text and the initial feature information of the information in the first answer. Alternatively, if the first text information includes initial feature information of text input by the user, the first text information may include the initial feature information of the text input by the user and the initial feature information of the information in the first answer.
[0246] In an embodiment of the present application, two situations in which the first text information may include information are listed, which improves the implementation flexibility of the present solution; when the first text information includes information of a preset text corresponding to the first task performed by the deep learning model, it is beneficial to quickly determine the first text information after determining the first task, so as to improve the efficiency of obtaining the second information; when the first text information includes information of a text input by the user, that is, questions can be answered based on the user's questions, which is beneficial to improve the user stickiness of the present solution and makes it more convenient for users to understand the behavior of the vehicle during the automatic driving process.
[0247] For example, step 1003 may include: inputting first information corresponding to the traffic scene surrounding the vehicle into the deep learning model by the execution device, thereby obtaining first prediction information generated by the deep learning model; and determining second information corresponding to the first information based on the first prediction information. The specific implementation of the aforementioned steps can be found in the description of the embodiment corresponding to FIG3 and is not further described here. The specific representation of the "first prediction information" is similar to the specific representation of the "third prediction information" in the embodiment corresponding to FIG7 and can be understood by reference thereto, so it is not further described here.
[0248] Exemplarily, the first information is the second question (or the updated second question), and the second information is the second answer corresponding to the second question (the updated second question). The second answer corresponds to any of the following tasks: identifying risk obstacles around the vehicle, identifying the behavior of objects around the vehicle, predicting the behavior of objects around the vehicle, predicting the trajectory of objects around the vehicle, making decisions on the behavior of the vehicle, planning the trajectory of the vehicle, controlling the vehicle, or other tasks in the process of autonomous driving, etc., which are not limited here. In the embodiment of the present application, a variety of tasks performed by a deep learning model are provided, which improves the implementation flexibility of the present solution.
[0249] Specifically, in one implementation, the execution device inputs a second question into the deep learning model, and obtains a first prediction information corresponding to the aforementioned second question generated by the deep learning model; and then, a second answer corresponding to the second question can be determined based on the aforementioned first prediction information.
[0250] Optionally, in one case, the second answer is included in a second preset answer set corresponding to the second question, and the second answer is obtained from at least one alternative answer generated from the deep learning model based on the second preset answer set. Exemplarily, a second preset answer set may be deployed in the execution device, and the second preset answer set includes one or more answers. After the execution device obtains the first prediction information, it may determine the intersection between the at least one alternative answer indicated by the first prediction information and the second preset answer set, where the aforementioned intersection includes at least one first alternative answer; the execution device may determine the second answer from at least one first alternative answer based on the first probability value corresponding to each first alternative answer in the at least one first alternative answer, and the first probability value corresponding to the second answer is the largest among the at least one first alternative answer.
[0251] In order to further understand the concept of "the intersection between at least one alternative answer and the second preset answer set", for example, the semantics of the second question is: Based on the given first feature information, what is the next lateral action decision of the vehicle? (); wherein, the content in the aforementioned () is the content that needs to be predicted by the deep learning model. The second preset answer set may include: left lane change, left bypass, hold, right bypass and right lane change; the at least one alternative answer indicated by the first prediction information may include: acceleration, left lane change, takeoff, hold and right lane change; the intersection between at least one alternative answer and the second preset answer set may include: left lane change, hold and right lane change. It should be understood that the examples here are only for the convenience of understanding this solution and are not used to limit this solution.
[0252] In an embodiment of the present application, since the deep learning model may sometimes give irrelevant answers or speak nonsense, that is, at least one alternative answer determined by the deep learning model may contain an answer that is completely irrelevant to the question, in order to avoid the occurrence of the aforementioned problem, a second preset answer set corresponding to the second question can be deployed, and at least one answer included in the second preset answer set corresponding to the second question is an answer related to the second question. The second answer corresponding to the second question is finally determined based on the aforementioned second preset answer set and at least one alternative answer determined based on the deep learning model, which is conducive to avoiding the problem that the final second answer is irrelevant to the second question, thereby improving the accuracy of the second answer obtained by the deep learning model.
[0253] Optionally, in another case, the second answer is obtained by performing a classification network that has been trained, and the second feature information includes feature information of the start flag [CLS] bit in the feature information of the alternative answers. The aforementioned classification network is used to determine the second answer corresponding to the second feature information from at least one preset answer; the aforementioned classification network is used to determine a predicted category to which the second feature information belongs in M categories, and the M categories can be the at least one preset answer, that is, the aforementioned classification network is used to determine a second answer to which the second feature information belongs in at least one answer. Exemplarily, a trained classification network can be deployed in the execution device. After obtaining the first prediction information corresponding to the second question generated by the deep learning model, the execution device can determine the feature information of the second alternative answer from the first prediction information, and the first probability value of the second alternative answer is the highest among the at least one alternative answer indicated by the first prediction information; the execution device can obtain the second feature information from the feature information of the second alternative answer, input the second feature information into the classification network, and obtain the second answer generated by the classification network.
[0254] In an embodiment of the present application, since the characteristic information of the starting flag [CLS] bit (that is, the second characteristic information) can represent the overall meaning of the alternative answers, that is, it is reasonable to use the characteristic information of the starting flag [CLS] bit to replace the characteristic information of the alternative answers, the classification network is used to determine a second answer belonging to the second characteristic information from at least one preset answer, and the second answer is limited to at least one preset answer, which is conducive to avoiding the problem that the final second answer is irrelevant to the second question, thereby improving the accuracy of the second answer obtained by the deep learning model.
[0255] In another case, after obtaining the first prediction information corresponding to the second question, that is, obtaining at least one alternative answer indicated by the first prediction information and the first probability value corresponding to each alternative answer, the execution device can obtain a second answer from at least one alternative answer, and the first probability value corresponding to the second answer is the highest among the at least one alternative answer indicated by the first prediction information.
[0256] In another implementation, the execution device inputs first information corresponding to the traffic scene around the vehicle into the deep learning model, including: the execution device inputs multiple second questions into the deep learning model; illustratively, the multiple second questions correspond to the same task, that is, the semantics of the multiple second questions are similar, and different second questions in the multiple second questions carry different prompts. The execution device obtains the second information corresponding to the first information, including: the execution device obtains multiple reference answers corresponding to the multiple second questions one by one, that is, the execution device can obtain a reference answer corresponding to each second question, and the multiple reference answers are all obtained through the deep learning model. The second answer can then be determined by the execution device or the first device based on the multiple reference answers. The relationship between the "first device" and the "execution device" can be found in the description of the embodiment corresponding to Figure 3, which will not be repeated here.
[0257] The process of "determining a second answer based on multiple reference answers" can include: determining the second answer based on the multiple reference answers using a majority voting strategy, i.e., the final second answer is included in the multiple reference answers. For a more intuitive understanding of this solution, please refer to Figure 11, which is a schematic diagram of multiple second questions provided in an embodiment of the present application. Figure 11 takes the first task of "making a decision about the vehicle" performed by the deep learning model as an example. Figure 11 shows three second questions, with the prompts for the three second questions being "Because...therefore, [CLS] the vehicle should ()," "Because...therefore, [CLS] the vehicle needs ()," and "Because...therefore, [CLS] the vehicle needs ()." As can be seen from the above description, the prompts for different second questions are different. The reference answers corresponding to the three second questions are "Change lane to overtake on the left," "Slow down and brake to stop," and "Change lane to overtake on the left." Based on the majority voting strategy, the final second answer is "Change lane to overtake on the left." It should be understood that the example in Figure 11 is only for facilitating understanding of this solution and is not intended to limit this solution.
[0258] Exemplarily, after inputting each of the multiple second questions into the deep learning model, the execution device can obtain multiple first prediction information generated by the deep learning model that corresponds one to one to the multiple second questions, and can determine a reference answer based on each first prediction information. It should be noted that the process of "determining a reference answer based on each first prediction information" can refer to the above-mentioned process of "determining the second answer based on the first prediction information", the difference being that the above-mentioned "second answer" is replaced by "alternative answer".
[0259] That is, in this implementation, in one case, the reference answer (i.e., the second answer) is included in the second preset answer set corresponding to the second question, and the reference answer is obtained from at least one alternative answer generated from the deep learning model based on the second preset answer set. In another case, the reference answer (i.e., the second answer) is obtained by performing a classification network that has been trained, and the second feature information includes feature information of the start flag [CLS] bit in the feature information of the alternative answer. The aforementioned classification network is used to determine the reference answer corresponding to the second feature information from at least one preset answer. In another case, the reference answer (i.e., the second answer) is obtained directly from at least one alternative answer indicated by the first prediction information.
[0260] In an embodiment of the present application, multiple second questions are input into the deep learning model, and different second questions carry different prompts. Multiple reference answers are obtained by using multiple different prompts, and then the final second answer can be obtained from the multiple reference answers, which improves the rigor of obtaining the final second answer and is conducive to improving the accuracy of the final two questions.
[0261] On the basis of the embodiments corresponding to Figures 1 to 11, in order to better implement the above-mentioned solutions of the embodiments of the present application, relevant equipment for implementing the above-mentioned solutions is also provided below. Please refer to Figure 12 in detail, which is a structural diagram of the information processing device provided in the embodiment of the present application. The information processing device 1200 includes: an input module 1201, which is used to input first information corresponding to the traffic scene around the vehicle into the deep learning model, and the first information includes first text information; an acquisition module 1202, which is used to obtain second information corresponding to the first information, wherein the second information is obtained based on the deep learning model, and the second information corresponds to any of the following tasks: making decisions on the behavior of the vehicle, planning the trajectory of the vehicle, or controlling the vehicle.
[0262] Optionally, the first text information includes information of a preset text corresponding to the task, or the first text information includes information of a text input by a user.
[0263] Optionally, the input module 1201 is also used to input a first question corresponding to the traffic scenario into the deep learning model, and the first question is used to obtain a first answer corresponding to the first question, and the first answer is obtained through the deep learning model, wherein the first information is the second question, the second information is the second answer corresponding to the second question, and the information in the first answer exists in the first text information.
[0264] Optionally, the first information further includes first characteristic information, the first characteristic information includes characteristic information of environmental information around the vehicle, and the environmental information includes physical property information of objects around the vehicle.
[0265] Optionally, the first feature information is obtained based on a feature extraction network, wherein the feature extraction network belongs to a first neural network, and the first neural network is used to perform at least two of the following tasks: trajectory prediction of objects around the vehicle, decision-making on the vehicle's behavior, trajectory planning for the vehicle, prediction of the vehicle's speed range, or control of the vehicle.
[0266] Optionally, the first information is the second question, and the second information is the second answer corresponding to the second question, wherein the second answer is included in a preset answer set corresponding to the second question, and the second answer is obtained from at least one alternative answer generated from the deep learning model based on the preset answer set.
[0267] Optionally, the first information is the second question, the second information is the second answer corresponding to the second question, the second answer is obtained through the classification network, the second feature information includes the feature information of the starting flag [CLS] bit in the feature information of the second answer, and the classification network is used to determine the second answer corresponding to the second feature information from at least one preset answer.
[0268] Optionally, the first information is the second question, and the second information is the second answer corresponding to the second question; the input module 1201 is specifically used to input multiple second questions into the deep learning model, and different second questions in the multiple second questions carry different prompts; the acquisition module 1202 is specifically used to obtain multiple reference answers corresponding one-to-one to the multiple second questions, and determine the second answer based on the multiple reference answers, and the multiple reference answers are all obtained through the deep learning model.
[0269] It should be noted that the information interaction, execution process, etc. between the modules / units in the information processing device 1200 are based on the same concept as the various method embodiments corresponding to Figures 3 to 11 in this application. For specific contents, please refer to the description in the method embodiments shown above in this application, and will not be repeated here.
[0270] Please continue to refer to Figure 13, which is another schematic diagram of the structure of the information processing device provided in an embodiment of the present application. Information processing device 1300 includes: an input module 1301 for inputting a first question into a deep learning model and obtaining a first answer corresponding to the first question; a determination module 1302 for determining a second question based on the first answer, wherein the second question carries information from the first answer; input module 1301 is also used to input a second question into the deep learning model and obtain a second answer corresponding to the second question, wherein the first question and the second question correspond to the same traffic scenario.
[0271] Optionally, the second answer corresponds to any of the following tasks: identifying risk obstacles around the vehicle, identifying the behavior of objects around the vehicle, predicting the behavior of objects around the vehicle, predicting the trajectory of objects around the vehicle, making decisions on the behavior of the vehicle, planning the trajectory of the vehicle, or controlling the vehicle.
[0272] It should be noted that the information interaction, execution process, etc. between the modules / units in the information processing device 1300 are based on the same concept as the various method embodiments corresponding to Figures 3 to 11 in this application. For specific contents, please refer to the description in the method embodiments shown above in this application, and will not be repeated here.
[0273] Please continue to refer to Figure 14, which is another structural diagram of the information processing device provided in an embodiment of the present application. The information processing device 1400 is used to train a deep learning model, and the deep learning model includes at least two training stages, and the at least two training stages include a first training stage and a second training stage. The information processing device 1400 includes: an input module 1401, which is used to input a third question into the deep learning model in the first training stage to obtain a predicted answer corresponding to the third question; the input module 1401 is also used to input a fourth question into the deep learning model in the second training stage to obtain a predicted answer corresponding to the fourth question; wherein, a first loss function is used when training the deep learning model. In the first training stage, the first loss function indicates the similarity between the predicted answer corresponding to the third question and the expected answer. In the second training stage, the first loss function indicates the similarity between the predicted answer corresponding to the fourth question and the expected answer. The third question and the fourth question are both related to traffic scenes, and the third question and the fourth question are used to obtain information at different levels.
[0274] Optionally, the third question includes first feature information, the first feature information includes feature information of environmental information around the vehicle, and the environmental information includes physical property information of objects around the vehicle.
[0275] Optionally, the information processing device 1400 also includes: a feature extraction module 1402, which is used to input environmental information into a feature extraction network to obtain second feature information generated by the feature extraction network, and the second feature information is used to obtain the first feature information; a feature processing module 1403, which is used to input the first feature information into a feature processing network to obtain prediction information generated by the feature processing network, and the feature extraction network and the feature processing network belong to the same neural network, and the neural network is used to perform at least the following tasks: trajectory prediction of objects around the vehicle, decision-making on the behavior of the vehicle, trajectory planning of the vehicle, prediction of the speed range of the vehicle, or control of the vehicle; wherein, the training stage of the deep learning model includes using a first loss function and a second loss function to train the deep learning model and the neural network, and the second loss function indicates the similarity between the prediction information corresponding to the environmental information and the second expected information.
[0276] It should be noted that the information interaction, execution process, etc. between the modules / units in the information processing device 1400 are based on the same concept as the various method embodiments corresponding to Figures 3 to 11 in this application. For specific contents, please refer to the description in the method embodiments shown above in this application, and will not be repeated here.
[0277] Next, a device provided in an embodiment of the present application will be described. Please refer to Figure 15, which is a schematic structural diagram of a device provided in an embodiment of the present application. Specifically, the device 1500 includes: a receiver 1501, a transmitter 1502, a processor 1503, and a memory 1504 (wherein the number of processors 1503 in the device 1500 may be one or more, and Figure 15 uses one processor as an example). The processor 1503 may include an application processor 15031 and a communication processor 15032. In some embodiments of the present application, the receiver 1501, the transmitter 1502, the processor 1503, and the memory 1504 may be connected via a bus or other means.
[0278] Memory 1504 may include read-only memory and random access memory, and provides instructions and data to processor 1503. A portion of memory 1504 may also include non-volatile random access memory (NVRAM). Memory 1504 stores processor and operation instructions, executable modules, or data structures, or subsets or extended sets thereof. The operation instructions may include various operation instructions for implementing various operations.
[0279] Processor 1503 controls the operation of the device. In specific applications, the various components of the device are coupled together via a bus system. In addition to a data bus, the bus system may also include a power bus, a control bus, and a status signal bus. However, for clarity, all bus systems are referred to as a bus system in the figure.
[0280] The methods disclosed in the above embodiments of the present application can be applied to or implemented by processor 1503. Processor 1503 can be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above method can be completed by hardware integrated logic circuits or software instructions in processor 1503. The above processor 1503 can be a general-purpose processor, a digital signal processor (DSP), a microprocessor, or a microcontroller, and can further include an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. The processor 1503 can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the methods disclosed in conjunction with the embodiments of the present application can be directly implemented as being executed by a hardware decoding processor, or can be executed by a combination of hardware and software modules in the decoding processor. The software module can be located in a storage medium well-known in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, etc. The storage medium is located in memory 1504, and processor 1503 reads the information in memory 1504 and, in conjunction with its hardware, completes the steps of the above method.
[0281] Receiver 1501 can be used to receive input digital or character information and generate signal input related to device settings and function control. Transmitter 1502 can be used to output digital or character information through the first interface. Transmitter 1502 can also be used to send instructions to the disk pack through the first interface to modify data on the disk pack. Transmitter 1502 can also include a display device such as a display screen.
[0282] In the embodiment of the present application, in one case, the processor 1503 is used to execute the first device in the embodiment corresponding to Figures 3 to 11 and the method executed by the execution device; in another case, the processor 1503 is used to execute the method executed by the training device in the embodiment corresponding to Figures 3 to 11. It should be noted that the specific manner in which the application processor 15031 in the processor 1503 executes the aforementioned steps is based on the same concept as the various method embodiments corresponding to Figures 3 to 11 in the present application, and the technical effects it brings are the same as the various method embodiments corresponding to Figures 3 to 11 in the present application. For specific content, please refer to the description of the method embodiments shown above in the present application, and no further details will be given here.
[0283] The present application also provides a vehicle. Please refer to FIG16 , which is a schematic structural diagram of a vehicle provided by the present application. Vehicle 100 is configured for a fully or partially autonomous driving mode. For example, vehicle 100 can control itself while in autonomous driving mode, and can determine the current state of the vehicle and its surrounding environment through human operation, determine the possible behavior of at least one other vehicle in the surrounding environment, and determine the confidence level corresponding to the possibility of the other vehicle performing the possible behavior, and control vehicle 100 based on the determined information. When vehicle 100 is in autonomous driving mode, vehicle 100 can also be set to operate without human interaction.
[0284] The vehicle 100 may include various subsystems, such as a travel system 102, a sensor system 104, a control system 106, one or more peripheral devices 108, a power source 110, a computer system 112, and a user interface 116. Alternatively, the vehicle 100 may include more or fewer subsystems, and each subsystem may include multiple components. In addition, each subsystem and component of the vehicle 100 may be interconnected via wired or wireless connections.
[0285] Travel system 102 may include components that provide powered movement for vehicle 100. In one embodiment, travel system 102 may include engine 118, power source 119, transmission 120, and wheels / tires 121.
[0286] The engine 118 may be an internal combustion engine, an electric motor, an air compression engine, or a combination of other types of engines, such as a hybrid engine consisting of a gasoline engine and an electric motor, or a hybrid engine consisting of an internal combustion engine and an air compression engine. The engine 118 converts the energy source 119 into mechanical energy. Examples of the energy source 119 include gasoline, diesel, other petroleum-based fuels, propane, other compressed gas-based fuels, ethanol, solar panels, batteries, and other sources of electricity. The energy source 119 may also provide energy for other systems of the vehicle 100. The transmission 120 may transmit the mechanical power from the engine 118 to the wheels 121. The transmission 120 may include a gearbox, a differential, and a drive shaft. In one embodiment, the transmission 120 may also include other devices, such as a clutch. The drive shaft may include one or more shafts that can be coupled to one or more wheels 121.
[0287] Sensor system 104 may include several sensors that sense information about the environment surrounding vehicle 100. For example, sensor system 104 may include a positioning system 122 (the positioning system may be a global positioning system (GPS), a BeiDou system, or other positioning systems), an inertial measurement unit (IMU) 124, a radar 126, a laser rangefinder 128, and a camera 130. Sensor system 104 may also include sensors for internal systems of monitored vehicle 100 (e.g., an in-vehicle air quality monitor, a fuel gauge, an oil temperature gauge, etc.). Sensor data from one or more of these sensors may be used to detect objects and their corresponding characteristics (position, shape, direction, speed, etc.). This detection and recognition is a key function for the safe operation of autonomous vehicle 100.
[0288] Among them, the positioning system 122 can be used to estimate the geographic location of the vehicle 100. The IMU 124 is used to sense the position and orientation changes of the vehicle 100 based on inertial acceleration. In one embodiment, the IMU 124 can be a combination of an accelerometer and a gyroscope. The radar 126 can use radio signals to sense objects in the surrounding environment of the vehicle 100, and can specifically be a millimeter wave radar or a laser radar. In some embodiments, in addition to sensing objects, the radar 126 can also be used to sense the speed and / or direction of travel of objects. The laser rangefinder 128 can use lasers to sense objects in the environment in which the vehicle 100 is located. In some embodiments, the laser rangefinder 128 may include one or more laser sources, a laser scanner, and one or more detectors, as well as other system components. The camera 130 can be used to capture multiple images of the surrounding environment of the vehicle 100. The camera 130 can be a still camera or a video camera.
[0289] Control system 106 controls the operation of vehicle 100 and its components. Control system 106 may include various components, including a steering system 132 , a throttle 134 , a brake unit 136 , a computer vision system 140 , a lane control system 142 , and an obstacle avoidance system 144 .
[0290] The steering system 132 is operable to adjust the direction of travel of the vehicle 100. For example, in one embodiment, it may be a steering wheel system. The throttle 134 is used to control the operating speed of the engine 118 and, in turn, the speed of the vehicle 100. The brake unit 136 is used to control the deceleration of the vehicle 100. The brake unit 136 may use friction to slow the wheels 121. In other embodiments, the brake unit 136 may convert the kinetic energy of the wheels 121 into electrical current. The brake unit 136 may also take other forms to slow the rotation speed of the wheels 121 to control the speed of the vehicle 100. The computer vision system 140 is operable to process and analyze images captured by the camera 130 to identify objects and / or features in the environment surrounding the vehicle 100. These objects and / or features may include traffic signs, road boundaries, and obstacles. The computer vision system 140 may use object recognition algorithms, structure from motion (SFM) algorithms, video tracking, and other computer vision techniques. In some embodiments, the computer vision system 140 can be used to map the environment, track objects, estimate their speed, and so on. The route control system 142 is used to determine the route and speed of the vehicle 100. In some embodiments, the route control system 142 may include a lateral planning module 1421 and a longitudinal planning module 1422, which are respectively used to determine the route and speed for the vehicle 100 by combining data from the obstacle avoidance system 144, GPS 122, and one or more predetermined maps. The obstacle avoidance system 144 is used to identify, evaluate, and avoid or otherwise navigate obstacles in the environment of the vehicle 100. The aforementioned obstacles can specifically be represented by actual obstacles and virtual moving objects that may collide with the vehicle 100. In one embodiment, the control system 106 may include additional or alternative components other than those shown and described. Alternatively, some of the components shown above may be reduced.
[0291] Vehicle 100 interacts with external sensors, other vehicles, other computer systems, or users via peripheral devices 108. Peripheral devices 108 may include a wireless communication system 146, an onboard computer 148, a microphone 150, and / or a speaker 152. In some embodiments, peripheral devices 108 provide a means for the user of vehicle 100 to interact with user interface 116. For example, onboard computer 148 may provide information to the user of vehicle 100. User interface 116 may also operate onboard computer 148 to receive user input. Onboard computer 148 may be operated via a touchscreen. In other cases, peripheral devices 108 may provide a means for vehicle 100 to communicate with other devices located within the vehicle. For example, microphone 150 may receive audio (e.g., voice commands or other audio input) from the user of vehicle 100. Similarly, speaker 152 may output audio to the user of vehicle 100. Wireless communication system 146 may wirelessly communicate with one or more devices directly or via a communication network. For example, the wireless communication system 146 may utilize 3G cellular communications, such as CDMA, EVDO, GSM / GPRS, or 4G cellular communications, such as LTE. Or 5G cellular communications. The wireless communication system 146 may utilize wireless local area network (WLAN) communications. In some embodiments, the wireless communication system 146 may utilize infrared links, Bluetooth, or ZigBee to communicate directly with devices. Other wireless protocols, such as various vehicle communication systems, may include one or more dedicated short range communications (DSRC) devices, which may include public and / or private data communications between vehicles and / or roadside stations.
[0292] Power source 110 can provide power to various components of vehicle 100. In one embodiment, power source 110 can be a rechargeable lithium-ion or lead-acid battery. One or more battery packs of such batteries can be configured as a power source to provide power to various components of vehicle 100. In some embodiments, power source 110 and energy source 119 can be implemented together, such as in some all-electric vehicles.
[0293] Some or all functions of vehicle 100 are controlled by computer system 112. Computer system 112 may include at least one processor 113 that executes instructions 115 stored in a non-transitory computer-readable medium, such as memory 114. Computer system 112 may also be a plurality of computing devices that control individual components or subsystems of vehicle 100 in a distributed manner. Processor 113 may be any conventional processor, such as a commercially available central processing unit (CPU). Alternatively, processor 113 may be a specialized device, such as an application-specific integrated circuit (ASIC) or other hardware-based processor. Although FIG. 16 functionally illustrates the processor, memory, and other components of computer system 112 in the same block, those skilled in the art will appreciate that the processor or memory may actually include multiple processors or memories that are not stored in the same physical housing. For example, memory 114 may be a hard drive or other storage medium located in a different housing than computer system 112. Therefore, references to processor 113 or memory 114 should be understood to include references to a collection of processors or memories that may or may not operate in parallel. Rather than using a single processor to perform the steps described herein, some components, such as the steering assembly and the retarding assembly, may each have its own processor that performs only calculations related to the functionality of the component specific component.
[0294] In various aspects described herein, the processor 113 may be located remotely from the vehicle 100 and in wireless communication with the vehicle 100. In other aspects, some of the processes described herein are performed on the processor 113 disposed within the vehicle 100 while others are performed by the remote processor 113, including taking the necessary steps to perform a single maneuver.
[0295] In some embodiments, memory 114 may contain instructions 115 (e.g., program logic) that are executable by processor 113 to perform various functions of vehicle 100, including those described above. Memory 114 may also contain additional instructions, including instructions for sending data to, receiving data from, interacting with, and / or controlling one or more of travel system 102, sensor system 104, control system 106, and peripherals 108. In addition to instructions 115, memory 114 may also store data such as road maps, route information, the vehicle's location, direction, speed, and other such vehicle data, as well as other information. This information may be used by vehicle 100 and computer system 112 during operation of vehicle 100 in autonomous, semi-autonomous, and / or manual modes. A user interface 116 is provided for providing information to or receiving information from a user of vehicle 100. Optionally, user interface 116 may include one or more input / output devices within the set of peripherals 108, such as wireless communication system 146, onboard computer 148, microphone 150, and speaker 152.
[0296] Computer system 112 may control functions of vehicle 100 based on input received from various subsystems (e.g., travel system 102, sensor system 104, and control system 106) and from user interface 116. For example, computer system 112 may utilize input from control system 106 to control steering system 132 to avoid obstacles detected by sensor system 104 and obstacle avoidance system 144. In some embodiments, computer system 112 may be operable to provide control over many aspects of vehicle 100 and its subsystems.
[0297] Alternatively, one or more of the above components may be installed or associated separately from the vehicle 100. For example, the memory 114 may be partially or completely separate from the vehicle 100. The above components may be communicatively coupled together in a wired and / or wireless manner.
[0298] Optionally, the above components are just an example. In actual applications, the components in the above modules may be added or deleted according to actual needs. Figure 16 should not be understood as limiting the embodiments of the present application. A vehicle traveling on a road, such as vehicle 100 above, can identify objects in its surrounding environment to determine adjustments to the current speed. The objects can be other vehicles, traffic control devices, or other types of objects. In some examples, each identified object can be considered independently, and based on the respective characteristics of the object, such as its current speed, acceleration, distance from the vehicle, etc., it can be used to determine the speed to be adjusted for the vehicle.
[0299] Optionally, the vehicle 100 or a computing device associated with the vehicle 100, such as the computer system 112, computer vision system 140, and memory 114 of Figure 16, can predict the behavior of the identified objects based on the characteristics of the identified objects and the state of the surrounding environment (e.g., traffic, rain, ice on the road, etc.). Optionally, each identified object depends on the behavior of each other, so all identified objects can be considered together to predict the behavior of a single identified object. The vehicle 100 can adjust its speed based on the predicted behavior of the identified objects. In other words, the vehicle 100 can determine what stable state the vehicle will need to adjust to (e.g., accelerate, decelerate, or stop) based on the predicted behavior of the objects. In this process, other factors can also be considered to determine the speed of the vehicle 100, such as the lateral position of the vehicle 100 on the road it is traveling on, the curvature of the road, the proximity of static and dynamic objects, etc. In addition to providing instructions to adjust the speed of the vehicle, the computing device may also provide instructions to modify the steering angle of the vehicle 100 so that the vehicle 100 follows a given trajectory and / or maintains a safe lateral and longitudinal distance from objects near the vehicle 100 (e.g., cars in adjacent lanes on the road).
[0300] The vehicle 100 may be a car, truck, motorcycle, bus, ship, airplane, helicopter, lawn mower, recreational vehicle, amusement park vehicle, construction equipment, tram, golf cart, train, etc., and the embodiments of the present application do not impose any particular limitation thereto.
[0301] In the embodiment of the present application, the processor 113 in the vehicle 10 is used to execute the method executed by the first device and / or the execution device in the embodiments corresponding to Figures 3 to 11. It should be noted that the specific manner in which the processor 113 executes the aforementioned steps is based on the same concept as the various method embodiments corresponding to Figures 3 to 11 in this application, and the technical effects it brings are the same as the various method embodiments corresponding to Figures 3 to 11 in this application. For specific details, please refer to the description of the method embodiments shown above in this application, and will not be repeated here.
[0302] A computer-readable storage medium is also provided in an embodiment of the present application. The computer-readable storage medium stores a program, which, when executed on a computer, enables the computer to execute the steps executed by the first device and / or the execution device in the method described in the embodiments shown in Figures 3 to 11 above, or enables the computer to execute the steps executed by the training device in the method described in the embodiments shown in Figures 3 to 11 above.
[0303] Also provided in an embodiment of the present application is a computer program product comprising a program which, when executed on a computer, enables the computer to execute the steps executed by the first device and / or the execution device in the method described in the embodiments shown in the aforementioned Figures 3 to 11, or enables the computer to execute the steps executed by the training device in the method described in the embodiments shown in the aforementioned Figures 3 to 11.
[0304] A circuit system is also provided in an embodiment of the present application, which includes a processing circuit, and the processing circuit is configured to execute the steps performed by the first device and / or the execution device in the method described in the embodiments shown in Figures 3 to 11 above, or the processing circuit is configured to execute the steps performed by the training device in the method described in the embodiments shown in Figures 3 to 11 above.
[0305] The information processing device, equipment or vehicle provided in the embodiments of the present application may specifically be a chip, and the chip includes: a processing unit and a communication unit. The processing unit may be, for example, a processor, and the communication unit may be, for example, an input / output interface, a pin or a circuit. The processing unit may execute computer-executable instructions stored in the storage unit so that the chip in the server executes the method described in the embodiments shown in Figures 3 to 11 above. Optionally, the storage unit is a storage unit in the chip, such as a register, a cache, etc. The storage unit may also be a storage unit located outside the chip in the wireless access device, such as a read-only memory (ROM) or other types of static storage devices that can store static information and instructions, a random access memory (RAM), etc.
[0306] Specifically, see Figure 17 , which is a schematic diagram of the structure of a chip provided in an embodiment of the present application. The chip may be a neural network processor (NPU) 170. NPU 170 is mounted on a host CPU (host CPU) as a coprocessor, with tasks assigned by the host CPU. The core of the NPU is arithmetic circuit 1703, which is controlled by controller 1704 to extract matrix data from memory and perform multiplication operations.
[0307] In some implementations, arithmetic circuit 1703 includes multiple processing units (PEs). In some implementations, arithmetic circuit 1703 is a two-dimensional systolic array. Arithmetic circuit 1703 can also be a one-dimensional systolic array or other electronic circuitry capable of performing mathematical operations such as multiplication and addition. In some implementations, arithmetic circuit 1703 is a general-purpose matrix processor.
[0308] For example, assume there are input matrix A, weight matrix B, and output matrix C. The arithmetic circuit retrieves the corresponding data of matrix B from weight memory 1702 and caches it on each PE in the arithmetic circuit. The arithmetic circuit retrieves the data of matrix A from input memory 1701 and performs a matrix operation on matrix B. The partial or final matrix result is stored in accumulator 1708.
[0309] Unified memory 1706 is used to store input and output data. Weight data is directly transferred to weight memory 1702 through the Direct Memory Access Controller (DMAC) 1705. Input data is also transferred to unified memory 1706 through the DMAC.
[0310] BIU stands for Bus Interface Unit 1710 , which is used for interaction between the AXI bus, DMAC, and instruction fetch buffer (IFB) 1709 .
[0311] The bus interface unit 1710 (BIU) is used for the instruction fetch memory 1709 to obtain instructions from the external memory, and is also used for the storage unit access controller 1705 to obtain the original data of the input matrix A or the weight matrix B from the external memory.
[0312] DMAC is mainly used to move input data in the external memory DDR to the unified memory 1706 or to move weight data to the weight memory 1702 or to move input data to the input memory 1701.
[0313] The vector calculation unit 1707 includes multiple operation processing units. When necessary, it further processes the output of the operation circuit, such as vector multiplication, vector addition, exponential operation, logarithmic operation, size comparison, etc. It is mainly used for non-convolutional / fully connected layer network calculations in neural networks, such as batch normalization, pixel-level summation, and upsampling of feature planes.
[0314] In some implementations, the vector calculation unit 1707 can store the processed output vector to the unified memory 1706. For example, the vector calculation unit 1707 can apply a linear function and / or a nonlinear function to the output of the operation circuit 1703, such as linear interpolation of the feature plane extracted by the convolution layer, or accumulate a vector of values to generate an activation value. In some implementations, the vector calculation unit 1707 generates a normalized value, a pixel-level summed value, or both. In some implementations, the processed output vector can be used as an activation input to the operation circuit 1703, for example, for use in a subsequent layer in a neural network.
[0315] An instruction fetch buffer 1709 connected to the controller 1704 is used to store instructions used by the controller 1704;
[0316] Unified memory 1706, input memory 1701, weight memory 1702, and instruction fetch memory 1709 are all on-chip memories. External memories are private to the NPU hardware architecture.
[0317] Among them, the operations of each layer in the deep learning model, the first feature extraction network and the second feature extraction network shown in Figures 3 to 11 can be performed by the operation circuit 1703 or the vector calculation unit 1707.
[0318] The processor mentioned in any of the above places can be a general-purpose central processing unit, a microprocessor, an ASIC, or one or more integrated circuits for controlling the execution of the program of the above-mentioned first aspect method.
[0319] It should also be noted that the device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separate, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed across multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the present embodiment. In addition, in the drawings of the device embodiments provided in this application, the connection relationship between the modules indicates that there is a communication connection between them, which can be specifically implemented as one or more communication buses or signal lines.
[0320] Through the description of the above embodiments, those skilled in the art can clearly understand that the present application can be implemented by means of software plus necessary general hardware, and of course can also be implemented by special hardware including application-specific integrated circuits, special CPUs, special memories, special components, etc. In general, all functions performed by computer programs can be easily implemented with corresponding hardware, and the specific hardware structures used to implement the same function can also be diverse, such as analog circuits, digital circuits or special circuits, etc. However, for the present application, software program implementation is a better implementation method in most cases. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art can be embodied in the form of a software product, which is stored in a readable storage medium, such as a computer's floppy disk, USB flash drive, mobile hard disk, ROM, RAM, magnetic disk or optical disk, etc., and includes a number of instructions to enable a computer device (which can be a personal computer, training equipment, or network equipment, etc.) to execute the methods described in each embodiment of the present application.
[0321] In the above embodiments, all or part of the embodiments may be implemented by software, hardware, firmware, or any combination thereof. When implemented by software, all or part of the embodiments may be implemented in the form of a computer program product.
[0322] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function described in the embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from a website, a computer, a training device or a data center by wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) mode to another website, a computer, a training device or a data center. The computer-readable storage medium can be any available medium that a computer can store or a data storage device such as a training device, a data center, etc. that includes one or more available media integrations. The available medium can be a magnetic medium, (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state drive (SSD)).
Claims
1. An information processing method, characterized in that: The method comprises: Inputting first information corresponding to a traffic scene around the vehicle into the deep learning model, wherein the first information includes first text information; Obtain second information corresponding to the first information, wherein the second information is obtained based on the deep learning model, and the second information corresponds to any one of the following tasks: making decisions on the behavior of the vehicle, planning the trajectory of the vehicle, or controlling the vehicle.
2. The method according to claim 1, characterized in that The first text information includes information of a preset text corresponding to the task, or the first text information includes information of a text input by a user.
3. The method according to claim 1 or 2, characterized in that: The method further comprises: A first question corresponding to the traffic scenario is input into the deep learning model, the first question is used to obtain a first answer corresponding to the first question, the first answer is obtained by the deep learning model, wherein the first information is a second question, the second information is a second answer corresponding to the second question, and the information in the first answer is present in the first text information.
4. The method according to claim 1 or 2, characterized in that: The first information also includes first feature information, and the first feature information includes feature information of environmental information around the vehicle, and the environmental information includes physical attribute information of objects around the vehicle.
5. The method according to claim 4, characterized in that The first feature information is obtained based on a feature extraction network, wherein the feature extraction network belongs to a first neural network, and the first neural network is used to perform at least two of the following tasks: trajectory prediction of objects around the vehicle, decision-making on the behavior of the vehicle, trajectory planning of the vehicle, prediction of the speed range of the vehicle, or control of the vehicle.
6. The method according to claim 1 or 2, characterized in that: The first information is a second question, and the second information is a second answer corresponding to the second question, wherein the second answer is included in a preset answer set corresponding to the second question, and the second answer is obtained from at least one alternative answer generated from the deep learning model based on the preset answer set.
7. The method according to claim 1 or 2, characterized in that: The first information is the second question, the second information is the second answer corresponding to the second question, the second answer is obtained through a classification network, the second feature information includes feature information of a start flag [CLS] bit in the feature information of the second answer, and the classification network is used to determine the second answer corresponding to the second feature information from at least one preset answer.
8. The method according to claim 1 or 2, characterized in that: The first information is a second question, the second information is a second answer corresponding to the second question, and the inputting the first information corresponding to the traffic scene around the vehicle into the deep learning model includes: Inputting a plurality of the second questions into the deep learning model, wherein different second questions among the plurality of second questions carry different prompts; The acquiring second information corresponding to the first information includes: Obtaining a plurality of reference answers corresponding one-to-one to the plurality of second questions, wherein the plurality of reference answers are all obtained through the deep learning model; The second answer is determined based on the multiple reference answers.
9. An information processing method, characterized in that: The method comprises: Inputting a first question into the deep learning model to obtain a first answer corresponding to the first question; Determine a second question according to the first answer, wherein the second question carries the information in the first answer; The second question is input into the deep learning model to obtain a second answer corresponding to the second question, wherein the first question and the second question correspond to the same traffic scenario.
10. The method according to claim 9, characterized in that The second answer corresponds to any of the following tasks: identifying risk obstacles around the vehicle, identifying the behavior of objects around the vehicle, predicting the behavior of objects around the vehicle, predicting the trajectory of objects around the vehicle, making decisions on the behavior of the vehicle, planning the trajectory of the vehicle, or controlling the vehicle.
11. An information processing method, characterized in that: The method is used to train a deep learning model, the deep learning model includes at least two training stages, the at least two training stages include a first training stage and a second training stage, and the method includes: In the first training stage, a third question is input into the deep learning model to obtain a predicted answer corresponding to the third question; In the second training stage, a fourth question is input into the deep learning model to obtain a predicted answer corresponding to the fourth question; In which, a first loss function is used when training the deep learning model. In the first training stage, the first loss function indicates the similarity between the predicted answer corresponding to the third question and the expected answer. In the second training stage, the first loss function indicates the similarity between the predicted answer corresponding to the fourth question and the expected answer. The third question and the fourth question are both related to traffic scenes, and the third question and the fourth question are used to obtain information at different levels.
12. The method according to claim 11, characterized in that The third question includes first feature information, where the first feature information includes feature information of environmental information around the vehicle, and the environmental information includes physical property information of objects around the vehicle.
13. The method according to claim 12, characterized in that Before inputting the third question into the deep learning model, the method further includes: Inputting the environmental information into a feature extraction network to obtain second feature information generated by the feature extraction network, wherein the second feature information is used to obtain the first feature information; Inputting the first feature information into a feature processing network to obtain prediction information generated by the feature processing network, wherein the feature extraction network and the feature processing network belong to the same neural network, and the neural network is used to perform at least the following tasks: predicting trajectories of objects around the vehicle, making decisions on the behavior of the vehicle, planning trajectories of the vehicle, predicting a speed range of the vehicle, or controlling the vehicle; Among them, the training stage of the deep learning model includes using the first loss function and the second loss function to train the deep learning model and the neural network, and the second loss function indicates the similarity between the predicted information corresponding to the environmental information and the second expected information.
14. An information processing device, characterized in that: The device comprises: An input module, configured to input first information corresponding to a traffic scene around the vehicle into the deep learning model, wherein the first information includes first text information; An acquisition module is used to acquire second information corresponding to the first information, wherein the second information is obtained based on the deep learning model, and the second information corresponds to any of the following tasks: making decisions on the behavior of the vehicle, planning the trajectory of the vehicle, or controlling the vehicle.
15. The device according to claim 14, characterized in that The first text information includes information of a preset text corresponding to the task, or the first text information includes information of a text input by a user.
16. The device according to claim 14 or 15, characterized in that The input module is also used to input a first question corresponding to the traffic scenario into the deep learning model, the first question is used to obtain a first answer corresponding to the first question, the first answer is obtained by the deep learning model, wherein the first information is the second question, the second information is the second answer corresponding to the second question, and the information in the first answer is present in the first text information.
17. The device according to claim 14 or 15, characterized in that The first information also includes first feature information, and the first feature information includes feature information of environmental information around the vehicle, and the environmental information includes physical attribute information of objects around the vehicle.
18. The device according to claim 17, characterized in that The first feature information is obtained based on a feature extraction network, wherein the feature extraction network belongs to a first neural network, and the first neural network is used to perform at least two of the following tasks: trajectory prediction of objects around the vehicle, decision-making on the behavior of the vehicle, trajectory planning of the vehicle, prediction of the speed range of the vehicle, or control of the vehicle.
19. The device according to claim 14 or 15, characterized in that The first information is a second question, and the second information is a second answer corresponding to the second question, wherein the second answer is included in a preset answer set corresponding to the second question, and the second answer is obtained from at least one alternative answer generated from the deep learning model based on the preset answer set.
20. The device according to claim 14 or 15, characterized in that The first information is the second question, the second information is the second answer corresponding to the second question, the second answer is obtained through a classification network, the second feature information includes feature information of a start flag [CLS] bit in the feature information of the second answer, and the classification network is used to determine the second answer corresponding to the second feature information from at least one preset answer.
21. The device according to claim 14 or 15, characterized in that The first information is a second question, and the second information is a second answer corresponding to the second question; The input module is specifically used to input a plurality of the second questions into the deep learning model, and different second questions in the plurality of the second questions carry different prompts; The acquisition module is specifically used to obtain multiple reference answers corresponding to the multiple second questions one by one, and determine the second answer based on the multiple reference answers, and the multiple reference answers are all obtained through the deep learning model.
22. An information processing device, characterized in that: The device comprises: An input module, configured to input a first question into the deep learning model to obtain a first answer corresponding to the first question; a determination module, configured to determine a second question according to the first answer, wherein the second question carries information in the first answer; The input module is further used to input the second question into the deep learning model to obtain a second answer corresponding to the second question. The first question and the second question correspond to the same traffic scenario.
23. The device according to claim 22, characterized in that The second answer corresponds to any of the following tasks: identifying risk obstacles around the vehicle, identifying the behavior of objects around the vehicle, predicting the behavior of objects around the vehicle, predicting the trajectory of objects around the vehicle, making decisions on the behavior of the vehicle, planning the trajectory of the vehicle, or controlling the vehicle.
24. An information processing device, characterized in that: The device is used to train a deep learning model, the deep learning model includes at least two training stages, the at least two training stages include a first training stage and a second training stage, and the device includes: An input module, configured to input a third question into the deep learning model in the first training phase to obtain a predicted answer corresponding to the third question; The input module is further configured to input a fourth question into the deep learning model during the second training phase to obtain a predicted answer corresponding to the fourth question; In which, a first loss function is used when training the deep learning model. In the first training stage, the first loss function indicates the similarity between the predicted answer corresponding to the third question and the expected answer. In the second training stage, the first loss function indicates the similarity between the predicted answer corresponding to the fourth question and the expected answer. The third question and the fourth question are both related to traffic scenes, and the third question and the fourth question are used to obtain information at different levels.
25. The device according to claim 24, characterized in that The third question includes first feature information, where the first feature information includes feature information of environmental information around the vehicle, and the environmental information includes physical property information of objects around the vehicle.
26. The device according to claim 25, characterized in that The device also includes: A feature extraction module, used for inputting the environmental information into a feature extraction network to obtain second feature information generated by the feature extraction network, wherein the second feature information is used to obtain the first feature information; a feature processing module, configured to input the first feature information into a feature processing network to obtain prediction information generated by the feature processing network, wherein the feature extraction network and the feature processing network belong to the same neural network, and the neural network is configured to perform at least the following tasks: trajectory prediction of objects around the vehicle, decision making on the behavior of the vehicle, trajectory planning of the vehicle, prediction of the speed range of the vehicle, or control of the vehicle; Among them, the training stage of the deep learning model includes using the first loss function and the second loss function to train the deep learning model and the neural network, and the second loss function indicates the similarity between the predicted information corresponding to the environmental information and the second expected information.
27. A device, characterized in that The method comprises a processor, wherein the processor is coupled to a memory, wherein the memory stores program instructions, and when the program instructions stored in the memory are executed by the processor, the method according to any one of claims 1 to 13 is implemented.
28. A vehicle, characterized in that: The method comprises a processor, wherein the processor is coupled to a memory, wherein the memory stores program instructions, and when the program instructions stored in the memory are executed by the processor, the method according to any one of claims 1 to 10 is implemented.
29. A computer-readable storage medium comprising a program, which, when executed on a computer, causes the computer to execute the method according to any one of claims 1 to 13.
30. A circuit system, characterized in that: The circuit system comprises a processing circuit configured to perform the method of any one of claims 1 to 13 .
Citation Information
Patent Citations
Information processing method and related equipment
CN120039268A
Decision planning method and device, and computer readable storage medium
CN108960432A
Automatic driving method, system and device and readable storage medium
CN109358614A
Training method of automatic driving decision model and vehicle control method and device
CN116279589A
Deep learning model training method, target object detection method and device
CN116663650A
Cited By
Motion state planning method and device based on time domain convolutional neural network
CN120963757A
Vehicle cloud collaborative automatic driving model iterative training method and system
CN122332933A