Information processing method and related equipment

By using deep learning models that can process text information in autonomous vehicles, the problem of poor interpretability in model operation process is solved, and more transparent decision-making and behavioral understanding is achieved.

CN120039268APending Publication Date: 2025-05-27HUAWEI TECH CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202311594214.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-24
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

The operation process of deep learning models in autonomous vehicles is poorly interpretable, which makes it difficult for users to understand the decisions and behaviors of the vehicles.

Method used

A deep learning model that can process text information is introduced, and information on bicycle behavior decisions, trajectory planning or control is generated by inputting text information corresponding to traffic scenes around the bicycle.

Benefits of technology

The interpretation of the operation process of the deep learning model is improved, making the decisions and behaviors of autonomous vehicles more transparent, and users can understand the behavior of vehicles more intuitively.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120039268A_ABST
    Figure CN120039268A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses an information processing method and related equipment. The method can be applied to the field of automatic driving in artificial intelligence. The method comprises the following steps: inputting first information corresponding to a traffic scene around a vehicle to a deep learning model, and obtaining second information corresponding to the first information; the first information comprises first text information, the second information is obtained based on a deep learning model, and the second information corresponds to any one of the following tasks: making a decision on the behavior of the vehicle, planning the trajectory of the vehicle, or controlling the vehicle; the input information of the deep learning model provided by the invention has the text information convenient for the user to understand, so that the interpretability of the running process of the deep learning model is improved, and the decision, trajectory planning or control process of the automatic driving vehicle is more transparent; a user can more intuitively understand the behavior of the autonomous vehicle.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence, and in particular to an information processing method and related equipment. Background Art

[0002] Artificial Intelligence (AI) is the theory, method, technology and application system that uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results. In other words, artificial intelligence is a branch of computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a similar way to human intelligence. Artificial intelligence is also the study of the design principles and implementation methods of various intelligent machines, so that machines have the functions of perception, reasoning and decision-making. Research in the field of artificial intelligence includes robotics, natural language processing, computer vision, decision-making and reasoning, human-computer interaction, recommendation and search, basic AI theory, etc.

[0003] The field of autonomous driving is an application field of a scenario in the field of artificial intelligence. For example, it can obtain environmental information around the vehicle, generate prediction information through one or more neural networks, and then make decisions on the vehicle's behavior and plan its trajectory based on the aforementioned prediction information; for example, when there is a car illegally parked in the road ahead, decide whether to change lanes, and plan when and how to change lanes.

[0004] However, since what is obtained is the environmental information around the vehicle, the prediction information obtained through the neural network is directly used to determine the behavior of the vehicle, and the entire operation process of the neural network has poor interpretability. Summary of the invention

[0005] The present application provides an information processing method and related equipment, which improve the interpretability of the operation process of the deep learning model, that is, make the decision-making, trajectory planning or control process of the autonomous driving vehicle more transparent, so that users can understand the behavior of the autonomous driving vehicle more intuitively.

[0006] This application provides the following technical solutions:

[0007] In the first aspect, the present application provides an information processing method that can be used in the field of autonomous driving in the field of artificial intelligence, in which a first device can input first information corresponding to a traffic scene around a self-vehicle into a deep learning model that has performed a training operation, thereby obtaining second information corresponding to the first information. The aforementioned first information includes first text information, the second information is obtained based on the deep learning model that has performed a training operation, and the second information corresponds to any of the following tasks: making decisions on the behavior of the self-vehicle, planning the trajectory of the self-vehicle, or controlling the self-vehicle.

[0008] The first device may be specifically a vehicle or a cloud server. Exemplarily, in one case, the first device is a vehicle. If the above-mentioned deep learning model is deployed in the vehicle, the first device inputs the first information corresponding to the traffic scene around the vehicle to the deep learning model that has performed the training operation, which may include: the first device inputs the first information corresponding to the traffic scene around the vehicle to the locally deployed deep learning model. The first device obtains the second information corresponding to the first information, which may include: the first device determines the second information corresponding to the first information based on the first prediction information generated by the deep learning model.

[0009] In another case, the first device is a vehicle, and if the deep learning model that has performed the training operation is deployed on a cloud server, the first device inputs the first information corresponding to the traffic scene around the vehicle to the deep learning model that has performed the training operation, which may include: the first device sends the first information corresponding to the traffic scene around the vehicle to the cloud server that deploys the deep learning model. The first device obtains the second information corresponding to the first information, which may include: the first device receives the first prediction information sent by the cloud server, and then determines the second information corresponding to the first information according to the first prediction information.

[0010] In another case, the first device is a cloud server that deploys a deep learning model, and the first device inputs the first information corresponding to the traffic scene around the vehicle to the deep learning model that has performed the training operation, which may include: after receiving the first information sent by the vehicle, the first device inputs the first information corresponding to the traffic scene around the vehicle to the locally deployed deep learning model. The first device obtains the second information corresponding to the first information, which may include: the first device sends the first prediction information to the vehicle, and the first prediction information is used for the vehicle to determine the aforementioned second information.

[0011] Among them, the first information includes first text information, and the first text information can correspond to a task performed by the deep learning model (for the convenience of description, hereinafter referred to as the "first task").

[0012] For example, "making decisions on the behavior of the vehicle" can be understood as the second information being used to indicate what behavior the vehicle should perform next; "planning the trajectory of the vehicle" can be understood as the second information may include the next trajectory information of the vehicle; and "controlling the vehicle" can be understood as the second information may include the vehicle control instructions.

[0013] In this implementation, a deep learning model capable of processing text information is introduced into the field of autonomous driving. Based on the second information generated by the deep learning model that processes text information, a decision is made on the behavior of the vehicle, the trajectory of the vehicle is planned, or the vehicle is controlled. Since the input information of the deep learning model provided by this application contains text information that is easy for users to understand, the interpretability of the operation process of the deep learning model is improved, and the decision-making, trajectory planning or control process of the autonomous driving vehicle is made more transparent, so that users can understand the behavior of the autonomous driving vehicle more intuitively.

[0014] In one possible implementation, the first text information includes information of a preset text corresponding to the first task performed by the deep learning model. Exemplarily, a preset text corresponding to each of the at least one second task may be deployed in the vehicle, and the vehicle obtains the preset text corresponding to the first task from the preset text corresponding to the aforementioned at least one second task, thereby obtaining the first text information. Alternatively, the first text information includes information of a text input by a user.

[0015] Exemplarily, the first text information may include initial feature information of a preset text (or text input by a user) obtained after feature extraction of the preset text (or text input by a user), and the aforementioned "feature extraction" may also be understood as "initial encoding", or the aforementioned "feature extraction" may also be understood as vectorization (embedding); exemplarily, the first text information may be specifically expressed in the form of a token.

[0016] In this implementation, two situations in which the first text information may include information are listed, which improves the implementation flexibility of this solution; when the first text information includes information of a preset text corresponding to the first task performed by the deep learning model, then after the first task is determined, it is beneficial to quickly determine the first text information to improve the efficiency of obtaining the second information; when the first text information includes text information input by the user, that is, questions can be answered based on the user's questions, which is beneficial to improving the user stickiness of this solution and making it easier for users to understand the behavior of the vehicle during the autonomous driving process.

[0017] In a possible implementation, in the method, the first device may also input a first question corresponding to a traffic scenario into the deep learning model, the first question is used to obtain a first answer corresponding to the first question, and the first answer is obtained by the deep learning model. The first information is the second question, the second information is the second answer corresponding to the second question, and the first text information contains information in the first answer.

[0018] In this implementation, a first answer to the first question is first obtained, and then a second question is generated based on the first answer, and then the answer to the second question is obtained. The first question and the second question correspond to the same traffic scenario, that is, the questions asked to the deep learning model are asked in a progressive manner, which is conducive to reducing the difficulty of the deep learning model in answering the second question, and is also conducive to improving the ability of the deep learning model in dealing with problems corresponding to complex traffic scenarios; in addition, the use of a progressive questioning method is also conducive to introducing the logical thinking of gradual reasoning into the deep learning model used in the field of autonomous driving, thereby improving the human-likeness of the second answer finally obtained.

[0019] In a possible implementation, the first information further includes first feature information, the first feature information includes feature information of environmental information around the vehicle, and the environmental information includes physical property information of objects around the vehicle. For example, the objects around the vehicle may include dynamic obstacles around the vehicle, static obstacles around the vehicle, traffic markings around the vehicle, traffic signs around the vehicle, or other types of objects.

[0020] Optionally, the first characteristic information also includes characteristic information of the vehicle's driving behavior information and / or characteristic information of navigation information.

[0021] In this implementation, the first information includes not only the first text information, but also the first feature information. The first feature information includes feature information of the environmental information around the vehicle, so that the deep learning model can fully combine the feature information of the environmental information around the vehicle to determine the second information, which is conducive to improving the rationality of the second information output by the deep learning model.

[0022] In one possible implementation, the first feature information is obtained based on a feature extraction network (hereinafter referred to as the "first feature extraction network" for the convenience of description), wherein the first feature extraction network belongs to a first neural network, and the first neural network is used to perform at least two of the following tasks: trajectory prediction of objects around the vehicle, decision-making on the behavior of the vehicle, trajectory planning of the vehicle, prediction of the speed range of the vehicle, or control of the vehicle.

[0023] In this implementation, since the first neural network is used to perform trajectory prediction of objects around the vehicle, make decisions on the vehicle's behavior, plan the vehicle's trajectory, predict the vehicle's speed range, or control at least two of the multiple tasks, the first feature information needs to cover richer information, that is, the first information containing the first feature information will cover more information, which is conducive to the deep learning model obtaining more information, and further conducive to improving the accuracy of the information output by the deep learning model.

[0024] In one possible implementation, the vehicle can input the environmental information around the vehicle into the first feature extraction network. Optionally, the vehicle can also input the driving behavior information and / or navigation information of the vehicle into the first feature extraction network to obtain the third feature information generated by the first feature extraction network, and then obtain the first feature information.

[0025] Exemplarily, in one case, "third feature information" and "first feature information" have the same meaning, that is, the vehicle can directly use the third feature information generated by the first feature extraction network as the first feature information.

[0026] In another case, after obtaining the third feature information, the vehicle can also input the third feature information into the second feature extraction network, and update the features of the third feature information through the second feature extraction network to obtain the first feature information. Exemplarily, the third feature information can be specifically expressed in the form of a feature map, and the first feature information can be specifically expressed in the form of a token, that is, the third feature information in the form of a feature map can be converted into the first feature information in the form of a token through the second feature extraction network. Since the deep learning model is a machine learning model that processes text information, after obtaining the first feature information in the form of a token and then inputting it into the deep learning model, the deep learning model can more easily understand the first feature information, thereby reducing the difficulty of the deep learning model in understanding the first feature information, which is conducive to improving the accuracy of the information output by the deep learning model.

[0027] In one possible implementation, the first information is the second question, and the second information is the second answer corresponding to the second question, wherein the second answer is included in a preset answer set corresponding to the second question, and the second answer is obtained from at least one alternative answer generated from a deep learning model based on the preset answer set. Exemplarily, a preset answer set may be deployed in the execution device, and the preset answer set includes one or more answers. After the execution device obtains the first prediction information, the intersection between at least one alternative answer indicated by the first prediction information and the preset answer set may be determined, and the aforementioned intersection includes at least one first alternative answer; the execution device may determine the second answer from at least one first alternative answer based on the first probability value corresponding to each first alternative answer in the at least one first alternative answer, and the first probability value corresponding to the second answer is the largest among the at least one first alternative answer.

[0028] In this implementation, since the deep learning model may sometimes give irrelevant answers or speak nonsense, that is, at least one alternative answer determined by the deep learning model may contain an answer that is completely irrelevant to the question, in order to avoid the occurrence of the aforementioned problem, a second preset answer set corresponding to the second question can be deployed, and at least one answer included in the second preset answer set corresponding to the second question is an answer related to the second question. The second answer corresponding to the second question is finally determined based on the aforementioned second preset answer set and at least one alternative answer determined based on the deep learning model. This is conducive to avoiding the problem that the final second answer is irrelevant to the second question, thereby improving the accuracy of the second answer obtained by the deep learning model.

[0029] In one possible implementation, the first information is the second question, the second information is the second answer corresponding to the second question, the second answer is obtained through a classification network, the second feature information includes feature information of a start flag [CLS] bit in the feature information of the second answer, and the classification network is used to determine the second answer corresponding to the second feature information from at least one preset answer. Exemplarily, a trained classification network may be deployed in the execution device, and after obtaining the first prediction information corresponding to the second question generated by the deep learning model, the execution device may determine the feature information of the second candidate answer from the first prediction information, and the first probability value of the second candidate answer is the highest among the at least one candidate answer indicated by the first prediction information; the execution device may obtain the second feature information from the feature information of the second candidate answer, input the second feature information into the classification network, and obtain the second answer generated by the classification network.

[0030] In this implementation, since the characteristic information of the starting mark [CLS] bit (that is, the second characteristic information) can represent the overall meaning of the alternative answers, that is, it is reasonable to use the characteristic information of the starting mark [CLS] bit to replace the characteristic information of the alternative answers. The classification network is used to determine a second answer belonging to the second characteristic information from at least one preset answer, and the second answer is limited to at least one preset answer, which is conducive to avoiding the problem that the second answer finally obtained is irrelevant to the second question, thereby improving the accuracy of the second answer obtained by the deep learning model.

[0031] In one possible implementation, the first information is the second question, and the second information is the second answer corresponding to the second question. The first device inputs the first information corresponding to the traffic scene around the vehicle into the deep learning model, which may include: the first device inputs multiple second questions into the deep learning model; illustratively, the multiple second questions correspond to the same task, that is, the semantics of the multiple second questions are similar, and different second questions in the multiple second questions carry different prompts. The first device obtains the second information corresponding to the first information, which may include: the first device obtains multiple reference answers corresponding to the multiple second questions one by one, and then determines the second answer based on the multiple reference answers, wherein the multiple reference answers are all obtained through the deep learning model.

[0032] In this implementation, multiple second questions are input into the deep learning model, and different second questions carry different prompts. Multiple reference answers are obtained by utilizing multiple different prompts, and then the final second answer can be obtained from the multiple reference answers, thereby improving the rigor of obtaining the final second answer and facilitating improving the accuracy of the second questions finally obtained.

[0033] In one possible implementation, the first device determines the second answer based on multiple reference answers, which may include: the first device determines the second answer based on the multiple reference answers using a majority voting strategy, that is, the ultimately determined second answer is included in the multiple reference answers.

[0034] In a second aspect, the present application provides an information processing method that can be used in the field of autonomous driving in the field of artificial intelligence, in which a first device inputs a first question into a deep learning model to obtain a first answer corresponding to the first question; a second question is determined based on the first answer, and the second question carries information in the first answer. The first device inputs a second question into the deep learning model to obtain a second answer corresponding to the second question, wherein the first question and the second question correspond to the same traffic scenario.

[0035] Exemplarily, the first question may include a prompt, and optionally, the first question may also include first characteristic information, the first characteristic information including characteristic information of environmental information around the vehicle, the environmental information including physical property information of objects around the vehicle. Optionally, the first characteristic information also includes characteristic information of driving behavior information of the vehicle and / or characteristic information of navigation information.

[0036] Optionally, the first question and the second question can be used to obtain different levels of information from the target traffic scene. For example, the first question is used to obtain basic attribute information and semantic information in the traffic scene; for another example, the first question is used to obtain the behavior of objects in the target traffic scene or the relationship between different objects.

[0037] In one possible implementation, the second answer corresponds to any of the following tasks: identifying risk obstacles around the vehicle, identifying the behavior of objects around the vehicle, predicting the behavior of objects around the vehicle, predicting the trajectory of objects around the vehicle, making decisions on the behavior of the vehicle, planning the trajectory of the vehicle, or controlling the vehicle. In the embodiment of the present application, a variety of tasks performed by a deep learning model are provided, which improves the implementation flexibility of the solution.

[0038] In the second aspect, the first device is also used to execute the steps performed by the first device in the first aspect and various possible implementation methods of the first aspect. For the meanings of the nouns in the second aspect of this application and various possible implementation methods of the second aspect, the specific implementation methods of the steps, and the beneficial effects brought about by each possible implementation method, please refer to the description in the various possible implementation methods of the first aspect, and will not be repeated here one by one.

[0039] In the third aspect, the present application provides an information processing method that can be used in the field of autonomous driving in the field of artificial intelligence. The method is used to train a deep learning model, and the deep learning model includes at least two training stages, and the at least two training stages include a first training stage and a second training stage. In this method, in the first training stage, the training device inputs a third question to the deep learning model to obtain a predicted answer corresponding to the third question; in the second training stage, the training device inputs a fourth question to the deep learning model to obtain a predicted answer corresponding to the fourth question. Among them, a first loss function is used when training the deep learning model. In the first training stage, the first loss function indicates the similarity between the predicted answer corresponding to the third question and the expected answer. In the second training stage, the first loss function indicates the similarity between the predicted answer corresponding to the fourth question and the expected answer. The third question and the fourth question are both related to traffic scenes, and the third question and the fourth question are used to obtain information at different levels.

[0040] In this implementation, at least two training stages are used to train the deep learning model. The at least two training stages respectively use the third problem and the fourth problem. The third problem and the fourth problem are used to obtain information of different levels from traffic scenes. That is, the training process of the deep learning model adopts a step-by-step training method, which is conducive to enabling the deep learning model to learn logical thinking of step-by-step reasoning, thereby improving the human-likeness of the deep learning model, and also helping the trained deep learning model to be more reasonable when performing tasks related to autonomous driving.

[0041] In a possible implementation, the third question includes first feature information, the first feature information includes feature information of environmental information around the vehicle, and the environmental information includes physical property information of objects around the vehicle.

[0042] In one possible implementation, before the training device inputs the third question into the deep learning model, the method also includes: the training device inputs the environmental information into the feature extraction network to obtain second feature information generated by the feature extraction network, and the second feature information is used to obtain the first feature information; the first feature information is input into the feature processing network to obtain prediction information generated by the feature processing network, the feature extraction network and the feature processing network belong to the same neural network, and the neural network is used to perform at least the following tasks: trajectory prediction of objects around the vehicle, decision-making on the behavior of the vehicle, trajectory planning for the vehicle, prediction of the speed range of the vehicle, or control of the vehicle; wherein the training stage of the deep learning model includes training the deep learning model and the neural network using a first loss function and a second loss function, and the second loss function indicates the similarity between the prediction information corresponding to the environmental information and the second expected information.

[0043] In the third aspect, the training device is also used to execute the steps performed by the first device in the first aspect and various possible implementation methods of the first aspect. For the meanings of the nouns in the third aspect of this application and various possible implementation methods of the third aspect, the specific implementation methods of the steps, and the beneficial effects brought about by each possible implementation method, please refer to the description of the various possible implementation methods in the first aspect, and will not be repeated here one by one.

[0044] In a fourth aspect, the present application provides an information processing device that can be used in the field of autonomous driving in the field of artificial intelligence. The information processing device includes: an input module for inputting first information corresponding to the traffic scene around the vehicle into a deep learning model, where the first information includes first text information; an acquisition module for acquiring second information corresponding to the first information, where the second information is obtained based on the deep learning model, and the second information corresponds to any of the following tasks: making decisions on the behavior of the vehicle, planning the trajectory of the vehicle, or controlling the vehicle.

[0045] In the fourth aspect, the information processing device is also used to execute the steps performed by the first device in the first aspect and various possible implementation methods of the first aspect. For the meaning of the nouns in the fourth aspect of this application and various possible implementation methods of the fourth aspect, the specific implementation methods of the steps, and the beneficial effects brought about by each possible implementation method, please refer to the description of the various possible implementation methods in the first aspect, and will not be repeated here one by one.

[0046] In a fifth aspect, the present application provides an information processing device that can be used in the field of autonomous driving in the field of artificial intelligence. The information processing device includes: an input module, used to input a first question into a deep learning model to obtain a first answer corresponding to the first question; a determination module, used to determine a second question based on the first answer, the second question carries information in the first answer; the input module is also used to input a second question into the deep learning model to obtain a second answer corresponding to the second question, wherein the first question and the second question correspond to the same traffic scenario.

[0047] In the fifth aspect, the information processing device is also used to execute the steps performed by the first device in the first aspect and various possible implementation methods of the first aspect. For the meaning of the nouns in the fifth aspect of this application and various possible implementation methods of the fifth aspect, the specific implementation methods of the steps, and the beneficial effects brought about by each possible implementation method, please refer to the description of the various possible implementation methods in the first aspect, and will not be repeated here one by one.

[0048] In a sixth aspect, the present application provides an information processing device that can be used in the field of autonomous driving in the field of artificial intelligence. The device is used to train a deep learning model. The deep learning model includes at least two training stages, and the at least two training stages include a first training stage and a second training stage. The device includes: an input module, which is used to input a third question into the deep learning model in the first training stage to obtain a predicted answer corresponding to the third question; the input module is also used to input a sixth question into the deep learning model in the second training stage to obtain a predicted answer corresponding to the sixth question; wherein, a first loss function is used when training the deep learning model. In the first training stage, the first loss function indicates the similarity between the predicted answer corresponding to the third question and the expected answer. In the second training stage, the first loss function indicates the similarity between the predicted answer corresponding to the sixth question and the expected answer. The third question and the sixth question are both related to traffic scenes, and the third question and the sixth question are used to obtain information at different levels.

[0049] In the sixth aspect, the information processing device is also used to execute the steps performed by the training device in the third aspect and various possible implementation methods of the third aspect. For the meaning of the nouns in the sixth aspect of this application and various possible implementation methods of the sixth aspect, the specific implementation methods of the steps, and the beneficial effects brought about by each possible implementation method, please refer to the description of the various possible implementation methods in the third aspect, and will not be repeated here one by one.

[0050] In the seventh aspect, an embodiment of the present application provides a device, including a processor and a memory, where the processor is coupled to the memory, the memory is used to store programs; the processor is used to execute the programs in the memory, so that the device executes the method described in the first aspect, the second aspect or the third aspect above.

[0051] In an eighth aspect, an embodiment of the present application provides a vehicle, including a processor and a memory, wherein the processor is coupled to the memory, the memory is used to store programs; the processor is used to execute the programs in the memory, so that the vehicle executes the method described in the first or second aspect above.

[0052] In a ninth aspect, the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, which, when executed on a computer, enables the computer to execute the method described in the first, second or third aspect above.

[0053] In a tenth aspect, the present application provides a circuit system, which includes a processing circuit, and the processing circuit is configured to execute the method described in the first aspect, the second aspect or the third aspect above.

[0054] In an eleventh aspect, the present application provides a computer program product, which includes a computer program, and when the computer program is run on a computer, it enables the computer to execute the method described in the first aspect, the second aspect or the third aspect above.

[0055] In a twelfth aspect, the present application provides a chip system, which includes a processor for supporting the implementation of the functions involved in the above aspects, for example, sending or processing the data and / or information involved in the above methods. In one possible design, the chip system also includes a memory, which is used to store program instructions and data necessary for a server or communication device. The chip system can be composed of a chip, or it can include a chip and other discrete devices. BRIEF DESCRIPTION OF THE DRAWINGS

[0056] Figure 1 A structural diagram of an artificial intelligence main framework provided for an embodiment of the present application;

[0057] Figure 2 A system architecture diagram of a data processing system provided in an embodiment of the present application;

[0058] Figure 3 A schematic diagram of a flow chart of an information processing method provided in an embodiment of the present application;

[0059] Figure 4 A schematic diagram of a deep learning model provided in an embodiment of the present application;

[0060] Figure 5a A schematic diagram of first information provided in an embodiment of the present application;

[0061] Figure 5b Another schematic diagram of the first information provided in the embodiment of the present application;

[0062] Figure 6 A schematic diagram of a vehicle obtaining first information provided by an embodiment of the present application;

[0063] Figure 7 Another flowchart of the information processing method provided in the embodiment of the present application;

[0064] Figure 8 A schematic diagram of alternative answers provided for an embodiment of the present application;

[0065] Fig. 9 A schematic diagram of a traffic scenario provided in an embodiment of the present application;

[0066] Fig.10 Another flowchart of the information processing method provided in the embodiment of the present application;

[0067] Fig.11 A schematic diagram of multiple second problems provided in an embodiment of the present application;

[0068] Fig.12 A schematic diagram of the structure of an information processing device provided in an embodiment of the present application;

[0069] Fig.13 Another structural schematic diagram of an information processing device provided in an embodiment of the present application;

[0070] Fig.14 Another structural schematic diagram of an information processing device provided in an embodiment of the present application;

[0071] Fig.15 A schematic diagram of the structure of a device provided in an embodiment of the present application;

[0072] Fig.16 A schematic diagram of the structure of a vehicle provided in an embodiment of the present application;

[0073] Fig.17 A schematic diagram of the structure of a chip provided in an embodiment of the present application. DETAILED DESCRIPTION

[0074] The embodiments of the present application are described below in conjunction with the accompanying drawings. Those skilled in the art will appreciate that, with the development of technology and the emergence of new scenarios, the technical solutions provided in the embodiments of the present application are also applicable to similar technical problems.

[0075] The terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and need not be used to describe a specific order or sequential order. It should be understood that the terms used in this way can be interchangeable under appropriate circumstances, which is only to describe the distinction mode adopted by the objects of the same attributes when describing in the embodiments of the present application. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, so that the process, method, system, product or equipment comprising a series of units need not be limited to those units, but may include other units that are not clearly listed or inherent to these processes, methods, products or equipment.

[0076] "Send" and "receive" in the embodiments of the present application indicate the direction of signal transmission. For example, "send information to XX device" can be understood as the destination of the information is XX device, which can include direct transmission through the air interface, and also include indirect transmission through the air interface by other units or modules. "Receive information from YY device" can be understood as the source of the information is YY device, which can include direct reception from YY device through the air interface, and can also include indirect reception from YY device through other units or modules through the air interface. "Send" can also be understood as the "output" of the chip interface, and "receive" can also be understood as the "input" of the chip interface. In other words, sending and receiving can be carried out between devices or within the device, for example, sending or receiving between components, modules, chips, software modules or hardware modules within the device through a bus, a wiring or an interface. It is understandable that the information may be processed as necessary between the source and the destination of the information transmission, such as encoding, modulation, etc., but the destination can understand the valid information from the source. Similar expressions in this application can be understood in a similar way and will not be repeated.

[0077] In the embodiments of the present application, "indication" may include direct indication and indirect indication, and may also include explicit indication and implicit indication. The information indicated by a certain information (such as the indication information described below) is called information to be indicated. In the specific implementation process, there are many ways to indicate the information to be indicated, such as but not limited to, directly indicating the information to be indicated, such as the information to be indicated itself or the index of the information to be indicated. The information to be indicated may also be indirectly indicated by indicating other information, wherein there is an association between the other information and the information to be indicated; it may also be possible to indicate only a part of the information to be indicated, while the other part of the information to be indicated is known or agreed in advance, for example, the indication of specific information can be realized by means of the arrangement order of each information agreed in advance (such as predefined by the protocol), thereby reducing the indication overhead to a certain extent. The present application does not limit the specific method of indication. It is understandable that, for the sender of the indication information, the indication information can be used to indicate the information to be indicated, and for the receiver of the indication information, the indication information can be used to determine the information to be indicated.

[0078] First, the overall workflow of the artificial intelligence system is described. Figure 1 , Figure 1 The figure shows a structural diagram of the main framework of artificial intelligence. The following is an explanation of the above artificial intelligence theme framework from the two dimensions of "intelligent information chain" (horizontal axis) and "IT value chain" (vertical axis). Among them, the "intelligent information chain" reflects a series of processes from data acquisition to processing. For example, it can be a general process of intelligent information perception, intelligent information representation and formation, intelligent reasoning, intelligent decision-making, intelligent execution and output. In this process, the data has undergone a condensation process of "data-information-knowledge-wisdom". The "IT value chain" reflects the value that artificial intelligence brings to the information technology industry from the underlying infrastructure of human intelligence, information (providing and processing technology implementation) to the industrial ecology process of the system.

[0079] (1) Infrastructure

[0080] The infrastructure provides computing power support for the artificial intelligence system, enables communication with the outside world, and is supported by the basic platform. It communicates with the outside world through sensors; computing power is provided by smart chips, which can specifically use hardware acceleration chips such as central processing units (CPU), embedded neural network processing units (NPU), graphics processing units (GPU), application specific integrated circuits (ASIC) or field programmable gate arrays (FPGA); the basic platform includes distributed computing frameworks and networks and other related platform guarantees and support, which can include cloud storage and computing, interconnected networks, etc. For example, sensors communicate with the outside world to obtain data, and these data are provided to the smart chips in the distributed computing system provided by the basic platform for calculation.

[0081] (2) Data

[0082] The data on the upper layer of the infrastructure is used to represent the data sources in the field of artificial intelligence. The data involves graphics, images, voice, text, and IoT data of traditional devices, including business data of existing systems and perception data such as force, displacement, liquid level, temperature, and humidity.

[0083] (3) Data processing

[0084] Data processing usually includes data training, machine learning, deep learning, search, reasoning, decision-making and other methods.

[0085] Among them, machine learning and deep learning can symbolize and formalize data for intelligent information modeling, extraction, preprocessing, and training.

[0086] Reasoning refers to the process of simulating human intelligent reasoning in computers or intelligent systems, using formalized information to perform machine thinking and solve problems based on reasoning control strategies. Typical functions are search and matching.

[0087] Decision-making refers to the process of making decisions after intelligent information is reasoned, usually providing functions such as classification, sorting, and prediction.

[0088] (4) General ability

[0089] After the data has undergone the data processing mentioned above, some general capabilities can be further formed based on the results of the data processing, such as an algorithm or a general system, for example, translation, text analysis, computer vision processing, speech recognition, image recognition, etc.

[0090] (5) Smart products and industry applications

[0091] Smart products and industry applications refer to the products and applications of artificial intelligence systems in various fields. They are the encapsulation of the overall artificial intelligence solution, which productizes intelligent information decision-making and realizes practical applications. Its application areas mainly include: smart terminals, smart manufacturing, smart transportation, smart homes, smart medical care, smart security, autonomous driving, smart cities, etc.

[0092] The method provided in this application can be applied in the field of autonomous driving. For example, in the method provided in this application, a deep learning model capable of processing text information is applied in the field of autonomous driving. Figure 2 , Figure 2 A system architecture diagram of a data processing system provided in an embodiment of the present application, in Figure 2 In the embodiment, the data processing system 200 includes a training device 210 , a database 220 , an execution device 230 , a data storage system 240 and a client device 250 , and the execution device 230 includes a computing module 231 .

[0093] Among them, the database 220 stores a training data set. In the training stage of the deep learning model 201, the training device 210 generates the deep learning model 201, and uses the training data set to iteratively train the deep learning model 201 to obtain the deep learning model 201 that has performed the training operation. The deep learning model 201 can be specifically expressed as a neural network or a non-neural network model. In the embodiment of the present application, only the deep learning model 201 expressed as a neural network is used as an example for explanation.

[0094] The deep learning model 201 that has been trained and obtained by the training device 210 can be deployed in the computing module 231 of the execution device 230. The execution device 230 can call the data, code, etc. in the data storage system 240, or store the data, instructions, etc. in the data storage system 240. The data storage system 240 can be placed in the execution device 230, or the data storage system 240 can be an external memory relative to the execution device 230.

[0095] During the application stage of the deep learning model 201 of the U-shaped over-training operation, after the execution device 230 inputs the first information into the deep learning model 201 in the computing module 231, the second information output by the deep learning model 201 can be obtained, wherein the first information includes the first text information, and the specific information carried by the second information is related to the type of task executed by the deep learning model 201 (hereinafter referred to as the "first task" for the convenience of description).

[0096] In some embodiments of this application, please refer to Figure 2 The execution device 230 and the client device 250 can be independent devices. The execution device 230 is configured with an input / output (I / O) interface to exchange data with the client device 250. After determining the first information, the client device sends the first information to the execution device 230 through the I / O interface. After the execution device 230 generates second information corresponding to the first information through the deep learning model 201 in the computing module 231, the second information can be returned to the client device through the I / O interface.

[0097] It is worth noting that Figure 2 It is only an architectural diagram of the data processing system provided by an embodiment of the present invention, and the positional relationship between the devices, components, modules, etc. shown in the figure does not constitute any limitation. For example, in other embodiments of the present application, the execution device 230 and the client device can be integrated in the same device, and the user can interact directly with the execution device 230. Exemplarily, the execution device 230 can be a module in the host processor (Host CPU) of the client device that uses a deep learning model to process data. The execution device 230 can also be a graphics processing unit (GPU) or a neural network processor (NPU) in the client device. The GPU or NPU is mounted on the host processor as a coprocessor, and the tasks are assigned by the host processor.

[0098] In combination with the above description, this application provides an information processing method. For details, please refer to Figure 3 , Figure 3 A flow chart of an information processing method provided in an embodiment of the present application.

[0099] 301. Input first information corresponding to a traffic scene around the vehicle into a deep learning model, where the first information includes first text information.

[0100] In the embodiment of the present application, before the first device inputs the first information corresponding to the traffic scene around the vehicle into the deep learning model, the first information needs to be obtained by the vehicle first. In one case, the first device is a vehicle, and if the deep learning model that has performed the training operation is deployed in the vehicle (that is, the vehicle is also the execution device of the deep learning model that has performed the training operation), then step 301 includes: the vehicle inputs the first information corresponding to the traffic scene around the vehicle to the locally deployed deep learning model.

[0101] In another case, the first device is a vehicle. If the deep learning model that has performed the training operation is deployed on a cloud server (that is, the device that performs the deep learning model that has performed the training operation is a cloud server), step 301 may include: the first device sends first information corresponding to the traffic scene around the vehicle to the cloud server where the deep learning model is deployed.

[0102] In another case, the first device is a cloud server that deploys a deep learning model, and step 301 may include: after receiving the first information sent by the vehicle corresponding to the traffic scene around the vehicle, the first device inputs the first information into the deep learning model.

[0103] Exemplarily, the above-mentioned deep learning model can specifically adopt a machine learning model based on an attention mechanism, or the above-mentioned deep learning model can also adopt a recurrent neural network, a convolutional neural network, a fully connected neural network or other types of machine learning models, etc., which are not limited here.

[0104] Exemplarily, when the deep learning model specifically adopts a machine learning model based on an attention mechanism, the deep learning model may include N first neural network modules, where N is an integer greater than or equal to 1, and each first neural network module is a neural network module based on an attention mechanism. The deep learning model may also include other neural network layers, for example, other neural network layers may include linear neural network layers, neural network layers for normalization, or other neural network layers, etc. It should be noted that the specific composition form of the deep learning model can be flexibly determined in combination with the actual application scenario. The examples given here are only for the convenience of understanding this solution and are not used to limit this solution.

[0105] For a more intuitive understanding of this solution, please refer to Figure 4 , Figure 4 A schematic diagram of a deep learning model provided in an embodiment of the present application, such as Figure 4 As shown, the deep learning model may include N first neural network modules, and the deep learning model may also include a normalized exponential activation function (Softmax) layer, a linear fully connected (Linear) layer, and other neural network layers.

[0106] Among them, each first neural network module is a neural network module based on the attention mechanism, such as Figure 4 As shown, each first neural network module may include 3 different Linear layers, a neural network layer based on an attention mechanism (Attention), a residual link and normalization (Add&Norm) layer, a feedforward neural network layer and a Softmax layer. Figure 4 The order in which information is processed in each first neural network module is also shown. It should be understood that Figure 4 The examples are only for facilitating the understanding of this solution and are not intended to limit this solution.

[0107] The first information includes first text information, which may correspond to a task performed by a deep learning model (hereinafter referred to as the "first task" for the convenience of description), and the first text information may be used to prompt the deep learning model to generate prediction information corresponding to the first task. Exemplarily, the first information may be specifically expressed as a second question, and the first text information may be specifically expressed as the prompt information in the aforementioned second question; it should be understood that the term "first question" will appear in subsequent steps and will not be introduced here.

[0108] Optionally, the first information further includes first feature information, the first feature information including feature information of environmental information around the vehicle, the environmental information around the vehicle including physical property information of objects around the vehicle. Exemplarily, the environmental information around the vehicle may include physical property information of objects around the vehicle in each frame of one or more frames of images (or point cloud data).

[0109] Exemplarily, the objects around the vehicle may include dynamic obstacles around the vehicle, static obstacles around the vehicle, traffic markings around the vehicle, traffic signs around the vehicle, or other types of objects; for example, the aforementioned dynamic obstacles may be other vehicles, pedestrians, electric vehicles, or other dynamic obstacles around the vehicle, the aforementioned static obstacles may be houses, fences, grass, or other static obstacles around the vehicle, the aforementioned traffic markings may be stop lines, crosswalks, or other traffic sign lines, the aforementioned traffic signs may be traffic lights, speed limit signs, or other traffic signs, etc., all of which may be determined in combination with the actual application scenario and are not limited here.

[0110] For example, the physical property information of dynamic obstacles may include the category, position, speed, direction, height, color, shape or other types of physical property information of the dynamic obstacles; the physical property information of static obstacles may include the category, position, direction, height, material, shape, color or other physical property information of the static obstacles; the physical property information of traffic markings may include the category, position, shape, color or other physical property information of traffic markings; the physical properties of traffic signs may include the category, position, shape, color, content of traffic signs or other physical property information of traffic signs. It should be noted that the examples of various physical property information here are only for the convenience of understanding of this scheme and are not used to limit this scheme.

[0111] Optionally, the first characteristic information also includes characteristic information of the vehicle's driving behavior information and / or characteristic information of navigation information; illustratively, the vehicle's driving behavior information may indicate information about the vehicle's driving behavior in each frame of one or more frames of images (or point cloud data), and the navigation information also includes navigation information in each frame of one or more frames of images (or point cloud data). Exemplarily, the vehicle's driving behavior information may include the vehicle's speed, acceleration, orientation, position, or other information, and the navigation information may include the position of the navigation line, the direction indicated by the navigation line, the shape of the navigation line, or other information, which may be determined in combination with actual conditions and are not limited here.

[0112] The process of the vehicle acquiring the first information may include: the vehicle acquiring the first feature information. With respect to the specific implementation process of the vehicle acquiring the first feature information, illustratively, the first feature information may be obtained based on a feature extraction network (hereinafter referred to as the "first feature extraction network" for the convenience of description). For example, the vehicle may input the environmental information around the vehicle into the first feature extraction network, and optionally, may also input the driving behavior information and / or navigation information of the vehicle into the first feature extraction network to obtain the third feature information generated by the first feature extraction network, and then obtain the first feature information. "Third feature information" may also be understood as the hidden feature information of the environmental information around the vehicle (optionally, it may also include the driving behavior information and / or navigation information of the vehicle).

[0113] Furthermore, in one case, "third feature information" and "first feature information" have the same meaning, that is, the vehicle can directly use the third feature information generated by the first feature extraction network as the first feature information.

[0114] In another case, after obtaining the third feature information, the vehicle can also input the third feature information into the second feature extraction network, and update the features of the third feature information through the second feature extraction network to obtain the first feature information. Exemplarily, the third feature information can be specifically expressed in the form of a feature map, and the first feature information can be specifically expressed in the form of a token, that is, the third feature information in the form of a feature map can be converted into the first feature information in the form of a token through the second feature extraction network. Since the deep learning model is a machine learning model that processes text information, after obtaining the first feature information in the form of a token and then inputting it into the deep learning model, the deep learning model can more easily understand the first feature information, thereby reducing the difficulty of the deep learning model in understanding the first feature information, which is conducive to improving the accuracy of the information output by the deep learning model.

[0115] Exemplarily, the second feature extraction network can be specifically expressed as a neural network based on an attention mechanism, a convolutional neural network, a fully connected neural network or other types of neural networks, etc. The specific form of expression of the second feature extraction network can be determined according to the actual application scenario and is not limited in the embodiments of the present application.

[0116] Among them, the first feature extraction network belongs to the first neural network, and the first neural network can also include a first feature processing network. Exemplarily, the first neural network can be specifically manifested as a neural network based on an attention mechanism, a convolutional neural network, a fully connected neural network, or other types of neural networks, etc., which can be determined in combination with actual application scenarios and are not limited here. Exemplarily, the first feature extraction network can also be called an encoder, and the first feature processing network can also be called a decoder.

[0117] Exemplarily, the first neural network is used to perform at least one of the following tasks: predicting the trajectory of objects around the vehicle, making decisions on the vehicle's behavior, planning the trajectory of the vehicle, predicting the speed range of the vehicle, controlling the vehicle, or other tasks; optionally, the first neural network is used to perform at least two of the aforementioned multiple tasks.

[0118] In an embodiment of the present application, the first information not only includes the first text information, but also includes the first feature information. The first feature information includes the feature information of the environmental information around the vehicle, so that the deep learning model can fully combine the feature information of the environmental information around the vehicle to determine the second information, which is conducive to improving the rationality of the second information output by the deep learning model.

[0119] In addition, if the first neural network is used to perform at least two of the following tasks: predicting the trajectory of objects around the vehicle, making decisions on the vehicle's behavior, planning the trajectory of the vehicle, predicting the speed range of the vehicle, or controlling the vehicle, the first feature information needs to cover richer information, that is, the first information containing the first feature information will cover more information, which will help the deep learning model obtain more information, and thus help improve the accuracy of the information output by the deep learning model.

[0120] The process of the vehicle acquiring the first information includes: the vehicle acquiring the first text information.

[0121] Regarding the specific implementation process of a vehicle acquiring the first text information, in one case, the first text information includes information about a preset text of a first task performed by a deep learning model.

[0122] Exemplarily, a preset text corresponding to each second task in at least one second task may be deployed in the vehicle; optionally, in the case where the first information is specifically a second question, each preset text may also be understood as a preset prompt. The aforementioned at least one second task may include any one or more of the following tasks: making decisions on the behavior of the vehicle, planning the trajectory of the vehicle, controlling the vehicle, or other tasks, etc., which are not exhaustive here.

[0123] In order to further understand the relationship between the words "decision-making", "trajectory planning" and "control", the aforementioned nouns are further explained here. For example, "making decisions on the behavior of the vehicle" refers to making decisions at the vehicle behavior level; for example, "making decisions on the behavior of the vehicle" can be specifically manifested as turning left at the intersection ahead. For another example, "making decisions on the behavior of the vehicle" can be specifically manifested as overtaking the vehicle ahead or other behaviors, etc., and the examples are not exhaustive here.

[0124] "Trajectory planning for the vehicle" can also be called "path planning for the vehicle". "Trajectory planning for the vehicle" refers to determining the path of the behavior that implements the decision based on the decision on the vehicle's behavior (it can also be understood as "determining the trajectory to implement the behavior"). For example, after the decision on the vehicle's behavior is to turn left at the intersection ahead, trajectory planning for the vehicle can be to determine what path the vehicle should take to implement the behavior of "turning left at the intersection ahead"; for another example, the decision on the vehicle's behavior is to overtake the vehicle ahead, trajectory planning for the vehicle can be to determine what path the vehicle should take to implement the behavior of "overtaking the vehicle ahead", etc.

[0125] "Controlling the vehicle" refers to control instructions for vehicle components to achieve the planned trajectory of the vehicle. The aforementioned "control instructions for vehicle components" can also be called "vehicle control instructions". For example, "controlling the vehicle" can be specifically manifested as controlling the angle of the steering wheel, controlling the amount of accelerator pedal, controlling the amount of brake pedal, or control instructions for other components in the vehicle, etc., which are not listed here exhaustively.

[0126] For example, the preset text corresponding to the task of "making decisions about the vehicle" may be: "What interactive behavior will the vehicle perform next and what is the reason for performing the aforementioned behavior", "What behavior does the vehicle need to perform for what reason", "What behavior is the vehicle going to perform soon" or other preset text corresponding to the task of "making decisions about the vehicle", etc.

[0127] The preset text corresponding to the task of "planning a path for the vehicle" may be: "What is the path planned for the vehicle?", "How should the vehicle go next?", "What is the next trajectory of the vehicle?" or other preset text corresponding to the task of "planning a path for the vehicle".

[0128] The preset text corresponding to the task of "controlling the vehicle" can be: "Please confirm the vehicle control instructions", "How to control the vehicle's movement", "How to control the various components of the vehicle", "Can you give a set of vehicle control instructions?" or other preset texts corresponding to the task of "controlling the vehicle", etc. It should be noted that the examples of various preset texts here are only for the convenience of understanding this solution and are not used to limit this solution.

[0129] When the vehicle needs to perform a first task in the at least one second task, the vehicle will be triggered to automatically obtain the first information, that is, the vehicle will be triggered to automatically obtain the first text information. The vehicle obtaining the first text information may include: the vehicle obtains one or more preset texts corresponding to the first task from the preset texts corresponding to the at least one second task, and then obtains one or more first text information; each first text information includes information of a preset text.

[0130] Exemplarily, in this scenario, the first text information may include initial feature information of the preset text obtained after feature extraction of the preset text, and the aforementioned "feature extraction" may also be understood as "initial encoding", or the aforementioned "feature extraction" may also be understood as vectorization (embedding); exemplarily, the initial feature information of the preset text may be specifically expressed in the form of a token.

[0131] In another case, the first text information may also include information of text input by the user (hereinafter referred to as "first text" for the convenience of distinction), that is, the first text input by the user may trigger the vehicle to obtain the first information. Exemplarily, the vehicle may receive any text input by the user (hereinafter referred to as "second text" for the convenience of distinction), and when it is determined that the second text input by the user corresponds to the at least one second task, the second text may be determined to be the first text, and the vehicle may be triggered to obtain the first information, and then step 301 may be executed.

[0132] Exemplarily, in this scenario, the first text information may be the initial feature information of the first text obtained after feature extraction of the first text input by the user, and the aforementioned "feature extraction" may also be understood as "initial encoding"; exemplarily, the initial feature information of the first text may be specifically expressed in the form of a token.

[0133] Regarding the specific implementation method of the vehicle obtaining the first text input by the user, exemplarily, in one implementation method, the vehicle can provide the user with a receiving icon corresponding to a first function (which can also be replaced by a "button"), and the first function is used for the user to understand the decision-making situation, trajectory planning situation, or vehicle control situation during the automatic driving process. Then, when the user clicks the aforementioned receiving icon (or "button") corresponding to the first function, the vehicle can receive the first text input by the user, and determine that the first text input by the user corresponds to the above-mentioned at least one second task, thereby triggering the vehicle to obtain the first information.

[0134] In another implementation, a first machine learning model that has performed training operations may be pre-deployed in the vehicle, and the first machine learning model is used to determine whether the text input by the user corresponds to one or more tasks in at least one second task; illustratively, the first machine learning model may be specifically manifested as a machine learning model for performing classification functions in N categories, and the aforementioned N categories may include each second task and irrelevant ones. The vehicle can receive any text (i.e., the second text) input by the user, input the aforementioned second text into the first machine learning model, and obtain prediction information output by the first machine learning model, which indicates which category of the N categories the second text input by the user belongs to; if it is determined that the second text belongs to any one of the at least one second task according to the prediction information output by the first machine learning model, then the aforementioned second text is determined to be the first text, thereby triggering the vehicle to obtain the first information.

[0135] It should be noted that other methods may also be used to trigger the vehicle to obtain the first information based on the first text input by the user. The example here is only to prove the feasibility of this solution and is not used to limit this solution.

[0136] After obtaining the first text information and the first feature information, the vehicle can combine the first text information and the first feature information to obtain combined information; illustratively, the aforementioned "combination" can be done by splicing, in which case there is an obvious separation between the first text information and the first feature information; alternatively, the first text information in token form and the first feature information in token form can be combined together in a mixed manner, that is, there is no obvious separation between the first text information and the first feature information; alternatively, the first text information and the first feature information can be combined in other ways, which are not limited in the embodiments of the present application.

[0137] In one case, the aforementioned "combined information" can be directly used as the "first information". In another case, the vehicle can perform type encoding on the aforementioned combined information to obtain the first information; the purpose of performing type encoding is to indicate that the first text information and the first feature information are different types (or can also be called "different modes") of information.

[0138] For a more intuitive understanding of this solution, please refer to Figure 5a , Figure 5b and Figure 6 , Figure 5a A schematic diagram of the first information provided in the embodiment of the present application, Figure 5b Another schematic diagram of the first information provided in the embodiment of the present application, Figure 6 A schematic diagram of a vehicle obtaining first information provided in an embodiment of the present application, first refer to Figure 5a As shown in 5a, the first text information in the token form and the first feature information in the token form are completely separated. Figure 5a In the method in which the first text information in token form is placed first and the first feature information in token form is placed later, the first text information in token form and the first feature information in token form are concatenated to obtain the combined information; and then the combined information is type encoded to obtain the first information. It should be understood that Figure 5a The examples are only for facilitating the understanding of this solution and are not intended to limit this solution.

[0139] Continue reading Figure 5b , Figure 5b In the example, the first text information in the form of token and the first feature information in the form of token are combined in a mixed manner. Figure 5bThe method adopted is to mix the first text information and the first feature information together according to the logic of the language. The first text information is in the form of token, "Based on the given..., please tell me how to cross this intersection." Then the first information after mixing the first text information and the first feature information can be understood as: Based on the given <first feature information>, please tell me how to cross this intersection. It should be noted that the method of text description here is to facilitate understanding of how the first text information in token form and the first feature information in token form are mixed, and is not used to limit this solution.

[0140] Continue reading Figure 6 ,like Figure 6 As shown, a training sample including environmental information around the vehicle, driving behavior information of the vehicle and navigation information is input into the first feature extraction network to obtain third feature information generated by the first feature extraction network. The first feature extraction network and the first feature processing network both belong to the first neural network. The first neural network is used to perform tasks including: trajectory prediction of objects around the vehicle, decision-making on the vehicle's behavior, and prediction of the vehicle's speed range. The third feature information is input into the second feature extraction network to obtain the first feature information generated by the second feature extraction network. The preset text corresponding to the first task is obtained to obtain the first text information. Based on the first text information and the first feature information, a type encoding operation is performed to obtain the first information, and the first information is then input into the deep learning model. It should be understood that Figure 6 The examples are only for facilitating the understanding of this solution and are not intended to limit this solution.

[0141] 302. Obtain second information corresponding to the first information, wherein the second information is obtained based on a deep learning model, and the second information corresponds to any of the following tasks: making decisions on the behavior of the vehicle, planning the trajectory of the vehicle, or controlling the vehicle.

[0142] In an embodiment of the present application, after the first device inputs the first information corresponding to the traffic scene around the vehicle into the deep learning model, it can obtain the first prediction information generated by the deep learning model, and obtain the second information corresponding to the first information based on the aforementioned first prediction information. Among them, the second information corresponds to any of the following tasks: making decisions on the behavior of the vehicle, planning the trajectory of the vehicle, controlling the vehicle, or other tasks; that is, the second information is used to indicate what behavior the vehicle will perform next, or the second information may include the next trajectory information of the vehicle, or the second information may include the vehicle control instructions of the vehicle, etc. The specific information included in the second information can be determined according to the specific application scenario of the second information, and is not limited here.

[0143] In one case, the first device is a vehicle, and if the deep learning model is deployed in the vehicle, step 302 includes: the vehicle obtains first prediction information output by the locally deployed deep learning model, and then determines the second information based on the first prediction information. In another case, the first device is a vehicle, and if the deep learning model is deployed on a cloud server, step 302 may include: the vehicle receives first prediction information sent by the server that deploys the deep learning model, and then determines the second information based on the first prediction information.

[0144] In another case, the first device is a server that deploys a deep learning model, and step 302 may include: the first device obtains first prediction information output by the locally deployed deep learning model, determines second information based on the first prediction information, and sends the second information to the vehicle.

[0145] In an embodiment of the present application, a deep learning model capable of processing text information is introduced into the field of autonomous driving. Based on the second information generated by the deep learning model that processes the text information, a decision is made on the behavior of the vehicle, the trajectory of the vehicle is planned, or the vehicle is controlled. Since the input information of the deep learning model provided by the present application contains text information that is easy for users to understand, the interpretability of the operation process of the deep learning model is improved, and the decision-making, trajectory planning or control process of the autonomous driving vehicle is made more transparent, so that users can understand the behavior of the autonomous driving vehicle more intuitively.

[0146] In the above Figure 3 On the basis of the corresponding embodiment, the specific implementation process of the training phase and application phase of the deep learning model provided in the embodiment of the present application is described below.

[0147] 1. Training Phase

[0148] In the embodiment of the present application, referring to the above description, it can be known that after completing the training operation of the deep learning model, the first task performed by the deep learning model can be to make decisions on the behavior of the vehicle, to plan the trajectory of the vehicle, or to control the vehicle. Optionally, the entire training process of the deep learning model can include at least two training stages, and the aforementioned at least two training stages are used to train the machine learning model to learn information at different levels. The present application will disclose four different training stages of the deep learning model in subsequent steps (i.e., the first training stage, the second training stage, the third training stage, and the fourth training stage in the subsequent description).

[0149] Exemplarily, in the first training stage, the purpose is to allow the deep learning model to learn basic attribute information and semantic information, for example, the deep learning model can learn basic attribute information such as speed, position or orientation, and for example, the deep learning model can learn obstacle labels, the concept of "frame", the concept of "second" or other basic semantic information.

[0150] In the second training stage, the purpose is to allow the deep learning model to learn the semantics of the directional behavior and / or action behavior of the vehicle or obstacles in multiple frames of images (or point cloud data). For example, the aforementioned "directional behavior" may refer to lateral behavior, longitudinal behavior or other behaviors, and the aforementioned "action behavior" may refer to acceleration behavior, steering behavior, lane changing behavior or other behaviors; and / or, to allow the deep learning model to learn the semantics of the relative spatial relationship between different objects in a single frame image (or point cloud data), for example, learning that obstacle 1 and obstacle 2 are the front and rear vehicle relationship, obstacle 1 is within road topology 1, etc.

[0151] The third training stage is to allow the deep learning model to learn the semantics of the interaction behaviors between different objects. The aforementioned "interaction behavior" can also be called "game behavior", "interaction relationship" or other names with equivalent meanings. For example, the aforementioned "interaction behavior" can be obstacle 1 pressing on obstacle 2, and obstacle 2 avoiding it by moving sideways to the left. For another example, the aforementioned "interaction behavior" can be the interaction behavior of the vehicle changing lanes and overtaking with obstacle 3.

[0152] In the fourth training stage, the purpose is to let the deep learning model learn the cause and effect and logical relationship of the interaction between different objects. For example, because the obstacle 3 in front of the vehicle is moving slowly, and the lane on the left of the vehicle is idle, in order to improve driving efficiency, the vehicle changes lanes to overtake, etc. It should be noted that the examples here are only for the convenience of understanding this solution and are not used to limit this solution.

[0153] For details, please refer to Figure 7 , Figure 7 Another flowchart of the information processing method provided in the embodiment of the present application is provided. The information processing method provided in the present application may include:

[0154] 701. In the first training stage, the training device obtains the first-level questions corresponding to the traffic scene.

[0155] In the embodiment of the present application, for example, the "first level of questions" includes at least the initial characteristic information of the first prompt, and the "first level of questions" may also include the first characteristic information, the first characteristic information includes the characteristic information of the environmental information around the vehicle, and the environmental information includes the physical attribute information of the objects around the vehicle. The meaning of "first characteristic information" can be found in the above Figure 3The description in the corresponding embodiment is not repeated here. For example, the first prompt may include at least one mask (MASK), and the content of the MASK is predicted by the deep learning model, that is, the content of the MASK in the first prompt is completed by the deep learning model. Since the first training stage is to allow the deep learning model to learn basic attribute information and semantic information, the content of the MASK predicted by the deep learning model each time may include at least one of the following data frames, vehicle or obstacle labels, attribute names or attribute values.

[0156] For example, the "first prompt" can be specifically expressed as "In frame 0, the vehicle's driving speed is ()", and the part in "()" is a MASK, which is the part that needs to be predicted by the deep learning model. In the above example, the attribute value of the attribute "driving speed" is predicted by the deep learning model. For another example, the "first prompt" can be specifically expressed as "In frame 0, the x-coordinate position of the vehicle is ()", and the attribute value of the attribute "x-coordinate position" is predicted by the deep learning model in the above example.

[0157] For another example, the "first prompt" can be specifically expressed as "In the -1 frame, the speed of () is 10m / s", and the aforementioned example predicts the vehicle or obstacle label through the deep learning model. For another example, the "first prompt" can be specifically expressed as "In the () frame, the y coordinate position of obstacle 1 is 2.9m", and the aforementioned example predicts the frame of data through the deep learning model. For another example, the "first prompt" can be specifically expressed as "In the -1 frame, the speed of the vehicle () is 10m / s", and the aforementioned example predicts the attribute name through the deep learning model. It should be noted that the various examples here are only for the convenience of understanding the concept of "first prompt" and are not used to limit this solution.

[0158] Step 701 may include: the training device obtains a first prompt, performs feature extraction on the first prompt to obtain initial feature information of the first prompt, and obtains a first-level question corresponding to the traffic scene based on the initial feature information of the first prompt and the first feature information. The aforementioned "feature extraction" can also be understood as "initial encoding"; illustratively, the initial feature information of the first prompt can be specifically expressed in the form of a token. The specific implementation method of the training device "obtaining a first-level question corresponding to the traffic scene based on the initial feature information of the first prompt and the first feature information" is the same as Figure 3 The specific implementation method of "generating the first information according to the first text information and the first feature information" in the corresponding embodiment is similar, except that Figure 3 The "first text information" in the corresponding embodiment is replaced by the "initial feature information of the first prompt" in step 701. Figure 3The "first information" in the corresponding embodiment is replaced by the "first-level problem" in step 701, and the specific implementation method of the aforementioned steps will not be repeated here.

[0159] A specific implementation method for obtaining the first feature information by the training device. A training data set may be deployed in the training device, and the training data set includes multiple training samples; wherein each training sample may include environmental information around the vehicle, and optionally, each training sample may also include driving behavior information of the vehicle and / or navigation information of the vehicle. The training device inputs the training sample into the first feature extraction network to obtain the third feature information generated by the first feature extraction network, and the third feature information is used to obtain the first feature information.

[0160] It should be noted that the meanings of “environmental information around the vehicle”, “driving behavior information of the vehicle”, “navigation information of the vehicle”, “third characteristic information” and “first characteristic information” can all be referred to in Figure 3 In the description of the corresponding embodiment, the specific implementation method of "obtaining the first characteristic information based on the third characteristic information" can also be referred to Figure 3 The description in the corresponding embodiment will not be repeated here.

[0161] Optionally, the training device can also input the first feature information into the first feature processing network to obtain second prediction information generated by the first feature processing network corresponding to the training sample (that is, the second prediction information corresponding to the environmental information around the vehicle). The first feature extraction network and the first feature processing network belong to the same first neural network, and "the second prediction information generated by the first feature processing network" can also be understood as "the second prediction information generated by the first neural network."

[0162] The first neural network is used to perform at least one of the following tasks: predicting the trajectory of objects around the vehicle, making decisions on the vehicle's behavior, planning the trajectory of the vehicle, predicting the speed range of the vehicle, controlling the vehicle, or other tasks; optionally, the first neural network is used to perform at least two of the aforementioned tasks.

[0163] A specific implementation method for obtaining the first prompt for the training device. In one implementation method, a second neural network may be pre-deployed in the training device, and the training device inputs the training sample into the second neural network to obtain a first prompt corresponding to the first level generated by the second neural network. Exemplarily, the second neural network can be understood as a generator, and the second neural network can specifically adopt a neural network based on an attention mechanism, a convolutional neural network, a recurrent neural network, or other types of neural networks, etc., which are not limited here.

[0164] Optionally, the second neural network can be used to generate prompts in the first training stage, the second training stage, the third training stage, and the fourth training stage. In the first training stage, the training device can input the training sample and the first parameter into the second neural network to obtain the first prompt generated by the second neural network, and the first parameter is used to instruct the second neural network to generate the aforementioned first prompt corresponding to the first level. In the subsequent second training stage, the training device can input the training sample and the second parameter into the second neural network to obtain the second prompt generated by the second neural network, and the second parameter is used to instruct the second neural network to generate the aforementioned second prompt corresponding to the second level. In the subsequent third training stage, the training device can input the training sample and the third parameter into the second neural network to obtain the third prompt generated by the second neural network, and the third parameter is used to instruct the second neural network to generate the aforementioned third prompt corresponding to the third level. In the subsequent fourth training stage, the training device can input the training sample and the fourth parameter into the second neural network to obtain the fourth prompt generated by the second neural network, and the fourth parameter is used to instruct the second neural network to generate the aforementioned fourth prompt corresponding to the fourth level.

[0165] Exemplarily, the first parameter, the second parameter, the third parameter, and the fourth parameter are all different. For example, the first parameter may be 1111, the second parameter may be 2222, the third parameter may be 3333, and the fourth parameter may be 4444; for another example, the first parameter may be the first level, the second parameter may be the second level, the third parameter may be the third level, and the fourth parameter may be the fourth level, etc. It should be noted that the examples given here are only for the convenience of understanding the differences between the first parameter, the second parameter, the third parameter, and the fourth parameter, and are not used to limit this solution.

[0166] It should be noted that in other implementations, different neural networks may be used in different training stages of the deep learning model to generate prompts at different levels. In this case, it is only necessary to input training samples into the neural network used to generate prompts, and it is no longer necessary to input additional parameters (such as the first parameter mentioned above) into the neural network used to generate prompts.

[0167] In an embodiment of the present application, during the training stage of the deep learning model, a second neural network is used to generate questions that need to be answered by the deep learning model, and a rich variety of questions can be generated. In order to be able to answer the aforementioned rich and diverse questions, it is beneficial for the deep learning model to fully understand various information and to improve the deep learning model's ability to understand traffic scenes. In addition, training samples containing environmental information around the vehicle are input into the second neural network to obtain the questions output by the second neural network, that is, the traffic scene corresponding to the questions generated by the second neural network, and the traffic scene that needs to be understood by the deep learning model is the same traffic scene, so that the questions output by the second neural network can be closer to the traffic scene that needs to be understood by the deep learning model, thereby helping to improve the deep learning model's ability to understand traffic scenes.

[0168] In one implementation, multiple prompts corresponding to the first level may also be pre-deployed in the training device. Then, each time the training device obtains a question of the first level, it may obtain a first prompt corresponding to the first level from the multiple prompts corresponding to the first level.

[0169] 702. The training device inputs a first-level question into the deep learning model to obtain a first predicted answer corresponding to the first-level question.

[0170] In an embodiment of the present application, after inputting a first-level question into a deep learning model, the training device may obtain third prediction information corresponding to the first-level question generated by the deep learning model; and obtain a first predicted answer corresponding to the first-level question based on the third prediction information generated by the deep learning model. Exemplarily, the third prediction information may indicate at least one alternative answer corresponding to the first-level question and a first probability value corresponding to each alternative answer.

[0171] Exemplarily, each alternative answer may be specifically expressed in the form of a token. The alternative answer in the form of a token may also be understood as feature information of the alternative answer. The alternative answer in the form of a token is used to indicate one or more words.

[0172] Optionally, a start mark [CLS] bit may be inserted into the head of each alternative answer in token form. The start mark [CLS] bit in the alternative answer in token form can not only identify the beginning of the alternative answer, but also represent the overall meaning of the alternative answer. The characteristic information of each alternative answer also includes characteristic information of the [CLS] bit. Optionally, an end mark [SEP] bit may be inserted into the tail of each alternative answer in token form. The characteristic information of each alternative answer also includes characteristic information of the [SEP] bit.

[0173] For a more intuitive understanding of this solution, please refer to Figure 8 , Figure 8 A schematic diagram of alternative answers provided in the embodiment of the present application, such as Figure 8 As shown, the first part of the alternative answer in token form has a start flag [CLS] bit, and the tail part of the alternative answer in token form has an end flag [SEP] bit. It should be understood that Figure 8 The examples are only for facilitating the understanding of this solution and are not intended to limit this solution.

[0174] In one implementation, after obtaining the third prediction information, the training device can determine at least one alternative answer indicated by the third prediction information and a first probability value corresponding to each alternative answer, and then can determine a first predicted answer from the at least one alternative answer, wherein the first probability value corresponding to the first predicted answer is the highest among the at least one alternative answer.

[0175] In another implementation, the first predicted answer is included in a first preset answer set corresponding to the question of the first level, and the first predicted answer is obtained from at least one alternative answer indicated by the third prediction information based on the first preset answer set.

[0176] Exemplarily, a first preset answer set may be deployed in the training device, and the first preset answer set includes one or more answers. After the training device obtains the third prediction information, the intersection between at least one alternative answer indicated by the third prediction information and the first preset answer set may be determined, and the aforementioned intersection includes at least one first alternative answer; the training device may determine the first predicted answer from at least one first alternative answer based on the first probability value corresponding to each first alternative answer in the at least one first alternative answer, and the first probability value corresponding to the first predicted answer is the largest among the at least one first alternative answer.

[0177] In order to further understand the concept of "the intersection between at least one alternative answer and the first preset answer set", for example, the first level of the problem is specifically manifested as follows: based on the given first feature information, in the -1 frame, the speed of () is 10m / s, the content in the aforementioned () is the content that needs to be predicted by the deep learning model, and the first preset answer set may include: the vehicle, obstacle x, social vehicles, other vehicles, electric vehicles and pedestrians; obstacle x means obstacle + label, for example, obstacle 1, obstacle 2, obstacle 3 or obstacle 4, etc. all meet the requirements of obstacle + label, that is, they all meet the requirements of obstacle x. After the training device obtains the third prediction information output by the deep learning model, it determines at least one alternative answer indicated by the third prediction information. For example, at least one alternative answer may include: obstacle 1, street lamp, vehicle, Chinese textbook and obstacle 2, then the intersection between at least one alternative answer and the first preset answer set may include: obstacle 1, vehicle and obstacle 2. It should be understood that the examples here are only for the convenience of understanding this solution and are not used to limit this solution.

[0178] In another implementation, the first predicted answer can be obtained through a classification network, which is used to determine the first predicted answer corresponding to the second feature information from at least one preset answer. The first predicted answer can be one of the at least one preset answer. The second feature information may include a second alternative answer in token form (also referred to as feature information of the second alternative answer) including feature information of a start flag [CLS] bit. The first probability value corresponding to the second alternative answer is the largest among at least one alternative answer.

[0179] Exemplarily, a trained classification network may be deployed in the training device. After obtaining the third prediction information, the training device may determine the feature information of the second candidate answer from the third prediction information, obtain the second feature information from the feature information of the second candidate answer, input the second feature information into the classification network, and obtain the first prediction answer generated by the classification network. The classification network is used to determine a prediction category to which the second feature information belongs in M ​​categories, and the M categories may be at least one of the above-mentioned preset answers.

[0180] Exemplarily, the above-mentioned “classification network” may also be referred to as a “classifier”. The classification network may be specifically expressed as a recurrent neural network, a fully connected neural network, a convolutional neural network or other types of neural networks, etc., which is not limited in the embodiments of the present application.

[0181] 703. The training device trains the deep learning model according to the first loss function. In the first training stage, the first loss function indicates the similarity between the first predicted answer and the first expected answer corresponding to the first level question.

[0182] In an embodiment of the present application, exemplarily, in the first training stage, the training device may obtain the first expected answer corresponding to the above-mentioned first-level question, and the aforementioned "first expected answer corresponding to the first-level question" may also be referred to as "the first true value corresponding to the first-level question", "the first annotation information corresponding to the first-level question", "the first label corresponding to the first-level question" or other names with equivalent meanings. The training device may generate a function value of a first loss function based on the first predicted answer and the first expected answer. The goal of training using the first loss function includes improving the similarity between the first predicted answer and the first expected answer corresponding to the first-level question. Based on the function value of the first loss function, the training device uses a back-propagation algorithm to update the weight parameters of the deep learning model to achieve one-time training of the deep learning model.

[0183] Optionally, step 703 may include: the training device trains the deep learning model and the first neural network according to a first loss function and a second loss function, the second loss function indicates the similarity between the second prediction information corresponding to the training sample and the second expected information, and the goal of training using the second loss function includes improving the similarity between the second prediction information corresponding to the training sample and the second expected information.

[0184] Exemplarily, what content the second expected information carries is related to the type of task performed by the first neural network.

[0185] Exemplarily, the training device can generate a function value of a first loss function according to the first predicted answer and the first expected answer; and generate a function value of a second loss function according to the second predicted information and the second expected information corresponding to the training sample. According to the function value of the first loss function and the function value of the second loss function, a total loss function value can be determined; according to the total loss function value, a back propagation algorithm is used to update the weight parameters of the deep learning model and the first neural network to achieve one-time training of the deep learning model and the first neural network.

[0186] The training device can repeat steps 701 to 703 multiple times until the first convergence condition is met, so as to train the deep learning model multiple times in the first training stage; wherein the first convergence condition can be that steps 701 to 703 reach the first number, or the first convergence condition can be the convergence condition of satisfying the first loss function (optionally, also including the second loss function).

[0187] It should be noted that steps 701 to 703 are optional steps. If steps 701 to 703 are not performed, step 704 may be performed directly.

[0188] 704. In the second training phase, the training device obtains second-level questions corresponding to the traffic scenario.

[0189] In the embodiment of the present application, for example, the "second-level question" includes at least the initial characteristic information of the second prompt, and the "second-level question" may also include the first characteristic information. The meaning of "first characteristic information" can be found in the above Figure 3 The description in the corresponding embodiment is not repeated here. Exemplarily, the second prompt may include at least one mask (MASK), and the content of the MASK is predicted by the deep learning model, that is, the content of the MASK in the second prompt is completed by the deep learning model.

[0190] Since the second training stage is to let the deep learning model learn: the semantics of the directional behavior of the ego vehicle or obstacles, the semantics of the action behavior of the ego vehicle or obstacles, or the relative spatial relationship between different objects; the content of the MASK predicted by the deep learning model may include at least one of the following: directional behavior, action behavior, or the relative spatial relationship between different objects. Optionally, in the second training stage, the content of the MASK predicted by the deep learning model may also include data frames, ego vehicle or obstacle labels, or other content.

[0191] For example, the "second prompt" can be specifically expressed as "In frame 0, the vehicle has undergone () in the horizontal direction". The part in "()" is a MASK, that is, the part that needs to be predicted by the deep learning model. The aforementioned example predicts action behaviors through the deep learning model, for example, the content in "()" can be a left lane change. For another example, the "second prompt" can be specifically expressed as "In frame 0, the vehicle has undergone (left lane change) in (). The aforementioned example predicts directional behaviors through the deep learning model, for example, the content in "()" can be a horizontal direction, etc. It should be noted that the various examples here are only for the convenience of understanding the concept of the "second prompt" and are not used to limit this solution.

[0192] The specific implementation method of step 704 is similar to the specific implementation method of step 701, the difference is that the "first prompt" in step 701 is replaced by the "second prompt" in step 704, and the "first-level question" in step 701 is replaced by the "second-level question" in step 704. The specific implementation method of step 704 will not be repeated here.

[0193] 705. The training device inputs the second-level question into the deep learning model to obtain a second predicted answer corresponding to the second-level question.

[0194] 706. The training device trains the deep learning model according to the first loss function. In the second training stage, the first loss function indicates the similarity between the second predicted answer corresponding to the second level question and the second expected answer.

[0195] In the embodiment of the present application, the meanings of the nouns in steps 705 and 706 and the specific implementation methods of the steps can refer to the descriptions in steps 702 and 703. The difference is that the "first-level question" in steps 702 and 703 is replaced by "second-level question", the "first predicted answer" in steps 702 and 703 is replaced by "second predicted answer", and the "first expected answer" in steps 702 and 703 is replaced by "second expected answer". The specific implementation methods of steps 705 and 706 will not be repeated here.

[0196] It should be noted that steps 704 to 706 are optional steps. If steps 704 to 706 are not executed, step 707 can be directly executed after executing step 703; if steps 701 to 707 are not executed, step 707 can be directly executed; if steps 701 to 707 are executed, steps 701 to 703 can also be executed crosswise during the process of the training device executing steps 704 to 706 multiple times.

[0197] 707. In the third training stage, the training device obtains the third level of questions corresponding to the traffic scene.

[0198] In the embodiment of the present application, for example, the "third level question" includes at least the initial characteristic information of the third prompt, and the "third level question" may also include the first characteristic information. The meaning of "first characteristic information" can be found in the above Figure 3 The description in the corresponding embodiment is not repeated here. Exemplarily, the third prompt may include at least one mask (MASK), and the content of the MASK is predicted by the deep learning model, that is, the content of the MASK in the third prompt is completed by the deep learning model.

[0199] Since the third training stage is to allow the deep learning model to learn the semantics of the interaction behaviors between different objects, the content of the MASK predicted by the deep learning model may include interaction behaviors; optionally, in the third training stage, the content of the MASK predicted by the deep learning model may also include directional behaviors and / or action behaviors.

[0200] For example, the "third prompt" can be specifically expressed as "the self-car and social car 3 have interactive game behavior, the category is ()", and the part in "()" is MASK, that is, the part that needs to be predicted by the deep learning model. The above example predicts the interactive behavior through the deep learning model. For example, the content in "()" can be overtaking, etc. It should be noted that the various examples here are only for the convenience of understanding the concept of the "third prompt" and are not used to limit this solution.

[0201] The specific implementation method of step 707 is similar to the specific implementation method of step 701, the difference is that the "first prompt" in step 701 is replaced by the "third prompt" in step 707, and the "first-level problem" in step 701 is replaced by the "third-level problem" in step 707. The specific implementation method of step 707 is not repeated here.

[0202] 708. The training device inputs the third-level question into the deep learning model to obtain a third predicted answer corresponding to the third-level question.

[0203] 709. The training device trains the deep learning model according to the first loss function. In the third training stage, the first loss function indicates the similarity between the third predicted answer corresponding to the third level question and the third expected answer.

[0204] In the embodiment of the present application, the meanings of the nouns in steps 708 and 709 and the specific implementation methods of the steps can refer to the descriptions in steps 702 and 703. The difference is that the "first-level question" in steps 702 and 703 is replaced by "third-level question", the "first predicted answer" in steps 702 and 703 is replaced by "third predicted answer", and the "first expected answer" in steps 702 and 703 is replaced by "third expected answer". The specific implementation methods of steps 705 and 709 will not be repeated here.

[0205] It should be noted that steps 707 to 709 are optional steps. If steps 707 to 709 are not executed, step 710 can be directly executed after executing step 707; if steps 701 to 709 are not executed, step 710 can be directly executed; if steps 701 to 709 are executed, then in the process of the training device executing steps 707 to 709 multiple times, steps 701 to 703 can be executed crosswise, and / or steps 704 to 706 can be executed crosswise.

[0206] 710. In the fourth training stage, the training device obtains the fourth level of questions corresponding to the traffic scenario.

[0207] In the embodiment of the present application, illustratively, the "fourth level question" includes at least the initial characteristic information of the fourth prompt, and the "fourth level question" may also include the first characteristic information. The meaning of "first characteristic information" can be found in the above Figure 3 The description in the corresponding embodiment is not repeated here. Exemplarily, the fourth prompt may include at least one mask (MASK), and the content of the MASK is predicted by the deep learning model, that is, the content of the MASK in the fourth prompt is completed by the deep learning model.

[0208] Since the fourth training stage is to allow the deep learning model to learn the causes, consequences and logical relationships of interactive behaviors between different objects, the content of the MASK predicted by the deep learning model may include the reasons for the interactive behaviors; optionally, in the fourth training stage, the content of the MASK predicted by the deep learning model may also include interactive behaviors.

[0209] For example, the "fourth prompt" can be specifically expressed as "Social car 1 and social car 3 have an interactive game behavior, the category is giving way, because ()", the part in "()" is MASK, that is, the part that needs to be predicted by the deep learning model. The above example predicts the reason for the interactive behavior of giving way through the deep learning model. It should be noted that the various examples here are only for the convenience of understanding the concept of the "fourth prompt" and are not used to limit this solution.

[0210] The specific implementation method of step 710 is similar to the specific implementation method of step 701, the difference is that the "first prompt" in step 701 is replaced by the "fourth prompt" in step 710, and the "first-level question" in step 701 is replaced by the "fourth-level question" in step 710. The specific implementation method of step 710 will not be repeated here.

[0211] 711. The training device inputs a fourth-level question into the deep learning model to obtain a fourth predicted answer corresponding to the fourth-level question.

[0212] 712. The training device trains the deep learning model according to the first loss function. In the fourth training stage, the first loss function indicates the similarity between the fourth predicted answer and the fourth expected answer corresponding to the fourth level question.

[0213] In the embodiment of the present application, the meanings of the nouns in steps 711 and 712 and the specific implementation methods of the steps can refer to the descriptions in steps 702 and 703. The difference is that the "first-level question" in steps 702 and 703 is replaced by "fourth-level question", the "first predicted answer" in steps 702 and 703 is replaced by "fourth predicted answer", and the "first expected answer" in steps 702 and 703 is replaced by "fourth expected answer". The specific implementation methods of steps 705 and 712 will not be repeated here.

[0214] It should be noted that steps 710 to 712 are optional steps. If steps 710 to 712 are not executed, then after executing step 710, the re-training of the deep learning model can be directly stopped; if steps 701 to 712 are all executed, then in the process of the training device executing steps 710 to 712 multiple times, steps 701 to 703 can be cross-executed, and / or steps 704 to 706 can be cross-executed, and / or steps 707 to 709 can be cross-executed.

[0215] In order to more intuitively understand the "first prompt", "second prompt", "third prompt" and "fourth prompt", the following is an example of a traffic scenario. Fig. 9 , Fig. 9 A schematic diagram of a traffic scenario provided in an embodiment of the present application. Fig. 9 As shown, Fig. 9 The traffic scene of "the vehicle changing lanes to overtake" is shown in FIG. For example, in this traffic scene, the first prompt used in the first training stage may include: in the (0) frame, the (driving speed) of the (vehicle) is (15m / s). For another example, the first prompt may include: in the (-2) frame, the (x-coordinate position) of (obstacle 1) is (16.8m). It should be noted that the content in "()" can be set to [MASK], and the content of [MASK] is the content that needs to be predicted by the deep learning model. In each of the above examples of the first prompt, there are multiple "()". In each actual training process, the content in one of the multiple "()" can be set to [MASK].

[0216] For example, the second prompt used in the second training stage may include: from the (-10)th frame to the (0)th frame, (the vehicle) performed (lane change to the left) in (lateral behavior); for another example, from the (-10)th frame to the (0)th frame, (the vehicle) performed (slight acceleration) in (longitudinal behavior). The content of one of the multiple "()"s included in each of the aforementioned second prompts can be set to [MASK], and the deep learning model is used to predict the content of [MASK]. It can be seen from the aforementioned examples that the second training stage is to allow the machine learning model to learn the behavior of objects and the relationship between different objects.

[0217] For example, the third prompt used in the third training stage may include: from the (-9)th frame to the (-2)th frame, (the self-car) and (the social car 1) had an interactive game behavior, the category is (lane change and overtaking), specifically: (the self-car changes lanes to the left and overtakes the social car 1), the content of one of the multiple "()"s included in each of the aforementioned second prompts can be set to [MASK], and the deep learning model is used to predict the content of [MASK]. From the aforementioned examples, it can be seen that the third training stage is to allow the machine learning model to learn the interactive behaviors between different objects.

[0218] For example, the fourth prompt used in the fourth training stage may include: In the past 3 seconds, the self-vehicle changed lanes to overtake because (the social vehicle 1 in front of the self-vehicle was driving slowly, and the left lane was idle. In order to improve driving efficiency, the self-vehicle changed lanes to overtake). The content in "()" can be set to [MASK]. From the above examples, it can be seen that the fourth training stage is to allow the machine learning model to learn the causal relationship of interactive behavior. It should be understood that Fig. 9 The examples are only for facilitating the understanding of this solution and are not intended to limit this solution.

[0219] In an embodiment of the present application, the training process of the deep learning model may include at least two training stages, and the aforementioned at least two training stages may include at least two of the first training stage, the second training stage, the third training stage, and the fourth training stage. The specific training stages used for the deep learning model can be flexibly determined in combination with the actual application scenario, and are not limited in the embodiment of the present application. Exemplarily, the "first training stage" in the present application may refer to any one of the above-mentioned first training stage, the second training stage, the third training stage, and the fourth training stage, and the "second training stage" and the "first training stage" in the present application are different training stages.

[0220] The third question represents the question used in the first training stage. After determining which training stage the "first training stage" is, it can be determined which of the first-level questions, second-level questions, third-level questions, and fourth-level questions the third question represents. For example, if the first training stage is the first training stage, the third question represents the first-level question; for another example, if the first training stage is the second training stage, the third question represents the second-level question; for another example, if the first training stage is the third training stage, the third question represents the third-level question; for another example, if the first training stage is the fourth training stage, the third question represents the fourth-level question.

[0221] Correspondingly, the fourth question represents the question used in the second training stage. After determining which training stage the "second training stage" is, it can be determined which of the first-level questions, second-level questions, third-level questions, and fourth-level questions the third question represents.

[0222] Exemplarily, at least two training stages may include a first training stage and a second training stage; or, at least two training stages may include a second training stage and a third training stage; or, at least two training stages may include a third training stage and a fourth training stage; or, at least two training stages may include a first training stage, a second training stage, and a third training stage; or, at least two training stages may include a first training stage, a second training stage, and a fourth training stage; or, at least two training stages may include a second training stage, a third training stage, and a fourth training stage; or, at least two training stages may include a first training stage, a second training stage, a third training stage, and a fourth training stage, etc. The training stages used in the training process of a specific deep learning model can be flexibly determined in combination with the actual application scenario, and are not limited in the embodiments of the present application.

[0223] In an embodiment of the present application, at least two training stages are used to train the deep learning model. The at least two training stages respectively use the third question and the fourth question. The third question and the fourth question are used to obtain information of different levels from traffic scenes. That is, the training process of the deep learning model adopts a step-by-step training method, which is conducive to enabling the deep learning model to learn logical thinking of step-by-step reasoning, thereby improving the human-likeness of the deep learning model, and also helping the trained deep learning model to be more reasonable when performing tasks related to autonomous driving.

[0224] 2. Application Phase

[0225] In the embodiments of this application, please refer to Fig.10 , Fig.10Another flowchart of the information processing method provided in the embodiment of the present application is provided. The information processing method provided in the present application may include:

[0226] 1001. The execution device inputs a first question into the deep learning model to obtain a first answer corresponding to the first question.

[0227] In the embodiment of the present application, step 1001 is an optional step. Exemplarily, the first information in the present application may be a first question, and the second information in the present application may be a second answer corresponding to the second question. Optionally, before inputting the first information into the deep learning model, the training device may also input the first question into the deep learning model to obtain a first answer corresponding to the first question.

[0228] Exemplarily, the first question may include a prompt. Optionally, the first question may also include first characteristic information. The first characteristic information includes characteristic information of the environment information around the vehicle. For a detailed description of the meaning of "first characteristic information" and "vehicle generating first characteristic information", please refer to Figure 3 The description in the corresponding embodiment is not repeated here.

[0229] The first question and the second question may correspond to the same traffic scene (hereinafter referred to as the "target traffic scene" for convenience of description). Optionally, the first question and the second question may be used to obtain different levels of information from the target traffic scene. For the concept of "different levels", please refer to the above Figure 7 For example, the first question is used to obtain basic attribute information and semantic information in the target traffic scene. The meaning of the aforementioned “basic attribute information and semantic information” can be found in the above Figure 7 The description in the corresponding embodiment; for example, the first question is used to obtain the behavior of the object in the target traffic scene or the relationship between different objects, etc. The specific expression of the "first question" can be determined in combination with the actual application scenario and is not limited here.

[0230] Exemplarily, the execution device refers to a device deployed with a machine learning model that has performed a trained operation. In one case, the execution device can be specifically manifested as a vehicle, that is, a trained deep learning model is deployed in the vehicle; in another case, the execution device can be a device other than a vehicle, and the vehicle can send the first question to the execution device, and the execution device inputs the first question to the deep learning model; in another case, the execution device can be a device other than a vehicle, and after the vehicle sends the first information to the execution device, the execution device can generate the first question and input the first question to the deep learning model. Exemplarily, after inputting the first question to the deep learning model, the execution device can obtain the prediction information generated by the deep learning model, and determine the first answer corresponding to the first question based on the prediction information generated by the deep learning model.

[0231] 1002. The execution device determines a second question based on the first answer, and the second question contains information in the first answer.

[0232] In the embodiment of the present application, step 1002 is an optional step. After determining the first answer corresponding to the first question, the execution device may generate a second question based on the first answer. Exemplarily, after determining the first answer, the execution device may update the second question generated for the vehicle (which may also be understood as the "second question before the update") based on the first answer to obtain an updated second question, and the "updated second question" may also be understood as the "first information".

[0233] Exemplarily, the execution device may perform feature extraction on all or part of the information included in the first answer, and after obtaining initial feature information of all or part of the information included in the first answer, add the aforementioned initial feature information to the prompt of the second question before the update to obtain the updated second question.

[0234] Exemplarily, the second answer corresponding to the second question may correspond to any of the following tasks: identifying risk obstacles around the vehicle, identifying the behavior of objects around the vehicle, predicting the behavior of objects around the vehicle, predicting the trajectory of objects around the vehicle, making decisions on the behavior of the vehicle, planning the trajectory of the vehicle, or controlling the vehicle. That is, the second question is used to obtain information corresponding to the aforementioned task, and the second question may correspond to any of the aforementioned multiple tasks.

[0235] For example, the second answer corresponds to the task of "identifying the behavior of objects around the vehicle", and the semantics of the prompt of the second question before the update generated by the vehicle may include: from the ()th frame to the ()th frame, the obstacle that has an interactive game relationship with the vehicle is (), and the specific behavior of the obstacle is (). The information carried in the first answer may include: from the -26th frame to the -5th frame, the obstacle that has an interactive game relationship with the vehicle is obstacle 3. The semantics of the prompt in the updated second question may include: from the -26th frame to the -5th frame, the obstacle that has an interactive game relationship with the vehicle is obstacle 3, and the specific behavior of the obstacle is (); the content of "()" in the updated second question is the content that needs to be predicted by the deep learning model.

[0236] For another example, the second answer corresponds to the task of "predicting the behavior of objects around the vehicle". The semantics of the prompt for the second question generated by the vehicle before the update may include: in frame 0, the topology where obstacle 1 is located is (); based on the topological relationship of obstacle 1 and the historical frame behavior, the navigation behavior that obstacle 1 will take is (), because (). The information carried in the first answer may include: the topology where obstacle 1 is located is topology 1. The semantics of the prompt in the updated second question may include: in frame 0, the topology where obstacle 1 is located is topology 1; based on the topological relationship of obstacle 1 and the historical frame behavior, the navigation behavior that obstacle 1 will take is (), because (); the content of "()" in the updated second question is the content that needs to be predicted by the deep learning model.

[0237] For another example, the second answer corresponds to the task of "predicting the behavior of objects around the vehicle". The semantics of the prompt for the second question generated by the vehicle before the update may include: in frame 0, the object that interacts with obstacle 1 is (). According to the topological relationship between obstacle 1 and () and the historical frame behavior, the interactive behavior that obstacle 1 will have is (). The information carried in the first answer may include: the object that interacts with obstacle 1 is obstacle 2. The semantics of the prompt in the updated second question may include: in frame 0, the object that interacts with obstacle 1 is obstacle 2. According to the topological relationship between obstacle 1 and obstacle 2 and the historical frame behavior, the interactive behavior that obstacle 1 will have is (); the content of "()" in the updated second question is the content that needs to be predicted by the deep learning model.

[0238] For another example, the second answer corresponds to the task of "making decisions on the behavior of the ego vehicle". The semantics of the prompt of the second question before the update generated by the vehicle may include: in frame 0, the topology of the ego vehicle is (), and the obstacles that have an interactive relationship with the ego vehicle are (). According to the topological relationship between the ego vehicle and the obstacle with an interactive relationship () and the historical frame behavior, the next behavior performed by the ego vehicle is (), and the reason for deciding to perform the aforementioned interactive behavior is (). The information carried in the first answer may include: in frame 0, the topology of the ego vehicle is topology 2, and the obstacles that have an interactive relationship with the ego vehicle are obstacle 4. The semantics of the prompt in the updated second question may include: in frame 0, the topology of the ego vehicle is topology 2, and the obstacles that have an interactive relationship with the ego vehicle are obstacle 4. According to the topological relationship between the ego vehicle and the obstacle with an interactive relationship (obstacle 4) and the historical frame behavior, the next behavior performed by the ego vehicle is (), and the reason for deciding to perform the aforementioned interactive behavior is (); the content of "()" in the updated second question is the content that needs to be predicted by the deep learning model.

[0239] It should be noted that the above examples of "the first text in the first information (which can also be understood as the prompt in the second question)", "the information carried in the first answer" and "the updated second question" are only for the convenience of understanding the relationship between the "first text information", "the information in the first answer" and "the updated second question", and are not used to limit this solution.

[0240] In an embodiment of the present application, a first answer to a first question is first obtained, and then a second question is generated based on the first answer, and then the answer to the second question is obtained. The first question and the second question correspond to the same traffic scenario, that is, the questions asked to the deep learning model are asked in a progressive manner, which is conducive to reducing the difficulty of the deep learning model in answering the second question, and is also conducive to improving the ability of the deep learning model in dealing with problems corresponding to complex traffic scenarios; in addition, the use of a progressive questioning method is also conducive to introducing the logical thinking of step-by-step reasoning into the deep learning model used in the field of autonomous driving, thereby improving the human-likeness of the second answer finally obtained.

[0241] 1003. The execution device inputs first information corresponding to the traffic scene around the vehicle into the deep learning model, and obtains second information corresponding to the first information, wherein the first information includes first text information.

[0242] In the embodiment of the present application, since steps 1001 and 1002 are optional steps, if steps 1001 and 1002 are not performed, in step 1003, the first text information in the "first information" is generated by the vehicle; if steps 1001 and 1002 are performed, in step 1003, the "first information" can be understood as the "updated second question", and the first text information in the first information is determined based on the second question and the first answer before the update generated by the vehicle.

[0243] Exemplarily, if the first text information includes initial feature information of a preset text corresponding to the first task performed by the deep learning model, the first text information may include the initial feature information of the preset text and the initial feature information of the information in the first answer. Alternatively, the first text information includes initial feature information of a text input by a user, the first text information may include initial feature information of the text input by the user and the initial feature information of the information in the first answer.

[0244] In an embodiment of the present application, two situations in which the first text information may include information are listed, thereby improving the implementation flexibility of the present solution; when the first text information includes information of a preset text corresponding to the first task performed by the deep learning model, then after the first task is determined, it is beneficial to quickly determine the first text information to improve the efficiency of obtaining the second information; when the first text information includes text information input by the user, that is, questions can be answered based on the user's questions, which is beneficial to improving the user stickiness of the present solution and making it more convenient for users to understand the behavior of the vehicle during the autonomous driving process.

[0245] Exemplarily, step 1003 may include: the execution device inputs first information corresponding to the traffic scene around the vehicle into the deep learning model to obtain first prediction information generated by the deep learning model; according to the first prediction information, second information corresponding to the first information is determined; the specific implementation of the above steps can be referred to in Figure 3 The description in the corresponding embodiment is not repeated here. The specific expression form of the "first prediction information" is the same as Figure 7 The specific expression form of the "third prediction information" in the corresponding embodiment is similar, which can be referred to for understanding and will not be elaborated here.

[0246] Exemplarily, the first information is the second question (or the updated second question), and the second information is the second answer corresponding to the second question (the updated second question). The second answer corresponds to any of the following tasks: identifying risk obstacles around the vehicle, identifying the behavior of objects around the vehicle, predicting the behavior of objects around the vehicle, predicting the trajectory of objects around the vehicle, making decisions on the behavior of the vehicle, planning the trajectory of the vehicle, controlling the vehicle, or other tasks in the process of autonomous driving, etc., which are not limited here. In the embodiment of the present application, a variety of tasks performed by a deep learning model are provided, which improves the implementation flexibility of the present solution.

[0247] Specifically, in one implementation, the execution device inputs a second question into the deep learning model, and obtains a first prediction information generated by the deep learning model corresponding to the aforementioned second question; and then a second answer corresponding to the second question can be determined based on the aforementioned first prediction information.

[0248] Optionally, in one case, the second answer is included in a second preset answer set corresponding to the second question, and the second answer is obtained from at least one alternative answer generated from the deep learning model based on the second preset answer set. Exemplarily, a second preset answer set may be deployed in the execution device, and the second preset answer set includes one or more answers. After the execution device obtains the first prediction information, the intersection between at least one alternative answer indicated by the first prediction information and the second preset answer set may be determined, and the aforementioned intersection includes at least one first alternative answer; the execution device may determine the second answer from at least one first alternative answer based on the first probability value corresponding to each first alternative answer in the at least one first alternative answer, and the first probability value corresponding to the second answer is the largest among the at least one first alternative answer.

[0249] In order to further understand the concept of "the intersection between at least one alternative answer and the second preset answer set", illustratively, the semantics of the second question is: Based on the given first feature information, what is the next lateral action decision of the vehicle? (); wherein the content in the aforementioned () is the content that needs to be predicted by the deep learning model. The second preset answer set may include: left lane change, left bypass, hold, right bypass and right lane change; at least one alternative answer indicated by the first prediction information may include: acceleration, left lane change, take-off, hold and right lane change; the intersection between at least one alternative answer and the second preset answer set may include: left lane change, hold and right lane change. It should be understood that the examples here are only for the convenience of understanding this solution and are not used to limit this solution.

[0250] In an embodiment of the present application, since the deep learning model may sometimes give irrelevant answers or speak nonsense, that is, at least one alternative answer determined by the deep learning model may contain an answer that is completely irrelevant to the question, in order to avoid the occurrence of the aforementioned problem, a second preset answer set corresponding to the second question can be deployed, and at least one answer included in the second preset answer set corresponding to the second question is an answer related to the second question. The second answer corresponding to the second question is finally determined based on the aforementioned second preset answer set and at least one alternative answer determined based on the deep learning model. This is conducive to avoiding the problem that the final second answer is irrelevant to the second question, thereby improving the accuracy of the second answer obtained by the deep learning model.

[0251] Optionally, in another case, the second answer is obtained by performing a classification network that has been trained, and the second feature information includes feature information of the start flag [CLS] bit in the feature information of the candidate answer, and the aforementioned classification network is used to determine the second answer corresponding to the second feature information from at least one preset answer; the aforementioned classification network is used to determine a prediction category to which the second feature information belongs in M ​​categories, and the M categories can be at least one of the preset answers, that is, the aforementioned classification network is used to determine a second answer to which the second feature information belongs in at least one answer. Exemplarily, a trained classification network may be deployed in the execution device, and after obtaining the first prediction information corresponding to the second question generated by the deep learning model, the execution device may determine the feature information of the second candidate answer from the first prediction information, and the first probability value of the second candidate answer is the highest among the at least one candidate answer indicated by the first prediction information; the execution device may obtain the second feature information from the feature information of the second candidate answer, and input the second feature information into the classification network to obtain the second answer generated by the classification network.

[0252] In an embodiment of the present application, since the characteristic information of the starting mark [CLS] bit (that is, the second characteristic information) can represent the overall meaning of the alternative answers, that is, it is reasonable to use the characteristic information of the starting mark [CLS] bit to replace the characteristic information of the alternative answers. The classification network is used to determine a second answer belonging to the second characteristic information from at least one preset answer, and the second answer is limited to at least one preset answer, which is conducive to avoiding the problem that the second answer finally obtained is irrelevant to the second question, thereby improving the accuracy of the second answer obtained by the deep learning model.

[0253] In another case, after obtaining the first prediction information corresponding to the second question, that is, obtaining at least one alternative answer indicated by the first prediction information and the first probability value corresponding to each alternative answer, the execution device can obtain a second answer from at least one alternative answer, and the first probability value corresponding to the second answer is the highest among the at least one alternative answer indicated by the first prediction information.

[0254] In another implementation, the execution device inputs first information corresponding to the traffic scene around the vehicle into the deep learning model, including: the execution device inputs multiple second questions into the deep learning model; exemplarily, the multiple second questions correspond to the same task, that is, the semantics of the multiple second questions are similar, and different second questions in the multiple second questions carry different prompts. The execution device obtains second information corresponding to the first information, including: the execution device obtains multiple reference answers corresponding to the multiple second questions one by one, that is, the execution device can obtain a reference answer corresponding to each second question, and the multiple reference answers are all obtained through the deep learning model. The execution device or the first device can then determine the second answer based on the multiple reference answers. The relationship between the "first device" and the "execution device" can be found in Figure 3 The description in the corresponding embodiment is not repeated here.

[0255] The process of "determining the second answer based on multiple reference answers" may include: determining the second answer based on multiple reference answers using a majority voting strategy, that is, the final second answer is included in the multiple reference answers. Fig.11 , Fig.11 A schematic diagram of multiple second questions provided in an embodiment of the present application, Fig.11 In this example, the first task performed by the deep learning model is "making decisions for the vehicle". Fig.11 The three second questions are shown in the figure. The prompts for the three second questions are "Since...therefore, [CLS] the car should ()," "Since...therefore, [CLS] the car needs ()," and "Since...therefore, [CLS] the car needs ()." It can be seen from the above description that the prompts for different second questions are different. The reference answers corresponding to the three second questions are "change lanes to the left to overtake", "slow down and stop", and "change lanes to the left to overtake". Based on the majority voting strategy, the final second answer is "change lanes to the left to overtake". It should be understood that Fig.11 The examples are only for facilitating the understanding of this solution and are not intended to limit this solution.

[0256] Exemplarily, after inputting each of the multiple second questions into the deep learning model, the execution device can obtain multiple first prediction information generated by the deep learning model and corresponding to the multiple second questions one by one, and can determine a reference answer based on each first prediction information. It should be noted that the process of "determining a reference answer based on each first prediction information" can refer to the above-mentioned process of "determining the second answer based on the first prediction information", the difference being that the above-mentioned "second answer" is replaced by an "alternative answer".

[0257] That is, in this implementation, in one case, the reference answer (i.e., the second answer) is included in the second preset answer set corresponding to the second question, and the reference answer is obtained from at least one alternative answer generated from the deep learning model based on the second preset answer set. In another case, the reference answer (i.e., the second answer) is obtained by performing a classification network that has been trained, and the second feature information includes feature information of the start flag [CLS] bit in the feature information of the alternative answer, and the aforementioned classification network is used to determine the reference answer corresponding to the second feature information from at least one preset answer. In another case, the reference answer (i.e., the second answer) is obtained directly from at least one alternative answer indicated by the first prediction information.

[0258] In an embodiment of the present application, multiple second questions are input into the deep learning model, different second questions carry different prompts, and multiple reference answers are obtained by utilizing multiple different prompts, and then the final second answer can be obtained from the multiple reference answers, thereby improving the rigor of obtaining the final second answer and facilitating improving the accuracy of the second questions finally obtained.

[0259] exist Figures 1 to 11 On the basis of the corresponding embodiments, in order to better implement the above solutions of the embodiments of the present application, the following also provides related devices for implementing the above solutions. Fig.12 , Fig.12 A structural diagram of an information processing device provided in an embodiment of the present application. The information processing device 1200 includes: an input module 1201, which is used to input first information corresponding to the traffic scene around the vehicle into the deep learning model, and the first information includes first text information; an acquisition module 1202, which is used to acquire second information corresponding to the first information, wherein the second information is obtained based on the deep learning model, and the second information corresponds to any of the following tasks: making decisions on the behavior of the vehicle, planning the trajectory of the vehicle, or controlling the vehicle.

[0260] Optionally, the first text information includes information of a preset text corresponding to the task, or the first text information includes information of a text input by a user.

[0261] Optionally, the input module 1201 is also used to input a first question corresponding to the traffic scene into the deep learning model, and the first question is used to obtain a first answer corresponding to the first question, and the first answer is obtained through the deep learning model, wherein the first information is the second question, the second information is the second answer corresponding to the second question, and the information in the first answer exists in the first text information.

[0262] Optionally, the first information also includes first characteristic information, the first characteristic information includes characteristic information of environmental information around the vehicle, and the environmental information includes physical property information of objects around the vehicle.

[0263] Optionally, the first feature information is obtained based on a feature extraction network, wherein the feature extraction network belongs to a first neural network, and the first neural network is used to perform at least two of the following tasks: trajectory prediction of objects around the vehicle, decision-making on the vehicle's behavior, trajectory planning for the vehicle, prediction of the vehicle's speed range, or control of the vehicle.

[0264] Optionally, the first information is the second question, and the second information is the second answer corresponding to the second question, wherein the second answer is included in a preset answer set corresponding to the second question, and the second answer is obtained from at least one alternative answer generated from the deep learning model based on the preset answer set.

[0265] Optionally, the first information is the second question, the second information is the second answer corresponding to the second question, the second answer is obtained through a classification network, the second feature information includes feature information of a start flag [CLS] bit in the feature information of the second answer, and the classification network is used to determine the second answer corresponding to the second feature information from at least one preset answer.

[0266] Optionally, the first information is the second question, and the second information is the second answer corresponding to the second question; the input module 1201 is specifically used to input multiple second questions into the deep learning model, and different second questions in the multiple second questions carry different prompts; the acquisition module 1202 is specifically used to obtain multiple reference answers corresponding one by one to the multiple second questions, and determine the second answer based on the multiple reference answers, and the multiple reference answers are all obtained through the deep learning model.

[0267] It should be noted that the information interaction, execution process, etc. between the modules / units in the information processing device 1200 are the same as those in the present application. Figures 3 to 11 The corresponding method embodiments are based on the same concept. For specific contents, please refer to the description in the method embodiments shown above in this application, which will not be repeated here.

[0268] Please continue reading Fig.13 , Fig.13Another structural schematic diagram of an information processing device provided in an embodiment of the present application. The information processing device 1300 includes: an input module 1301, which is used to input a first question into a deep learning model to obtain a first answer corresponding to the first question; a determination module 1302, which is used to determine a second question based on the first answer, wherein the second question carries information in the first answer; the input module 1301 is also used to input a second question into the deep learning model to obtain a second answer corresponding to the second question, wherein the first question and the second question correspond to the same traffic scene.

[0269] Optionally, the second answer corresponds to any of the following tasks: identifying risk obstacles around the vehicle, identifying the behavior of objects around the vehicle, predicting the behavior of objects around the vehicle, predicting the trajectory of objects around the vehicle, making decisions on the behavior of the vehicle, planning the trajectory of the vehicle, or controlling the vehicle.

[0270] It should be noted that the information interaction, execution process, etc. between the modules / units in the information processing device 1300 are the same as those in the present application. Figures 3 to 11 The corresponding method embodiments are based on the same concept. For specific contents, please refer to the description in the method embodiments shown above in this application, which will not be repeated here.

[0271] Please continue reading Fig.14 , Fig.14 Another structural schematic diagram of an information processing device provided in an embodiment of the present application. The information processing device 1400 is used to train a deep learning model, and the deep learning model includes at least two training stages, and the at least two training stages include a first training stage and a second training stage. The information processing device 1400 includes: an input module 1401, which is used to input a third question to the deep learning model in the first training stage to obtain a predicted answer corresponding to the third question; the input module 1401 is also used to input a fourth question to the deep learning model in the second training stage to obtain a predicted answer corresponding to the fourth question; wherein, a first loss function is used when training the deep learning model, and in the first training stage, the first loss function indicates the similarity between the predicted answer corresponding to the third question and the expected answer, and in the second training stage, the first loss function indicates the similarity between the predicted answer corresponding to the fourth question and the expected answer, and the third question and the fourth question are both related to traffic scenes, and the third question and the fourth question are used to obtain information at different levels.

[0272] Optionally, the third question includes first feature information, the first feature information includes feature information of environmental information around the vehicle, and the environmental information includes physical property information of objects around the vehicle.

[0273] Optionally, the information processing device 1400 also includes: a feature extraction module 1402, which is used to input environmental information into a feature extraction network to obtain second feature information generated by the feature extraction network, and the second feature information is used to obtain the first feature information; a feature processing module 1403, which is used to input the first feature information into a feature processing network to obtain prediction information generated by the feature processing network, the feature extraction network and the feature processing network belong to the same neural network, and the neural network is used to perform at least the following tasks: trajectory prediction of objects around the vehicle, decision-making on the behavior of the vehicle, trajectory planning of the vehicle, prediction of the speed range of the vehicle, or control of the vehicle; wherein, the training stage of the deep learning model includes training the deep learning model and the neural network using a first loss function and a second loss function, and the second loss function indicates the similarity between the prediction information corresponding to the environmental information and the second expected information.

[0274] It should be noted that the information interaction, execution process, etc. between the modules / units in the information processing device 1400 are the same as those in the present application. Figures 3 to 11 The corresponding method embodiments are based on the same concept. For specific contents, please refer to the description in the method embodiments shown above in this application, which will not be repeated here.

[0275] Next, a device provided by an embodiment of the present application is introduced. Please refer to Fig.15 , Fig.15 A schematic diagram of a structure of a device provided in an embodiment of the present application, specifically, the device 1500 includes: a receiver 1501, a transmitter 1502, a processor 1503 and a memory 1504 (wherein the number of processors 1503 in the device 1500 can be one or more, Fig.15 In the example of FIG. 1501 , the processor 1503 may include an application processor 15031 and a communication processor 15032. In some embodiments of the present application, the receiver 1501, the transmitter 1502, the processor 1503 and the memory 1504 may be connected via a bus or other means.

[0276] The memory 1504 may include a read-only memory and a random access memory, and provides instructions and data to the processor 1503. A portion of the memory 1504 may also include a non-volatile random access memory (NVRAM). The memory 1504 stores processor and operation instructions, executable modules or data structures, or subsets thereof, or extended sets thereof, wherein the operation instructions may include various operation instructions for implementing various operations.

[0277] The processor 1503 controls the operation of the device. In a specific application, the various components of the device are coupled together through a bus system, wherein the bus system includes not only a data bus but also a power bus, a control bus, and a status signal bus, etc. However, for the sake of clarity, various buses are referred to as bus systems in the figure.

[0278] The method disclosed in the above embodiment of the present application can be applied to the processor 1503, or implemented by the processor 1503. The processor 1503 can be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by the hardware integrated logic circuit or software instructions in the processor 1503. The above processor 1503 can be a general processor, a digital signal processor (digital signal processing, DSP), a microprocessor or a microcontroller, and can further include an application specific integrated circuit (application specific integrated circuit, ASIC), a field programmable gate array (field-programmable gate array, FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components. The processor 1503 can implement or execute the various methods, steps and logic block diagrams disclosed in the embodiments of the present application. The general processor can be a microprocessor or the processor can also be any conventional processor, etc. The steps of the method disclosed in the embodiment of the present application can be directly embodied as a hardware decoding processor to execute, or the hardware and software modules in the decoding processor can be combined and executed. The software module may be located in a storage medium mature in the art, such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. The storage medium is located in the memory 1504, and the processor 1503 reads the information in the memory 1504 and completes the steps of the above method in combination with its hardware.

[0279] The receiver 1501 can be used to receive input digital or character information and generate signal input related to the relevant settings and function control of the device. The transmitter 1502 can be used to output digital or character information through the first interface; the transmitter 1502 can also be used to send instructions to the disk group through the first interface to modify the data in the disk group; the transmitter 1502 can also include a display device such as a display screen.

[0280] In one embodiment of the present application, the processor 1503 is configured to execute Figures 3 to 11 The first device and the method executed by the execution device in the corresponding embodiment; in another case, the processor 1503 is used to execute Figures 3 to 11The method executed by the training device in the corresponding embodiment. It should be noted that the specific manner in which the application processor 15031 in the processor 1503 executes the above steps is the same as that in the present application. Figures 3 to 11 The corresponding method embodiments are based on the same concept, and the technical effects they bring are the same as those in this application. Figures 3 to 11 The corresponding method embodiments are the same. For specific contents, please refer to the description in the method embodiments shown above in this application, which will not be repeated here.

[0281] The present application also provides a vehicle, see Fig.16 , Fig.16 A structural schematic diagram of a vehicle provided in an embodiment of the present application, wherein the vehicle 100 is configured in a fully or partially automatic driving mode, for example, the vehicle 100 can control itself while in the automatic driving mode, and can determine the current state of the vehicle and its surrounding environment through human operation, determine the possible behavior of at least one other vehicle in the surrounding environment, and determine the confidence level corresponding to the possibility of other vehicles performing possible behaviors, and control the vehicle 100 based on the determined information. When the vehicle 100 is in the automatic driving mode, the vehicle 100 can also be set to operate without human interaction.

[0282] The vehicle 100 may include various subsystems, such as a travel system 102, a sensor system 104, a control system 106, one or more peripheral devices 108, and a power source 110, a computer system 112, and a user interface 116. Optionally, the vehicle 100 may include more or fewer subsystems, and each subsystem may include multiple components. In addition, each subsystem and component of the vehicle 100 may be interconnected by wire or wirelessly.

[0283] Travel system 102 may include components that provide powered movement to vehicle 100. In one embodiment, travel system 102 may include engine 118, power source 119, transmission 120, and wheels / tires 121.

[0284] Among them, the engine 118 can be an internal combustion engine, an electric motor, an air compression engine, or a combination of other types of engines, for example, a hybrid engine consisting of a gasoline engine and an electric motor, and a hybrid engine consisting of an internal combustion engine and an air compression engine. The engine 118 converts the energy source 119 into mechanical energy. Examples of the energy source 119 include gasoline, diesel, other petroleum-based fuels, propane, other compressed gas-based fuels, ethanol, solar panels, batteries, and other sources of electricity. The energy source 119 can also provide energy for other systems of the vehicle 100. The transmission 120 can transmit the mechanical power from the engine 118 to the wheels 121. The transmission 120 may include a gearbox, a differential, and a drive shaft. In one embodiment, the transmission 120 may also include other devices, such as a clutch. Among them, the drive shaft may include one or more shafts that can be coupled to one or more wheels 121.

[0285] The sensor system 104 may include several sensors that sense information about the environment around the vehicle 100. For example, the sensor system 104 may include a positioning system 122 (the positioning system may be a global positioning GPS system, or a Beidou system or other positioning systems), an inertial measurement unit (IMU) 124, a radar 126, a laser rangefinder 128, and a camera 130. The sensor system 104 may also include sensors of the internal systems of the monitored vehicle 100 (e.g., an in-vehicle air quality monitor, a fuel gauge, an oil temperature gauge, etc.). The sensing data from one or more of these sensors may be used to detect objects and their corresponding characteristics (position, shape, direction, speed, etc.). Such detection and recognition are key functions for the safe operation of the autonomous vehicle 100.

[0286] Among them, the positioning system 122 can be used to estimate the geographic location of the vehicle 100. The IMU 124 is used to sense the position and orientation changes of the vehicle 100 based on inertial acceleration. In one embodiment, the IMU 124 can be a combination of an accelerometer and a gyroscope. The radar 126 can use radio signals to sense objects in the surrounding environment of the vehicle 100, and can be specifically expressed as a millimeter wave radar or a laser radar. In some embodiments, in addition to sensing objects, the radar 126 can also be used to sense the speed and / or direction of travel of the object. The laser rangefinder 128 can use lasers to sense objects in the environment where the vehicle 100 is located. In some embodiments, the laser rangefinder 128 may include one or more laser sources, a laser scanner, and one or more detectors, as well as other system components. The camera 130 can be used to capture multiple images of the surrounding environment of the vehicle 100. The camera 130 can be a static camera or a video camera.

[0287] The control system 106 controls the operation of the vehicle 100 and its components. The control system 106 may include various components, including a steering system 132 , a throttle 134 , a brake unit 136 , a computer vision system 140 , a lane control system 142 , and an obstacle avoidance system 144 .

[0288] Among them, the steering system 132 can be operated to adjust the forward direction of the vehicle 100. For example, it can be a steering wheel system in one embodiment. The throttle 134 is used to control the operating speed of the engine 118 and thus control the speed of the vehicle 100. The brake unit 136 is used to control the deceleration of the vehicle 100. The brake unit 136 can use friction to slow down the wheel 121. In other embodiments, the brake unit 136 can convert the kinetic energy of the wheel 121 into electric current. The brake unit 136 can also take other forms to slow down the rotation speed of the wheel 121 to control the speed of the vehicle 100. The computer vision system 140 can be operated to process and analyze the images captured by the camera 130 in order to identify objects and / or features in the surrounding environment of the vehicle 100. The objects and / or features may include traffic signals, road boundaries and obstacles. The computer vision system 140 can use object recognition algorithms, Structure from Motion (SFM) algorithms, video tracking and other computer vision technologies. In some embodiments, the computer vision system 140 can be used to map the environment, track objects, estimate the speed of objects, and so on. The route control system 142 is used to determine the route and speed of the vehicle 100. In some embodiments, the route control system 142 may include a lateral planning module 1421 and a longitudinal planning module 1422, which are respectively used to determine the route and speed of the vehicle 100 in combination with data from the obstacle avoidance system 144, GPS122, and one or more predetermined maps. The obstacle avoidance system 144 is used to identify, evaluate, and avoid or otherwise cross obstacles in the environment of the vehicle 100, and the aforementioned obstacles can be specifically manifested as actual obstacles and virtual moving bodies that may collide with the vehicle 100. In one example, the control system 106 may include components other than those shown and described in addition or in an alternative manner. Alternatively, a portion of the components shown above may be reduced.

[0289] The vehicle 100 interacts with external sensors, other vehicles, other computer systems, or users through the peripheral device 108. The peripheral device 108 may include a wireless communication system 146, an onboard computer 148, a microphone 150, and / or a speaker 152. In some embodiments, the peripheral device 108 provides a means for the user of the vehicle 100 to interact with the user interface 116. For example, the onboard computer 148 may provide information to the user of the vehicle 100. The user interface 116 may also operate the onboard computer 148 to receive user input. The onboard computer 148 may be operated through a touch screen. In other cases, the peripheral device 108 may provide a means for the vehicle 100 to communicate with other devices located in the vehicle. For example, the microphone 150 may receive audio (e.g., voice commands or other audio input) from the user of the vehicle 100. Similarly, the speaker 152 may output audio to the user of the vehicle 100. The wireless communication system 146 may communicate wirelessly with one or more devices directly or via a communication network. For example, the wireless communication system 146 may use 3G cellular communication, such as CDMA, EVDO, GSM / GPRS, or 4G cellular communication, such as LTE. Or 5G cellular communication. The wireless communication system 146 may communicate using a wireless local area network (WLAN). In some embodiments, the wireless communication system 146 may communicate directly with the device using an infrared link, Bluetooth, or ZigBee. Other wireless protocols, such as various vehicle communication systems, for example, the wireless communication system 146 may include one or more dedicated short range communications (DSRC) devices, which may include public and / or private data communications between vehicles and / or roadside stations.

[0290] The power source 110 can provide power to various components of the vehicle 100. In one embodiment, the power source 110 can be a rechargeable lithium-ion or lead-acid battery. One or more battery packs of such batteries can be configured as a power source to provide power to various components of the vehicle 100. In some embodiments, the power source 110 and the energy source 119 can be implemented together, such as in some all-electric vehicles.

[0291] Some or all of the functions of the vehicle 100 are controlled by a computer system 112. The computer system 112 may include at least one processor 113 that executes instructions 115 stored in a non-transitory computer-readable medium such as a memory 114. The computer system 112 may also be a plurality of computing devices that control individual components or subsystems of the vehicle 100 in a distributed manner. The processor 113 may be any conventional processor, such as a commercially available central processing unit (CPU). Alternatively, the processor 113 may be a dedicated device such as an application specific integrated circuit (ASIC) or other hardware-based processor. Although Fig.16 The processor, memory, and other components of the computer system 112 in the same block are functionally illustrated, but it will be appreciated by those skilled in the art that the processor, or memory, may actually include multiple processors, or memories that are not stored in the same physical housing. For example, the memory 114 may be a hard drive or other storage medium located in a housing different from the computer system 112. Therefore, references to the processor 113 or memory 114 will be understood to include references to a collection of processors or memories that may or may not operate in parallel. Different from using a single processor to perform the steps described herein, some components such as the steering assembly and the deceleration assembly may each have their own processor that performs only calculations related to the functions specific to the component.

[0292] In various aspects described herein, the processor 113 may be located remotely from the vehicle 100 and in wireless communication with the vehicle 100. In other aspects, some of the processes described herein are performed on a processor 113 disposed within the vehicle 100 while others are performed by the remote processor 113, including taking the necessary steps to perform a single maneuver.

[0293] In some embodiments, the memory 114 may include instructions 115 (e.g., program logic) that can be executed by the processor 113 to perform various functions of the vehicle 100, including those described above. The memory 114 may also include additional instructions, including instructions to send data to, receive data from, interact with, and / or control one or more of the travel system 102, the sensor system 104, the control system 106, and the peripheral devices 108. In addition to the instructions 115, the memory 114 may also store data such as road maps, route information, the vehicle's location, direction, speed, and other such vehicle data, as well as other information. This information can be used by the vehicle 100 and the computer system 112 during the operation of the vehicle 100 in autonomous, semi-autonomous, and / or manual modes. A user interface 116 is used to provide information to or receive information from a user of the vehicle 100. Optionally, the user interface 116 may include one or more input / output devices within the set of peripheral devices 108, such as a wireless communication system 146, an onboard computer 148, a microphone 150, and a speaker 152.

[0294] The computer system 112 may control functions of the vehicle 100 based on input received from various subsystems (e.g., the travel system 102, the sensor system 104, and the control system 106) and from the user interface 116. For example, the computer system 112 may utilize input from the control system 106 in order to control the steering system 132 to avoid obstacles detected by the sensor system 104 and the obstacle avoidance system 144. In some embodiments, the computer system 112 may be operable to provide control over many aspects of the vehicle 100 and its subsystems.

[0295] Alternatively, one or more of the above-mentioned components may be installed or associated separately from the vehicle 100. For example, the memory 114 may exist partially or completely separately from the vehicle 100. The above-mentioned components may be communicatively coupled together in a wired and / or wireless manner.

[0296] Optionally, the above components are only examples. In actual applications, the components in the above modules may be added or deleted according to actual needs. Fig.16 It should not be construed as limiting the embodiments of the present application. A vehicle traveling on a road, such as vehicle 100 above, can identify objects in its surrounding environment to determine an adjustment to the current speed. The object can be another vehicle, a traffic control device, or another type of object. In some examples, each identified object can be considered independently, and based on the respective characteristics of the object, such as its current speed, acceleration, spacing from the vehicle, etc., it can be used to determine the speed of the vehicle to be adjusted.

[0297] Optionally, the vehicle 100 or a computing device associated with the vehicle 100 may be Fig.16 The computer system 112, computer vision system 140, and memory 114 can predict the behavior of the identified objects based on the characteristics of the identified objects and the state of the surrounding environment (e.g., traffic, rain, ice on the road, etc.). Optionally, each of the identified objects depends on the behavior of each other, so all the identified objects can also be considered together to predict the behavior of a single identified object. The vehicle 100 can adjust its speed based on the predicted behavior of the identified objects. In other words, the vehicle 100 can determine what stable state the vehicle will need to adjust to (e.g., accelerate, decelerate, or stop) based on the predicted behavior of the object. In this process, other factors can also be considered to determine the speed of the vehicle 100, such as the lateral position of the vehicle 100 in the road it is traveling on, the curvature of the road, the proximity of static and dynamic objects, etc. In addition to providing instructions to adjust the speed of the vehicle, the computing device can also provide instructions to modify the steering angle of the vehicle 100 so that the vehicle 100 follows a given trajectory and / or maintains a safe lateral and longitudinal distance from objects near the vehicle 100 (e.g., cars in adjacent lanes on the road).

[0298] The vehicle 100 may be a car, a truck, a motorcycle, a bus, a ship, an airplane, a helicopter, a lawn mower, an amusement vehicle, an amusement park vehicle, construction equipment, a tram, a golf cart, a train, etc., and the embodiments of the present application do not make any particular limitation.

[0299] In the embodiment of the present application, the processor 113 in the vehicle 10 is used to execute Figures 3 to 11 The method executed by the first device and / or the execution device in the corresponding embodiment. It should be noted that the specific manner in which the processor 113 executes the above steps is the same as that in the present application. Figures 3 to 11 The corresponding method embodiments are based on the same concept, and the technical effects they bring are the same as those in this application. Figures 3 to 11 The corresponding method embodiments are the same. For specific contents, please refer to the description in the method embodiments shown above in this application, which will not be repeated here.

[0300] The present application also provides a computer-readable storage medium in which a program is stored. When the program is run on a computer, the computer executes the above-mentioned Figures 3 to 11 The steps performed by the first device and / or the execution device in the method described in the embodiment shown, or the computer is caused to perform the above Figures 3 to 11 The illustrated embodiment describes the steps performed by the training device in the method.

[0301] The present application also provides a computer program product, which includes a program, which, when executed on a computer, enables the computer to execute the above-mentioned Figures 3 to 11The steps performed by the first device and / or the execution device in the method described in the embodiment shown, or the computer is caused to perform the above Figures 3 to 11 The illustrated embodiment describes the steps performed by the training device in the method.

[0302] The present application also provides a circuit system, which includes a processing circuit, wherein the processing circuit is configured to perform the above Figures 3 to 11 The steps performed by the first device and / or the execution device in the method described in the embodiment shown, or the processing circuit is configured to perform the steps as described above Figures 3 to 11 The illustrated embodiment describes the steps performed by the training device in the method.

[0303] The information processing device, equipment or vehicle provided in the embodiment of the present application may be a chip, which includes: a processing unit and a communication unit. The processing unit may be, for example, a processor, and the communication unit may be, for example, an input / output interface, a pin or a circuit. The processing unit may execute the computer execution instructions stored in the storage unit to enable the chip in the server to execute the above Figures 3 to 11 Optionally, the storage unit is a storage unit in the chip, such as a register, a cache, etc. The storage unit may also be a storage unit located outside the chip in the wireless access device, such as a read-only memory (ROM) or other types of static storage devices that can store static information and instructions, a random access memory (RAM), etc.

[0304] For details, please refer to Fig.17 , Fig.17 A schematic diagram of the structure of a chip provided in an embodiment of the present application, the chip can be expressed as a neural network processor NPU 170, NPU 170 is mounted on the host CPU (Host CPU) as a coprocessor, and the host CPU assigns tasks. The core part of the NPU is the operation circuit 1703, which is controlled by the controller 1704 to extract matrix data in the memory and perform multiplication operations.

[0305] In some implementations, the operation circuit 1703 includes multiple processing units (Process Engine, PE) inside. In some implementations, the operation circuit 1703 is a two-dimensional systolic array. The operation circuit 1703 can also be a one-dimensional systolic array or other electronic circuits capable of performing mathematical operations such as multiplication and addition. In some implementations, the operation circuit 1703 is a general-purpose matrix processor.

[0306] For example, assume there is an input matrix A, a weight matrix B, and an output matrix C. The operation circuit takes the corresponding data of matrix B from the weight memory 1702 and caches it on each PE in the operation circuit. The operation circuit takes the matrix A data from the input memory 1701 and performs matrix operation with matrix B, and the partial result or final result of the matrix is ​​stored in the accumulator 1708.

[0307] The unified memory 1706 is used to store input data and output data. The weight data is directly transferred to the weight memory 1702 through the direct memory access controller (DMAC) 1705. The input data is also transferred to the unified memory 1706 through the DMAC.

[0308] BIU stands for Bus Interface Unit 1710 , which is used for the interaction between the AXI bus and the DMAC and the instruction fetch buffer (IFB) 1709 .

[0309] The bus interface unit 1710 (Bus Interface Unit, BIU for short) is used for the instruction fetch memory 1709 to obtain instructions from the external memory, and is also used for the storage unit access controller 1705 to obtain the original data of the input matrix A or the weight matrix B from the external memory.

[0310] DMAC is mainly used to transfer input data in the external memory DDR to the unified memory 1706 or to transfer weight data to the weight memory 1702 or to transfer input data to the input memory 1701.

[0311] The vector calculation unit 1707 includes multiple operation processing units, and further processes the output of the operation circuit when necessary, such as vector multiplication, vector addition, exponential operation, logarithmic operation, size comparison, etc. It is mainly used for non-convolutional / fully connected layer network calculations in neural networks, such as Batch Normalization, pixel-level summation, upsampling of feature planes, etc.

[0312] In some implementations, the vector calculation unit 1707 can store the processed output vector to the unified memory 1706. For example, the vector calculation unit 1707 can apply a linear function and / or a nonlinear function to the output of the operation circuit 1703, such as linear interpolation of the feature plane extracted by the convolution layer, and then, for example, a vector of accumulated values ​​to generate an activation value. In some implementations, the vector calculation unit 1707 generates a normalized value, a pixel-level summed value, or both. In some implementations, the processed output vector can be used as an activation input to the operation circuit 1703, for example, for use in a subsequent layer in a neural network.

[0313] An instruction fetch buffer 1709 connected to the controller 1704 is used to store instructions used by the controller 1704;

[0314] Unified memory 1706, input memory 1701, weight memory 1702 and instruction fetch memory 1709 are all on-chip memories. External memories are private to the NPU hardware architecture.

[0315] in, Figures 3 to 11 The operations of each layer in the deep learning model, the first feature extraction network, and the second feature extraction network shown can be performed by the operation circuit 1703 or the vector calculation unit 1707.

[0316] The processor mentioned in any of the above places may be a general-purpose central processing unit, a microprocessor, an ASIC, or one or more integrated circuits for controlling the execution of the program of the above-mentioned first aspect method.

[0317] It should also be noted that the device embodiments described above are merely schematic, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed over multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. In addition, in the drawings of the device embodiments provided by the present application, the connection relationship between the modules indicates that there is a communication connection between them, which may be specifically implemented as one or more communication buses or signal lines.

[0318] Through the description of the above implementation mode, the technicians in the field can clearly understand that the present application can be implemented by means of software plus necessary general hardware, and of course, it can also be implemented by special hardware including special integrated circuits, special CPUs, special memories, special components, etc. In general, all functions completed by computer programs can be easily implemented by corresponding hardware, and the specific hardware structure used to implement the same function can also be various, such as analog circuits, digital circuits or special circuits. However, for the present application, software program implementation is a better implementation mode in more cases. Based on such an understanding, the technical solution of the present application is essentially or the part that contributes to the prior art can be embodied in the form of a software product, which is stored in a readable storage medium, such as a computer floppy disk, a U disk, a mobile hard disk, a ROM, a RAM, a disk or an optical disk, etc., including a number of instructions to enable a computer device (which can be a personal computer, a training device, or a network device, etc.) to execute the methods described in each embodiment of the present application.

[0319] In the above embodiments, all or part of the embodiments may be implemented by software, hardware, firmware or any combination thereof. When implemented by software, all or part of the embodiments may be implemented in the form of a computer program product.

[0320] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function described in the embodiment of the present application is generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions may be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions may be transmitted from a website site, a computer, a training device, or a data center by wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) mode to another website site, computer, training device, or data center. The computer-readable storage medium may be any available medium that a computer can store or a data storage device such as a training device, a data center, etc. that includes one or more available media integrations. The available medium may be a magnetic medium, (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state drive (SSD)), etc.

Claims

1. An information processing method, characterized in that, the method includes: inputting first information corresponding to the traffic scene around the host vehicle into the deep learning model, where the first information includes first text information; obtaining second information corresponding to the first information, where the second information is obtained based on the deep learning model, and the second information corresponds to any one of the following tasks: making a decision on the behavior of the host vehicle, performing trajectory planning on the host vehicle, or controlling the host vehicle.

2. The method according to claim 1, characterized in that, the first text information includes information of a preset text corresponding to the task, or the first text information includes information of text input by the user.

3. The method according to claim 1 or 2, characterized in that, the method further includes: inputting a first question corresponding to the traffic scene into the deep learning model, where the first question is used to obtain a first answer corresponding to the first question, and the first answer is obtained through the deep learning model. Wherein, the first information is a second question, the second information is a second answer corresponding to the second question, and the information in the first answer exists in the first text information.

4. The method according to claim 1 or 2, characterized in that, the first information further includes first feature information, where the first feature information includes feature information of the environmental information around the host vehicle, and the environmental information includes physical attribute information of the objects around the host vehicle.

5. The method according to claim 4, characterized in that, the first feature information is obtained based on a feature extraction network, where the feature extraction network belongs to a first neural network, and the first neural network is used to perform at least two of the following tasks: predicting the trajectory of the objects around the host vehicle, making a decision on the behavior of the host vehicle, performing trajectory planning on the host vehicle, predicting the speed range of the host vehicle, or controlling the host vehicle.

6. The method according to claim 1 or 2, characterized in that, the first information is a second question, the second information is a second answer corresponding to the second question, where the second answer is included in a preset answer set corresponding to the second question, and the second answer is obtained from at least one alternative answer generated by the deep learning model based on the preset answer set.

7. The method according to claim 1 or 2, characterized in that, the first information is a second question, the second information is a second answer corresponding to the second question, the second answer is obtained through a classification network, and the second feature information includes the feature information of the starting flag [CLS] bit in the feature information of the second answer. The classification network is used to determine the second answer corresponding to the second feature information from at least one preset answer.

8. The method according to claim 1 or 2, characterized in that, the first information is a second question, the second information is a second answer corresponding to the second question, and inputting the first information corresponding to the traffic scene around the host vehicle into the deep learning model includes: Input multiple second questions into the deep learning model, where different second questions among the multiple second questions carry different hints; The obtaining the second information corresponding to the first information includes: Obtain multiple reference answers corresponding one by one to the multiple second questions, and all the multiple reference answers are obtained through the deep learning model; Determine the second answer according to the multiple reference answers.

9. An information processing method, characterized in that, the method includes: Input a first question into a deep learning model, and obtain a first answer corresponding to the first question; Determine a second question according to the first answer, and the second question carries the information in the first answer; Input the second question into the deep learning model, and obtain a second answer corresponding to the second question, where the first question and the second question correspond to the same traffic scenario.

10. The method according to claim 9, characterized in that, The second answer corresponds to any one of the following tasks: identifying risk obstacles around the host vehicle, identifying the behaviors of objects around the host vehicle, predicting the behaviors of objects around the host vehicle, predicting the trajectories of objects around the host vehicle, making decisions on the behaviors of the host vehicle, planning the trajectory of the host vehicle, or controlling the host vehicle.

11. An information processing method, characterized in that, The method is used to train a deep learning model, the deep learning model includes at least two training stages, the at least two training stages include a first training stage and a second training stage, and the method includes: In the first training stage, input a third question into the deep learning model, and obtain a predicted answer corresponding to the third question; In the second training stage, input a fourth question into the deep learning model, and obtain a predicted answer corresponding to the fourth question; Wherein, when training the deep learning model, a first loss function is adopted. In the first training stage, the first loss function indicates the similarity between the predicted answer corresponding to the third question and the expected answer. In the second training stage, the first loss function indicates the similarity between the predicted answer corresponding to the fourth question and the expected answer. The third question and the fourth question are both related to the traffic scenario, and the third question and the fourth question are used to obtain information at different levels.

12. The method according to claim 11, characterized in that, The third question includes first feature information, and the first feature information includes feature information of the environmental information around the host vehicle, and the environmental information includes physical attribute information of the objects around the host vehicle.

13. The method according to claim 12, characterized in that, Before inputting the third question into the deep learning model, the method further includes: Input the environmental information into a feature extraction network, and obtain second feature information generated by the feature extraction network, and the second feature information is used to obtain the first feature information; Input the first feature information into a feature processing network to obtain prediction information generated by the feature processing network. The feature extraction network and the feature processing network belong to the same neural network, and the neural network is used to perform at least multiple of the following tasks: predicting the trajectories of objects around the host vehicle, making decisions on the behavior of the host vehicle, planning the trajectory of the host vehicle, predicting the speed range of the host vehicle, or controlling the host vehicle; Wherein, the training phase of the deep learning model includes training the deep learning model and the neural network using the first loss function and the second loss function, and the second loss function indicates the similarity between the prediction information corresponding to the environment information and the second expected information.

14. An information processing device, characterized in that, the device includes: an input module, configured to input first information corresponding to the traffic scenario around the host vehicle into the deep learning model, where the first information includes first text information; an acquisition module, configured to acquire second information corresponding to the first information, where the second information is obtained based on the deep learning model, and the second information corresponds to any one of the following tasks: making decisions on the behavior of the host vehicle, planning the trajectory of the host vehicle, or controlling the host vehicle.

15. The device according to claim 14, characterized in that, the first text information includes information of a preset text corresponding to the task, or the first text information includes information of a text input by a user.

16. The device according to claim 14 or 15, characterized in that, the input module is further configured to input a first question corresponding to the traffic scenario into the deep learning model, where the first question is used to obtain a first answer corresponding to the first question, and the first answer is obtained through the deep learning model. Wherein, the first information is a second question, the second information is a second answer corresponding to the second question, and the information in the first answer exists in the first text information.

17. The device according to claim 14 or 15, characterized in that, the first information further includes first feature information, where the first feature information includes feature information of the environment information around the host vehicle, and the environment information includes physical attribute information of an object around the host vehicle.

18. The device according to claim 17, characterized in that, the first feature information is obtained based on a feature extraction network, where the feature extraction network belongs to a first neural network, and the first neural network is used to perform at least two of the following tasks: predicting the trajectories of objects around the host vehicle, making decisions on the behavior of the host vehicle, planning the trajectory of the host vehicle, predicting the speed range of the host vehicle, or controlling the host vehicle.

19. The device according to claim 14 or 15, characterized in that, The first information is a second question, and the second information is a second answer corresponding to the second question. Among them, the second answer is included in a preset answer set corresponding to the second question, and the second answer is obtained from at least one alternative answer generated by the deep learning model based on the preset answer set.

20. The device according to claim 14 or 15, wherein, the first information is a second question, the second information is a second answer corresponding to the second question, the second answer is obtained through a classification network, the second feature information includes the feature information of the starting flag [CLS] bit in the feature information of the second answer, and the classification network is used to determine the second answer corresponding to the second feature information from at least one preset answer.

21. The device according to claim 14 or 15, wherein, the first information is a second question, and the second information is a second answer corresponding to the second question; The input module is specifically configured to input a plurality of the second questions to the deep learning model, and different second questions among the plurality of second questions carry different prompts; The obtaining module is specifically configured to obtain a plurality of reference answers corresponding to the plurality of second questions one by one, and determine the second answer according to the plurality of reference answers, and the plurality of reference answers are all obtained through the deep learning model.

22. An information processing device, wherein, the device includes: An input module, configured to input a first question to a deep learning model and obtain a first answer corresponding to the first question; A determination module, configured to determine a second question according to the first answer, and the second question carries information in the first answer; The input module is further configured to input the second question to the deep learning model and obtain a second answer corresponding to the second question, wherein the first question and the second question correspond to the same traffic scenario.

23. The device according to claim 22, wherein, The second answer corresponds to any one of the following tasks: identifying risk obstacles around the vehicle, identifying the behaviors of objects around the vehicle, predicting the behaviors of objects around the vehicle, predicting the trajectories of objects around the vehicle, making decisions on the behaviors of the vehicle, planning the trajectory of the vehicle, or controlling the vehicle.

24. An information processing device, wherein, the device is used to train a deep learning model, the deep learning model includes at least two training stages, the at least two training stages include a first training stage and a second training stage, and the device includes: An input module, configured to input a third question to the deep learning model and obtain a predicted answer corresponding to the third question in the first training stage; The input module is further configured to input a fourth question to the deep learning model and obtain a predicted answer corresponding to the fourth question in the second training stage; Among them, when training the deep learning model, a first loss function is adopted. In the first training stage, the first loss function indicates the similarity between the predicted answer corresponding to the third question and the expected answer. In the second training stage, the first loss function indicates the similarity between the predicted answer corresponding to the fourth question and the expected answer. Both the third question and the fourth question are related to the traffic scene, and the third question and the fourth question are used to obtain information at different levels.

25. The apparatus according to claim 24, wherein, the third question includes first feature information, and the first feature information includes feature information of the environment information around the vehicle itself. The environment information includes physical attribute information of the objects around the vehicle itself.

26. The apparatus according to claim 25, wherein, the apparatus further includes: a feature extraction module, configured to input the environment information into a feature extraction network to obtain second feature information generated by the feature extraction network, and the second feature information is used to obtain the first feature information; a feature processing module, configured to input the first feature information into a feature processing network to obtain prediction information generated by the feature processing network. The feature extraction network and the feature processing network belong to the same neural network, and the neural network is used to perform at least multiple of the following tasks: predicting the trajectory of the objects around the vehicle itself, making decisions on the behavior of the vehicle itself, planning the trajectory of the vehicle itself, predicting the speed range of the vehicle itself, or controlling the vehicle itself; Among them, the training stage of the deep learning model includes training the deep learning model and the neural network by using the first loss function and the second loss function. The second loss function indicates the similarity between the prediction information corresponding to the environment information and the second expected information.

27. A device, wherein, it includes a processor, the processor is coupled with a memory, and the memory stores program instructions. When the program instructions stored in the memory are executed by the processor, the method according to any one of claims 1 to 13 is implemented.

28. A vehicle, wherein, it includes a processor, the processor is coupled with a memory, and the memory stores program instructions. When the program instructions stored in the memory are executed by the processor, the method according to any one of claims 1 to 10 is implemented.

29. A computer-readable storage medium, including a program, when it runs on a computer, enables the computer to execute the method according to any one of claims 1 to 13.

30. A circuit system, wherein, the circuit system includes a processing circuit, and the processing circuit is configured to execute the method according to any one of claims 1 to 13.

Citation Information

Cited By

  • Information processing method and related device

    EP4796407A1

  • Information processing method and related device

    WO2025107911A1