Request processing method and device of intelligent cabin and vehicle

Through the comprehensive prompt word-driven target request processing model, the inference thinking chain is used to disassemble complex tasks, and the accuracy and efficiency of the intelligent cockpit request processing method is solved, and more efficient interactive response is achieved.

CN120409685APending Publication Date: 2025-08-01GUANGZHOU XIAOPENG MOTORS TECH CO LTD

Patent Information

Application Number
CN202510513938.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-22
Publication Date
2025-08-01

AI Technical Summary

Technical Problem

The request processing method of smart cockpits in the prior art is poorly accurate and inefficient, and it is impossible to effectively understand user interaction requests in an overall manner.

Method used

The comprehensive prompt word-driven target request processing model is adopted. Through the task disassembly of the inference thinking chain, complex tasks are disassembled into inference tasks with multiple dependencies, and the target request processing model is used for processing to generate interactive responses.

Benefits of technology

It improves the accuracy and efficiency of request processing, can quickly generate accurate interactive responses, and solves the problems of poor accuracy and low efficiency of request processing methods for complex tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120409685A_ABST
    Figure CN120409685A_ABST
Patent Text Reader

Abstract

The invention discloses a request processing method and device of an intelligent cabin and a vehicle. The method comprises the steps that request data are acquired, and the request data are used for representing a user interaction request received by an intelligent cabin of a vehicle; the target request processing model is driven by the comprehensive cue word to perform disassembly reasoning on the request data to obtain a reasoning result, the comprehensive cue word is used for determining a reasoning standard, and the target request processing model is used for constructing and processing a plurality of reasoning tasks according to a task disassembly mode of a reasoning thinking chain. The reasoning result is used for determining response data corresponding to the user interaction request; and generating an interaction response according to a reasoning result. According to the method and the device, the technical problems of poor accuracy and low efficiency of a request processing method for a complex task in related technologies are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of data processing and intelligent driving, and in particular, to a method, device and vehicle for processing requests in an intelligent cockpit. Background Art

[0002] With the continuous increase of cockpit scenarios, the complexity of intelligent cockpit interaction requests to be processed has increased significantly. In related technologies, a method of processing interaction requests using multiple models and their corresponding prompt words respectively is usually adopted. However, on the one hand, this method cannot enable the model to comprehensively understand the interaction requests, resulting in poor accuracy of the request processing method. On the other hand, this method requires relying on multiple prompt words and multiple models to process the interaction requests, resulting in low efficiency of the request processing method.

[0003] In view of the above problems, no effective solution has been proposed yet. Summary of the Invention

[0004] Embodiments of the present invention provide a method, device and vehicle for processing requests in an intelligent cockpit, so as to at least solve the technical problems of poor accuracy and low efficiency of the request processing method for complex tasks in related technologies.

[0005] According to one aspect of the embodiments of the present invention, a method for processing requests in an intelligent cockpit is provided, including: obtaining request data, where the request data is used to represent a user interaction request received by the intelligent cockpit of the vehicle; driving a target request processing model to disassemble and reason about the request data using a comprehensive prompt word to obtain a reasoning result, where the comprehensive prompt word is used to determine a reasoning criterion, the target request processing model is used to construct and process multiple reasoning tasks according to the task disassembling method of the reasoning thought chain, and the reasoning result is used to determine response data corresponding to the user interaction request; generating an interaction response according to the reasoning result.

[0006] Optionally, the task disassembling method of the reasoning thought chain is used to represent the dependency relationship between multiple reasoning tasks. Driving a target request processing model to disassemble and reason about the request data using a comprehensive prompt word to obtain a reasoning result includes: driving a target request processing model to disassemble the request data using a comprehensive prompt word to construct multiple reasoning tasks, and performing reasoning processing on the multiple reasoning tasks to obtain a reasoning result.

[0007] Optionally, the comprehensive prompt includes a list of task types. The target request processing model is driven by the comprehensive prompt to perform inference processing on multiple inference tasks, and the obtained inference results include: driving the target request processing model by the comprehensive prompt to perform inference calculation processing on multiple inference tasks to obtain the calculation processing results corresponding to the multiple inference tasks respectively, selecting the target task type corresponding to the request data from the list of task types according to the multiple calculation processing results, and performing inference transformation processing on the request data based on the target task type to obtain the inference results, where the target task type is used to determine the processing strategy corresponding to the request data.

[0008] Optionally, the target task type includes a first task type, and the first task type is used to represent the task type of responding to the request data by controlling in-vehicle devices or functions. Performing inference transformation processing on the request data based on the target task type, the obtained inference results include: in response to the target task type being the first task type, performing inference transformation on the request data based on the comprehensive prompt to obtain a transformation result; obtaining the inference results according to the comprehensive prompt and the transformation result.

[0009] Optionally, obtaining the inference results according to the comprehensive prompt and the transformation result includes: performing position recognition on the transformation result based on the comprehensive prompt to obtain a position recognition result; using the position recognition result and the position mapping relationship to determine the control object position information, where the position mapping relationship is used to represent the corresponding relationship between the position recognition result and the control object position information; obtaining the inference results according to the control object position information and the transformation result.

[0010] Optionally, the target request processing model is obtained by training an initial request processing model with multiple groups of training data. The multiple groups of training data are constructed based on the task types corresponding to multiple inference tasks respectively. The training data includes original data, the reply results corresponding to the original data, and the thought chains corresponding to the original data.

[0011] Optionally, the request data further includes historical conversation data, and the above request processing method for the intelligent cockpit further includes: determining whether there is information to be supplemented in the request data based on the historical conversation data; in response to the existence of information to be supplemented in the historical conversation data, updating the request data with the information to be supplemented.

[0012] Optionally, the interaction response includes an interaction reply result and a control instruction. Generating the interaction response according to the inference results includes: analyzing the inference results to generate an interaction reply result and a control instruction; sending the interaction reply result to the in-vehicle multimedia system corresponding to the intelligent cockpit; sending the control instruction to the vehicle control system to control the vehicle to perform the target action.

[0013] According to another aspect of the embodiments of the present invention, there is also provided a processing device for cockpit instructions, including: an acquisition module, configured to acquire request data, where the request data is used to represent a user interaction request received by the intelligent cockpit of the vehicle; a processing module, configured to drive a target request processing model with a comprehensive prompt word, disassemble and reason about the request data to obtain a reasoning result, where the comprehensive prompt word is used to determine a reasoning criterion, the target request processing model is used to construct and process multiple reasoning tasks according to the task disassembling method of the reasoning thinking chain, and the reasoning result is used to determine response data corresponding to the user interaction request; and an execution module, configured to generate an interaction response according to the reasoning result.

[0014] According to another aspect of the embodiments of the present invention, there is also provided a vehicle, including: a memory storing an executable program; a processor configured to run the program, where when the program runs, it executes the request processing method of the intelligent cockpit in any one of the above.

[0015] In the embodiments of the present invention, request data is acquired, where the request data is used to represent a user interaction request received by the intelligent cockpit of the vehicle; the target request processing model is driven with a comprehensive prompt word to disassemble and reason about the request data to obtain a reasoning result, where the comprehensive prompt word is used to determine a reasoning criterion, the target request processing model is used to construct and process multiple reasoning tasks according to the task disassembling method of the reasoning thinking chain, and the reasoning result is used to determine response data corresponding to the user interaction request; and an interaction response is generated according to the reasoning result.

[0016] It is easy to note that in the embodiments of the present application, request data is acquired through a user interaction request received by the intelligent cockpit of the vehicle; and then the target request processing model is driven with a comprehensive prompt word, and the request data is disassembled and reasoned according to the task disassembling method of the reasoning thinking chain, that is, by combining the comprehensive prompt word and the task disassembling method of the reasoning thinking chain, the request data containing complex tasks is decomposed into multiple reasoning tasks with dependency relationships, which can more accurately process complex tasks, thereby improving the accuracy of the reasoning result and the efficiency of obtaining the reasoning result; further, according to the reasoning result, a more accurate interaction response can be quickly generated, achieving the purpose of quickly and accurately responding to request data containing complex tasks, thereby realizing the technical effect of improving the accuracy and efficiency of processing request data containing complex tasks, and further solving the technical problems of poor accuracy and low efficiency of the request processing method for complex tasks in the related art. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] The drawings described herein are used to provide a further understanding of the present invention, and constitute a part of this application. The illustrative embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation to the present invention. In the drawings:

[0018] Figure 1 is a hardware block diagram of a computing terminal for an optional method of handling requests in an intelligent cockpit according to an embodiment of the present application;

[0019] Figure 2 is a flowchart of a method for handling requests in an intelligent cockpit according to an embodiment of the present application;

[0020] Figure 3 is a schematic diagram of an optional directed acyclic graph according to an embodiment of the present application;

[0021] Figure 4 is a block diagram of a device for handling requests in an intelligent cockpit according to an embodiment of the present application. Detailed implementation manners

[0022] In order to enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0023] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above drawings are used to distinguish similar objects, and do not necessarily need to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present invention described herein can be implemented in an order different from those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.

[0024] According to an embodiment of the present invention, an embodiment of a method for handling requests in an intelligent cockpit is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.

[0025] First, the operating environment of the above method embodiment is exemplarily described. Figure 1 is a hardware block diagram of a computing terminal for an optional method of handling requests in an intelligent cockpit according to an embodiment of the present application, as Figure 1As shown, the computing terminal 10 (e.g., a computer terminal, a mobile intelligent terminal, a vehicle terminal, or a cloud computing virtual terminal, etc.) may include: one or more processors 102 (e.g., may include processors 102a, 102b, ……, 102n), a memory 104 for storing data, and a transmission device 106 for implementing communication functions. Among them, the processor 102 may include, but is not limited to, processing components such as a microcontroller unit (MCU) or a field programmable gate array (FPGA).

[0026] The above computing terminal 10 may further include: a display device 110, an input / output interface 108, a universal serial bus (USB) port (this USB port may be one of the ports of the computer bus, not shown in the figure), a network interface (not shown in the figure), a power supply (not shown in the figure), and a camera (not shown in the figure).

[0027] It should be noted that one or more processors 102 and / or other data processing circuits in the above computing terminal 10 may be embodied in whole or in part as software, hardware, firmware, or any other combination. In addition, the data processing circuit may be a single independent processing module, or may be wholly or partially incorporated into any one of the other elements in the computing terminal 10 (or mobile device).

[0028] The memory 104 can be used to store software programs and modules of application software, such as the program instructions and data storage devices corresponding to the request processing method of the intelligent cockpit in the embodiments of the present application. The processor 102 executes various functional applications and data processing by running the software programs and modules stored in the memory 104, that is, implements the above-mentioned request processing method of the intelligent cockpit. The memory 104 may include high-speed random access memory, and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memories. In some instances, the memory 104 may further include a memory remotely located relative to the processor 102, and these remote memories may be connected to the vehicle terminal through a network. Examples of the above network include, but are not limited to, the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.

[0029] The transmission device 106 is used to receive or send data via a network. Specific examples of the above network may include a wireless network provided by a communication provider of a vehicle terminal. In one example, the transmission device 106 includes a network adapter (Network Interface Controller, abbreviated as NIC) and a network interface. The network adapter can be connected to other network devices through a base station to communicate with the Internet. The transmission device 106 can perform data communication in a wired and / or wireless network connection manner. In one example, the transmission device 106 can be a Radio Frequency (RF) module, which is used to communicate with the Internet wirelessly.

[0030] The input / output interface can be connected to the input / output devices corresponding to the computing terminal 10 to implement input / output functions. The input / output devices may include, but are not limited to: cursor control devices, keyboards, displays, etc. The above input / output devices can be built into the computing terminal 10 or external external devices of the computing terminal 10.

[0031] Those of ordinary skill in the art can understand that Figure 1 The structure of the illustrated computing terminal 10 is only schematic and does not strictly limit the structure of the above computing terminal 10. For example, the computing terminal 10 may further include more or fewer components than those shown in Figure 1 or the computing terminal 10 may have different categories of components from those shown in Figure 1 the illustration.

[0032] An embodiment of the present application provides a method for processing requests in an intelligent cockpit. The method for processing requests in an intelligent cockpit can be used to provide a vehicle-human interaction function for a preset application scenario. The above preset application scenario may include the following scenarios in the vehicle field.

[0033] For example, in the commuting autonomous driving scenario, the driver requests to control vehicle navigation, turn on the entertainment system, set the vehicle interior environment, etc. through voice interaction. Using the technical solution of the embodiment of the present application, these voice commands can be accurately and quickly processed, thereby improving the driver's driving experience.

[0034] For example, in the Artificial Intelligence (AI) chauffeur scenario of a household vehicle, using the technical solution of the embodiment of the present application, interaction requests related to the chauffeur service, such as adjusting the vehicle interior atmosphere, etc., can be more accurately understood and executed, improving the quality of the chauffeur service.

[0035] For example, in the Navigation Guided Pilot (NGP) scenario in urban areas or highway areas, by using the technical solution of the embodiment of the present application, it is possible to more effectively process navigation-related interaction requests, such as "planning a route to avoid congestion", and provide more intelligent navigation support for the driver. In addition, the above-mentioned preset application scenarios may also include, but are not limited to: the human-vehicle interaction scenario of intelligent driving trucks or driverless trucks in the field of logistics transportation, the human-vehicle interaction scenario of autonomous agricultural vehicles in the field of agricultural machinery, the human-machine interaction scenario of unmanned aerial vehicles, and the human-machine interaction scenario of intelligent robots (such as cleaning robots, service robots, delivery robots, etc.).

[0036] When the above-mentioned preset application scenario is a scenario in other fields except the vehicle field, those skilled in the art should be able to understand that the vehicle in the request processing method of the above-mentioned intelligent cockpit can be replaced by other objects (such as agricultural machinery, unmanned aerial vehicles, robots, etc.). On this basis, in the embodiment of the present application, the household car field is taken as an example to exemplarily illustrate the specific implementation manner of the request processing method of the above-mentioned intelligent cockpit.

[0037] In the above operating environment, the embodiment of the present application provides a request processing method for an intelligent cockpit as follows Figure 2 shown, Figure 2 is a flowchart of a request processing method for an intelligent cockpit according to an embodiment of the present application, as Figure 2 shown, and the method includes the following implementation steps S21 to step S23:

[0038] Step S21, obtain request data, where the request data is used to represent a user interaction request received by the intelligent cockpit of the vehicle;

[0039] Step S22, use a comprehensive prompt word to drive a target request processing model to disassemble and reason about the request data to obtain a reasoning result, where the comprehensive prompt word is used to determine the reasoning standard, the target request processing model is used to construct and process multiple reasoning tasks according to the task disassembling method of the reasoning thinking chain, and the reasoning result is used to determine the response data corresponding to the user interaction request;

[0040] Step S23, generate an interaction response according to the reasoning result.

[0041] The above user interaction request can be used to represent the original input information in the form of natural language of the user. The original input information may include, but is not limited to: voice input request information, text input request information. The above voice input request information can be obtained through the in-vehicle microphone array. The above text input request information can be obtained through the in-vehicle touch screen. The user interaction request can be a request containing complex tasks. When a request contains at least one of the following tasks, it can be considered that the request is a request containing complex tasks: rejection recognition task, multi-round understanding task, multi-instruction understanding task, location reasoning task, intelligent understanding task combined with perception points.

[0042] The above request data can be data in the standard structured form corresponding to the user interaction request (for example, a form that distinguishes the user's location from the user's needs). In particular, when the above user interaction request is voice input request information, the in-vehicle voice system of the intelligent cockpit listens to the conversation of the user in the vehicle, receives the voice input request information in the natural form (that is, the user interaction request), and further converts the voice input request information into data in the standard structured form (that is, the request data).

[0043] It is easy to understand that in the embodiment of the present application, the user interaction request in the form of natural language is converted into request data in the standard structured form, and the request data helps the target request processing model to understand the request data, so that the subsequent target request processing model can more accurately disassemble and reason about the request data, improving the accuracy of the reasoning result.

[0044] The above comprehensive prompt word can be used to represent the logical framework of the disassembling and reasoning process of the target request processing model. The comprehensive prompt word can include judgment conditions and processing strategies. For a certain request data obtained, using the comprehensive prompt word can determine the reasoning standard for the target request processing model to disassemble and reason about the request data. The comprehensive prompt word can be constructed according to actual needs and stored in the corresponding storage area of the in-vehicle request processing system. Using the comprehensive prompt word can enhance the ability of the in-vehicle request processing system of the intelligent cockpit to process complex tasks.

[0045] The task disassembling methods of the above reasoning thinking chain may include, but are not limited to: sequential disassembling method, parallel disassembling method, conditional disassembling method, dialogue turn disassembling method. By using the task disassembling methods of the reasoning thinking chain, the target request processing model can convert a request containing complex tasks (that is, request data) into multiple interrelated reasoning tasks that are easy for the model to understand and process, execute each reasoning task in turn, determine the results corresponding to each reasoning task, and integrate the results corresponding to the multiple reasoning tasks respectively to respond to the request data, so as to ensure that the request data can be accurately responded to.

[0046] The above response data may refer to the feedback information generated after the vehicle-mounted request processing system parses the user interaction request. The response data may include, but is not limited to: interaction reply results, control instructions, error prompt messages, personalized recommendation information.

[0047] It is easy to understand that in the embodiments of the present application, by combining the comprehensive prompt word and the task decomposition method of the reasoning thought chain, the request data is decomposed into multiple independent but interrelated reasoning tasks, so that the model can more accurately respond to the request data, thereby improving the accuracy of the reasoning result. In addition, in the embodiments of the present application, during the process of processing the request data, only the comprehensive prompt word can be used to drive the target request processing model. Compared with the related art that relies on multiple prompt words and multiple models to process the interaction request, the embodiments of the present application improve the efficiency of human-computer interaction.

[0048] The above interaction response may include, but is not limited to: generating instructions, feedback results, controlling the vehicle to execute actions, and issuing error warnings. According to the reasoning result output by the target request processing model, the vehicle-mounted request processing system can choose to feedback the interaction reply result to the user and / or control the vehicle to execute the target action to complete the interaction response and enhance the experience of human-vehicle interaction.

[0049] It is easy to notice that through the above steps S21 to S23, in the embodiments of the present application, the request data is obtained through the user interaction request received by the intelligent cockpit of the vehicle; then, the comprehensive prompt word is used to drive the target request processing model, and the request data is disassembled and inferred according to the task decomposition method of the reasoning thought chain. That is, by combining the comprehensive prompt word and the task decomposition method of the reasoning thought chain, the request data containing complex tasks is decomposed into multiple reasoning tasks with dependency relationships, which can more accurately process complex tasks, thereby improving the accuracy of the reasoning result and the efficiency of obtaining the reasoning result; further, according to the reasoning result, a more accurate interaction response can be quickly generated, achieving the purpose of quickly and accurately responding to the request data containing complex tasks, thereby realizing the technical effect of improving the accuracy and efficiency of processing the request data containing complex tasks, and further solving the technical problems of poor accuracy and low efficiency of the request processing method for complex tasks in the related art.

[0050] As an optional implementation manner, the task decomposition method of the reasoning thought chain is used to represent the dependency relationship between multiple reasoning tasks. In the above step S22, using the comprehensive prompt word to drive the target request processing model to disassemble and infer the request data, the implementation steps for obtaining the reasoning result are as follows:

[0051] Step S220, using the comprehensive prompt word to drive the target request processing model to disassemble the request data to construct multiple reasoning tasks, and performing reasoning processing on the multiple reasoning tasks to obtain the reasoning result.

[0052] The above reasoning tasks may include at least one of the following: unclear tasks, harmful tasks, conflicting tasks, experience-based tasks, atomic instruction decomposition tasks, location reasoning tasks, and chatty tasks. Multiple reasoning steps may be included in the reasoning task. Each reasoning task may include different reasoning steps.

[0053] It should be noted that in this application, the specific types and quantities of reasoning tasks are not limited. Different types and / or quantities of reasoning tasks can be constructed for different request data. For example, for a certain request data A, using a comprehensive prompt to drive the target request processing model to decompose the request data, 2 reasoning tasks can be constructed (including an atomic instruction decomposition task and a location reasoning task); for another example, for a certain request data B, using a comprehensive prompt to drive the target request processing model to decompose the request data, 3 reasoning tasks can be constructed (including an unclear task, a harmful task, and a conflicting task).

[0054] In an exemplary application scenario, the above target request processing model implicitly learns the dependency relationships between multiple reasoning tasks (i.e., the associations and collaborations between multiple reasoning tasks). Using a comprehensive prompt to drive the target request processing model, the target request processing model combines the dependency relationships between multiple reasoning tasks to decompose the request data, constructs multiple reasoning tasks corresponding to the request data, and further performs reasoning processing on the constructed multiple reasoning tasks to obtain a reasoning result.

[0055] It is easy to understand that through the above step S220, in the embodiments of this application, using a comprehensive prompt to drive the target request processing model, the target request processing model combines the dependency relationships between multiple reasoning tasks, decomposes the request data containing complex tasks into multiple reasoning tasks, enabling the target request processing model to more accurately perform reasoning processing on each reasoning task in the multiple reasoning tasks, thereby improving the performance of the target request processing model in processing complex tasks and improving the accuracy of the reasoning result.

[0056] As an optional implementation manner, the comprehensive prompt includes a task type list. In the above step S220, using the comprehensive prompt to drive the target request processing model to perform reasoning processing on multiple reasoning tasks to obtain a reasoning result includes the following implementation steps:

[0057] Step S221: Use the comprehensive prompt to drive the target request processing model to perform inference calculation processing on multiple inference tasks to obtain calculation processing results corresponding to the multiple inference tasks respectively. Select the target task type corresponding to the request data from the task type list according to the multiple calculation processing results, and perform inference transformation processing on the request data based on the target task type to obtain an inference result, where the target task type is used to determine the processing strategy corresponding to the request data.

[0058] The above task type list may include multiple task types to be selected (such as, unclear task type, harmful task type, conflict task type, experience task type, simple instruction task type, complex instruction task type, chat task type, other task type). The multiple task types to be selected included in the task type list can be adjusted according to processing requirements (such as, adding task types, deleting task types).

[0059] The above calculation processing results can be used to represent the results corresponding to each inference task in the multiple inference tasks. The calculation processing results can be obtained using the execution strategy corresponding to each inference task.

[0060] It should be noted that each inference task corresponds to an execution strategy. For example, if an inference task is an unclear task, the unclear task execution strategy can be determined as the execution strategy of this inference task. Using this unclear task execution strategy to perform inference calculation on this inference task can obtain the calculation processing result corresponding to this inference task. Correspondingly, using the execution strategy corresponding to each inference task in the multiple inference tasks, the calculation processing results corresponding to each inference task in the multiple inference tasks can be obtained.

[0061] The above processing strategy may include but is not limited to: only reply strategy, atomic instruction disassembling task call strategy, location inference task call strategy. According to the calculation processing results corresponding to each inference task in the multiple inference tasks, match them with the task types to be selected in the task type list included in the comprehensive prompt in turn, and select the target task type that matches the request data. In particular, in one application scenario, the above target task type can only include one task type. Further, based on the target task type, determine the processing strategy corresponding to the request data, and use this processing strategy to perform inference transformation processing on the request data to obtain an inference result.

[0062] It is easy to understand that, through the above step S221, in the embodiment of the present application, by performing inference calculations on multiple inference tasks respectively, obtaining the calculation processing results corresponding to the multiple inference tasks respectively, and then determining the task type matching the user request including complex tasks according to the multiple calculation processing results, and processing the complex tasks according to the processing strategy corresponding to the matched task type, the complex tasks can be processed more accurately, thereby improving the accuracy of the inference results.

[0063] As an optional implementation manner, the target task type includes a first task type, and the first task type is used to represent the task type of responding to request data by controlling in-vehicle devices or functions. In the above step S221, the inference conversion processing of the request data is performed according to the target task type, and the implementation steps of obtaining the inference result are as follows:

[0064] Step S2201, in response to the target task type being the first task type, perform inference conversion on the request data based on the comprehensive prompt word to obtain a conversion result;

[0065] Step S2202, obtain the inference result according to the comprehensive prompt word and the conversion result.

[0066] The above first task type includes: experience task type, simple instruction task type, complex instruction task type.

[0067] The above conversion result is used to represent one or more atomic instructions corresponding to the request data. The above atomic instruction can be an instruction that includes an independent in-vehicle function operation or execution action and can be independently executed by the machine without relying on other instructions. The above simple instruction task can refer to the request data that includes one atomic instruction. The above complex instruction task can refer to an instruction that includes multiple atomic instructions. By converting the request data into one or more atomic instructions, it can assist the vehicle to execute these atomic instructions more accurately in sequence, which helps to perform interactive responses efficiently subsequently.

[0068] In an exemplary application scenario, it is judged whether the target task type corresponding to the request data belongs to the first task type (that is, whether it is necessary to control in-vehicle devices or functions to respond to the request data) according to the judgment conditions included in the comprehensive prompt word. When the target task type matched by the request data belongs to the first task type, based on the comprehensive prompt word, the execution strategy corresponding to the target task type is used to perform inference conversion on the request data, and the user request is converted into one or more atomic instructions; further, the conversion result is processed according to the comprehensive prompt word to obtain the inference result.

[0069] Still in the above application scenario, when the target task type matched by the request data does not belong to the first task type, it is determined that the processing strategy corresponding to the request data is the only reply strategy, and there is no need to perform inference conversion on the user request.

[0070] It is easy to understand that through the above steps S2201 to S2202, in the embodiments of the present application, first, it is determined whether the target task type matched by the request data belongs to the first task type. For the request data whose target task type belongs to the first task type, the request data is converted into one or more independent atomic instructions, which helps the vehicle to accurately execute these atomic instructions in sequence, thereby improving the accuracy of the inference result.

[0071] As an optional implementation manner, in the above step S2201, obtaining the inference result according to the comprehensive prompt word and the conversion result includes the following implementation steps:

[0072] Step S2211, perform position recognition on the conversion result based on the comprehensive prompt word to obtain a position recognition result;

[0073] Step S2221, use the position recognition result and the position mapping relationship to determine the control object position information, where the position mapping relationship is used to represent the corresponding relationship between the position recognition result and the control object position information;

[0074] Step S2231, obtain the inference result according to the control object position information and the conversion result.

[0075] The above position recognition result is used to determine whether the atomic instruction contains position information to be converted. The above position information to be converted can be used to represent the position information that does not belong to the control object position information. In particular, when the above vehicle is a vehicle with three rows of seats (that is, two seats in the first row, two seats in the second row, and two seats in the third row), the above control object position information may include: the driver's seat or the main driver's seat, the co-driver's seat, the left side of the second row, the right side of the second row or the boss's seat, the left side of the third row, the right side of the third row.

[0076] Perform position recognition on each atomic instruction in the multiple atomic instructions of the conversion result based on the comprehensive prompt word, determine whether the atomic instruction contains position information to be converted, and obtain the position recognition result corresponding to the atomic instruction. Further, if the atomic instruction contains position information to be converted, use the position mapping relationship to perform position inference on the position recognition result to determine the control object position information corresponding to the atomic instruction, and replace the position information to be converted included in the atomic instruction with the control object position information; if the atomic instruction does not contain position information to be converted, there is no need to perform position inference on the atomic instruction. Similarly, perform position recognition and position inference on each atomic instruction in the conversion result to obtain the inference result.

[0077] It should be noted that the italicized text content in this application is an optional exemplary content of prompt words (such as, comprehensive prompt words, process prompt words, inference task prompt words), so that those skilled in the art can understand the technical content of this application more clearly.

[0078] In an exemplary application scenario, the above comprehensive prompt words are as follows: "You are an in-vehicle voice assistant of an intelligent vehicle, named 'Xiaop', responsible for listening to all conversations among in-vehicle users. You need to analyze the current user interaction request input by the user in combination with the input historical conversation information and give a standard output. The standard structured form of the user interaction request is: 'User location: User requirement'. Among them, different users are distinguished by 'user location', including: driver's seat / passenger seat / left side of the second row / right side of the second row / left side of the third row / right side of the third row. The list of 'types' that the user interaction request matches: [unclear task type, harmful task type, conflict task type, experience task type, simple instruction task type, complex instruction task type, chat task type, other task type]. Please select the 'type' that best matches the current user interaction request input by the user from them and <answer> (Note: only select from the 'type' list, remember not to fabricate at will, and each user request can only match one type). When your <answer> belongs to one of 'experience instruction','simple instruction', 'complex instruction', you need to convert the user request into one or more atomic instructions ('atomic instruction' corresponds to an instruction that contains an independent in-vehicle function operation or execution action and can be independently executed by the machine without relying on other instructions); if the converted atomic instruction contains 'location information', then you need to infer the 'concrete location' of the object to be controlled by the user in combination with the mapping relationship between 'user location' and 'concrete location' (including: driver's seat, passenger seat, left side of the second row, right side of the second row, left side of the third row, right side of the third row), and replace the 'location information' in the atomic instruction with the inferred 'concrete location' and then output <instruction>." This comprehensive prompt word includes the judgment conditions for judging the inference task and the processing strategies for different types of inference tasks.

[0079] Still in the above application scenario, the above position mapping relationship may include: "The right side of the driver's seat refers to the co-driver's seat, the direct rear or the rear refers to the left side of the second row, the oblique rear or the diagonal or the right side of the rear row or the right rear refers to the right side of the second row, and the rear row refers to the second row; the left side of the co-driver's seat refers to the driver's seat, the direct rear or the rear refers to the right side of the second row, the oblique rear or the diagonal or the left side of the rear row or the left rear refers to the left side of the second row, and the rear row refers to the second row; the right side of the left side of the second row refers to the right side of the second row, the direct front or the front refers to the driver's seat, the direct rear or the rear refers to the left side of the third row, the oblique front or the diagonal or the right side of the front row or the right front refers to the co-driver's seat, the oblique rear or the right side of the rear row or the right rear refers to the right side of the third row, the front row refers to the first row, and the rear row refers to the third row; the left side of the right side of the second row refers to the left side of the second row, the direct front or the front refers to the co-driver's seat, the direct rear or the rear refers to the right side of the third row, the oblique front or the diagonal or the left side of the front row or the left front refers to the driver's seat, the oblique rear or the left side of the rear row or the left rear refers to the left side of the third row, the front row refers to the first row, and the rear row refers to the third row; the right side of the left side of the third row refers to the right side of the third row, the direct front or the front refers to the left side of the second row, the oblique front or the diagonal or the right side of the front row or the right front refers to the right side of the second row, and the front row refers to the second row; the left side of the right side of the third row refers to the left side of the third row, the direct front or the front refers to the right side of the second row, the oblique front or the diagonal or the left side of the front row or the left front refers to the left side of the second row, and the front row refers to the second row." This position mapping relationship can be set according to the type of vehicle.

[0080] Still in the above application scenario, based on the comprehensive prompt words, perform position recognition on the conversion result (i.e., perform a position reasoning task) to obtain a position recognition result. Further, judge the position recognition result. If the judgment result indicates that the conversion result contains the position information to be converted, use the position mapping relationship to perform position reasoning on the position recognition result to determine the control object position information. For example, if a certain atomic instruction in the conversion result contains the position information to be converted "behind the driver's seat", then perform position recognition on the conversion result based on the comprehensive prompt words to obtain a position recognition result, and use the position mapping relationship to perform position reasoning on this position recognition result to determine the control object position information (i.e., "the left side of the second row"). Further, replace the position information to be converted "behind the driver's seat" contained in the atomic instruction in the above conversion result with this control object position information "the left side of the second row", and determine the conversion result after the position information replacement as the reasoning result.

[0081] It is easy to understand that through the above steps S2211 to S2231, in the embodiment of the present application, position reasoning is performed on each atomic instruction in the conversion result, and the position information to be converted contained in each atomic instruction is replaced with the control object position information, enhancing the clarity of the atomic instruction and making it easier for the system to understand and process the atomic instruction, thereby improving the accuracy of the interaction response.

[0082] As an alternative implementation, in the above request processing method for the intelligent cockpit, the target request processing model is obtained by training the initial request processing model with multiple sets of training data. The multiple sets of training data are constructed based on the task types respectively corresponding to multiple inference tasks. The training data includes the original data, the reply results corresponding to the original data, and the thought chains corresponding to the original data.

[0083] The above task types include: unclear task type, harmful task type, conflict task type, experience task type, atomic instruction disassembling task type, location reasoning task type, chatting task type.

[0084] The above multiple sets of training data may include subtask training data. The subtask training data may include unclear task training data, harmful task training data, conflict task training data, experience task training data, atomic instruction disassembling task training data, location reasoning task training data, chatting task training data. This training data can be constructed by disassembling the original data according to the task disassembling method of the thought chain. The task disassembling method of the above thought chain can be determined by using the process prompt words and the inference task prompt words respectively corresponding to multiple inference tasks. The above process prompt words can be used to determine the execution order of multiple inference tasks. This process prompt word can represent the dependency relationship between multiple inference tasks corresponding to the directed acyclic graph. The above inference task prompt words can be used to determine the execution strategy of the inference task corresponding to this inference task prompt word. The inference task prompt words can include unclear task prompt words, harmful task prompt words, conflict task prompt words, experience task prompt words, atomic instruction disassembling task prompt words, location reasoning task prompt words, chatting task prompt words. The above original data can be obtained through a predefined data set. The training data may also include comprehensive training data. This comprehensive training data is constructed based on the subtask training data.

[0085] The above initial request processing model may be a model that has not been optimized at the beginning of training. This initial request processing model may refer to a deep learning model with a large number of model parameters (i.e., a large model), which usually can contain hundreds of millions, tens of billions, hundreds of billions, trillions or even more than one quadrillion model parameters. The large model can also be called a foundation model. Through the pre-training of the large model with a large amount of unlabeled corpus, a pre-trained model with more than one billion parameters is produced. This kind of model can adapt to a wide range of downstream tasks and has good generalization ability. For example, large language models (LLMs), multi-modal pre-training models, etc.

[0086] The above model parameters can be learnable variables in a machine learning model, which can assist the model in processing inputs and generating outputs. For example, in a request processing model, the model parameters can include, but are not limited to: the weights of a neural network, the parameters of an attention mechanism. The optimization of the model parameters of the request processing model is achieved through a training process, during which the model parameters of the dialogue quality inspection model are continuously adjusted to enable the target request processing model to learn the dependencies between multiple task types and the execution strategies to be adopted for each task type.

[0087] In an exemplary application scenario, Figure 3 is a schematic diagram of an optional directed acyclic graph according to an embodiment of the present application. As Figure 3 shown, the directed acyclic graph can be used to determine the execution order of tasks of multiple task types (unclear task type, harmful task type, conflict task type, experience-based task type, atomic instruction disassembling task type, location reasoning task type, chitchat task type, where the atomic instruction disassembling task type can be used to determine whether the request data belongs to a simple instruction task type or a complex instruction task). A process prompt can be constructed to represent the directed acyclic graph.

[0088] The constructed process prompt is as follows: "complex_task_prompt = You are an in-vehicle voice assistant (assistant) of an intelligent vehicle, named 'Xiaop', responsible for listening to all conversations among in-vehicle users (user). You need to combine the historical conversations between the user (user) and the in-vehicle voice assistant (assistant), the content of the responses from 'Xiaop', and the current user interaction request, and strictly follow the following steps to gradually analyze the request input by the user currently:

[0089] 1. If the input does not contain historical conversation data between the user (user) and the in-vehicle voice assistant (assistant), skip this step; otherwise, please determine whether the current user interaction request has information to be supplemented (such as an entity or an action) based on the semantic coherence or the relationship between entities in the historical conversation data and the current user interaction request, and supplement the current user interaction request with the information to be supplemented from the historical conversation information. If the current user interaction request depends on historical conversation for completion, output 'Dependency on historical turn for completion: Yes', and at the same time replace the current user interaction request with the completed request for the next judgment; if it does not depend on historical turn for completion, proceed to the next judgment based on the input user interaction request;

[0090] 2. Determine whether the input user interaction request is clear. Unclear means: the sentence components or words are arranged in a disorderly manner; or the sentence contains a lot of messy or meaningless information; or the sentence lacks components resulting in ambiguous semantics; or the sentence only contains single entity information, and single entity information includes but is not limited to: multimedia resource nouns such as singers, songs, etc., vehicle-related equipment or function nouns (e.g., air conditioners, refrigerators, car windows, cameras, etc.). If the user interaction request matches the unclear situation, please do not make assumptions or speculations about the unclear text and directly answer 'unclear';

[0091] 3. If the input user interaction request is clear, you need to determine whether the request contains harmful information (harmful information includes but is not limited to: vulgarity, abuse, insult, malicious attack, violation of social morality, violation of laws and regulations, and malicious evaluation of a certain person, event, brand or product, etc.). If it contains such information, answer 'harmful';

[0092] 4. When the input user interaction request contains two opposite operation actions for the same in-vehicle device or function at the same time, and these two actions are mutually exclusive logically and cannot be executed simultaneously (e.g., 'open' and 'close', 'turn on' and 'turn off', 'rise' and 'fall', 'turn up' and 'turn down', etc.), then answer 'conflict';

[0093] 5. When none of the above steps match, you need to determine whether the user interaction request is a clear vehicle-mounted instruction. A clear vehicle-mounted instruction needs to include an 'operation action' and a specific'vehicle-mounted device' or'vehicle-mounted function' to be operated. The'vehicle-mounted functions' include, but are not limited to: control of sound zones, control of sound effects volume, playback and control of multimedia resources such as audio / video / radio stations, opening / closing / downloading of apps, selection / confirmation / cancellation of interface lists, global controls such as general exit / closing, making phone calls, navigation services, time setting, mathematical calculations, knowledge-related query functions such as ancient poems, vocabulary, exchange rates, stocks, weather, time / date, news, multilingual translation, vehicle restrictions, traffic, vehicle usage tutorials, creating story picture books, painting, and multimodal Q&A using images of various positions of the vehicle. When the user interaction request matches a clear vehicle-mounted instruction, if the vehicle-mounted instruction consists of a single atomic instruction, answer 'Simple Instruction'; if it consists of two or more atomic instructions, answer 'Complex Instruction' (multiple atomic instructions can be independent of each other and / or related to each other), and split out multiple atomic instructions in sequence. If the atomic instruction contains relative position information (relative positions include, but are not limited to: front row, back row, left side, right side, directly in front, right rear, left front, diagonal, rear diagonal, etc.), please infer the 'control object position information' of the object the user wants to control based on the'seat layout' and 'user position' of the intelligent vehicle ('control object position information' can be the user position, limited to: 'driver's seat, passenger seat, left side of the second row, right side of the second row, left side of the third row, right side of the third row'; or it can be the row where the user is located, limited to: 'first row, second row, third row'), and replace the position information in the atomic instruction with the inferred 'control object position information'. After completing the splitting and position inference of the vehicle-mounted instruction, output one or more semantically complete atomic instructions;

[0094] 6. When the request does not match a clear vehicle-mounted instruction and expresses the user's experience or feeling in the vehicle usage scenario, you need to determine whether it can be associated with the control of one or more vehicle devices or functions to improve the user's experience or meet the user's needs. If it can be associated, answer 'Experience-based Instruction', and at the same time output one or more associated standard <Instructions>;

[0095] 7. When the request does not match a clear vehicle-mounted instruction, you can combine the historical conversation to determine whether the current request is chatting with 'Xiaop'. The chat content includes, but is not limited to: sharing of hobbies, personal experiences, personal recent situations, emotional expressions, discussion of news events, food, weather, travel, fitness, work and study topics, etc., as well as role-playing exchanges with the name 'Xiaop', such as discussing Xiaop's gender, age, appearance, personality, greeting Xiaop, complaining about Xiaop, etc. If the user is chatting with 'Xiaop', answer 'Chatting'.

[0096] 8. When the request does not match any explicit in-vehicle instructions, and no in-vehicle device / function can be associated to improve the user experience, and it is not a casual chat between the user and 'Xiaop', directly answer 'Other'.

[0097] Please note that this is a process of gradually analyzing the current user interaction request, and it is necessary to strictly judge and answer according to the above steps. If a certain step is matched, there is no need to continue to judge the next step, and just directly answer the corresponding result. Please strictly follow the unclear task type, harmful task type, conflicting task type, experience task type, simple instruction task type, complex instruction task type, casual chat task type, other task type to answer. Remember, when your answer belongs to one of the experience task type, simple instruction task type, complex instruction task type, it is necessary to output the corresponding standard <instruction>. To perceive and distinguish users in different positions inside the vehicle, the user input format is: 'User location: User demand', where the user location includes: driver's seat / passenger seat / left side of the second row / right side of the second row / left side of the third row / right side of the third row.

[0098] According to the above process prompt words, the execution process of multiple tasks for a certain user interaction request can be determined.

[0099] Still in the above application scenario, for this directed acyclic graph, respectively construct the inference task prompt words corresponding to multiple task types.

[0100] The constructed unclear task prompt words are as follows: "You are an in-vehicle voice assistant of an intelligent vehicle, named 'Xiaop', responsible for listening to all conversations between users in the vehicle. You need to analyze the request input by the user currently in combination with the input historical conversation data and give a standard output.

[0101] The standard structured form of the user interaction request is: 'User location: User demand'. Among them, different users are distinguished by 'user location', including: driver's seat / passenger seat / left side of the second row / right side of the second row / left side of the third row / right side of the third row.

[0102] The definition of an 'unclear' user interaction request usually refers to:

[0103] 1. The components or words of the sentence are arranged in a disorderly manner;

[0104] 2. The sentence contains a lot of messy or meaningless information.

[0105] If the 'user interaction request' matches one or more of the above 'unclear' definitions, answer 'unclear', otherwise answer 'not matched'.

[0106] The constructed harmful task prompt words are as follows: "You are an in-vehicle voice assistant for a smart car named 'Xiaop', responsible for listening to all conversations among in-vehicle users. You need to analyze the current request input by the user in combination with the input historical conversation data and give a standardized output."

[0107] The standard structured form of the user interaction request is: 'User location: User requirement'. Among them, different users are distinguished by 'user location', including: driver's seat / passenger seat / left side of the second row / right side of the second row / left side of the third row / right side of the third row.

[0108] The definition of a 'harmful' user request includes but is not limited to:

[0109] 1. Vulgar language: including but not limited to swear words, obscenities, insults, personal attacks and other words that make people feel uncomfortable or offended;

[0110] 2. Violating laws and regulations: including but not limited to illegal activities, abetting crime, disrupting social order and other illegal acts that endanger the safety of society and others;

[0111] 3. Involving insult: directly or indirectly insulting, humiliating or demeaning others, damaging the self-esteem and dignity of others;

[0112] 4. Involving discrimination: discrimination or discriminatory terms against specific groups or individuals;

[0113] 5. Topics related to affecting unity: including but not limited to beliefs, customs and traditions, etc.;

[0114] 6. Remarks that may lead to social instability;

[0115] 7. Sensitive topics related to culture, or controversial social views, etc.;

[0116] 8. Topics violating social morality;

[0117] 9. Involving personal privacy or sensitive information: including but not limited to personal identity information such as mobile phone numbers, bank card numbers or passwords, social insurance numbers and other personal sensitive information;

[0118] 10. Involving vehicle and driving safety: remarks that require bypassing the vehicle safety system and encourage speeding, drunk driving, fatigue driving, violating traffic signals and other behaviors that endanger driving safety;

[0119] 11. Related topics about brands: Personal subjective feelings or evaluations of a certain brand or product (e.g., 'I like car XL'), comparisons between brands or products (e.g., 'Which is better, car XP or car XM?', 'Which electric vehicle is the best?'), guiding 'Xiaop' to evaluate or slander a certain product or brand maliciously (e.g., 'Do you think highly of car XM?') and other remarks that tend to incite disputes between brands or products and damage the image of 'Xiaop'.

[0120] If the 'user interaction request' matches one or more of the above 'harmful' definitions, answer 'harmful'; otherwise, answer 'not matched'.

[0121] The constructed conflict task prompt is as follows: "You are an in-vehicle voice assistant of an intelligent car, named 'Xiaop', responsible for listening to all conversations among in-vehicle users. You need to analyze the request input by the user currently in combination with the historical conversation data and give a standardized output."

[0122] The standard structured form of the user interaction request is: 'User location: User demand'. Among them, different users are distinguished by 'user location', including: driver's seat / passenger seat / left side of the second row / right side of the second row / left side of the third row / right side of the third row.

[0123] The definition of a 'conflicting' user request usually refers to:

[0124] 1. The request contains two or more operation actions on a certain device or function at the same time, and these actions have an obvious logical conflict (e.g., 'open' and 'close', 'turn on' and 'turn off', 'rise' and 'fall', 'turn up' and 'turn down', 'adjust forward' and 'adjust backward', 'raise the temperature' and 'lower the temperature' and other opposite operations, and cannot perform operations on a device or function at the same time).

[0125] If the 'user interaction request' matches the above 'conflict' definition, answer 'conflict'; otherwise, answer 'not matched'.

[0126] The constructed atomic instruction decomposition task prompt is as follows: "You are an in-vehicle voice assistant of an intelligent car, named 'Xiaop', responsible for listening to all conversations among in-vehicle users. You need to analyze the request input by the user currently in combination with the historical conversation data and give a standardized output."

[0127] The standard structured form of the user interaction request is: 'User location: User demand'. Among them, different users are distinguished by 'user location', including: driver's seat / passenger seat / left side of the second row / right side of the second row / left side of the third row / right side of the third row.

[0128] If the input user interaction request does not match 'unclear', 'harmful', 'conflicting', and 'experience - type instructions', and it is a clear instruction from the user to the intelligent vehicle (including but not limited to: vehicle equipment control, sound volume control, multimedia resource playback and control, opening / closing / downloading of APPs, navigation services, chatting, and queries about ancient poems, calculators, exchange rates, stocks, weather, calendars, news, multilingual translation, vehicle traffic restrictions, transportation, encyclopedic knowledge, etc., creating story picture books, drawing, image generation, and multi - modal Q&A using images of all aspects of the vehicle), then you need to determine whether the user interaction request is a'simple instruction' or a 'complex instruction'.

[0129] An 'atomic instruction' corresponds to a clear and independent in - vehicle function or action, which can be executed independently without relying on other instructions. A'simple instruction' means that the user interaction request consists of one atomic instruction; a 'complex instruction' means that the user interaction request consists of multiple atomic instructions. If the user interaction request matches a simple instruction, answer'simple instruction'; if it matches a complex instruction, answer 'complex instruction' and split out multiple atomic <instructions> in sequence. Using this atomic instruction disassembling task prompt, it is possible to determine whether the user interaction request belongs to the simple instruction task type or the complex instruction task type.

[0130] The constructed location reasoning task prompt is as follows: "You are an in - vehicle voice assistant of an intelligent vehicle (assistant), named 'Xiaop', responsible for listening to all conversations among in - vehicle users (user). You need to analyze the user interaction request input by the user currently, combined with the input historical conversation data, and give a standardized output.

[0131] The standard structured form of the user interaction request is: 'User location: User requirement'. Among them, different users are distinguished by the 'user location', including: driver's seat / passenger seat / left side of the second row / right side of the second row / left side of the third row / right side of the third row.

[0132] The seat layout of the intelligent vehicle: There are a total of three rows of seats (i.e., two seats in the first row, two seats in the second row, and two seats in the third row), and six 'control object position information' are marked. The detailed description of each position is as follows:

[0133] 1. 'Driver's seat': Refers to the position of the driver, located on the left side of the first row, also known as the 'driver's position';

[0134] 2. 'Passenger seat': Refers to the position of the co - driver, located on the right side of the first row;

[0135] 3. 'Left side of the second row': Refers to the left position in the second row (i.e., the middle row) of seats;

[0136] 4. 'Right side of the second row': refers to the right position in the second row (i.e., the middle row) of seats, also known as the 'boss seat';

[0137] 5. 'Left side of the third row': refers to the left position in the third row of seats;

[0138] 6. 'Right side of the third row': refers to the right position in the third row of seats.

[0139] Among them, the mapping relationships between the six 'control object position information' are as follows:

[0140] The right side of the driver's seat refers to the co-driver's seat, directly behind or at the back refers to the left side of the second row, diagonally behind or at an oblique angle or the right side of the back row or the right rear refers to the right side of the second row, and the back row refers to the second row;

[0141] The left side of the co-driver's seat refers to the driver's seat, directly behind or at the back refers to the right side of the second row, diagonally behind or at an oblique angle or the left side of the back row or the left rear refers to the left side of the second row, and the back row refers to the second row;

[0142] The right side of the left side of the second row refers to the right side of the second row, directly in front or in front refers to the driver's seat, directly behind or at the back refers to the left side of the third row, diagonally in front or at an oblique angle or the right side of the front row or the right front refers to the co-driver's seat, diagonally behind or the right side of the back row or the right rear refers to the right side of the third row, the front row refers to the first row, and the back row refers to the third row;

[0143] The left side of the right side of the second row refers to the left side of the second row, directly in front or in front refers to the co-driver's seat, directly behind or at the back refers to the right side of the third row, diagonally in front or at an oblique angle or the left side of the front row or the left front refers to the driver's seat, diagonally behind or the left side of the back row or the left rear refers to the left side of the third row, the front row refers to the first row, and the back row refers to the third row;

[0144] The right side of the left side of the third row refers to the right side of the third row, directly in front or in front refers to the left side of the second row, diagonally in front or at an oblique angle or the right side of the front row or the right front refers to the right side of the second row, the front row refers to the second row;

[0145] The left side of the right side of the third row refers to the left side of the third row, directly in front or in front refers to the right side of the second row, diagonally in front or at an oblique angle or the left side of the front row or the left front refers to the left side of the second row, the front row refers to the second row.

[0146] If the input user interaction request is an 'atomic instruction' of a smart vehicle that contains 'location information' ('atomic instruction' corresponds to a clear and independent in-vehicle function or action that can be executed independently without relying on other instructions), please combine the seat layout of the smart vehicle and strictly follow the steps below to infer the 'controlled object location information' of the object to be controlled by the user (the controlled object location information includes: driver's seat, front passenger seat, left side of the second row, right side of the second row, left side of the third row, right side of the third row):

[0147] 1. If the instruction only contains the 'controlled object location information', no location inference is required, and the input atomic instruction can be directly output;

[0148] 2. Otherwise, based on the mapping relationship between the user's location information and the 'controlled object location information', infer the 'controlled object location information' of the object to be controlled by the user, and replace the 'location information' in the atomic instruction with the inferred 'controlled object location information' and then output.

[0149] The constructed experience-based task prompt words are as follows: "You are an in-vehicle voice assistant (assistant) of a smart vehicle named 'Xiaop', responsible for listening to all conversations among in-vehicle users (users). You need to analyze the current input request of the user by combining the input historical conversation data and give a standard output.

[0150] The standard structured form of the user interaction request is: 'User Location: User Requirement'. Among them, different users are distinguished by the 'User Location', including: driver's seat / front passenger seat / left side of the second row / right side of the second row / left side of the third row / right side of the third row.

[0151] If the input user interaction request does not match 'unclear', 'harmful', 'conflicting','simple instruction' and complex instructions, it is necessary to determine whether the user request matches 'experience-based instruction'. 'Experience-based instruction' generally refers to: in the vehicle usage scenario, the user interaction request does not contain the operation of vehicle equipment or functions, but fully expresses the user's experience or feeling in the vehicle usage scenario, and can improve the user's vehicle usage feeling or experience by associating with the control of vehicle equipment or functions. If it matches, answer 'experience-based instruction', and at the same time, it is necessary to output one or more standard <instructions> associated with vehicle equipment or functions; otherwise, answer 'not matched' ('not matched' includes: the user request contains the operation of vehicle equipment or functions; or although the user interaction request expresses the user's experience and feeling (such as 'The sun is really nice today', 'The grass is really green', etc.), but it cannot be associated with the operation of some vehicle equipment or functions, etc.)."

[0152] The constructed prompt words for the chat task are as follows: "You are an in-vehicle voice assistant for a smart car named 'Xiaop', responsible for listening to all conversations between users in the car. You need to analyze the user interaction requests in the current input based on the input historical conversation data and give a standard output."

[0153] The standard structured form of the user interaction request is: 'User location: User requirement'. Among them, different users are distinguished by 'user location', including: driver's seat / passenger seat / left side of the second row / right side of the second row / left side of the third row / right side of the third row.

[0154] If the input user interaction request does not match 'unclear', 'harmful', 'conflicting', 'experience instructions','simple instructions' and complex instructions, it is necessary to determine whether the user request matches 'chat'. 'Chat' generally refers to: in the context of using a car, the user is chatting with 'Xiaop' you, and the chat content includes but is not limited to discussions on hobbies, personal experiences, sharing of personal updates, emotional expressions, news events, food, weather, travel, fitness, work and study topics, etc., as well as role-playing exchanges with the address of 'Xiaop', such as discussing Xiaop's gender, age, appearance, personality, greeting Xiaop, complaining about Xiaop, etc. If the user interaction request matches the chat, reply 'chat'. Otherwise, reply 'other'.

[0155] Still in the above application scenario, after constructing the process prompt words and the inference task prompt words corresponding to multiple task types, the original data is disassembled using the execution order of multiple inference tasks determined by the process prompt words to construct multiple inference tasks corresponding to the original data, and each inference task corresponds to a task type. Furthermore, for the inference task of a certain task type, the inference task is processed using the inference task prompt word corresponding to the inference task to obtain the reply result corresponding to the inference task and the thought chain corresponding to the inference task. Further, the training data corresponding to the task type is constructed using the original data, the reply result corresponding to the inference task, and the thought chain corresponding to the inference task. Thus, for each task type, a set of training data corresponding to the task type can be constructed, thereby constructing seven sets of subtask training data, and using the seven sets of subtask training data to construct comprehensive task training data.

[0156] Still in the above application scenario, the training process of the target request processing model can be as follows: Adopting a cold start training method, input the constructed seven groups of subtask training data and comprehensive task training data (i.e., a total of eight groups of training data) into the initial request processing model, perform hybrid training on the initial request processing model, calculate the loss value corresponding to each group of training data respectively, use multiple loss value and total loss calculation methods (such as weighted average method, average value calculation method, direct addition method) to calculate the total loss value, and use the total loss value to update the parameters of the initial request processing model to obtain the target request processing model. The target request processing model obtained by performing hybrid training on the initial request processing model using these eight groups of training data can enable the target request processing model to implicitly learn the dependency relationships between multiple reasoning tasks, ensuring that in the subsequent process of using the target request processing model to process request data, there is no need to rely on the above process prompt words. In addition, the target request processing model can learn the execution strategies to be adopted for each task type, so as to ensure that in the subsequent process of using the target request processing model to process request data, there is also no need to rely on the above multiple reasoning task prompt words. Thus, a target request processing model that can be driven only by comprehensive prompt words is trained, improving the efficiency of the model in processing complex tasks.

[0157] It should be noted that the specific examples of the comprehensive prompt words, process prompt words, and reasoning task prompt words listed in the embodiments of the present application are only optional setting methods in the exemplary application scenarios, and different prompt words from those in the embodiments of the present application can also be adopted according to the scenario requirements in other application scenarios. That is to say, the present application does not limit the specific setting methods of the above comprehensive prompt words, process prompt words, and reasoning task prompt words.

[0158] It is easy to understand that through the above request processing method for the intelligent cockpit, in the embodiments of the present application, by constructing multiple groups of training data corresponding to multiple task types, the training data not only includes the original data, but also includes the reply results corresponding to the original data and the thought chains corresponding to the original data. The training data contains more comprehensive information, which helps the model to more accurately learn the feature information of complex tasks, thereby improving the performance of the model in processing complex tasks. In addition, in the embodiments of the present application, the target request processing model obtained through the above request processing method for the intelligent cockpit does not need to rely on process prompt words or multiple reasoning task prompt words during the stage of reasoning and processing request data. Compared with the related art that relies on multiple prompt words and multiple models to process interaction requests, it can improve the efficiency of the target request processing model in outputting reasoning results, thus improving the efficiency of human-computer interaction.

[0159] As an optional implementation manner, the request data further includes historical conversation data, and the above request processing method for the intelligent cockpit further includes the following implementation steps:

[0160] Step S24: Determine whether there is information to be supplemented in the request data based on the historical conversation data.

[0161] Step S25: In response to the existence of information to be supplemented in the historical conversation data, update the request data using the information to be supplemented.

[0162] The above historical conversation data may include conversation context information and timestamp information. During the interaction between the user and the intelligent cockpit system, the system can automatically record the conversation content to form historical conversation data. The above information to be supplemented may include entity information and action information.

[0163] In an exemplary application scenario, still as Figure 3 shown, before disassembling and reasoning the user request, first determine whether it is necessary to rely on the historical conversation data to complete the user request. Specifically, judge the request data based on the historical conversation data to determine whether there is information to be supplemented in the request data (such as missing entity information, missing action information). In response to the existence of information to be supplemented in the historical conversation data, judge whether there is corresponding information to be supplemented in the historical conversation data. If there is information to be supplemented in the historical conversation data, then use the information to be supplemented in the historical conversation data to rewrite the request data and update the request data.

[0164] It is easy to understand that through the above steps S24 to S25, in the embodiment of the present application, using the historical conversation data to supplement information to the request data and update the request data can enable the target request processing model to process the request data more accurately, thereby improving the accuracy of the interaction response.

[0165] As an optional implementation manner, the interaction response includes an interaction reply result and a control instruction. In the above step S23, according to the reasoning result, generating the interaction response includes the following implementation steps:

[0166] Step S231: Analyze the reasoning result to generate an interaction reply result and a control instruction.

[0167] Step S232: Send the interaction reply result to the in-vehicle multimedia system corresponding to the intelligent cockpit.

[0168] Step S233: Send the control instruction to the control system of the vehicle to control the vehicle to perform the target action.

[0169] The above interaction reply result is used to represent the feedback result corresponding to the user interaction request. The above control instruction is used to control the vehicle to perform the target action. The above in-vehicle multimedia system may include, but is not limited to: in-vehicle audio, in-vehicle display screen.

[0170] In an exemplary application scenario, the above in-vehicle multimedia system may be an in-vehicle audio system. By analyzing the inference result, the response data corresponding to the user interaction request is determined. According to the response data, an interaction reply result and a control instruction are generated. Further, the interaction reply result is sent to the in-vehicle audio system corresponding to the intelligent cockpit, and the interaction reply result is played, enabling the user to directly know the feedback result. Furthermore, the control instruction is sent to the vehicle control system, enabling the control system to control the vehicle to perform the target action according to the control instruction.

[0171] It is easy to understand that through the above steps S231 to S233, in the embodiments of the present application, the interaction reply result is fed back to the user and the vehicle is controlled to perform the target action according to the control instruction, enhancing the user's vehicle-human interaction experience and improving user satisfaction.

[0172] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the present application are all information and data authorized by the user or fully authorized by all parties. And the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of relevant countries and regions, and corresponding operation entrances are provided for users to choose to authorize or refuse.

[0173] According to the embodiments of the present application, a request processing device for an intelligent cockpit is also provided. Please refer to Figure 4 . The device includes:

[0174] An acquisition module 401, configured to acquire request data, where the request data is used to represent a user interaction request received by the intelligent cockpit of the vehicle;

[0175] A processing module 402, configured to drive a target request processing model by using a comprehensive prompt word, disassemble and infer the request data to obtain an inference result, where the comprehensive prompt word is used to determine an inference criterion, the target request processing model is used to construct and process multiple inference tasks according to the task disassembling method of an inference thinking chain, and the inference result is used to determine response data corresponding to the user interaction request;

[0176] An execution module 403, configured to generate an interaction response according to the inference result.

[0177] It should be noted here that the above acquisition module 401, processing module 402, and execution module 403 correspond to steps S21 to S23 in the method embodiments. The functions of these three modules and the corresponding steps are the same in terms of implementation examples and application scenarios, but are not limited to the content disclosed in the above method embodiments.

[0178] It should be noted that each module mentioned in the above device embodiments can be implemented by software, hardware, or a combination of software and hardware. For example, when implementing the above modules in hardware, each module can be set in the same processor, or each module can be set in different processors in any combination. For another example, the above modules can be hardware components or software components stored in a memory (for example, memory 104) and processed by one or more processors (for example, processors 102a, 102b,..., 102n), and the above modules can also operate as part of a device in a Figure 1 computing terminal 10 as shown.

[0179] According to an embodiment of the present application, an embodiment of a vehicle is further provided. The vehicle includes a memory storing an executable program; and a processor for running the program, wherein when the program runs, it executes the request processing method of the intelligent cockpit described in any one of the above.

[0180] According to an embodiment of the present application, a computer-readable storage medium is further provided. The computer-readable storage medium includes a stored program, wherein when the program runs, it controls the device where the computer-readable storage medium is located to implement the request processing method of the above intelligent cockpit.

[0181] Optionally, the above computer storage medium may include, but is not limited to: a hard disk drive (HDD), a solid state drive (SSD), a USB flash drive, an optical disc, a memory card, a cloud storage medium, and a network-attached storage (NAS), etc.

[0182] Optionally, the above computer-readable storage medium can be set to store a computer program for executing the following steps: obtaining request data, where the request data is used to represent a user interaction request received by the intelligent cockpit of the vehicle; using a comprehensive prompt word to drive a target request processing model to disassemble and reason about the request data to obtain a reasoning result, where the comprehensive prompt word is used to determine the reasoning standard, the target request processing model is used to construct and process multiple reasoning tasks according to the task disassembling method of the reasoning thinking chain, and the reasoning result is used to determine response data corresponding to the user interaction request; generating an interaction response according to the reasoning result.

[0183] According to an embodiment of the present application, a computer program product is further provided. The computer program product includes a computer program that can implement the request processing method of the above intelligent cockpit when executed by a processor.

[0184] Optionally, the above computer program product may provide a service for improving the accuracy and efficiency of processing requests of the intelligent cockpit based on the above request processing method of the intelligent cockpit.

[0185] Optionally, in this embodiment, the above computer program product may be a set of instructions and codes pre-written according to the above request processing method of the intelligent cockpit. The computer program product can run on various different computer platforms, including personal computers, servers, mobile devices, etc.

[0186] Optionally, in this embodiment, the instructions and codes corresponding to the computer program product are used to implement the following method steps: obtaining request data, where the request data is used to represent a user interaction request received by the intelligent cockpit of the vehicle; using a comprehensive prompt to drive a target request processing model to disassemble and reason about the request data to obtain a reasoning result, where the comprehensive prompt is used to determine the reasoning criterion, the target request processing model is used to construct and process multiple reasoning tasks according to the task disassembling method of the reasoning thinking chain, and the reasoning result is used to determine response data corresponding to the user interaction request; generating an interaction response according to the reasoning result.

[0187] In the above embodiments of the present invention, the descriptions of each embodiment have their own emphases. For parts not detailed in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.

[0188] In several embodiments provided in the present application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only illustrative. For example, the division of the units can be a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the couplings or direct couplings or communication connections shown or discussed with each other can be through some interfaces. The indirect couplings or communication connections of the units or modules can be in electrical or other forms.

[0189] The units described as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they can be located in one place, or can be distributed to multiple units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0190] In addition, in each embodiment of the present invention, the functional units can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.

[0191] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, read-only memories (ROM, short for Read-Only Memory), random access memories (RAM, short for Random Access Memory), mobile hard disks, magnetic disks, or optical discs that can store program codes.

[0192] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.

Claims

1. A request processing method for an intelligent cockpit, characterized in that, Including: Obtain request data, where the request data is used to represent a user interaction request received by the intelligent cockpit of the vehicle; Use a comprehensive prompt to drive a target request processing model to disassemble and reason about the request data to obtain a reasoning result, where the comprehensive prompt is used to determine a reasoning criterion, the target request processing model is used to construct and process multiple reasoning tasks according to the task disassembling method of a reasoning thinking chain, and the reasoning result is used to determine response data corresponding to the user interaction request; Generate an interaction response according to the reasoning result.

2. The request processing method for the intelligent cockpit according to claim 1, characterized in that, The task disassembling method of the reasoning thinking chain is used to represent the dependency relationship between the multiple reasoning tasks. Using the comprehensive prompt to drive the target request processing model to disassemble and reason about the request data to obtain the reasoning result includes: Use the comprehensive prompt to drive the target request processing model to disassemble the request data to construct the multiple reasoning tasks, and perform reasoning processing on the multiple reasoning tasks to obtain the reasoning result.

3. The request processing method for an intelligent cockpit according to claim 2, wherein The comprehensive prompt includes a task type list. Using the comprehensive prompt to drive the target request processing model to perform reasoning processing on the multiple reasoning tasks to obtain the reasoning result includes: Use the comprehensive prompt to drive the target request processing model to perform reasoning calculation processing on the multiple reasoning tasks to obtain calculation processing results corresponding to the multiple reasoning tasks respectively, select a target task type corresponding to the request data from the task type list according to the multiple calculation processing results, and perform reasoning transformation processing on the request data according to the target task type to obtain the reasoning result, where the target task type is used to determine a processing strategy corresponding to the request data.

4. The request processing method for the intelligent cockpit according to claim 3, characterized in that, The target task type includes a first task type, and the first task type is used to represent a task type for responding to the request data by controlling in-vehicle devices or functions. Performing reasoning transformation processing on the request data according to the target task type to obtain the reasoning result includes: In response to the target task type being the first task type, perform reasoning transformation on the request data based on the comprehensive prompt to obtain a transformation result; Obtain the reasoning result according to the comprehensive prompt and the transformation result.

5. The request processing method for the intelligent cockpit according to claim 4, characterized in that, Obtain the reasoning result according to the comprehensive prompt and the transformation result includes: Perform position recognition on the transformation result based on the comprehensive prompt to obtain a position recognition result; Use the position recognition result and a position mapping relationship to determine control object position information, where the position mapping relationship is used to represent the corresponding relationship between the position recognition result and the control object position information; Obtain the reasoning result according to the control object position information and the transformation result.

6. The request processing method for the intelligent cockpit according to claim 1, wherein The target request processing model is obtained by training an initial request processing model with multiple groups of training data. The multiple groups of training data are constructed based on the task types corresponding to the multiple reasoning tasks respectively. The training data includes original data, a reply result corresponding to the original data, and a thinking chain corresponding to the original data.

7. The request processing method for the intelligent cockpit according to claim 1, wherein, The request data further includes historical conversation data, and the request processing method of the intelligent cockpit further includes: Determining whether there is information to be supplemented in the request data according to the historical conversation data; In response to the existence of the information to be supplemented in the historical conversation data, updating the request data with the information to be supplemented.

8. The request processing method for the intelligent cockpit according to claim 1, characterized in that, The interaction response includes an interaction reply result and a control instruction. Generating the interaction response according to the reasoning result includes: Analyzing the reasoning result to generate the interaction reply result and the control instruction; Sending the interaction reply result to the in-vehicle multimedia system corresponding to the intelligent cockpit; Sending the control instruction to the control system of the vehicle to control the vehicle to perform a target action.

9. A processing device for cockpit instructions, characterized in that, Including: An acquisition module, configured to acquire request data, where the request data is used to represent a user interaction request received by the intelligent cockpit of the vehicle; A processing module, configured to drive a target request processing model by using a comprehensive prompt word, disassemble and reason the request data to obtain a reasoning result, where the comprehensive prompt word is used to determine a reasoning criterion, the target request processing model is used to construct and process multiple reasoning tasks according to the task disassembling method of a reasoning thinking chain, and the reasoning result is used to determine response data corresponding to the user interaction request; An execution module, configured to generate an interaction response according to the reasoning result.

10. A vehicle, characterized in that, Including: A memory, storing an executable program; A processor, configured to run the program, where when the program runs, it executes the request processing method of the intelligent cockpit according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Response data reasoning method and device, electronic equipment and storage medium

    CN118410135A

  • Request statement processing method and device, storage medium and electronic equipment

    CN119558320A

Cited By

  • Input system, device and method for artificial intelligence generation content of intelligent cabin and storage medium

    CN121075321A