Human-computer interaction method, apparatus and system, device, vehicle, medium, and program product
By introducing systems of intent identification module, task processing module, service module and generative UI module into intelligent devices, the problem that the existing intelligent interaction mode is restricted by prefabricated programs is solved, and efficient and intelligent human-computer interaction is achieved.
Patent Information
- Application Number
- PCT/CN2024/140123
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-19
- Filing Date
- 2024-12-17
- Publication Date
- 2025-06-26
AI Technical Summary
The existing intelligent interaction method is limited by prefabricated programs when processing user demand instructions, and cannot respond to unprefabricated instructions. It has low interaction efficiency and requires multiple human-computer interactions to complete a single task.
The human-computer interaction system is adopted, including an intention identification module, a task processing module, a service module and a generative UI module. By identifying the integrated information, the task to be executed is disassembled as the target subtask, the task execution service execution subtask is called, and the task results are output.
It realizes intelligent human-computer interaction without being restricted by prefabricated programs, improves interaction efficiency, and completes the interactive process through mutual calls between modules.
Smart Images

Figure CN2024140123_26062025_PF_FP_ABST
Abstract
Description
Human-computer interaction method, device, system, equipment, vehicle, medium and program product
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This disclosure claims that the Chinese patent application number is 2023117553896 filed on December 19, 2023, and the applicant is Beijing Rockwell Technology Co., Ltd., and the application name is “Human-computer interaction method, device, system and vehicle”; the Chinese patent application number is 2023117561712 filed on December 19, 2023, and the applicant is Beijing Rockwell Technology Co., Ltd., and the application name is “Human-computer interaction method and device, vehicle, electronic device and storage medium”; The Chinese patent application number is 2023117564528, the applicant is Beijing Rockwell Technology Co., Ltd., and the application name is "Human-computer interaction display method and device, system, vehicle, electronic device", and the priority is Chinese patent application number 2023117574214 filed on December 19, 2023, the applicant is Beijing Rockwell Technology Co., Ltd., and the application name is "Vehicle interaction method and device, vehicle, electronic device and storage medium". The full text of the above application is incorporated into this disclosure by reference. Technical Field
[0003] The present disclosure relates to the field of vehicle technology, and relates to, but is not limited to, a human-computer interaction method, apparatus, system, device, vehicle, storage medium, and program product. Background Art
[0004] Human-computer interaction, in simple terms, is the process by which humans and smart devices exchange information through certain interactions. This interaction directly impacts the user experience. For example, in the automotive sector, interaction methods have evolved from simple physical knobs to intelligent interactions like touchscreens.
[0005] With the current intelligent interaction method, smart devices can only respond to pre-made programs for user demand instructions. For example, the temperature adjustment function has a pre-made corresponding program. When the demand instruction is to adjust the temperature to 23 degrees, the smart device can execute it. When the demand instruction does not correspond to a pre-made program, although the intention is clear, the smart device cannot respond due to the limitations of the pre-made program, resulting in a low level of intelligent human-computer interaction. In addition, the execution of each demand instruction requires a deep touch screen interface layer and multiple human-computer interactions to complete the execution, resulting in low efficiency of human-computer interaction. Summary of the Invention
[0006] The present disclosure provides a human-computer interaction method, apparatus, system, device, vehicle, storage medium, and program product, the main purpose of which is to improve the efficiency of human-computer interaction.
[0007] The present disclosure first provides a human-computer interaction system, comprising: an intention recognition module, a task processing module, a service module, and a generative UI module;
[0008] The intention recognition module recognizes the fused information in the received interaction request, obtains a to-be-executed task corresponding to the interaction request, and decomposes the to-be-executed task into at least one target subtask, wherein the fused information includes at least one interactive input form;
[0009] The task processing module sends task request information for scheduling a task execution service corresponding to the at least one target subtask to the service module;
[0010] The service module responds to the task request information, searches for a task execution service of at least one target subtask, and controls each task execution service to execute the corresponding target subtask to obtain a task execution result;
[0011] The generative UI module outputs and displays the task execution result corresponding to the at least one target subtask.
[0012] The present disclosure also provides a human-computer interaction method, comprising:
[0013] Identifying fused information in a received interaction request to obtain a task to be executed corresponding to the interaction request, wherein the fused information includes at least one interactive input form;
[0014] Decomposing the task to be performed into at least one target subtask;
[0015] Searching for a task execution service corresponding to each of the at least one target subtask, and calling the task execution service to execute the corresponding target subtask to obtain a task execution result;
[0016] Output and display the task execution result of the at least one target subtask.
[0017] The present disclosure also provides a human-computer interaction device, comprising:
[0018] an identification unit configured to identify fused information in a received interaction request to obtain a to-be-executed task corresponding to the interaction request, wherein the fused information includes at least one interactive input form;
[0019] a disassembly unit, configured to disassemble the task to be executed into at least one target subtask;
[0020] a search unit configured to search for a task execution service corresponding to each of the at least one target subtask;
[0021] A calling unit configured to call the task execution service to execute the corresponding target subtask and obtain the target subtask execution result;
[0022] The output unit is configured to output and display the task execution result of the at least one target subtask.
[0023] The present disclosure also provides a vehicle, comprising any of the aforementioned devices or any of the aforementioned systems.
[0024] The present disclosure also provides a human-computer interaction device, the device comprising:
[0025] at least one processor;
[0026] The storage device is configured to store at least one program, and when the at least one program is executed by the at least one processor, the device implements the human-computer interaction method as described above.
[0027] The present disclosure also provides a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to enable the computer to execute the human-computer interaction method as described above.
[0028] The present disclosure also provides a computer program product, including a computer program, which implements the human-computer interaction method as described above when executed by a processor.
[0029] The human-computer interaction system provided by the present disclosure includes: an intention recognition module, a task processing module, a service module and a generative UI module; the intention recognition module recognizes the fused information in the received interaction request, obtains the to-be-executed task corresponding to the interaction request, and decomposes the to-be-executed task into at least one target subtask, and the fused information includes at least one interactive input form; the task processing module sends task request information for scheduling the task execution service corresponding to the at least one target subtask to the service module; the service module responds to the task request information, searches for the task execution service of at least one target subtask, and controls each of the task execution services to execute the corresponding target subtask to obtain a task execution result; the generative UI module outputs and displays the task execution result corresponding to the at least one target subtask. Compared with the related art, the present invention identifies the fused information in the received human-computer interaction request through the intention recognition module, obtains the to-be-executed task corresponding to the interaction request, and further decomposes the to-be-executed task corresponding to the interaction request into at least one target subtask, queries and calls the task execution service corresponding to each target subtask to execute the corresponding target subtask, and obtains the task execution result; the entire execution process is not restricted by the program corresponding to the pre-made demand instruction, which greatly improves the intelligence of human-computer interaction; and after the user inputs the human-computer interaction request, the entire process of responding to the interaction request is completed by mutual calls between various modules, and the user does not need to intervene in the deeper touch screen interface level, making the entire interaction process automated and intelligent, greatly improving the efficiency of human-computer interaction.
[0030] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it intended to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] The accompanying drawings are provided to facilitate a better understanding of the present invention and do not constitute a limitation of the present disclosure.
[0032] FIG1 is a schematic structural diagram of a human-computer interaction system provided by an embodiment of the present disclosure;
[0033] FIG2 is a schematic diagram of an interface display provided by an embodiment of the present disclosure;
[0034] FIG3 is a structural diagram of a task processing module 2 provided in an embodiment of the present disclosure;
[0035] FIG4 is a schematic diagram of another human-computer interaction system provided by an embodiment of the present disclosure;
[0036] FIG5 is a schematic diagram of a generative user interface (UI) module 4 provided in an embodiment of the present disclosure;
[0037] FIG6 is a flow chart of a human-computer interaction method provided by an embodiment of the present disclosure;
[0038] FIG7A is a flow chart of another human-computer interaction method provided by an embodiment of the present disclosure;
[0039] FIG7B is a flow chart of another human-computer interaction method provided by an embodiment of the present disclosure;
[0040] FIG7C is a flowchart illustrating a task execution service provided by an embodiment of the present disclosure executing the corresponding target subtasks to obtain task execution results;
[0041] FIG7D is a schematic diagram of a flow chart of a method for displaying human-computer interaction provided in an embodiment of the present disclosure;
[0042] FIG7E is a flow chart of another method for displaying human-computer interaction provided by an embodiment of the present disclosure;
[0043] FIG7F is a schematic flow chart of a vehicle interaction method provided in an embodiment of the present disclosure;
[0044] FIG7G is a flow chart of a plug-in registration method provided in an embodiment of the present disclosure;
[0045] FIG8 is a schematic structural diagram of a human-computer interaction device provided by an embodiment of the present disclosure;
[0046] FIG9 is a schematic structural diagram of another human-computer interaction device provided by an embodiment of the present disclosure;
[0047] FIG10 is a schematic block diagram of an example human-computer interaction device provided by an embodiment of the present disclosure. DETAILED DESCRIPTION
[0048] The following description of exemplary embodiments of the present disclosure is made in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding. These details should be considered as merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.
[0049] The human-computer interaction method, system, and vehicle according to embodiments of the present disclosure are described below with reference to the accompanying drawings.
[0050] FIG1 is a schematic structural diagram of a human-computer interaction system provided by an embodiment of the present disclosure.
[0051] As shown in Figure 1, the system includes the following modules: intention recognition module 1, task processing module 2, service module 3 and generative UI module 4; the human-computer interaction system has the capabilities of perception, understanding, memory, execution, and decision-making power. Among them, the intention recognition module 1 provides perception and understanding capabilities, the task processing module 2 provides memory and decision-making power capabilities, the service module 3 provides execution capabilities, and the generative UI module 4 displays the direct results of the task for users to view.
[0052] The intention recognition module 1 identifies the fused information in the received interaction request, obtains the task to be executed corresponding to the interaction request, and decomposes the task to be executed into at least one target subtask and a corresponding attribute category. The fused information includes at least one form of interaction input; the user inputs the unstructured fused information into the intention recognition module, and the fused information includes but is not limited to at least one of image information, acoustic information, touch interaction, geographic location information, vehicle status information, environmental information (light information, etc.), etc. The fused information is used to trigger the interaction request during human-computer interaction between people and vehicles.
[0053] The intention recognition module 1 takes natural language as input and provides the ability to interact with the user in natural language dialogue. After receiving the interaction request, the intention recognition module 1 identifies the fused information in the interaction request through a pre-established recognition model, obtains the task to be executed corresponding to the interaction request, and decomposes the task to be executed to obtain at least one target subtask. The input form of the fused information includes at least one multimodal information of pictures, voice, gestures, videos, touch, and gaze. The embodiment of the present disclosure does not limit the input form of the fused information.
[0054] The pre-established recognition model described in the embodiment of the present disclosure is a diversified platform that can recognize any type of input information including image information, acoustic information, touch interaction, geographic location information, status information, environmental information (light information, etc.), that is, the user only needs to give the vehicle computer a target demand information, and the recognition model in the intention recognition module 1 will analyze and process it, understand the environment and task requirements according to the requirements corresponding to the target demand information, and convert the recognition results into tasks to be executed. Its purpose is to recognize the user intention as an executable task that the machine can recognize and execute, and to decompose the task to be executed into at least one target sub-task. Its purpose is to have at least one target sub-task jointly complete the user's needs, and send the at least one decomposed target sub-task to the task processing module 2.
[0055] As a possible implementation method, the pre-established recognition model may include, for example, an artificial intelligence (AI) large model. AI large models refer to high-performance artificial intelligence models built through huge training samples and computing resources. They can learn a large amount of language knowledge, image features and speech patterns, and can reason and generate outputs similar to humans. They have wide applications in natural language processing, image recognition, speech recognition and other fields. Artificial intelligence large models may include, for example, large language models (LLM), ChatGPT (Chat Generative Pre-trained Transformer), multimodal large models and multimodal cognitive large models. For example, it is possible to recognize fused information in interactive requests.
[0056] As a feasible method of an embodiment of the present disclosure, in addition to recognizing the user intention as a task to be executed through a pre-established recognition model, the intention recognition module 1 can also recognize the user intention through a recognition algorithm corresponding to the interactive input form. For example, when the interactive input form is image input, the recognition algorithm is an image recognition algorithm; when the interactive input form is voice input, the recognition algorithm is a sound recognition algorithm, and so on. For example, the embodiment of the present disclosure does not limit the processing method of the intention recognition module 1.
[0057] In some embodiments, the target subtask is described using a domain-specific language (DSL). The target subtask described using the DSL is then transmitted to the task processing module 2 for processing. The DSL protocol uses the DCC format and includes information about the domain, command, and content. For example, when a user inputs "I want to travel to city B with user A on X month X day," the corresponding command is search and the content is travel. The disclosed embodiments do not specifically limit the interactive content.
[0058] The task processing module 2 sends task request information for scheduling a task execution service corresponding to at least one target subtask to the service module; wherein each subtask corresponds to a different task execution service according to different attribute categories; the task execution service is implemented by calling the service module in the system, and the task execution service in the service module includes but is not limited to applications, plug-ins, and applications that require plug-in assistance to run, etc.
[0059] For example, when the user voice inputs "I want to travel to city B with user A on X month X day", the target subtasks that can be disassembled include but are not limited to purchasing subway tickets, booking air tickets and / or train tickets, booking hotels, searching for local food and / or attractions, booking tickets, etc. The disassembled target subtasks each correspond to an attribute type. The purpose is to send task request information for scheduling the corresponding task execution service to the service module according to the attribute category, so as to determine the corresponding task execution service, and the task execution service executes the corresponding target subtask.
[0060] The service module 3 responds to the task request information, searches for task execution services corresponding to at least one attribute category, and controls each task execution service to execute the target subtask corresponding to the task request information to obtain a task execution result.
[0061] The service module 3 described in the embodiment of the present disclosure provides computing power to the task processing module 2, expands the functions of the task processing module 2, and the auxiliary capabilities provided by the service module 3 help the task processing module 2 complete the interaction process. During the interaction process, it can help developers use various application programming interfaces (Application Programming Interface, API) more conveniently. For example, during the interaction process, you can describe the function or algorithm you want to implement to the task execution service, and the task execution service can provide a code example that uses a specific AI Framework to implement the function, thereby helping developers to complete the task assembly and orchestration of the interaction process more efficiently.
[0062] After the task execution service is started, at least one task execution service executes the corresponding target subtasks in parallel or serially.
[0063] The generative UI module 4 combines the contents to be displayed of the task execution results corresponding to the at least one task execution service, generates a user interface view including all task execution results, and displays the user interface view.
[0064] The generative UI module described in the embodiment of the present disclosure dynamically builds an interactive interface according to the task execution results output by the task processing module 2.
[0065] In actual applications, there are two types of task execution results of task execution services. One is to directly execute the target subtask (also called a single-round task). For example, when the target subtask is to open the car window, the task execution service is a car window driver plug-in, which directly controls the opening of the car window to complete the execution of this target subtask. This type of task execution result does not require the display of the task execution result; the other is a scenario where the task execution result needs to be displayed. For example, when the user plays music or navigates to a certain location, the music playback interface or the navigation interface needs to be displayed on the display screen.
[0066] For the obtained task execution results, the execution results to be displayed of all task execution services are combined, with the purpose of synchronously displaying the processing status of the user's intention to be executed by multiple tasks in parallel in the display interface for user viewing.
[0067] Exemplarily, the target subtasks include purchasing subway tickets, booking air tickets and / or train tickets, booking hotels, inquiring about local food and / or attractions, and booking tickets. The task execution services (5 in total) corresponding to each target subtask (5 in total) respectively execute subway ticket inquiries, air ticket and / or train ticket inquiries, hotel inquiries, local food and / or attractions inquiries, and ticket inquiries. The inquiries serve as the task execution results of the task execution services, and the queried content is the content to be displayed. The disclosed embodiment combines the content to be displayed, and the combination form includes but is not limited to displaying the content to be displayed of each task execution result side by side on the display desktop, or randomly displaying the content to be displayed of each task execution result, etc. Furthermore, the content to be displayed of each task execution result can also be displayed according to the display format of the priority setting application.
[0068] The embodiments of the present disclosure do not limit the method of combination. It should be noted that the combined content to be displayed needs to be displayed at the same display level. The purpose is to facilitate users to simultaneously view the content to be displayed of the task execution results corresponding to all task execution services.
[0069] For example, as shown in Figure 2, which is a schematic diagram of an interface display provided by an embodiment of the present disclosure, assuming that there are three task execution services and the corresponding user interface views of the three task execution services are loaded simultaneously. Figure 2 is merely an example. The interface display dynamically adjusts the user interface view based on the number of different task execution services and the task execution results of different target subtasks, rather than constructing a user interface UI view with a consistent layout every time.
[0070] The human-computer interaction system provided by the present disclosure includes: an intention recognition module 1, a task processing module 2, a service module and a generative UI module; the intention recognition module 1 recognizes the fused information in the received interaction request, obtains the to-be-executed task corresponding to the interaction request, and decomposes the to-be-executed task into at least one target subtask, and the fused information includes at least one interactive input form; the task processing module 2 sends task request information for scheduling the task execution service corresponding to the at least one target subtask to the service module; the service module 3 responds to the task request information, searches for the task execution service of at least one target subtask, and controls each of the task execution services to execute the corresponding target subtask to obtain a task execution result; the generative UI module 4 outputs and displays the task execution result corresponding to the at least one target subtask. Compared with the related art, the present invention identifies the fused information in the received human-computer interaction request through the intention recognition module, obtains the to-be-executed task corresponding to the interaction request, and further decomposes the to-be-executed task corresponding to the interaction request into at least one target subtask, queries and calls the task execution service corresponding to each target subtask to execute the corresponding target subtask, and obtains the task execution result; the entire execution process is not restricted by the program corresponding to the pre-made demand instruction, which greatly improves the intelligence of human-computer interaction; and after the user inputs the human-computer interaction request, the entire process of responding to the interaction request is completed by mutual calls between various modules, and the user does not need to intervene in the deeper touch screen interface level, making the entire interaction process automated and intelligent, greatly improving the efficiency of human-computer interaction.
[0071] As shown in FIG3 , FIG3 is a structural diagram of a human-computer interaction system provided by an embodiment of the present disclosure. The task processing module 2 includes: an engine unit 21 and a scheduling unit 22, wherein:
[0072] The engine unit 21 determines whether each target subtask needs to call the task execution service. The task execution service includes the target application and the target plug-in. After the intent module disassembles at least one target subtask based on the pre-established recognition model, it will call the Manifest description file: this file uses natural language to explain how to call the API, which API to call in different scenarios (such as the API of the task execution service corresponding to the target subtask), what parameters need to be passed in when calling the API, and what are the output parameters.
[0073] For example, the engine unit 21 may determine whether the target subtask needs to call the task execution service based on the Manifest description file. Since target subtasks are different, the results of determining whether to call the task execution service are also different.
[0074] When the scheduling unit 22 determines that the target application and / or target plug-in needs to be called, it sends task request information for scheduling the corresponding target application and / or target plug-in to the preset application service and / or preset plug-in service based on the attribute category of each target subtask. When the engine unit 21 recognizes that the system preset plug-in capability needs to be called, it sends the task request information to the plug-in proxy in the preset plug-in service, which responds to the task request information sent by the scheduling unit 22 to schedule the target plug-in.
[0075] The Preset Plugin Service also provides the ability to manage the lifecycle of each plugin. Considering that plugins in the entire software ecosystem come from different developers, including but not limited to generative UI plugins, vertical data plugins, and third-party plugins.
[0076] Please refer to Figure 4. The service module 3 includes a receiving unit 31, a search engine 32, and a processing unit 33. The receiving unit 31 receives the task request information of the target application and / or target plug-in corresponding to at least one target subtask sent by the scheduling unit; the preset plug-in service receives the task request information sent by the engine unit 21 through the plug-in proxy Plugin Proxy.
[0077] The search engine 32 queries the corresponding target application and / or target plug-in according to the attribute category of the at least one target subtask; in the preset application service, the search engine 32 can query the corresponding target application according to the Manifest description file based on the attribute category of the at least one target subtask.
[0078] In the pre-configured plug-in service, the plug-in proxy uses the PME (Proxy Match Engine) to dispatch the target subtask to the corresponding target plug-in. The PME engine records the address information of all plug-ins in the pre-configured plug-in service. After determining the target plug-in to call, the PME engine can query the target plug-in corresponding to the target subtask. The plug-in proxy performs logic execution through remote calls with the plug-in and the business platform.
[0079] The processing unit 33 controls the target application and / or target plug-in to execute the corresponding target subtask, and obtains the task execution result of the target application and / or target plug-in.
[0080] As shown in FIG5 , FIG5 is a schematic diagram of a generative UI module 4 provided in an embodiment of the present disclosure, wherein the generative UI module 4 includes: a parser 41 , a loader 42 , a layout container 43 and a rendering engine 44 , wherein:
[0081] The parser 41 parses the interactive interface layout description information to determine the dynamic layout information of all target applications; the interactive interface layout description information is generated according to the number and application category of the target applications, and is used to describe the layout information of the corresponding target applications in the interactive interface; the target application calls the data function of the target plug-in.
[0082] The interactive interface layout description information is generated after the task processing module 2 decomposes the task to be executed into at least one target subtask and determines to schedule the target subtask to the corresponding target application. The interactive interface layout description information uses DSL to describe the interactive interface layout rules. After determining the target application, the number of target applications and service categories are counted. The purpose is to configure the layout description information based on the number of target applications and service categories. For ease of understanding, each target application can be understood as a card. Each card has position information, size information, style information, etc. on the interactive interface. Configuring the layout description information is to configure the dynamic layout information of all target applications (cards) on the display interface. For example, the position information, size information, and style information (collectively referred to as dynamic layout information) corresponding to different service categories may be different. Different numbers of task execution services may also have different layouts in the display interface. Therefore, due to the different target applications involved in each interaction or the different target plug-ins called, the constructed user interface UI view may be different. Therefore, the constructed UI view is a dynamically changing one.
[0083] In practical applications, since the generation and parsing of layout description information are transmitted in two different modules, secure data transmission can be achieved by formulating security standards for data transmission, such as encryption and authentication.
[0084] The loader 42 loads at least one corresponding target plug-in according to at least one target subtask; the at least one target plug-in executes the corresponding target subtask respectively to obtain the task execution result; when the loader executes the loading of the target plug-in, if the target plug-in Plugin has been loaded, the data and service proxy Proxy of the target plug-in are directly provided to the task processing module 2; if the target plug-in Plugin has not been loaded, the target plug-in Plugin needs to be registered first, and after the registration is completed, it is loaded and then the data and service proxy Proxy are provided.
[0085] The layout container 43 combines the contents to be displayed of the task execution results corresponding to the at least one target plug-in to generate a user interface view including all task execution results.
[0086] In some embodiments, when generating a user interface view containing all task execution results, it can be implemented in but not limited to the following manner: based on the layout description information obtained by parsing the parser, dynamic layout information containing all target applications is determined, and the content to be displayed of the task execution results corresponding to the at least one target plug-in is combined according to the dynamic layout information to build and generate a user interface view containing all task execution results.
[0087] For ease of understanding, assuming that there are three task execution services, load layout description information containing the three task execution services in the layout container, determine the dynamic typesetting information containing all the task execution services based on the layout description information, and combine the content to be displayed of the task execution results corresponding to at least one task execution service according to the dynamic typesetting information, that is, display the content 1 to be displayed of the task execution result of task execution service 1 in the interface of task execution service 1, display the content 2 to be displayed of the task execution result of task execution service 2 in the interface of task execution service 2, and display the content 3 to be displayed of the task execution result of task execution service 3 in the interface of task execution service 3, so as to build and generate a user interface view containing all task execution results.
[0088] The rendering engine 44 renders and displays the user interface view. It can be understood that the task execution service is a node tree, and the rendering engine dynamically constructs the UI view based on the task execution service and presents the UI view through the generative UI container.
[0089] Furthermore, the generative UI module 4 includes an event processing engine 45, which receives interaction events in the user interface view and sends them to the task processing module 2. In response to user interaction events, user clicks, slides, drags, and other operation actions are called back. The interaction events are processed by the event processing engine, and the generative UI service event processing engine passes the interaction task to the scheduling unit of the task processing module 2 for scheduling, completing the closed loop of the interaction link.
[0090] The human-computer interaction system provided by the disclosed embodiments uses natural language as input and output, enabling natural language conversational interaction with users. However, in smart cockpit applications, it is also necessary for the task processing module 2 to be able to interact with users by invoking service modules and constructing a graphical interface. This allows the task processing module 2 to efficiently present information, accurately obtain user feedback, and improve the efficiency of human-computer interaction.
[0091] Corresponding to the above-mentioned human-computer interaction system, the present disclosure also proposes a human-computer interaction method.
[0092] FIG6 is a flow chart of a human-computer interaction method provided in an embodiment of the present disclosure. As shown in FIG6 , the method can be applied to the human-computer interaction system described in the above embodiment, including:
[0093] Step 101: Identify fused information in a received interaction request to obtain a task to be executed corresponding to the interaction request, wherein the fused information includes at least one interactive input form.
[0094] The user inputs unstructured fused information into the human-computer interaction system. The fused information includes but is not limited to at least one of image information, acoustic information, touch interaction, geographic location information, vehicle status information, environmental information (light information, etc.), etc. The fused information is used to trigger the interaction request between the human and the vehicle during human-computer interaction.
[0095] The human-computer interaction system described in the embodiment of the present disclosure uses natural language as input and provides the ability to interact with users in natural language dialogue. After receiving an interaction request, the human-computer interaction system identifies the fused information in the interaction request through a pre-established recognition model, obtains the task to be executed corresponding to the interaction request, and decomposes the task to be executed to obtain at least one target subtask. The input form of the fused information includes at least one multimodal information of pictures, voice, gestures, videos, touch, and gaze. The embodiment of the present disclosure does not limit the input form of the fused information.
[0096] The pre-established recognition model described in the embodiment of the present disclosure is a diversified platform that can recognize any input information including image information, acoustic information, touch interaction, geographic location information, status information, environmental information (light information, etc.), that is, the user only needs to give the vehicle computer a target demand information, and the recognition model in the intention recognition module 1 will analyze and process it, understand the environment and task requirements according to the requirements corresponding to the target demand information, and convert the recognition results into tasks to be executed. Its purpose is to recognize the user intention as an executable task that the machine can recognize and execute, and to decompose the task to be executed into at least one target sub-task. Its purpose is to have at least one target sub-task jointly complete the user needs.
[0097] In some embodiments, the task to be executed is described using the DSL protocol, and the task to be executed described using the DSL protocol is transmitted to the next node for corresponding processing. The DSL protocol is in the DCC format, which includes the domain, command, and content information.
[0098] Step 102: Decompose the task to be executed into at least one target subtask.
[0099] Continue to decompose the task to be performed into at least one target subtask based on the pre-established recognition model described in step 101.
[0100] For example, when the user voice inputs "I want to travel to city B with user A on X month X day", the target subtasks that can be disassembled include but are not limited to purchasing subway tickets, booking air tickets and / or train tickets, booking hotels, searching for local food and / or attractions, booking tickets, etc.; for example, the embodiment of the present disclosure does not limit the disassembly of target subtasks, which varies according to the differences in the interaction requests.
[0101] Step 103: Find the task execution service corresponding to each of the at least one target subtask, and call the task execution service to execute the corresponding target subtask to obtain the target subtask execution result.
[0102] The auxiliary capabilities provided by the Task Execution Service help target subtasks complete the interactive process, making it easier for developers to use various APIs. For example, during the interaction, you can describe the desired function or algorithm to the Task Execution Service, and the Task Execution Service will provide code examples that implement the function using a specific AI Framework, helping developers to more efficiently assemble and orchestrate tasks during the interactive process.
[0103] As a feasible method of the embodiment of the present disclosure, when searching for the task execution service corresponding to each of the at least one target subtask, the reasoning unit in the human-computer interaction system can be used to determine what the task execution service executed by at least one target subtask is, and used to autonomously execute the corresponding control strategy and make decisions on behalf of the user. The reasoning engine can be a reasoning engine developed by the human-computer interaction system itself, or it can be any reasoning engine used in the existing technology, and the embodiment of the present disclosure does not limit it.
[0104] After the task execution service is started, at least one task execution service executes each corresponding target subtask in parallel or serially to obtain a task execution result. The task execution service described in the disclosed embodiment can be an application corresponding to the target subtask, or a plug-in corresponding to the target subtask. It can also be an application corresponding to the target subtask, but the application requires data support from the plug-in. The disclosed embodiment does not limit the task execution service.
[0105] Step 104: Output and display the task execution result of the at least one target subtask.
[0106] In actual applications, there are two types of task execution results of task execution services. One is to directly execute the target subtask (also called a single-round task). For example, when the target subtask is to open the car window, the task execution service is a car window driver plug-in, which directly controls the opening of the car window to complete the execution of this target subtask. This type of task execution result does not require the display of the task execution result; the other is a scenario where the task execution result needs to be displayed. For example, when the user plays music or navigates to a certain location, the music playback interface or the navigation interface needs to be displayed on the display screen, and the interaction continues based on the displayed content (also called a multi-round task).
[0107] For the obtained task execution results, the execution results of all task execution services to be displayed are output and displayed. The purpose is to synchronously display the processing status of the user's intention to be executed by multiple tasks in parallel in the display interface for user viewing.
[0108] Exemplarily, the target subtasks include purchasing subway tickets, booking air tickets and / or train tickets, booking hotels, inquiring about local food and / or attractions, and booking tickets. The task execution services (5 in total) corresponding to each target subtask (5 in total) respectively execute subway ticket inquiries, air ticket and / or train ticket inquiries, hotel inquiries, local food and / or attractions inquiries, and ticket inquiries. The query results serve as the task execution results of the task execution service, and the queried content is the content to be displayed. The embodiment of the present disclosure displays the task execution results of the target subtasks, and the display format includes but is not limited to displaying each task execution result side by side on the display desktop, or randomly displaying the content to be displayed of each task execution result, etc. Furthermore, each task execution result can also be displayed according to the pre-set display format of the task execution service. The embodiment of the present disclosure does not limit the display method of the task execution result of the task execution service.
[0109] The embodiment of the present disclosure does not limit the display method. It should be noted that the task execution results need to be displayed at the same display level. The purpose is to facilitate users to view the task execution results corresponding to all task execution services at the same time.
[0110] The human-computer interaction method provided by the present disclosure identifies the fusion information in the received interaction request, obtains the task to be executed corresponding to the interaction request, and the fusion information includes at least one interactive input form; decomposes the task to be executed into at least one target subtask; searches for the task execution service corresponding to each of the at least one target subtask, and calls the task execution service to execute the corresponding target subtask to obtain the target subtask execution result; and outputs and displays the task execution result of the at least one target subtask. Compared with the related art, the present disclosure identifies the fusion information in the received human-computer interaction request, obtains the task to be executed corresponding to the interaction request, and further decomposes the task to be executed corresponding to the interaction request into at least one target subtask, queries and calls the task execution service corresponding to each target subtask to execute the corresponding target subtask, and obtains the task execution result; the entire execution process is not restricted by the corresponding program of the prefabricated demand instruction, which greatly improves the intelligence of human-computer interaction; and after the user inputs the human-computer interaction request, the entire process of responding to the interaction request is completed by mutual calls between the various modules, without the user needing to intervene in the deeper touch screen interface layer, making the entire interaction process automated and intelligent, greatly improving the efficiency of human-computer interaction.
[0111] The task execution service described in the embodiment of the present disclosure includes a target plug-in, a target application, and a target application that calls the target plug-in. The following describes the task execution service according to different categories.
[0112] In some embodiments, when the task execution service is a target plug-in, the target plug-in corresponding to the at least one target subtask is searched in the preset plug-in service according to the attribute category of the at least one target subtask, and the at least one target plug-in executes the corresponding target subtask to obtain the task execution result.
[0113] As an implementation method of an embodiment of the present disclosure, when the service module queries the target plug-in corresponding to at least one target subtask, the query can be performed through the Manifest file. The Manifest file records different plug-ins divided according to attribute categories. The essence of calling the plug-in is to call the plug-in through the API, that is, to determine which API to call in different scenarios (under different attribute categories), the parameter information required to call the API, and the output parameter information.
[0114] During the specific application process, the service module obtains the interface address of at least one target plug-in based on the correspondence between the attribute category and the plug-in, and searches for the corresponding target plug-in based on each interface address. It should be noted that the correspondence between tasks and plug-ins is not a strong mapping relationship. The correspondence between tasks and plug-ins includes human-oriented prompts, such as informing users of the plug-in and its purpose, to facilitate user understanding. There are also prompts for the processing engine, which can help the processing engine understand the purpose of the plug-in, etc.
[0115] After the target plug-in is started, the target plug-in executes the target subtasks corresponding to the task request information in parallel or serially.
[0116] As another implementation method of the embodiment of the present disclosure, a generative pre-trained transformer (GPT) is called in a preset plug-in service to search for a target plug-in corresponding to at least one target subtask; and according to the correspondence between tasks and plug-ins in the generative pre-trained model, the target plug-ins corresponding to the target subtask are searched respectively.
[0117] The embodiment of the present disclosure does not specifically limit the process of searching for the target plug-in.
[0118] In some embodiments, when the task execution service is a target application, a target application corresponding to the at least one target subtask is searched in a preset application service, and the at least one target application executes the corresponding target subtask respectively to obtain a task execution result.
[0119] Pre-configured application services, also known as intelligent co-pilots, provide AI assistant services to humans. Through natural language interaction, these services intelligently acquire certain assistance capabilities, thereby improving productivity. In traditional software, humans independently complete tasks based on pre-set program rules through commands, keystrokes, and touch, with virtually no AI assistance. In contrast, with pre-configured application services, AI becomes involved in tasks, requiring humans to interact with the AI, obtain information, and collaborate to complete tasks. Currently, Windows, Microsoft 365, GitHub, Bing, and others have implemented pre-configured application services. This model significantly changes human-machine collaboration and improves task processing efficiency.
[0120] The preset application service provides an interactive entry point, takes text, images, voice, and other integrated information as input, calls functions in the AI Framework API library to use the preset application service, and uses the preset application service of the AI Framework to implement session processing, operation orchestration, and data auditing capabilities, and obtain data access capabilities through the domain.
[0121] After obtaining at least one target subtask based on the pre-established recognition model, the manifest file will be called: this file uses natural language to explain how to call the API, which API to call in different scenarios (such as the API of the task execution service corresponding to the target subtask), what parameters need to be passed in when calling the API, and what are the output parameters.
[0122] The Preset Application Service introduces an AI Framework API library, which helps developers more conveniently use various AI Frameworks. For example, during an interaction, you can describe the desired function or algorithm to the Preset Application Service, which will then provide code examples that implement this function using a specific AI Framework, helping developers complete the interaction more efficiently.
[0123] The target application found by the query executes the corresponding target subtask. The result returned by the target application may be code, text, audio or video files. The target application needs to display these results in the UI and provide feedback to the user.
[0124] In some embodiments, when the task execution service includes a target application and the target application calls the data function of the target plug-in, the target application corresponding to the at least one target subtask is searched for in the preset application service, and the target plug-in corresponding to the at least one target subtask is loaded to provide the data function for the target application. The at least one target plug-in executes the corresponding target subtask to obtain the task execution result.
[0125] After determining at least one target application corresponding to the target subtask, and loading a target plug-in corresponding to at least one target application, the target plug-in is used to execute the corresponding target subtask. The target plug-in can be provided by the application, the system, or an independent service. As long as the module that provides atomic capabilities follows the standard plug-in protocol, it can be regarded as a plug-in. Each target plug-in triggers a service with the target subtask processing capability, executes the target subtask in parallel, and obtains the task execution result. The call of the target plug-in can improve the match with the user's intention and improve the user experience of the vehicle system in various interactive scenarios.
[0126] The present disclosure also provides a human-computer interaction method, as shown in FIG7A , including:
[0127] Step 201 : Identify the fused information in the received interaction request to obtain a task to be executed corresponding to the interaction request, wherein the fused information includes at least one interactive input form.
[0128] Step 202: Decompose the task to be executed into at least one target subtask.
[0129] Step 203: Find the task execution service corresponding to each of the at least one target subtask.
[0130] For the description of steps 201 to 203 , please refer to the detailed description of FIG6 , which is illustrative and will not be described in detail in the embodiment of the present disclosure.
[0131] Step 204: Configure execution parameters of at least one target subtask according to the target requirements in the interaction request.
[0132] For example, in the above embodiment, if a user inputs "I want to travel to City B with User A on X month X day," and the target subtasks include, but are not limited to, purchasing subway tickets, booking air and / or train tickets, booking a hotel, researching local food and / or attractions, and booking tickets, then the corresponding execution parameters for booking air and / or train tickets are X month X day, destination City B, and User A's identity information; the execution parameter for booking a hotel is City B, and so on. The above example only obtains the current environment information based on the user input. However, tasks such as booking a hotel, researching local food and / or attractions, and booking tickets need to be configured based on the user's preferences to better meet the user's needs.
[0133] Configuring the execution parameters of at least one target subtask separately specifically includes: obtaining current environmental information, as well as first historical data (short-term memory) of a first time period and second historical data (long-term memory) of a second time period; wherein, the second time period is longer than the first time period; illustratively, the second time period includes historical data of the past year and the past six months, and the first time period includes historical data of the past month and the past two months; configuring the execution parameters of the at least one target subtask according to a pre-established context learning model, the context learning model is used to configure the execution parameters by learning the current environmental information, the first historical data of the first time period and the second historical data of the second time period. The pre-established context learning model is obtained by training the first historical data (short-term memory) and the second historical data (long-term memory) of the second time period.
[0134] Step 205: Call each of the task execution services to execute the target subtask corresponding to the task request information according to the corresponding execution parameters.
[0135] The execution parameters are execution parameters that meet user needs. The task execution results obtained based on the execution parameters can be closer to or meet the actual needs of the user, thereby improving the user experience.
[0136] Step 206: Output and display the task execution result of the at least one target subtask.
[0137] As a feasible solution of the embodiment of the present disclosure, after the task to be executed is decomposed into at least one target subtask and the target subtask is determined to be scheduled to the corresponding task execution service, layout description information is generated. The layout description information uses a specific layout language (DSL) to describe the interface layout rules, which is usually output by understanding user behavior and intention. After the task execution service is determined, the number of task execution services and service categories are counted. The purpose is to configure the layout description information according to the number and service category of task execution services. For ease of understanding, each task execution service application can be understood as a card. Each card has position information, size information, style information, etc. in the display area. Configuring the layout description information is to configure the dynamic layout information of all task execution services (cards) on the display interface. For example, the position information, size information, and style information (collectively referred to as dynamic layout information) corresponding to different service categories may be different. There may also be differences in the layout of the display interface for different numbers of task execution services. Therefore, due to the different target applications involved in each interaction or the different target plug-ins called, the user interface UI view constructed may be different. Therefore, the constructed UI view is a dynamically changing one.
[0138] In some embodiments, when generating a user interface view containing all task execution results, it can be implemented in but not limited to the following manner: based on the layout description information, dynamic layout information containing all target applications is determined, and the task execution results corresponding to the at least one target plug-in are displayed according to the dynamic layout information to generate a user interface UI view containing all task execution results.
[0139] As another feasible solution of the embodiment of the present disclosure, the task execution result of at least one target subtask can also be displayed through text, video, sound, vibration, digital human animation, etc. For example, the embodiment of the present disclosure does not limit the display effect. However, the preferred method is to display the task execution results of all target subtasks for user viewing.
[0140] Step 207: respond to the interaction event triggered in the user interface view and dispatch the interaction event to the corresponding application or plug-in for execution.
[0141] In response to user interaction events, user clicks, slides, drags and other operation behaviors are called back, and the interaction events are processed by the event processing engine. The generative UI service event processing engine passes the interaction task to the scheduling unit of the task processing module 2 for execution, completing the closed loop of the interaction link.
[0142] This system framework is a pre-built tool library and software collection for designing, training, and validating AI models for human-computer interaction systems. It provides high-level APIs for programming languages (such as Python or JavaScript) to create and operate system atomics, integrating a series of algorithms and function libraries to accelerate model training and deployment. Common system frameworks include, but are not limited to, the following:
[0143] TensorFlow: Applied in deep learning and machine learning.
[0144] PyTorch: An open-source framework developed by Facebook's AI research team that provides a large number of tools and libraries for building and training neural networks.
[0145] Keras: A high-level neural network API that can use TensorFlow, CNTK, or Theano as a backend.
[0146] Caffe: is a fast, open-source deep learning framework particularly suitable for image classification and convolutional models, including real-world products and hardware.
[0147] MXNet: is a deep learning framework optimized for efficiency, flexibility, and portability.
[0148] In addition, the AI framework provides powerful functions for AI developers, allowing developers to focus more on model design and optimization rather than underlying computing details.
[0149] In some embodiments, the interactive interface layout description information may include layout parameters of at least one target subtask execution result information, and the outputting and displaying of the task execution result of the at least one target subtask may include: laying out the task execution result of each target subtask according to the layout parameters of the at least one target subtask execution result information, generating a user interface view containing all target subtask execution results; rendering the user interface view, and displaying the rendered user interface view.
[0150] Here, the layout parameters are used to determine the display interface of the execution result when displaying it to the user, such as interface size, transparency, display position, etc. The above are only illustrative examples, and the embodiments of the present disclosure do not limit the layout parameters.
[0151] It should be noted that different target subtasks correspond to different layout parameters. These layout parameters can be set in advance based on the output of the target subtask. For example, if the output of the target subtask may contain a lot of text, the interface size in the layout parameters of the target subtask can be set larger to reduce the user's reading pressure. The above is merely an illustrative example, and the present disclosure does not limit the setting of layout parameters.
[0152] Here, the layout parameters may include the aforementioned layout information. For each acquired task execution result of a target subtask, the execution result is laid out and typeset according to the corresponding layout parameters, generating a user interface view containing all target subtask execution results. Furthermore, the task execution results of each target subtask are displayed through the user interface view for the user to review, select, etc.
[0153] It should be noted that the user interface views need to be displayed at the same display level. The purpose is to facilitate users to display the task execution results of all target subtasks at the same time.
[0154] In some embodiments, the task execution result of each target subtask is laid out and typeset according to the layout parameters of the at least one target subtask execution result information to generate a user interface view containing all target subtask execution results, which may include: parsing the layout description information, and obtaining the layout parameters corresponding to each target subtask execution result in the layout description information; the layout parameters include attribute information of the target subtask execution result, multiple display position information and display position parameter information corresponding to each display position information; wherein the attribute information includes identification information representing the importance of the target subtask execution result: according to the attribute information of the target subtask execution result, the multiple display position information and the display position parameter information corresponding to each display position information, the task execution result of at least one target subtask is laid out and typeset to generate a user interface view containing all target subtask execution results.
[0155] Exemplarily, the layout description information can be parsed by a parser and converted into a data structure that can be processed by the program. In one feasible method of the embodiment of the present disclosure, the layout description information is saved through DSL, and a corresponding DSL parser can be used to parse the layout description information, such as a generative UI engine, and the layout parameters of the execution results of each target subtask are obtained based on the parsing results.
[0156] In order to make the layout of the task execution results more in line with the user's usage habits, the layout parameters described in the embodiment of the present disclosure include attribute information of the target subtask execution results, multiple display position information and display position parameter information corresponding to each display position information; the attribute information includes identification information that characterizes the importance of the target subtask execution results.
[0157] For ease of understanding, assuming that when the target subtasks include three, the corresponding task execution results also include three, namely task execution result 1, task execution result 2 and task execution result 3. The layout description information describes the attribute information corresponding to task execution result 1, task execution result 2 and task execution result 3 respectively. For example, the attribute information can be presented through three different weights, or through multiple rounds of clicks on the task execution results. Specifically, the embodiment of the present disclosure does not limit the presentation form of the attribute information of the task execution results.
[0158] In addition, the layout description information also includes multiple display position information and display position parameter information corresponding to each display position information, wherein the display position information usually corresponds to the number of task execution results, and records the display position of each task execution result in the display interface, and the display position parameter information includes but is not limited to interface size, transparency, display style, display size, etc. Specifically, the embodiment of the present disclosure does not limit the layout description information.
[0159] In the embodiment of the present disclosure, the task execution service not only provides data and functions to the generative UI interactive interface, but also provides a dynamic card for the task execution result of the target subtask. The task execution result of each target subtask corresponds to a card, which is displayed in the form of a card. The corresponding card information varies with the type of the target subtask. For example, when the target subtask is a video task, the corresponding card information corresponds to the video display form, including but not limited to the display position, card size, video content, duration, video layout, etc. When the target subtask is a text task, the corresponding card information corresponds to the text display form, including but not limited to the size, layout, color, font, etc. of the text. Specifically, the specific content of the card information in the embodiment of the present disclosure is not limited.
[0160] After determining the layout parameters of the card, the execution results of each target subtask are laid out according to the layout parameters of the at least one card to generate a user interface view containing the execution results of all target subtasks. Specifically, the embodiment of the present disclosure does not limit the implementation method of generating the user interface view.
[0161] As an implementable method of an embodiment of the present disclosure, displaying the user interface view may include: determining whether the target subtask contained in the user interface view is a task completed by the scheduling task execution service; when the target subtask is completed by the scheduling task execution service, displaying the user interface view in a preset user interface view generation area, and the preset user interface view generation area includes a desktop generation area and an application generation area.
[0162] In the embodiment of the present disclosure, the preset user interface view generation area is a configurable generation area, but the preset display position in the desktop or application needs to be preset in advance. The embodiment of the present disclosure does not limit the specific location of the desktop generation area and the application generation area.
[0163] In some embodiments, calling the task execution service of each target subtask to execute the corresponding target subtask and obtain the task execution result of the target subtask can include: determining the task execution service corresponding to each target subtask based on the pre-established correspondence between the subtask and the task execution service; scheduling the at least one target subtask to the corresponding task execution service respectively, and having each target task execution service execute the corresponding target subtask to obtain the task execution result of the target subtask.
[0164] According to the pre-established correspondence between subtasks and task execution services, the task execution services corresponding to the target subtasks are obtained respectively, and the at least one target subtask is dispatched to the corresponding task execution service respectively, and each target task execution service executes the corresponding target subtask. It should be noted that the pre-established correspondence between subtasks and task execution services is not a strong mapping relationship. The pre-established correspondence between subtasks and task execution services has human-oriented prompts, such as informing the user what this task execution service is and what it is used for, so that the user can understand it easily. There are also prompts for the task execution service, which can enable the task execution service to understand this subtask, etc.
[0165] As an implementation method of the embodiment of the present disclosure, the task execution service for executing at least one target subtask can be determined through the reasoning unit in the human-computer interaction system, and used to autonomously execute the corresponding control strategy and make decisions on behalf of the user. The reasoning engine can be a reasoning engine developed by the human-computer interaction system itself, or it can be any reasoning engine used in the existing technology. Specifically, the embodiment of the present disclosure does not limit it.
[0166] After identifying the target requirement according to a pre-established preset recognition model and obtaining the to-be-executed task associated with the target requirement, the to-be-executed task is disassembled and processed to obtain at least one target subtask; the reasoning unit determines the task execution service corresponding to each of the target subtasks and the corresponding service type, and the service type is used to determine the task execution service that needs to be called by each target subtask.
[0167] In the embodiment of the present disclosure, two service categories are included. One is the application type, and the corresponding task execution service includes the application service; the other is the plug-in type, and the task execution service includes the plug-in service.
[0168] When the service type includes an application type and the task execution service includes an application service, the target subtasks are dispatched to their respective corresponding application services according to the application type, and each application service executes its respective corresponding target subtask.
[0169] The application service is stored in the preset application service, which is also called the intelligent co-pilot. It provides AI assistant services for humans. The preset application service obtains certain auxiliary capabilities intelligently through natural language interaction, thereby improving production efficiency.
[0170] In some embodiments, outputting and displaying the task execution result of the at least one target subtask includes: combining the task execution results of the at least one target subtask in a UI; and outputting and displaying the combined UI.
[0171] In order to facilitate users in viewing the task execution results of the target task execution service, the task execution results of at least one target subtask are combined in a UI, and the combined UI is output and displayed. Before the target task execution service corresponding to at least one target subtask executes its corresponding target subtask, layout description information is generated based on the number of target task execution services and the task execution service category. The layout description information is used to describe the dynamic layout information of the corresponding target task execution service in the user interface view; after outputting the execution results of each target task execution service, the content to be displayed of the task execution results corresponding to all plug-in services is combined according to the dynamic layout information to generate a user interface view containing all task execution results, and the user interface UI view is output.
[0172] In some embodiments, in addition to being displayed through a UI view, the task execution results can also be displayed in the form of a human-computer interaction interface, which includes at least one of sound, vibration, digital human motion effects, task execution service controls, text images, and videos.
[0173] In some embodiments, searching for at least one target plug-in corresponding to the at least one target subtask in a preset plug-in service may include: calling a generative pre-trained model in the preset plug-in service to search for a target scene corresponding to the at least one target subtask; and searching for at least one target plug-in corresponding to the at least one target subtask based on the correspondence between the scene and the plug-in in the generative pre-trained model.
[0174] As an implementation method of an embodiment of the present disclosure, when executing the above steps to call at least one target plug-in corresponding to the at least one target subtask in the preset plug-in service, the following method can be adopted but not limited to: calling a generative pre-trained model (Generative Pre-trained Transformer, GPT) in the preset plug-in service to search for a target scene corresponding to the at least one target subtask; and then searching for at least one target plug-in corresponding to the at least one target subtask based on the correspondence between the scene and the plug-in in the generative pre-trained model.
[0175] The preset plug-in service includes different preset scenarios and their corresponding plug-ins, and different scenarios correspond to the same or different plug-ins. The embodiment of the present disclosure does not specifically limit the search process for the target plug-in.
[0176] In some embodiments, outputting and displaying the task execution result of the at least one target subtask may include: obtaining the input category of the fused information, the input category of the fused information including at least one of pictures, voice, gestures, videos, and gazes; and displaying the processing result of the at least one target plug-in in a display style that matches the input category.
[0177] In the embodiment of the present disclosure, the input category and its matching display style form an editable correspondence, and users can flexibly configure them according to their preferences.
[0178] For ease of understanding, when the input type of the fusion information is a gesture, the display styles that can be matched with the processing result of the target plug-in include but are not limited to voice and UI interface. Specifically, the embodiments of the present disclosure are not limited to this.
[0179] In some embodiments, outputting and displaying the task execution result of the at least one target subtask may include: parsing the processing result of the at least one target plug-in, determining the marking information carried in the processing result, and the processing result carries marking information of directly controlling the vehicle state / indirectly controlling the vehicle state; if it is determined that the marking information is directly controlling the vehicle state, outputting the execution result information corresponding to the controlled vehicle state; if it is determined that the marking information is indirect controlling the vehicle state, calling the user interface container, and constructing the user interface view corresponding to the processing result based on the user interface container; and outputting and displaying the user interface view.
[0180] To facilitate understanding of direct control of vehicle status / indirect control of vehicle status, the following example is used to illustrate. When the target subtask is to open the window, the target plug-in is the window driver plug-in. Directly controlling the opening of the window can complete the execution of this target subtask. The window driver plug-in directly controls the vehicle status and does not need to display the task execution results. The other is the scenario where the task execution results need to be displayed (indirect control of vehicle status). For example, when the user plays music or navigates to a certain location, the music playback interface or the navigation interface needs to be displayed on the display screen. The target plug-in does not directly control the vehicle status and needs to display the processing results.
[0181] For scenarios where the tag information does not directly control the vehicle state, the target plug-in's processing results are converted into corresponding layout description information. The layout description information is used to describe the processing results and layout attribute information of each target plug-in. The layout attribute information includes the target plug-in's identification information and display information. In actual applications, the layout description information can use a specific layout language (DSL language) to describe the display information of the interface layout. The display information includes but is not limited to the position information on the display screen or application, the display window size, display style, transparency, and other information. Specifically, the embodiments of this disclosure do not limit this.
[0182] FIG7B is a flow chart of another human-computer interaction method provided by an embodiment of the present disclosure.
[0183] As shown in FIG7B , the method comprises the following steps:
[0184] Step 401: Receive the user's target needs.
[0185] The user inputs unstructured fused information into the vehicle. This fused information includes, but is not limited to, at least one of image information, acoustic information, touch interaction, geographic location information, vehicle status information, and environmental information (such as light). This fused information is used to trigger interaction requests during human-machine interaction. The input form of this fused information includes at least one multimodal information such as image, voice, gesture, video, touch, and gaze. The present embodiment does not limit the input form of this fused information.
[0186] Step 402: Identify the target requirement according to a pre-established preset recognition model to obtain the to-be-executed task associated with the target requirement, wherein the to-be-executed task includes at least one target subtask, each subtask corresponding to a different task execution service; the preset recognition model is a model with natural language dialogue interaction capabilities, and converts the target requirement recognition result into an associated to-be-executed task by analyzing the environment and task requirements of the target requirement.
[0187] For example, when the user's target needs are understood to be traveling, reading, taking a nap, etc., it is necessary to first identify the user's target needs as executable tasks to be performed, and identify the target needs through a pre-established recognition model to obtain the tasks to be performed associated with the target needs. The pre-established recognition model described in the embodiment of the present disclosure is a diversified platform that can recognize any type of input information such as image information, acoustic information, touch interaction, geographic location information, vehicle status information, environmental information (light information, etc.), that is, the user only needs to provide an interactive target need to the human-computer interaction system, and the pre-established recognition model analyzes and processes the target need, understands the environment and task requirements according to the needs corresponding to the target need interaction request, and converts the recognition results into executable tasks to be performed associated with the target need in the human-computer interaction system. The task to be performed includes at least one target subtask, and the target subtask corresponds to different task execution services.
[0188] In order to facilitate a better understanding of the target needs and at least one target subtask, the following is explained in an example manner. When the user's target need is a reading need, the target need can be identified as at least one target subtask, and the target subtasks include but are not limited to turning off music / video, adjusting reading lights, air conditioning, seat / table adjustment, etc.; when the user's target need is a nap need, it can be identified as at least one target subtask, and the target subtasks include but are not limited to turning off music / video, air conditioning, setting alarms, door lock settings, etc.; when the user's target need is a short trip, the target subtasks include but are not limited to music / video, air conditioning, map settings, etc. The above examples are only illustrative. It can be seen from the above examples that the target need is associated with a task to be performed, and the task to be performed is associated with at least one target subtask.
[0189] In another implementation, the target requirement and the pending tasks associated with it are not strictly mapped. Instead, the relationship between the target requirement and the pending tasks includes human-oriented prompts, such as informing the user of the associated pending task and how to adjust it, to facilitate user understanding. There are also system-oriented prompts to help the system understand the user's actual needs.
[0190] Step 403: Call the task execution service of each target subtask to execute the corresponding target subtask and obtain the task execution result of the target subtask.
[0191] In the disclosed embodiments, to increase the system's computing power and expand its functionality, the Task Execution Service provides auxiliary capabilities to assist the system in completing the interaction process. During this interaction, developers can more conveniently use various APIs. For example, during the interaction, the Task Execution Service can describe the desired function or algorithm to be implemented. The Task Execution Service will then provide code examples that implement that function using a specific AI Framework, thereby helping developers more efficiently assemble and orchestrate tasks during the interaction process.
[0192] After the task execution service is started, at least one task execution service executes the corresponding target subtasks in parallel or serially.
[0193] Taking the example of step 402 as an example, when the user's target demand is reading demand, the target subtasks include but are not limited to turning off music / video, adjusting reading lights, air conditioning, seat / table adjustment, etc., and the corresponding target task execution services are player, light array task execution service, air conditioning drive task execution service, and seat / table drive task execution service.
[0194] Step 404: Display the task execution result of the at least one target subtask in the form of a user interface UI.
[0195] In actual applications, there are two types of task execution results of task execution services. One is to directly execute the target subtask (also called a single-round task). For example, when the target subtask is to open the car window, the task execution service is a car window driver plug-in, which directly controls the opening of the car window to complete the execution of this target subtask. This type of task execution result does not require the display of the task execution result; the other is a scenario where the task execution result needs to be displayed. For example, when the user plays music or navigates to a certain location, the music playback interface or the navigation interface needs to be displayed on the display screen.
[0196] For the task execution results corresponding to at least one target subtask obtained, the execution results of all task execution services to be displayed are combined. The purpose is to synchronously display the processing status of the user's intention to be executed by multiple tasks in parallel in the display interface for user viewing.
[0197] Taking step 402 as an example, when the user's target demand is reading demand, the execution results of each target task execution service include: turning off the player, adjusting the light to reading mode based on the light array task execution service, adjusting to the user's long-set temperature based on the air-conditioning drive task execution service, and adjusting the seat and / or table and chairs to a height that matches the user's height based on the seat / table drive task execution service.
[0198] The human-computer interaction method provided by the present disclosure receives a user's target demand; identifies the target demand according to a pre-established preset recognition model to obtain a task to be executed associated with the target demand, wherein the task to be executed includes at least one target subtask, and each target subtask corresponds to a different task execution service; calls the task execution service of each target subtask to execute the corresponding target subtask to obtain a target subtask execution result; and displays the task execution result of the at least one target subtask in the form of a user interface UI. The embodiment of the present disclosure identifies the user's target demand through a preset recognition model, and decomposes a target demand into at least one target subtask, wherein these target subtasks are associated with the target demand, and executes at least one associated target subtask to achieve the target demand. This realizes the input of one target demand, the identification of multiple tasks, and the execution of at least one related target subtask, without the user having to face multiple tasks and input them one by one to achieve the execution of multiple tasks, thereby improving the efficiency of human-computer interaction.
[0199] As a refinement of the above embodiment, in step 402, the target requirement is identified according to a pre-established preset recognition model, and the task to be executed associated with the target requirement is obtained, including: inputting the target requirement into the preset recognition model to obtain the target intention corresponding to the target requirement; searching for at least one target subtask associated with the target intention; outputting the at least one target subtask to obtain the task to be executed associated with the target requirement containing at least one target subtask.
[0200] The pre-established preset recognition model takes natural language as input and provides the ability to interact with users in natural language dialogue. After receiving the target demand, the fused information in the target demand is identified by the pre-established recognition model to obtain the task to be executed associated with the target demand. The pre-established recognition model described in the embodiment of the present disclosure is a diversified platform that can recognize any type of input information such as image information, acoustic information, touch interaction, geographic location information, vehicle status information, environmental information (light information, etc.), that is, the user only needs to provide a target demand information to the human-computer interaction system, which will be analyzed and processed by the pre-established recognition model, and the environment and task requirements will be understood according to the requirements corresponding to the target demand information, and the recognition result will be converted into an associated task to be executed, which contains at least one target subtask.
[0201] Specifically, the AI master model identifies a target requirement as a corresponding target intent. It then decomposes the task to be performed, identifying at least one target subtask associated with the target intent. In summary, users will minimize interaction with the product, handing over operational decision-making to the AI master model. Upon receiving the user's request, the AI master model autonomously infers the required task execution services, such as search, web applications, or plugins, and then uses these services to cascade through the target subtasks, without requiring user intervention.
[0202] The access standards between task execution services include but are not limited to:
[0203] Access protocol: Determine the access protocol, such as HTTP / HTTPS, RESTful API, etc.
[0204] Data format: Determine the data exchange format, such as JSON, XML, etc.;
[0205] Security: Establish security standards for data transmission, such as encryption and authentication.
[0206] Based on the above description of the task execution service system, the task execution service of each target subtask is called to execute the corresponding target subtask, and the task execution result of the target subtask is obtained, as shown in FIG7C , including:
[0207] Step 501 : determining the task execution service corresponding to each target subtask based on the pre-established correspondence between subtasks and task execution services.
[0208] According to the pre-established correspondence between subtasks and task execution services, the task execution services corresponding to the target subtasks are obtained respectively, and the at least one target subtask is dispatched to the corresponding task execution service respectively, and each target task execution service executes the corresponding target subtask. It should be noted that the pre-established correspondence between subtasks and task execution services is not a strong mapping relationship. The pre-established correspondence between subtasks and task execution services has human-oriented prompts, such as informing the user what this task execution service is and what it is used for, so that the user can understand it easily. There are also prompts for the task execution service, which can enable the task execution service to understand this subtask, etc.
[0209] As an implementation method of the embodiment of the present disclosure, the task execution service for executing at least one target subtask can be determined through the reasoning unit in the human-computer interaction system, and used to autonomously execute the corresponding control strategy and make decisions on behalf of the user. The reasoning engine can be a reasoning engine developed by the human-computer interaction system itself, or it can be any reasoning engine used in the existing technology. Specifically, the embodiment of the present disclosure does not limit it.
[0210] Step 502: dispatch the at least one target subtask to its corresponding task execution service, and each target task execution service executes its corresponding target subtask.
[0211] After identifying the target requirement according to a pre-established preset recognition model and obtaining the to-be-executed task associated with the target requirement, the to-be-executed task is disassembled and processed to obtain at least one target subtask; the reasoning unit determines the task execution service corresponding to each of the target subtasks and the corresponding service type, and the service type is used to determine the task execution service that needs to be called by each target subtask.
[0212] In the embodiment of the present disclosure, two service categories are included. One is the application type, and the corresponding task execution service includes the application service; the other is the plug-in type, and the task execution service includes the plug-in service.
[0213] When the service type includes an application type and the task execution service includes an application service, the target subtasks are dispatched to their respective corresponding application services according to the application type, and each application service executes its respective corresponding target subtask.
[0214] Application services are stored in pre-set application services, also known as intelligent co-pilots, which provide AI assistant services to humans. Pre-set application services intelligently acquire certain auxiliary capabilities through natural language interaction, thereby improving production efficiency. In traditional software, humans need to independently complete tasks based on pre-set program rules through commands, keystrokes, touch, etc., and the AI assistance capability is zero. In contrast, in the form of pre-set application services, AI begins to be involved in tasks, and humans need to interact with AI, obtain information, and then collaborate to complete the tasks. Currently, Windows, Microsoft 365, GitHub, Bing, etc. have implemented the capabilities of pre-set application services. The pre-set application service form has a significant change in human-computer collaboration and improved the efficiency of processing tasks.
[0215] The preset application service provides an interactive entry point, takes text, images, voice, and other integrated information as input, calls functions in the AI Framework API library to use the preset application service, and uses the preset application service of the AI Framework to implement session processing, operation orchestration, and data auditing capabilities, and obtain data access capabilities through the domain.
[0216] When dispatching the target subtasks to their corresponding application services, a manifest file is invoked. This file uses natural language to describe how to call an API, which API to call in different scenarios (such as the API of the application corresponding to the target subtask), and what parameters are required and what the output parameters are. Based on the manifest file, the target subtasks are dispatched to their corresponding application services. The application services return results, which may be code, text, audio, or video files. The application services need to display these results in the UI and provide feedback to the user.
[0217] The Preset Application Service introduces an AI Framework API library, which helps developers more conveniently use various AI Frameworks. For example, during an interaction, you can describe the desired function or algorithm to the Preset Application Service, which will then provide code examples that implement this function using a specific AI Framework, helping developers complete the interaction more efficiently.
[0218] In some embodiments, when the task execution service is a plug-in service, the plug-in type corresponding to each target subtask is identified, and the target subtask is executed through the plug-in service corresponding to each plug-in type.
[0219] During the specific application process, each plug-in will carry the plug-in type of the plug-in when it is registered, so that the reasoning unit can obtain the plug-in type of the plug-in and select and call the plug-in according to different user needs. After the reasoning unit infers the task execution service corresponding to each target subtask, it identifies the plug-in type corresponding to the target subtask and sends a task request information to the plug-in proxy Plugin Proxy in the preset plug-in service to call the plug-in corresponding to the plug-in category. The plug-in proxy Plugin Proxy responds to the task request information sent to schedule the plug-in service. The plug-in proxy Plugin Proxy is responsible for scheduling the target subtask to the corresponding plug-in service Plugin through the PME (Proxy Match Engine) engine. The PME engine records the address information of all plug-ins in the preset plug-in service. Based on this address information, the plug-in service corresponding to each plug-in type can be queried, and the target subtask is executed through the plug-in service.
[0220] In another implementation of the embodiment of the present disclosure, the plug-in service is stored in the business platform. When the plug-in in the business platform needs to be called, the plug-in proxy Plugin Proxy is responsible for scheduling the target subtask to the corresponding business platform through the PME (Proxy Match Engine) engine. The PME engine records the address information of multiple business platforms and all plug-ins in multiple business platforms.
[0221] For example, the business platform can correspond one-to-one to the plug-in type. For example, when the plug-in type is a temperature control type, the corresponding business platform stores plug-ins related to temperature control. When the plug-in type is a display device type, the corresponding business platform stores plug-ins related to the display device.
[0222] Corresponding to the configuration relationship between the plug-in category and the business platform, the preset plug-in matching rules described in the embodiment of the present disclosure may include but are not limited to first determining the matching business platform based on the plug-in type, and then matching the corresponding plug-in service (or the address of the plug-in service) from the business platform. Specifically, the embodiment of the present disclosure does not limit the setting of the preset plug-in matching rules.
[0223] The PME engine uses pre-defined plugin matching rules based on the plugin type to find the plugin service corresponding to the target subtask. This plugin service then dispatches the target subtask to the business platform for execution. The Plugin Proxy, plugin service, and business platform perform logical execution via remote calls.
[0224] In order to facilitate users in viewing the task execution results of the target task execution service, the task execution results of at least one target subtask are combined in a UI, and the combined UI is output and displayed. Before the target task execution service corresponding to the at least one target subtask executes its corresponding target subtask, layout description information is generated based on the number of target task execution services and the task execution service category. The layout description information is used to describe the dynamic layout information of the corresponding target task execution service in the user interface view. After outputting the execution results of each target task execution service, the task execution results corresponding to all plug-in services are combined according to the dynamic layout information to be displayed, generating a user interface view containing all task execution results, and outputting the user interface UI view.
[0225] The task to be executed is decomposed into at least one target subtask, and the target subtask is determined to be dispatched to the corresponding target task execution service, and then the layout description information is generated. The layout description information uses a specific layout language (DSL) to describe the interface layout rules, and the layout description language is usually output by understanding user behavior and intention. After determining the target task execution service to be called, the number of target task execution services and the category of task execution services are counted. The purpose is to configure the layout description information according to the number of target task execution services and the category of task execution services. For ease of understanding, each target task execution service can be understood as a card. Each card has position information, size information, style information, etc. in the display area. Configuring the layout description information means configuring the dynamic layout information of all target task execution services (cards) on the display interface. Specifically, the position information, size information, and style information (collectively referred to as dynamic layout information) corresponding to different task execution service categories may be different. Different numbers of target task execution services may also result in differences in the layout of the display interface. Therefore, since the target task execution services involved in each interaction are different, or the plug-in services called are different, the constructed user interface UI views may be different. Therefore, the constructed UI view is a dynamically changing one.
[0226] In some embodiments, in addition to being displayed through a UI view, the task execution results can also be displayed in the form of a human-computer interaction interface, which includes at least one of sound, vibration, digital human motion effects, task execution service controls, text images, and videos.
[0227] FIG7D is a flow chart of a method for displaying human-computer interaction provided in an embodiment of the present disclosure.
[0228] As shown in FIG7D , the method comprises the following steps:
[0229] Step 601, obtain at least one target subtask and interactive interface layout description information corresponding to the intent information, wherein the task to be executed representing the intent information is decomposed into at least one target subtask, and the interactive interface layout description information includes layout parameters of at least one target subtask execution result information.
[0230] After receiving the user's intention information, the pre-established recognition model is used to identify the intention information as a task to be executed, and the task to be executed is decomposed into multiple target subtasks, each of which is assigned to the corresponding target application for execution.
[0231] For example, when a user inputs "I want to travel to city B with user A on X month X day (intention information)", the target subtasks that can be broken down include, but are not limited to, purchasing subway tickets, booking air / train tickets, booking hotels, researching local food / attractions, and booking tickets. The above examples are merely illustrative, and the present disclosure does not limit the intent information or the individual target subtasks.
[0232] The interactive interface layout description information includes layout parameters corresponding to the execution results of each target subtask. These parameters are used to determine the display interface of the execution results when they are presented to the user, such as interface size, transparency, and display position. The above distances are merely exemplary, and the present disclosure does not limit the layout parameters.
[0233] It should be noted that different target subtasks correspond to different layout parameters. These layout parameters can be set in advance based on the output of the target subtask. For example, if the output of the target subtask may contain a lot of text, the interface size in the layout parameters of the target subtask can be set larger to reduce the user's reading pressure. The above is merely an illustrative example, and the present disclosure does not limit the setting of layout parameters.
[0234] Step 602: Load the at least one target subtask separately to obtain a task execution result.
[0235] The target subtask is loaded into the corresponding task execution service for execution, resulting in the task execution result. The task execution service can be provided by the application, the system, or an independent service. Modules that adhere to standard service protocols and provide atomic capabilities can be considered services. Each task execution service triggers a service capable of handling the target subtask, executing the target subtasks in parallel to obtain the task execution result. Calling the task execution service can improve the match with intent information and enhance the user experience in various interaction scenarios.
[0236] Step 603 : Layout the execution result of each target subtask according to the layout parameters of the at least one target subtask execution result information, and generate a user interface view including the execution results of all target subtasks.
[0237] The execution results of all target plug-ins in step 602 are combined according to their corresponding layout parameters, with the purpose of synchronously displaying the processing status of the intent information being executed in parallel by multiple tasks in the display interface for user viewing.
[0238] In the example of step 701, the target subtasks include purchasing subway tickets, booking air tickets / train tickets, booking hotels, searching for local food / attractions, and booking tickets. The target plug-ins (5 in total) corresponding to each target subtask (5 in total) respectively execute the query of subway tickets, the query of air tickets / train tickets, the query of hotels, the query of local food / attractions, and the query of tickets. The query serves as the task execution result of the target plug-in, and the queried content is the content to be displayed. The embodiment of the present disclosure, based on the layout parameters of the task execution result of each target plug-in, after determining the display interface of each content to be displayed, lays out the execution results according to the layout parameters to generate a user interface view containing the execution results of all target subtasks.
[0239] Step 604: display the user interface view.
[0240] According to the user interface view in step 603 , the execution structure of each target subtask is displayed for the user to refer to, select, etc.
[0241] It should be noted that the user interface views need to be displayed at the same display level. The purpose is to facilitate users to display the execution results of all target subtasks at the same time.
[0242] The present disclosure provides a method for displaying human-computer interaction, which obtains at least one target subtask and interactive interface layout description information corresponding to intent information, wherein the task to be executed representing the intent information is decomposed into at least one target subtask, and the interactive interface layout description information includes layout parameters of at least one target subtask execution result information; the at least one target subtask is loaded separately to obtain the task execution result; the execution result of each target subtask is laid out and typeset according to the layout parameters of the at least one target subtask execution result information, generating a user interface view containing all the target subtask execution results; and the user interface view is displayed. Compared with the prior art, by obtaining at least one target subtask and interactive interface layout description information corresponding to intent information, the corresponding target subtask can be loaded, and the returned execution result of the at least one target subtask can be typeset and displayed in a user interface view, without having to start the corresponding applications one by one according to a single demand and displaying the corresponding content in the user interface views of different applications. This reduces the complexity of the user's operation of obtaining information, improves the timeliness of obtaining information, and thus improves the user experience.
[0243] For further explanation of step 603, please refer to FIG. 7E , which is a flowchart of a human-computer interaction display method provided by an embodiment of the present disclosure, including:
[0244] Step 701: parse the layout description information and obtain layout parameters corresponding to each target subtask execution result in the layout description information; the layout parameters include attribute information of the target subtask execution result, multiple display position information and display position parameter information corresponding to each display position information; wherein, the attribute information includes identification information representing the importance of the target subtask execution result.
[0245] The layout description information is parsed and converted into a data structure that can be processed by the program. In one feasible manner of the embodiment of the present disclosure, the layout description information is saved via a DSL (Domain-Specific Language). A corresponding DSL parser can be used to parse the layout description information, such as a generative UI engine, and layout parameters of the execution results of each target subtask can be obtained based on the parsing results.
[0246] In order to make the layout of the task execution results more in line with the user's usage habits, the layout parameters described in the embodiment of the present disclosure include attribute information of the target subtask execution results, multiple display position information and display position parameter information corresponding to each display position information; the attribute information includes identification information that characterizes the importance of the target subtask execution results.
[0247] For ease of understanding, assuming that when the target subtasks include three, the corresponding task execution results also include three, namely task execution result 1, task execution result 2 and task execution result 3. The layout description information describes the attribute information corresponding to task execution result 1, task execution result 2 and task execution result 3 respectively. For example, the attribute information can be presented through three different weights, or through multiple rounds of clicks on the task execution results. Specifically, the embodiment of the present disclosure does not limit the presentation form of the attribute information of the task execution results.
[0248] In addition, the layout description information also includes multiple display position information and display position parameter information corresponding to each display position information, wherein the display position information usually corresponds to the number of task execution results, and records the display position of each task execution result in the display interface, and the display position parameter information includes but is not limited to interface size, transparency, display style, display size, etc. Specifically, the embodiment of the present disclosure does not limit the layout description information.
[0249] Step 702: Layout the task execution results of at least one target subtask according to the attribute information of the target subtask execution results, the multiple display position information, and the display position parameter information corresponding to each display position information, and generate a user interface view containing all target subtask execution results.
[0250] In some embodiments, the generation of a user interface view including the execution results of all target subtasks can be implemented in two ways:
[0251] Method 1: According to the attribute information of the target subtask execution result and the display position parameter information corresponding to each display position information, match the display position information corresponding to each target subtask execution result, obtain the target display position information corresponding to each target subtask execution result, construct the target typesetting format for displaying the user interface view according to the target display position information and the target subtask execution result, layout and typeset the task execution result of at least one target subtask according to the target typesetting format, and generate a user interface view containing the task execution results of all the target subtasks.
[0252] As a feasible method of the embodiment of the present disclosure, when matching the display position information corresponding to each target subtask execution result according to the attribute information of the target subtask execution result and the display position parameter information corresponding to each display position information, and obtaining the target display position information corresponding to each target subtask execution result, it can be implemented in but not limited to the following ways, for example: configuring corresponding display position parameter information for each display position of the display interface according to the user's historical usage frequency, determining the display position of the target subtask execution result according to the attribute information of the target subtask execution result and the display position parameter information, wherein the attribute information of the target subtask execution result and the display position parameter information are positively correlated. The purpose of such setting is to determine the display position of the target subtask execution result, which is closer to the user's usage habits and increases user stickiness.
[0253] For ease of understanding, for example, assuming that the display interface includes 4 display positions (four grids), the 4 historical usage frequencies of the 4 display positions are determined in turn according to the user's historical frequency of use of the screen. For example, in the past month, the user's historical usage frequency of display position 1 is 200 times (the display position parameter information ranks first), the historical usage frequency of display position 2 is 180 times (the display position parameter information ranks third), the historical usage frequency of display position 3 is 185 times (the display position parameter information ranks second), and the historical usage frequency of display position 4 is 40 times (the display position parameter information ranks fourth). According to the importance of the target subtask execution results, they are task execution result 2, task execution result 3, task execution result 1 and task execution result 4, and according to the importance of the target subtask execution results (attribute information), display position 1 is determined as the display position of task execution result 2, display position 3 is determined as the display position of task execution result 3, display position 2 is determined as the display position of task execution result 1, and display position 4 is determined as the display position of task execution result 4.
[0254] The above is merely an exemplary explanation for facilitating the understanding of determining the display position of the target subtask execution result, and is not a specific limitation on the target subtask, task execution result, display position parameter information, and display position of the target subtask execution result.
[0255] In addition to the four-grid example above, the display position of the task execution results is not limited to the nine-grid, and can be displayed in a pentagonal display, a star-shaped display, or other shapes and layouts.
[0256] Method 2: The target subtask execution result information is in the form of a card. The card information corresponding to each target subtask execution result is obtained. The card information is dynamically constructed by the task execution service corresponding to the target subtask according to the task execution result; the execution result of each target subtask is laid out and typeset according to the layout parameters of the at least one card to generate a user interface view containing all target subtask execution results.
[0257] In the embodiment of the present disclosure, the task execution service not only provides data and functions to the generative UI interactive interface, but also provides a dynamic card for the task execution result of the target subtask. The task execution result of each target subtask corresponds to a card, which is displayed in the form of a card. The corresponding card information varies with the type of the target subtask. For example, when the target subtask is a video task, the corresponding card information corresponds to the video display form, including but not limited to the display position, card size, video content, duration, video layout, etc. When the target subtask is a text task, the corresponding card information corresponds to the text display form, including but not limited to the size, layout, color, font, etc. of the text. Specifically, the specific content of the card information in the embodiment of the present disclosure is not limited.
[0258] After determining the layout parameters of the card, the execution results of each target subtask are laid out according to the layout parameters of the at least one card to generate a user interface view containing the execution results of all target subtasks. Specifically, the embodiment of the present disclosure does not limit the implementation method of generating the user interface view.
[0259] As an implementable method of an embodiment of the present disclosure, when displaying the user interface view in step 604, it is determined whether the target subtask contained in the user interface view is a task completed by the scheduling task execution service; when the target subtask is completed by the scheduling task execution service, the user interface view is displayed in a preset user interface view generation area, and the preset user interface view generation area includes a desktop generation area and an application generation area.
[0260] In the embodiment of the present disclosure, the preset user interface view generation area is a configurable generation area, but the preset display position in the desktop or application needs to be preset in advance. The embodiment of the present disclosure does not limit the specific location of the desktop generation area and the application generation area.
[0261] After displaying the user interface view, it receives interaction events triggered in the user interface view and dispatches the interaction events to the corresponding application or plug-in for execution. Responding to user interaction events, i.e., callbacks to user actions such as clicks, slides, and drags, completes the interaction loop by dispatching the interaction events to the corresponding application or plug-in for execution.
[0262] In some embodiments, as a further explanation of step 602, when executing the loading of the at least one target subtask separately and obtaining the task execution result, it can be implemented in but not limited to the following ways, including: matching the corresponding plug-in agent for the at least one target subtask through the plug-in matching service, the plug-in matching service is used to manage the registration and scheduling of plug-ins; calling the plug-in agent to execute the target subtask and obtain the corresponding task execution result.
[0263] In the plug-in matching service, a corresponding plug-in proxy is matched for each target subtask. The plug-in proxy is responsible for scheduling the target subtask to the corresponding target plug-in through the PME (Proxy Match Engine) engine. The PME engine records the address information of all plug-ins in the plug-in matching service. After determining the target plug-in to be called, the target plug-in corresponding to the target subtask can be queried through the PME engine.
[0264] FIG7F is a flow chart of a vehicle interaction method provided in an embodiment of the present disclosure.
[0265] As shown in FIG7F , the method comprises the following steps:
[0266] Step 801: Receive a human-computer interaction request.
[0267] The system described in the embodiment of the present disclosure uses natural language as input and provides the ability to interact with users in natural language dialogue. The user inputs unstructured fused information into the vehicle's interactive system. The fused information includes but is not limited to at least one of image information, acoustic information, touch interaction, geographic location information, vehicle status information, and environmental information (light information, etc.). The fused information is used to trigger interaction requests during human-computer interaction between people and vehicles.
[0268] Step 802: Identify the fused information in the interaction request to obtain a task to be executed corresponding to the interaction request, wherein the fused information includes at least one type of interaction input information.
[0269] As an implementation method of the embodiment of the present disclosure, the fused information in the interaction request is identified through a pre-established recognition model to obtain the task to be executed corresponding to the interaction request. The pre-established recognition model described in the embodiment of the present disclosure is a diversified platform, which can recognize any type of input information such as image information, acoustic information, touch interaction, geographic location information, vehicle status information, environmental information (light information, etc.), that is, the user only needs to give an interaction to the vehicle system, and the pre-established recognition model will analyze and process it, understand the environment and task requirements according to the needs corresponding to the interaction request, and convert the recognition results into executable tasks to be executed by the vehicle system.
[0270] As another implementation method of the embodiment of the present disclosure, the fused information in the interaction request can also be identified by any recognition algorithm in the existing technology. For example, when the fused information includes an image, the recognition algorithm used is any image recognition algorithm in the existing technology. When the fused information includes voice, the recognition algorithm used is any acoustic recognition algorithm in the existing technology, and so on. The embodiment of the present disclosure will not elaborate on the specific recognition algorithms in the existing technology that are called.
[0271] Step 803: Call at least one target plug-in corresponding to the task to be executed in the preset plug-in service, and process the corresponding task to be executed in parallel through the at least one target plug-in.
[0272] The preset plug-in service described in the embodiment of the present disclosure is intended to enhance and customize the natural language processing capabilities of the vehicle's interactive system. First, it can process and change the user's input information and provide additional context to the system, thereby optimizing the model's output results. The plug-in allows the system to access and obtain the latest information, access and use third-party services, improve the match between actual application requirements, and improve the system's performance in various scenarios. Secondly, the plug-in can also empower developers, allowing developers to better control the behavior of the system, allowing developers to optimize, customize and expand the system's capabilities according to specific needs and preferences, and upgrade the system to a powerful and diversified platform.
[0273] In the embodiments of the present disclosure, a plug-in can be provided by an application, a system, or an independent service. Any module that complies with the standard plug-in protocol and provides atomic capabilities can be considered a plug-in.
[0274] After obtaining the executable task to be executed, the preset plug-in service matches at least one target plug-in corresponding to the task to be executed in the preset plug-in service, and schedules the task to be executed to the corresponding target plug-in. Different target plug-ins execute the task to be executed in parallel.
[0275] Step 804: Output and display the task execution result of the task to be executed.
[0276] The execution results of each target plug-in can be presented in the form of vibration, light, voice, text, etc. Specifically, the embodiment of the present disclosure does not limit the presentation form of the execution results.
[0277] The vehicle interaction method provided by the present disclosure receives an interaction request for human-computer interaction, identifies the fused information in the interaction request according to a pre-established preset processing model, obtains a to-be-executed task corresponding to the interaction request, wherein the fused information includes at least one interactive input information, calls at least one target plug-in corresponding to the to-be-executed task in a preset plug-in service, processes the corresponding to-be-executed task in parallel through the at least one target plug-in, and outputs and displays the task execution result of the to-be-executed task. Compared with the related art, the embodiment of the present disclosure simplifies the interaction operation steps from the user's perspective by identifying the user's interaction request as an executable to-be-executed task, determining at least one target plug-in corresponding to the to-be-executed task by calling a preset plug-in service, and having at least one plug-in execute the corresponding subtask in parallel. Only one human-computer interaction operation with the vehicle computer needs to be performed once to meet the vehicle use demand of performing multiple tasks in one interaction.
[0278] As a refinement of the above embodiment, when executing the call of at least one target plug-in corresponding to the task to be performed in the preset plug-in service in step 803, the following method can be adopted but not limited to: calling a generative pre-trained model (Generative Pre-trained Transformer, GPT) in the preset plug-in service to search for the target scene corresponding to the task to be performed; the preset plug-in service contains different preset scenes and their corresponding plug-ins, and different scenes correspond to the same or different plug-ins; according to the correspondence between scenes and plug-ins in the generative pre-trained model, search for at least one target plug-in corresponding to the task to be performed. In the embodiment of the present disclosure, one task to be performed corresponds to one target scene, and one scene corresponds to at least one target plug-in. These target plug-ins jointly respond to a user's interaction request, realizing the user's car use needs of completing multiple tasks in one interaction.
[0279] In order to enhance the accuracy of GPT search, different scenarios and corresponding at least one target plug-in can be trained separately to improve the accuracy of GPT. The training process is not the focus of the embodiment of this disclosure, so it will not be described in detail.
[0280] Before using a plug-in, it is necessary to register the plug-in. When registering the plug-in, the plug-in interface address is carried. When searching for the target plug-in in the future, the corresponding target plug-in can be found directly through the addressing method. The specific search process includes: obtaining the interface address of at least one target plug-in from the correspondence between the scenario and the plug-in in the generative pre-trained model, and searching for the corresponding target plug-in according to each interface address. Each target plug-in corresponds to a unique interface (Application Programming Interface, API) address, which makes it more convenient for users to use various plug-ins.
[0281] The above embodiment describes the plug-in usage process in detail. As shown in FIG7G , FIG7G is a flowchart of a plug-in registration method provided by an embodiment of the present disclosure, including:
[0282] Step 901: In response to a request message for registering a plug-in, the preset plug-in service sends the request message for registering a plug-in to the generative pre-trained model according to the plug-in protocol. The request message includes the life cycle of the registered plug-in and the corresponding scenario.
[0283] Plugin registration includes, but is not limited to, system-internal and third-party plugin registration. Plugin registration must adhere to the standard plugin protocol. The plugin protocol includes, but is not limited to, the plugin interface, the API interface, and the manifest file. The plugin interface defines the system interface, which plugins must implement. The API interface contains multiple functions that define data input and output in different scenarios. The manifest file uses natural language prompts to instruct the system how to call the API, allowing the system to learn which API to call in different scenarios, what parameters to pass in, and what the output parameters are.
[0284] The preset plug-in service has a standard registration mechanism. When a registered plug-in is registered with the plug-in system, it carries the (specific domain) of the registered plug-in. The specific domain can be divided into scenarios. In the preset plug-in service, the request information for registering the plug-in is received, and the preset plug-in service registers the plug-in with the GPT plug-in interface.
[0285] In step 902, the generative pre-trained model obtains the scene to which it belongs, and adds the registered plug-in and the scene to which it belongs to the corresponding relationship between the plug-in and the scene.
[0286] Step 903: Monitor the registered plug-in according to the life cycle.
[0287] Step 904: After the life cycle is reached, the registered plug-in is deregistered and the registered plug-in is deleted from the corresponding relationship between the plug-in and the scene.
[0288] Each pre-configured plug-in service includes key steps such as registration, creation, invocation, and deregistration. The plug-in management service manages the lifecycle of each plug-in through these steps. When a plug-in is created or deregistered, the corresponding relationship between the plug-in and the scenario must be updated synchronously to avoid GPT invocation errors that may affect the user experience.
[0289] After the target plug-in completes the corresponding subtask, the target plug-in's processing results must be fed back to the user so that the user can view them. This embodiment of the disclosure uses two methods for display:
[0290] Method 1: Obtain the input category of the fused information, where the input category of the fused information includes at least one of picture, voice, gesture, video, and gaze; and display the processing result of the at least one target plug-in in a display style that matches the input category.
[0291] In the embodiment of the present disclosure, the input category and its matching display style form an editable correspondence, and users can flexibly configure them according to their preferences.
[0292] For ease of understanding, when the input type of the fusion information is a gesture, the display styles that can be matched with the processing result of the target plug-in include but are not limited to voice and UI interface. Specifically, the embodiments of the present disclosure are not limited to this.
[0293] Method 2: parse the processing result of the at least one target plug-in, determine the tag information carried in the processing result, and the processing result carries the tag information of directly controlling the vehicle state / indirectly controlling the vehicle state; if it is determined that the tag information is directly controlling the vehicle state, output the execution result information corresponding to the controlled vehicle state; if it is determined that the tag information is indirect controlling the vehicle state, call the user interface container, and construct the user interface view corresponding to the processing result based on the user interface container, and output and display the user interface view.
[0294] To facilitate understanding of direct control of vehicle status / indirect control of vehicle status, the following example is used to illustrate. When the target subtask is to open the window, the target plug-in is the window driver plug-in. Directly controlling the opening of the window can complete the execution of this target subtask. The window driver plug-in directly controls the vehicle status and does not need to display the task execution results. The other is the scenario where the task execution results need to be displayed (indirect control of vehicle status). For example, when the user plays music or navigates to a certain location, the music playback interface or the navigation interface needs to be displayed on the display screen. The target plug-in does not directly control the vehicle status and needs to display the processing results.
[0295] For scenarios where the tag information does not directly control the vehicle state, the target plug-in's processing results are converted into corresponding layout description information. The layout description information is used to describe the processing results and layout attribute information of each target plug-in. The layout attribute information includes the target plug-in's identification information and display information. In actual applications, the layout description information can use a specific layout language (DSL language) to describe the display information of the interface layout. The display information includes but is not limited to the position information on the display screen or application, the display window size, display style, transparency, and other information. Specifically, the embodiments of this disclosure do not limit this.
[0296] The system calls the user interface UI container to parse the acquired layout description information to obtain the identification information and display information of the target plug-in. In the specific application process, the identification information of the target plug-in obtained by parsing is regarded as the identification information of a card, that is, one target plug-in corresponds to one card node, such as target plug-in 1 corresponds to card node 1, target plug-in 2 corresponds to card node 2, and so on.
[0297] Based on the identification information of the target plug-in, the target plug-ins corresponding to the multiple target subtasks are loaded separately. Based on the obtained display information, the layout information of the loaded target plug-ins within the display desktop or application is determined. Based on the layout information, user interface (UI) views corresponding to the multiple sub-target tasks are generated and constructed, and the UI views are rendered and displayed. This can be understood as the layout information being treated as a node tree, and the rendering engine dynamically constructing the UI view based on the node tree, presenting the UI view through a generative UI container.
[0298] Based on the displayed UI view, in response to the interaction instruction triggered by the user in the user interface UI view, the interaction instruction is determined as an execution event and sent to the scheduling module, and the scheduling module responds to the interaction instruction to complete multiple rounds of interaction. The scheduling module interacts by calling the preset plug-in service.
[0299] Respond to user interaction instructions, call back user clicks, slides, drags and other operation behaviors, process interaction events through the scheduling module, and the scheduling module passes the interaction events to the preset plug-in service for execution scheduling, completing the closed loop of the interaction link.
[0300] Corresponding to the above-mentioned human-computer interaction method, the present disclosure also provides a human-computer interaction device that can be installed in a vehicle. Since the device embodiments of the present disclosure correspond to the above-mentioned method embodiments, details not disclosed in the device embodiments can be referred to the above-mentioned method embodiments and will not be further described in this disclosure.
[0301] The present disclosure also provides a human-computer interaction device, as shown in FIG8 , including:
[0302] The identification unit 301 is configured to identify the fused information in the received interaction request and obtain a to-be-executed task corresponding to the interaction request, wherein the fused information includes at least one interactive input form;
[0303] A decomposition unit 302 is configured to decompose the task to be executed into at least one target subtask;
[0304] A searching unit 303 is configured to search for a task execution service corresponding to each of the at least one target subtask;
[0305] The calling unit 304 is configured to call the task execution service to execute the corresponding target subtask and obtain the target subtask execution result;
[0306] The output unit 305 is configured to output and display the task execution result of the at least one target subtask.
[0307] The human-computer interaction device described in the embodiment of the present disclosure is a virtual execution device of the human-computer interaction system shown in Figure 1. The recognition unit 301 and the disassembly unit 302 correspond to the functions that can be implemented by the intention recognition module 1 in Figure 1, the search unit 303 corresponds to the functions that can be implemented by the task processing module 2 in Figure 1, the calling unit 304 corresponds to the functions that can be implemented by the service module 3 in Figure 1, and the output unit 305 corresponds to the functions that can be implemented by the generative UI module 3 in Figure 1. The exemplary implementation process will not be repeated here in the embodiment of the present disclosure.
[0308] The human-computer interaction device provided by the present disclosure identifies the fused information in the received interaction request, obtains the task to be executed corresponding to the interaction request, and the fused information includes at least one interactive input form; decomposes the task to be executed into at least one target subtask; searches for the task execution service corresponding to each of the at least one target subtask, and calls the task execution service to execute the corresponding target subtask to obtain the target subtask execution result; and outputs and displays the task execution result of the at least one target subtask. Compared with the related art, the present disclosure identifies the fused information in the received human-computer interaction request, obtains the task to be executed corresponding to the interaction request, and further decomposes the task to be executed corresponding to the interaction request into at least one target subtask, queries and calls the task execution service corresponding to each target subtask to execute the corresponding target subtask to obtain the task execution result; the entire execution process is not restricted by the corresponding program of the pre-made demand instruction, which greatly improves the intelligence of human-computer interaction; and after the user inputs the human-computer interaction request, the entire process of responding to the interaction request is completed by calling each other between the various units, without the user needing to intervene in the deeper touch screen interface layer, making the entire interaction process automated and intelligent, greatly improving the efficiency of human-computer interaction.
[0309] Furthermore, in a possible implementation of this embodiment, as shown in FIG9 , the task execution service includes a target application;
[0310] The search unit 303 includes:
[0311] The first search module 3031 is configured to search for the target application corresponding to each of the at least one target subtask in the preset application service;
[0312] The calling unit 304 is further configured to call the at least one target application to execute the corresponding target subtask respectively, and obtain a task execution result.
[0313] Furthermore, in a possible implementation of this embodiment, as shown in FIG9 , the task execution service includes a target application, and the target application calls a data function of a target plug-in;
[0314] The search unit 303 includes:
[0315] The second search module 3032 is configured to search for the target application corresponding to each of the at least one target subtask in the preset application service;
[0316] A loading module 3033 is configured to load a target plug-in corresponding to the at least one target subtask to provide data functions for the target application;
[0317] The calling unit 304 is further configured to call the at least one target application, and the target plug-ins of each target application respectively execute the corresponding target subtask to obtain a task execution result.
[0318] Furthermore, in a possible implementation of this embodiment, as shown in FIG9 , the searching unit 303 further includes:
[0319] The determining module 3034 is configured to determine whether the target plug-in corresponding to the at least one target subtask has been loaded before loading the target plug-in corresponding to the at least one target subtask to provide data functions for the target application;
[0320] The registration module 3035 is configured to, when it is determined that the target plug-in corresponding to the at least one target subtask is not loaded, register the target plug-in corresponding to the at least one target subtask with the preset plug-in service after the preset plug-in service is started;
[0321] The loading module 3033 is further configured to load a target plug-in corresponding to the at least one target subtask to provide data functions for the target application.
[0322] Furthermore, in a possible implementation of this embodiment, as shown in FIG9 , the task execution service includes a target plug-in;
[0323] The search unit 303 includes:
[0324] A third search module 3036 is configured to search for a target plug-in corresponding to the at least one target subtask in a preset plug-in service;
[0325] The calling unit 304 is further configured to call the at least one target plug-in to execute the corresponding target subtask respectively, and obtain a task execution result.
[0326] Furthermore, in a possible implementation of this embodiment, as shown in FIG9 , the apparatus further includes:
[0327] The scheduling unit 306 is configured to respond to an interaction event triggered in the user interface view and schedule the interaction event to a corresponding application or plug-in for execution after the output unit 305 outputs and displays the task execution result of the at least one target subtask.
[0328] Furthermore, in a possible implementation of this embodiment, as shown in FIG9 , the apparatus includes:
[0329] The configuration unit 307 is configured to configure execution parameters of at least one target subtask according to the target requirements in the interaction request before the calling unit 304 calls the task execution service to execute the corresponding target subtask and obtains the target subtask execution result;
[0330] The calling unit 304 is further configured to call each of the task execution services to execute the target subtask corresponding to the task request information according to the corresponding execution parameters.
[0331] Furthermore, in a possible implementation of this embodiment, as shown in FIG9 , the configuration unit 307 is further configured to:
[0332] Acquire current environment information, as well as first historical data of a first time period and second historical data of a second time period; wherein the second time period is longer than the first time period;
[0333] The execution parameters of the at least one target subtask are configured according to a pre-established context learning model, and the context learning model is used to configure the execution parameters through the learning results of the current environmental information, the first historical data of the first time period and the second historical data of the second time period.
[0334] Furthermore, in a possible implementation of this embodiment, the interactive interface layout description information includes at least one layout parameter of target subtask execution result information; the output unit 305 is further configured to:
[0335] Laying out the task execution result of each target subtask according to the layout parameters of the at least one target subtask execution result information, and generating a user interface view including all target subtask execution results;
[0336] The user interface view is rendered, and the rendered user interface view is displayed.
[0337] Furthermore, in a possible implementation of this embodiment, the output unit 305 is further configured to:
[0338] Parse the layout description information to obtain layout parameters corresponding to each target subtask execution result in the layout description information; the layout parameters include attribute information of the target subtask execution result, multiple display position information, and display position parameter information corresponding to each display position information; wherein the attribute information includes identification information representing the importance of the target subtask execution result:
[0339] According to the attribute information of the target subtask execution result, the multiple display position information and the display position parameter information corresponding to each display position information, the task execution result of at least one target subtask is laid out and typeset to generate a user interface view containing all target subtask execution results.
[0340] Furthermore, in a possible implementation of this embodiment, the target subtask execution result information is in the form of a card, and the output unit 305 is further configured to:
[0341] Obtain the card information corresponding to the execution results of each target subtask, where the card information is dynamically constructed by the task execution service corresponding to the target subtask according to the task execution results;
[0342] The execution result of each target subtask is laid out and typeset according to the layout parameters of the at least one card to generate a user interface view including the execution results of all target subtasks.
[0343] Furthermore, in a possible implementation of this embodiment, the output unit 305 is further configured to:
[0344] Determining whether the target subtask included in the user interface view is a task completed by the scheduling task execution service;
[0345] When the target subtask is completed for the scheduling task execution service, the user interface view is displayed in a preset user interface view generation area, and the preset user interface view generation area includes a desktop generation area and an application generation area.
[0346] Furthermore, in a possible implementation of this embodiment, the calling unit 304 is further configured to:
[0347] Determine the task execution service corresponding to each target subtask based on the pre-established correspondence between subtasks and task execution services;
[0348] The at least one target subtask is dispatched to the corresponding task execution service respectively, and each target task execution service executes the corresponding target subtask to obtain a task execution result.
[0349] Furthermore, in a possible implementation of this embodiment, after decomposing the to-be-executed task into at least one target subtask, the calling unit 304 is further configured to:
[0350] A service type corresponding to each of the at least one target subtask is determined, where the service type is used to determine a task execution service that needs to be called by each target subtask.
[0351] Furthermore, in a possible implementation of this embodiment, the output unit 305 is further configured to:
[0352] Combining the task execution results of the at least one target subtask in a UI;
[0353] The combined UI is output and displayed.
[0354] Furthermore, in a possible implementation of this embodiment, the first search module 3031 is further configured to:
[0355] Invoke a generative pre-trained model in a preset plug-in service to search for a target scenario corresponding to at least one target subtask; the preset plug-in service includes different preset scenarios and their corresponding plug-ins, and different scenarios correspond to the same or different plug-ins;
[0356] At least one target plug-in corresponding to the at least one target subtask is searched according to the correspondence between scenarios and plug-ins in the generative pre-trained model.
[0357] Furthermore, in a possible implementation of this embodiment, the output unit 305 is further configured to:
[0358] Acquire an input category of the fused information, where the input category of the fused information includes at least one of picture, voice, gesture, video, and gaze;
[0359] The processing result of the at least one target plug-in is displayed in a display style that matches the input category.
[0360] Furthermore, in a possible implementation of this embodiment, the output unit 305 is further configured to:
[0361] Parsing a processing result of the at least one target plug-in to determine flag information carried in the processing result, wherein the processing result carries flag information of directly controlling a vehicle state or indirectly controlling a vehicle state;
[0362] If it is determined that the flag information is a direct vehicle control state, outputting execution result information corresponding to the vehicle control state;
[0363] If it is determined that the marking information is a non-direct vehicle control state, calling a user interface container and constructing a user interface view corresponding to the processing result based on the user interface container;
[0364] The user interface view is output and displayed.
[0365] Furthermore, in a possible implementation of this embodiment, as shown in FIG9 , the disassembling unit 302 is further configured to identify the fused information according to a pre-established recognition model to obtain at least one target subtask corresponding to the task to be executed.
[0366] Furthermore, an embodiment of the present disclosure also provides a vehicle, which includes the human-computer interaction system or human-computer interaction device described in any of the above embodiments.
[0367] It should be noted that the application of the human-computer interaction system or human-computer interaction device described in any of the above embodiments in a vehicle is only an illustration of an application. In fact, the human-computer interaction system or human-computer interaction device described in any of the above embodiments can also be applied to all intelligent devices that include human-computer interaction functions, such as intelligent robots, intelligent terminals, wearable devices, etc. For example, the embodiments of the present disclosure do not limit the application fields of the human-computer interaction system or human-computer interaction device described in any of the above embodiments.
[0368] According to an embodiment of the present disclosure, the present disclosure further provides a human-computer interaction device, a readable storage medium, and a computer program product.
[0369] An embodiment of the present disclosure also provides a human-computer interaction device, comprising at least one processor; a storage device configured to store at least one program, wherein when the at least one program is executed by the at least one processor, the human-computer interaction device implements the human-computer interaction method provided in the above-mentioned embodiments.
[0370] Figure 10 shows a schematic diagram of a computer device suitable for implementing a human-computer interaction device according to an embodiment of the present disclosure. It should be noted that the computer system 1000 of the human-computer interaction device shown in Figure 10 is only an example and should not limit the functions and scope of use of the embodiments of the present disclosure.
[0371] As shown in Figure 10, the computer system 1000 includes a central processing unit (CPU) 1001, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 1002 or the program loaded from the storage part 1008 into the random access memory (RAM) 1003, such as executing the method described in the above embodiment. Various programs and data required for system operation are also stored in the RAM 1003. The CPU 1001, ROM 1002, and RAM 1003 are connected to each other via a bus 1004. An input / output (I / O) interface 1005 is also connected to the bus 1004.
[0372] The following components are connected to the I / O interface 1005: an input section 1006 including a keyboard, a mouse, etc.; an output section 1007 including devices such as a cathode ray tube (CRT), a liquid crystal display (LCD), and a speaker; a storage section 1008 including a hard disk; and a communication section 1009 including a network interface card such as a LAN (Local Area Network) card or a modem. The communication section 1009 performs communication processing via a network such as the Internet. A drive 1010 is also connected to the I / O interface 1005 as needed. A removable medium 1011, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed in the drive 1010 as needed so that a computer program read therefrom can be installed into the storage section 1008 as needed.
[0373] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes a computer program for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network via the communication section 1009, and / or installed from a removable medium 1011. When the computer program is executed by the central processing unit (CPU) 1001, the various functions defined in the apparatus of the present disclosure are performed.
[0374] It should be noted that the computer-readable medium shown in the embodiments of the present disclosure may be a computer-readable signal medium or a computer-readable storage medium or any combination of the two. The computer-readable storage medium may be, for example, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or component, or any combination of the above. More exemplary examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a flash memory, an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries a computer-readable computer program. This propagated data signal may take a variety of forms, including but not limited to an electromagnetic signal, an optical signal, or any suitable combination of the above. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in connection with an instruction execution system, apparatus, or device. A computer program embodied on a computer-readable medium may be transmitted using any suitable medium, including but not limited to wireless, wired, or any suitable combination thereof.
[0375] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions and operations of the devices, methods and computer program products according to various embodiments of the present disclosure. Each box in the flowchart or block diagram can represent a module, program segment, or part of the code, and the above-mentioned module, program segment, or part of the code contains one or more executable instructions for implementing the specified logical functions. It should also be noted that in some alternative implementations, the functions marked in the boxes can also occur in an order different from that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram or flowchart, and the combination of boxes in the block diagram or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0376] The units involved in the embodiments described in this disclosure may be implemented in software or hardware, and the units described may also be provided in a processor. In some cases, the names of these units do not constitute limitations on the units themselves.
[0377] Another aspect of the present disclosure provides a computer program product or computer program, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the human-computer interaction method provided in each of the above embodiments.
[0378] The above embodiments are merely illustrative of the principles and effects of the present disclosure and are not intended to limit the present disclosure. Anyone skilled in the art may modify or alter the above embodiments without departing from the spirit and scope of the present disclosure. Therefore, any equivalent modifications or alterations made by a person skilled in the art without departing from the spirit and technical concepts disclosed herein shall be covered by the claims of the present disclosure. Industrial Applicability
[0379] The present disclosure relates to a human-computer interaction method, apparatus, system, device, vehicle, storage medium, and program product. An intent recognition module identifies fused information in a received human-computer interaction request to obtain a pending task corresponding to the interaction request, further decomposes the pending task corresponding to the interaction request into at least one target subtask, and a task processing module queries and calls a task execution service corresponding to each target subtask to execute the corresponding target subtask. The service module executes the target subtask to obtain a task execution result, which is then displayed by a generative UI module.
Claims
1. A human-computer interaction system, comprising: Intent recognition module, task processing module, service module and generative UI module; The intention recognition module recognizes the fused information in the received interaction request, obtains the task to be executed corresponding to the interaction request, and decomposes the task to be executed into at least one target subtask, wherein the fused information includes at least one interactive input form; The task processing module sends task request information for scheduling a task execution service corresponding to the at least one target subtask to the service module; The service module responds to the task request information, searches for a task execution service of at least one target subtask, and controls each of the task execution services to execute the corresponding target subtask to obtain a task execution result; The generative UI module outputs and displays the task execution result corresponding to the at least one target subtask.
2. The system according to claim 1, wherein: The task processing module includes: an engine unit and a scheduling unit, wherein: The engine unit determines whether each target subtask needs to call a service module, wherein the service module includes a target application and a target plug-in; When determining that the target application and / or target plug-in needs to be called, the scheduling unit sends task request information for scheduling the corresponding target application and / or target plug-in to the preset application service and / or preset plug-in service according to the attribute category of each target subtask.
3. The system according to claim 2, wherein: The service module includes a receiving unit, a search engine, and a processing unit, wherein: The receiving unit receives the scheduling correspondence sent by the scheduling unit and obtains task request information of a target application and / or a target plug-in respectively corresponding to at least one target subtask; The search engine queries the corresponding target application and / or target plug-in according to the attribute category of the at least one target subtask; The processing unit controls the target application and / or target plug-in to execute the corresponding target subtask to obtain a task execution result.
4. The system according to claim 2 or 3, wherein: The generative UI module includes: a parser, a loader, a layout container and a rendering engine, wherein: The parser parses the interactive interface layout description information to determine the layout information of all target applications; the interactive interface layout description information is generated according to the number and application category of the target applications and is used to describe the layout information of the corresponding target application in the interactive interface; the target application calls the data function of the target plug-in; The loader loads at least one corresponding target plug-in according to at least one target subtask; the at least one target plug-in executes the corresponding target subtask respectively to obtain a task execution result; The layout container displays the task execution result corresponding to the at least one target plug-in in a layout; The rendering engine renders and displays the user interface view.
5. The system according to any one of claims 1 to 4, wherein: The generative UI module includes: an event processing engine, The event processing engine receives interaction events in the user interface view and sends the interaction events to the task processing module.
6. A human-computer interaction method, comprising: Identifying fused information in the received interaction request to obtain a task to be executed corresponding to the interaction request, wherein the fused information includes at least one interactive input form; Decomposing the task to be performed into at least one target subtask; Finding a task execution service corresponding to each of the at least one target subtask, and calling the task execution service to execute the corresponding target subtask, to obtain an execution result of the target subtask; The task execution result of the at least one target subtask is output and displayed.
7. The method according to claim 6, wherein: The task execution service includes a target application; The searching for the task execution service corresponding to each of the at least one target subtask includes: Searching for target applications corresponding to the at least one target subtask in the preset application service respectively; The calling of the task execution service to execute the corresponding target subtask and obtaining the target subtask execution result includes: The at least one target application is called to execute corresponding target subtasks respectively to obtain task execution results.
8. The method according to claim 6 or 7, wherein: The task execution service includes a target application, and the target application calls a data function of a target plug-in; The searching for the task execution service corresponding to each of the at least one target subtask includes: Searching for target applications corresponding to the at least one target subtask in the preset application service respectively; Loading a target plug-in corresponding to the at least one target subtask to provide data functions for the target application; The calling of the task execution service to execute the corresponding target subtask and obtaining the target subtask execution result includes: The at least one target application is called, and the target plug-ins of each target application respectively execute corresponding target subtasks to obtain task execution results.
9. The method according to claim 8, wherein: Before loading the target plug-in corresponding to the at least one target subtask to provide data functions for the target application, the method further includes: Determine whether the target plug-in corresponding to the at least one target subtask has been loaded; The loading of the target plug-in corresponding to the at least one target subtask to provide a data function for the target application includes: When it is determined that the target plug-in corresponding to the at least one target subtask is not loaded, after the preset plug-in service is started, the target plug-in corresponding to the at least one target subtask is registered with the preset plug-in service.
10. The method according to any one of claims 6 to 9, wherein: The task execution service includes a target plug-in; The searching for the task execution service corresponding to each of the at least one target subtask includes: Searching for at least one target plug-in corresponding to the at least one target subtask in a preset plug-in service; The calling of the task execution service to execute the corresponding target subtask and obtain the task execution result includes: The at least one target plug-in is called to execute the corresponding target subtask respectively to obtain the task execution result.
11. The method according to any one of claims 6 to 10, wherein: After outputting and displaying the task execution result of the at least one target subtask, the method further includes: Respond to interaction events triggered in the user interface view, and dispatch the interaction events to corresponding applications or plug-ins for execution.
12. The method according to any one of claims 6 to 11, wherein: Before calling the task execution service to execute the corresponding target subtask and obtaining the target subtask execution result, the method further includes: According to the target requirements in the interaction request, respectively configure execution parameters of at least one target subtask; The calling of the task execution service to execute the corresponding target subtask and obtaining the target subtask execution result includes: Each of the task execution services is called to execute the target subtask corresponding to the task request information according to the corresponding execution parameters.
13. The method according to claim 12, wherein: The configuring the execution parameters of at least one target subtask respectively according to the target requirement in the interaction request includes: Acquire current environment information, as well as first historical data of a first time period and second historical data of a second time period; wherein the second time period is longer than the first time period; The execution parameters of the at least one target subtask are configured according to a pre-established context learning model, and the context learning model is used to configure the execution parameters through the learning results of the current environment information, the first historical data of the first time period, and the second historical data of the second time period.
14. The method according to any one of claims 6 to 13, wherein: The interactive interface layout description information includes layout parameters of at least one target subtask execution result information; and outputting and displaying the task execution result of the at least one target subtask includes: Laying out and typeset the task execution result of each target subtask according to the layout parameters of the at least one target subtask execution result information, and generating a user interface view including all target subtask execution results; The user interface view is rendered, and the rendered user interface view is displayed.
15. The method according to claim 14, wherein: The step of laying out the task execution result of each target subtask according to the layout parameter of the at least one target subtask execution result information to generate a user interface view including all target subtask execution results includes: Parse the layout description information, and obtain layout parameters corresponding to each of the target subtask execution results in the layout description information; the layout parameters include attribute information of the target subtask execution result, multiple display position information, and display position parameter information corresponding to each display position information; wherein the attribute information includes identification information representing the importance of the target subtask execution result: According to the attribute information of the target subtask execution result, the multiple display position information and the display position parameter information corresponding to each display position information, the task execution result of at least one target subtask is laid out and typeset to generate a user interface view containing all target subtask execution results.
16. The method according to claim 14 or 15, wherein: The target subtask execution result information is information in the form of a card, and the method further includes: Obtain card information corresponding to the execution results of each target subtask, where the card information is dynamically constructed by the task execution service corresponding to the target subtask according to the task execution results; The step of laying out the task execution result of each target subtask according to the layout parameter of the at least one target subtask execution result information to generate a user interface view including all target subtask execution results includes: The execution result of each target subtask is laid out according to the layout parameters of the at least one card to generate a user interface view including the execution results of all target subtasks.
17. The method according to any one of claims 11 to 16, wherein: Displaying the user interface view includes: Determining whether the target subtask included in the user interface view is a task completed by the scheduling task execution service; When the target subtask is completed for the scheduling task execution service, the user interface view is displayed in a preset user interface view generation area, and the preset user interface view generation area includes a desktop generation area and an application generation area.
18. The method according to any one of claims 6 to 17, wherein: The calling of the task execution service to execute the corresponding target subtask and obtaining the target subtask execution result includes: Determine the task execution service corresponding to each target subtask according to the pre-established correspondence between the subtask and the task execution service; The at least one target subtask is dispatched to the corresponding task execution service respectively, and each target task execution service executes the corresponding target subtask respectively to obtain the task execution result of the target subtask.
19. The method according to any one of claims 6 to 18, wherein: After decomposing the to-be-performed task into at least one target subtask, the method further includes: A service type corresponding to each of the at least one target subtask is determined, where the service type is used to determine a task execution service that each target subtask needs to call.
20. The method according to any one of claims 6 to 19, wherein: The outputting and displaying the task execution result of the at least one target subtask includes: combining the task execution results of the at least one target subtask in a UI; The combined UI output is displayed.
21. The method according to claim 10, wherein: The searching the preset plug-in service for at least one target plug-in corresponding to the at least one target subtask includes: Invoke a generative pre-trained model in a preset plug-in service to search for a target scenario corresponding to at least one target subtask; the preset plug-in service includes different preset scenarios and their corresponding plug-ins, and different scenarios correspond to the same or different plug-ins; At least one target plug-in corresponding to the at least one target subtask is searched according to the correspondence between the scenarios and the plug-ins in the generative pre-trained model.
22. The method according to any one of claims 6 to 21, wherein: The outputting and displaying the task execution result of the at least one target subtask includes: Acquire an input category of the fused information, where the input category of the fused information includes at least one of a picture, a voice, a gesture, a video, and a gaze; The processing result of the at least one target plug-in is displayed in a display style matching the input category.
23. The method according to any one of claims 10 to 22, wherein: The outputting and displaying the task execution result of the at least one target subtask includes: Parsing the processing result of the at least one target plug-in to determine the marking information carried in the processing result, wherein the processing result carries the marking information of directly controlling the vehicle state / indirectly controlling the vehicle state; If it is determined that the marking information is a direct control vehicle state, outputting execution result information corresponding to the control vehicle state; If it is determined that the marking information is a non-direct vehicle control state, calling a user interface container, and constructing a user interface view corresponding to the processing result based on the user interface container; The user interface view is output and displayed.
24. A human-computer interaction method, characterized in that: include: Receive users' target needs; The target requirement is identified according to a pre-established preset identification model to obtain a task to be executed associated with the target requirement, wherein the task to be executed includes at least one target subtask, and each target subtask corresponds to a different task execution service; The preset recognition model is a model with natural language dialogue interaction capability, which converts the target demand recognition result into an associated task to be executed by analyzing the environment and task requirements of the target demand; Calling the task execution service of each target subtask to execute the corresponding target subtask and obtain the task execution result of the target subtask; The task execution result of the at least one target subtask is displayed in the form of a user interface UI.
25. A human-computer interaction display method, characterized in that: include: Acquire at least one target subtask and interactive interface layout description information corresponding to the intent information, wherein the task to be executed representing the intent information is decomposed into at least one target subtask, and the interactive interface layout description information includes layout parameters of at least one target subtask execution result information; Loading the at least one target subtask respectively to obtain a task execution result; Laying out and typeset the execution result of each target subtask according to the layout parameters of the at least one target subtask execution result information, and generating a user interface view including the execution results of all target subtasks; The user interface view is displayed.
26. A human-computer interaction device, comprising: an identification unit, configured to identify fused information in a received interaction request to obtain a to-be-executed task corresponding to the interaction request, wherein the fused information includes at least one interactive input form; A disassembly unit, used for disassembling the task to be executed into at least one target subtask; A searching unit, used for searching for a task execution service corresponding to each of the at least one target subtask; A calling unit, used for calling the task execution service to execute the corresponding target subtask and obtain the target subtask execution result; The output unit is used to output and display the task execution result of the at least one target subtask.
27. A human-computer interaction device, the device comprising: at least one processor; A storage device for storing at least one program, which, when executed by the at least one processor, enables the device to implement the method as claimed in any one of claims 6 to 25.
28. A vehicle, comprising the human-computer interaction system according to any one of claims 1 to 5, the human-computer interaction apparatus according to claim 26, or the human-computer interaction device according to claim 27.
29. A non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause the computer to implement the method according to any one of claims 6 to 25 when executed.
30. A computer program product comprising a computer program, which, when executed by a processor, implements the method of any one of claims 6 to 25.
Citation Information
Patent Citations
Service task execution method, device and system
CN111597318A
Task generation method, system and equipment based on large language model and storage medium
CN116775183A
Task execution method, device, system and equipment and storage medium
CN117112082A
Interaction method and device, electronic equipment and storage medium
CN117130524A
Interactive timeline for presenting and organizing tasks
US20140012574A1