Vehicle control method, server, and computer-readable storage medium

By receiving in-vehicle voice requests, displaying content, and perceptual information, and using a preset model to determine the target action, the problem of limited interaction depth and efficiency of in-vehicle voice assistants has been solved, improving user experience and system intelligence.

CN119741923BActive Publication Date: 2025-12-09GUANGZHOU XIAOPENG MOTORS TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411858551.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-13
Publication Date
2025-12-09
Estimated Expiration
2044-12-13

AI Technical Summary

Technical Problem

Existing in-vehicle voice assistant designs mainly focus on general functions, which limits the depth and efficiency of user interaction with the in-vehicle voice assistant, affecting user experience, and lacks the ability to process the vehicle's current perception information.

Method used

By receiving voice requests, displayed content information, and perceived information forwarded by vehicles, the system uses a pre-defined model to determine the target action to be performed, including rewriting, analysis, and generating follow-up information, thereby improving the system's understanding and response capabilities.

Benefits of technology

It improves driving convenience and safety, reduces driver distraction, enhances user experience and system intelligence, and strengthens the understanding and interactivity of user needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119741923B_ABST
    Figure CN119741923B_ABST
Patent Text Reader

Abstract

The application discloses a vehicle control method, a server and a computer readable storage medium. The method comprises the following steps: receiving a first voice request forwarded by a vehicle, current display content information of a vehicle display component, and current perception information of the vehicle. Based on a preset model, a target execution action is determined according to the first voice request, the current display content information and the current perception information. The target execution action is issued to the vehicle to control the vehicle to execute the target execution action. In this way, by analyzing the content of the first voice request, understanding the current displayed content and the perception information of the vehicle, and determining the target execution action to be executed, the convenience and safety of driving are improved, the distraction of the driver is reduced, and the overall user experience is improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of voice interaction, and in particular to a vehicle control method, a server and a computer readable storage medium. BACKGROUND

[0002] In related technologies, a vehicle-mounted voice assistant is used to interact with a user to facilitate user operation. However, the design and training of the vehicle-mounted voice assistant mainly focus on general functions, which limits the depth and efficiency of user interaction with the vehicle-mounted voice assistant and affects user experience. SUMMARY

[0003] The present application provides a vehicle control method, a server and a computer readable storage medium.

[0004] The present application provides a vehicle control method, a server and a computer readable storage medium.

[0005] receiving a first voice request forwarded by a vehicle, current display content information of a vehicle display component, and current perception information of the vehicle;

[0006] determining a target execution action based on a preset model according to the first voice request, the current display content information and the current perception information;

[0007] issuing the target execution action to the vehicle to control the vehicle to execute the target execution action.

[0008] In this way, the server receives a first voice request forwarded by a vehicle, current display content information of a vehicle display component, and current perception information of the vehicle. Then, based on a preset model, the server determines a target execution action according to the first voice request, the current display content information and the current perception information. Finally, the server issues the target execution action to the vehicle to control the vehicle to execute the target execution action. In this way, by analyzing the content of the first voice request, understanding the currently displayed content and the perception information of the vehicle, and determining the target execution action to be executed, the convenience and safety of driving are improved, the distraction of the driver is reduced, and the overall user experience is improved.

[0009] In some embodiments, the determining of the target execution action based on the first voice request, the current display content information and the current perception information comprises:

[0010] performing rewriting processing on the current display content information, the first voice request and the current perception information based on a preset task rewriting template to obtain first input information;

[0011] determining the target execution action according to the first input information.

[0012] Thus, based on the preset task rewriting template, the server performs rewriting processing on the current display content information, the first voice request and the current perception information to obtain first input information. Then, the server determines a target execution action according to the first input information. In this way, through semantic understanding of the first voice request, analysis and interpretation of the display content, and integration and interpretation of the perception information, the server can convert the original information into a form that can be understood and processed by the system, i.e., the first input information, so as to determine the target execution action to be executed.

[0013] In some embodiments, the determining the target execution action according to the first input information comprises:

[0014] performing analysis processing on the first input information;

[0015] determining a target execution task and a task type of the target execution task according to a result of the analysis processing;

[0016] determining the target execution action according to the task type, the target execution task and the current display content information.

[0017] Thus, the server performs analysis processing on the first input information. Then, the server determines a target execution task and a task type of the target execution task according to a result of the analysis processing. Finally, the server determines the target execution action according to the task type, the target execution task and the current display content information. In this way, by performing analysis processing on the first input information and processing the result of the analysis processing, the target execution task and the task type of the target execution task are determined to determine the target execution action, so as to understand and respond to user demand and improve user experience.

[0018] In some embodiments, the determining the target execution action according to the task type, the target execution task and the current display content information comprises:

[0019] in a case where the task type is a first task type, determining the target execution action according to the target execution task and the current display content information, wherein the first task type is a task type with complete semantic information and a unique execution path.

[0020] Thus, in a case where the task type is a first task type, the server determines the target execution action according to the target execution task and the current display content information, wherein the first task type is a task type with complete semantic information and a unique execution path. In this way, in a case where the task type is the first task type, the server can quickly and accurately respond to user demand, improve user satisfaction and loyalty, and provide better service for users.

[0021] In some embodiments, the determining the target execution action according to the task type, the target execution task and the current display content information comprises:

[0022] In a case where the task type is a second task type, processing the target execution task and the current display content information to determine missing slot information, the second task type being a task type in which semantic information is missing;

[0023] determining the target execution action according to the missing slot information.

[0024] In this way, in a case where the task type is a second task type, the server processes the target execution task and the current display content information to determine missing slot information, the second task type being a task type in which semantic information is missing. Then, the server determines the target execution action according to the missing slot information. In this way, in a case where the task type is a second task type, the ability of the server to ask follow-up questions is improved by determining missing slot information and requesting the user to provide relevant information, thereby better understanding the user's demand and improving the intelligence and interactivity of the system.

[0025] In some embodiments, the processing the target execution task and the current display content information to determine missing slot information comprises:

[0026] determining, based on a preset task database, a target preset task and a target preset task slot corresponding to the first input information according to the first input information;

[0027] determining the missing slot information according to the first input information and the target preset task slot.

[0028] In this way, the server determines, based on a preset task database, a target preset task and a target preset task slot corresponding to the first input information according to the first input information. Then, the server determines the missing slot information according to the first input information and the target preset task slot. In this way, by using the preset task database, the system can determine a target preset task and a target preset task slot corresponding to the first input information, thereby improving the accuracy of task processing.

[0029] In some embodiments, the determining the target execution action according to the missing slot information comprises:

[0030] generating first follow-up information according to the missing slot information, and sending the follow-up information to the vehicle;

[0031] receiving second voice requests forwarded by the vehicle in response to the first follow-up information, the current display content information and the current perception information;

[0032] generating the target execution action based on the first voice request, the second voice request, the current display content information and the current perception information according to the preset model.

[0033] In this way, the server generates first follow-up information according to missing slot information, and issues the follow-up information to the vehicle. Then, the server receives the second voice request forwarded by the vehicle for the first follow-up information, the current display content information and the current perception information. Finally, based on the preset model, the server generates the target execution action according to the first voice request, the second voice request, the current display content information and the current perception information. In this way, by asking and collecting missing information, the system can more accurately understand and execute the user's task, thereby improving the accuracy of task completion.

[0034] In some embodiments, the determining the target execution action according to the task type, the target execution task and the current display content information comprises:

[0035] In the case where the task type is a third task type, processing the target execution task and the current display content information to determine second follow-up information, and issuing the second follow-up information to the vehicle, wherein the third task type is a task type with complete semantic information and non-unique execution path.

[0036] receiving third voice request forwarded by the vehicle for the second follow-up information, the current display content information and the current perception information;

[0037] generating the target execution action based on the first voice request, the third voice request, the current display content information and the current perception information according to the preset model.

[0038] In this way, the server processes the target execution task and the current display content information to determine the second follow-up information in the case where the task type is the third task type, and issues the second follow-up information to the vehicle, wherein the third task type is a task type with complete semantic information and non-unique execution path. Then, the server receives the third voice request forwarded by the vehicle for the second follow-up information, the current display content information and the current perception information. Finally, based on the preset model, the server generates the target execution action according to the first voice request, the third voice request, the current display content information and the current perception information. In this way, by processing the target execution task and the current display content information to determine the second follow-up information, the user's selection or preference is obtained through interaction, thereby improving the flexibility of task completion and improving the user experience.

[0039] In some embodiments, the method further comprises:

[0040] In the case where the vehicle completes the target execution action, a preset feedback generation template is used to generate voice broadcast information based on the current perception information, the voice broadcast information including execution completion state information of the target execution action and / or action suggestions.

[0041] In this way, in the case where the vehicle completes the target execution action, a preset feedback generation template is used to generate voice broadcast information based on the current perception information, the voice broadcast information including execution completion state information of the target execution action and / or action suggestions. In this way, by providing real-time execution completion state information and action suggestions, the system can help the driver better cope with various situations during driving, thereby improving the user experience.

[0042] The embodiments of the present application provide a server, comprising a processor and a memory, wherein the memory stores a computer program, and when the computer program is executed by the processor, the vehicle control method described above is implemented.

[0043] The embodiments of the present application provide a computer readable storage medium, which stores a computer program, and when the computer program is executed by a processor, the steps of the vehicle control method described above are implemented.

[0044] Additional aspects and advantages of the embodiments of the present application will be in part apparent and in part pointed out hereinafter. BRIEF DESCRIPTION OF DRAWINGS

[0045] The above and / or additional aspects and advantages of the present application will become apparent and be readily appreciated from the following description, including the drawings, in which:

[0046] Figure 1 is one of the flowcharts of the vehicle control method according to some embodiments of the present application;

[0047] Figure 2 is a flowchart of processing of voice requests according to some embodiments of the present application;

[0048] Figure 3 is another flowchart of the vehicle control method according to some embodiments of the present application;

[0049] Figure 4 is a third flowchart of the vehicle control method according to some embodiments of the present application;

[0050] Figure 5 is a fourth flowchart of the vehicle control method according to some embodiments of the present application;

[0051] Figure 6Fig. 5 is a flowchart of a vehicle control method according to some embodiments of the present application;

[0052] Figure 7 Fig. 6 is a flowchart of a vehicle control method according to some embodiments of the present application;

[0053] Figure 8 Fig. 7 is a flowchart of a vehicle control method according to some embodiments of the present application;

[0054] Figure 9 Fig. 8 is a flowchart of a vehicle control method according to some embodiments of the present application;

[0055] Figure 10 Fig. 9 is a flowchart of a vehicle control method according to some embodiments of the present application. DETAILED DESCRIPTION

[0056] The embodiments of the present application are described in detail below with reference to the accompanying drawings. Examples of the embodiments are shown in the drawings, wherein the same or similar reference numerals are used throughout to designate the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the drawings are exemplary and are for the purpose of explaining the embodiments of the present application and should not be understood as limiting the embodiments of the present application.

[0057] In the field of intelligent vehicles, in-vehicle voice assistants have become an important technology to improve driving experience and safety. These assistants can usually understand and respond to users' voice commands, perform general functions such as navigation, phone calls, music playback, etc. However, although these functions provide convenience for users, the design and training of in-vehicle voice assistants are often mainly focused on these general functions, which to some extent limits the depth and efficiency of user interaction with in-vehicle voice assistants, and thus affects the user experience.

[0058] First, since in-vehicle voice assistants are mainly trained for general functions, they cannot understand or respond to more complex or specific instructions from users, resulting in poor user experience due to poor interactivity between in-vehicle voice assistants and users.

[0059] Second, in-vehicle voice assistants lack the ability to process current perception information of the vehicle. For example, they may not be able to adjust the interaction mode or provide more appropriate suggestions based on the state information of the vehicle or changes in the driving environment.

[0060] Based on the above problems, please refer to Figure 1 The embodiments of the present application provide a vehicle control method, the method comprising:

[0061] 01: receiving a first voice request forwarded by a vehicle, current display content information of a vehicle display component, and current perception information of the vehicle;

[0062] 02: determining, based on the preset model, the target execution action according to the first voice request, the current display content information, and the current perception information;

[0063] 03: issuing the target execution action to the vehicle to control the vehicle to execute the target execution action.

[0064] The vehicle control method of the present application embodiment can be implemented by the server of the present application embodiment. Specifically, the memory stores a computer program, and the processor is configured to receive the first voice request forwarded by the vehicle, the current display content information of the display component of the vehicle, and the current perception information of the vehicle. Based on the preset model, the target execution action is determined according to the first voice request, the current display content information, and the current perception information. The target execution action is issued to the vehicle to control the vehicle to execute the target execution action.

[0065] The vehicle control method of the present application embodiment can be implemented by the vehicle control device of the present application embodiment. Specifically, the vehicle control device includes a receiving module, a determining module, and an issuing module. The receiving module is configured to receive the first voice request forwarded by the vehicle, the current display content information of the display component of the vehicle, and the current perception information of the vehicle. The determining module is configured to determine the target execution action based on the preset model according to the first voice request, the current display content information, and the current perception information. The issuing module is configured to issue the target execution action to the vehicle to control the vehicle to execute the target execution action.

[0066] Specifically, the first voice request refers to the voice instruction issued by the user through the vehicle-mounted voice assistant for the first time in a certain round of voice interaction, including a certain function control of the vehicle, information query, or service request. For example, "turn on the air conditioner", "what's the weather like today?", and "I need the nearest gas station".

[0067] The current display content information of the display component of the vehicle refers to the information content currently presented on the vehicle-mounted display screen or other display devices. For example, navigation-related content such as the current driving route, destination, estimated arrival time, and traffic conditions, or entertainment system-related content such as the music playlist, track information, and radio frequency, and communication-related content such as displaying incoming call information, SMS content, and contact list. Display content information is an important component in the vehicle intelligent control system, which provides key driving information and entertainment services for the driver, and is also an important interface for system and user interaction. It should be noted that the vehicle control method is explained based on the vehicle-mounted display screen as the display component of the vehicle in the present application embodiment, i.e., the current display content information is a screenshot of the vehicle-mounted display screen.

[0068] Current perception information refers to the current vehicle's environmental state information collected by its sensors and the state perception information of the vehicle's on-board device functions. Environmental state information includes data about the vehicle's surroundings, such as traffic conditions, road conditions, weather conditions, and geographic location. Traffic conditions refer to the road traffic situation collected by sensors such as cameras, radars, or LiDARs, including the positions and speeds of other vehicles. Road conditions refer to information such as the degree of wetness, potholes, and obstacles on the road. Weather conditions refer to weather information such as temperature, rainfall, and wind speed obtained through weather sensors or data exchange with external services. Geographic location refers to the current location information of the vehicle obtained through positioning systems such as GPS. On-board device function state perception information includes state data of the vehicle's internal systems and devices, such as the state of the on-board entertainment system, the settings of the navigation system, the air conditioning configuration, the engine state, the battery level, the oil level, and the tire pressure.

[0069] Target execution actions refer to specific operations determined and executed by the system based on the first voice request, current display content information, and current perception information, including click operations, swipe operations, and typing operations. Click operations refer to operations such as clicking on icons or buttons to activate specific functions or services. For example, if the user says "open music," the system may need to perform the operation of clicking on the music application icon in the on-board entertainment system. Swipe operations refer to operations such as swiping the screen or flipping pages to view more information when the current page information list is too long or folded. For example, when the navigation system displays too much route information, the system may need to perform a screen swipe operation to display the complete route. Typing operations refer to operations such as clicking on the search box and typing content to find specific information or services in the system. For example, if the user says "search for nearby restaurants," the system may need to perform the operation of clicking on the search box and entering "nearby restaurants." Typing operations can also indicate the input and sending of text information in the vehicle system, such as sending messages or making social media comments. For example, if the user says "send a message to Zhang San," the system may need to perform the operation of opening the message application, selecting the contact Zhang San, and opening the chat window.

[0070] The preset model refers to a pre-trained vision-language model (VLM) that can understand and process the relationship between images and text, enabling machines to understand both visual content and natural language text.

[0071] Please refer to Figure 2 The server receives the first voice request (i.e., the voice request) forwarded by the vehicle, the current display content information of the vehicle's display components, and the current perception information of the vehicle.

[0072] Then, the server uses the information as input, processes the information based on a preset model, i.e., analyzes the content of the first voice request, understands the current display content information and the current perception information of the vehicle, and then determines the target execution action to be performed.

[0073] Finally, the server issues the determined target execution action to the vehicle, and the system on the vehicle performs the corresponding operation after receiving the instruction.

[0074] To sum up, in the vehicle control method and the server provided by the embodiments of the present application, the server receives the first voice request forwarded by the vehicle, the current display content information of the display component of the vehicle, and the current perception information of the vehicle. Then, based on a preset model, the server determines the target execution action according to the first voice request, the current display content information, and the current perception information. Finally, the server issues the target execution action to the vehicle to control the vehicle to perform the target execution action. In this way, by analyzing the content of the first voice request, understanding the current display content and the perception information of the vehicle, and determining the target execution action to be performed, the convenience and safety of driving are improved, the distraction of the driver is reduced, and the overall user experience is improved.

[0075] For reference Figure 3 In some embodiments, step 02 (determining the target execution action according to the first voice request, the current display content information, and the current perception information) comprises:

[0076] 021: rewriting the current display content information, the first voice request, and the current perception information based on a preset task rewriting template to obtain first input information;

[0077] 022: determining the target execution action according to the first input information.

[0078] In some embodiments, the determining module is further configured to rewrite the current display content information, the first voice request, and the current perception information based on a preset task rewriting template to obtain first input information, and determine the target execution action according to the first input information.

[0079] In some embodiments, the processor is further configured to rewrite the current display content information, the first voice request, and the current perception information based on a preset task rewriting template to obtain first input information, and determine the target execution action according to the first input information.

[0080] Specifically, the preset task rewriting template refers to a pre-designed template or rule for converting or mapping the voice request of the user, the current display content information, and the current perception information into a task description that can be understood and executed by the system. In some embodiments, the preset task rewriting template is as follows:

[0081] Assuming you are an intelligent assistant, I will give you some information. Please rewrite the content based on these information. Output the rewriting task:

[0082] User task: <input request>

[0083] Current perception information: <environment state information and state perception information of vehicle-mounted device functions>

[0084] Current display content information: <screenshot of vehicle-mounted display screen>

[0085] Please reorganize the task into 20 words or less based on the above information:

[0086] Please note that if the server generates follow-up information, the server will also rewrite the second voice request based on the user's second voice request, that is, based on the preset task rewriting template, the current display content information, the first voice request, the current perception information and the second voice request are rewritten. Of course, the current display content information and the current perception information need to be re-acquired when receiving the second voice request.

[0087] Rewriting processing refers to the process of reorganizing and expressing the task based on the current perception information, current display content information and user input request.

[0088] Based on the preset task rewriting template, the server rewrites the current display content information, the first voice request and the current perception information to obtain the first input information, including semantic understanding of the voice request, analysis and interpretation of the display content, and integration and interpretation of the perception information. Then, the server determines the target execution action based on the first input information.

[0089] In this way, through semantic understanding of the first voice request, analysis and interpretation of the display content, and integration and interpretation of the perception information, the server can convert the original information into a form that the system can understand and process, i.e. the first input information, in order to determine the target execution action that needs to be executed.

[0090] Please refer to Figure 4 In some embodiments, step 022 (determining the target execution action based on the first input information) includes:

[0091] 0221: Analyze and process the first input information;

[0092] 0222: Determine the target execution task and the task type of the target execution task based on the results of the analysis and processing;

[0093] 0223: Determine the target execution action based on the task type, the target execution task and the current display content information.

[0094] In some embodiments, the determining module is further configured to analyze the first input information, and determine the target execution task and the task type of the target execution task according to the analysis result. The processor is further configured to determine the target execution action according to the task type, the target execution task and the current display content information.

[0095] In some embodiments, the determining module is further configured to analyze the first input information, and determine the target execution task and the task type of the target execution task according to the analysis result. The processor is further configured to determine the target execution action according to the task type, the target execution task and the current display content information.

[0096] Specifically, the task type refers to a way of classifying tasks according to the integrity of semantic information and the uniqueness of execution path of user requests, including a first task type, a second task type and a third task type. The first task type is a task type with complete semantic information and unique execution path. The second task type is a task type with missing semantic information. The third task type is a task type with complete semantic information and non-unique execution path.

[0097] The analysis refers to understanding the first input information to determine the user's demand and intention, i.e., to determine the target execution task and the task type of the target execution task.

[0098] The server analyzes the first input information. Then, the server determines the target execution task and the task type of the target execution task according to the analysis result. For example, if the analysis result shows that the user requests to navigate to a certain place, the target execution task may be "to provide navigation service", and the task type may be "the first task type".

[0099] Finally, the server determines the target execution action according to the task type, the target execution task and the current display content information.

[0100] For example, the driver says: "I need to go to A gas station", and the current display of the vehicle is a navigation interface. At this time, the server receives the voice request and the current display content information, and obtains the first input information "user requests to navigate to A gas station and select the nearest route, and the current display is the navigation interface" through rewriting processing. The server analyzes the first input information, determines that the target execution task is "to provide navigation service", and the task type is "the first task type". Finally, the server determines the target execution action as "enter A gas station in the search box of the navigation interface and select the nearest route".

[0101] Thus, by analyzing and processing the first input information and processing the result of the analysis and processing, the target execution task and the task type of the target execution task are determined to determine the target execution action, so as to understand and respond to the user demand and improve the user experience.

[0102] Referring to Figure 5 In some embodiments, the step 0223 (determining the target execution action according to the task type, the target execution task, and the current display content information) comprises:

[0103] 02231: In the case where the task type is the first task type, the target execution action is determined according to the target execution task and the current display content information.

[0104] In some embodiments, the determining module is further configured to, in the case where the task type is the first task type, determine the target execution action according to the target execution task and the current display content information.

[0105] In some embodiments, the processor is further configured to, in the case where the task type is the first task type, determine the target execution action according to the target execution task and the current display content information.

[0106] Specifically, the target execution action refers to a specific operation determined and executed by the system according to the user's request or the current context.

[0107] The server first determines the task type. If the task type is the first task type, that is, the demand of the task is very clear, and there is only one clear execution mode. Then the server determines the target execution task according to the first input information. For example, if the user request is "navigate to the nearest gas station", the target execution task is to "provide a route to the nearest gas station".

[0108] Then, the current display content information (i.e. the screenshot of the vehicle display screen) is analyzed by picture UI, and the target execution action is determined according to the result of the analysis and the target execution task. That is, how to integrate the target execution task into the current display content. For example, if the user request is "navigate to the nearest gas station", the current display content information is a song playing interface. Then the target execution action is "close the song playing interface", "open the navigation program", and "enter the gas station in the navigation interface".

[0109] Thus, in the case where the task type is the first task type, the server can quickly and accurately respond to the user demand, improve the user satisfaction and loyalty, and provide better service for the user.

[0110] Referring to Figure 6In some embodiments, the step 0223 (determining the target execution action according to the task type, the target execution task and the current display content information) comprises:

[0111] 02232: In the case where the task type is the second task type, processing the target execution task and the current display content information to determine the missing slot information;

[0112] 02233: determining the target execution action according to the missing slot information.

[0113] In some embodiments, the determining module is further configured to, in the case where the task type is the second task type, process the target execution task and the current display content information to determine the missing slot information, and determine the target execution action according to the missing slot information.

[0114] In some embodiments, the processor is further configured to, in the case where the task type is the second task type, process the target execution task and the current display content information to determine the missing slot information, and determine the target execution action according to the missing slot information.

[0115] Specifically, the missing slot information refers to the information that the system identifies as needing to be further obtained or confirmed in order to fully understand and respond to the user's intent when processing the user's request.

[0116] The server first determines the task type. If the task type is the "second task type", i.e., the task type in which there is missing semantic information, the server will process the target execution task and the current display content information to determine the missing slot information. For example, if the user's request is "navigate to the nearest", but the specific destination is not specified, the missing slot information may be "destination".

[0117] The server determines the target execution action according to the missing slot information. For example, if the missing slot information is "destination", the server will determine the target execution action according to the obtained missing slot information and other data information after obtaining the "destination".

[0118] In this way, in the case where the task type is the second task type, the server's follow-up ability is improved by confirming the missing slot information and requesting the user to provide related information, so that the user's needs can be better understood, and the intelligence and interactivity of the system are improved.

[0119] For reference Figure 7 In some embodiments, the step 02232 (processing the target execution task and the current display content information to determine the missing slot information) comprises:

[0120] 022321: Based on the preset task database, according to the first input information, confirm the target preset task and the target preset task slot corresponding to the first input information;

[0121] 022322: According to the first input information and the target preset task slot, determine the missing slot information.

[0122] In some embodiments, the determining module is configured to, based on the preset task database, according to the first input information, confirm the target preset task and the target preset task slot corresponding to the first input information; and according to the first input information and the target preset task slot, determine the missing slot information.

[0123] In some embodiments, the processor is further configured to, based on the preset task database, according to the first input information, confirm the target preset task and the target preset task slot corresponding to the first input information; and according to the first input information and the target preset task slot, determine the missing slot information.

[0124] Specifically, the preset task database refers to a database that stores a variety of preset tasks and corresponding slot information, which helps the system to understand and respond to the user's request. For example, if the user's target execution task is "navigation", the preset task slot may include "destination" (such as location A) and "route preference" (such as shortest distance or least red light, etc.). If the user's target execution task is "send a message", the preset task slot may include "sender" and "sending content". If the user's target execution task is "find nearby [facility type]", the preset task slot may include "facility type".

[0125] The server first confirms the target preset task corresponding to the received first input information based on the preset task database. At the same time, it identifies each slot in the target preset task, which represents all the necessary information required to complete the task.

[0126] The server then analyzes and determines which slot information is missing according to the first input information and the confirmed target preset task slot. For example, if the target preset task is "navigate to [destination]", and the specific [destination] information is not provided in the first input information, then [destination] is a missing slot. In some embodiments, the vehicle-mounted intelligent assistant is prompted by a certain template, as follows:

[0127] Suppose you are an agent planning assistant, please judge the execution type of the next action of the task according to the given first input information and reference the preset task slot. The execution type has two types: "ask again" and "direct execution", and the supplementary information is as follows:

[0128] 1. "Re-ask" refers to the current task lacking semantic information, and the sentence information does not meet the task slot information, and the next step cannot be executed.

[0129] 2. "Direct execution" refers to the current task being relatively clear, and the information in the sentence information can extract the task slot information, and it can be judged how to execute the next step.

[0130] Task: <first input information>

[0131] Task slot: <preset task slot>

[0132] Execution type:

[0133] In this way, by using the preset task database, the system can confirm the target preset task and the target preset task slot corresponding to the first input information, thereby improving the accuracy of task processing.

[0134] See Figure 8 In some embodiments, step 02233 (determining a target execution action according to missing slot information) includes:

[0135] 022331: generating first follow-up information according to missing slot information, and issuing the follow-up information to the vehicle;

[0136] 022332: receiving the second voice request forwarded by the vehicle for the first follow-up information, the current display content information and the current perception information;

[0137] 022333: generating a target execution action based on a preset model according to the first voice request, the second voice request, the current display content information and the current perception information.

[0138] In some embodiments, the determining module is configured to generate first follow-up information according to missing slot information, and issue the follow-up information to the vehicle. And receive the second voice request forwarded by the vehicle for the first follow-up information, the current display content information and the current perception information. And generate a target execution action based on a preset model according to the first voice request, the second voice request, the current display content information and the current perception information.

[0139] In some embodiments, the processor is further configured to generate first follow-up information according to missing slot information, and issue the follow-up information to the vehicle. And receive the second voice request forwarded by the vehicle for the first follow-up information, the current display content information and the current perception information. And generate a target execution action based on a preset model according to the first voice request, the second voice request, the current display content information and the current perception information.

[0140] Specifically, the first follow-up information refers to a question asked by the system to the user in order to obtain missing slot information or further clarify the user's intention when processing the user's request. Assuming the user says, "Help me find a nearby restaurant." The system identifies that this is a "find nearby facility" task and finds that the "facility type" slot information needs to be further obtained. Therefore, the system generates the first follow-up information: "What kind of restaurant do you want?" After the user answers, the system can obtain complete slot information and then perform the corresponding operation.

[0141] The second voice request refers to the user's response to the system's follow-up or further instruction after the first voice request. For example, the user requests: Help me find a nearby restaurant. The system asks: What kind of restaurant do you want? The user answers: I want to eat Chinese food. In this example, the user's first request is "Help me find a nearby restaurant," and the system obtains more information by asking "What kind of restaurant do you want?" The user's answer "I want to eat Chinese food" is the second voice request, which provides the key information needed for the system to complete the task. It should be noted that the current display content information and the current perception information received at this time are the display content information and the perception information reacquired when the second voice request is received.

[0142] The server generates the first follow-up information according to the determined missing slot information. For example, if the missing slot information is "destination", the follow-up information may be "Please tell me where you want to navigate to?". Then, the server sends the generated follow-up information to the vehicle. Then, the user will make a supplementary answer according to the follow-up information, that is, give the second voice request. Subsequently, the server receives the second voice request for the first follow-up information, the current display content information and the current perception information forwarded by the vehicle. Finally, based on the preset model, the server generates the target execution action in combination with the first voice request, the second voice request, the current display content information and the current perception information. For example, if the user provides destination information in the second voice request, the server will generate instructions to navigate to the destination according to these information.

[0143] In this way, by asking and collecting missing information, the system can more accurately understand and execute the user's task, thereby improving the accuracy of task completion.

[0144] Please refer to Figure 9 In some embodiments, step 0223 (determining the target execution action according to the task type, the target execution task and the current display content information) comprises:

[0145] 02234: In the case where the task type is the third task type, processing the target execution task and the current display content information, confirming the second follow-up information, and sending the second follow-up information to the vehicle;

[0146] 02235: receiving a third voice request for the second follow-up information forwarded by the vehicle, the current display content information, and the current perception information;

[0147] 02236: generating a target execution action based on the preset model according to the first voice request, the third voice request, the current display content information, and the current perception information.

[0148] In some embodiments, the determining module is configured to, in a case where the task type is a third task type, process the target execution task and the current display content information, confirm the second follow-up information, and issue the second follow-up information to the vehicle. The third voice request for the second follow-up information forwarded by the vehicle, the current display content information, and the current perception information are received. The target execution action is generated based on the preset model according to the first voice request, the third voice request, the current display content information, and the current perception information.

[0149] In some embodiments, the processor is further configured to, in a case where the task type is a third task type, process the target execution task and the current display content information, confirm the second follow-up information, and issue the second follow-up information to the vehicle. The third voice request for the second follow-up information forwarded by the vehicle, the current display content information, and the current perception information are received. The target execution action is generated based on the preset model according to the first voice request, the third voice request, the current display content information, and the current perception information.

[0150] Specifically, the second follow-up information refers to a situation in which, when processing a user request, there are multiple possible execution paths, and the system needs to ask the user further questions to determine which specific execution path should be taken. For example, a user requests: Help me find a nearby restaurant. The system asks: What kind of restaurant do you want? The user answers: I want to eat Chinese food. The system will ask again: Do you want traditional Chinese food or modern Chinese food? In this example, the user's first request is "Help me find a nearby restaurant", and the system obtains more information through the first follow-up information "What kind of restaurant do you want?" After the user's answer "I want to eat Chinese food", the system finds that there are many different choices under the category of Chinese food, so it further clarifies the user's specific needs through the second follow-up information "Do you want traditional Chinese food or modern Chinese food?"

[0151] The server processes the target execution task and the current display content information in the case of the task type being a third task type, confirms second follow-up information, and issues the second follow-up information to the vehicle, where the third task type is a task type with complete semantic information and non-unique execution path. Then, the server receives a third voice request for the second follow-up information forwarded by the vehicle, current display content information, and current perception information. Finally, the server generates a target execution action based on a preset model according to the first voice request, the third voice request, the current display content information, and the current perception information.

[0152] In this way, by processing the target execution task and the current display content information, confirming the second follow-up information, and obtaining the selection or preference of the user through interaction, the flexibility of task completion is improved, and the user experience is improved.

[0153] Referring to Figure 10 In some embodiments, the method further includes:

[0154] 04: In the case where the vehicle completes the target execution action, a preset feedback generation template is generated based on the current perception information, and voice broadcast information is generated.

[0155] In some embodiments, the generation module is configured to generate voice broadcast information based on a preset feedback generation template in the case where the vehicle completes the target execution action, and based on the current perception information.

[0156] In some embodiments, the processor is further configured to generate voice broadcast information based on a preset feedback generation template in the case where the vehicle completes the target execution action, and based on the current perception information.

[0157] Specifically, the preset feedback generation template refers to a series of pre-designed and personalized voice feedback templates generated by the system based on the current perception information of the vehicle and the completed task after the target execution action is executed. For example, the user task is: help me navigate to the Fragrant Hills and book a hotel nearby. The current perception information is: the outdoor temperature is 10 degrees / the air conditioner heating is turned on. The output voice broadcast information tts is: successfully navigated to the Fragrant Hills and booked a hotel nearby. The outdoor temperature is 10 degrees, the air conditioner heating is turned on, and the outdoor temperature is relatively cold. If you go outside, remember to keep warm. In some embodiments, the preset feedback generation template is as follows:

[0158] Suppose you are a car machine intelligent assistant, please you according to the following tasks and car machine state, please you according to these information, output suitable task broadcast tts:

[0159] Examples:

[0160] User task: help me navigate to the Fragrant Hills and book a hotel nearby

[0161] Current perception information: outside temperature 10 degrees / air conditioning heating on

[0162] Output tts: You have successfully navigated to Fragrant Hills and booked a nearby hotel. The outside temperature is 10 degrees, and the air conditioning heating is on. It is quite cold outside, so remember to keep warm if you go outside.

[0163] User task: <input voice request>

[0164] Current perception information: <input current perception information>

[0165] Please output the task completion broadcast tts based on the above information, with lively language style and within 60 words:

[0166] Output tts:

[0167] The server first confirms that the vehicle has completed the target execution action issued previously. Then, based on the preset feedback generation template, the server generates voice broadcast information according to the current perception information, which describes different execution completion states and provides action suggestions. For example, if the target execution action is to navigate to a certain location, the voice broadcast information may be "You have arrived at the destination, please pay attention to the safety of the surrounding environment when parking."

[0168] In this way, by providing real-time execution completion state information and action suggestions, the system can help the driver better cope with various situations during driving, thereby improving the user experience.

[0169] The present application also provides a computer readable storage medium having a computer program stored thereon. When the computer program processor is executed, the steps of the vehicle control method as described above are implemented.

[0170] It can be understood that the computer program includes computer program code. The computer program code can be in the form of source code, object code, executable files, or some intermediate forms, etc. The computer readable storage medium can include any entity or device capable of carrying computer program code, recording medium, U disk, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), and software distribution medium, etc.

[0171] In the description of the specification, the description of the terms "specifically", "further", "particularly", "understandably" and the like means that the specific features, structures, materials or characteristics described in connection with the embodiments or examples are included in at least one embodiment or example of the present application. In the present specification, the illustrative expression of the above terms does not intend to refer to the same embodiment or example. Also, the specific features, structures, materials or characteristics described can be combined in any one or more embodiments or examples in a suitable manner. Furthermore, the person skilled in the art can combine and combine the different embodiments or examples described in the specification and the features of the different embodiments or examples without contradiction.

[0172] Any process or method descriptions or descriptions of the flow diagrams in the flow charts described herein and elsewhere can be understood as representing the steps of the executable request code of the modules, segments or portions for performing specific logic functions or steps in the processes, and the scope of the preferred embodiments of the present application includes additional implementation in which the functions can be performed in the same order as shown or discussed, in a substantially simultaneous manner or in a reverse order, as appropriate, depending on the functionality involved, as would be understood by those reasonably skilled in the art.

[0173] Although the embodiments of the present application have been shown and described above, it is understood that the above-described embodiments are exemplary and are not to be construed as limiting the present application, and those ordinarily skilled in the art can make changes, modifications, substitutions and variations to the above-described embodiments within the scope of the present application.

Claims

1. A vehicle control method characterized by, The method comprises: receiving a first voice request forwarded by a vehicle, current display content information of a vehicle display component, and current perception information of the vehicle; determining a target execution action based on a preset model according to the first voice request, the current display content information, and the current perception information; issuing the target execution action to the vehicle to control the vehicle to execute the target execution action; the determining of the target execution action according to the first voice request, the current display content information, and the current perception information comprises: rewriting the current display content information, the first voice request, and the current perception information based on a preset task rewriting template to obtain first input information; determining the target execution action according to the first input information.

2. The vehicle control method according to claim 1, characterized by, the determining of the target execution action according to the first input information comprises: analyzing and processing the first input information; determining a target execution task and a task type of the target execution task according to a result of the analysis and processing; determining the target execution action according to the task type, the target execution task, and the current display content information.

3. The vehicle control method according to claim 2, characterized by, the determining of the target execution action according to the task type, the target execution task, and the current display content information comprises: in a case where the task type is a first task type, determining the target execution action according to the target execution task and the current display content information, wherein the first task type is a task type with complete semantic information and a unique execution path.

4. The vehicle control method according to claim 2, characterized by the determining of the target execution action according to the task type, the target execution task, and the current display content information comprises: in a case where the task type is a second task type, processing the target execution task and the current display content information to confirm missing slot information, wherein the second task type is a task type with missing semantic information; determining the target execution action according to the missing slot information.

5. The vehicle control method according to claim 4, characterized by the processing of the target execution task and the current display content information to confirm the missing slot information comprises: confirming a target preset task and a target preset task slot corresponding to the first input information based on a preset task database according to the first input information; determining the missing slot information according to the first input information and the target preset task slot.

6. The vehicle control method according to claim 4, characterized by the determining of the target execution action according to the missing slot information comprises: generating first follow-up information according to the missing slot information, and issuing the follow-up information to the vehicle; receiving a second voice request forwarded by the vehicle in response to the first follow-up information, the current display content information, and the current perception information; generating the target execution action based on the preset model according to the first voice request, the second voice request, the current display content information, and the current perception information.

7. The vehicle control method according to claim 2, characterized by the determining of the target execution action according to the task type, the target execution task, and the current display content information comprises: In a case where the task type is a third task type, the target execution task and the current display content information are processed, second follow-up information is confirmed, and the second follow-up information is sent to the vehicle, where the third task type is a task type with complete semantic information and non-unique execution path; receiving a third voice request for the second follow-up information, the current display content information, and the current perception information forwarded by the vehicle; based on the preset model, generating the target execution action according to the first voice request, the third voice request, the current display content information, and the current perception information.

8. The vehicle control method according to claim 1, characterized by The method further comprises: in a case where the vehicle completes the target execution action, generating a voice broadcast information based on a preset feedback generation template and according to the current perception information, the voice broadcast information including execution completion state information of the target execution action and / or action suggestion.

9. A server, characterized by The server comprises a processor and a memory, and the memory stores a computer program, which, when executed by the processor, implements the vehicle control method of any one of claims 1-8.

10. A computer-readable storage medium having stored thereon a computer program, characterized in that, The program, when executed by the processor, implements the steps of the method of any one of claims 1-8.

Citation Information

Patent Citations

  • Voice interaction method, vehicle and computer storage medium

    CN112017667A

  • Voice interaction method and device, electronic equipment and vehicle

    CN117198281A