Voice interaction method, server and computer readable storage medium
By introducing voice interaction methods in the vehicle voice assistant, including obtaining voice requests, determining command categories and target operation instructions, the problem of insufficient reasoning capabilities in the prior art is solved and the user experience is improved.
Patent Information
- Application Number
- CN202510214153.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-25
- Publication Date
- 2025-05-13
AI Technical Summary
When existing vehicle voice assistants process voice requests that users express their feelings, they lack reasoning ability and cannot provide accurate and applicable feedback, resulting in poor user experience.
By obtaining voice requests, determining the instruction category, determining the target operation instructions based on the preset large language model and knowledge base, and sending instructions to the vehicle to realize voice interaction.
Improves the user experience and provides natural and convenient user interaction by intelligently responding to user voice commands.
Smart Images

Figure CN119993148A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of voice interaction technology, and in particular to a voice interaction method, a server and a computer-readable storage medium. Background Art
[0002] In the related art, the in-vehicle voice assistant is used to interact with the user to facilitate the user's operation. However, for the voice request of the user to express his feelings, the reasoning ability of the in-vehicle voice assistant is often unable to support the provision of accurate and applicable feedback, resulting in a poor user experience. Summary of the invention
[0003] The present application provides a voice interaction method, a server and a computer-readable storage medium.
[0004] The present application provides a voice interaction method, the method comprising:
[0005] Get voice request;
[0006] Determining a command category according to the voice request;
[0007] Determining a target operation instruction according to the voice request and the instruction category;
[0008] A target operation instruction is sent to the vehicle, so that the vehicle performs the voice interaction according to the target operation instruction.
[0009] In this way, the server receives the voice request. Next, the server determines the instruction category based on the voice request. Then, the server determines the target operation instruction based on the voice request and the instruction category. Finally, the server sends the target operation instruction to the vehicle so that the vehicle performs voice interaction according to the target operation instruction. In this way, by understanding and responding to the received voice request, the server can determine the instruction category of the voice request and provide different voice request processing methods according to different instruction categories, thereby improving the user experience.
[0010] In some implementations, determining the instruction category according to the voice request includes:
[0011] Based on a preset large language model, the instruction category is determined according to the voice request.
[0012] In this way, based on the preset large language model, the server determines the instruction category according to the voice request. In this way, the server can understand the user's voice request and classify it into the appropriate instruction category for subsequent processing and response, so as to intelligently respond to the user's voice instruction and provide a natural and convenient user experience.
[0013] In some implementations, determining the target operation instruction according to the voice request and the instruction category includes:
[0014] In the case where the instruction category is the first target category, determining a reference function point according to the voice request based on a preset knowledge base;
[0015] The target operation instruction is determined according to the voice request and the reference function point.
[0016] In this way, when the instruction category is the first target category, based on the preset knowledge base, the server determines the reference function point according to the voice request. Then, the server determines the target operation instruction according to the voice request and the reference function point. In this way, based on the preset knowledge base, by analyzing and identifying the voice request, determining the reference function point, and according to the determined reference function point and voice request, the server can generate a target operation instruction that accurately meets the user's needs, thereby improving the user experience.
[0017] In some implementations, determining the target operation instruction according to the voice request and the reference function point includes:
[0018] According to the reference function point, obtaining current state perception information of the vehicle components corresponding to the reference function point;
[0019] The target operation instruction is determined according to the voice request, the reference function point and the current state perception information.
[0020] In this way, the server obtains the current state perception information of the vehicle parts corresponding to the reference function point based on the reference function point. Then, the server determines the target operation instruction based on the voice request, the reference function point and the current state perception information. In this way, by obtaining the current state perception information and performing a comprehensive analysis based on the voice request and the reference function point, the server can accurately understand the user's instruction target and select the appropriate target operation instruction to meet the user's needs, thereby improving the user experience.
[0021] In some implementations, determining the target operation instruction according to the voice request, the reference function point, and the current state perception information includes:
[0022] Determining a target function point according to the reference function point and the current state perception information;
[0023] The target operation instruction is determined according to the voice request and the target function point.
[0024] In this way, the server determines the target function point based on the reference function point and the current state perception information. Then, the server determines the target operation instruction based on the voice request and the target function point. In this way, by determining the target function point, the server can accurately focus on the function that can meet the user's needs, and perform detailed analysis and reasoning, so as to generate appropriate target operation instructions to meet the user's needs and improve the user experience.
[0025] In some implementations, determining the target operation instruction according to the voice request and the target function point includes:
[0026] Perform slot recognition according to the voice request to obtain a first entity;
[0027] According to the first entity and the target function point, target function point parameter filling is performed to determine the target operation instruction.
[0028] In this way, the server performs slot recognition according to the voice request and obtains the first entity. Then, the server performs parameter filling of the target function point according to the first entity and the target function point and determines the target operation instruction. In this way, through slot recognition and parameter filling, the server can accurately understand the user's instruction intention and generate a target operation instruction that meets the user's expectations, thereby improving the user experience.
[0029] In certain embodiments, the method further comprises:
[0030] When the instruction category is the first target category, a temporary voice broadcast feedback is generated and the temporary voice broadcast feedback is sent to the vehicle.
[0031] In this way, when the instruction category is the first target category, the server generates a temporary voice broadcast feedback and sends the temporary voice broadcast feedback to the vehicle. In this way, by generating temporary voice broadcast feedback, the server can alleviate the delay in the instruction processing process and improve the user experience.
[0032] In some implementations, generating a target operation instruction according to the instruction category includes:
[0033] When the instruction category is the second target category, performing slot recognition on the voice request to determine a second entity;
[0034] Performing application program interface prediction on the voice request;
[0035] According to the second entity and the predicted application program interface, application program interface parameter filling is performed to determine the target operation instruction.
[0036] In this way, when the instruction category is the second target category, the server performs slot identification on the voice request and determines the second entity. Then, the server performs application program interface prediction on the voice request. Finally, according to the second entity and the predicted application program interface, the application program interface parameter filling is performed to determine the target operation instruction. In this way, through slot identification, application program interface prediction and application program interface parameter filling, the server can efficiently execute user instructions and generate target operation instructions that meet user expectations, thereby improving the user experience.
[0037] An embodiment of the present application provides a server, which includes a processor and a memory, wherein a computer program is stored in the memory, and when the computer program is executed by the processor, the above-mentioned voice interaction method is implemented.
[0038] An embodiment of the present application provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the steps of the voice interaction method described above are implemented.
[0039] Additional aspects and advantages of the embodiments of the present application will be given in part in the description below, and in part will become apparent from the description below, or will be learned through the practice of the embodiments of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] The above and / or additional aspects and advantages of the present application will become apparent and easily understood from the description of the embodiments in conjunction with the following drawings, in which:
[0041] Figure 1 It is one of the flowcharts of the voice interaction method of certain embodiments of the present application;
[0042] Figure 2 is a schematic diagram of a processing flow of a voice request in certain embodiments of the present application;
[0043] Figure 3 This is the second flow chart of the voice interaction method of certain implementation modes of the present application;
[0044] Figure 4 This is a schematic diagram of the processing flow of a traditional large language model;
[0045] Figure 5 It is a schematic diagram of the processing flow of a preset large language model in certain embodiments of the present application;
[0046] Figure 6 This is the third flow chart of the voice interaction method of certain implementation modes of the present application;
[0047] Figure 7 This is a fourth flowchart of a voice interaction method according to certain embodiments of the present application;
[0048] Figure 8 This is a fifth flow chart of a voice interaction method according to certain embodiments of the present application;
[0049] Fig. 9 This is the sixth flow chart of the voice interaction method of certain implementation modes of the present application;
[0050] Fig.10 This is the seventh flow chart of the voice interaction method of certain implementation modes of the present application;
[0051] Fig.11 This is the eighth flow chart of the voice interaction method of certain embodiments of the present application. DETAILED DESCRIPTION
[0052] The embodiments of the present application are described in detail below, and examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals represent the same or similar elements or elements having the same or similar functions from beginning to end. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the embodiments of the present application, and cannot be understood as limiting the embodiments of the present application.
[0053] In smart car systems, voice assistants, as a convenient way of human-computer interaction, have greatly improved the driving experience and safety. However, traditional in-vehicle voice assistants often have insufficient reasoning capabilities when processing voice requests from users to express their feelings, resulting in an inability to provide accurate and applicable feedback, which affects the user experience. For example, when a user says "I feel a little cold", the voice assistant may simply lower the air conditioning temperature without considering factors such as whether the windows are open or the seat heating function is turned on, resulting in no real improvement in the user's feelings and a poor user experience.
[0054] Based on the above questions, please refer to Figure 1 , the present application embodiment provides a voice interaction method, the method comprising:
[0055] 01: Receive voice request;
[0056] 02: Determine the command category based on the voice request;
[0057] 03: Determine the target operation instruction based on the voice request and instruction category;
[0058] 04: Send target operation instructions to the vehicle so that the vehicle can perform voice interaction according to the target operation instructions.
[0059] The embodiment of the present application also provides a server, including a memory and a processor. The voice interaction method of the embodiment of the present application can be implemented by the server of the embodiment of the present application. Specifically, a computer program is stored in the memory, and the processor is used to receive a voice request. And determine the instruction category according to the voice request. The processor is also used to determine the target operation instruction according to the voice request and the instruction category. And send the target operation instruction to the vehicle so that the vehicle performs voice interaction according to the target operation instruction.
[0060] The embodiment of the present application also provides a voice interaction device. The voice interaction method of the embodiment of the present application can be implemented by the voice interaction device of the embodiment of the present application. Specifically, the voice interaction device includes a receiving module, a determining module and a sending module. The receiving module is used to receive a voice request. The determining module is used to determine the instruction category according to the voice request. The determining module is also used to determine the target operation instruction according to the voice request and the instruction category. The sending module is used to send the target operation instruction to the vehicle so that the vehicle performs voice interaction according to the target operation instruction.
[0061] Specifically, a voice request refers to a voice request issued by a user to the in-vehicle voice assistant through voice, which can cover various functions, such as control functions, query functions, and setting functions, including perceptual voice requests. Perceptual voice requests refer to the user's perceptual description of the in-vehicle environment perceived by the senses (such as vision, hearing, touch, etc.), such as in-vehicle temperature, in-vehicle humidity, in-vehicle light, seat comfort, and noise level. In-vehicle temperature refers to the in-vehicle air temperature perceived by the user, such as: cold, hot, warm, etc. In-vehicle humidity refers to the in-vehicle air humidity perceived by the user, such as: dry, damp, etc. In-vehicle light refers to the in-vehicle light intensity perceived by the user, such as: bright, dim, etc. Seat comfort refers to the seat comfort perceived by the user, such as: soft, hard, comfortable, etc. Noise level refers to the in-vehicle noise level perceived by the user, such as: quiet, noisy, etc. The current environmental perception information can help the server better understand the user's command intention. For example, when the user says "I'm a little cold", the server can judge that the user needs to increase the air conditioning temperature based on the temperature inside the car perceived by the user.
[0062] The instruction category includes a first target category and a second target category, which are obtained by classifying the intention of the user's voice request. By classifying the instructions, the user's instruction intention can be better understood and more accurate services can be provided.
[0063] Target operation instructions refer to specific operation instructions generated in the vehicle intelligent system based on voice requests and instruction categories, which can guide the vehicle to perform specific operations or services to meet the needs of users. For example, if the user perception information is "I feel too hot", the target operation instruction generated is "increase the air volume of the air conditioner". It should be noted that there may be multiple target operation instructions, which are not limited here.
[0064] See also Figure 2 The vehicle collects voice requests and forwards them to the server. Then, the server determines the category of the command based on the received voice request. For example, when the user says "I feel too hot" in an environment with the air-conditioning temperature at 22.5 degrees, the air-conditioning volume at level 5, and the seat ventilation turned off, the server will analyze the information and determine that the command category of the current user's needs is the first target category.
[0065] Then, once the instruction category is determined, the server further determines the target operation instruction based on the voice request and the instruction category. Continuing with the above example, the server combines the voice request and the instruction category, and the target operation instruction determined may be "turn on the seat ventilation".
[0066] Finally, the server sends the target operation instructions to the vehicle. After receiving the instructions, the vehicle performs corresponding operations, such as adjusting the air-conditioning temperature, thereby completing the voice interaction.
[0067] In summary, in the voice interaction method and server provided in the embodiment of the present application, the server receives a voice request forwarded by a vehicle. Next, the server determines the instruction category based on the voice request. Then, the server determines the target operation instruction based on the voice request and the instruction category. Finally, the server sends the target operation instruction to the vehicle to complete the voice interaction. In this way, by understanding and responding to the received voice request, the server can determine the instruction category of the voice request and provide services based on the instruction category, thereby improving the user experience.
[0068] See also Figure 3 In some embodiments, step 02 (determining the instruction category according to the voice request) includes:
[0069] 021: Based on the preset large language model and voice request, determine the command category.
[0070] In some implementations, the determination module is further configured to determine the instruction category based on a preset large language model and according to the voice request.
[0071] In some embodiments, the processor is further configured to determine the instruction category according to the voice request based on a preset large language model.
[0072] Specifically, the preset large language model refers to a pre-trained large language model that can be used to process user voice requests, including identifying the instruction category of the user's voice request, obtaining reference function points, determining target function points, and performing natural language understanding and reasoning on the user's voice request. In some embodiments, the functions implemented by the above-mentioned preset large language model can be implemented by multiple pre-trained large language models respectively. For example, the preset model can be used to implement the two functions of identifying the instruction category of the user's voice request and performing natural language understanding and reasoning on the user's voice request, and the perception model can be used to implement the two functions of obtaining reference function points and determining target function points.
[0073] See also Figure 4 , Figure 4 This is a schematic diagram of the processing flow of a traditional large language model for perceptual voice requests. In related technologies, the large language model processes perceptual voice requests in the same way as non-perceptual voice requests, which results in the inference ability of the in-vehicle voice assistant often being unable to support the provision of accurate and applicable feedback, resulting in a poor user experience. Figure 5 , Figure 5 The schematic diagram of the processing flow of the preset large language model for perceptual voice requests. The preset large language model provided in the implementation mode of the present application can obtain the current state perception information of vehicle parts and generate multiple target operation instructions to facilitate the determination of subsequent steps, thereby enhancing the reasoning ability of perceptual voice requests and improving the user experience.
[0074] First, the preset large language model is used to parse the user's voice request to understand the user's intention. Then, the category of the instruction is determined in combination with environmental perception information, such as the temperature inside the car, the weather outside the car, etc. In some embodiments, the user's voice tone, historical preferences and other information are also combined to jointly determine the instruction category of the user's voice request. For example, if the user says "I feel a little cold", the system will determine the instruction category as the "first target category". If the user says "turn on the air conditioner", the system will determine the instruction category as the "second target category".
[0075] It should be noted that in some implementations, the server will also determine the command category of the user's voice request based on environmental information of the user's current environment obtained by various sensors on the vehicle.
[0076] In this way, the server can understand the user's voice request and classify it into the appropriate command category for subsequent processing and response, so that it can intelligently respond to the user's voice command and provide a natural and convenient user experience.
[0077] See also Figure 6 In some embodiments, step 03 (determining the target operation instruction according to the voice request and the instruction category) includes:
[0078] 031: When the instruction category is the first target category, based on the preset knowledge base and according to the voice request, determine the reference function point;
[0079] 032: Determine the target operation instructions based on the voice request and reference function points.
[0080] In some embodiments, the determination module is further configured to determine a reference function point based on a preset knowledge base and a voice request when the instruction category is the first target category, and to determine a target operation instruction based on the voice request and the reference function point.
[0081] In some embodiments, the processor is further configured to determine a reference function point based on a preset knowledge base and a voice request when the instruction category is the first target category, and determine a target operation instruction based on the voice request and the reference function point.
[0082] Specifically, the first target category refers to the perceptual instruction category, that is, after the preset large language model recognizes the user's voice request, it is confirmed that the user's voice request is the user's feeling about the current environment, rather than a specific demand. For example, "I'm a little cold", "It's a little too cold", "My back is stuffy" and "I'm sweating" are all voice requests of the first target category.
[0083] The preset knowledge base refers to a database containing a large amount of structured information, which can support the decision-making and operation of the system, that is, it provides the information basis required to process user needs and determine reference function points. By analyzing the user perception information and the current environmental perception information of the user's environment, the server can determine the appropriate reference function points from the preset knowledge base for subsequent processes. In some embodiments, the data in the preset knowledge base may include:
[0084] Auditory perception category: [Voice request: It’s too noisy; reference function points: close the windows, turn down the media volume, turn down the navigation volume, turn down the voice broadcast volume], [Voice request: The sound is too loud; reference function points: turn down the media volume, turn down the navigation volume, turn down the voice broadcast volume, close the windows], [Voice request: I can’t hear the navigation clearly; reference function points: turn up the navigation volume, turn down the media volume, turn down the voice broadcast volume, turn up and close the windows], [Voice request: I want it louder; reference function points: turn up the navigation volume, turn up the media volume, turn up the voice broadcast volume], [Voice request: The music is too noisy; reference function points: turn down the music volume, pause music playback, switch music type].
[0085] Temperature perception category: [Voice request: It’s so cold; reference function points: close the windows, increase the air conditioning temperature, increase the air conditioning air volume, turn on the seat heating, turn up the seat heating], [Voice request: It’s too cold; reference function points: close the windows, increase the air conditioning temperature, increase the air conditioning air volume, turn on the seat heating, turn up the seat heating], [Voice request: It’s so hot; reference function points: lower the air conditioning temperature, increase the air conditioning air volume, turn off the seat heating, turn on the seat ventilation, open the windows], [Voice request: My butt is so hot; reference function points: turn down the seat heating, turn on the seat ventilation, turn up the seat ventilation], [Voice request: My butt is so cold; reference function points: turn on the seat ventilation, turn down the seat ventilation, turn up the seat heating].
[0086] Olfactory perception category: [Voice request: It stinks; reference function points: open the car window, turn on the air purifier, turn on the car aromatherapy system, turn on the air conditioning external circulation], [Voice request: The air is not fresh; reference function points: open the car window, open the air purifier, turn on the car aromatherapy system, turn on the air conditioning external circulation], [Voice request: I want some fragrance; reference function points: turn on the car aromatherapy system, turn on the air purifier, open the car window, turn on the air conditioning external circulation], [Voice request: No need for deodorization; reference function points: open the air purifier, close the car window, turn off the air conditioning external circulation, turn off the car aromatherapy system].
[0087] Visual perception category: [Voice request: The road is too dark; reference function points: turn on the headlights, adjust the headlight brightness, turn on the fog lights or auxiliary lights], [Voice request: It is too bright; reference function points: open the sun visor, turn off the reading light, turn off the ceiling light, adjust the dashboard brightness to the minimum, turn off the headlights], [Voice request: The view behind is not clear; reference function points: use the reversing image system, turn on the rear windshield defogger, turn on the rear windshield defrost, turn on the headlights], [Voice request: The reflection in the rearview mirror is not clear; reference function points: turn on the air conditioning defogger, turn on the rearview mirror heating (if available), adjust the air conditioning wind direction to the rearview mirror, and turn on the vehicle's external circulation].
[0088] Likes and dislikes: [Voice request: I don’t like these songs very much; Reference function points: Change the music playlist, switch the music style, enable random play mode, ask the user about music preferences].
[0089] Comfort perception category: [Voice request: I want to take a break; reference function points: adjust the seat angle to a comfortable lying position, turn off the lights in the car, close the windows, adjust the air-conditioning temperature to a suitable temperature for rest, and adjust the air-conditioning volume to a moderate level], [Voice request: I am so tired; reference function points: adjust the seat angle to a comfortable lying position, turn off the lights in the car, close the windows, and adjust the air-conditioning temperature], [Voice request: It is too crowded; reference function points: adjust the seat back and the backrest backward], [Voice request: The steering wheel is so far away; reference function points: adjust the steering wheel back, adjust the seat forward, and adjust the seat higher].
[0090] The reference function point refers to the vehicle function and related operation information obtained by the preset large language model from the preset knowledge base according to the voice request, which can be used to process or realize the current user intention.
[0091] Please refer to Figure 2 , when the instruction category is the first target category, based on the preset knowledge base, the server determines the reference function point according to the voice request. For example, continuing the above example, the user's voice request is "I feel too hot", and after the server recognizes the voice request "I feel too hot", it confirms that the user's voice request is the first target category. Subsequently, the server will search for relevant reference function points based on the preset knowledge base, and the reference function points obtained are "lower the air conditioning temperature, increase the air conditioning volume, turn off the seat heating, turn on the seat ventilation and open the window".
[0092] Next, the server determines the target operation instruction based on the voice request and reference function points.
[0093] In this way, based on the preset knowledge base, by analyzing and identifying the voice request, the reference function points are determined, and according to the determined reference function points and voice requests, the server can generate target operation instructions that accurately meet the user's needs, thereby improving the user experience.
[0094] See also Figure 7 In some implementations, step 032 (determining a target operation instruction according to a voice request and a reference function point) includes:
[0095] 0321: According to the reference function point, obtain the current state perception information of the vehicle parts corresponding to the reference function point;
[0096] 0322: Determine the target operation instructions based on the voice request, reference function points and current state perception information.
[0097] In some embodiments, the voice interaction device further includes an acquisition module, which is further used to acquire the current state perception information of the vehicle components corresponding to the reference function point according to the reference function point, and determine the target operation instruction according to the voice request, the reference function point and the current state perception information.
[0098] In some embodiments, the processor is further configured to obtain, based on the reference function point, current state perception information of the vehicle components corresponding to the reference function point, and determine the target operation instruction based on the voice request, the reference function point and the current state perception information.
[0099] Specifically, the current state perception information of vehicle components refers to the specific working state or performance data of each vehicle component or system at the current moment collected by the vehicle's built-in sensors and monitoring systems. For example, the air conditioning system status: including the current temperature setting, wind speed, internal and external circulation mode, etc., or the seat status: including seat position, temperature, massage function status, etc.
[0100] Based on the reference function points, the server obtains the current state perception information of the vehicle components corresponding to the reference function points. Continuing with the above example, the user voice request is "I feel too hot". After the server recognizes the voice request "I feel too hot", it confirms that the user voice request is the first target category. Subsequently, the server will search for relevant reference function points based on the preset knowledge base, and the reference function points obtained are "lower the air conditioning temperature, increase the air conditioning air volume, turn off the seat heating, turn on the seat ventilation, and open the window". Then, based on the determined reference function points, the server queries the current state perception information of the corresponding components of the vehicle, that is, queries the current state of the air conditioner, the current state of the seat, and the current state of the window. The current state perception information obtained is:
[0101] "Front air conditioning air volume": 5,"Whether driver's seat ventilation is supported": "TRUE","Whether passenger seat ventilation is supported": "TRUE","Whether rear seat heating is supported": "TRUE","Query driver's air conditioning temperature": 22.5,"Query passenger seat air conditioning temperature": 22.5,"Query whether driver's seat heating is currently supported, return boolean":"TRUE",Query whether passenger seat heating is currently supported, return boolean":"TRUE","Query maximum air conditioning temperature, return double":"32","Query minimum air conditioning temperature, return double":"18","Query vehicle interior temperature": 22.0,"Query window status":"closed","Query maximum air volume, return int":" 10","Query the minimum air volume and return int":"1","Get the driver's seat heating level":0","Get the driver's seat ventilation level":0","Get the passenger seat heating level":0","Get the passenger seat ventilation level":0","Get the right rear seat heating current gear":0","Get the status of the blowing mode (for example: blowing head, blowing window, blowing foot)":1","Get the status of four air outlets":"on","Get the current gear of the left rear seat heating":0","Get the maximum value of seat heating":"3","Get the minimum value of seat heating":"1","Get the maximum value of seat ventilation":"3","Get the minimum value of seat ventilation":"1","Judge the window addition status (fully open, fully closed, intermediate state)":NaN","Whether the window is currently controllable":"yes".
[0102] Then, after obtaining the current state perception information of the vehicle parts corresponding to the reference function point, the server determines the target operation instruction based on the voice request, the reference function point and the current state perception information. By considering the current state of the vehicle parts, the system can generate more accurate operation instructions to better meet the needs of users.
[0103] In this way, by obtaining current state perception information and conducting comprehensive analysis based on voice requests and reference function points, the server can accurately understand the user's command goals and select appropriate target operation instructions to meet user needs, thereby improving the user experience.
[0104] See also Figure 8 In some embodiments, step 0322 (determining a target operation instruction according to a voice request, a reference function point, and current state perception information) includes:
[0105] 03221: Determine the target function point based on the reference function point and current state perception information;
[0106] 03222: Determine the target operation instructions based on the voice request and target function point.
[0107] In some implementations, the determination module is further configured to determine a target function point based on the reference function point and the current state perception information. The determination module is further configured to determine a target operation instruction based on the voice request and the target function point.
[0108] In some implementations, the processor is further configured to determine a target function point based on the reference function point and the current state perception information, and to determine a target operation instruction based on the voice request and the target function point.
[0109] Specifically, the target function point refers to the specific function that the system determines needs to be executed or adjusted based on user needs and the current status of the vehicle. It is usually related to various equipment and systems of the vehicle, including but not limited to air conditioning, seats, audio, navigation, etc.
[0110] The server will further filter out the function points that can meet the user's needs, namely the target function points, based on the reference function points and the current vehicle perception information. In some embodiments, the server will not only filter out the target function points based on the reference function points and the current vehicle perception information, but also combine the interaction and influence between the reference function points. Continuing with the above example, the user feels too hot, which may be caused by the high air conditioning temperature setting, small air volume, seat heating turned on, or low ventilation level. According to the perception point information, the current air conditioning temperature is 22.5 degrees, the air volume is 5 (the maximum value is 10), the main driver's seat heating level is 0, and the ventilation level is also 0. Considering that the current air volume is not the maximum value, you can try to increase the air volume; at the same time, since both seat ventilation and heating functions are supported, you can consider turning on seat ventilation to improve comfort. The target function point obtained in this way is "increase the air volume of the air conditioner and turn on seat ventilation."
[0111] Next, the server will conduct comprehensive analysis and reasoning based on the user's voice request and the selected target function points to determine the final target operation instruction. Continuing with the above example, based on the user's voice request and the above-obtained target function points, the target operation instruction is determined to be "lower the air conditioning temperature to 18°C and turn on the seat ventilation."
[0112] In this way, by determining the target functional points, the server can accurately focus on the functions that can meet user needs, and perform detailed analysis and reasoning, so as to select appropriate target operation instructions to meet user needs and improve user experience.
[0113] See also Fig. 9 In some implementations, step 03222 (determining a target operation instruction according to a voice request and a target function point) includes:
[0114] 032221: Perform slot recognition according to the voice request and obtain the first entity;
[0115] 032222: According to the first entity and the target function point, execute target function point parameter filling and determine the target operation instruction.
[0116] In some implementations, the acquisition module is further used to perform slot recognition according to the voice request to acquire the first entity. The determination module is further used to perform target function point parameter filling according to the first entity and the target function point to determine the target operation instruction.
[0117] In some implementations, the processor is further configured to perform slot recognition according to the voice request to obtain the first entity, and to perform target function point parameter filling according to the first entity and the target function point to determine the target operation instruction.
[0118] Specifically, slot recognition refers to extracting specific information fragments from the user's input, which are usually called "slots". Slots are usually key information required to complete a task or request, such as time, place, object, etc. Taking the user's voice request "What will be the temperature tomorrow" as an example, the slot information that can be obtained through slot recognition includes ["tomorrow" - date (Date)], that is, the slot information includes slot value and slot type, where "tomorrow" is the slot value and date (Date) is the slot type. Taking the user's voice request "Navigate to address A" as an example, the slot information that can be obtained through slot recognition is ["address A" - place name (Place)], where "address A" is the slot value and place name (Place) is the slot type. Slot recognition can help the server understand the key information in the user's voice request and provide a basis for the subsequent generation of operation instructions.
[0119] The first entity refers to a named entity obtained by performing slot recognition on a voice request, such as the slot information ["tomorrow" - date (Date)] and slot information ["address A" - place name (Place)] mentioned above.
[0120] Parameter filling refers to filling the first entity as a parameter into the corresponding parameter position in the reference function point set. For example, the slot information ["Address A" - Place name (Place)] is filled into the "Destination" parameter position of the "Navigation" function point as a parameter.
[0121] The server will perform slot recognition based on the voice request and extract key information from the instruction. Then, based on the first entity and the target function point, the server will fill the value of the first entity into the corresponding parameter of the target function point to generate a complete target operation instruction.
[0122] In this way, through slot identification and parameter filling, the server can accurately understand the user's command intention and generate target operation instructions that meet the user's expectations, thereby improving the user experience.
[0123] See also Fig.10 In some embodiments, the method further comprises:
[0124] 05: When the instruction category is the first target category, a temporary voice broadcast feedback is generated and sent to the vehicle.
[0125] In some embodiments, the voice interaction device further includes a generation module, which is used to generate temporary voice broadcast feedback when the instruction category is the first target category, and send the temporary voice broadcast feedback to the vehicle.
[0126] In certain embodiments, the processor is further configured to generate temporary voice broadcast feedback when the instruction category is the first target category, and send the temporary voice broadcast feedback to the vehicle.
[0127] Specifically, temporary voice broadcast feedback refers to a voice prompt that the system will broadcast during the voice interaction process. Since large language model inference takes a certain amount of time, in order to alleviate the user's discomfort during the waiting process, the system will broadcast the temporary voice broadcast feedback to inform the user that the instruction is being processed and please wait a moment. In some embodiments, the temporary voice broadcast feedback can be a pre-set voice prompt or a voice prompt generated based on a specific prompt and user voice request. For example, "Processing for you, please wait a moment", or "The air conditioning temperature is being adjusted for you, please wait a moment".
[0128] When the command category is the first target category, the server will generate a temporary voice broadcast feedback based on the command category and the user's voice request. Then, the server will send the generated temporary voice broadcast feedback to the vehicle and play it to the user through the vehicle's audio system. In this way, the user's anxiety while waiting for the command to be processed can be alleviated, improving the user experience.
[0129] In this way, by generating temporary voice broadcast feedback, the server can alleviate the delay in command processing and improve the user experience.
[0130] See also Fig.11 In some implementations, generating a target operation instruction according to the instruction category includes:
[0131] 06: When the instruction category is the second target category, slot recognition is performed on the voice request to determine the second entity;
[0132] 07: API prediction for voice requests;
[0133] 08: According to the second entity and the predicted application program interface, execute application program interface parameter filling and determine the target operation instruction.
[0134] In some embodiments, the determination module is used to perform slot recognition on the voice request and determine the second entity when the instruction category is the second target category, and perform application program interface prediction on the voice request, and perform application program interface parameter filling according to the second entity and the predicted application program interface to determine the target operation instruction.
[0135] In some embodiments, the processor is further configured to, when the instruction category is the second target category, perform slot identification on the voice request to determine the second entity, and perform application program interface prediction on the voice request, and perform application program interface parameter filling according to the second entity and the predicted application program interface to determine the target operation instruction.
[0136] Specifically, when the instruction category is the second target category, first, the server will perform slot recognition on the user's voice request, extract key information in the instruction, such as temperature, air volume, location, etc., and use it as the second entity.
[0137] Next, the server will predict the application program interface that needs to be called based on the user's voice request and the second entity. For example, if the instruction is "adjust the air conditioner temperature to 24 degrees", the server will predict that the application interface "air conditioner control" needs to be called.
[0138] Then, the server will fill the value of the second entity into the corresponding parameter of the application interface according to the second entity and the predicted application interface, thereby generating a complete operation instruction. For example, if the application interface is the application interface "air conditioning control" and the temperature value in the second entity is "24 degrees", the server will fill the temperature value "24 degrees" into the temperature parameter of the application interface "air conditioning control".
[0139] Finally, the server will generate the final target operation instruction based on the filled application interface parameters. For example, if the application interface is air conditioning control, and the temperature value in the second entity is "24 degrees", the final target operation instruction is "raise the air conditioning temperature to 24 degrees".
[0140] In this way, through slot identification, application interface prediction and parameter filling, the server can efficiently execute user instructions and generate target operation instructions that meet user expectations, thereby improving the user experience.
[0141] The present application also provides a computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, the steps of the above-mentioned voice interaction method are implemented.
[0142] It is understood that a computer program includes computer program code. The computer program code may be in source code form, object code form, executable file or some intermediate form. Computer readable storage media may include: any entity or device capable of carrying computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), and software distribution medium.
[0143] In the description of this specification, the descriptions with reference to the terms "specifically", "further", "particularly", "understandably", etc. are intended to mean that the specific features, structures, materials or characteristics described in conjunction with the embodiments or examples are included in at least one embodiment or example of the present application. In this specification, the schematic representations of the above terms are not intended to refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art may combine and combine the different embodiments or examples described in this specification and the features of the different embodiments or examples, unless they are contradictory.
[0144] Any process or method description in a flowchart or otherwise described herein may be understood to represent a module, fragment or portion of code that includes one or more executable requests for implementing specific logical functions or steps of a process, and the scope of the preferred embodiments of the present application includes alternative implementations in which functions may not be performed in the order shown or discussed, including performing functions in a substantially simultaneous manner or in the reverse order depending on the functions involved, which should be understood by technicians in the technical field to which the embodiments of the present application belong.
[0145] Although the embodiments of the present application have been shown and described above, it can be understood that the above embodiments are exemplary and cannot be understood as limitations on the present application. Ordinary technicians in this field can change, modify, replace and modify the above embodiments within the scope of the present application.
Claims
1. A voice interaction method, characterized in that: The method comprises: receiving a voice request; Determining a command category according to the voice request; Determining a target operation instruction according to the voice request and the instruction category; A target operation instruction is sent to the vehicle, so that the vehicle performs the voice interaction according to the target operation instruction.
2. The voice interaction method according to claim 1, characterized in that: The step of determining the instruction category according to the voice request includes: Based on a preset large language model, the instruction category is determined according to the voice request.
3. The voice interaction method according to claim 1, characterized in that: The step of determining the target operation instruction according to the voice request and the instruction category includes: In the case where the instruction category is the first target category, determining a reference function point according to the voice request based on a preset knowledge base; The target operation instruction is determined according to the voice request and the reference function point.
4. The voice interaction method according to claim 3, characterized in that: The determining the target operation instruction according to the voice request and the reference function point includes: According to the reference function point, obtaining current state perception information of the vehicle components corresponding to the reference function point; The target operation instruction is determined according to the voice request, the reference function point and the current state perception information.
5. The voice interaction method according to claim 4, characterized in that: The determining the target operation instruction according to the voice request, the reference function point and the current state perception information includes: Determining a target function point according to the reference function point and the current state perception information; The target operation instruction is determined according to the voice request and the target function point.
6. The voice interaction method according to claim 5, characterized in that: The step of determining the target operation instruction according to the voice request and the target function point includes: Perform slot recognition according to the voice request to obtain a first entity; According to the first entity and the target function point, target function point parameter filling is performed to determine the target operation instruction.
7. The voice interaction method according to claim 1, characterized in that: The method further comprises: When the instruction category is the first target category, a temporary voice broadcast feedback is generated and the temporary voice broadcast feedback is sent to the vehicle.
8. The voice interaction method according to claim 1, characterized in that: Generating a target operation instruction according to the instruction category includes: When the instruction category is the second target category, performing slot recognition on the voice request to determine a second entity; Performing application program interface prediction on the voice request; According to the second entity and the predicted application program interface, application program interface parameter filling is performed to determine the target operation instruction.
9. A server, characterized in that: The server includes a processor and a memory, and a computer program is stored in the memory. When the computer program is executed by the processor, the voice interaction method described in any one of claims 1 to 8 is implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps of the method according to any one of claims 1 to 8 are implemented.
Citation Information
Patent Citations
Vehicle-mounted voice interaction method, system and computer readable memory medium
CN106992009A
Vehicle control method and system and vehicle
CN112435660A
Voice interaction method, device, equipment, medium and product
CN117198289A
Voice interaction method, server and storage medium
CN117524221A
Vehicle control method and device, electronic equipment and storage medium
CN118230727A