Voice interaction method, server and computer readable storage medium

By combining large language model and streaming parallel technology, synchronous processing of voice requests and generating target operation instructions, the problem of on-board voice assistant waiting time is solved and the user experience is improved.

CN120089135APending Publication Date: 2025-06-03GUANGZHOU XIAOPENG MOTORS TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510183216.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-18
Publication Date
2025-06-03

AI Technical Summary

Technical Problem

The existing car voice assistant takes a certain amount of time to obtain relevant knowledge from the external knowledge base, resulting in the user waiting time for too long and the experience is poor.

Method used

Through the combination of large language model and streaming parallel technology, natural language processing of voice requests and target operation instructions are realized, reducing resource and time waiting, and shortening server response time.

Benefits of technology

It greatly shortens the server response time, improves the user experience, and provides a more natural and convenient voice interaction experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120089135A_ABST
    Figure CN120089135A_ABST
Patent Text Reader

Abstract

The invention discloses a voice interaction method, a server and a computer readable storage medium. The method comprises the steps of obtaining a voice request; and determining an instruction category according to the voice request. And under the condition that the instruction category is the first target category, determining a reference function point set according to the voice request. And performing natural language processing on the voice request according to the reference function point set, and determining a candidate operation instruction set. And determining a target operation instruction according to the reference function point set and the candidate operation instruction set. And sending the target operation instruction to the vehicle, so that the vehicle completes voice interaction according to the target operation instruction. Thus, through combination of the large language model and the streaming parallel technology, synchronization of natural language processing of the large language model on the voice request and determination of the target operation instruction is realized, resources and time required for waiting for completion of the whole voice request are reduced, the response time of a server is greatly shortened, and the user experience is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of voice interaction, and particularly relates to a voice interaction method, a server, and a computer-readable storage medium. Background Art

[0002] In the related art, human-machine interaction is carried out with the user through an in-vehicle voice assistant, and the voice request expressing the user's feelings can be inferred according to an external knowledge base to accurately understand the user's intention. However, obtaining relevant knowledge from the external knowledge base takes a certain amount of time, resulting in too long waiting time for the user and poor experience. Summary of the Invention

[0003] The present application provides a voice interaction method, a server, and a computer-readable storage medium.

[0004] An embodiment of the present application provides a voice interaction method, the method including:

[0005] Obtain a voice request;

[0006] Determine an instruction category according to the voice request;

[0007] When the instruction category is a first target category, determine a reference function point set according to the voice request;

[0008] Perform natural language processing on the voice request according to the reference function point set to determine a candidate operation instruction set;

[0009] Determine a target operation instruction according to the reference function point set and the candidate operation instruction set;

[0010] Send the target operation instruction to the vehicle so that the vehicle completes the voice interaction according to the target operation instruction.

[0011] In this way, the server obtains a voice request. Then, the server determines an instruction category according to the voice request. Next, when the instruction category is a first target category, the server determines a reference function point set according to the voice request. Subsequently, the server performs natural language processing on the voice request according to the reference function point set to determine a candidate operation instruction set. At the same time, the server also determines a target operation instruction according to the reference function point set and the candidate operation instruction set. Finally, the server sends the target operation instruction to the vehicle so that the vehicle completes the voice interaction according to the target operation instruction. In this way, by combining the large language model with the streaming parallel technology, the natural language processing of the voice request by the large language model and the determination of the target operation instruction are realized in parallel, reducing the resources and time required to wait for the entire voice request to be completed, thereby greatly shortening the response time of the server and improving the user experience.

[0012] In some embodiments, determining the instruction category according to the voice request includes:

[0013] Based on a preset large language model, determining the instruction category according to the voice request.

[0014] In this way, based on a preset large language model, the server determines the instruction category according to the voice request. Thus, the server can understand the user's voice request and classify it into an appropriate instruction category for subsequent processing and response, so as to be able to intelligently respond to the user's voice instruction and provide a natural and convenient user experience.

[0015] In some embodiments, when the instruction category is the first target category, determining the reference function point set according to the voice request includes:

[0016] Based on a preset database, determining the reference function point set according to the voice request.

[0017] In this way, based on a preset database, the server determines the reference function point set according to the voice request. Thus, based on a preset knowledge base, by analyzing and identifying the voice request, determining the reference function point, and according to the determined reference function point and the voice request, the server can generate a target operation instruction that precisely meets the user's needs, thereby improving the user experience.

[0018] In some embodiments, performing natural language processing on the voice request according to the reference function point set to determine a candidate operation instruction set includes:

[0019] Performing slot recognition according to the voice request to obtain a first entity;

[0020] According to the first entity and the reference function point set, filling parameters for each reference function point in the reference function point set to determine the candidate operation instruction set.

[0021] In this way, the server performs slot recognition according to the voice request to obtain a first entity. Then, the server fills parameters for each reference function point in the reference function point set according to the first entity and the reference function point set to determine the candidate operation instruction set. Thus, through slot recognition and parameter filling, the server can accurately understand the instruction intention of the user and generate a candidate operation instruction set that meets the user's expectations for subsequent determination of the target operation instruction.

[0022] In some embodiments, determining the target operation instruction according to the reference function point set and the candidate operation instruction set includes:

[0023] Obtain the current state perception information of vehicle components corresponding to each reference function point in the reference function point set;

[0024] Determine the target operation instruction according to the reference function point set, the current state perception information, and the candidate operation instruction set.

[0025] In this way, the server obtains the current state perception information of vehicle components corresponding to each reference function point in the reference function point set according to the reference function point set. Then, the server determines the target operation instruction according to the reference function point set, the current state perception information, and the candidate operation instruction set. In this way, by obtaining the current state perception information and comprehensively analyzing it in combination with the candidate operation instruction set and the reference function point set, the server can accurately understand the instruction target of the user, determine the target operation instruction that meets the user's needs and the current state of the vehicle to realize the user's needs, thereby improving the user experience.

[0026] In some embodiments, the determining the target operation instruction according to the reference function point set, the current state perception information, and the candidate operation instruction set includes:

[0027] Determine the target function point according to the reference function point set and the current state perception information;

[0028] Determine the candidate operation instruction corresponding to the target function point in the candidate operation instruction set as the target operation instruction.

[0029] In this way, the server determines the target function point according to the reference function point set and the current state perception information. Then, the server determines the candidate operation instruction corresponding to the target function point in the candidate operation instruction set as the target operation instruction. In this way, by determining the target function point, the server can accurately focus on the function that can meet the user's needs, conduct detailed analysis and reasoning, and thus select the appropriate target operation instruction to realize the user's needs and improve the user experience.

[0030] In some embodiments, the method further includes:

[0031] Generate a temporary voice broadcast feedback and send the temporary voice broadcast feedback to the vehicle.

[0032] In this way, the server generates a temporary voice broadcast feedback and sends the temporary voice broadcast feedback to the vehicle. In this way, by generating the temporary voice broadcast feedback, the server can alleviate the delay in the instruction processing process and improve the user experience.

[0033] In some embodiments, the method further includes:

[0034] When the instruction category is the second target category, perform slot recognition on the voice request to determine a second entity;

[0035] Perform application programming interface prediction on the voice request;

[0036] According to the second entity and the predicted application programming interface, perform application programming interface parameter filling to determine the target operation instruction.

[0037] In this way, when the instruction category is the second target category, the server performs slot recognition on the voice request to determine a second entity. Then, the server performs application programming interface prediction on the voice request. Finally, according to the second entity and the predicted application programming interface, perform application programming interface parameter filling to determine the target operation instruction. In this way, through slot recognition, application interface prediction, and application interface parameter filling, the server can efficiently execute user instructions and determine the target operation instruction that meets the user's expectations, thereby enhancing the user experience.

[0038] An embodiment of the present application provides a server, which includes a processor and a memory. A computer program is stored on the memory. When the computer program is executed by the processor, the above-mentioned voice interaction method is implemented.

[0039] An embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the above-mentioned voice interaction method are implemented.

[0040] Additional aspects and advantages of the embodiments of the present application will be given in part in the following description, will become apparent in part from the following description, or will be understood through the practice of the embodiments of the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] The above and / or additional aspects and advantages of the present application will become apparent and easy to understand from the description of the embodiments in conjunction with the following drawings, where:

[0042] Figure 1 is one of the flow diagrams of the voice interaction method of some embodiments of the present application;

[0043] Figure 2 is the processing flow diagram of the voice request of some embodiments of the present application;

[0044] Figure 3 is the second flow diagram of the voice interaction method of some embodiments of the present application;

[0045] Figure 4 is the processing flow diagram of the traditional large language model;

[0046] Figure 5 It is a schematic diagram of the processing flow of the preset large language model in some embodiments of the present application;

[0047] Figure 6 It is the third schematic diagram of the process of the voice interaction method in some embodiments of the present application;

[0048] Figure 7 It is the fourth schematic diagram of the process of the voice interaction method in some embodiments of the present application;

[0049] Figure 8 It is the fifth schematic diagram of the process of the voice interaction method in some embodiments of the present application;

[0050] Figure 9 It is the sixth schematic diagram of the process of the voice interaction method in some embodiments of the present application;

[0051] Figure 10 It is the seventh schematic diagram of the process of the voice interaction method in some embodiments of the present application;

[0052] Figure 11 It is the eighth schematic diagram of the process of the voice interaction method in some embodiments of the present application. Specific embodiments

[0053] The following details the embodiments of the present application. The examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals denote the same or similar elements or elements with the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the embodiments of the present application and should not be construed as a limitation on the embodiments of the present application.

[0054] In the existing in-vehicle voice assistant technology, it is usually possible to reason about the voice requests of users expressing feelings based on an external knowledge base, so as to accurately understand the user's intention. For example, if the user says "I feel a bit cold", the voice assistant can obtain information about the in-vehicle temperature adjustment from the external knowledge base and infer that the user may need to raise the in-vehicle temperature.

[0055] However, obtaining relevant information from the external knowledge base takes a certain amount of time, which will lead to too long waiting time for users and poor user experience. In addition, during the reasoning process of the in-vehicle voice assistant, a large amount of information needs to be processed, which will result in slow reasoning speed and further extend the user's waiting time. In this way, the user needs to wait for a long time after sending a request to get a response, thus affecting the user experience.

[0056] Based on the above problems, please refer to Figure 1 , the embodiments of the present application provide a voice interaction method, which includes:

[0057] 011: Obtain a voice request;

[0058] 012: Determine the instruction category according to the voice request;

[0059] 013: When the instruction category is the first target category, determine a reference function point set according to the voice request;

[0060] 014: Perform natural language processing on the voice request according to the reference function point set to determine a candidate operation instruction set;

[0061] 015: Determine a target operation instruction according to the reference function point set and the candidate operation instruction set;

[0062] 016: Send the target operation instruction to the vehicle so that the vehicle completes voice interaction according to the target operation instruction.

[0063] The embodiment of the present application also provides a server, including a memory and a processor. The voice interaction method of the embodiment of the present application can be implemented by the server of the embodiment of the present application. Specifically, a computer program is stored in the memory, and the processor is used to obtain a voice request, determine the instruction category according to the voice request, and when the instruction category is the first target category, determine a reference function point set according to the voice request. The processor is also used to perform natural language processing on the voice request according to the reference function point set to determine a candidate operation instruction set, determine a target operation instruction according to the reference function point set and the candidate operation instruction set, and send the target operation instruction to the vehicle so that the vehicle completes voice interaction according to the target operation instruction.

[0064] The embodiment of the present application also provides a voice interaction device. The voice interaction method of the embodiment of the present application can be implemented by the voice interaction device of the embodiment of the present application. Specifically, the voice interaction device includes an obtaining module, a determining module, and a sending module. The obtaining module is used to obtain a voice request. The determining module is used to determine the instruction category according to the voice request. The determining module is also used to, when the instruction category is the first target category, determine a reference function point set according to the voice request. The determining module is also used to perform natural language processing on the voice request according to the reference function point set to determine a candidate operation instruction set. The determining module is also used to determine a target operation instruction according to the reference function point set and the candidate operation instruction set. The sending module is used to send the target operation instruction to the vehicle so that the vehicle completes voice interaction according to the target operation instruction.

[0065] Specifically, in the related art, the processing of perception-based voice requests (i.e., voice requests of the first target instruction category) can be divided into four stages, including determining the instruction category of the voice request, determining a set of candidate function points based on the voice request, determining the target function point based on the set of candidate function points, and performing natural language processing on the voice request based on the target function point. Among them, the time required to determine the instruction category of the voice request may be A1, the time required to determine the set of candidate function points based on the voice request may be A2, the time required to determine the target function point based on the set of candidate function points may be A3, and the time required to perform natural language processing on the voice request based on the target function point may be A4. The time for the server to complete this human-computer interaction is A1 + A2 + A3 + A4.

[0066] The voice interaction method provided by the embodiments of the present application can, after determining the set of candidate function points based on the voice request, while determining the target function point based on the set of candidate function points, simultaneously perform natural language processing on the voice request according to each candidate function point in the set of candidate function points to obtain a set of candidate operation instructions. Then, after determining the target function point, directly determine the target operation instruction from the set of candidate operation instructions according to the target function point. In this way, the time required for the entire interaction process can be reduced. That is, the time for the server to determine the instruction category of the voice request is also A1, the time to determine the set of candidate function points based on the voice request is also A2, the time required to determine the target function point based on the set of candidate function points is also A3, and the time required to determine the target operation instruction based on the target function point and the set of candidate operation instructions may be A5 (A5 < A4). The time for the server to complete this human-computer interaction is A1 + A2 + A3 + A5, and the response speed is faster.

[0067] A voice request refers to a voice request sent by a user to the in-vehicle voice assistant in a voice manner, which can cover various functions, such as control functions, query functions, and setting functions, including environmental perception voice requests. An environmental perception voice request refers to a perceptive description by the user of the in-vehicle environment sensed by the senses (such as vision, hearing, touch, etc.), such as the in-vehicle temperature, in-vehicle humidity, in-vehicle light, seat comfort, and noise level. The in-vehicle temperature refers to the air temperature in the vehicle perceived by the user, for example: cold, hot, warm, etc. The in-vehicle humidity refers to the air humidity in the vehicle perceived by the user, for example: dry, humid, etc. The in-vehicle light refers to the light intensity in the vehicle perceived by the user, for example: bright, dim, etc. The seat comfort refers to the seat comfort perceived by the user, for example: soft, hard, comfortable, etc. The noise level refers to the noise level in the vehicle perceived by the user, for example: quiet, noisy, etc. The current environmental perception information can help the server better understand the user's instruction intention. For example, when the user says "I'm a bit cold", the server can judge that the user needs to turn up the air conditioner temperature based on the in-vehicle temperature perceived by the user.

[0068] The instruction categories include the first target category and the second target category, which are obtained by classifying the intention of the user's voice request. By classifying the instructions, the intention of the user's instructions can be better understood, and more accurate services can be provided. By classifying the voice requests, the server can quickly understand the user's intention and provide an accurate response. Different categories of voice requests require different processing methods. The voice requests of the first target category need to be inferred by combining the external knowledge base and the current status information of the vehicle components, while the voice requests of the second target category can be directly processed by natural language processing.

[0069] The first target category refers to the perception-based instruction category, that is, after the preset large language model recognizes the user's voice request, it is confirmed that the user's voice request is the user's feeling of the current environment, rather than a specific requirement. For example, "I'm a bit cold", "It's a bit too cold", "My back is very stuffy", and "I'm sweating" are all voice requests of the first target category.

[0070] The reference function point set refers to the set of vehicle function points and related operation information obtained from the external knowledge base that can be used to process the current user intention or implement the current user intention. For example, for the voice request "It stinks", the corresponding reference function points are "Open the window, turn on the air purifier, turn on the in-vehicle aromatherapy system, and turn on the external air circulation of the air conditioner". For the voice request "It's too hot", the reference function points are "Lower the air conditioner temperature, increase the air conditioner fan speed, turn off the seat heating, turn on the seat ventilation, and open the window".

[0071] The candidate operation instruction set refers to the set of candidate operation instructions that can initially meet the user's needs, obtained by performing natural language processing on the user's voice request according to each reference function point in the reference function point set. It should be noted that each reference function point in the reference function point set generates a candidate operation instruction.

[0072] The target operation instruction refers to the specific operation instruction generated in the vehicle intelligent system according to the voice request and the instruction category, which can guide the vehicle to perform specific operations or services to meet the user's needs. For example, when the user's perception information is "I feel too hot", the generated target operation instruction is "Increase the air conditioner fan speed". It should be noted that there may be multiple target operation instructions, which are not limited here.

[0073] Please refer to Figure 2 ., The vehicle collects the voice requests and forwards this information to the server. Then, the server determines the category of the instruction according to the received voice request. For example, in an environment where the air conditioner temperature is 22.5 degrees, the air conditioner fan speed is at level 5, and the seat ventilation is turned off, when the user verbally expresses "I feel too hot", the server will analyze this information and determine that the instruction category of the current user's need is the first target category.

[0074] Then, for a voice request belonging to the first target category, the server further analyzes the voice request to determine the vehicle function points related thereto, i.e., the reference function point set. Continuing with the above example, the server combines the voice request "I feel too hot" and determines the reference function point set as "lower the air conditioner temperature, increase the air conditioner air volume, turn off the seat heating, turn on the seat ventilation, and open the window".

[0075] After determining the reference function point set, the server converts the user's voice request into an executable candidate operation instruction set according to the reference function point set. Continuing with the above example, if the determined reference function point set is "lower the air conditioner temperature, increase the air conditioner air volume, turn off the seat heating, turn on the seat ventilation, and open the window", then the determined candidate operation instruction set is "lower the air conditioner temperature to 18°C, increase the air conditioner air volume to gear 7, turn off the seat heating, turn on the seat ventilation, and open the window".

[0076] At the same time, based on the streaming parallel technology, the server determines the target function points according to the reference function set. And according to the target function points and the candidate operation instruction set, the server determines the target operation instruction. Continuing with the above example, the finally determined target operation instruction is "lower the air conditioner temperature to 18°C and turn on the seat ventilation".

[0077] Finally, the server sends the target operation instruction to the vehicle, and the vehicle executes the corresponding operation after receiving the instruction, thus completing the voice interaction.

[0078] In summary, the server obtains the voice request. Then, the server determines the instruction category according to the voice request. Then, when the instruction category is the first target category, the server determines the reference function point set according to the voice request. Subsequently, the server performs natural language processing on the voice request according to the reference function point set to determine the candidate operation instruction set. At the same time, the server also determines the target operation instruction according to the reference function point set and the candidate operation instruction set. Finally, the server sends the target operation instruction to the vehicle so that the vehicle completes the voice interaction according to the target operation instruction. In this way, by combining the large language model and the streaming parallel technology, the natural language processing of the voice request by the large language model and the determination of the target operation instruction are synchronized, reducing the resources and time required to wait for the entire voice request to complete, thus greatly shortening the server response time and improving the user experience.

[0079] Please refer to Figure 3 , in some embodiments, step 012 (determine the instruction category according to the voice request) includes:

[0080] 0121: Based on a preset large language model, determine the instruction category according to the voice request.

[0081] In some embodiments, the determination module is further configured to determine an instruction category based on a preset large language model according to a voice request.

[0082] In some embodiments, the processor is further configured to determine an instruction category based on a preset large language model according to a voice request.

[0083] Specifically, the preset large language model refers to a pre-trained large language model that can be used to process a user's voice request, including identifying the instruction category of the user's voice request, obtaining reference function points, determining target function points, and performing natural language understanding and reasoning on the user's voice request. In some embodiments, the functions implemented by the above preset large language model can be respectively implemented by multiple pre-trained large language models. For example, the functions of identifying the instruction category of the user's voice request and performing natural language understanding and reasoning on the user's voice request can be implemented through a preset model, and the functions of obtaining reference function points and determining target function points can be implemented through a perception model.

[0084] Please refer to Figure 4 , Figure 4 FIG. [FIGURE NUMBER] is a schematic diagram of the processing flow of a traditional large language model for perception-based voice requests. In the related art, the processing method of the large language model for perception-based voice requests is the same as that for non-perception-based voice requests, which may result in the inference ability of the in-vehicle voice assistant often being unable to support providing accurate and applicable feedback, and the user experience is poor. Please refer to Figure 5 , Figure 5 FIG. [FIGURE NUMBER] is a schematic diagram of the processing flow of the preset large language model for perception-based voice requests. The preset large language model provided by the embodiments of the present application can obtain the current state perception information of vehicle components and generate multiple target operation instructions for subsequent steps to determine, enhancing the inference ability for perception-based voice requests and improving the user experience.

[0085] First, use the preset large language model to parse the user's voice request and understand the user's intention. Then, combine the environmental perception information, such as the temperature inside the vehicle and the weather outside the vehicle, to determine the category of the instruction. In some embodiments, information such as the user's voice intonation and historical preferences will also be combined to jointly determine the instruction category of the user's voice request. For example, if the user says "I feel a bit cold", the system will determine the instruction category as "First Target Category". If the user says "Turn on the air conditioner", the system will determine the instruction category as "Second Target Category".

[0086] It should be noted that, in some embodiments, the server will also combine the environmental information of the user's current location obtained by various sensors on the vehicle to determine the instruction category of the user's voice request.

[0087] In this way, the server can understand the user's voice request, classify it into an appropriate instruction category for subsequent processing and response, so as to intelligently respond to the user's voice instruction and provide a natural and convenient user experience.

[0088] Please refer to Figure 6 , in some embodiments, step 013 (when the instruction category is the first target category, determining a reference function point set according to the voice request) includes:

[0089] 0131: Based on a preset database, determine a reference function point set according to the voice request.

[0090] In some embodiments, the determination module is further configured to determine a reference function point set based on a preset database according to the voice request.

[0091] In some embodiments, the processor is further configured to determine a reference function point set based on a preset database according to the voice request.

[0092] Specifically, the preset knowledge base refers to a database including a large amount of structured information, which can support the decision-making and operation of the system, that is, it provides the information basis required for processing user requirements and determining reference function points. By analyzing the user perception information and the current environmental perception information of the user's environment, the server can determine appropriate reference function points from the preset knowledge base for subsequent processes. In some embodiments, the data in the preset knowledge base may include:

[0093] Auditory perception category: [Voice request: It's too noisy; Reference function points: Close the window, Turn down the media volume, Turn down the navigation volume, Turn down the voice announcement volume], [Voice request: The sound is too loud; Reference function points: Turn down the media volume, Turn down the navigation volume, Turn down the voice announcement volume, Close the window], [Voice request: I can't hear the navigation clearly; Reference function points: Turn up the navigation volume, Turn down the media volume, Turn down the voice announcement volume, Raise and close the window], [Voice request: I want it louder; Reference function points: Turn up the navigation volume, Turn up the media volume, Turn up the voice announcement volume], [Voice request: The music is too noisy; Reference function points: Turn down the music volume, Pause the music playback, Switch the music type].

[0094] Temperature Sensation Category: [Voice Request: So cold; Reference Function Points: Close the windows, turn up the air conditioner temperature, turn up the air conditioner fan speed, turn on the seat heater, turn up the seat heater], [Voice Request: Too cold; Reference Function Points: Close the windows, turn up the air conditioner temperature, turn up the air conditioner fan speed, turn on the seat heater, turn up the seat heater], [Voice Request: So hot; Reference Function Points: Turn down the air conditioner temperature, turn up the air conditioner fan speed, turn off the seat heater, turn on the seat ventilation, open the windows], [Voice Request: My butt is so hot; Reference Function Points: Turn down the seat heater, turn on the seat ventilation, turn up the seat ventilation], [Voice Request: My butt is so cold; Reference Function Points: Turn on the seat ventilation, turn down the seat ventilation, turn up the seat heater].

[0095] Olfactory Sensation Category: [Voice Request: So stinky; Reference Function Points: Open the windows, turn on the air purifier, turn on the in-vehicle aromatherapy system, turn on the air conditioner external circulation], [Voice Request: The air is not fresh; Reference Function Points: Open the windows, turn on the air purifier, turn on the in-vehicle aromatherapy system, turn on the air conditioner external circulation], [Voice Request: Want some fragrance; Reference Function Points: Turn on the in-vehicle aromatherapy system, turn on the air purifier, open the windows, turn on the air conditioner external circulation], [Voice Request: Don't need to remove the odor anymore; Reference Function Points: Turn on the air purifier, close the windows, turn off the air conditioner external circulation, turn off the in-vehicle aromatherapy system].

[0096] Visual Sensation Category: [Voice Request: The road is too dark; Reference Function Points: Turn on the headlights, adjust the headlight brightness, turn on the fog lights or auxiliary lights], [Voice Request: Too bright; Reference Function Points: Open the sun visor, turn off the reading light, turn off the ceiling light, adjust the dashboard brightness to the lowest, turn off the headlights], [Voice Request: Can't see clearly behind; Reference Function Points: Use the reverse image system, turn on the rear windshield defogger, turn on the rear windshield defrost, turn on the headlights], [Voice Request: Can't see clearly due to the rearview mirror reflection; Reference Function Points: Turn on the air conditioner defogger, turn on the rearview mirror heating (if available), adjust the air conditioner air direction to the rearview mirror, turn on the vehicle external circulation].

[0097] Liking and Disliking Sensation Category: [Voice Request: Don't like these songs very much; Reference Function Points: Change the music playlist, switch the music style, enable the random play mode, ask the user about their music preferences].

[0098] Comfort Sensation Category: [Voice Request: Want to take a break; Reference Function Points: Adjust the seat angle to a comfortable reclining position, turn off the interior lights, close the windows, adjust the air conditioner temperature to a suitable temperature for rest, adjust the air conditioner fan speed to medium], [Voice Request: I'm so tired; Reference Function Points: Adjust the seat angle to a comfortable reclining position, turn off the interior lights, close the windows, adjust the air conditioner temperature], [Voice Request: It's too crowded; Reference Function Points: Move the seat backward, recline the backrest], [Voice Request: The steering wheel is too far away; Reference Function Points: Move the steering wheel backward, move the seat forward, raise the seat].

[0099] Please refer to again Figure 2 Based on a preset knowledge base, the server determines a set of reference function points according to the voice request. For example, continuing with the above example, the user's voice request is "I feel too hot". After the server recognizes the voice request "I feel too hot", it confirms that the user's voice request belongs to the first target category. Subsequently, the server will search for a relevant set of reference function points based on the preset knowledge base, and the obtained set of reference function points is "lower the air conditioner temperature, increase the air volume of the air conditioner, turn off the seat heating, turn on the seat ventilation, and open the window".

[0100] In this way, based on the preset knowledge base, by analyzing and recognizing the voice request, determining the reference function points, and according to the determined reference function points and the voice request, the server can generate target operation instructions that precisely meet the user's needs, thereby improving the user experience.

[0101] Please refer to Figure 7 In some embodiments, step 014 (performing natural language processing on the voice request according to the set of reference function points to determine a set of candidate operation instructions) includes:

[0102] 0141: Performing slot recognition according to the voice request to obtain a first entity;

[0103] 0142: Filling in parameters for each reference function point in the set of reference function points according to the first entity and the set of reference function points to determine a set of candidate operation instructions.

[0104] In some embodiments, the determination module is further configured to perform slot recognition according to the voice request to obtain a first entity. And filling in parameters for each reference function point in the set of reference function points according to the first entity and the set of reference function points to determine a set of candidate operation instructions.

[0105] In some embodiments, the processor is further configured to perform slot recognition according to the voice request to obtain a first entity. And filling in parameters for each reference function point in the set of reference function points according to the first entity and the set of reference function points to determine a set of candidate operation instructions.

[0106] Specifically, slot recognition refers to extracting specific information fragments from the user's input, which are usually called "slots". Slots are usually the key information necessary to complete a certain task or request, such as time, location, object, etc. Taking the user's voice request "What's the temperature tomorrow?" as an example, the slot information obtained through slot recognition includes ["tomorrow" - Date], that is, the slot information includes the slot value and the slot type, where "tomorrow" is the slot value and Date is the slot type. Taking the user's voice request "Navigate to Address A" as an example, the slot information obtained through slot recognition is ["Address A" - Place], where "Address A" is the slot value and Place is the slot type. Slot recognition can help the server understand the key information in the user's voice request and provide a basis for generating subsequent operation instructions.

[0107] The first entity refers to the named entity obtained by performing slot recognition on the voice request. Such as the above slot information ["tomorrow" - Date] and slot information ["Address A" - Place], etc.

[0108] Parameter filling refers to using the first entity as a parameter and filling it into the corresponding parameter position in the reference function point set. For example, filling the slot information ["Address A" - Place] as a parameter into the "destination" parameter position of the "navigation" function point.

[0109] The server will perform slot recognition based on the voice request and extract the key information in the instruction, that is, the first entity. Then, the server will fill the value of the first entity into the corresponding parameters of the reference function point according to the first entity and each reference function point in the reference function point set, so as to generate a complete target operation instruction. Continuing with the above example, according to the reference function point set "Lower the air conditioner temperature, Increase the air conditioner wind volume, Turn off the seat heating, Turn on the seat ventilation, and Open the window", the candidate operation instruction set is determined as "Lower the air conditioner temperature to 18°C, Increase the air conditioner wind volume to gear 7, Turn off the seat heating, Turn on the seat ventilation, and Open the window".

[0110] In this way, through slot recognition and parameter filling, the server can accurately understand the instruction intention of the user and generate a candidate operation instruction set that meets the user's expectations for subsequent determination of the target operation instruction.

[0111] Please refer to Figure 8 , in some embodiments, step 015 (Determine the target operation instruction according to the reference function point set and the candidate operation instruction set) includes:

[0112] 0151: Obtain the current state perception information of vehicle components corresponding to each reference function point in the reference function point set according to the reference function point set;

[0113] 0152: Determine the target operation instruction according to the reference function point set, the current state perception information, and the candidate operation instruction set.

[0114] In some embodiments, the voice interaction device further includes an acquisition module. The acquisition module is further configured to obtain the current state perception information of vehicle components corresponding to each reference function point in the reference function point set according to the reference function point set, and determine the target operation instruction according to the reference function point set, the current state perception information, and the candidate operation instruction set.

[0115] In some embodiments, the processor is further configured to obtain the current state perception information of vehicle components corresponding to each reference function point in the reference function point set according to the reference function point set, and determine the target operation instruction according to the reference function point set, the current state perception information, and the candidate operation instruction set.

[0116] Specifically, the current state perception information of vehicle components refers to the specific working state or performance data of each vehicle component or system at the current moment collected by the sensors and monitoring systems built in the vehicle. For example, the air conditioning system state: including the current temperature setting, wind speed, internal and external circulation modes, etc., or the seat state: including the seat position, temperature, massage function state, etc.

[0117] The server obtains the current state perception information of vehicle components corresponding to each reference function point in the reference function point set according to the reference function point set. Continuing with the above example, the user voice request is "I feel too hot". After the server recognizes the voice request "I feel too hot", it confirms that the user voice request belongs to the first target category. Subsequently, the server will search for relevant reference function points based on the preset knowledge base, and the obtained reference function point set is "lower the air conditioning temperature, increase the air conditioning air volume, turn off the seat heating, turn on the seat ventilation, open the window". Then, the server queries the current state perception information of the corresponding vehicle components according to the determined reference function points, that is, queries the current state of the air conditioning, the current state of the seat, and the current state of the window. The obtained current state perception information is:

[0118] "Front row air volume": 5, "Whether the driver's seat ventilation is supported": "TRUE", "Whether the passenger seat ventilation is supported": "TRUE", "Whether the rear seat heating is supported": "TRUE", "Query the driver's seat air temperature": 22.5, "Query the passenger seat air temperature": 22.5, "Query whether the driver's seat heating is currently supported, return boolean": "TRUE", "Query whether the passenger seat heating is currently supported, return boolean": "TRUE", "Query the maximum air temperature value, return double": "32", "Query the minimum air temperature value, return double": "18", "Query the vehicle interior temperature": 22.0, "Query the window status": "closed", "Query the maximum air volume value, return int": "10", "Query the minimum air volume value, return int": "1", "Get the driver's seat heating level": 0, "Get the driver's seat ventilation level": 0, "Get the passenger seat heating level": 0, "Get the passenger seat ventilation level": 0, "Get the current gear of the right rear seat heating": 0, "Get the status of the blowing mode (e.g., head blowing, window blowing, foot blowing)": 1, "Get the status of the four air outlets": "on", "Get the current gear of the left rear seat heating": 0, "Get the maximum value of the seat heating": "3", "Get the minimum value of the seat heating": "1", "Get the maximum value of the seat ventilation": "3", "Get the minimum value of the seat ventilation": "1", "Window increase status judgment (fully open, fully closed, intermediate state)": NaN, "Whether the window is currently controllable": "yes".

[0119] Next, the server determines the target operation instruction according to the reference function point set, the current state perception information, and the candidate operation instruction set. That is, the server determines the target function point according to the reference function set. And according to the target function point and the candidate operation instruction set, the server determines the target operation instruction.

[0120] In this way, by obtaining the current state perception information and comprehensively analyzing it in combination with the candidate operation instruction set and the reference function point set, the server can accurately understand the user's instruction target, determine the target operation instruction that meets the user's needs and the current state of the vehicle to achieve the user's needs, thereby improving the user experience.

[0121] Please refer to Figure 9 , in some embodiments, step 0152 (determine the target operation instruction according to the reference function point set, the current state perception information, and the candidate operation instruction set) includes:

[0122] 01521: Determine the target function point according to the reference function point set and the current state perception information;

[0123] 01522: Determine the candidate operation instruction corresponding to the target function point from the candidate operation instruction set as the target operation instruction.

[0124] In some embodiments, the determination module is further configured to determine the target function point according to the reference function point set and the current state perception information. The determination module is further configured to determine the candidate operation instruction corresponding to the target function point from the candidate operation instruction set as the target operation instruction.

[0125] In some embodiments, the processor is further configured to determine the target function point according to the reference function point set and the current state perception information. And determine the candidate operation instruction corresponding to the target function point from the candidate operation instruction set as the target operation instruction.

[0126] Specifically, the target function point refers to the specific function determined by the system that needs to be executed or adjusted according to the user's needs and the current state of the vehicle, usually related to various devices and systems of the vehicle, including but not limited to air conditioners, seats, stereos, navigation, etc.

[0127] The server will further screen out the function points that can meet the user's needs, that is, the target function points, according to the reference function points and the current vehicle perception information. In some embodiments, the server will not only screen according to the reference function points and the current vehicle perception information, but also combine the interactions and influences between the reference function points to screen out the target function points. Continuing the above example, the user feels too hot, which may be caused by reasons such as a high air conditioner temperature setting, a small air volume, the seat heating being turned on, or a low ventilation level. According to the perception point information, the current air conditioner temperature is 22.5 degrees, the air volume is 5 (the maximum is 10), the main driver's seat heating level is 0, and the ventilation level is also 0. Considering that the current air volume is not the maximum, the air volume can be tried to be increased; at the same time, since both the seat ventilation and heating functions are supported, the seat ventilation can be considered to be turned on to improve comfort. The target function points obtained in this way are "increase the air conditioner air volume and turn on the seat ventilation".

[0128] Next, the server determines the candidate operation instruction corresponding to the target function point from the candidate operation instruction set as the target operation instruction. Continuing the above example, according to the determined target function points "increase the air conditioner air volume and turn on the seat ventilation", from the candidate operation instruction set "lower the air conditioner temperature to 18°C, increase the air conditioner air volume to 7 gears, turn off the seat heating, turn on the seat ventilation and open the window", "increase the air conditioner air volume to 7 gears and turn on the seat ventilation" is determined as the target operation instruction.

[0129] In this way, by determining the target function point, the server can accurately focus on the functions that can meet the user's needs, and conduct detailed analysis and reasoning, so as to select the appropriate target operation instruction to meet the user's needs and improve the user experience.

[0130] Please refer to Figure 10 , in some embodiments, the method further includes:

[0131] 017: Generate a temporary voice broadcast feedback and send the temporary voice broadcast feedback to the vehicle.

[0132] In some embodiments, the voice interaction device further includes a generation module, and the generation module is used to generate a temporary voice broadcast feedback and send the temporary voice broadcast feedback to the vehicle.

[0133] In some embodiments, the processor is further used to generate a temporary voice broadcast feedback and send the temporary voice broadcast feedback to the vehicle.

[0134] Specifically, the temporary voice broadcast feedback refers to a voice prompt that the system will broadcast during the voice interaction process. Since the inference of the large language model takes a certain amount of time, in order to relieve the discomfort of the user during the waiting process, the system will broadcast this temporary voice broadcast feedback to inform the user that the instruction is being processed and please wait a moment. In some embodiments, the temporary voice broadcast feedback can be a pre-set voice prompt or a voice prompt generated according to a specific Prompt and the user's voice request. For example, "Processing for you, please wait a moment", or "Adjusting the air conditioner temperature for you, please wait a moment".

[0135] The server will generate a temporary voice broadcast feedback according to the instruction category and the user's voice request. Then, the server will send the generated temporary voice broadcast feedback to the vehicle and play it to the user through the vehicle's audio system. In this way, the anxiety of the user during the waiting for the instruction processing can be relieved, and the user experience can be improved.

[0136] In this way, by generating the temporary voice broadcast feedback, the server can relieve the delay in the instruction processing process and improve the user experience.

[0137] Please refer to Figure 11 , in some embodiments, the method further includes:

[0138] 018: In the case where the instruction category is the second target category, perform slot recognition on the voice request to determine the second entity;

[0139] 019: Perform application programming interface prediction on the voice request;

[0140] 020: According to the second entity and the predicted application programming interface, perform application programming interface parameter filling to determine the target operation instruction.

[0141] In some embodiments, the determination module is configured to perform slot recognition on a voice request to determine a second entity when the instruction category is the second target category, perform application programming interface (API) prediction on the voice request, and perform API parameter filling based on the second entity and the predicted API to determine a target operation instruction.

[0142] In some embodiments, the processor is further configured to perform slot recognition on a voice request to determine a second entity when the instruction category is the second target category, perform API prediction on the voice request, and perform API parameter filling based on the second entity and the predicted API to determine a target operation instruction.

[0143] Specifically, when the instruction category is the second target category, first, the server performs slot recognition on the user's voice request, extracts key information in the instruction, such as temperature, air volume, location, etc., and uses it as the second entity.

[0144] Next, the server predicts the API to be called based on the user's voice request and the second entity. For example, if the instruction is "set the air conditioner temperature to 24 degrees", the server will predict that the API "air conditioner control" needs to be called.

[0145] Then, the server fills the values of the second entity into the corresponding parameters of the API based on the second entity and the predicted API, thereby generating a complete operation instruction. For example, if the API is the "air conditioner control" API and the temperature value in the second entity is "24 degrees", the server will fill the temperature value "24 degrees" into the temperature parameter of the "air conditioner control" API.

[0146] Finally, the server generates a final target operation instruction based on the filled API parameters. For example, if the API is air conditioner control and the temperature value in the second entity is "24 degrees", the final target operation instruction is "set the air conditioner temperature to 24 degrees".

[0147] In this way, through slot recognition, API prediction, and parameter filling, the server can efficiently execute user instructions and generate target operation instructions that meet user expectations, thereby improving the user experience.

[0148] The present application also provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the steps of the voice interaction method as described above are implemented.

[0149] It can be understood that a computer program includes computer program code. The computer program code can be in the form of source code, object code, executable file or some intermediate form, etc. The computer-readable storage medium can include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), and software distribution medium, etc.

[0150] In the description of this specification, the descriptions referring to terms such as "specifically", "further", "specially", "understandably", etc. mean that the specific features, structures, materials or characteristics described in connection with the embodiments or examples are included in at least one embodiment or example of the present application. In this specification, the schematic expressions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in a suitable manner in any one or more embodiments or examples. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.

[0151] Any process or method description shown in the flowchart or described in other ways herein can be understood as representing a module, segment or part of the code of an executable request including one or more steps for implementing a specific logical function or process, and the scope of the preferred embodiments of the present application includes additional implementations, where the functions can be executed in a substantially simultaneous manner or in the reverse order according to the involved functions, rather than in the order shown or discussed, which should be understood by those skilled in the technical field to which the embodiments of the present application belong.

[0152] Although the embodiments of the present application have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present application. Those of ordinary skill in the art can make changes, modifications, substitutions and variations to the above embodiments within the scope of the present application.

Claims

1. A voice interaction method, characterized in that: The method comprises: Get voice request; Determining a command category according to the voice request; In a case where the instruction category is a first target category, determining a reference function point set according to the voice request; Performing natural language processing on the voice request according to the reference function point set to determine a candidate operation instruction set; Determining a target operation instruction according to the reference function point set and the candidate operation instruction set; A target operation instruction is sent to the vehicle so that the vehicle completes the voice interaction according to the target operation instruction.

2. The voice interaction method according to claim 1, characterized in that: The step of determining the instruction category according to the voice request includes: Based on a preset large language model, the instruction category is determined according to the voice request.

3. The voice interaction method according to claim 1, characterized in that: When the instruction category is the first target category, determining the reference function point set according to the voice request includes: Based on a preset database and according to the voice request, the reference function point set is determined.

4. The voice interaction method according to claim 1, characterized in that: The performing natural language processing on the voice request according to the reference function point set to determine a candidate operation instruction set includes: Perform slot recognition according to the voice request to obtain a first entity; According to the first entity and the reference function point set, parameters are filled for each reference function point in the reference function point set to determine the candidate operation instruction set.

5. The voice interaction method according to claim 1, characterized in that: The step of determining a target operation instruction according to the reference function point set and the candidate operation instruction set includes: According to the reference function point set, obtaining current state perception information of vehicle components corresponding to each reference function point of the reference function point set; The target operation instruction is determined according to the reference function point set, the current state perception information and the candidate operation instruction set.

6. The voice interaction method according to claim 5, characterized in that: The determining the target operation instruction according to the reference function point set, the current state perception information and the candidate operation instruction set includes: Determining a target function point according to the reference function point set and the current state perception information; A candidate operation instruction corresponding to the target function point is determined from the candidate operation instruction set as the target operation instruction.

7. The voice interaction method according to claim 1, characterized in that: The method further comprises: A temporary voice broadcast feedback is generated, and the temporary voice broadcast feedback is sent to the vehicle.

8. The voice interaction method according to claim 1, characterized in that: The method further comprises: When the instruction category is the second target category, performing slot recognition on the voice request to confirm the second entity; Performing application program interface prediction on the voice request; According to the second entity and the predicted application program interface, application program interface parameter filling is performed to determine the natural language understanding result.

9. A server, characterized in that: The server includes a processor and a memory, and a computer program is stored in the memory. When the computer program is executed by the processor, the voice interaction method described in any one of claims 1 to 8 is implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps of the method according to any one of claims 1 to 8 are implemented.