Voice interaction method, server and computer readable storage medium
Patent Information
- Application Number
- CN202410790345.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-06-18
- Publication Date
- 2026-08-28
- Estimated Expiration
- 2044-06-18
AI Technical Summary
然而,由于这种简短的语音指令可指向多种车载功能,导致用户意图的识别难度较高,进而难以被准确地理解和执行
Smart Images

Figure CN119170012B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of voice interaction technology, and in particular to a voice interaction method, a server, and a computer-readable storage medium. Background Technology
[0002] As users become more accustomed to in-vehicle voice interaction functions, they may use short voice commands to control the vehicle, such as saying "adjust to 28 degrees" instead of "set the air conditioning temperature to 28 degrees Celsius." However, because such short voice commands can refer to multiple in-vehicle functions, it is difficult to recognize the user's intent, making it difficult to be accurately understood and executed. Summary of the Invention
[0003] This application provides a voice interaction method, a server, and a computer-readable storage medium.
[0004] This application provides a voice interaction method, including:
[0005] Receive the first voice request of the current dialogue round forwarded by the vehicle;
[0006] Based on the first voice request, the second voice request, and the natural language understanding results of the second voice request, the first voice request is subjected to application interface prediction and interface parameter filling, wherein the second voice request is the voice request of the previous dialogue round of the current dialogue round;
[0007] The execution result of the interface parameter filling is sent to the vehicle to complete the voice interaction.
[0008] In the voice interaction method provided in this application, the server can receive the first voice request of the current dialogue round forwarded by the vehicle, and perform application interface prediction and interface parameter filling for the first voice request based on the first voice request, the second voice request and the natural language understanding result of the second voice request, and send the execution result of the interface parameter filling to the vehicle so that the vehicle can perform corresponding operations according to the execution result, thereby completing the voice interaction with the user.
[0009] Thus, in this embodiment, the processing of the first voice request in the current dialogue round can be based on the second voice request from the previous dialogue round and the natural language understanding result of the second voice request. This achieves context-based voice interaction, thereby ensuring the reliability of the application prediction results and parameter filling execution results of the first voice request to a certain extent. This guarantees the reliable execution of voice interaction and ensures the user's experience with the voice interaction function. Furthermore, when the voice request in the current dialogue round is relatively short, the voice request and natural language understanding result from the previous dialogue round can be tracked to determine the user intent and perform natural language understanding corresponding to the voice request in the current dialogue round. This also ensures the reliable processing of the voice request in the current dialogue round to a certain extent and improves the applicability of the voice interaction function in complex multi-turn dialogue scenarios.
[0010] In some embodiments of this application, the step of performing application programming interface (API) prediction and API parameter filling on the first voice request based on the first voice request, the second voice request, and the natural language understanding results of the second voice request includes:
[0011] The process involves concatenating the first voice request, the second voice request, and the natural language understanding result.
[0012] Based on the result of the splicing process, the first voice request is subjected to application interface prediction and interface parameter filling.
[0013] Thus, in this embodiment of the application, the first voice request, the second voice request, and the natural language understanding result of the second voice request can be concatenated, and the application programming interface (API) prediction and interface parameter filling of the first voice request can be performed based on the concatenation result.
[0014] In some embodiments of this application, the step of performing application programming interface (API) prediction and API parameter filling on the first voice request based on the first voice request, the second voice request, and the natural language understanding results of the second voice request includes:
[0015] Configure the dialogue turn label for each parameter in the result of the splicing process.
[0016] Thus, in this embodiment of the application, when the natural language understanding results of the first voice request, the second voice request, and the second voice request are concatenated to obtain the concatenation result, the dialogue turn labels of each parameter in the concatenation result can be configured so that the parameters belonging to the current dialogue turn and the parameters belonging to the previous dialogue turn in the concatenation result can be distinguished based on the dialogue turn labels.
[0017] In some embodiments of this application, the method further includes:
[0018] Slot identification is performed on the first voice request;
[0019] The concatenation process based on the first voice request, the second voice request, and the natural language understanding result includes:
[0020] The first voice request, the slot recognition result, the second voice request, and the natural language understanding result are concatenated to obtain the concatenation result.
[0021] Thus, in this embodiment of the application, slot identification can be performed on the first voice request, and the first voice request, the result of slot identification of the first voice request, and the natural language understanding result of the second voice request can be spliced together to obtain the splicing result.
[0022] In some embodiments of this application, the step of performing application programming interface (API) prediction and API parameter filling on the first voice request based on the first voice request, the second voice request, and the natural language understanding results of the second voice request includes:
[0023] Configure the position label for each parameter in the result of the splicing process.
[0024] Thus, in this embodiment of the application, the prediction result of the application interface of the first voice request can be parameter-filled based on the slot identification result of the first voice request and the interface filling parameters of the application interface of the second voice request.
[0025] In some embodiments of this application, the natural language understanding result of the second voice request includes the application prediction corresponding to the second voice request and the target parameters for filling the application prediction.
[0026] Thus, in this embodiment of the application, when the natural language understanding results of the first voice request, the second voice request, and the second voice request are spliced together to obtain the splicing result, the position labels of each parameter in the splicing result can be configured so that the parameters in the splicing result can be distinguished and associated based on the position labels, thereby improving the accuracy of application interface prediction and parameter filling to a certain extent.
[0027] In some embodiments of this application, the method further includes:
[0028] Based on the first voice request, the second voice request, and the natural language understanding result, the dialogue association attribute information of the first voice request relative to the second voice request is determined, wherein the dialogue association attribute information is used to indicate whether the first voice request and the second voice request are associated.
[0029] Thus, in this embodiment of the application, the dialogue association attribute information of the first voice request relative to the second voice request can be determined based on the first voice request, the second voice request, and the natural language understanding result, thereby determining the relationship between the current round of dialogue and the previous round of dialogue.
[0030] In some embodiments of this application, sending the execution result of the interface parameter filling to the vehicle to complete the voice interaction includes:
[0031] The execution result and the dialogue association attribute information are sent to the vehicle to complete the voice interaction.
[0032] Thus, in this embodiment of the application, the execution result of the filling process of the application interface parameters and the dialogue association attribute information of the first voice request relative to the second voice request can be sent to the vehicle, so that the vehicle can perform corresponding operations based on the execution result and the dialogue association attribute information.
[0033] This application provides a server including a memory and a processor. The memory stores a computer program, which, when executed by the processor, implements the above-described voice interaction method.
[0034] This application provides a computer-readable storage medium storing a computer program that, when executed by one or more processors, implements the above-described voice interaction method.
[0035] The server and computer-readable storage medium provided in this application enable the processing of the first voice request in the current dialogue round to be executed based on the second voice request in the previous dialogue round and the natural language understanding result of the second voice request. This achieves context-based voice interaction, thereby ensuring the reliability of the application prediction results and parameter filling execution results of the first voice request to a certain extent. This ensures the reliable execution of voice interaction and guarantees the user's experience with the voice interaction function. Furthermore, when the voice request in the current dialogue round is relatively short, the voice request and natural language understanding result of the previous dialogue round can be tracked to determine the user intent and perform natural language understanding corresponding to the voice request in the current dialogue round. This further ensures the reliable processing of the voice request in the current dialogue round and improves the applicability of the voice interaction function in complex multi-turn dialogue scenarios.
[0036] Additional aspects and advantages of embodiments of this application will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of embodiments of this application. Attached Figure Description
[0037] The above and / or additional aspects and advantages of this application will become apparent and readily understood from the description of the embodiments taken in conjunction with the following drawings, wherein:
[0038] Figure 1 This is a flowchart illustrating the voice interaction method in some embodiments of this application;
[0039] Figure 2 This is a schematic diagram illustrating application scenarios in some embodiments of this application;
[0040] Figure 3 This is a flowchart illustrating the voice interaction method in some embodiments of this application;
[0041] Figure 4 This is a schematic diagram illustrating application scenarios in some embodiments of this application;
[0042] Figure 5 This is a schematic diagram illustrating the splicing process results in certain embodiments of this application;
[0043] Figure 6 This is a schematic diagram illustrating the splicing process results in certain embodiments of this application;
[0044] Figure 7 This is a schematic diagram illustrating application scenarios in some embodiments of this application;
[0045] Figure 8 This is a schematic diagram illustrating the splicing process results in certain embodiments of this application;
[0046] Figure 9 This is a schematic diagram illustrating the splicing process results in certain embodiments of this application;
[0047] Figure 10 This is a schematic diagram illustrating application scenarios in some embodiments of this application. Detailed Implementation
[0048] The embodiments of this application are described in detail below. Examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the embodiments of this application, and should not be construed as limiting the embodiments of this application.
[0049] As users become more familiar with in-vehicle voice interaction functions, they expect to engage in long and coherent multi-turn dialogues with their vehicles to execute complex commands or answer questions. However, during these multi-turn dialogues, the voice commands uttered by the user may become more concise as the number of dialogue rounds increases.
[0050] For example, in a multi-turn dialogue concerning a vehicle's air conditioning, the user might initially express their intention with a relatively complete voice command such as "Turn on the air conditioning" or "Turn on the air conditioning." However, as the dialogue progresses, the user might gradually omit components from the voice command. For instance, the user might express a voice command like "Set it to 28 degrees," which lacks both subject and object, instead of a more complete voice command like "Set the air conditioning temperature to 18 degrees Celsius."
[0051] It is understandable that when a user utters a short voice command with omitted sentence components, it is usually because these omitted sentence components have been explicitly or implicitly indicated in previous rounds of dialogue. For example, in the above example, the user has clearly defined the subject (or object) as "air conditioner" at the beginning of the dialogue. Therefore, in subsequent dialogues, the user defaults to "air conditioner" as the subject (or object), thus omitting the subject "air conditioner" and "air conditioner temperature" from "set the air conditioner to 28 degrees" or "set the air conditioner temperature to 28 degrees Celsius," thereby uttering a short voice command like "set to 28 degrees."
[0052] It is also understandable that the aforementioned short voice commands can refer to multiple in-vehicle functions. For example, "set to 28 degrees" can refer to "set the seat heating temperature to 28 degrees" and "set the air conditioning temperature to 28 degrees". Therefore, there are situations where the direction is unclear and the user's intention is ambiguous. Thus, when directly performing natural language understanding on the aforementioned short voice commands, the dialogue system (or natural language processing system) may not be able to accurately understand such short voice commands, causing the vehicle to have difficulty responding to the user's voice commands correctly, or to understand the wrong result and cause the vehicle to perform the wrong operation.
[0053] Therefore, the dialogue system can perform natural language understanding and processing for each round of dialogue based on the contextual information of each round. It's understandable that for in-vehicle voice interaction functions, or more specifically, for the dialogue system, as the number of dialogue rounds between the user and the vehicle increases and the voice interaction deepens, the contextual information of each round of dialogue also increases. Consequently, to ensure consistency of previous dialogue scenarios and smooth transitions between rounds of dialogue, the dialogue system needs to accurately understand and retain the information from previous conversations.
[0054] Traditional solutions typically maintain a context state to track the progress of the dialogue, but this is prone to failure in complex dialogue scenarios and requires the dialogue system to parse and correlate various information points in multiple rounds of dialogue, placing high demands on the system. Furthermore, as the number of dialogue rounds increases, the difficulty and complexity of processing the context information for each round also increase, leading to higher requirements for hardware computing power and storage.
[0055] Furthermore, this traditional context-state-based approach struggles to distinguish contextual information across different rounds, causing dialogue systems to easily forget previous interactions or fail to interpret new user input within the correct context.
[0056] Based on the issues mentioned above, please refer to Figure 1 This application provides a voice interaction method, including:
[0057] 01: Receive the first voice request of the current dialogue round forwarded by the vehicle;
[0058] 02: Based on the first voice request, the second voice request, and the natural language understanding result of the second voice request, perform application interface prediction and interface parameter filling for the first voice request, wherein the second voice request is the voice request of the previous dialogue round of the current dialogue round;
[0059] 03: Send the execution result of the interface parameter filling to the vehicle to complete the voice interaction.
[0060] This application provides a voice interaction device. The voice interaction method of this application can be implemented by the voice interaction device of this application. Specifically, the voice interaction device includes a receiving module, a processing module, and an interaction module. The receiving module receives a first voice request from the current dialogue round forwarded by the vehicle. The processing module performs application programming interface (API) prediction and interface parameter filling on the first voice request based on the first voice request, a second voice request, and the natural language understanding result of the second voice request, wherein the second voice request is the voice request from the previous dialogue round. The interaction module sends the execution result of the interface parameter filling to the vehicle to complete the voice interaction.
[0061] This application also provides a server, which includes a memory and a processor. The voice interaction method of this application can be implemented by the server of this application. Specifically, the memory stores a computer program, the processor is used to receive a first voice request of the current dialogue round forwarded by the vehicle, and to perform application programming interface prediction and interface parameter filling on the first voice request based on the first voice request, a second voice request, and the natural language understanding result of the second voice request, wherein the second voice request is the voice request of the previous dialogue round of the current dialogue round, and to send the execution result of the interface parameter filling to the vehicle to complete the voice interaction.
[0062] Specifically, in the embodiments of this application, when a user engages in multiple rounds of dialogue with a vehicle in order to cause the vehicle to perform corresponding actions, the vehicle can send each round of dialogue expressed by the user to the server. Therefore, for the first voice request expressed by the user in the current round of dialogue, the vehicle can forward or report the first voice request to the server.
[0063] When the server receives the first voice request forwarded by the vehicle for the current dialogue round, it can obtain the second voice request forwarded by the vehicle for the previous dialogue round, and obtain the natural language understanding result of the second voice request.
[0064] Furthermore, based on the first voice request, the second voice request, and the natural language understanding results of the second voice request, the server can predict the application interface (API) for the first voice request to implement the vehicle's function in fulfilling the first voice request (or user intent). Simultaneously, since API calls typically require corresponding input parameters, the server can also perform parameter filling processing on the API prediction results to obtain the parameter-filled results.
[0065] Additionally, the server can send the parameter filling results to the vehicle, so that the vehicle can perform corresponding operations based on the parameter filling results, satisfy the user's intentions, and complete the voice interaction with the user.
[0066] For a clearer illustration of the implementation methods of this application, please refer to [link / reference]. Figure 2 , Figure 2 This is a schematic diagram illustrating application scenarios in some embodiments of this application. Specifically, in one example, let the first voice request be denoted as now_query and its value be "set to 28 degrees," and let the second voice request be denoted as last_query and its value be "turn on the air conditioner." Then:
[0067] After receiving now_query, the server reads last_query and its natural language understanding result from the previous dialogue round, based on a pre-set history tracker that stores voice requests and their natural language understanding results for each dialogue round.
[0068] Next, based on the pre-trained natural language processing model, the server performs application interface prediction and parameter filling for now_query according to the natural language understanding results of now_query, last_query, and last_query. This determines that the API corresponding to now_query is AcSet. If the device parameter of AcSet is filled with "air conditioner" and the value parameter is filled with "28 degrees", the execution result of the application interface parameter filling will be represented in a preset format. The obtained execution result is: [{"apiName":"AcSet","arguments":[{"name":"device","type":"string","value":"air conditioner},{"name":"value","type":"string","value":"28 degrees"}]}].
[0069] Finally, the server can send the execution result [{"apiName":"AcSet","arguments":[{"name":"device","type":"string","value":"air conditioning"},{"name":"value","type":"string","value":"28 degrees"}]}] to the vehicle. The vehicle can parse the received execution result into corresponding control commands and execute them, setting the air conditioning temperature to 28 degrees in response to the user's voice request and the server.
[0070] It is understood that, in the embodiments of this application, the natural language understanding result of the first voice request may include the application prediction result corresponding to the first voice request and the execution result of filling in the application interface parameters.
[0071] Therefore, in some embodiments of this application, the natural language understanding result of the second voice request includes the application prediction corresponding to the second voice request and the target parameters for filling the application prediction. For example, in one example, the second voice request last_query is "turn on the air conditioner". The natural language understanding result of the second voice request last_query may include {last_api:AcOpen,last_arguments:[{"device":"air conditioner"}]}, that is, the API prediction corresponding to "turn on the air conditioner" is AcOpen, the parameters of AcOpen include device, and the target parameter filled into device is "air conditioner".
[0072] It is also understandable that historical turn information (i.e., the second voice request and its natural language understanding result) is crucial for maintaining the coherence and contextual relevance of the dialogue, helping the dialogue system adjust its current response based on preceding information. Therefore, the server in this embodiment can read pre-stored second voice requests and their natural language understanding results from the previous dialogue turn to perform natural language processing on the first voice request of the current dialogue turn.
[0073] Furthermore, it is understood that the embodiments of this application can predict the application interface of the voice request and perform application interface parameter filling based on a pre-trained natural language processing model, and the natural language processing model used in the embodiments of this application can be set according to the actual situation.
[0074] For example, in some embodiments of this application, the natural language processing model used in the embodiments of this application includes a generative model that has been pre-trained and is capable of performing application interface prediction and application interface parameter filling of the first voice request sample based on the received first voice request sample, second voice request sample and second voice request sample.
[0075] Furthermore, in some embodiments of this application, the generative model is obtained based on pre-trained weights of BERT (Bidirectional Encoder Representations from Transformers).
[0076] Optionally, in some embodiments of this application, the vehicle may also broadcast specific voice messages such as "OK, executed" to prompt the user when executing the corresponding instructions based on the execution result.
[0077] Optionally, in some embodiments of this application, the vehicle may also send corresponding response information to the server when executing corresponding instructions based on the execution result, so that the server can determine the instruction execution status on the vehicle side.
[0078] In summary, in this embodiment, the processing of the first voice request in the current dialogue round can be based on the second voice request from the previous dialogue round and the natural language understanding result of the second voice request. This achieves context-based voice interaction, thereby ensuring the reliability of the application's prediction results and parameter filling execution results of the first voice request to a certain extent. This guarantees the reliable execution of voice interaction and ensures a good user experience for the voice interaction function. Furthermore, when the voice request in the current dialogue round is relatively short, the voice request and natural language understanding result from the previous dialogue round can be tracked to determine the user intent and perform natural language understanding corresponding to the voice request in the current dialogue round. This further ensures the reliable processing of the voice request in the current dialogue round and improves the applicability of the voice interaction function in complex multi-turn dialogue scenarios.
[0079] Please see Figure 3 In some embodiments of this application, step 02 includes:
[0080] 020: Perform splicing processing based on the first voice request, the second voice request, and the natural language understanding results;
[0081] 021: Based on the splicing process, perform application interface prediction and interface parameter filling for the first voice request.
[0082] The processing module in this application embodiment is further configured to perform splicing processing on the first voice request, the second voice request, and the natural language understanding result, and to perform application interface prediction and interface parameter filling on the first voice request based on the splicing processing result.
[0083] The processor in this embodiment is further configured to perform concatenation processing on the first voice request, the second voice request, and the natural language understanding result, and to perform application interface prediction and interface parameter filling on the first voice request based on the result of the concatenation processing.
[0084] Specifically, in the embodiments of this application, the server can concatenate the first voice request, the second voice request, and the natural language understanding results of the second voice request, and then call the corresponding tools, programs, or natural language processing models to predict the application interface and fill in the interface parameters of the first voice request through the concatenation process.
[0085] Optional, please refer to Figure 4 and Figure 5 , Figure 4 This is a schematic diagram illustrating application scenarios in some embodiments of this application. Figure 5This is a schematic diagram illustrating the result of splicing processing in certain embodiments of this application. That is, in certain embodiments of this application, the server can splice the first voice request (i.e., "set to 28 degrees"), the second voice request ("turn on the air conditioner"), and the natural language understanding result of the second voice request (i.e., the API prediction "AcOpen" and the target parameter "air conditioner" used to fill in "AcOpen") to form a result similar to... Figure 5 After the splicing process is shown, the result is input into a pre-trained natural language processing model so that the natural language processing model can predict the API of the first voice request and fill in the parameters of the API.
[0086] What is understandable is that Figure 5 In the splicing result shown, "the result of the splicing process" is a serialized vector, therefore CLS indicates the starting position of this serialized vector. Also, it should be noted that... Figure 5 The splicing results shown are as follows: <value> This indicates the starting position of the target parameters of the second voice request in the serialization vector.< / value> This indicates the "end position of the target parameter of the second voice request" in the serialization vector.
[0087] It's also understandable that CLS, <value> and< / value> It can be used to enable natural language processing models to determine the meaning and boundaries of each data point in the input serialized vector, thereby enabling the model to reliably and accurately predict the API of the first voice request and fill in the parameters of the API based on each data point in the serialized vector.
[0088] Furthermore, it is also understood that, in the embodiments of this application, the server concatenates the first voice request, the second voice request, and the natural language understanding result of the second voice request to form a... Figure 5 Before the concatenation process shown, the server can perform over-segmentation (tokenization) on the first voice request, the second voice request, and the natural language understanding results of the second voice request to convert them into semantic units (tokens). Then, when performing the concatenation process in the subsequent process, the server can concatenate the first voice request, the second voice request, and the natural language understanding results of the second voice request into a serialized vector.
[0089] It is understandable that, in cases like Figure 5 In the splicing results shown, the data within each box can be understood as a semantic unit, such as "AcOpen" and "open".
[0090] Optionally, in some embodiments of this application, since the server has already completed tokenization of both the second voice request and its natural language understanding result, the server can then perform tokenization only on the first voice request, and then concatenate the tokenized first voice request with the second voice request and its natural language understanding result to form a tokenized message. Figure 5 The splicing result is shown.
[0091] Optionally, in some embodiments of this application, after receiving a voice request forwarded by a vehicle, the server may perform the aforementioned word segmentation processing on the voice request to obtain all semantic units contained in the voice request, and store all semantic units contained in the voice request. Furthermore, after predicting the application programming interface (API) of the voice request and completing the parameter filling process for the API, the server may also store the predicted API and the filled target parameters of the voice request.
[0092] Furthermore, when the server reads the voice request and its natural language understanding result from the previous dialogue round in the current dialogue round, it can read multiple semantic units of the voice request. Therefore, in subsequent concatenation processing, the voice request from the current dialogue round, the voice request from the previous dialogue round, and the natural language understanding result of the voice request from the previous dialogue round can be directly concatenated to form a structure similar to... Figure 5 The splicing result is shown.
[0093] Thus, in this embodiment of the application, the first voice request, the second voice request, and the natural language understanding result of the second voice request can be concatenated, and the application programming interface (API) prediction and interface parameter filling of the first voice request can be performed based on the concatenation result.
[0094] In some embodiments of this application, step 02 includes:
[0095] Configure the dialogue turn label for each parameter in the result of the splicing process.
[0096] The processing module in this application embodiment is also used to configure the dialogue turn label for each parameter in the result of the splicing process.
[0097] The processor in this embodiment is also used to configure the dialogue turn label for each parameter in the result of the splicing process.
[0098] Specifically, in order to distinguish parameters of different dialogue rounds in the result of splicing, the embodiments of this application may also configure a dialogue round label for each parameter in the result of splicing.
[0099] For example, please refer to Figure 6 , Figure 6 This is a schematic diagram illustrating the splicing process in certain embodiments of this application. Specifically, the dialog round labels for the parameters of the current dialog round (such as semantic units 'set', 'as', 'two', 'ten', 'eight', and 'degree') are configured to 0, and the parameters of the previous dialog round (such as semantic units 'AcOpen', 'open', 'open', 'empty', 'adjust', '...') are... <value> '、'air conditioning' and '< / value> The dialogue turn label is configured to 1.
[0100] Optionally, in some embodiments of this application, the dialogue turn label is represented as a preset turn-embedding vector.
[0101] Understandably, based on Figure 2 or Figure 4 The method shown is that, through a pre-trained natural language processing model, ... Figure 6 When the concatenation result is used for inference to predict the application interface and fill in the application interface parameters for the first voice request, the natural language processing model can identify the connection and difference between the current dialogue turn and the previous dialogue turn after receiving the concatenation result because the "concatenation result" is configured with dialogue turn labels.
[0102] Thus, in this embodiment of the application, when the natural language understanding results of the first voice request, the second voice request, and the second voice request are concatenated to obtain the concatenation result, the dialogue turn labels of each parameter in the concatenation result can be configured so that the parameters belonging to the current dialogue turn and the parameters belonging to the previous dialogue turn in the concatenation result can be distinguished based on the dialogue turn labels.
[0103] Furthermore, when performing application interface prediction and parameter filling using the results of splicing processing, the connection and differences between the current dialogue round and the previous dialogue round can be identified, thereby improving the accuracy of application interface prediction and parameter filling to a certain extent.
[0104] In some embodiments of this application, the voice interaction method further includes:
[0105] Slot identification is performed on the first voice request.
[0106] Therefore, step 021 above includes:
[0107] The first voice request, the slot recognition result, the second voice request, and the natural language understanding result are concatenated to obtain the concatenated result.
[0108] The voice interaction device according to this application embodiment further includes a slot recognition module. The slot recognition module is used to perform slot recognition on the first voice request. The processing module according to this application embodiment is also used to perform concatenation processing on the first voice request, the result of slot recognition, the second voice request, and the natural language understanding result to obtain a concatenation processing result.
[0109] The processor in this embodiment is further configured to perform slot identification on the first voice request, and to perform splicing processing on the first voice request, the result of slot identification, the second voice request, and the natural language understanding result to obtain the splicing processing result.
[0110] Understandably, slot information (or named entities) can be used to understand user voice requests. Furthermore, in in-vehicle voice interaction scenarios, the dialogue system needs to quickly and accurately extract vehicle-related entities from user requests, such as temperature, volume, and navigation destination.
[0111] Therefore, in this embodiment, the server can also perform slot identification on the first voice request before predicting the API of the first voice request, so as to obtain the corresponding slot identification result. For example, in one example, the first voice request is "set to 28 degrees", then the slot identification result of the first voice request may include [{"value":"28 degrees","pos":[3,6]}]. Wherein, "pos":[3,6] indicates that the sequence numbers of each semantic unit in the entity "28 degrees" in "set to 28 degrees" are 3, 4, 5 and 6.
[0112] It is also understood that, upon obtaining the slot recognition result of the first voice request, the server can concatenate the first voice request, the slot recognition result of the first voice request, and the natural language understanding result of the second voice request to obtain the concatenated result. For a clearer illustration of the implementation method of this application, please refer to... Figure 7 and Figure 8 , Figure 7 This is a schematic diagram illustrating application scenarios in some embodiments of this application. Figure 8 This is a schematic diagram of the splicing process results in some embodiments of this application.
[0113] That is, such as Figure 7 As shown, the server in this embodiment can perform named entity recognition on the first voice request using a pre-trained named entity recognition model to determine the named entities of each slot type in the first voice request, thereby obtaining the slot recognition result. For example, for the first voice request "set to 28 degrees", the slot recognition result may include [{"value":"28 degrees","pos":[3,6]}].
[0114] Furthermore, in the embodiments of this application, the server, upon obtaining the slot recognition result of the first voice request, can concatenate the first voice request, the slot recognition result of the first voice request, and the natural language understanding result of the second voice request to enable... Figure 7 Natural language processing models in the language can be based on, for example, Figure 8 The splicing result shown is used for API prediction and API parameter filling of the first voice request.
[0115] Furthermore, it is understandable that Figure 8 The numbers in the second row can be understood as dialogue turn labels, used to indicate the dialogue turn to which each parameter (or semantic unit) in the first row belongs.
[0116] For example, a dialogue turn label of 0 indicates that the parameters above belong to the current dialogue turn. Semantic units such as 'set', 'for', 'two', 'ten', 'eight', and 'degree' belong to the first voice request of the current dialogue turn.
[0117] A dialogue turn label of 1 indicates that the parameter above belongs to the previous dialogue turn, such as 'AcOpen', 'open', 'open', 'empty', 'adjust', 'open'. <value> '、'air conditioning' and '< / value> 'Semantic units such as ' belong to the second voice request of the previous dialogue round.'
[0118] Thus, in this embodiment of the application, slot identification can be performed on the first voice request, and the first voice request, the result of slot identification of the first voice request, and the natural language understanding result of the second voice request can be spliced together to obtain the splicing result.
[0119] Optionally, in this embodiment, the natural language understanding result of the second voice request includes the interface padding parameters of the second voice request, and step 03 includes:
[0120] Based on the slot identification results and interface filling parameters, parameter filling processing is performed on the application interface prediction results to obtain the execution result of parameter filling processing.
[0121] The processing module in this embodiment is further configured to perform parameter filling processing on the application interface prediction result based on the slot identification result and the interface filling parameter, so as to obtain the execution result of the parameter filling processing.
[0122] The processor in this embodiment is further configured to perform parameter filling processing on the application interface prediction result based on the slot identification result and the interface filling parameter, so as to obtain the execution result of the parameter filling processing.
[0123] Specifically, in this embodiment of the application, the server can perform parameter filling of the API of the first voice request based on the result of slot identification of the first voice request.
[0124] For example, the API prediction for the first voice request is AcSet, and the slot recognition result for the first voice request includes [{"value":"twenty-eight degrees"}]. The API for the second voice request is AcOpen, and the interface filling parameters for AcOpen include [{"device":"air conditioner"}]. Therefore, the device parameter of AcSet can be filled with "air conditioner", and the value parameter of AcSet can be filled with "twenty-eight degrees".
[0125] Thus, in this embodiment of the application, the prediction result of the application interface of the first voice request can be parameter-filled based on the slot identification result of the first voice request and the interface filling parameters of the application interface of the second voice request.
[0126] In some embodiments of this application, step 02 includes:
[0127] Configure the position label for each parameter in the result of the splicing process.
[0128] The processing module in this application embodiment is also used to configure the position label of each parameter in the result of the splicing process.
[0129] The processor in this embodiment is also used to configure the position label of each parameter in the result of the splicing process.
[0130] Specifically, in this application embodiment, the server can also configure a position label for each parameter in the splicing result, so that each parameter in the splicing result can be associated and distinguished based on the position label.
[0131] Optionally, in some embodiments of this application, the server may set location encoding in the result of the splicing process to configure location tags.
[0132] For a clearer illustration of the implementation methods of this application, please refer to [link / reference]. Figure 9 , Figure 9 This is a schematic diagram illustrating the splicing process in certain embodiments of this application. That is, when each semantic unit in the first voice request and each semantic unit in the second voice request are spliced together and the CLS marker is set, the splicing result can be a sequence composed of eleven semantic units: 'CLS', 'set', 'as', 'two', 'ten', 'eight', 'degree', 'open', 'empty', and 'tone', ordered and combined in sequence.
[0133] Furthermore, since the slot recognition result of the first voice request includes [{"value":"twenty-eight degrees"}], and the natural language understanding result of the second voice request includes [{"device":"air conditioner"}], and the entities "twenty-eight degrees" and "air conditioner" have already appeared in the "sequence of eleven semantic units ordered and combined in sequence", the parameter corresponding to the slot type "value" can be set in the sequence based on the positions of "twenty-eight degrees" and "air conditioner" in that sequence (i.e., "value"). <value> and< / value> ") and the parameters corresponding to the device slot type (i.e., " <device> and< / device> (), and configure the position labels corresponding to these parameters, that is, to <value> Configured as 3< / value> The value is configured as 6 to indicate that the semantic unit corresponding to the entity of the slot type "value" is at the beginning and end of the sequence, with positions 3 and 6.
[0134] Similarly, <device> Configured to 10< / device> The configuration is set to 11 to indicate that the semantic unit corresponding to the entity of the slot type device is 10 and 11 at the beginning and end of the sequence.
[0135] Understandably, compared to directly concatenating the slot recognition result of the first voice request and the natural language understanding result of the second voice request, or rather, compared to... Figure 5 , 6 Regarding the splicing process shown in Figure 8, Figure 9 The information in the document is more concise.
[0136] Furthermore, in the case of natural language processing models (such as...) Figure 2 , 4 7) When reasoning about the results of the splicing process, the natural language processing model is based on, for example... Figure 9 The splicing results shown can capture and understand the association between "slot recognition result of the first voice request" and "first voice request", and the association between "natural speech understanding result of the second voice request" and "second voice request", thereby improving the accuracy of reasoning to a certain extent.
[0137] Thus, in this embodiment of the application, when the natural language understanding results of the first voice request, the second voice request, and the second voice request are spliced together to obtain the splicing result, the position labels of each parameter in the splicing result can be configured so that the parameters in the splicing result can be distinguished and associated based on the position labels, thereby improving the accuracy of application interface prediction and parameter filling to a certain extent.
[0138] In some embodiments of this application, the voice interaction method further includes:
[0139] Based on the first voice request, the second voice request, and the natural language understanding results, the dialogue association attribute information of the first voice request relative to the second voice request is determined, wherein the dialogue association attribute information is used to indicate whether the first voice request and the second voice request are related.
[0140] The voice interaction device according to the embodiments of this application further includes an attribute information determination module. The attribute information determination module is used to determine the dialogue association attribute information of the first voice request relative to the second voice request based on the first voice request, the second voice request, and the natural language understanding result, wherein the dialogue association attribute information is used to indicate whether the first voice request and the second voice request are associated.
[0141] The processor in this embodiment is further configured to determine dialogue association attribute information of the first voice request relative to the second voice request based on the first voice request, the second voice request, and the natural language understanding result, wherein the dialogue association attribute information is used to indicate whether the first voice request and the second voice request are associated.
[0142] Specifically, in this embodiment of the application, the server can also determine the dialogue association attribute information of the first voice request relative to the second voice request based on the first voice request, the second voice request and the natural language understanding result, that is, determine whether the current round of dialogue is semantically continuous with the previous round of dialogue.
[0143] Optionally, for a clearer illustration of the embodiments of this application, please refer to [link to relevant documentation]. Figure 10 , Figure 10 This is a schematic diagram illustrating application scenarios in certain embodiments of this application. Specifically, when the server performs inference based on a pre-trained natural language processing model according to the first voice request, the second voice request, and the natural language understanding result; or, when receiving a result concatenated from the first voice request, the slot recognition result of the first voice request, the second voice request, and the natural language understanding result, the natural language processing model can output dialogue-related attribute information (corresponding to CLS tags) based on the CLS tags. Figure 10 The `inheritLabel` in the model (and the hidden state of the natural language processing model) can be used for API prediction and API parameter filling for the first voice request.
[0144] Furthermore, when inheritLabel is 1, it indicates that the first voice request and the second voice request are semantically continuous. Conversely, when inheritLabel is 0, it indicates that the first voice request and the second voice request are not semantically continuous.
[0145] Thus, in this embodiment of the application, the dialogue association attribute information of the first voice request relative to the second voice request can be determined based on the first voice request, the second voice request, and the natural language understanding result, thereby determining the relationship between the current round of dialogue and the previous round of dialogue.
[0146] In some embodiments of this application, step 03 includes:
[0147] Send the execution result and dialogue-related attribute information to the vehicle to complete the voice interaction.
[0148] The interaction module in this embodiment is also used to send execution results and dialogue-related attribute information to the vehicle to complete voice interaction.
[0149] The processor in this embodiment is also used to send execution results and dialogue-related attribute information to the vehicle to complete voice interaction.
[0150] It is understandable that when the first voice request and the second voice request are semantically consecutive, it indicates that the processing of the first voice request can utilize information from previous dialogue rounds. Conversely, when the first voice request and the second voice request are not semantically consecutive, it indicates that the processing of the first voice request can ignore information from previous dialogue rounds.
[0151] Therefore, the server in this application implementation can also send dialogue association attribute information to the vehicle, so that the vehicle can determine whether the relevant processing of the first voice request can use the information from the previous rounds based on the received dialogue association attribute information.
[0152] To more clearly illustrate the implementation methods of this application, please refer again. Figure 10 That is, in the embodiments of this application, the server can send the output of the natural language processing model [{"apiName":"AcSet","arguments":[{"name":"device","type":"string","value":"air conditioning"},{"name":"value","type":"string","value":"twenty-eight degrees"}],"inheritLabel":1}] to the vehicle, so that the vehicle can perform the relevant processing of the first voice request according to the value of the inheritLabel label (which is 1).
[0153] It should also be noted that, in the embodiments of this application, "the processing of the first voice request" can be understood as the vehicle's TTS (Text To Speech) prompts. For example, if a user says "turn on the air conditioner" and then immediately says "set to 28 degrees," the vehicle, based on the inheritLabel, confirms that "turn on the air conditioner" and "set to 28 degrees" are semantically continuous, and can therefore provide the user with a "response that is also semantically continuous," thereby achieving a natural dialogue flow.
[0154] Thus, in this embodiment of the application, the execution result of the filling process of the application interface parameters and the dialogue association attribute information of the first voice request relative to the second voice request can be sent to the vehicle, so that the vehicle can perform corresponding operations based on the execution result and the dialogue association attribute information.
[0155] Optionally, in the embodiments of this application, Figure 2 , 4 The process of obtaining natural language processing models 7 and 10 may include: fine-tuning the generative model using the pre-trained weights of BERT.
[0156] Understandably, fine-tuning generative models can effectively improve a system's natural language understanding capabilities. Furthermore, combining pre-trained models such as BERT for fine-tuning can significantly optimize model performance.
[0157] This application also provides a computer-readable storage medium storing a computer program that, when executed by one or more processors, implements the above-described voice interaction method.
[0158] In this specification, the terms "specifically," "furthermore," "particularly," "understandably," etc., refer to specific features, structures, materials, or characteristics described in connection with embodiments or examples that are included in at least one embodiment or example of this application. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.
[0159] Any process or method described in the flowchart or otherwise herein can be understood as representing a module, segment, or portion of code comprising one or more executable instructions for implementing a particular logical function or process, and the scope of the preferred embodiments of this application includes additional implementations in which functions may be performed not in the order shown or discussed, including substantially simultaneously or in reverse order depending on the function involved, as will be understood by those skilled in the art to which embodiments of this application pertain.
[0160] Although embodiments of this application have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting this application. Those skilled in the art can make changes, modifications, substitutions and variations to the above embodiments within the scope of this application.
Claims
1. A voice interaction method, characterized in that, include: Receive the first voice request of the current dialogue round forwarded by the vehicle; Based on the first voice request, the second voice request, and the natural language understanding results of the second voice request, the first voice request is subjected to application interface prediction and interface parameter filling, wherein the second voice request is the voice request of the previous dialogue round of the current dialogue round; The execution result of the interface parameter filling is sent to the vehicle to complete the voice interaction; The natural language understanding result of the second voice request includes the application prediction corresponding to the second voice request and the target parameters for filling the application prediction; The step of performing application programming interface (API) prediction and API parameter filling on the first voice request based on the first voice request, the second voice request, and the natural language understanding results of the second voice request includes: The process involves concatenating the first voice request, the second voice request, and the natural language understanding result. Based on the result of the splicing process, the first voice request is subjected to application interface prediction and interface parameter filling.
2. The method according to claim 1, characterized in that, The step of performing application programming interface (API) prediction and API parameter filling on the first voice request based on the first voice request, the second voice request, and the natural language understanding results of the second voice request includes: Configure the dialogue turn label for each parameter in the result of the splicing process.
3. The method according to claim 1, characterized in that, The method further includes: Slot identification is performed on the first voice request; The concatenation process based on the first voice request, the second voice request, and the natural language understanding result includes: The first voice request, the slot recognition result, the second voice request, and the natural language understanding result are concatenated to obtain the concatenation result.
4. The method according to claim 3, characterized in that, The step of performing application programming interface (API) prediction and API parameter filling on the first voice request based on the first voice request, the second voice request, and the natural language understanding results of the second voice request includes: Configure the position label for each parameter in the result of the splicing process.
5. The method according to claim 1, characterized in that, The method further includes: Based on the first voice request, the second voice request, and the natural language understanding result, the dialogue association attribute information of the first voice request relative to the second voice request is determined, wherein the dialogue association attribute information is used to indicate whether the first voice request and the second voice request are associated.
6. The method according to claim 5, characterized in that, The step of sending the execution result of the interface parameter filling to the vehicle to complete the voice interaction includes: The execution result and the dialogue association attribute information are sent to the vehicle to complete the voice interaction.
7. A server, characterized in that, The method includes a memory and a processor, wherein the memory stores a computer program, which, when executed by the processor, implements the method according to any one of claims 1-6.
8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by one or more processors, implements the method of any one of claims 1-6.
Citation Information
Patent Citations
Voice conversation interaction method and device thereof, vehicle and medium
CN114005447A
Vehicle interaction method and device, model training method and device, server and storage medium
CN116758913A