Voice interaction method, vehicle, and computer-readable storage medium

Through slot recognition and semantic unit correlation analysis, complex speech requests are processed in sentences, which solves the problem that traditional systems are difficult to understand short continuous reading instructions, and achieves precise control of the vehicle and improves user experience.

CN118262718BActive Publication Date: 2025-05-30GUANGZHOU XIAOPENG MOTORS TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202410391180.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-04-01
Publication Date
2025-05-30
Estimated Expiration
2044-04-01

AI Technical Summary

Technical Problem

Traditional natural language understanding systems have difficulty understanding short continuous reading voice control instructions, resulting in a decrease in vehicle response accuracy and affecting the user's voice interaction experience.

Method used

Through slot recognition and semantic unit correlation analysis, complex speech requests are processed in sentences, and application program interface prediction and parameter filling are used to achieve precise control of the vehicle.

Benefits of technology

It improves the accuracy of natural language understanding, simplifies the system architecture, reduces processing difficulty, and improves the user's voice interaction experience and driving experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118262718B_ABST
    Figure CN118262718B_ABST
Patent Text Reader

Abstract

The present application discloses a voice interaction method, a vehicle, and a computer-readable storage medium. The method includes obtaining a current voice request, performing slot recognition on the current voice request, performing clause splitting processing on the current voice request according to the correlation degree between any two semantic units in the current voice request to determine a plurality of target clauses, performing application program interface prediction and application program interface parameter filling on each target clause according to the target clause and the result of slot recognition to obtain a target execution result for each target clause, and performing control according to the target execution result to complete voice interaction. In this way, the present application can split the voice request based on the correlation degree between any two semantic units in the voice request to obtain a plurality of clauses. The natural language processing of the voice request can be completed through each clause, the difficulty of natural language processing is reduced, the accuracy of the natural language understanding result can be ensured, and the voice interaction experience and driving experience of the user can be guaranteed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of voice interaction technology, and particularly to a voice interaction method, a vehicle, and a computer-readable storage medium. Background Art

[0002] As users gradually get used to performing operations through voice interaction systems, they tend to use simple instructions to control the system to perform complex operations, such as using the sentence "Open the air conditioner and windows" to replace the two sentences "Open the air conditioner" and "Open the windows". However, traditional natural language understanding systems are difficult to understand such simple instructions for performing complex operations, thus affecting the user's voice interaction experience. Summary of the Invention

[0003] This application provides a voice interaction method, a vehicle, and a computer-readable storage medium.

[0004] A voice interaction method provided by an embodiment of this application for a vehicle includes:

[0005] Obtain a current voice request;

[0006] Perform slot recognition on the current voice request;

[0007] Perform clause splitting on the current voice request according to the correlation degree between any two semantic units in the current voice request to determine a plurality of target clauses;

[0008] According to the plurality of target clauses and the result of the slot recognition, perform application programming interface prediction and application programming interface parameter filling on each target clause to obtain a target execution result of the application programming interface parameter filling for each target clause;

[0009] Control the vehicle according to the target execution result to complete voice interaction.

[0010] In the voice interaction method provided by the embodiment of this application, the vehicle can obtain a current voice request, perform slot recognition on the current voice request to obtain a corresponding slot recognition result, perform clause splitting on the current voice request according to the correlation degree between any two semantic units in the current voice request to determine a plurality of target clauses, perform prediction of the application programming interface and filling of the application programming interface parameters on each target clause according to the determined target clauses and the result of the slot recognition, so as to obtain a target execution result of the application programming interface parameter filling for each target clause, and execute corresponding functions according to the target execution result to complete voice interaction.

[0011] Thus, in the embodiments of the present application, the vehicle can segment the current voice request based on the correlation degree between any two semantic units in the current voice request to obtain multiple target clauses, enabling the segmentation of the current voice request. Furthermore, when the natural language processing difficulty of the current voice request is relatively high, the natural language processing of the current voice request can be converted into the natural language processing of multiple target clauses, thereby reducing the natural language processing difficulty of the voice request to a certain extent. As a result, the accuracy of the natural language understanding result can be ensured, and then the interaction with the user can be completed based on the natural language processing result with guaranteed accuracy, ensuring the user's voice interaction experience and driving experience.

[0012] In some embodiments of the present application, the method further includes:

[0013] Determine the correlation degree between any two semantic units in the current voice request according to the entity in the current voice request.

[0014] Thus, in the embodiments of the present application, the vehicle can determine the correlation degree between any two semantic units in the current voice request according to the entity in the current voice request, to a certain extent making the credibility of the determined correlation degree result.

[0015] In some embodiments of the present application, the determining the correlation degree between any two semantic units in the current voice request according to the entity in the current voice request includes:

[0016] When the two semantic units are adjacent in the current voice request and correspond to one entity, determine that the two semantic units are correlated at a first degree.

[0017] Thus, in the embodiments of the present application, the vehicle can determine the correlation degree between two semantic units in the current voice request as the first degree when the two semantic units in the current voice request are adjacent and correspond to the same entity.

[0018] In some embodiments of the present application, the determining the correlation degree between any two semantic units in the current voice request according to the entity in the current voice request includes:

[0019] When the two entities corresponding to the two semantic units in the current voice request are consecutive, determine that the two semantic units are correlated at a first degree.

[0020] Thus, in the embodiments of the present application, the vehicle can determine the correlation degree between two semantic units as the first degree when the two semantic units in the current voice request are adjacent and the two entities corresponding to them are also adjacent in the current voice request.

[0021] In some embodiments of the present application, the two semantic units include a first unit and a second unit, and determining the degree of correlation between any two semantic units in the current voice request according to the entity in the current voice request includes:

[0022] When the two semantic units correspond to the first boundary entity and the second boundary entity in the current voice request, the first unit is located at the first position of the first boundary entity, and the second unit is located at the last position of the second boundary entity, it is determined that the two semantic units are related to the second degree.

[0023] Thus, in an embodiment of the present application, the vehicle can confirm that the correlation between the first unit and the second unit is the second degree when the first unit in the current voice request is located at the first position of the first boundary entity and the second unit is located at the last position of the second boundary entity.

[0024] In certain embodiments of the present application, the sentence processing of the current voice request and determining a plurality of target sentences according to the correlation between any two semantic units in the current voice request include:

[0025] Constructing a semantic unit relationship matrix according to the correlation degree;

[0026] The plurality of target sentences are determined according to the semantic unit relationship matrix.

[0027] Thus, in the implementation mode of the present application, the vehicle can construct a semantic unit relationship matrix through the correlation between any two semantic units in the current voice request and use the semantic unit relationship matrix to determine the target sentence, so that the credibility of the target sentence can be guaranteed to a certain extent.

[0028] In certain embodiments of the present application, the step of performing application program interface prediction and application program interface parameter filling on each target sentence according to the multiple target sentences and the results of the slot identification, and obtaining a target execution result of application program interface parameter filling for each target sentence includes:

[0029] Determine a slot identification sub-result of each target sentence according to the plurality of target sentences and the slot identification results;

[0030] According to each of the target clauses and the slot identification sub-result of each of the target clauses, the application program interface prediction and the application program interface parameter filling are performed on each of the target clauses to obtain the target execution result of each of the target clauses.

[0031] Thus, in the embodiments of the present application, the vehicle can perform application programming interface (API) prediction and parameter filling of the API for each target clause through the target clause and the slot recognition sub-result of the target clause, which to a certain extent ensures the reliable execution of the API prediction and the parameter filling of the API.

[0032] In some embodiments of the present application, the slot recognition of the current voice request includes:

[0033] Performing sequence labeling processing on the current voice request based on a pre-trained recognition model to complete the slot recognition.

[0034] Thus, in the embodiments of the present application, the vehicle can perform sequence labeling processing on the current voice request through a pre-trained generative model to obtain the slot recognition result of the current voice request, and further, the credibility of the slot recognition result can be guaranteed to a certain extent.

[0035] An embodiment of the present application provides a vehicle, including a memory and a processor. A computer program is stored in the memory. When the computer program is executed by the processor, the above-mentioned voice interaction method is implemented.

[0036] An embodiment of the present application provides a computer-readable storage medium. The computer-readable storage medium stores a computer program. When the computer program is executed by one or more processors, the above-mentioned voice interaction method is implemented.

[0037] The vehicle and the computer-readable storage medium provided by the embodiments of the present application can segment the current voice request based on the correlation degree of any two semantic units in the current voice request to obtain multiple target clauses, so as to realize the segmentation of the current voice request. Furthermore, when the natural language processing difficulty of the current voice request is relatively high, the natural language processing of the current voice request can be converted into the natural language processing of multiple target clauses. Therefore, the natural language processing difficulty of the voice request is reduced to a certain extent. As a result, the accuracy of the natural language understanding result can be guaranteed, and then the interaction with the user can be completed based on the natural language processing result with guaranteed accuracy, so as to ensure the user's voice interaction experience and driving experience.

[0038] The additional aspects and advantages of the embodiments of the present application will be partially given in the following description, partially become obvious from the following description, or be understood through the practice of the embodiments of the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] The above and / or additional aspects and advantages of the present application will become obvious and easy to understand from the description of the embodiments in conjunction with the following drawings, where:

[0040] Figure 1 It is a schematic flow chart of a voice interaction method in some embodiments of the present application;

[0041] Figure 2 It is a schematic flow chart of a voice interaction method in some embodiments of the present application;

[0042] Figure 3 It is a schematic diagram of a semantic unit relationship matrix in some embodiments of the present application;

[0043] Figure 4 It is a schematic diagram of an application scenario in some embodiments of the present application;

[0044] Figure 5 It is a schematic diagram of an application scenario in some embodiments of the present application;

[0045] Figure 6 It is a schematic diagram of an application scenario in some embodiments of the present application;

[0046] Figure 7 It is a schematic flow chart of a voice interaction method in some embodiments of the present application;

[0047] Figure 8 It is a schematic flow chart of a voice interaction method in some embodiments of the present application;

[0048] Figure 9 It is a schematic flow chart of a voice interaction method in some embodiments of the present application;

[0049] Figure 10 It is a schematic flow chart of a voice interaction method in some embodiments of the present application. Detailed implementation manners

[0050] The following details the implementation manners of the present application. Examples of the implementation manners are shown in the accompanying drawings, where the same or similar reference numerals denote the same or similar elements or elements with the same or similar functions throughout. The implementation manners described below by referring to the accompanying drawings are exemplary and are only used to explain the implementation manners of the present application, and should not be construed as a limitation to the implementation manners of the present application.

[0051] As the application time of the voice interaction function in the vehicle scenario gradually increases, the challenges faced by the in-vehicle voice interaction system also gradually increase. One of them is how to cope with the usage habits gradually revealed by users during the actual use of the voice interaction function.

[0052] Specifically, after users initially start using the voice interaction function and become familiar with it, they often hope to express complex operation intentions through short sentences and expect the vehicle to correctly understand and process these sentences. For example, during the voice interaction with the vehicle, a user may use the sentence "Open the air conditioner and the window" to replace the two sentences "Open the air conditioner" and "Open the window". Ideally, the vehicle should be able to execute the two functions of "starting the vehicle air conditioner" and "opening the window" based on the sentence "Open the air conditioner and the window".

[0053] It can be understood that for the natural language understanding system (or rather, the natural language processing system) relied on by the vehicle, it is more difficult to understand a short sentence like "Open the air conditioner and the window" that includes a continuous speech control instruction, and it is necessary to accurately distinguish and process each instruction in this sentence.

[0054] However, for traditional natural language understanding systems, when faced with the above-mentioned "short sentences that include continuous speech control instructions" (hereinafter referred to as "complex instruction sentences"), they often can only recognize and understand one intention in the sentence while ignoring other intentions. For example, when faced with "Open the air conditioner and the window", it may only recognize "Open the window" and then cause the vehicle to execute the function of "starting the vehicle air conditioner", while ignoring the 'open the window' in "Open the air conditioner and the window", resulting in the failure to execute the function of "opening the window". This defect in understanding ability will affect the response accuracy of the vehicle to the user's voice instructions, and thus affect the user's voice interaction experience.

[0055] It can also be understood that during the process of the user uttering the above complex instruction sentence to control the vehicle, due to the shortness of the sentence and the few pauses when the user utters the sentence, when using VAD (Voice activity detection) technology to segment the sentence, a good segmentation effect cannot be achieved. That is, it is difficult to reliably or effectively segment the sentence into several voice control instructions, such as segmenting "Open the air conditioner and the window" into "Open the air conditioner" and "Open the window".

[0056] Furthermore, when it is difficult to segment the above complex instruction sentence into several independent parts through VAD technology and it is desired to enable the natural language understanding system to reasonably process the sentence, a higher requirement is imposed on the natural language understanding system's ability to understand the relationships between entities and words, such as being able to recognize the nested and discontinuous entities in the complex instruction sentence.

[0057] Regarding the above problems, the traditional solution is a pipeline-based design approach, which designs complex control steps and a system operation with several modules to ensure the understanding ability and interaction effect of the natural language understanding system. For example, the natural language understanding system can first perform intent segmentation of complex instruction statements through sentence boundary detection to identify the boundaries of each sentence in the compound sentence, and then process each independent intent separately. Subsequently, the natural language understanding system will perform natural language understanding on each independent intent one by one to identify and summarize these intents, so as to form a comprehensive response.

[0058] It can be understood that although this solution is feasible, it faces many challenges, including but not limited to overlapping entity recognition, long processes, error accumulation, and service timeouts.

[0059] Among them, overlapping entity recognition refers to the situation where when the system faces complex instruction statements, since complex instruction statements may contain several voice control instructions, and different voice control instructions may contain the same components. For example, "turn on the air conditioner and the window" includes two parts: "turn on the air conditioner" and "turn on the window", and both "turn on the air conditioner" and "turn on the window" contain 'turn on'. Or rather, "turn on the air conditioner" and "turn on the window" share the 'turn on' that only appears once in "turn on the air conditioner and the window".

[0060] A long process refers to the situation where when designing a process based on a pipeline, it is necessary to design and maintain multiple interdependent modules and services, resulting in a long complete processing flow of the system, that is, the maintenance and expansion costs of the system are relatively high.

[0061] Error accumulation refers to the fact that in the processing flow implemented based on a pipeline, such as the above-mentioned segmentation, understanding, and synthesis processes, errors are easily accumulated. That is, the errors in the previous step will be carried over to the next step, resulting in the final output result of the system being affected by the errors of each step, and the accuracy is difficult to guarantee, thereby affecting the user experience.

[0062] Service timeout refers to the situation where in the processing flow designed based on a pipeline, a single service response may time out, resulting in the processing time of the entire processing flow being difficult to meet the expectation, and the stability of the system and the user experience are both difficult to guarantee.

[0063] Based on the above possible problems, please refer to Figure 1 , the embodiments of the present application provide a voice interaction method for a vehicle, including:

[0064] 01: Obtain the current voice request;

[0065] 02: Perform slot recognition on the current voice request;

[0066] 03: According to the correlation between any two semantic units in the current voice request, the current voice request is processed into sentences to determine multiple target sentences;

[0067] 04: According to the results of multiple target clauses and slot identification, perform application program interface prediction and application program interface parameter filling for each target clause, and obtain the target execution result of the application program interface parameter filling for each target clause;

[0068] 05: Control the vehicle to complete voice interaction based on the target execution results.

[0069] The embodiment of the present application provides a voice interaction device for a vehicle. The voice interaction method of the embodiment of the present application can be implemented by the voice interaction device of the embodiment of the present application. Specifically, the voice interaction device includes an acquisition module for acquiring the current voice request. The recognition module is used to perform slot recognition on the current voice request. The processing module is used to perform sentence processing on the current voice request according to the degree of correlation between any two semantic units in the current voice request, and determine multiple target sentences. The filling module is used to perform application interface prediction and application interface parameter filling on each target sentence according to the results of multiple target sentences and slot recognition, and obtain the target execution result of the application interface parameter filling of each target sentence. The interaction module is used to control the vehicle according to the target execution result to complete the voice interaction.

[0070] The embodiment of the present application also provides a vehicle, which includes a memory and a processor. The voice interaction method of the embodiment of the present application can be implemented by the vehicle of the embodiment of the present application. Specifically, a computer program is stored in the memory, and the processor is used to obtain the current voice request, and to perform slot recognition on the current voice request, and to perform sentence processing on the current voice request according to the degree of correlation between any two semantic units in the current voice request, determine multiple target sentences, and perform application program interface prediction and application program interface parameter filling for each target sentence according to the results of multiple target sentences and slot recognition, obtain the target execution result of the application program interface parameter filling of each target sentence, and control the vehicle according to the target execution result to complete the voice interaction.

[0071] Specifically, in the implementation manner of the present application, when a user wants to control a vehicle to perform a specific function or action at the current moment and thus utters a corresponding statement to complete the triggering of the current voice request, the vehicle can obtain the current voice request and perform slot recognition on the current voice request to obtain the slot recognition result of the current voice request. Moreover, the vehicle can also split the current voice request into several sentences that can represent the complete intention, that is, multiple target clauses, based on the relevance between any two voice requests in the current voice request. Furthermore, the vehicle can perform separate API (Application Interface) prediction and API parameter filling on each target clause based on the multiple target clauses of the current voice request and the slot recognition result of the current voice request, thereby obtaining the target execution result of the API parameter filling for each target clause. Based on the target execution result of each target clause, the above-mentioned specific function or action is executed, thereby completing the voice interaction with the user.

[0072] For example, when a user in the cockpit space wants to turn on the seat ventilation and seat heating of the vehicle at the current moment and thus triggers a query of "turn on seat heating and ventilation" (i.e., the current voice request), the vehicle can perform slot recognition on this query, and the obtained slot recognition result slots may include: [{'name': 'device', 'pos': ['2, 3'], 'value':'seat'}, {'name': 'target_function', 'pos': ['4, 5'], 'value': 'heating'}, {'name': 'target_function', 'pos': ['7, 8'], 'value': 'ventilation'}].

[0073] Meanwhile, the vehicle can also split the query into two target clauses according to the relevance between any two semantic units in the above query. One is "turn on seat heating", and the other is "turn on seat ventilation". In one example, the specific forms of "turn on seat heating" and "turn on seat ventilation" output by the vehicle are: ['sent1': 'turn on seat heating','sent2': 'turn on seat ventilation'].

[0074] After obtaining the slot recognition result slots and the target clauses, for each target clause, the vehicle can perform API prediction, that is, predict the API used to implement the target clause. For example, it is predicted that the API corresponding to "turn on seat heating" is "SeatOpen", and it is also predicted that the API corresponding to "turn on seat ventilation" is "SeatOpen".

[0075] In addition, since corresponding parameters (i.e., the input parameters of the API) need to be used in the process of calling the API, the vehicle will also fill in the parameters for the APIs of the two target clauses. For example, fill in the parameters for "SeatOpen" corresponding to "Turn on seat heating", and the execution result of the parameter filling for "SeatOpen" corresponding to "Turn on seat heating" can be: {"apiname":"SeatOpen","arguments":[{'name':'device','value':'seat'},{'name':'target_function','value':'heating'}]}.

[0076] Correspondingly, the vehicle fills in the parameters for "SeatOpen" corresponding to "Turn on seat ventilation", and the execution result of the parameter filling for "SeatOpen" corresponding to "Turn on seat heating" can be: {"apiname":"SeatOpen","arguments":[{'name':'device','value':'seat'},{'name':'target_function','value':'ventilation'}]}.

[0077] Finally, the vehicle can perform corresponding control based on the execution results of the parameter filling of the APIs of the above two target clauses. For example, call SeatOpen twice through the above execution results to execute the two functions of "Turn on seat heating" and "Turn on seat ventilation".

[0078] Among them, it can be understood that the semantic unit in the embodiment of the present application can be understood as a token, which can be any one of words, characters, and phrases in a sentence, and is used to represent the basic unit constituting the sentence. For example, the semantic units of the above query may include: "da", "kai", "zuo", "yi", "jia", "re", "he", "tong", and "feng".

[0079] It can also be understood that the semantic units may vary in different language environments and natural language processing tasks. For example, in a Chinese-oriented translation task, the semantic unit can be understood as Chinese characters in a sentence or text. In an English-oriented translation task, the semantic unit can be understood as a word in a sentence or text.

[0080] It should also be noted that in a voice request, the degree of correlation between two semantic units can to a certain extent characterize "whether the two semantic units can be combined". For example, in the sentence "Turn on the seat heating and ventilation", the degree of correlation between 'turn' and 'on' may be higher than that between 'turn' and'ventilation'. Therefore, compared with 'turn' and'ventilation', the probability that 'turn' and 'on' form a pair of combinations is higher. Another example is that since 'turn' and 'on' are two adjacent semantic units in the sentence "Turn on the seat heating and ventilation", the probability that 'turn' and 'on' form a pair of combinations is higher compared with 'turn' and'seat'.

[0081] Based on this, the vehicle according to the embodiment of the present application can split and combine the current voice request based on the degree of correlation between any two semantic units in the current voice request, so as to determine, through the combination of each semantic unit in the current voice request, the target clauses that can be split from the current voice request and contain complete intentions.

[0082] In summary, in the embodiment of the present application, the current voice request can be segmented based on the degree of correlation between any two semantic units in the current voice request to obtain multiple target clauses, so that the segmentation of the current voice request can be realized. Furthermore, when the natural language processing difficulty of the current voice request is relatively high, the natural language processing of the current voice request can be converted into the natural language processing of multiple target clauses. Therefore, to a certain extent, the natural language processing difficulty of the voice request is reduced, thereby ensuring the accuracy of the natural language understanding result. Furthermore, based on the natural language processing result with guaranteed accuracy, the interaction with the user can be completed, so that the voice interaction experience and driving experience of the user can be guaranteed.

[0083] Moreover, compared with the above-mentioned natural language processing solution based on the pipeline, the embodiment of the present application can complete sentence splitting and realize natural language processing based on clauses through the degree of correlation between semantic units, thereby simplifying the architecture of the corresponding natural language understanding system. At the same time, since the number of modules in the system is smaller, the cooperation efficiency between modules can be higher, and the system is easier to expand and maintain. In addition, the service scheduling can be optimized, the situations of service timeout and failure are reduced, and the user experience is improved.

[0084] In addition, it can be understood that the voice request in the embodiment of the present application can be understood as a sentence or voice "for causing the vehicle to perform a certain action", such as sentences or voices including only a single intention like "Turn on the air conditioner", "Navigate to the restaurant", and "Close the window".

[0085] In addition, in the embodiments of the present application, the voice request can also be understood as the above-mentioned complex instruction statement, and it is difficult to divide this statement into clauses through voice activity detection. For example, when the user quickly says "Turn on the seat heating and ventilation", it is difficult to use voice activity detection to divide "Turn on the seat heating and ventilation" into 'Turn on the seat heating' and 'Turn on the seat ventilation'.

[0086] It can also be understood that the above two voice requests can both be applied and processed by the embodiments of the present application. Therefore, to avoid repetition, only the above-mentioned complex instruction statement will be used as an example for subsequent description, such as the above query.

[0087] In some embodiments of the present application, the method further includes:

[0088] Determine the correlation degree between any two semantic units in the current voice request according to the entity in the current voice request.

[0089] The voice interaction device according to the embodiment of the present application further includes a determination module. The determination module is used to determine the correlation degree between any two semantic units in the current voice request according to the entity in the current voice request.

[0090] The processor according to the embodiment of the present application is further used to determine the correlation degree between any two semantic units in the current voice request according to the entity in the current voice request.

[0091] Specifically, to ensure the reliable splitting of the current voice request and the credibility of the target clause, the vehicle according to the embodiment of the present application can determine the correlation degree between any two semantic units in the current voice request based on the named entity in the current voice request, so as to ensure that the splitting process performed based on this correlation degree can be reliably executed, and further ensure the credibility of the split target clause.

[0092] For example, taking the above query as an example, the entities of the query can include: "Turn on", "seat", "heating", and "ventilation". In addition, the semantic units of the query can include: "da", "kai", "zuo", "yi", "jia", "re", "he", "tong", and "feng".

[0093] Based on this, the vehicle according to the embodiment of the present application can determine the combinations that can be formed by each semantic unit or each entity according to the relationship between entities and the relationship between semantics, and thus obtain the target clause. For example, after arranging the semantic units of the above query in order, "‘hit’ and ‘open’" belong to the entity "open", "‘seat’ and ‘chair’" belong to the entity "seat", "‘heat’ and ‘up’" belong to the entity "heating", and "‘hot’" and "‘and’" do not belong to the same entity. Therefore, the combination / clause "open the seat heating" corresponding to "‘hit’", "‘open’", "‘seat’", "‘chair’", "‘heat’", and "‘up’" can be determined as a target clause.

[0094] In this way, in the embodiment of the present application, the vehicle can determine the degree of correlation between any two semantic units in the current voice request according to the entities in the current voice request, and to a certain extent, improve the credibility of the determined degree of correlation result.

[0095] In some embodiments of the present application, the step of determining the degree of correlation between any two semantic units in the current voice request according to the entities in the current voice request includes:

[0096] When two semantic units are adjacent in the current voice request and correspond to one entity, it is determined that the two semantic units are correlated to the first degree.

[0097] The determination module according to the embodiment of the present application is further configured to determine that two semantic units are correlated to the first degree when the two semantic units are adjacent in the current voice request and correspond to one entity.

[0098] The processor according to the embodiment of the present application is further configured to determine that two semantic units are correlated to the first degree when the two semantic units are adjacent in the current voice request and correspond to one entity.

[0099] Specifically, in the embodiment of the present application, when the vehicle can determine that two semantic units are adjacent in the current voice request and the two semantic units correspond to or belong to the same entity, the degree of correlation between the two semantic units is determined to be the first degree.

[0100] For example, taking the above query as an example, "‘hit’ and ‘open’" belong to the entity "open", "‘seat’ and ‘chair’" belong to the entity "seat", "‘heat’ and ‘up’" belong to the entity "heating", "‘ventilate’ and ‘air’" belong to the entity "ventilation". Therefore, there are a total of 4 semantic unit combinations, namely "‘hit’ and ‘open’", "‘seat’ and ‘chair’", "‘heat’ and ‘up’", and "‘ventilate’ and ‘air’", and the two semantic units in each combination are correlated to the first degree.

[0101] Thus, in the embodiments of the present application, when two semantic units in the current voice request are adjacent in the current voice request and correspond to the same entity, the vehicle may determine that the correlation degree of the two semantic units is the first degree.

[0102] In some embodiments of the present application, the step of determining the correlation degree of any two semantic units in the current voice request according to the entity in the current voice request includes:

[0103] When two entities corresponding to two semantic units within the current voice request are consecutive, it is determined that the two semantic units are correlated to the first degree.

[0104] The determination module in the embodiments of the present application is further configured to determine that two semantic units are correlated to the first degree when two entities corresponding to the two semantic units within the current voice request are consecutive.

[0105] The processor in the embodiments of the present application is further configured to determine that two semantic units are correlated to the first degree when two entities corresponding to the two semantic units within the current voice request are consecutive.

[0106] Specifically, the vehicle in the embodiments of the present application is further configured to finally confirm that the correlation degree of two semantic units is the first degree when it is determined that the two semantic units are adjacent in the current voice request, and it is confirmed that the two voice requests respectively correspond to two entities, and it is confirmed that the two entities are adjacent in the current voice request.

[0107] It should be noted that, in the embodiments of the present application, "adjacent" can be understood as adjacency in position. For example, in the above query, "‘hit’ and ‘open’", "‘open’ and ‘seat’", "‘seat’ and ‘chair’", "‘chair’ and ‘add’", "‘add’ and ‘heat’", "‘heat’ and ‘and’", "‘and’ and ‘turn’", and "‘turn’ and ‘ventilate’" are all adjacent semantic units.

[0108] It should also be noted that, in the embodiments of the present application, "consecutive" can be understood as semantic connection and smoothness. For example, in the above query, (open, seat), (seat, ventilate), (seat, heat) can be understood as satisfying the "consecutive" three entity pairs.

[0109] Based on this, taking the above query as an example, the combinations of two semantic units of the first degree may include: "‘open’ and ‘seat’", "‘chair’ and ‘ventilate’", "‘chair’ and ‘add’".

[0110] Thus, in the embodiments of the present application, when two semantic units in the current voice request are adjacent in the current voice request, and the corresponding two entities are also adjacent in the current voice request, the relevance degree of the two semantic units is determined to be the first degree.

[0111] In some embodiments of the present application, the two semantic units include a first unit and a second unit. Further, the step of determining the relevance degree of any two semantic units in the current voice request according to the entities in the current voice request includes:

[0112] When the two semantic units correspond to a first boundary entity and a second boundary entity in the current voice request, the first unit is at the beginning of the first boundary entity, and the second unit is at the end of the second boundary entity, it is determined that the two semantic units are relevant to the second degree.

[0113] In the embodiments of the present application, the determination module determines that the two semantic units are relevant to the second degree when the two semantic units correspond to a first boundary entity and a second boundary entity in the current voice request, the first unit is at the beginning of the first boundary entity, and the second unit is at the end of the second boundary entity.

[0114] In the embodiments of the present application, the processor determines that the two semantic units are relevant to the second degree when the two semantic units correspond to a first boundary entity and a second boundary entity in the current voice request, the first unit is at the beginning of the first boundary entity, and the second unit is at the end of the second boundary entity.

[0115] It should be noted that, in the embodiments of the present application, a boundary entity can be understood as an entity that only satisfies a continuous relationship with another entity.

[0116] For example, taking the above query as an example, (open, seat), (seat, ventilation), (seat, heating) can be understood as satisfying the "continuous" three entity pairs. Further, it can be known that "open", "ventilation" and "heating" are all only continuous with "seat", so it can be determined that "open", "ventilation" and "heating" are all boundary entities.

[0117] Furthermore, for the semantic unit pair (hit, wind), 'hit' belongs to "open" and is at the beginning, 'wind' belongs to "ventilation" and is at the end. Therefore, (hit, wind) can be determined to be relevant to the second degree and can be used to define the boundary of a certain clause.

[0118] Similarly, for the semantic unit pair (hit, heat), 'hit' belongs to "open" and is at the beginning, 'heat' belongs to "heating" and is at the end. Therefore, (hit, heat) can be determined to be relevant to the second degree and can be used to define the boundary of another clause.

[0119] Thus, in the embodiments of the present application, when the first unit in the current voice request is at the beginning of the first boundary entity and the second unit is at the end of the second boundary entity, it is confirmed that the degree of relevance between the first unit and the second unit is the second degree.

[0120] Please refer to Figure 2 , in some embodiments of the present application, step 03 includes:

[0121] 030: Construct a semantic unit relationship matrix according to the degree of relevance;

[0122] 031: Determine multiple target clauses according to the semantic unit relationship matrix.

[0123] The processing module in the embodiments of the present application is further configured to construct a semantic unit relationship matrix according to the degree of relevance, and to determine multiple target clauses according to the semantic unit relationship matrix.

[0124] The processor in the embodiments of the present application is further configured to construct a semantic unit relationship matrix according to the degree of relevance, and to determine multiple target clauses according to the semantic unit relationship matrix.

[0125] For a clearer illustration of the embodiments of the present application, please refer to Figure 3 , Figure 3 which is a schematic diagram of the semantic unit relationship matrix in some embodiments of the present application. It should be noted that in Figure 3 , the two semantic units corresponding to the row and column where R1 is located are relevant to the first degree. For example, since the semantic units 'hit' and 'open' are adjacent and belong to the same entity, the semantic unit pair (hit, open) satisfies R1. Another example is that since the entity corresponding to the semantic unit 'open', which is 'open', is continuous with the'seat' corresponding to'seat', the semantic unit pair (open, seat) satisfies R1.

[0126] And, in Figure 3 , the two semantic units corresponding to the row and column where R2 is located are relevant to the second degree. For example, 'hit' belongs to the boundary entity 'open' and is at the beginning of this entity, and 'wind' belongs to the boundary entity'ventilate' and is at the end of this entity. Therefore, it can be considered that the semantic unit pair (hit, wind) satisfies R2. Another example is that 'hit' belongs to 'open' and is at the beginning, and 'hot' belongs to 'heat' and is at the end. Therefore, it can be considered that the semantic unit pair (hit, hot) satisfies R2. It can be understood that the semantic unit pairs that satisfy R1 only appear in the upper triangular position in the semantic unit relationship matrix.

[0127] Furthermore, based on the semantic unit relationship matrix, the embodiments of the present application can perform decoding processing on the semantic unit relationship matrix to obtain target clauses.

[0128] For example, in some embodiments of the present application, a vehicle may construct a group of consecutive semantic units according to the semantic units satisfying R1, and then determine the boundary of the group of semantic units through the semantic units satisfying R2. Specifically, please refer to Figure 4 , Figure 5 and Figure 6 , Figure 4 , Figure 5 and Figure 6 are all schematic diagrams of application scenarios in some embodiments of the present application, and please refer to Figure 3 again. That is, as shown in Figure 3 , Figure 4 and Figure 5 , the vehicle can also construct a group of semantic units through (turn on, ventilation) satisfying R2, and (turn on, open), (open, seat), (seat, chair), (chair, ventilation), and (ventilation, on) satisfying R1, so as to obtain the target clause "turn on seat ventilation".

[0129] Similarly, as shown in Figure 3 , Figure 4 and Figure 6 , the vehicle can construct a group of semantic units through (turn on, heat) satisfying R2, and (turn on, open), (open, seat), (seat, chair), (chair, heating), and (heating, on) satisfying R1, so as to obtain the target clause "turn on seat heating".

[0130] In this way, in the embodiments of the present application, the vehicle can construct a semantic unit relationship matrix based on the correlation degree between any two semantic units in the current voice request and use the semantic unit relationship matrix to determine the target clause, so that the credibility of the target clause can be guaranteed to a certain extent.

[0131] Please refer to Figure 7 , in some embodiments of the present application, step 04 includes:

[0132] 040: Determine the slot recognition sub-results of each target clause according to the results of multiple target clauses and slot recognition;

[0133] 041: Perform application programming interface prediction and application programming interface parameter filling on each target clause according to each target clause and the slot recognition sub-results of each target clause, and obtain the target execution results of each target clause.

[0134] The filling module in the embodiments of the present application is further configured to determine the slot recognition sub-results of each target clause according to the results of multiple target clauses and slot recognition, and to perform application programming interface prediction and application programming interface parameter filling on each target clause according to each target clause and the slot recognition sub-results of each target clause, so as to obtain the target execution results of each target clause.

[0135] The processor according to the embodiment of the present application is further configured to determine the slot recognition sub-results of each target clause according to the results of the recognition of multiple target clauses and slots, and to perform application programming interface prediction and application programming interface parameter filling on each target clause according to each target clause and the slot recognition sub-results of each target clause, so as to obtain the target execution results of each target clause.

[0136] Specifically, to ensure the reliable understanding and processing of each target clause, the vehicle according to the embodiment of the present application may split the complete slot recognition result of the current voice request into slot recognition sub-results corresponding to each target clause respectively, and then complete the corresponding API prediction and API parameter filling through the target clause and the slot recognition sub-result corresponding to the target clause.

[0137] Exemplarily, in some embodiments of the present application, taking the above query, slots and two target clauses corresponding to the query ("turn on the seat heating" and "turn on the seat ventilation") as an example, that is, letting the target clause "turn on the seat heating" be sent1, and letting the target clause "turn on the seat ventilation" be sent2, then: the slot recognition sub-result slot1 determined based on slots and sent1 may be: {'name': 'device', 'pos': ['2, 3'], 'value':'seat'}, {'name': 'target_function', 'pos': ['4, 5'], 'value': 'heating'}.

[0138] Similarly, the slot recognition sub-result slot2 determined based on slots and sent2 may be: {'name': 'device', 'pos': ['2, 3'], 'value':'seat'}, {'name': 'target_function', 'pos': ['7, 8'], 'value': 'ventilation'}.

[0139] Optionally, in some other embodiments of the present application, the form of the above slot1 may also be: {'name': 'device', 'value':'seat'}, {'name': 'target_function', 'value': 'heating'}.

[0140] Correspondingly, in some other embodiments of the present application, the form of the above slot2 may also be: {'name': 'device', 'value':'seat'}, {'name': 'target_function', "value":'ventilation'}

[0141] It can also be understood that after the target clause and the slot recognition sub-results of the target clause are determined, the vehicle can predict the API that can implement the target clause based on the target clause and the slot recognition sub-results of the target clause, and fill in the parameters of the predicted API.

[0142] For example, please refer to Figure 8 , Figure 8 which is a schematic flowchart of the voice interaction method in some embodiments of the present application. Based on the above-mentioned sent1 and slot1, the obtained API can be "SeatOpen". Then, by using sent1 and slot1 to fill in the parameters of "SeatOpen", the execution result of the parameter filling can include: {"apiname":"SeatOpen","arguments":[{'name':'device','value':'seat'},{'name':'target_function','value':'heating'}]}.

[0143] Similarly, based on the above-mentioned sent2 and slot2, the obtained API can be "SeatOpen". Then, by using sent2 and slot2 to fill in the parameters of "SeatOpen", the execution result of the parameter filling can include: {"apiname":"SeatOpen","arguments":[{'name':'device','value':'seat'},{'name':'target_function','value':'ventilation'}].

[0144] In this way, in the embodiments of the present application, the vehicle can perform application programming interface prediction and parameter filling of the application programming interface for each target clause through the target clause and the slot recognition sub-results of the target clause, which to a certain extent ensures the reliable execution of the application programming interface prediction and the parameter filling of the application programming interface.

[0145] In some embodiments of the present application, step 02 includes:

[0146] Based on the pre-trained recognition model, perform sequence labeling processing on the current voice request to complete slot recognition.

[0147] The recognition module in the embodiments of the present application is further configured to perform sequence labeling processing on the current voice request based on the pre-trained recognition model to complete slot recognition.

[0148] The processor in the embodiments of the present application is further configured to perform sequence labeling processing on the current voice request based on the pre-trained recognition model to complete slot recognition.

[0149] Specifically, to ensure the accurate slot recognition, the vehicle according to the embodiments of the present application can also call a pre-trained recognition model to complete the slot recognition of the current voice request. Or, the current voice request is input into the recognition model for the recognition model to perform sequence tagging on the current voice request, so as to identify, label, and extract named entities of each type in the current voice request.

[0150] In this way, in the embodiments of the present application, the vehicle can perform sequence tagging processing on the current voice request through a pre-trained generative model, so as to obtain the slot recognition result of the current voice request, and further, the credibility of the slot recognition result can be guaranteed to a certain extent.

[0151] Optionally, please refer to Figure 9 and Figure 10 , Figure 9 and Figure 10 are all flow diagrams of the voice interaction method in some embodiments of the present application. Specifically, as Figure 9 shown, in the vehicle according to the embodiments of the present application, when obtaining the current voice request (i.e., the above-mentioned query) triggered by the user, the query can be input into a pre-set joint named entity recognition module to perform slot recognition and sentence splitting through this model.

[0152] Further, as Figure 10 shown, the vehicle can encode the query at the semantic unit level based on the joint named entity recognition module to obtain the encoded representation of each semantic unit of the query.

[0153] Subsequently, the encoded semantic units can be input into a pre-trained label classifier (such as the above-mentioned recognition model), thereby performing decoding of the query, or performing slot recognition / sequence tagging of the query, and further obtaining the slot recognition result of the query, that is, the above-mentioned slots.

[0154] Meanwhile, the vehicle can also input the encoded semantic units into a pre-trained relationship classifier to predict the correlation degree between any two semantic units in the current voice request, that is, to predict that any two semantic units in the current voice request belong to any one of R0, R1, and R2, where R0 represents "unrelated", R1 represents the above-mentioned first correlation degree, and R2 represents the above-mentioned second correlation degree. Finally, according to the prediction result of the correlation degree between any two semantic units, a semantic unit relationship matrix as Figure 3 shown is formed.

[0155] Next, decode the semantic unit relationship matrix to obtain the clause splitting result of the query, namely, "Turn on the seat heating" and "Turn on the seat ventilation". It can be understood that Figure 9 the sub_index in

[0156] can be understood as the serial number or identifier of the target clause. Subsequently, the vehicle can splice each target clause of the query and the corresponding part in the slot recognition result slots to form the output result of the joint named entity recognition module, that is: [{'sub_index':'1', 'query': 'Turn on the seat heating','slots': [{'name': 'device', 'value':'seat'}, {'name': 'target_function', 'value': 'heating'}], {{'sub_index': '2', 'query': 'Turn on the seat ventilation','slots': [{'name': 'device', 'value':'seat'}, {'name': 'target_function', 'value': 'ventilation'}]}.

[0157] Then, the vehicle can perform API prediction and API parameter filling based on the above "output result of the joint named entity recognition module" to obtain the filling result of the API parameters, that is: {"apiname": "SeatOpen", "arguments": [{"name": "device", "value": "seat"}, {"name": "target_function", "value": "ventilation"}], "apiname": "SeatOpen", "arguments": [{"name": "device", "value": "seat"}, {"name": "target_function", "value": "ventilation"}]}.

[0158] Finally, the vehicle can parse the filling result to generate corresponding control instructions and execute the actions corresponding to the query through the control instructions, that is, turn on the ventilation function of the seat and turn on the heating function of the seat, thereby realizing the voice interaction feedback of the user.

[0159] Optionally, in some embodiments of the present application, the vehicle can input the "output result of the joint named entity recognition module" into a pre-trained generative model (Generative Models) to perform API prediction and API parameter filling through the generative model.

[0160] It should be noted that there are some syntax errors and inconsistent expressions in the original text, such as the incorrect use of "{{" in "[{{'sub_index':'2'...}}]" in ID=5 and the repeated "apiname" in ID=8. The translation is based on the original text as accurately as possible while maintaining the original format.Optionally, in some embodiments of the present application, the process of obtaining the generative model may include: based on the pre-trained BERT (Bidirectional Encoder Representations from Transformers) model, performing fine-tuning training on BERT to make BERT applicable to the downstream task processing capabilities required by the embodiments of the present application, and then the recognition model in the embodiments of the present application can be obtained.

[0161] Optionally, in some embodiments of the present application, when the vehicle is performing the function indicated by the current voice request, it may play a corresponding voice prompt or display a corresponding text prompt message for the user to know the function performed by the vehicle through the current voice request.

[0162] In addition, it should be noted that the voice interaction method provided by the embodiments of the present application can be applied not only to vehicles but also to servers communicatively connected to the vehicles. It can be understood that when the voice interaction method provided by the embodiments of the present application is applied to the server, the server can perform steps 02 to 04 based on the current voice request forwarded by the vehicle. Then, when the execution result of the parameter filling of the corresponding application is obtained, the execution result can be sent to the vehicle for the vehicle to perform the corresponding function, or a corresponding vehicle control instruction can be generated according to the execution result and then sent to the vehicle to control the vehicle to perform the corresponding function.

[0163] The present application also provides a computer-readable storage medium storing a computer program, which, when executed by one or more processors, implements the above voice interaction method.

[0164] In the description of this specification, the descriptions referring to terms such as "specifically", "further", "specially", "understandably", etc. mean that the specific features, structures, materials or characteristics described in connection with the embodiments or examples are included in at least one embodiment or example of the present application. In this specification, the schematic expressions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in a suitable manner in any one or more embodiments or examples. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.

[0165] Any process or method description, whether in a flowchart or otherwise described herein, can be understood to represent a module, segment, or portion of code including one or more executable instructions for implementing a specific logical function or process. The scope of the preferred embodiments of the present application includes additional implementations where functions may be executed not in the order shown or discussed, including in a substantially simultaneous manner according to the functions involved or in the reverse order, which should be understood by those skilled in the art to which the embodiments of the present application pertain.

[0166] Although the embodiments of the present application have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present application. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present application.

Claims

1. A voice interaction method for a vehicle, characterized in that: include: Get the current voice request; Performing slot identification on the current voice request; Determining, according to the entities in the current voice request, a correlation degree between any two semantic units in the current voice request, wherein the two semantic units include a first unit and a second unit; According to the correlation degree between any two semantic units in the current voice request, the current voice request is processed by sentence segmentation to determine a plurality of target sentences; According to the results of the multiple target clauses and the slot identification, perform application program interface prediction and application program interface parameter filling on each target clause to obtain a target execution result of the application program interface parameter filling of each target clause; Controlling the vehicle to complete voice interaction according to the target execution result; The determining, according to the entity in the current voice request, the degree of correlation between any two semantic units in the current voice request includes: When the two semantic units are adjacent in the current voice request and correspond to one entity, determining that the two semantic units are related to a first degree; In a case where two entities corresponding to the two semantic units in the current voice request are continuous, determining that the two semantic units are related to a first degree; When the two semantic units correspond to the first boundary entity and the second boundary entity in the current voice request, the first unit is located at the first position of the first boundary entity, and the second unit is located at the last position of the second boundary entity, it is determined that the two semantic units are related to the second degree.

2. The method according to claim 1, characterized in that The sentence processing of the current voice request according to the correlation between any two semantic units in the current voice request to determine a plurality of target sentences includes: Constructing a semantic unit relationship matrix according to the correlation degree; The plurality of target sentences are determined according to the semantic unit relationship matrix.

3. The method according to claim 1, characterized in that The step of performing application program interface prediction and application program interface parameter filling on each target sentence according to the results of the multiple target sentences and the slot identification, and obtaining a target execution result of the application program interface parameter filling of each target sentence, includes: Determine a slot identification sub-result of each target sentence according to the plurality of target sentences and the slot identification results; According to each of the target clauses and the slot identification sub-result of each of the target clauses, the application program interface prediction and the application program interface parameter filling are performed on each of the target clauses to obtain the target execution result of each of the target clauses.

4. The method according to claim 1, characterized in that: The performing slot identification on the current voice request includes: Based on the pre-trained recognition model, sequence labeling is performed on the current voice request to complete the slot recognition.

5. A vehicle, characterized in that: The method comprises a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the method according to any one of claims 1 to 4 is implemented.

6. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by one or more processors, the method according to any one of claims 1 to 4 is implemented.

Citation Information

Patent Citations

  • Multi-intention recognition method and device based on robot

    CN114266240A

  • Multi-intention statement segmentation method and device, storage medium and electronic device

    CN115482378A

  • Voice interaction method and device, server and computer readable storage medium

    CN116665667A