Multi-intention recognition method and device, electronic equipment, storage medium and vehicle
By obtaining voice input information and applying target intent rule information for text extraction and intent recognition, and combining arbitration strategies to process intent recognition results, the problems of low accuracy and low efficiency of multi-intention recognition in the prior art are solved, and more efficient multi-intention recognition is achieved.
Patent Information
- Application Number
- CN202411871422.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-18
- Publication Date
- 2025-05-23
AI Technical Summary
In the prior art, multi-intention recognition has low accuracy and low efficiency, especially when the voice spoken by a user contains multiple intent instructions, it is difficult to recognize or has low recognition accuracy.
By obtaining voice input information, determining the target intention rule information matching the voice input information, text extraction of the voice input information based on the target intention rule information, obtaining text content information and text position information, and intent recognition is performed based on the text content information and text position information, obtaining at least two intent recognition results. Then, based on the target arbitration strategy corresponding to the target intention rule information, each intent recognition result is carried out to obtain the target intention information corresponding to the voice input information.
It effectively improves the accuracy and efficiency of multi-intention recognition, and can more accurately identify and process voice inputs containing multiple intent instructions.
Smart Images

Figure CN120031048A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of automobile technology, and in particular to a multi-intention recognition method, device, electronic equipment, storage medium and vehicle. Background Art
[0002] With the rapid development of intelligent connected vehicle technology, the in-vehicle dialogue system has become one of the key technologies to improve driving experience and enhance the level of vehicle intelligence. At present, the in-vehicle dialogue system can recognize and execute the user's voice commands, such as adjusting the temperature, playing music, navigation, etc., which greatly facilitates the driver's operation. In actual applications, it is usually determined by comparing the user's input voice with the preset voice command library, where the commands in the preset voice command library are only common basic commands, such as turning on the air conditioner, opening the car window, adjusting the air conditioner temperature, etc. When the user's voice contains multiple intention commands, it will be difficult to recognize or the recognition accuracy will be low, that is, there is a problem of low accuracy and low efficiency in multi-intention recognition. Summary of the invention
[0003] The present application provides a multi-intent recognition method, device, electronic device, storage medium and vehicle to solve the problems of low accuracy and low efficiency of multi-intent recognition in the existing related technologies.
[0004] In a first aspect, the present application provides a multi-intent recognition method, comprising:
[0005] Get voice input information;
[0006] Determining target intent rule information matching the speech input information;
[0007] Based on the target intention rule information, performing text extraction on the voice input information to obtain text content information and text position information;
[0008] Performing intent recognition based on the text content information and the text position information to obtain at least two intent recognition results;
[0009] Based on the target arbitration strategy corresponding to the target intent rule information, intent arbitration is performed on each of the intent recognition results to obtain the target intent information corresponding to the speech input information.
[0010] Optionally, the voice input information is voice information collected by an in-vehicle device, and the determining of target intention rule information matching the voice input information includes:
[0011] Obtaining a multi-intention rule set preset for the vehicle based on the voice input information;
[0012] Performing text recognition on the voice input information to obtain voice text information;
[0013] Matching the voice text information with the preset intent rule information in the multi-intent rule set to obtain a first matching result;
[0014] In the case where the first matching result is a rule matching result, the preset intention rule information is used to determine the target intention rule information.
[0015] Optionally, matching the voice text information with preset intent rule information in the multi-intent rule set to obtain a first matching result includes:
[0016] Extracting action keywords and execution object keywords from the preset intention rule information;
[0017] Matching the voice text information with the action keyword to obtain a second matching result;
[0018] Matching the voice text information with the execution object keyword to obtain a third matching result;
[0019] When the second matching result is an action keyword mismatch result, and / or the third matching result is an object keyword mismatch result, the rule mismatch result is determined as the first matching result;
[0020] In the case where the second matching result is an action keyword matching result and the third matching result is an object keyword matching result, the rule matching result is determined as the first matching result.
[0021] Optionally, performing intent recognition based on the text content information and the text position information to obtain at least two intent recognition results includes:
[0022] extracting target sentence structure information from the target intention rule information;
[0023] Based on the target sentence structure information, the sentence is rewritten in combination with the text content information and the text position information to obtain target sentence text information;
[0024] The target sentence text information is determined as the intention recognition result.
[0025] Optionally, the target arbitration strategy corresponding to the target intent rule information is based on which the intent arbitration is performed on each of the intent recognition results to obtain the target intent information corresponding to the speech input information, including:
[0026] Based on the target arbitration strategy, extracting target sentence text information from each of the intent recognition results;
[0027] Determine the starting offset information and the ending offset information corresponding to each target sentence text information;
[0028] According to the start offset information and the end offset information, performing intention arbitration on each target sentence text information to obtain arbitration text information;
[0029] Based on the arbitration text information, the target intention information is generated.
[0030] Optionally, performing intention arbitration on each target sentence text information according to the start offset information and the end offset information to obtain arbitration text information includes:
[0031] Determining the intention boundary information corresponding to each of the target sentence text information according to the starting offset information and the ending offset information;
[0032] Comparing the intention boundary information between each two target sentence text information to obtain an intention boundary comparison result;
[0033] Based on the intention boundary comparison result, intention arbitration is performed on each target sentence text information to obtain the arbitration text information.
[0034] Optionally, performing intention arbitration on each target sentence text information based on the intention boundary comparison result to obtain the arbitration text information includes:
[0035] In the case where the intention boundary comparison result is a boundary overlap result, merging the two target sentence text information corresponding to the boundary overlap result to obtain the arbitration text information;
[0036] When the intention boundary comparison result is a boundary start overlap result or a boundary end overlap result, the arbitration text information corresponding to each target sentence text information is generated.
[0037] Optionally, performing intention arbitration on each target sentence text information based on the intention boundary comparison result to obtain the arbitration text information includes:
[0038] When the intention boundary comparison result is a boundary inclusion result, comparing the boundary ranges of the two target sentence text information corresponding to the boundary inclusion result according to the start offset information and the end offset information to obtain a boundary range comparison result;
[0039] The arbitration text information is determined according to the boundary range comparison result.
[0040] Optionally, after performing intent arbitration on each of the intent recognition results based on the target arbitration strategy corresponding to the target intent rule information to obtain the target intent information corresponding to the speech input information, the method further includes:
[0041] Performing semantic analysis on the target intention information to obtain a voice execution instruction;
[0042] The vehicle is controlled to execute the command operation corresponding to the voice execution command, and a voice control result corresponding to the voice input information is obtained.
[0043] In a second aspect, the present application provides a multi-intent recognition method, comprising:
[0044] An acquisition module, used to acquire the voice input information of the vehicle;
[0045] A determination module, configured to determine target intent rule information matching the voice input information according to a preset multi-intention rule set corresponding to the vehicle;
[0046] An extraction module, used to perform text extraction on the voice input information based on the target intention rule information to obtain text content information and text position information;
[0047] An intention recognition module, used to perform intention recognition based on the text content information and the text position information to obtain at least two intention recognition results;
[0048] An arbitration module is used to perform intent arbitration on each of the intent recognition results based on a target arbitration strategy corresponding to the target intent rule information, so as to obtain target intent information corresponding to the speech input information.
[0049] In a third aspect, an electronic device is provided, comprising a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory communicate with each other via the communication bus;
[0050] Memory, used to store computer programs;
[0051] The processor is used to implement the multi-intention recognition method described in any one of the first aspects when executing the program stored in the memory.
[0052] In a fourth aspect, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the multi-intention recognition method as described in any one of the first aspects is implemented.
[0053] In a fifth aspect, a vehicle is provided, the vehicle comprising the multi-intention recognition device described in the fourth aspect.
[0054] The embodiment of the present application obtains voice input information to determine target intent rule information that matches the voice input information, performs text extraction on the voice input information based on the target intent rule information to obtain text content information and text position information, and performs intent recognition based on the text content information and text position information to obtain at least two intent recognition results, and then performs intent arbitration on each intent recognition result based on a target arbitration strategy corresponding to the target intent rule information to obtain target intent information corresponding to the voice input information; thereby solving the problems of low accuracy and low efficiency in multi-intent recognition in the existing related technologies and being able to effectively improve the accuracy and efficiency of multi-intent recognition. BRIEF DESCRIPTION OF THE DRAWINGS
[0055] Figure 1 A flowchart of a multi-intent recognition method provided in an embodiment of the present application;
[0056] Figure 2 A schematic diagram of a scenario of a multi-intent recognition method provided in an embodiment of the present application;
[0057] Figure 3 A schematic diagram of another scenario of a multi-intent recognition method provided in an embodiment of the present application;
[0058] Figure 4 A schematic diagram of another scenario of a multi-intent recognition method provided in an embodiment of the present application;
[0059] Figure 5 A schematic diagram of the structure of a multi-intention recognition device provided in an embodiment of the present application;
[0060] Figure 6 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0061] The following will describe the embodiments of the present invention with reference to the accompanying drawings and preferred embodiments. Those skilled in the art can easily understand other advantages and effects of the present invention from the contents disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and the details in this specification can also be modified or changed in various ways based on different viewpoints and applications without departing from the spirit of the present invention. It should be understood that the preferred embodiments are only for illustrating the present invention, not for limiting the scope of protection of the present invention.
[0062] In the existing related technologies, in order to recognize and process multi-intention speech, syntactic analysis or training model analysis is usually used. Among them, the method of syntactic analysis is to analyze the subject-predicate-object structure and recognize clauses of the sentence. A subject-predicate-object structure is considered to be an intention, and clauses such as attributive clauses and adverbial clauses are used as additional information to supplement the intention. However, in real life, the speech input by the user is seriously colloquial, and the syntax is often incomplete, or it does not follow the subject-predicate-object structure, resulting in the problem of low accuracy of the syntactic analysis method. Through the method of training model analysis, a training model is usually used to perform multi-intention recognition on the input speech, and multi-intention recognition often requires text extraction and semantic recognition of more complex speech. At this time, not only does the training model need to have a large number of parameters, but also a large number of training samples and training time. That is, the use of training models for multi-intention recognition will have the problem of high difficulty and low efficiency.
[0063] In order to solve the problems of low accuracy and low efficiency of multi-intent recognition in the existing related technologies, the present application provides a multi-intent recognition method, device, electronic device, storage medium and vehicle, which obtains voice input information to determine target intent rule information that matches the voice input information, and performs text extraction on the voice input information based on the target intent rule information to obtain text content information and text position information, and performs intent recognition based on the text content information and text position information to obtain at least two intent recognition results, and then performs intent arbitration on each intent recognition result based on the target arbitration strategy corresponding to the target intent rule information to obtain the target intent information corresponding to the voice input information; thereby solving the problems of low accuracy and low efficiency of multi-intent recognition in the existing related technologies, and can effectively improve the accuracy and efficiency of multi-intent recognition.
[0064] Figure 1 A flow chart of a multi-intent recognition method provided for an embodiment of the present application. This method can be applied to one or more electronic devices such as vehicles, clients, and servers. In addition, the execution subject of this method can be hardware or software. When the above-mentioned execution subject is hardware, the execution subject can be one or more of the above-mentioned electronic devices. For example, a single electronic device can execute this method, or multiple electronic devices can cooperate with each other to execute this method. When the above-mentioned execution subject is software, this method can be implemented as multiple software or software modules, or as a single software or software module. No specific limitation is made here.
[0065] like Figure 1 As shown, a multi-intent recognition method provided in an embodiment of the present application may specifically include the following steps:
[0066] Step S110: Acquire voice input information;
[0067] Among them, the voice input information can represent the voice used for voice control, such as the voice in the vehicle cockpit, the voice information sent by the server, etc., and the embodiment of the present application does not impose specific restrictions on this.
[0068] In addition, before obtaining the voice input information, it can also include determining whether a voice collection start instruction is received. The voice collection start instruction represents an instruction for starting voice collection. The voice collection instruction can be a preset start voice, key signal or other instruction. After receiving the voice collection start instruction, the voice input information can be obtained. Of course, the above is only an example for illustration, and this embodiment does not make any specific limitations on this.
[0069] Step S120: Determine target intention rule information that matches the voice input information.
[0070] Specifically, the present implementation may pre-configure multiple preset intent rule information, and the preset intent rule information may represent pre-configured rules for intent recognition and extraction. After obtaining voice input information, the target intent rule information matching the voice input information may be determined, wherein the target intent rule information may be preset intent rule information matching the voice input information.
[0071] Specifically, the target intent rule information may include keyword extraction rules and rewriting rules, wherein the keyword extraction rules represent rules for extracting keywords from voice input information, such as the extracted keyword type, keyword text, etc., and the rewriting rules may represent rules for rewriting the extracted keywords; for example, the target intent rule information may be:<rule_def> :=<sentence_pattern> '->'<rewrite_sentence> , which means that the rule can be defined as<sentence_pattern> That is, keyword extraction rules and
[0072] <rewrite_sentence> That is, the rewriting rule consists of two parts, separated by '->', indicating the causal relationship between the previous and the next.<sentence_pattern> The specific rules used to match and extract the text corresponding to the voice input information can be similar to regular expressions;<rewrite_sentence> The former is a rule that is rewritten after the keyword is successfully extracted. The specific example is as follows:
[0073] The target intention rule information is (?P <open>(open|on)).*(?P <device>(air conditioning|windows))->
[0074] In the case of open&device, the keyword extraction rule in the first half of the rule can be simplified to (open|turn on).*(air conditioning|window), which means recognizing the voice input information of 'xxx open / turn on xxx window / air conditioning xxx', where 'xxx' represents any character or an empty character, such as recognizing the voice input information of "open window", "turn on air conditioning", "open my window", "please turn on the air conditioning", etc.; and (?P <open>(open|open)) can indicate that when "open" or "open" appears in the voice input information, it will be captured as a keyword. The identifier name after the keyword is captured is 'open'.<rewrite_sentence> Rules can reference captured values according to open, (?P <device>(air conditioning|car window)) is the same. The rewrite rule 'open&device' means rewriting the voice input information into a concise intent, where open and device refer to the keyword content captured by the keyword extraction rules mentioned above, such as "open the car window". Through the above rules, it is possible to extract the intents in multiple intents that contain the above rule patterns and rewrite them into a single intent text with simple and clear intents. Of course, the above is only an example for explanation. In specific implementations, multiple target intent rule information can be configured. For example, in scenarios where inverted sentences frequently appear, (?P <device>(air conditioning|windows))->(?P <open>(open|enable)).*open&device's target intent rule information is not specifically limited in this embodiment.
[0075] Step S130: Based on the target intention rule information, perform text extraction on the voice input information to obtain text content information and text location information.
[0076] Specifically, after determining the target intent rule information that matches the voice input information, the target intent rule information can be used to perform text extraction on the voice input information to obtain text content information and text position information, wherein the text content information represents the text content obtained by performing text extraction on the voice input information based on the target intent rule information, and the text position information can represent the appearance position of each keyword in the text content information in the voice input information.
[0077] Step S140: performing intent recognition based on the text content information and the text position information to obtain at least two intent recognition results.
[0078] Specifically, after obtaining the text content information and text position information, the text content information can be subjected to intent recognition based on the target intent rule information to obtain at least two intent recognition results, wherein intent recognition can be a means of rewriting or extracting the text content based on the text content information and the text position information, and the intent recognition result can represent the result used to directly represent a single intent in the voice input information.
[0079] In one example, when the voice input information is "turn on that air conditioner, and the car windows", and the text content information is turn on, air conditioner, and car windows, the text position information is that turn on appears at the beginning of the voice input information, air conditioner appears in the middle of the voice input information, and car windows appears at the end of the voice input information. At this time, the text content information is rewritten according to the target intention rule information, and at least two intention recognition results are obtained, namely "turn on the air conditioner" and "open the car windows", that is, each intention recognition result can directly represent a single intention in the voice input information; of course, the basis is only for illustrative purposes, and this embodiment does not make specific limitations on this.
[0080] Step S150: Based on the target arbitration strategy corresponding to the target intent rule information, perform intent arbitration on each intent recognition result to obtain the target intent information corresponding to the speech input information.
[0081] Specifically, after obtaining the intent recognition result, the target arbitration strategy corresponding to the target intent rule information can be determined. The target arbitration strategy represents an arbitration strategy pre-configured to match the target intent rule information, and the target arbitration strategy can specifically perform intent arbitration on each intent recognition result corresponding to the target intent rule information. Intent arbitration can include but is not limited to deleting repeated intents, retaining intents with the same beginning but different endings, retaining intents with different beginnings but the same endings, and merging intents with inclusion relationships; of course, different target intent rule information can correspond to different target arbitration strategies, so as to play the role of intent arbitration for the intent recognition results corresponding to different target intent rule information using its adaptive target arbitration strategy; and the target intent information can represent the intent recognition result retained after intent arbitration, so that the target intent information can represent the specific intent corresponding to the voice input information. Of course, the target intent information can contain one or more intents, and this embodiment does not specifically limit this.
[0082] In one example, when the intent recognition results include turning on the air conditioner, opening the car window, and turning on the air conditioner, and the target arbitration strategy includes deleting repeated intents, each intent recognition result is subjected to intent arbitration based on the target arbitration strategy corresponding to the target intent rule information, and repeated turning on the air conditioner in the intent recognition results can be deleted, so that the target intent information corresponding to the voice input information is turning on the air conditioner and opening the car window. Of course, the above is only an example for illustration, and this embodiment does not make specific limitations on this.
[0083] It can be seen that this embodiment obtains voice input information to determine the target intent rule information that matches the voice input information, and based on the target intent rule information, performs text extraction on the voice input information to obtain text content information and text position information, and performs intent recognition based on the text content information and text position information to obtain at least two intent recognition results, which plays a role in determining the various intentions contained in the voice input information, and then performs intent arbitration on each intent recognition result based on the target arbitration strategy corresponding to the target intent rule information to obtain the target intent information corresponding to the voice input information, which plays a role in screening and arbitrating the various intentions in the intent recognition results to improve the accuracy of the target intent information in expressing the intent corresponding to the voice input information; thereby solving the problems of low accuracy and low efficiency in multi-intent recognition in the existing related technologies, and can effectively improve the accuracy and efficiency of multi-intent recognition.
[0084] In an optional embodiment of the present application, the voice input information in this embodiment may be voice information collected by an on-board device, wherein the on-board device may be a microphone arranged in a cockpit of the vehicle, and step S120 determines the target intention rule information that matches the voice input information, which may specifically include the following steps: obtaining a multi-intention rule set preset for the vehicle for the voice input information; performing text recognition on the voice input information to obtain voice text information; matching the voice text information with the preset intention rule information in the multi-intention rule set to obtain a first matching result; and when the first matching result is a rule matching result, determining the target intention rule information from the preset intention rule information.
[0085] After acquiring the voice input information, the present embodiment can acquire a vehicle preset multi-intention rule set for the voice input information, wherein the vehicle can represent the vehicle configured by the vehicle-mounted device corresponding to the collected voice input information, and the multi-intention rule set can represent a set of multiple preset intention rule information pre-configured for the vehicle, and the preset intention rule information can represent a pre-configured rule for intent extraction; then, text recognition can be performed on the voice input information to obtain voice text information, and the voice text information can represent the text corresponding to the voice input information, and the specific method can be automatic speech recognition (Automatic Speech Recognition). Recognition, ASR) technology, neural network, etc., which are not specifically limited in this embodiment; thereby, the voice text information can be matched with the preset intent rule information in the multi-intent rule set to obtain a first matching result, and the first matching result can be a rule mismatch result or a rule matching result, wherein the rule mismatch result indicates that the voice text information and the preset intent rule information do not have a matching relationship, and the rule matching result indicates that the voice text information and the preset intent rule information have a matching relationship; then, it can be determined whether the first matching result is a rule matching result. When the first matching result is a rule matching result, the preset intent rule information can be determined as the target intent rule information. Of course, one voice text information can correspond to one or more target intent rule information, which are not specifically limited in this embodiment.
[0086] It should be noted that if the preset intent rule information corresponding to the multi-intention rule set are all rule mismatch results, it means that the preset multi-intention rule set is unable to perform intent recognition and analysis on the voice input information. At this time, the voice input information can be output to other preset intent recognition models, and other intent recognition models can be used to perform intent recognition on the voice input information to obtain the target intent information corresponding to the voice input information.
[0087] In one example, the vehicle preset multi-intention rule set including preset intention rule information A, preset intention rule information B, and preset intention rule information C is obtained, and the voice text information is matched with the preset intention rule information A, preset intention rule information B, and preset intention rule information C respectively to obtain the first matching result corresponding to each preset intention rule information, that is, including the first matching result A corresponding to the preset intention rule information A, the first matching result B corresponding to the preset intention rule information B, and the first matching result C corresponding to the preset intention rule information C. At this time, it can be determined whether each first matching result is a rule matching result. For example, the first matching result A and the first matching result C are rule matching results, and the first matching result B is a rule non-matching result. At this time, the preset intention rule information A and the preset intention rule information C can be determined as the target intention rule information. Of course, the above is only an example for illustration, and this embodiment does not make specific limitations on this.
[0088] In an optional embodiment of the present application, the voice text information is matched with the preset intent rule information in the multi-intent rule set to obtain a first matching result, which may specifically include the following steps: extracting action keywords and execution object keywords from the preset intent rule information; matching the voice text information with the action keywords to obtain a second matching result; matching the voice text information with the execution object keywords to obtain a third matching result; when the second matching result is an action keyword mismatch result, and / or the third matching result is an object keyword mismatch result, determining the rule mismatch result as the first matching result; when the second matching result is an action keyword match result, and the third matching result is an object keyword match result, determining the rule matching result as the first matching result.
[0089] In the process of matching the voice text information with the preset intent rule information in the multi-intent rule set, the present embodiment can extract action keywords and execution object keywords from the preset intent rule information, wherein the action keyword represents a keyword pre-configured for extracting the execution action corresponding to the action, such as turn on, open, close, adjust, etc., and the execution object keyword can represent a keyword pre-configured for extracting the execution object, such as car window, air conditioner, seat, etc.; thereby, the voice text information can be matched with the action keyword to obtain a second matching result, which can be an action keyword mismatch result or an action keyword matching result, the action keyword mismatch result indicating that the voice text information does not contain text matching the action keyword, and the action keyword matching result indicating that the voice text information contains text matching the action keyword; and the voice text information can be matched with the execution object keyword to obtain a third matching result, which can be an object keyword mismatch result or an object keyword matching result, the object keyword mismatch result indicating that the voice text information does not contain text matching the execution object keyword, and the object keyword matching result indicating that the voice text information contains Contains text that matches the execution object keyword; when the second matching result is an action keyword mismatch result, and / or the third matching result is an object keyword mismatch result, it means that the voice text information does not contain text that matches the action keyword, and / or does not contain text that matches the execution object keyword, and at this time the rule mismatch result can be determined as the first matching result; when the second matching result is an action keyword matching result, and the third matching result is an object keyword matching result, it means that the voice text information contains text that matches the action keyword, and contains text that matches the execution object keyword, then the rule matching result is determined as the first matching result, so as to perform the subsequent step of determining the target intent rule information from the preset intent rule information, and in the process of performing text extraction on the voice input information based on the target intent rule information to obtain text content information and text position information, the voice input information can be subjected to action keyword extraction and execution object keyword extraction respectively according to the action keywords and execution object keywords in the target intent rule information to obtain text content information containing action keywords and execution object keywords.
[0090] In one example, if the preset intention rule information is (?P <open>(open|on)).*(?P <device>(air conditioning|windows))->open&device, at this time, the action keyword extracted from the preset intention rule information is "open|turn on" and the execution object keyword is "air conditioning|windows". If the voice text information is "turn on the radio and Bluetooth", the voice text information is matched with the action keyword, that is, "open" and "open" have a matching relationship, and the second matching result is the action keyword matching result, and the voice text information is matched with the execution object keyword. At this time, "air conditioning|windows" and "radio and Bluetooth" do not have a matching relationship, and the third matching result is the object keyword mismatch result, and the rule mismatch result can be determined as the first matching result; and if the voice text information is "help turn on the air conditioning and windows", the voice text information is matched with the action keyword, that is, "open" and "open" have a matching relationship, and the second matching result is the action keyword matching result, and the voice text information is matched with the execution object keyword. At this time, "air conditioning|windows" and "air conditioning and windows" have a matching relationship, and the third matching result is the object keyword matching result, and the rule matching result can be determined as the first matching result. Of course, the above is only an example for illustration. In specific implementation, a variety of different preset intention rule information can be pre-configured, and this embodiment does not make specific limitations on this.
[0091] It can be seen that this embodiment compares and judges the action keywords and execution object keywords in the preset intention rule information with the voice text information, and determines the rule matching result as the first matching result only when both the action keywords and the execution object keywords match the voice text information. The target intention rule information determined by the first matching result can accurately extract keywords from the voice text information, thereby improving the accuracy and efficiency of intent recognition.
[0092] In an optional embodiment of the present application, step S140 performs intent recognition based on text content information and text position information to obtain at least two intent recognition results, which may specifically include the following steps: extracting target sentence structure information from the target intent rule information; rewriting the sentence based on the target sentence structure information in combination with the text content information and text position information to obtain the target sentence text information; and determining the target sentence text information as the intent recognition result.
[0093] After obtaining the text content information and the text position information, the present embodiment can extract the target sentence structure information from the target intent rule information. The target sentence structure information can represent a pre-configured sentence format for rewriting the text content, such as an action keyword + execution object keyword sentence, an execution object keyword + action keyword sentence, etc. Of course, different target intent rule information can correspond to different target sentence structure information; based on the target sentence structure information, the sentence is rewritten in combination with the text content information and the text position information to obtain the target sentence text information. Since the text content information is the text content obtained by text extraction from the voice input information, the keyword attributes corresponding to each word in the text content information can be determined. The keyword attributes can be attributes such as execution objects and actions. The method for determining the number of keywords can be through a word analysis model, or it can be determined by adding keyword attributes when extracting text from the voice input information based on the action keywords and execution object keywords in the target intent rule information. This embodiment does not make specific restrictions on this; that is, the keyword attributes of each word in the text content information can be used to rewrite the sentence according to the target sentence structure information to obtain the target sentence text information, and then the target sentence text information can be determined as the intent recognition result.
[0094] It should be noted that, since a voice text information can correspond to one or more target intent rule information, in the process of executing the steps of this embodiment, the steps of this embodiment can be repeated for each target intent rule information corresponding to the same voice text information, so as to obtain the target sentence text information corresponding to each target intent rule information.
[0095] In one example, when extracting the target sentence structure information from the target intent rule information as a sentence of action keyword + execution object keyword, if the text content information contains open, air conditioning, and car window, the keyword attribute of open is the action attribute, the keyword attribute of air conditioning is the execution object attribute, and the keyword attribute of car window is the execution object attribute. The specific determination method can be based on the extracted action keywords and execution object keywords in the target intent rule information, and in the process of text extraction of the voice input information, the text extraction is performed based on the action keywords to obtain open, then it can be determined that the keyword attribute of open is the action attribute, and the text extraction is performed based on the execution object keywords to obtain air conditioning and car windows, the keyword attributes of air conditioning and car windows can be determined as execution object attributes, of course, other methods can also be used; then, according to the target sentence structure information of action keyword + execution object keyword, combined with the keyword attribute of open as action attribute, the keyword attribute of air conditioning as execution object attribute, and the keyword attribute of car windows as execution object attribute, the sentence can be rewritten to obtain the target sentence text information "turn on the air conditioner" and the target sentence text information "open the car windows", so as to determine the target sentence text information "turn on the air conditioner" and the target sentence text information "open the car windows" as intention recognition results respectively, that is, the target sentence text information and the intention recognition result can have a one-to-one correspondence.
[0096] It should also be noted that the intention recognition result in this embodiment may also include text position information corresponding to the target sentence text information. For example, if the intention recognition result includes the target sentence text information "turn on the air conditioner", the intention recognition result may also include text position information corresponding to "turn on" and text position information corresponding to "air conditioner".
[0097] In an optional embodiment of the present application, based on the target arbitration strategy corresponding to the target intent rule information, intent arbitration is performed on each intent recognition result to obtain the target intent information corresponding to the speech input information, which may specifically include the following steps: based on the target arbitration strategy, extracting the target sentence text information from each intent recognition result; determining the starting offset information and the ending offset information corresponding to each target sentence text information; based on the starting offset information and the ending offset information, performing intent arbitration on each target sentence text information to obtain arbitration text information; based on the arbitration text information, generating the target intent information.
[0098] In the present embodiment, during the process of performing intent arbitration on each intent recognition result, the target sentence text information can be extracted from each intent recognition result based on the target arbitration strategy, and the starting offset information and the ending offset information corresponding to each target sentence text information can be determined, wherein the starting offset information can indicate the position of the starting text in the target sentence text information in the voice input information, and the ending offset information can indicate the position of the ending text in the target sentence text information in the voice input information; then, intent arbitration can be performed on each target sentence text information based on the starting offset information and the ending offset information to obtain arbitration text information, wherein intent arbitration indicates the process of screening the target sentence text information according to the target arbitration strategy in combination with the starting offset information and the ending offset information, and the arbitration text information indicates the target sentence text information obtained after screening by the target arbitration strategy, and then the target intent information can be generated based on the arbitration text information.
[0099] In one example, based on the target arbitration strategy, the target sentence text information is extracted from each intent recognition result to obtain target sentence text information A, target sentence text information B, and target sentence text information C, and the starting offset information A1 and the ending offset information A2 corresponding to the target sentence text information A are determined, and the starting offset information B1 and the ending offset information B2 corresponding to the target sentence text information A are determined, and the starting offset information C1 and the ending offset information C2 corresponding to the target sentence text information A are determined. According to the starting offset information and the ending offset information, each target sentence text information is subjected to intent arbitration, and the arbitration text information is obtained as the target sentence text information A and the target sentence text information B. At this time, the target intent information can be generated based on the target sentence text information A and the target sentence text information B corresponding to the arbitration text information. Of course, the above is only an example for illustration, and this embodiment specifically defines this step.
[0100] In an optional embodiment of the present application, intent arbitration is performed on each target sentence text information based on the starting offset information and the ending offset information to obtain arbitration text information, which may specifically include the following steps: determining the intent boundary information corresponding to each target sentence text information based on the starting offset information and the ending offset information; comparing the intent boundary information between each two target sentence text information to obtain an intent boundary comparison result; based on the intent boundary comparison result, performing intent arbitration on each target sentence text information to obtain arbitration text information.
[0101] In the process of performing intention arbitration on each target sentence text information, the present embodiment can determine the intention boundary information corresponding to each target sentence text information based on the starting offset information and the ending offset information, and the intention boundary information indicates the position area where the starting text and the ending text contained in the target sentence text information span the voice input information; then the intention boundary information between each two target sentence text information can be compared to obtain the intention boundary comparison result, and the intention boundary comparison result can indicate the relationship between each two intention boundary information, such as the relationship of starting overlap and ending non-overlap, the inclusion relationship, etc.; then based on the intention boundary comparison result, perform intention arbitration on each target sentence text information to obtain arbitration text information, wherein the intention arbitration can be to screen each target sentence text information based on the intention boundary comparison result, and determine the screened target sentence text information as the arbitration text information.
[0102] In an optional embodiment of the present application, based on the intention boundary comparison result, intention arbitration is performed on each target sentence text information to obtain arbitration text information, which may specifically include the following steps: when the intention boundary comparison result is a boundary overlap result, the two target sentence text information corresponding to the boundary overlap result are merged to obtain arbitration text information; when the intention boundary comparison result is a boundary start overlap result or a boundary end overlap result, arbitration text information corresponding to each target sentence text information is generated.
[0103] In the process of performing intention arbitration on each target sentence text information based on the intention boundary comparison result, this embodiment can determine whether the intention boundary comparison result is a boundary overlap result. The boundary overlap result indicates that the starting offset information corresponding to the two target sentence text information overlaps and the ending offset information overlaps, indicating that the two target sentence text information are the same target sentence text information; for example, both target sentence text information are "turn on the air conditioner", at this time, the two target sentence text information corresponding to the boundary overlap result can be merged into one target sentence text information to obtain arbitration text information. Then, it can be determined whether the intention boundary comparison result is a boundary start overlap result or a boundary end overlap result, wherein the boundary start overlap result indicates that the starting offset information corresponding to the two target sentence text information overlaps but the ending offset information does not overlap, and the boundary end overlap result indicates that the starting offset information corresponding to the two target sentence text information does not overlap but the ending offset information overlaps; it can be seen that when the intention boundary comparison result is a boundary start overlap result or a boundary end overlap result, it can be explained that the two target sentence text information only have the same action or the same execution object, such as the voice input information "turn on the air conditioner". The intention boundary comparison result of the target sentence text information "turn on the air conditioner" corresponding to the "adjust the windows" and the target sentence text information "open the windows" is a boundary start overlap result, and the intention boundary comparison result of the target sentence text information "turn on Bluetooth" corresponding to the voice input information "turn on and display Bluetooth" and the target sentence text information "display Bluetooth" is a boundary end overlap result. At this time, the two target sentence text information can be retained, that is, when the intention boundary comparison result is a boundary start overlap result or a boundary end overlap result, arbitration text information corresponding to each target sentence text information can be generated.
[0104] In an optional embodiment of the present application, based on the intention boundary comparison result, intent arbitration is performed on each target sentence text information to obtain arbitration text information, which may specifically include the following steps: when the intention boundary comparison result is a boundary inclusion result, the boundary ranges of the two target sentence text information corresponding to the boundary inclusion result are compared based on the starting offset information and the ending offset information to obtain a boundary range comparison result; based on the boundary range comparison result, the arbitration text information is determined.
[0105] In the process of performing intent arbitration on each target sentence text information based on the intent boundary comparison result, this embodiment can determine whether the intent boundary comparison result is a boundary inclusion result. The boundary overlap result indicates that the starting offset information and the ending offset information of one of the two target sentence text information are included between the starting offset information and the ending offset information of the other target sentence text information. At this time, the boundary ranges of the two target sentence text information corresponding to the boundary inclusion result can be compared based on the starting offset information and the ending offset information to obtain the boundary range comparison result. The boundary range comparison result can indicate the inclusion relationship corresponding to the two boundary ranges. Subsequently, the arbitration text information can be determined based on the boundary range comparison result.
[0106] Specifically, if the two target sentence text information corresponding to the comparison boundary inclusion result are the first target sentence text information and the second target sentence text information, respectively, the first target sentence text information corresponds to the first boundary range, and the second target sentence text information corresponds to the second boundary range. At this time, the boundary range comparison result can represent the inclusion relationship between the first boundary range and the second boundary range. If the boundary range comparison result indicates that the first boundary range is included by the second boundary range, the second boundary range can be determined as the target boundary range, and the target sentence text information corresponding to the target boundary range can be determined as the arbitration text information, that is, the second target sentence text information can be determined as the arbitration text information. Conversely, if the boundary range comparison result indicates that the second boundary range is included by the first boundary range, the first boundary range can be determined as the target boundary range, and the target sentence text information corresponding to the target boundary range can be determined as the arbitration text information, that is, the first target sentence text information can be determined as the arbitration text information. Of course, the above is only an example for illustration, and this embodiment does not make specific limitations on this.
[0107] In one example, if the voice input information is "adjust the air conditioner to, uh, turn on the air conditioner, adjust it to 15 degrees", the two target sentence text information corresponding to the voice input information are the target sentence text information "adjust the air conditioner to 15 degrees" and the target sentence text information "turn on the air conditioner", and at this time, the starting offset information and the ending offset information of "turn on the air conditioner" are both included between the starting offset information and the ending offset information of "adjust the air conditioner to 15 degrees", and the target sentence text information "adjust the air conditioner to 15 degrees" can be determined as the arbitration text information. Of course, the above is only an example for illustration, and this embodiment does not make specific limitations on this.
[0108] In combination with the foregoing embodiments, it can be known that after obtaining voice input information, the voice input information can be used as a query, and based on the target intent rule information, text extraction is performed on the query to obtain text content information and text position information, so that the text content information and text position information generate at least two target sentence text information corresponding to the query, and then the starting offset information of the target sentence text information is used as start and the ending offset information is used as end. According to the target arbitration strategy corresponding to different target intent rule information, the target sentence text information corresponding to each intent recognition result can be combined with start and end to perform intent arbitration to obtain the target intent information corresponding to the voice input information. Figure 2-4 , for example:
[0109] 1. When there is partial overlap in the text information of multiple target sentences, that is, the start is the same but the end is different, or the start is different but the end is the same, such as Figure 2 As shown in the figure, result1 and result2 represent different target sentence text information, and it is considered that all the multiple results are valid;
[0110] 2. When multiple target sentence text information has an inclusion relationship, such as Figure 3 As shown, result1 and result2 in the figure represent different target sentence text information, and a larger starting range can be retained;
[0111] 3. When there is overlap in the text information of multiple target sentences, such as Figure 4 As shown in the figure, result1 and result2 represent different target sentence text information, which can ensure that at least one result of each rule is effective.
[0112] Of course, the above is only an example for illustration. In specific implementation, different target arbitration strategies can be configured for different target intent rule information, so as to perform intent arbitration on the target sentence text information corresponding to each intent recognition result.
[0113] In an optional embodiment of the present application, based on the target arbitration strategy corresponding to the target intention rule information, each intention recognition result is subjected to intention arbitration to obtain the target intention information corresponding to the voice input information. Specifically, the following steps may be included: performing semantic analysis on the target intention information to obtain a voice execution instruction; controlling the vehicle to execute the instruction operation corresponding to the voice execution instruction to obtain a voice control result corresponding to the voice input information.
[0114] After obtaining the target intention information corresponding to the voice input information, this embodiment can perform semantic analysis on the target intention information to obtain a voice execution instruction. The voice execution instruction can represent the control instruction corresponding to the voice input information. Since the target intention information in this embodiment is the intention recognition result retained after intention arbitration, that is, the target intention information can directly and simply represent the user's intention, a simple model or algorithm can be used for voice analysis at this time to accurately and efficiently obtain the voice execution instruction corresponding to the target intention information, thereby controlling the vehicle to execute the instruction operation corresponding to the voice execution instruction, and obtaining the voice control result corresponding to the voice input information. The voice control result can represent the completion status of the instruction operation.
[0115] It can be seen that the multi-intent recognition method of this embodiment can be applied to the command matching algorithm or semantic analysis before it is applied, so that the subsequent command matching algorithm or semantic analysis can accurately and efficiently obtain the voice input information and the corresponding voice execution instructions, which can effectively improve the accuracy and efficiency of multi-intent recognition.
[0116] like Figure 5 As shown, the present application also discloses an embodiment, providing a multi-intention recognition device, including:
[0117] An acquisition module 510 is used to acquire voice input information of the vehicle;
[0118] A determination module 520, configured to determine target intent rule information matching the voice input information according to a preset multi-intent rule set corresponding to the vehicle;
[0119] An extraction module 530 is used to perform text extraction on the voice input information based on the target intention rule information to obtain text content information and text position information;
[0120] An intention recognition module 540, configured to perform intention recognition based on the text content information and the text position information to obtain at least two intention recognition results;
[0121] The arbitration module 550 is used to perform intent arbitration on each of the intent recognition results based on the target arbitration strategy corresponding to the target intent rule information, and obtain the target intent information corresponding to the speech input information.
[0122] In an optional embodiment of the present application, the voice input information is voice information collected by the vehicle-mounted device, and the determination module 520 may include:
[0123] A first acquisition unit, configured to acquire a multi-intention rule set preset for the vehicle based on the voice input information;
[0124] A text recognition unit, used to perform text recognition on the voice input information to obtain voice text information;
[0125] A matching unit, configured to match the voice text information with the preset intention rule information in the multi-intention rule set to obtain a first matching result;
[0126] The first determining unit is used to determine the target intention rule information by using the preset intention rule information when the first matching result is a rule matching result.
[0127] In an optional embodiment of the present application, the matching unit may include:
[0128] A first extraction subunit, configured to extract action keywords and execution object keywords from the preset intention rule information;
[0129] A first matching subunit, used for matching the voice text information with the action keyword to obtain a second matching result;
[0130] A second matching subunit is used to match the voice text information with the execution object keyword to obtain a third matching result;
[0131] a first determining subunit, configured to determine the rule non-matching result as the first matching result when the second matching result is an action keyword non-matching result and / or the third matching result is an object keyword non-matching result;
[0132] The second determining subunit is configured to determine the rule matching result as the first matching result when the second matching result is an action keyword matching result and the third matching result is an object keyword matching result.
[0133] In an optional embodiment of the present application, the intention recognition module 540 may include:
[0134] A first extraction unit, configured to extract target sentence structure information from the target intention rule information;
[0135] A sentence rewriting unit, configured to rewrite the sentence based on the target sentence structure information, in combination with the text content information and the text position information, to obtain target sentence text information;
[0136] The second determination unit is used to determine the target sentence text information as the intention recognition result.
[0137] In an optional embodiment of the present application, the arbitration module 550 may include:
[0138] A second extraction unit, configured to extract target sentence text information from each of the intention recognition results based on the target arbitration strategy;
[0139] A third determining unit is used to determine the starting offset information and the ending offset information corresponding to each target sentence text information;
[0140] An intention arbitration unit, configured to perform intention arbitration on each target sentence text information according to the start offset information and the end offset information to obtain arbitration text information;
[0141] The first generating unit is used to generate the target intention information based on the arbitration text information.
[0142] In an optional embodiment of the present application, the intention arbitration unit may include:
[0143] A third determining subunit is used to determine the intention boundary information corresponding to each of the target sentence text information according to the starting offset information and the ending offset information;
[0144] A first comparison subunit is used to compare the intention boundary information between each two target sentence text information to obtain an intention boundary comparison result;
[0145] The intention arbitration subunit is used to perform intention arbitration on each target sentence text information based on the intention boundary comparison result to obtain the arbitration text information.
[0146] In an optional embodiment of the present application, the intention arbitration subunit may include:
[0147] A merging subunit, configured to merge the two target sentence text information corresponding to the boundary coincidence result to obtain the arbitration text information when the intention boundary comparison result is a boundary coincidence result;
[0148] A generating subunit is used to generate the arbitration text information corresponding to each target sentence text information when the intention boundary comparison result is a boundary start overlap result or a boundary end overlap result.
[0149] In an optional embodiment of the present application, the intention arbitration subunit may include:
[0150] A second comparison subunit is used for comparing the boundary ranges of the two target sentence text information corresponding to the boundary inclusion result according to the starting offset information and the ending offset information when the intention boundary comparison result is a boundary inclusion result, so as to obtain a boundary range comparison result;
[0151] The fourth determining subunit is used to determine the arbitration text information according to the boundary range comparison result.
[0152] In an optional embodiment of the present application, the multi-intention recognition device may further include:
[0153] A semantic analysis module, used to perform semantic analysis on the target intention information to obtain a voice execution instruction;
[0154] The control module is used to control the vehicle to execute the command operation corresponding to the voice execution command and obtain the voice control result corresponding to the voice input information.
[0155] The implementation process of the functions and effects of each module in the above-mentioned device is specifically described in the implementation process of the corresponding steps in the above-mentioned method, which will not be repeated here.
[0156] like Figure 6 As shown, an embodiment of the present application provides an electronic device, including a processor 610, a communication interface 620, a memory 630 and a communication bus 640, wherein the processor 610, the communication interface 620, and the memory 630 communicate with each other through the communication bus 640.
[0157] Memory 630, for storing computer programs;
[0158] In one embodiment of the present application, the processor 610, when executing the program stored in the memory 630, implements the multi-intent recognition method provided by any of the aforementioned method embodiments, by obtaining voice input information to determine the target intent rule information matching the voice input information, and based on the target intent rule information, performs text extraction on the voice input information to obtain text content information and text position information, and performs intent recognition based on the text content information and text position information to obtain at least two intent recognition results, and then performs intent arbitration on each intent recognition result based on the target arbitration strategy corresponding to the target intent rule information to obtain the target intent information corresponding to the voice input information; thereby solving the problems of low accuracy and low efficiency of multi-intent recognition in the existing related technologies, and can effectively improve the accuracy and efficiency of multi-intent recognition.
[0159] An embodiment of the present application also provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the steps of the multi-intent recognition method provided in any of the aforementioned method embodiments are implemented, by obtaining voice input information to determine target intent rule information matching the voice input information, performing text extraction on the voice input information based on the target intent rule information to obtain text content information and text position information, and performing intent recognition based on the text content information and the text position information to obtain at least two intent recognition results, and then performing intent arbitration on each intent recognition result based on a target arbitration strategy corresponding to the target intent rule information to obtain target intent information corresponding to the voice input information; thereby solving the problems of low accuracy and low efficiency of multi-intent recognition in the existing related technologies, and being able to effectively improve the accuracy and efficiency of multi-intent recognition.
[0160] An embodiment of the present application also provides a vehicle, which includes the multi-intention recognition device described in any of the aforementioned embodiments.
[0161] The device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0162] Through the description of the above implementation methods, those skilled in the art can clearly understand that each implementation method can be implemented by means of software plus a general hardware platform, and of course, by hardware. Based on this understanding, the above technical solution is essentially or the part that contributes to the relevant technology can be embodied in the form of a software product, and the computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a disk, an optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0163] It should be understood that the terms used herein are only for the purpose of describing specific example embodiments and are not intended to be limiting. Unless the context clearly indicates otherwise, the singular forms "one", "an" and "said" as used herein may also be meant to include plural forms. The terms "include", "comprise", "contain", and "have" are inclusive and therefore specify the existence of stated features, steps, operations, elements and / or parts, but do not exclude the existence or addition of one or more other features, steps, operations, elements, parts, and / or combinations thereof. The method steps, processes, and operations described herein are not interpreted as necessarily requiring them to be performed in the specific order described or illustrated, unless the execution order is clearly indicated. It should also be understood that additional or alternative steps may be used.
[0164] The foregoing is merely a specific embodiment of the present invention, which enables those skilled in the art to understand or implement the present invention. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but rather to the widest scope consistent with the principles and novel features claimed herein.< / device> < / open> < / open> < / device> < / device> < / open> < / device> < / open>
Claims
1. A multi-intention recognition method, characterized in that: include: Get voice input information; Determining target intent rule information matching the speech input information; Based on the target intention rule information, performing text extraction on the voice input information to obtain text content information and text position information; Performing intent recognition based on the text content information and the text position information to obtain at least two intent recognition results; Based on the target arbitration strategy corresponding to the target intent rule information, intent arbitration is performed on each of the intent recognition results to obtain the target intent information corresponding to the speech input information.
2. The multi-intention recognition method according to claim 1, characterized in that: The voice input information is voice information collected by the vehicle-mounted device, and the determining of target intention rule information matching the voice input information includes: Acquire a multi-intention rule set preset for the vehicle based on the voice input information; Performing text recognition on the voice input information to obtain voice text information; Matching the voice text information with the preset intent rule information in the multi-intent rule set to obtain a first matching result; In the case where the first matching result is a rule matching result, the preset intention rule information is used to determine the target intention rule information.
3. The multi-intention recognition method according to claim 2, characterized in that: The matching of the voice text information with the preset intention rule information in the multi-intention rule set to obtain a first matching result includes: Extracting action keywords and execution object keywords from the preset intention rule information; Matching the voice text information with the action keyword to obtain a second matching result; Matching the voice text information with the execution object keyword to obtain a third matching result; When the second matching result is an action keyword mismatch result, and / or the third matching result is an object keyword mismatch result, the rule mismatch result is determined as the first matching result; In the case where the second matching result is an action keyword matching result and the third matching result is an object keyword matching result, the rule matching result is determined as the first matching result.
4. The multi-intention recognition method according to claim 1, characterized in that: The performing of intent recognition based on the text content information and the text position information to obtain at least two intent recognition results includes: extracting target sentence structure information from the target intention rule information; Based on the target sentence structure information, the sentence is rewritten in combination with the text content information and the text position information to obtain target sentence text information; The target sentence text information is determined as the intention recognition result.
5. The multi-intention recognition method according to claim 1, characterized in that: The target arbitration strategy corresponding to the target intent rule information is based on which the intent arbitration is performed on each of the intent recognition results to obtain the target intent information corresponding to the speech input information, including: Based on the target arbitration strategy, extracting target sentence text information from each of the intent recognition results; Determine the starting offset information and the ending offset information corresponding to each target sentence text information; According to the start offset information and the end offset information, performing intention arbitration on each target sentence text information to obtain arbitration text information; Based on the arbitration text information, the target intention information is generated.
6. The multi-intention recognition method according to claim 5, characterized in that: The method of performing intention arbitration on each target sentence text information according to the start offset information and the end offset information to obtain arbitration text information includes: Determining the intention boundary information corresponding to each of the target sentence text information according to the starting offset information and the ending offset information; Comparing the intention boundary information between each two target sentence text information to obtain an intention boundary comparison result; Based on the intention boundary comparison result, intention arbitration is performed on each target sentence text information to obtain the arbitration text information.
7. The multi-intention recognition method according to claim 6, characterized in that: The performing intention arbitration on each target sentence text information based on the intention boundary comparison result to obtain the arbitration text information includes: In the case where the intention boundary comparison result is a boundary overlap result, merging the two target sentence text information corresponding to the boundary overlap result to obtain the arbitration text information; When the intention boundary comparison result is a boundary start overlap result or a boundary end overlap result, the arbitration text information corresponding to each target sentence text information is generated.
8. The multi-intention recognition method according to claim 6, characterized in that: The performing intention arbitration on each target sentence text information based on the intention boundary comparison result to obtain the arbitration text information includes: When the intention boundary comparison result is a boundary inclusion result, comparing the boundary ranges of the two target sentence text information corresponding to the boundary inclusion result according to the start offset information and the end offset information to obtain a boundary range comparison result; The arbitration text information is determined according to the boundary range comparison result.
9. The multi-intention recognition method according to any one of claims 1 to 8, characterized in that: After performing intention arbitration on each of the intention recognition results based on the target arbitration strategy corresponding to the target intention rule information to obtain the target intention information corresponding to the speech input information, the method further includes: Performing semantic analysis on the target intention information to obtain a voice execution instruction; The vehicle is controlled to execute the command operation corresponding to the voice execution command, and a voice control result corresponding to the voice input information is obtained.
10. A multi-intention recognition device, characterized in that: include: An acquisition module, used to acquire the voice input information of the vehicle; A determination module, configured to determine target intent rule information matching the voice input information according to a preset multi-intention rule set corresponding to the vehicle; An extraction module, used to perform text extraction on the voice input information based on the target intention rule information to obtain text content information and text position information; An intention recognition module, used to perform intention recognition based on the text content information and the text position information to obtain at least two intention recognition results; An arbitration module is used to perform intent arbitration on each of the intent recognition results based on a target arbitration strategy corresponding to the target intent rule information, so as to obtain target intent information corresponding to the speech input information.
11. An electronic device, characterized in that: It includes a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory communicate with each other through the communication bus; Memory, used to store computer programs; A processor, used to implement the multi-intention recognition method described in any one of claims 1-9 when executing a program stored in a memory.
12. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the multi-intention recognition method as described in any one of claims 1-9 is implemented.
13. A vehicle, characterized in that: The vehicle includes the multi-intention recognition device according to claim 10.