Method, apparatus, and electronic device for matching buttons according to voice
By converting user voice information into instruction text and performing multi-level matching, the problem of restricted rule definitions in the prior art is solved, and the recognition of natural expression statements and accurate matching of buttons is achieved, which improves the flexibility and efficiency of user interaction.
Patent Information
- Application Number
- CN202211647975.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-21
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2042-12-21
AI Technical Summary
In the prior art, the rule definition is limited, and the voice commands under the user's natural expression statement cannot be recognized, resulting in limited user interaction.
Three matching methods are used to convert user voice information into instruction text, character matching, first logical matching and second logical matching are performed, and combined with semantic slots, speech domains and semantic intent rules, a recall result is generated to identify the target button.
It improves the matching success rate of various buttons, supports users' natural language interaction, reduces maintenance and update costs, and improves user experience.
Smart Images

Figure CN116052675B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of autonomous driving, and in particular, to a method, an apparatus, and an electronic device for matching buttons according to voice. Background Art
[0002] Driven by the dual promotion of the change of the consumer group and the shaping of the new profit model, the "intelligentization" of automobiles is accelerating, and the boundary of "intelligentization" is also constantly expanding. In the whole vehicle intelligentization, in addition to autonomous driving, the "intelligent cockpit" occupies a core position. The physical buttons in the vehicle are gradually incorporated into the electronic large screen, and the "visible and speakable" human-computer interaction is realized by triggering "simulated touch" through "voice", that is, any text and button that the user sees on the vehicle computer can be directly interacted with by issuing a voice command, such as "click the Bluetooth switch" and "click the navigation button".
[0003] At present, the existing implementation schemes of "visible and speakable" are divided into two ways: one is through the way of full string matching, by determining whether the user instruction text completely matches the text displayed on the screen. The problem with this method is that the full match limits the user instructions and cannot recognize more natural expressions. The other is through the way of rule templates, by defining a large number of rule sentence templates and then combining them with the text displayed on the screen. To a certain extent, this method will bring more sentence support, but limited by the definition and maintenance of a large number of rule sentences, the rule definition is limited, and the maintenance cost and usage cost are high.
[0004] Therefore, the prior art has the problems that the rule definition is limited and the user voice under natural expression cannot be recognized. Summary of the Invention
[0005] The present application provides a method, an apparatus, and an electronic device for matching buttons according to voice, so as to at least solve the problems in the related art that the rule definition is limited and the user voice under natural expression cannot be recognized.
[0006] According to one aspect of the embodiments of the present application, a method for matching buttons according to voice is provided. The method includes:
[0007] Converting the user's voice information into an instruction text;
[0008] Performing character matching on the instruction text to determine whether a target button can be obtained;
[0009] In the case where the target button cannot be obtained, performing a first logical matching on the instruction text to obtain a first recall result;
[0010] Performing a second logical matching on the instruction text to obtain a second recall result;
[0011] Based on the first recall result and the second recall result, the target button is obtained.
[0012] According to another aspect of the embodiments of the present application, there is also provided a device for matching buttons according to voice, the device includes:
[0013] A conversion module, configured to convert the user's voice information into an instruction text;
[0014] A first matching module, configured to perform character matching on the instruction text to determine whether the target button can be obtained;
[0015] A second matching module, configured to perform a first logical matching on the instruction text to obtain a first recall result when the target button cannot be obtained;
[0016] A third matching module, configured to perform a second logical matching on the instruction text to obtain a second recall result;
[0017] An obtaining module, configured to obtain the target button according to the first recall result and the second recall result.
[0018] Optionally, the second matching module includes:
[0019] An input unit, configured to input the instruction text into a preset model to obtain the semantic slot, the speech field, and the semantic intention corresponding to the instruction text;
[0020] A first obtaining unit, configured to obtain a first preset button definition of a first preset type of button, where the first preset button definition includes a semantic slot rule, a speech field rule, and a semantic intention rule for each first preset type of button;
[0021] A first obtaining unit, configured to obtain the first recall result when the semantic slot, the speech field, and the semantic intention corresponding to the instruction text respectively match the semantic slot rule, the speech field rule, and the semantic intention rule of the same first preset type of button successfully.
[0022] Optionally, the first obtaining unit includes:
[0023] A first obtaining sub-module, configured to obtain the correct semantic slot type and name, and the incorrect semantic slot type and name corresponding to each first preset type of button, where the incorrect semantic slot type and name are used to avoid incorrect matching;
[0024] A first generating sub-module, configured to generate the semantic slot rule according to the correct semantic slot type and name, and the incorrect semantic slot type and name;
[0025] A second acquisition sub-module, configured to acquire the corresponding conversation domain rules and semantic intention rules of each of the first preset type of buttons;
[0026] A third acquisition sub-module, configured to acquire the first identifier of each of the first preset type of buttons;
[0027] A second generation sub-module, configured to generate the first preset button definition of each of the first preset type of buttons according to the first identifier, the semantic slot rules, the conversation domain rules, and the semantic intention rules.
[0028] Optionally, the third matching module includes:
[0029] A second acquisition unit, configured to acquire a dictionary file and a recall path file;
[0030] A construction unit, configured to construct a prefix tree according to the dictionary file;
[0031] A second obtaining unit, configured to perform word segmentation on the instruction text to obtain a word list;
[0032] A third obtaining unit, configured to obtain candidate main component words according to the word list, a preset matching method, the dictionary file, and the prefix tree, where the candidate main component words include a candidate main body name and a candidate action of a button;
[0033] A generation unit, configured to generate a candidate recall path set according to the candidate main component words, where the candidate recall path set includes a first preset number of candidate recall paths;
[0034] A judgment unit, configured to judge whether each of the candidate recall paths only includes one candidate main body name and one candidate action;
[0035] A deletion unit, configured to delete the candidate recall path when the candidate recall path does not only include one candidate main body name and one candidate action;
[0036] A fourth obtaining unit, configured to, when all the candidate recall paths only include one candidate main body name and one candidate action, match each candidate recall path in the candidate recall path set with a preset recall path in the recall path file to obtain the second recall result.
[0037] Optionally, the second acquisition unit includes:
[0038] The fourth acquisition sub-module is used to acquire the second preset button definitions of the second preset type of buttons, where the second preset button definitions include the second identifiers, the set of main body names, and the set of action actions of each of the second preset type of buttons. The set of main body names includes the alias names, standard names, and error names of the second preset type of buttons. The set of action actions includes the alias names, standard names, and error names of the actions. The error names of the second preset type of buttons and the error names of the actions are used to avoid matching errors. The alias names are the correct names other than the standard names.
[0039] The first construction sub-module is used to construct a dictionary file according to the alias names, standard names, error names of the second preset type of buttons, the alias names, standard names, error names of the actions, and the preset order.
[0040] The third generation sub-module is used to generate the second preset number of preset recall paths corresponding to each of the second preset type of buttons according to the second identifiers, the set of main body names, and the set of action actions.
[0041] The second construction sub-module is used to construct a recall path file according to the preset recall paths.
[0042] Optionally, the fourth acquisition sub-module includes:
[0043] The first acquisition sub-unit is used to acquire the alias names, standard names, and error names of each of the second preset type of buttons.
[0044] The first construction sub-unit is used to construct the set of main body names of each of the second preset type of buttons according to the alias names, standard names, and error names of the second preset type of buttons.
[0045] The second acquisition sub-unit is used to acquire the third preset number of actions supported by each of the second preset type of buttons.
[0046] The third acquisition sub-unit is used to acquire the alias names, standard names, and error names of each of the actions.
[0047] The second construction sub-unit is used to construct the set of action actions of each of the second preset type of buttons according to the alias names, standard names, and error names of the actions.
[0048] The fourth acquisition sub-unit is used to acquire the second identifiers of each of the second preset type of buttons.
[0049] The generation sub-unit is used to generate the second preset button definitions of each of the second preset type of buttons according to the second identifiers, the set of main body names, and the set of action actions.
[0050] Optionally, the third obtaining unit includes:
[0051] A fifth obtaining sub-module, configured to obtain the error name of each of the second preset type buttons according to the dictionary file;
[0052] A first obtaining sub-module, configured to obtain the hit words according to the word list, the preset matching method, the dictionary file, and the prefix tree;
[0053] A screening sub-module, configured to screen the hit words according to the error name to obtain the screened hit words;
[0054] A second obtaining sub-module, configured to obtain the candidate principal component words according to the screened hit words.
[0055] Optionally, the first recall result includes the length of the hit words, and the obtaining module includes:
[0056] A third obtaining unit, configured to obtain a third identifier of a button element in the current interface;
[0057] A fifth obtaining unit, configured to obtain a fourth identifier of a first intermediate button according to the first recall result and the length of the hit words;
[0058] A sixth obtaining unit, configured to obtain the target button according to the third identifier and the fourth identifier;
[0059] A seventh obtaining unit, configured to obtain a fifth identifier of a second intermediate button according to the second recall result;
[0060] An eighth obtaining unit, configured to obtain the target button according to the third identifier and the fifth identifier.
[0061] According to another aspect of the embodiments of the present application, there is also provided an electronic device, including a processor, a communication interface, a memory, and a communication bus. The processor, the communication interface, and the memory complete mutual communication through the communication bus. The memory is used to store a computer program. The processor is configured to execute the method steps in any of the above embodiments by running the computer program stored on the memory.
[0062] According to another aspect of the embodiments of the present application, there is also provided a computer-readable storage medium, in which a computer program is stored. The computer program is configured to execute the method steps in any of the above embodiments when running.
[0063] In the embodiments of the present application, the voice information of the user is converted into an instruction text; the instruction text is subjected to character matching to determine whether a target button can be obtained; in the case where the target button cannot be obtained, the instruction text is subjected to a first logical matching to obtain a first recall result; the instruction text is subjected to a second logical matching to obtain a second recall result; and a target button is obtained according to the first recall result and the second recall result. Through the above method, three different matching methods are used to match the instruction text, which improves the matching success rate for various buttons, so as to meet different user needs and break the rule definition limitations in the traditional technology. By using three different matching methods, the natural language of the user can be recognized, and at the same time, through the configuration of buttons, more convenient, faster, and more flexible generalization statements can be used, reducing the maintenance cost, update cost, and user learning cost of this method. The problem that there are rule definition limitations in the related technology and the user voice under natural expression cannot be recognized is solved. BRIEF DESCRIPTION OF THE DRAWINGS
[0064] The accompanying drawings herein are incorporated into the specification and constitute a part of the specification, showing embodiments consistent with the present invention, and are used together with the specification to explain the principles of the present invention.
[0065] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0066] Figure 1 It is a flowchart of an optional method for matching buttons according to voice according to the embodiments of the present application;
[0067] Figure 2 It is a schematic diagram of an optional visible-and-speakable matching module according to the embodiments of the present application;
[0068] Figure 3 It is a flowchart of an optional method for matching buttons according to voice according to the embodiments of the present application;
[0069] Figure 4 It is a flowchart of another optional second logical matching according to the embodiments of the present application;
[0070] Figure 5 It is a structural block diagram of an optional device for matching buttons according to voice according to the embodiments of the present application;
[0071] Figure 6 It is a structural block diagram of an optional electronic device according to the embodiments of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0072] To enable those skilled in the art to better understand the solution of this application, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in this application without creative efforts shall fall within the scope of protection of this application.
[0073] It should be noted that the terms "first", "second", etc. in the description and claims of this application and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of this application described here can be implemented in an order different from those illustrated or described here. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.
[0074] According to one aspect of the embodiments of this application, a method for matching buttons according to voice is provided. As Figure 1 shown, the process of this method may include the following steps:
[0075] Step S101, convert the user's voice information into an instruction text.
[0076] Optionally, in the in-vehicle system, when the user inputs voice, the user's voice is converted into an instruction text through voice recognition.
[0077] Step S102, perform character matching on the instruction text to determine whether a target button can be obtained.
[0078] Optionally, this application realizes pure string matching through substring matching enhancement and extension, and performs character matching on the instruction text, that is, uses the string substrings in the instruction text for matching.
[0079] Among them, the pure string matching method is used for general basic matching hits. A hit means a successful match, that is: in the in-vehicle system, the names and aliases of the screen buttons are predefined first, and then it is determined whether the user's instruction text exactly matches the predefined name as a string. If the match is successful, the hit is completed and the target button is obtained. On this basis, the present application enhances and extends substring matching, that is: the user can input a string substring of more than three characters of the text displayed on the voice input interface, and the target button can be obtained through this string substring match determination, without the need for the instruction text to exactly match the predefined name as a string. The above method is as Figure 2 shown: Pure string matching is implemented through enhanced extension of substring matching.
[0080] The extended substring matching method can reduce the user's voice input when there is long text on the screen, and support the hit of buttons with long names on the in-vehicle system.
[0081] Step S103, in the case where the target button cannot be obtained, perform a first logical match on the instruction text to obtain a first recall result;
[0082] Optionally, first perform the above text matching of the string substring. If the string substring matching condition is met and the target button is obtained, the result will be preferentially sent. If the hit fails, it will go through the mapping match of the Natural Language Understanding (NLU) model (i.e., the first logical match) and the principal component extraction rule match (i.e., the second logical match), match the instruction text, and generate the corresponding recall result.
[0083] As Figure 2 shown: Define rule constraints for semantic slots, domains, and semantic intents, and perform NLU model mapping matching.
[0084] Obtain the semantic slots, domains, and semantic intents of the instruction text through the NLU model, and then perform button matching according to the above semantic slots, domains, and semantic intents to obtain a first recall result. The first recall result contains information such as the identification ID of the hit button and the length of the hit word.
[0085] Step S104, perform a second logical match on the instruction text to obtain a second recall result.
[0086] Optionally, as Figure 2 shown: Perform flexible operation configuration on the button subject and control action, and perform principal component extraction rule matching.
[0087] Perform principal component extraction on the instruction text, perform matching based on the extraction results, and obtain the second recall result. The second recall result includes the identification ID of the hit button.
[0088] Step S105, obtain the target button according to the first recall result and the second recall result.
[0089] Optionally, the ID can be understood as the unique identifier of the button. After generating the recall result, the ID of the current interface element reported by the in-vehicle terminal will be matched with the IDs of each button in the first recall result and the second recall result. If the match is successful, the button with the successful match is the target button, and the button is triggered.
[0090] In the embodiment of the present application, the user's voice information is converted into an instruction text; the instruction text is subjected to character matching to determine whether the target button can be obtained; in the case where the target button cannot be obtained, the instruction text is subjected to the first logical matching to obtain the first recall result; the instruction text is subjected to the second logical matching to obtain the second recall result; the target button is obtained according to the first recall result and the second recall result. Through the above method, three different matching methods are used to match the instruction text, which improves the matching success rate of various buttons to meet different user needs and relaxes the rule definition limitations in the traditional technology. By using three different matching methods, the natural language of the user can be recognized, and at the same time, through the configuration of buttons, more convenient, faster, and more flexible generalization statements can be used, reducing the maintenance cost, update cost, and user learning cost of this method. It solves the problems in the related technology that the rule definition is restricted and the user's voice under the natural expression cannot be recognized.
[0091] As an optional embodiment, performing the first logical matching on the instruction text to obtain the first recall result includes:
[0092] Input the instruction text into a preset model to obtain the semantic slot, speech domain, and semantic intention corresponding to the instruction text;
[0093] Obtain the first preset button definition of the first preset type of buttons, where the first preset button definition includes the semantic slot rule, speech domain rule, and semantic intention rule of each first preset type of button;
[0094] When the semantic slot, speech domain, and semantic intention corresponding to the instruction text are respectively successfully matched with the semantic slot rule, speech domain rule, and semantic intention rule of the same first preset type of button, the first recall result is obtained.
[0095] Optionally, the NLU model mapping matching method is mainly used for the matching of the first preset type of buttons. The first preset type of buttons includes the following two types of buttons: one is the permanently resident buttons on the screen. For example, a fixed area is reserved at the bottom of the screen to place permanently resident button menu bars such as the back button, Home button, and menu button. Since such buttons always appear on the screen, the model has stronger capabilities. Compared with rule matching, it can avoid mis-hits and mis-triggering to a greater extent. The other is general-purpose buttons, such as "play", "return", "navigation" and other buttons. These buttons are very likely to cause mis-hits. For example Figure 3 as shown: "Select navigation" and "Select navigation to City B", the two user question intents are different and are likely to cause mis-hits.
[0096] In the in-vehicle system, after the user's voice input and after obtaining the user's command text through speech recognition, the NLU model (i.e., the preset model) is called to obtain the corresponding semantic slot S i , discourse domain D i , semantic intent I i , where the semantic slot refers to the important entities and corresponding entity types that appear in a sentence; the discourse domain refers to classifying a sentence into a specific domain, such as: general-purpose, multimedia, navigation, etc.; the semantic intent refers to classifying a sentence into a specific intent, such as: the intent to open the window, the intent to play music, the intent to navigate to a certain place, etc. As Figure 3 shown: there are two command texts, namely "Select navigation" and "Select navigation to City B". Among them, the semantic slot (Slots) of "Select navigation" is action Operation (select), action Operation (navigation), the discourse domain (Domain) is general-purpose (Common), and the semantic intent (Intents) is the general selection intent (common_selection); the semantic slot (Slots) of "Select navigation to City B" is action Operation (select), action Operation (navigation), destination Dest_POI (City B), the discourse domain (Domain) is navigation (Navigation), and the semantic intent (Intents) is the intent to navigate to (navi_go).
[0097] Match the above semantic slot S i , discourse domain D i and semantic intent I i respectively with the semantic slot rules of each button, such as discourse domain rules, such as and semantic intent rules, such as for matching, where g represents the first preset type of button g. If and only if: When it is true, the instruction text can hit button g, and finally the most suitable button is obtained through comparison of the hit path lengths.
[0098] Among them, obtaining the most suitable button through comparison of the hit path lengths at the end means that the names of some buttons may be subsets of the names of other buttons. For example, button A is "blow the face", and button B is "blow the face and feet"; instruction text query1: "Turn on the face blowing mode" and instruction text query2: "Turn on the face and feet blowing mode", their semantic intents and domains of discourse are the same. When the user says query2, the semantic slots are recognized as 2: "blow the face" and "blow the feet". Since it contains "blow the face", when button B is triggered, A will also be triggered. Therefore, here, the hit length comparison is adopted. Each time a semantic slot is hit, its length is increased by 1. Finally, the most suitable one or more buttons are selected for distribution, that is, the first recall result is obtained. The first recall result contains information such as the identification ID of the hit button and the length of the hit word.
[0099] In the embodiment of the present application, through NLU mapping and matching, the capabilities of the model are fully utilized, and there is a higher matching success rate for resident and general buttons, improving the generalization matching success rate for resident and general buttons.
[0100] As an optional embodiment, obtaining the first preset button definition of the first preset type of buttons includes:
[0101] Obtaining the correct semantic slot types and names, incorrect semantic slot types and names corresponding to each first preset type of button, where the incorrect semantic slot types and names are used to avoid incorrect matching;
[0102] Generating semantic slot rules according to the correct semantic slot types and names, incorrect semantic slot types and names;
[0103] Obtaining the corresponding domain of discourse rules and semantic intent rules for each first preset type of button;
[0104] Obtaining the first identifier of each first preset type of button;
[0105] Generating the first preset button definition of each first preset type of button according to the first identifier, semantic slot rules, domain of discourse rules, and semantic intent rules.
[0106] Optionally, it is necessary to pre-define the buttons supported by the NLU model, and set that the first preset type of button g can be defined as a quadruple R g :
[0107]
[0108] Among them, AssistName g represents the first identification ID of button g. Here, the ID can be understood as the unique identifier of the button. After generating the recall result, the ID of the current interface element reported by the in-vehicle terminal will be matched with the IDs of each button in the recall result. If the match is successful, the button will be triggered successfully.
[0109] represents the rule of the semantic slot (Slots) of button g. The semantic slot rule stipulates the mandatory Slots types and corresponding names, as well as the Slots types and names that cannot appear. For example, the Slots type is Operation, and its corresponding names (values) may include open, close, etc. If you want to recall "open xx", you have to clearly stipulate that the corresponding name of Operation can only be open to perform the recall. Similarly, if a Slots type that cannot appear is stipulated, when the model parsing result contains this type, it can be considered in advance that the recall condition of the button is not met, thus omitting a series of subsequent operations and avoiding the accidental triggering of the button.
[0110] In addition, and respectively represent the rules of the discourse domain (Domain) and semantic intent (Intent) of button g. and stipulate that each button will only have one discourse domain and one intent. The same applies to other first preset type buttons as button g, which will not be elaborated here.
[0111] In the embodiment of the present application, the buttons supported by the NLU model are predefined, and the semantic slot rule, discourse domain rule, and semantic intent rule are respectively set, providing a basis for the subsequent NLU model mapping and matching. By setting the Slots types and names that cannot appear, the accidental triggering of other buttons is avoided, improving the accuracy.
[0112] As an alternative embodiment, the instruction text is subjected to a second logical match to obtain a second recall result, including:
[0113] Obtain a dictionary file and a recall path file;
[0114] Construct a prefix tree according to the dictionary file;
[0115] Perform word segmentation on the instruction text to obtain a word list;
[0116] According to the word list, preset matching method, dictionary file, and prefix tree, obtain candidate main component words, where the candidate main component words include the candidate main body name and candidate action of the button;
[0117] Generate a candidate recall path set according to the candidate principal component words, where the candidate recall path set includes a first preset number of candidate recall paths;
[0118] Determine whether each candidate recall path contains only one candidate subject name and one candidate action;
[0119] When the candidate recall path does not contain only one candidate subject name and one candidate action, delete the candidate recall path;
[0120] In the case that all candidate recall paths contain only one candidate subject name and one candidate action, match each candidate recall path in the candidate recall path set with the preset recall paths in the recall path file respectively to obtain a second recall result.
[0121] Optionally, although the above NLU model has strong capabilities, it has a high computational cost and a long training model cycle, and cannot meet the requirement of flexible configuration. To overcome the problem of defining a large number of sentence patterns in common solutions and at the same time meet the requirement of flexible configuration, the present application proposes a solution for principal component extraction rule matching. The principal component extraction rule matching (i.e., the second logical matching) is mainly used for the matching of buttons of a second preset type, and the buttons of the second preset type are buttons other than the above-mentioned first preset type of buttons.
[0122] When the vehicle-mounted system runs the principal component extraction rule matching (i.e., the second logical matching), it is necessary to initialize and load the dictionary file D, construct a prefix tree, and while generating the prefix tree, retain the corresponding type of the dictionary. Load the recall path file P to cache the mapping relationship between buttons and recall paths.
[0123] Segment the instruction text, and segment the input instruction text query into a word list at the word granularity; use the maximum reverse matching algorithm for this list to obtain all possible candidate principal component words (including: the candidate subject name, i.e., the button Subject, abbreviated as subj, and the candidate action, i.e., Operation, abbreviated as op) and their standard statements that appear in the query, that is, obtain the candidate subject name subj and the candidate action op. Filter and screen the candidate principal component words according to the prefix matching and suffix matching methods in combination with the badcase (wrong alias name, also called wrong name) of the candidate principal component words to improve the probability that the candidate words expected by the user are hit. The advantage of using the maximum reverse matching here is that when the configured generalization name is very fine, it can hit the specified word to the greatest extent.
[0124] According to the candidate main component words appearing in the query, construct a candidate recall path set U: Arrange the standard statements of the candidate words in sequence without overlapping in ascending order to obtain the first preset number of candidate recall paths u, where the first preset number represents multiple. This application has amplified the early stopping mechanism, that is, it is stipulated that in each candidate recall path u, only 1 subj and 1 op can appear. When multiple subjs or ops appear, it is determined that the candidate recall path does not meet the basic candidate recall path conditions. When multiple subjs or ops appear in a certain candidate recall path, its meaning is not equivalent to the "simulated touch" operation of "visible and speakable". Through the early stopping mechanism, the scope of action can be effectively restricted, preventing mis-triggering and avoiding affecting other in-vehicle functions.
[0125] For example: A certain candidate recall path is: Open the window and turn off the air conditioner. Here, "Open the window and turn off the air conditioner", after splitting, corresponds to "Open" (op), "window" (subj), "Turn off" (op), "air conditioner" (subj), with multiple ops and multiple subjs appearing, which no longer meets the specified rule of 1 subj and 1 op. So this combination of the recall path can be eliminated in advance, that is, directly delete this candidate recall path.
[0126] Match each candidate recall path u in the candidate recall path set U with the standard recall paths (i.e., preset recall paths) cached in the recall path file P. If the match is successful, the basic information of the button, such as the button identification ID, can be obtained, completing the hit and obtaining the second recall result. The second recall result includes the identification IDs of the hit buttons, such as button1, button2, etc. If the match is unsuccessful, the matching is exited.
[0127] In the embodiment of this application, the main component words in the instruction text are extracted through the maximum reverse matching algorithm and screened by the error name. Then, candidate recall paths are generated and matched with the preset recall paths to obtain the final recall result. The above method has a high accuracy rate, can recognize user statements in natural language, and screen the candidate main component words according to bad cases, avoiding mis-triggering other buttons.
[0128] As an alternative embodiment, obtaining the dictionary file and the recall path file includes:
[0129] Obtain the second preset button definitions of the second preset type of buttons, where the second preset button definitions include the second identification of each second preset type of button, the set of main body names, and the set of action actions. The set of main body names includes the alias names, standard names, and error names of the second preset type of buttons. The set of action actions includes the alias names, standard names, and error names of the actions. The error names of the second preset type of buttons and the error names of the actions are used to avoid matching errors, and the alias name is the correct name other than the standard name;
[0130] Construct a dictionary file according to the alias name, standard name, error name of the second preset type button, the alias name, standard name, error name of the action, and the preset order.
[0131] Generate the second preset number of preset recall paths corresponding to each second preset type button according to the second identifier, the set of subject names, and the set of action actions.
[0132] Construct a recall path file according to the preset recall path.
[0133] Optionally, according to the second preset button definition, each second preset type button can be flexibly configured. After the configuration is completed, 2 configuration files will be generated for it: the dictionary file D and the recall path file P.
[0134] The dictionary file D archives the set of button subject names of all second preset type buttons, such as Subject k and the set of action actions, such as Operations k for constructing the dictionary tree for principal component extraction. Among them, the content in the dictionary file D is arranged in the preset order. For example: each line is in turn: alias name|standard name|error name|type (subj / op), and the example is as follows:
[0135] Play once|Play|Cancel play|op.
[0136] Music|My Music|QQ Music, KuGou Music|subj.
[0137] The recall path file P combines the second preset type button ID (such as AssistName k ), the standard statement of the button (such as stdSubj k ), and N action actions (such as Ops i ) in positive and negative order combinations. In this way, each button can generate 2N lines of recall paths (i.e., preset recall paths), and each line of recall path can be expressed as follows:
[0138] button1 (Play|op) _ (My Music|subj)
[0139] button1 (My Music|subj) _ (Play|op)
[0140] Among them, button1 is the second identifier ID, and "Play" and "My Music" are the standard statements. All second preset type buttons can generate a total of the second preset number of preset recall paths, and no specific quantity limit is set here.
[0141] In the embodiment of the present application, according to the definition of the second preset button, each second preset type of button is flexibly configured to generate a dictionary file and a recall path file, providing a basis for subsequent matching of the principal component extraction rules.
[0142] As an alternative embodiment, obtaining the definition of the second preset button for the second preset type of button includes:
[0143] Obtaining the alias name, standard name, and error name of each second preset type of button;
[0144] Constructing a set of main body names for each second preset type of button according to the alias name, standard name, and error name of the second preset type of button;
[0145] Obtaining the third preset number of actions supported by each second preset type of button;
[0146] Obtaining the alias name, standard name, and error name of each action;
[0147] Constructing a set of behavioral actions for each second preset type of button according to the alias name, standard name, and error name of the action;
[0148] Obtaining the second identifier of each second preset type of button;
[0149] Generating the definition of the second preset button for each second preset type of button according to the second identifier, the set of main body names, and the set of behavioral actions.
[0150] Optionally, the second preset type of button is a button other than the above-mentioned first preset type of button. In terms of the definition of the second preset type of button, it is agreed that the second preset type of button k can be defined as a triple R k :
[0151] R k ={AssistName k , Subject k , Operations k}
[0152] Among them, AssistName k represents the second identifier ID of button k; Subject k represents the set of button main body names of button k; Operations k represents the set of behavioral actions of button k.
[0153] Furthermore, Subject k is defined as: one standard name stdSubj k , and one set aliasSubjs containing n generalized alias names (alias names)k , and a set badcaseSubjs containing m error alias names (also called error names) k . n and m represent multiple, and badcaseSubjs k is to avoid mis-hitting and mis-triggering unmatched buttons.
[0154] Subject k = {stdSubj k , aliasSubjs k , badcaseSubjs k}
[0155] aliasSubjs k = {aliasS1 k ,..., aliasS n k}
[0156] badcaseSubjs k = {badcaseS1 k ,..., badcaseS m k}
[0157] Among them, aliasS1 k is the first alias name of button k, and badcaseS1 k is the first error name of button k.
[0158] Although a button has only two states: pressed and not pressed, there are many natural language expressions for these two states. For example: The user can say "press once", "click", "tap", "open", "activate", etc. to represent wanting to press the button. Therefore, one button can support multiple groups of behavioral actions. Operations k is defined as containing N groups of behavioral actions:
[0159] Operations k = {Ops1,..., Ops N}
[0160] Among them, Ops N represents the Nth group of behavioral actions. Each group of behavioral actions contains multiple actions with the same meaning but different expressions. For example: Actions such as "open", "turn on", "start", "activate" all represent opening, and they are grouped into one group such as Ops1. Actions such as "play", "listen to", "want to listen" all represent playing, and they are grouped into another group of actions such as Ops N . Such grouping can better reuse and simplify button configuration.
[0161] The i-th group of behavioral actions Ops i includes 1 standard name stdOp i , 1 set aliasOps containing M alternative names (alias names) i and 1 set badcaseOps containing T incorrect alias names (also called error names) i The incorrect alias names are used to avoid mis-hitting and mis-triggering non-matching buttons. Ops i can be defined as:
[0162] Ops i = {stdOp i , aliasOps i , badcaseOps i}
[0163]
[0164]
[0165] Among them, is the 1st alternative name of the i-th group of behavioral actions Ops1, is the 1st error name of the i-th group of behavioral actions Ops1.
[0166] All buttons of the second preset type are selected and configured with the behavioral actions Ops supported by the current button from the same behavioral action pool i : During configuration, the actions Operation of all buttons are shared. For example: open, select, close, click, etc. Then, during configuration, according to the actual situation of the current button, select the actions supported by the current button from this shared action list; instead of each button having to define its own independent Operation separately. That is to say, the "open" configured in button A; and the "open" configured in button B are the same open, and the corresponding aliasOps i and badcaseOps i are the same set.
[0167] The definitions of other buttons of the second preset type are the same as that of button k, which will not be elaborated here.
[0168] In the embodiments of the present application, by configuring the Subject and Operation of the corresponding button, the generalization statements of the button can be supported, and the visual configuration can also greatly reduce the maintenance cost, update cost, and the learning cost of users.
[0169] As an alternative embodiment, candidate main component words are obtained based on a word list, a preset matching method, a dictionary file, and a prefix tree, including:
[0170] Obtain the error name of each second preset type button according to the dictionary file;
[0171] Obtain the hit words according to the word list, the preset matching method, the dictionary file, and the prefix tree;
[0172] Filter the hit words according to the error name to obtain the filtered hit words;
[0173] Obtain candidate main component words according to the filtered hit words.
[0174] Optionally, prefix matching and suffix matching are used to filter the bad cases (error aliases, also known as error names) of the second preset type buttons.
[0175] For example: The word segmentation result (word list) of the instruction text "Cancel playing today's recommendations" is: "Cancel" (0) "Play" (1) "Today" (2) "Recommendations" (3). The above four words are also hit words at the same time. The numbers in the parentheses represent the positions of the hit words. Through the maximum reverse matching, "Play" will be hit and the bad case of "Play", which is "Cancel play", will be obtained. Prefix matching will take the sentence from the position of the hit word to the end of the complete instruction text and match it with the bad case to see if it starts with the bad case "Cancel play". Here, for example, "Play today's recommendations" does not meet the condition of starting with "Cancel play", so suffix matching is performed. Match from the beginning of the complete instruction text to the end position of the hit word to see if it ends with the bad case. Here, for example, "Cancel play" is found to meet the condition, so the hit word "Play" this time will be filtered out. Finally, candidate main component words are obtained according to the filtered hit words.
[0176] In the embodiment of the present application, filtering the hit words according to the bad cases to obtain candidate main component words avoids mis-triggering other buttons.
[0177] As an alternative embodiment, the first recall result includes the length of the hit words. According to the first recall result and the second recall result, the target button is obtained, including:
[0178] Obtain the third identifier of the button element in the current interface;
[0179] Obtain the fourth identifier of the first intermediate button according to the first recall result and the length of the hit words;
[0180] Obtain the target button according to the third identifier and the fourth identifier;
[0181] According to the second recall result, obtain the fifth identifier of the second intermediate button;
[0182] According to the third identifier and the fifth identifier, obtain the target button.
[0183] Optionally, for example, there are the following button elements in the current interface: menu bar buttons such as the return key, Home key, menu key, etc., and each key has an identifier ID, that is, the third identifier. Then, match the identifier ID of the intermediate button in the recall result with the third identifier, that is, determine whether there is an intermediate button in the recall result in the current interface. If so, this intermediate button is the target button.
[0184] According to the hit word length in the first recall result, determine the first intermediate button and obtain the fourth identifier of the first intermediate button. Some buttons may be subsets of another button. For example, button A is "blow the face" and button B is "blow the face and feet"; instruction text query1: "turn on the face blowing mode" and instruction text query2: "turn on the face and feet blowing mode", their intentions and speech fields are the same. When the user says query2, the semantic slots are recognized as 2: "blow the face" and "blow the feet". Since "blow the face" is included, when button B is triggered, A will also be triggered. Therefore, the hit word length comparison is adopted here. Each time a semantic slot Slots is hit, its length is increased by 1, and finally the most suitable one or more buttons are selected for distribution, that is, the first intermediate button is determined. Match the fourth identifier ID of the first intermediate button in the first recall result with the third identifier. If the match is successful, this first intermediate button is the target button. Match the fifth identifier ID of the second intermediate button in the second recall result with the third identifier. If the match is successful, this second intermediate button is the target button.
[0185] In the embodiment of the present application, by matching the identifier of the intermediate button in the recall result with the identifier of the button element existing in the current interface, the accuracy of the hit button is further improved, and a basis is provided for subsequent simulating a click on the target button.
[0186] As an alternative embodiment, Figure 4 It is a schematic flowchart of another alternative second logical matching according to the embodiment of the present application. The process includes:
[0187] User query; word segmentation; maximum reverse matching to obtain candidate recall words; in positive order, non-overlapping construction of candidate recall paths; determine whether it satisfies 1subj + 1ops. If not, trigger the early stop mechanism. If so, form a set U of candidate recall paths; compare each u with the standard path file P; determine whether the path u matches. If not, compare the next path. If it matches, obtain the hit button information; return the recall matching result.
[0188] Optionally, for the specific implementation manners of this application, refer to other embodiments of this application, which will not be elaborated here.
[0189] In the embodiments of this application, only by configuring the Subject and Operation of the corresponding button can the generalization statement of the button be supported. The visual configuration also greatly reduces the maintenance cost, update cost, and user learning cost.
[0190] According to another aspect of the embodiments of this application, there is also provided a button matching device according to voice for implementing the above method for matching buttons according to voice. Figure 5 It is a structural block diagram of an optional button matching device according to voice of the embodiments of this application, as Figure 5 shown. The device may include:
[0191] A conversion module 501, configured to convert the user's voice information into an instruction text;
[0192] A first matching module 502, configured to perform character matching on the instruction text to determine whether a target button can be obtained;
[0193] A second matching module 503, configured to perform a first logical matching on the instruction text to obtain a first recall result in the case where the target button cannot be obtained;
[0194] A third matching module 504, configured to perform a second logical matching on the instruction text to obtain a second recall result;
[0195] An obtaining module 505, configured to obtain a target button according to the first recall result and the second recall result.
[0196] Through the above modules, three different matching methods are used to match the instruction text, which improves the matching success rate for various buttons, so as to meet different user requirements and release the rule definition limitations in the traditional technology. By using three different matching methods, the natural language of the user can be recognized. At the same time, through the configuration of the button, the generalization statement is more convenient, fast, flexible and changeable, reducing the maintenance cost, update cost and user learning cost of this method. It solves the problem that there are rule definition limitations in the related technology and the user voice under the natural expression cannot be recognized.
[0197] As an optional embodiment, the second matching module includes:
[0198] An input unit, configured to input the instruction text into a preset model to obtain the semantic slot, speech domain, and semantic intention corresponding to the instruction text;
[0199] A first acquisition unit, configured to acquire a first preset button definition of a first preset type of button, where the first preset button definition includes a semantic slot rule, a speech domain rule, and a semantic intention rule for each first preset type of button;
[0200] A first obtaining unit, configured to obtain a first recall result when the semantic slot, speech domain, and semantic intention corresponding to the instruction text respectively match the semantic slot rule, speech domain rule, and semantic intention rule of the same first preset type of button.
[0201] As an optional embodiment, the first acquisition unit includes:
[0202] A first acquisition sub-module, configured to acquire the correct semantic slot type and name, and the incorrect semantic slot type and name corresponding to each first preset type of button, where the incorrect semantic slot type and name are used to avoid incorrect matching;
[0203] A first generation sub-module, configured to generate a semantic slot rule according to the correct semantic slot type and name, and the incorrect semantic slot type and name;
[0204] A second acquisition sub-module, configured to acquire the speech domain rule and semantic intention rule corresponding to each first preset type of button;
[0205] A third acquisition sub-module, configured to acquire a first identifier of each first preset type of button;
[0206] A second generation sub-module, configured to generate a first preset button definition for each first preset type of button according to the first identifier, semantic slot rule, speech domain rule, and semantic intention rule.
[0207] As an optional embodiment, the third matching module includes:
[0208] A second acquisition unit, configured to acquire a dictionary file and a recall path file;
[0209] A construction unit, configured to construct a prefix tree according to the dictionary file;
[0210] A second obtaining unit, configured to perform word segmentation on the instruction text to obtain a word list;
[0211] A third obtaining unit, configured to obtain candidate main component words according to the word list, a preset matching method, the dictionary file, and the prefix tree, where the candidate main component words include a candidate main body name and a candidate action of the button;
[0212] A generation unit, configured to generate a candidate recall path set according to the candidate main component words, where the candidate recall path set includes a first preset number of candidate recall paths;
[0213] A judgment unit, configured to judge whether each candidate recall path contains only one candidate subject name and one candidate action;
[0214] A deletion unit, configured to delete the candidate recall path when the candidate recall path does not contain only one candidate subject name and one candidate action;
[0215] A fourth obtaining unit, configured to, when all candidate recall paths contain only one candidate subject name and one candidate action, match each candidate recall path in the candidate recall path set with a preset recall path in the recall path file to obtain a second recall result.
[0216] As an optional embodiment, the second obtaining unit includes:
[0217] A fourth obtaining sub-module, configured to obtain a second preset button definition of a second preset type of button, where the second preset button definition includes a second identifier, a set of subject names, and a set of action actions of each second preset type of button. The set of subject names includes an alias name, a standard name, and an error name of the second preset type of button. The set of action actions includes an alias name, a standard name, and an error name of the action. The error name of the second preset type of button and the error name of the action are used to avoid matching errors, and the alias name is a correct name other than the standard name;
[0218] A first constructing sub-module, configured to construct a dictionary file according to the alias name, standard name, error name of the second preset type of button, the alias name, standard name, error name of the action, and a preset order;
[0219] A third generating sub-module, configured to generate a second preset number of preset recall paths corresponding to each second preset type of button according to the second identifier, the set of subject names, and the set of action actions;
[0220] A second constructing sub-module, configured to construct a recall path file according to the preset recall path.
[0221] As an optional embodiment, the fourth obtaining sub-module includes:
[0222] A first obtaining sub-unit, configured to obtain the alias name, standard name, and error name of each second preset type of button;
[0223] A first constructing sub-unit, configured to construct a set of subject names of each second preset type of button according to the alias name, standard name, and error name of the second preset type of button;
[0224] A second obtaining sub-unit, configured to obtain a third preset number of actions supported by each second preset type of button;
[0225] A third obtaining sub-unit, configured to obtain the alias name, standard name, and error name of each action;
[0226] A second construction subunit, configured to construct a set of behavioral actions for each second preset type of button according to the alias name, standard name, and error name of the action;
[0227] A fourth acquisition subunit, configured to acquire a second identifier for each second preset type of button;
[0228] A generation subunit, configured to generate a second preset button definition for each second preset type of button according to the second identifier, the set of main body names, and the set of behavioral actions.
[0229] As an optional embodiment, the third obtaining unit includes:
[0230] A fifth acquisition sub-module, configured to acquire the error name for each second preset type of button according to the dictionary file;
[0231] A first obtaining sub-module, configured to obtain the hit words according to the word list, the preset matching method, the dictionary file, and the prefix tree;
[0232] A screening sub-module, configured to screen the hit words according to the error name to obtain the screened hit words;
[0233] A second obtaining sub-module, configured to obtain the candidate main component words according to the screened hit words.
[0234] As an optional embodiment, the first recall result includes the length of the hit words, and the obtaining module includes:
[0235] A third acquisition unit, configured to acquire a third identifier of a button element in the current interface;
[0236] A fifth obtaining unit, configured to obtain a fourth identifier of a first intermediate button according to the first recall result and the length of the hit words;
[0237] A sixth obtaining unit, configured to obtain a target button according to the third identifier and the fourth identifier;
[0238] A seventh obtaining unit, configured to obtain a fifth identifier of a second intermediate button according to the second recall result;
[0239] An eighth obtaining unit, configured to obtain a target button according to the third identifier and the fifth identifier.
[0240] It should be noted here that the examples and application scenarios implemented by the above modules and the corresponding steps are the same, but are not limited to the content disclosed in the above embodiments.
[0241] According to another aspect of the embodiments of the present application, there is also provided an electronic device for implementing the above method for matching buttons according to voice. The electronic device may be a server, a terminal, or a combination thereof.
[0242] Figure 6 is a structural block diagram of an optional electronic device according to an embodiment of the present application. As shown in Figure 6 , it includes a processor 601, a communication interface 602, a memory 603, and a communication bus 604. Among them, the processor 601, the communication interface 602, and the memory 603 complete mutual communication through the communication bus 604. Among them,
[0243] The memory 603 is used to store computer programs;
[0244] The processor 601, when executing the computer program stored on the memory 603, realizes the following steps:
[0245] Convert the user's voice information into instruction text;
[0246] Perform character matching on the instruction text to determine whether a target button can be obtained;
[0247] In the case where the target button cannot be obtained, perform a first logical matching on the instruction text to obtain a first recall result;
[0248] Perform a second logical matching on the instruction text to obtain a second recall result;
[0249] Obtain the target button according to the first recall result and the second recall result.
[0250] Optionally, in this embodiment, the above-mentioned communication bus may be a PCI (Peripheral Component Interconnect) bus, or an EISA (Extended Industry Standard Architecture) bus, etc. This communication bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of simplicity of representation, Figure 6 only a thick line is used to represent it in the figure, but it does not mean that there is only one bus or one type of bus.
[0251] The communication interface is used for communication between the above-mentioned electronic device and other devices.
[0252] The memory may include a RAM, and may also include a non-volatile memory, for example, at least one disk memory. Optionally, the memory may also be at least one storage device located far from the aforementioned processor.
[0253] As an example, as shown in Figure 6As shown, the above-mentioned memory 603 may include, but is not limited to, the conversion module 501, the first matching module 502, the second matching module 503, the third matching module 504, and the obtaining module 505 in the above-mentioned device for matching buttons according to voice. In addition, it may also include, but is not limited to, other module units in the above-mentioned device for matching buttons according to voice, which will not be elaborated in this example.
[0254] The above-mentioned processor may be a general-purpose processor, which may include, but is not limited to: CPU (Central Processing Unit, central processing unit), NP (Network Processor, network processor), etc.; it may also be a DSP (Digital Signal Processing, digital signal processor), ASIC (Application Specific Integrated Circuit, application-specific integrated circuit), FPGA (Field-Programmable Gate Array, field programmable gate array) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.
[0255] Optionally, the specific examples in this embodiment may refer to the examples described in the above-mentioned embodiment, and will not be elaborated here.
[0256] Those of ordinary skill in the art can understand that Figure 6 the structure shown is only schematic. The device for implementing the method of matching buttons according to voice may be a terminal device, which may be a smart phone (such as an Android phone, an iOS phone, etc.), a tablet computer, a palm computer, and a mobile Internet device (Mobile Internet Devices, MID), a PAD and other terminal devices. Figure 6 It does not limit the structure of the above-mentioned electronic device. For example, the terminal device may further include more or fewer components (such as a network interface, a display device, etc.) than those shown in Figure 6 and may have a different configuration from that shown in Figure 6 the figure.
[0257] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above-mentioned embodiments can be completed by instructing the relevant hardware of the terminal device through a program, and the program can be stored in a computer-readable storage medium. The storage medium may include: a flash drive, a ROM, a RAM, a magnetic disk or an optical disc, etc.
[0258] According to another aspect of the embodiments of the present application, a storage medium is further provided. Optionally, in this embodiment, the above-mentioned storage medium may be used to store the program code for executing the method of matching buttons according to voice.
[0259] Optionally, in this embodiment, the above storage medium may be located on at least one of multiple network devices in the network shown in the above embodiment.
[0260] Optionally, in this embodiment, the storage medium is configured to store program code for performing the following steps:
[0261] Convert the user's voice information into instruction text;
[0262] Perform character matching on the instruction text to determine whether a target button can be obtained;
[0263] In the case where the target button cannot be obtained, perform a first logical match on the instruction text to obtain a first recall result;
[0264] Perform a second logical match on the instruction text to obtain a second recall result;
[0265] Obtain the target button according to the first recall result and the second recall result.
[0266] Optionally, specific examples in this embodiment may refer to the examples described in the above embodiment, and details are not described herein again.
[0267] Optionally, in this embodiment, the above storage medium may include but is not limited to: various media that can store program code such as USB flash drives, ROMs, RAMs, mobile hard disks, magnetic disks, or optical discs.
[0268] In the description of this specification, the description with reference to terms such as "this embodiment", "one embodiment", "some embodiments", "examples", "specific examples", or "some examples" means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present disclosure. In this specification, the schematic descriptions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described may be combined in a suitable manner in any one or more embodiments or examples. In addition, without contradiction, those skilled in the art may combine and combine different embodiments or examples described in this specification and the features of different embodiments or examples. In the description of the present disclosure, "a plurality" means at least two, such as two, three, etc., unless otherwise specifically defined.
[0269] Obviously, the above embodiments are merely examples for clear illustration and not limitations on the implementation manners. For those of ordinary skill in the art, other different forms of changes or modifications can be made based on the above description. It is not necessary and impossible to enumerate all implementation manners here. And the obvious changes or modifications derived therefrom still fall within the protection scope of the present invention.
Claims
1. A method for matching buttons according to voice, characterized in that The method includes: Converting the user's voice information into an instruction text; Performing character matching on the instruction text to determine whether a target button can be obtained; In the case where the target button cannot be obtained, performing a first logical matching on the instruction text to obtain a first recall result; Performing a second logical matching on the instruction text to obtain a second recall result; Obtaining the target button according to the first recall result and the second recall result; Among them, the performing a first logical matching on the instruction text to obtain a first recall result includes: inputting the instruction text into a preset model to obtain the semantic slot, speech domain, and semantic intention corresponding to the instruction text; obtaining the first preset button definition of the first preset type of buttons, where the first preset button definition includes the semantic slot rule, speech domain rule, and semantic intention rule of each first preset type of button; obtaining the first recall result when the semantic slot, speech domain, and semantic intention corresponding to the instruction text respectively match the semantic slot rule, the speech domain rule, and the semantic intention rule of the same first preset type of button; Among them, the performing a second logical matching on the instruction text to obtain a second recall result includes: obtaining a dictionary file and a recall path file; constructing a prefix tree according to the dictionary file; performing word segmentation processing on the instruction text to obtain a word list; obtaining candidate main component words according to the word list, a preset matching method, the dictionary file, and the prefix tree, where the candidate main component words include the candidate main body name and candidate action of the button; generating a candidate recall path set according to the candidate main component words, where the candidate recall path set includes a first preset number of candidate recall paths; determining whether each candidate recall path only includes one candidate main body name and one candidate action; deleting the candidate recall path when the candidate recall path does not only include one candidate main body name and one candidate action; obtaining the second recall result when all the candidate recall paths only include one candidate main body name and one candidate action, and respectively matching each candidate recall path in the candidate recall path set with a preset recall path in the recall path file; Among them, the first recall result includes the length of the hit word, and the obtaining the target button according to the first recall result and the second recall result includes: obtaining the third identifier of the button element in the current interface; obtaining the fourth identifier of the first intermediate button according to the first recall result and the length of the hit word; obtaining the target button according to the third identifier and the fourth identifier; obtaining the fifth identifier of the second intermediate button according to the second recall result; obtaining the target button according to the third identifier and the fifth identifier.
2. The method of matching buttons according to voice as claimed in claim 1, wherein The obtaining the first preset button definition of the first preset type of buttons includes: Obtain the correct semantic slot types and names, as well as the incorrect semantic slot types and names corresponding to each of the first preset type of buttons, where the incorrect semantic slot types and names are used to avoid matching errors; Generate the semantic slot rules based on the correct semantic slot types and names, and the incorrect semantic slot types and names; Obtain the speech domain rules and the semantic intent rules corresponding to each of the first preset type of buttons; Obtain the first identifier of each of the first preset type of buttons; Generate the first preset button definition for each of the first preset type of buttons according to the first identifier, the semantic slot rules, the speech domain rules, and the semantic intent rules; 3. The method of matching buttons according to voice as claimed in claim 1, wherein, The obtaining of the dictionary file and the recall path file includes: Obtain the second preset button definition of the second preset type of buttons, where the second preset button definition includes the second identifier, the set of main body names, and the set of action actions of each of the second preset type of buttons. The set of main body names includes the alias names, standard names, and incorrect names of the second preset type of buttons. The set of action actions includes the alias names, standard names, and incorrect names of the actions. The incorrect names of the second preset type of buttons and the incorrect names of the actions are used to avoid matching errors. The alias name is the correct name other than the standard name; Construct a dictionary file according to the alias names, standard names, incorrect names of the second preset type of buttons, the alias names, standard names, incorrect names of the actions, and the preset order; Generate the second preset number of preset recall paths corresponding to each of the second preset type of buttons according to the second identifier, the set of main body names, and the set of action actions; Construct a recall path file according to the preset recall paths; 4. The method of matching buttons according to voice as claimed in claim 3, wherein The obtaining of the second preset button definition of the second preset type of buttons includes: Obtain the alias name, standard name, and incorrect name of each of the second preset type of buttons; Construct the set of main body names of each of the second preset type of buttons according to the alias name, standard name, and incorrect name of the second preset type of buttons; Obtain the third preset number of actions supported by each of the second preset type of buttons; Obtain the alias name, standard name, and incorrect name of each of the actions; Construct the set of action actions of each of the second preset type of buttons according to the alias name, standard name, and incorrect name of the actions; Obtain the second identifier of each of the second preset type of buttons; Generate the second preset button definition for each of the second preset type of buttons according to the second identifier, the set of main body names, and the set of action actions; 5. The method for matching a button according to speech as claimed in claim 3, wherein, The obtaining of the candidate main component words according to the word list, the preset matching method, the dictionary file, and the prefix tree includes: Obtain the incorrect name of each of the second preset type of buttons according to the dictionary file; Obtain the hit words according to the word list, the preset matching method, the dictionary file, and the prefix tree; Filter the hit words according to the incorrect name to obtain the filtered hit words; Obtain the candidate main component words according to the filtered hit words; 6. A device for matching buttons according to voice, characterized in that, Includes: A conversion module for converting the user's voice information into an instruction text; The first matching module is used to perform character matching on the instruction text to determine whether a target button can be obtained; The second matching module is used to perform a first logical matching on the instruction text to obtain a first recall result when the target button cannot be obtained; The third matching module is used to perform a second logical matching on the instruction text to obtain a second recall result; The obtaining module is used to obtain the target button according to the first recall result and the second recall result; Among them, the second matching module includes: The input unit is used to input the instruction text into a preset model to obtain the semantic slot, the speech field, and the semantic intention corresponding to the instruction text; The first obtaining unit is used to obtain the first preset button definition of the first preset type of buttons, where the first preset button definition includes the semantic slot rule, the speech field rule, and the semantic intention rule of each first preset type of button; The first obtaining unit is used to obtain a first recall result when the semantic slot, the speech field, and the semantic intention corresponding to the instruction text respectively match successfully with the semantic slot rule, the speech field rule, and the semantic intention rule of the same first preset type of button; The third matching module includes: The second obtaining unit is used to obtain a dictionary file and a recall path file; The building unit is used to build a prefix tree according to the dictionary file; The second obtaining unit is used to perform word segmentation on the instruction text to obtain a word list; The third obtaining unit is used to obtain candidate main component words according to the word list, a preset matching method, the dictionary file, and the prefix tree, where the candidate main component words include the candidate main body name and the candidate action of the button; The generating unit is used to generate a candidate recall path set according to the candidate main component words, where the candidate recall path set includes a first preset number of candidate recall paths; The judging unit is used to judge whether each candidate recall path only includes one candidate main body name and one candidate action; The deleting unit is used to delete the candidate recall path when the candidate recall path does not only include one candidate main body name and one candidate action; The fourth obtaining unit is used to match each candidate recall path in the candidate recall path set with a preset recall path in the recall path file to obtain the second recall result when all the candidate recall paths only include one candidate main body name and one candidate action; The first recall result includes the hit word length, and the obtaining module includes: The third obtaining unit is used to obtain the third identifier of the button element on the current interface; The fifth obtaining unit is used to obtain the fourth identifier of the first intermediate button according to the first recall result and the hit word length; The sixth obtaining unit is used to obtain the target button according to the third identifier and the fourth identifier; The seventh obtaining unit is used to obtain the fifth identifier of the second intermediate button according to the second recall result; The eighth obtaining unit is used to obtain the target button according to the third identifier and the fifth identifier.
7. An electronic device, comprising a processor, a communication interface, a memory, and a communication bus, wherein, The processor, the communication interface, and the memory complete their mutual communication through the communication bus, characterized in that the memory is used for storing a computer program; the processor is used for executing the method according to the voice matching button described in any one of claims 1 to 5 by running the computer program stored on the memory.
Citation Information
Patent Citations
Recognition information determination method, device and equipment for target object, and storage medium
CN110781204A
Voice control method, device and equipment and computer storage medium
CN114067797A