Voice Control Method, Device, Electronic Device and Storage Medium

By identifying and executing the trigger condition information of voice commands, the problem that non-instant voice commands cannot be executed accurately in the prior art is solved, and the user experience is improved.

CN114121005BActive Publication Date: 2025-08-05GUANGDONG OPPO MOBILE TELECOMMUNICATIONS CORP LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111433111.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-11-29
Publication Date
2025-08-05
Estimated Expiration
2041-11-29

AI Technical Summary

Technical Problem

In the prior art, voice control methods cannot effectively support non-instant type voice commands, resulting in poor user experience.

Method used

By identifying the instruction type of the voice command, the trigger condition information and target execution information corresponding to the voice command of the non-instant type are obtained, and the corresponding operation interface operation is performed when the trigger condition is met.

Benefits of technology

It realizes accurate execution of non-instant type voice commands and improves user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114121005B_ABST
    Figure CN114121005B_ABST
Patent Text Reader

Abstract

This application discloses a voice control method, device, electronic device, and storage medium. The voice control method includes: obtaining a voice command; identifying the command type of the voice command; when the command type of the voice command is non-immediate, obtaining trigger condition information and target execution information corresponding to the voice command; and when the trigger condition corresponding to the trigger condition information is met, executing a target operation on a target operation interface corresponding to the target execution information. This method can implement non-immediate commands when controlling a graphical interface through voice, thereby improving the user experience.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of electronic devices, and more specifically, to a voice control method, device, electronic device, and storage medium. Background Art

[0002] With the rapid advancement of science and technology, voice recognition and natural language processing technologies can be combined to enable electronic devices to receive voice commands issued by users through auditory modalities and complete corresponding interactive tasks. As a result, users can complete interface interactive operations through voice input. However, in some cases, users may need to perform corresponding interface operations only when certain conditions are met, but the relevant technologies cannot effectively complete this type of non-instant triggering instructions, affecting the user experience. Summary of the Invention

[0003] In view of the above problems, the present application proposes a voice control method, device, electronic device and storage medium.

[0004] In a first aspect, an embodiment of the present application provides a voice control method, which includes: obtaining a voice instruction; identifying the instruction type of the voice instruction; when the instruction type of the voice instruction is a non-immediate type, obtaining trigger condition information and target execution information corresponding to the voice instruction; when the trigger condition corresponding to the trigger condition information is met, executing the target operation on the target operation interface corresponding to the target execution information.

[0005] In the second aspect, an embodiment of the present application provides a voice control device, which includes: an instruction acquisition module, an information acquisition module and an operation execution module, wherein the instruction acquisition module is used to acquire voice instructions; the information acquisition module is used to acquire trigger condition information and target execution information corresponding to the voice instructions when the instruction type of the voice instructions is a non-immediate type; the operation execution module is used to execute the target operation on the target operation interface corresponding to the target execution information when the trigger condition corresponding to the trigger condition information is met.

[0006] In a third aspect, an embodiment of the present application provides an electronic device comprising: one or more processors; a memory; and one or more applications, wherein the one or more applications are stored in the memory and configured to be executed by the one or more processors, and the one or more programs are configured to execute the voice control method provided in the first aspect above.

[0007] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, in which a program code is stored. The program code can be called by a processor to execute the voice control method provided in the first aspect above.

[0008] In a fifth aspect, an embodiment of the present application provides a computer program product, including a computer program / instruction, characterized in that when the computer program / instruction is executed by a processor, the voice control method provided in the first aspect above is implemented.

[0009] The solution provided by this application obtains a voice command. When the command type of the voice command is non-instantaneous, the trigger condition information and target execution information corresponding to the voice command are obtained. When the trigger condition corresponding to the trigger condition information is met, the target operation is executed on the target operation interface corresponding to the target execution information. In this way, it is possible to identify the trigger condition information of a non-instantaneous voice command input by the user, and then execute the corresponding interface operation based on the trigger condition information, thereby better completing the non-instantaneous voice command and improving the user experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For those skilled in the art, other drawings can be obtained based on these drawings without creative work.

[0011] Figure 1 A schematic diagram of a scenario provided in an embodiment of the present application is shown.

[0012] Figure 2 Another scenario schematic diagram provided by an embodiment of the present application is shown.

[0013] Figure 3 A schematic diagram of an application environment provided by an embodiment of the present application is shown.

[0014] Figure 4 Another schematic diagram of the application environment provided by the embodiment of the present application is shown.

[0015] Figure 5 A flow chart of a voice control method according to an embodiment of the present application is shown.

[0016] Figure 6 Another scenario schematic diagram provided by an embodiment of the present application is shown.

[0017] Figure 7 A flow chart of a voice control method according to another embodiment of the present application is shown.

[0018] Figure 8 A flow chart of a voice control method according to another embodiment of the present application is shown.

[0019] Figure 9A flow chart of a voice control method according to another embodiment of the present application is shown.

[0020] Figure 10 A schematic diagram showing the principles of the instruction recognition model provided in an embodiment of the present application is shown.

[0021] Figure 11 A flow chart of a voice control method according to yet another embodiment of the present application is shown.

[0022] Figure 12 A structural diagram of the instruction type recognition model provided in an embodiment of the present application is shown.

[0023] Figure 13 A flow chart of a voice control method according to yet another embodiment of the present application is shown.

[0024] Figure 14 A block diagram of a voice control device according to an embodiment of the present application is shown.

[0025] Figure 15 4 is a block diagram of an electronic device for executing a voice control method according to an embodiment of the present application.

[0026] Figure 16 It is a storage unit of an embodiment of the present application for storing or carrying program codes for implementing the voice control method according to an embodiment of the present application. DETAILED DESCRIPTION

[0027] In order to enable people skilled in the art to better understand the solution of the present application, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application.

[0028] The popularity of smart terminal devices has brought various conveniences to our lives. At the beginning of the birth of smart terminal devices, GUI (Graphic User Interface) has always been an important carrier for users to interact with smartphones. GUI can also be generally referred to as UI. Today, with the continuous development of voice interaction, using more convenient intelligent voice interaction to interact with smart terminal devices has become an important means of human-computer interaction. VGUI (Voice Graphic User Interface) can provide users with a more convenient and direct means of service interaction, or provide users with a seamless interaction solution in obstructed situations when it is inconvenient for users to touch the GUI.

[0029] However, after a long period of research, the inventors found that in the related art, the VGUI solution focuses on the immediate execution of user instructions and cannot provide good technical support for non-immediate instructions. Figure 1 The figure shows a scenario of voice control. The user inputs "Send a text message to mom" through voice. After parsing the voice message, the smart terminal device searches for the sending application on the GUI interface and then executes the operation to send the corresponding text message. The user's command is executed immediately. Figure 2 In the voice control scenario shown, the user inputs "When I get home, send a text message to Mom" through voice. At this time, the user adds a trigger condition "When I get home" based on the expression, turning the command into a non-instantaneous voice control graphical interface command that is triggered only when the condition is met. However, the smart terminal may not execute the command until the trigger condition is met, but execute it immediately. Therefore, it is impossible to implement non-instantaneous voice commands well, which brings inconvenience to the user.

[0030] To address the above issues, the inventors have proposed the voice control method, device, electronic device, and storage medium provided in the embodiments of this application. These methods can identify triggering condition information for non-immediate voice commands input by the user and then execute corresponding interface operations based on the triggering condition information, thereby effectively completing non-immediate voice commands and improving the user experience. The specific voice control method is described in detail in the subsequent embodiments.

[0031] The following first introduces the application scenarios involved in the embodiments of this application.

[0032] In the embodiment of the present application, the voice control method provided by the embodiment of the present application can be executed by an electronic device. In this way, all steps in the voice control method provided by the embodiment of the present application can be executed by the electronic device. For example, Figure 3 As shown, the voice collection device of the electronic device 100 can collect voice commands, and then transmit the collected voice commands and the current user interface to the processor, so that the processor identifies the instruction type of the voice command and then executes the steps involved in the voice control method provided in this application according to the identified instruction type.

[0033] Furthermore, the voice control method provided in the embodiments of the present application can also be executed by a server (cloud). Correspondingly, in this server-based execution mode, the electronic device can collect voice commands and send the collected voice commands and the current user interface to the server simultaneously. After the server recognizes the voice commands, the server triggers the electronic device to perform the target operation.

[0034] In addition, the electronic device and the server may collaborate to perform the operation. In this manner, some steps of the voice control method provided in the embodiment of the present application are performed by the electronic device, while other steps are performed by the server.

[0035] For example, Figure 4 As shown, the electronic device 100 can obtain a voice instruction, and then hand the voice instruction over to the server 200 to identify the instruction type of the voice instruction, and when the instruction type is a non-immediate type, identify the target execution information corresponding to the voice instruction and the trigger condition information corresponding to the target execution information, and then return it to the electronic device 100; the electronic device 100 then executes the target operation on the target interface corresponding to the target execution information according to the trigger condition information.

[0036] It should be noted that in this method of collaborative execution by the electronic device and the server, the steps respectively executed by the electronic device and the server are not limited to the methods introduced in the above examples. In actual applications, the steps respectively executed by the electronic device and the server can be dynamically adjusted according to actual conditions.

[0037] It should be noted that the electronic device 100 can be used for Figure 1 and Figure 2 In addition to the smartphone shown in FIG, it can also be a car device, wearable device, tablet computer, laptop computer, smart speaker, etc. The server 120 can be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server, etc., which is not limited here.

[0038] The voice control method provided in the embodiments of the present application is described in detail below with reference to the accompanying drawings.

[0039] See also Figure 5 , Figure 5 FIG. 1 is a flow chart showing a voice control method according to an embodiment of the present application. In a specific embodiment, the voice control method is applied to Figure 14 The voice control device 400 and the electronic device 100 ( Figure 15 ). The following will take electronic devices as an example to illustrate the specific process of this embodiment. Of course, it can be understood that the electronic devices used in this embodiment can be smart phones, tablet computers, smart watches, smart glasses, laptop computers, etc., which are not limited here. Figure 5 The process shown in FIG. 1 is described in detail. The voice control method may specifically include the following steps:

[0040] Step S110: Acquire voice instructions.

[0041] In an embodiment of the present application, a user can express his or her control intention by inputting voice into an electronic device. Correspondingly, the electronic device can use the user's voice as a voice command. The electronic device can collect the user's voice input through an audio acquisition device to obtain the voice command. The audio acquisition device is used to collect audio signals. Optionally, the audio acquisition device may include one or more audio acquisition devices, which may be microphones.

[0042] In some embodiments, the electronic device can detect the voice command input by the user when the voice control graphical interface function is turned on, and execute the steps involved in the voice control method provided in this application according to the detected voice command. Optionally, the electronic device can collect the voice command input by the user through a voice collection device when the voice assistant is turned on. For example, when the voice assistant of the electronic device is turned on, the user enters the voice command "Xiao Ou, help me turn on the NFC access control when I get home", the electronic device can collect the voice command.

[0043] In other embodiments, the electronic device may collect the voice input by the user when a voice control trigger operation is detected, thereby obtaining the voice command input by the user. Optionally, a control for a voice control graphical interface may be displayed on the screen of the electronic device, and when a corresponding operation on the control is detected, voice collection may be turned on to obtain the voice command input by the user. The operation may be a click operation, a press operation, a slide operation, etc., which are not limited here. Optionally, the electronic device may also collect the voice input by the user when an operation of a specified physical button is detected, thereby obtaining the voice command input by the user. Of course, the specific way in which the electronic device obtains the voice command may not be limited.

[0044] Step S120: When the command type of the voice command is a non-immediate type, the trigger condition information and target execution information corresponding to the voice command are obtained.

[0045] In an embodiment of the present application, after receiving a voice command, the electronic device can identify the command type of the voice command in order to identify the non-immediate control interface required by the user, thereby accurately implementing the control required by the user. The command type of the voice command can include an immediate type and a non-immediate type. The immediate type refers to the type of command that the electronic device needs to execute immediately after receiving the voice command; the non-immediate type refers to the type of command that the electronic device does not execute immediately after receiving the voice command, but executes only when corresponding conditions are met.

[0046] In some embodiments, the electronic device may recognize the acquired voice command and obtain the text content corresponding to the voice command. After obtaining the text content corresponding to the voice command, the electronic device may perform semantic recognition on the text content based on a pre-configured method to identify whether the voice command is of an immediate type or a non-immediate type. The electronic device may convert the voice command into the corresponding text content based on a pre-configured automatic speech recognition method.

[0047] As a possible implementation, the electronic device can identify the text content corresponding to the voice command based on a pre-trained command type recognition model, thereby obtaining the command type corresponding to the voice command. The command type recognition model can be trained based on text sample data labeled with command types.

[0048] In other embodiments, after obtaining the text content corresponding to the voice instruction, the electronic device may also segment the text content, and then obtain the keywords in the text content based on the segmentation results; the electronic device then matches the identified keywords with preset keywords, which are keywords corresponding to the pre-set non-immediate type of voice instructions; if any of the identified keywords matches the preset keywords, the instruction type of the voice instruction is determined to be a non-immediate type; if each identified keyword does not match the preset keyword, the instruction type of the voice instruction is determined to be an immediate type. Optionally, the preset keywords include keywords related to the trigger conditions, such as "at", "when", "if", "if", "when", "when", etc. Exemplarily, the text content corresponding to the voice instruction input by the user is "When the battery power is 10%, turn off the mobile data network", then the keyword "when" in the text content matches the preset keyword, and therefore the instruction type of the voice instruction is determined to be a non-immediate type.

[0049] Of course, the specific manner in which the electronic device recognizes the instruction type of the voice instruction is not limited.

[0050] In an embodiment of the present application, after the electronic device identifies the instruction type of the voice instruction, it can execute subsequent voice control steps according to the instruction type of the voice instruction. The electronic device can determine whether the instruction type of the voice instruction is a non-immediate type; if the instruction type of the voice instruction is a non-immediate type, it can identify the trigger condition information and target execution information corresponding to the voice instruction, so as to complete the voice control required by the user according to the trigger condition information and the target execution information. Among them, the target execution information can be understood as the control information obtained after the electronic device converts the voice instruction, which is used to characterize the user's control intention for the interface; the trigger condition information can be understood as the execution condition corresponding to the target execution information, that is, when what condition is met, the operation corresponding to the control information is executed, thereby completing the control required by the user.

[0051] In some implementations, the electronic device may obtain an instruction text corresponding to the voice instruction, and then obtain trigger condition information and target execution information contained in the instruction text.

[0052] As a possible implementation, the electronic device can perform semantic recognition on the text content obtained based on the conversion of the voice command using a preconfigured method, and then determine the trigger condition information and target execution information based on the semantic recognition result.

[0053] Optionally, the control intent, control object, object-attached information, and trigger conditions in the text content can be extracted based on natural language understanding (NLU) and integrated into a four-tuple of the form {action, object, information, condition}. In this method, the result of semantic recognition is the four-tuple. Among them, action represents the control intent, or can be understood as the control purpose, object represents the control object, information represents the object-attached information, and condition represents the trigger condition. Among them, the control intent, control object, and object-attached information can be used as target execution information, and the trigger condition is the trigger condition information mentioned above.

[0054] For example, the text content obtained by converting the voice command is "When Mom calls, reply to Mom's text message saying I'm temporarily unavailable." Based on natural language understanding, it can be understood that the user intention is "send a text message," the control object is "Mom," the object's attached information is "temporarily unavailable," and the control condition is "when Mom calls." This can be recorded as a four-tuple: {(send text message), (Mom), (temporarily unavailable), (when Mom calls)}.

[0055] As another possible implementation, the electronic device may also use a pre-trained command recognition model to recognize the text content corresponding to the voice command, thereby obtaining the target execution information corresponding to the voice command and the trigger condition information corresponding to the target execution information. The command recognition model can be trained based on text samples pre-labeled with the trigger condition information and the target execution information.

[0056] Of course, the specific manner in which the electronic device specifically identifies the target execution information and trigger condition information corresponding to the voice command may not be limited.

[0057] Step S130: When the trigger condition corresponding to the trigger condition information is satisfied, executing the target operation on the target operation interface corresponding to the target execution information.

[0058] In an embodiment of the present application, when the voice command is of a non-immediate type, the electronic device, upon recognizing the trigger condition information and target execution information corresponding to the voice command, can execute the target execution information according to the trigger condition information. Specifically, the electronic device can execute the target operation on the target operation interface corresponding to the target execution information when the trigger condition corresponding to the trigger condition information is satisfied.

[0059] For example, see Figure 6 , the user inputs "set flight mode when returning home" by voice, the trigger condition information is "when returning home", and the target execution information is "set flight mode". Then, when the electronic device determines that the device location is the home location based on the positioning information, in the device mode setting interface, the switch control corresponding to the flight mode is set to the on state.

[0060] For another example, in the example of the four-tuple information corresponding to the above-mentioned recognition voice command, the obtained four-tuple is: {(send text message), (Mom), (temporarily busy), (when Mom calls)}. Then, when the electronic device receives a call from Mom, it can switch to the sending interface of the SMS application, write "Mom" in the recipient edit box, write "temporarily busy" in the text edit box, and then execute the SMS sending.

[0061] In some embodiments, when an electronic device performs a target operation on a target operation interface corresponding to target execution information, it can generate a control instruction for the target interface corresponding to the target execution information by system injection (an operation method supported by Android) or by simulating a screen click. For example, when the trigger condition information is met, a user click operation can be simulated to switch to the target interface corresponding to the target execution information, and a user click operation can be simulated to execute the target control corresponding to the target execution information. In this way, the corresponding target operation is executed on the target interface corresponding to the target execution information.

[0062] In an embodiment of the present application, if the instruction type of the voice instruction is identified as an immediate type, the target operation information corresponding to the voice instruction can be identified to complete the real-time voice control required by the user. Among them, identifying the target operation information corresponding to the voice instruction can be identifying the target execution information corresponding to the voice instruction. The method of identifying the target execution information corresponding to the voice instruction can refer to the method of identifying the interface target execution information in the aforementioned embodiment, which will not be repeated here. It should be noted that when the instruction type of the voice instruction is an immediate type, it can be real-time voice control of the current user interface or real-time voice control of other interfaces of the electronic device.

[0063] In some embodiments, after the electronic device identifies the target operation information, it can match the target operation information with the interface operable elements, thereby obtaining the interface operable elements in the matching interface, and performing the operation corresponding to the interface operable elements in the interface. For details on this embodiment, please refer to the content of the aforementioned embodiment and will not be repeated here.

[0064] In one possible implementation, if the command type is immediate, it can be real-time voice control of the current user interface. The user's voice may be somewhat casual due to their pronunciation habits, but the voice command corresponding to such casual speech may not allow the electronic device to accurately determine the user's control intent. For example, if the voice command itself corresponds to "next," the meaning of "next" may be "next" or "download." For example, in an audio playback scenario, "next" may mean "next," such as "playing the next song." In a software download scenario, "next" may mean "download," such as "downloading an application." Therefore, to more accurately determine the user's true intent, the target operation information corresponding to the voice command can be updated based on the task scenario corresponding to the current user interface to obtain a scenario control command. The scenario control command is then matched with the interface operable elements of the current user interface to determine the target operable element from among the interface operable elements of the current user interface. In the above example, if the current application scenario is music playback, the target operation information "next music" can be updated to "play the next song." If the current application scenario is app downloading, the target operation information "next music" can be updated to "download a music playback app." This allows for more accurate voice control.

[0065] The voice control method provided in the embodiment of the present application can realize the recognition of the instruction type of the voice instruction input by the user. For the non-immediate type of voice instruction input by the user, after identifying its trigger condition information, the corresponding interface operation is performed according to the trigger condition information, thereby better completing the non-immediate type of voice instruction and thus improving the user experience.

[0066] See also Figure 7 , Figure 7 The flow chart of the voice control method provided by another embodiment of the present application is shown. The voice control method is applied to the above electronic device. Figure 7 The process shown in FIG. 1 is described in detail. The voice control method may specifically include the following steps:

[0067] Step S210: Acquire voice instructions.

[0068] Step S220: When the command type of the voice command is a non-immediate type, the trigger condition information and target execution information corresponding to the voice command are obtained.

[0069] In the embodiment of the present application, steps S210 to S220 can refer to the contents of other embodiments and will not be repeated here.

[0070] In the embodiment of the present application, step S210 and step S220 can refer to the contents of other embodiments and will not be repeated here.

[0071] Step S230: When the trigger condition corresponding to the trigger condition information is met, the target execution information is matched with the interface operable elements to obtain the target operation interface, and the target operation is performed on the target operation interface.

[0072] In an embodiment of the present application, when an electronic device performs voice control based on trigger condition information and target execution information, it can match the target execution information with the interface operable elements when the trigger condition corresponding to the trigger condition information is met, obtain the target operation interface, and perform the target operation on the target operation interface.

[0073] In some embodiments, the electronic device can pre-identify interface operable elements of multiple interfaces in the electronic device. The interface refers to the interface that the electronic device can run and display, which can include the system interface, the interfaces corresponding to the various installed applications, etc., which are not limited here. The electronic device can match the target execution information obtained by the above identification with the interface operable elements of the pre-identified interface, thereby obtaining the interface operable elements that match the target operation interface that matches the target execution information as the target operable elements, and then perform the operation corresponding to the target operable elements on the target interface. For example, for the SMS sending interface of the SMS application, it includes a recipient edit box and a text edit box. If the identified target execution information is "Send a text message to mom, saying I'm busy at the moment", the matching interface operation elements are: the recipient edit box, the text edit box and the send control in the SMS sending interface. When the electronic device performs the operation corresponding to the target operable element on the target interface, it can write "Mom" in the recipient edit box and "I'm busy at the moment" in the text edit box in the SMS sending interface, and then trigger the send control to send the SMS.

[0074] It is understandable that for an interface, there may be a variety of interface operable elements that can be operated by the user. The interface operable element may include a certain control in the interface, or it may be for the entire interface. For example, if the user's intention is to slide the page (for example, swipe up, swipe down, swipe left and swipe right), or the intention is to switch the interface, or to exit a certain interface, then the interface operable element is the entire interface. For another example, if the user's intention is to click on a certain position in the interface, then the interface operable element may be for a certain control in the interface. In the case where there can be multiple operations implemented on the interface, there can also be multiple interface operable elements corresponding to the interface.

[0075] In some embodiments, the electronic device identifies the interface operable elements of the interface, which may include at least one of the following identification methods: identifying the interface based on code parsing; identifying the interface based on graphic and text recognition; and identifying the interface based on a control classification model.

[0076] As a possible implementation method, identifying the interface based on code parsing to obtain the interface operable elements corresponding to the interface can be understood as identifying the components or components included in the interface based on code parsing, and the obtained interface operable elements can include the identification and description information of the identifiable components. Correspondingly, identifying the interface based on code parsing can be understood as obtaining the components included in the interface and the description information corresponding to the components based on code parsing. The description information may include information such as name, function, trigger operation, etc. Optionally, the interface can be identified based on code parsing based on Google Accessibility Service accessibility.

[0077] As a possible implementation method, the interface is recognized based on image and text recognition, which may include OCR (Optical Character Recognition) to identify components, controls, icons, etc. in the interface and obtain descriptive information of the identified components, controls, icons, etc. Specifically, the positions of components, controls, and icons in the user interface can be identified through OCR, and then a traversal can be performed to obtain all components, controls, icons, etc. in the user interface. Then, by analyzing the image content, descriptive information of the components, controls, icons, etc. can be determined.

[0078] As a possible implementation method, the training process of the control classification model includes: obtaining a user interface; obtaining controls classified from the user interface; and training a neural network to be trained using the classified controls to obtain a control classification model.

[0079] The electronic device can store the interface operable elements of multiple interfaces that have been pre-identified, so that when performing voice control, they can be matched with the target execution information corresponding to the voice command, thereby obtaining the interface operable elements in the corresponding target operation interface and using them as target operable elements.

[0080] In other embodiments, the electronic device may also, when the trigger condition corresponding to the trigger condition information is satisfied, identify interface operable elements of multiple interfaces in the electronic device, then match the target execution information with the interface operable elements to obtain a target operation interface, and perform the target operation on the target operation interface. The manner in which the electronic device identifies interface operable elements can be found in the above embodiments and will not be further described here.

[0081] In some embodiments, the trigger condition information may include the condition field and condition parameters of the trigger condition, and the target execution information may include the execution field and execution parameters. The condition field refers to the service type to which the trigger condition of the non-immediate type of voice instruction belongs, for example, the battery level of the electronic device, the location of the electronic device, the time of the electronic device, the date of the electronic device, the message notification received by the electronic device, the incoming call received, etc.; the condition parameter refers to the specific trigger state or parameter value of the service to which the non-immediate type of voice instruction belongs, for example, the specific power value, the specific time, the specific date, the specific location, the type of message received, the type of incoming call received, etc.; the execution field refers to the service field to which the operation specifically executed by the non-immediate type of voice instruction belongs, for example, controlling home appliances, controlling the parameters of electronic devices, operating application software, etc.; the execution parameter refers to the specific operation parameter corresponding to the operation in the execution field, for example, controlling the specific setting temperature of the smart air conditioner, the specific value of the device volume of the electronic device, etc.

[0082] Under this embodiment, when the electronic device matches the target execution information with the interface operable elements, it can match the execution domain and execution parameters with the interface operable elements, thereby obtaining the target operable elements in the target operation interface that matches the target execution information. Among them, the electronic device can determine the corresponding interface based on the execution domain, and then match the interface operable elements of the interface based on the execution parameters, thereby obtaining the target operable elements that match the target execution information. For example, if the execution domain is to control a smart air conditioner, the interface is the control interface of the smart air conditioner corresponding to the smart home application, and then based on the execution parameters, if the execution parameters are to lower the temperature of the smart air conditioner, it can be determined that the matching interface operable element is: a control for lowering the temperature of the smart air conditioner.

[0083] In this embodiment, by dividing the target execution information into execution fields and execution parameters, the electronic device can match the interface operable elements of the matching interface more accurately and improve the matching efficiency; similarly, dividing the trigger condition information into condition fields and condition parameters can enable the electronic device to execute the matched interface operable elements according to the trigger condition information, and can more accurately ensure that the target operation is performed on the target operation interface when the trigger condition information is met, thereby improving the accuracy of voice control.

[0084] The voice control method provided in the embodiment of the present application identifies the trigger condition information and target execution information corresponding to the non-instantaneous voice command input by the user, and when the trigger condition corresponding to the trigger condition information is met, matches the target execution information with the interface operable elements, thereby being able to quickly and accurately determine the interface operation required by the user, and then execute the corresponding interface operation, thereby being able to better complete the non-instantaneous voice command, thereby improving the user experience.

[0085] See also Figure 8 , Figure 8 A flow chart of a voice control method according to another embodiment of the present invention is shown. The voice control method is applied to the above electronic device. Figure 8 The process shown in FIG. 1 is described in detail. The voice control method may specifically include the following steps:

[0086] Step S310: Acquire voice instructions.

[0087] Step S320: When the command type of the voice command is a non-immediate type, the trigger condition information and target execution information corresponding to the voice command are obtained.

[0088] In the embodiment of the present application, step S310 and step S320 can refer to the contents of other embodiments and will not be repeated here.

[0089] Step S330: Match the target execution information with the interface operable elements to obtain the target operation interface.

[0090] Unlike the previous embodiment, in this embodiment of the present application, after the electronic device obtains the trigger condition information and target execution information corresponding to the voice command, it can match the target execution information with the interface operable elements to obtain the target operation interface, so that when the trigger condition corresponding to the trigger condition information is met, the target operation is performed on the target operation interface corresponding to the target execution information. The manner in which the electronic device matches the target execution information with the interface operable elements can be found in the previous embodiment and will not be repeated here.

[0091] Step S340: Generate corresponding control instructions according to the trigger condition information and the target operation interface.

[0092] In an embodiment of the present application, after the electronic device identifies the trigger condition information and the target operation interface, it can synthesize the control instructions based on the trigger condition information and the target operation interface, and pass the synthesized control instructions to the graphical interface for execution. Among them, the electronic device can generate corresponding control instructions based on the trigger condition information and the interface operable elements in the matched target operation interface. Optionally, the IFTTT instruction generation method can be used to generate corresponding control instructions based on the trigger condition information and the target operation interface.

[0093] Step S350: executing the control instruction, wherein the control instruction is used to execute a target operation on the target operation interface corresponding to the target execution information when the trigger condition corresponding to the trigger condition information is satisfied.

[0094] In an embodiment of the present application, since the above-mentioned control instructions are non-immediate type instructions, the control instructions exist in the application background according to the specific trigger condition information and the interface operable elements of the target operation interface; and, by real-time monitoring of the status of the trigger condition information, such as monitoring the actual status of the condition field and condition parameters in the aforementioned embodiment, and making the corresponding execution module (used to execute the target execution information) in a standby state (existing in the memory); when it is detected that the trigger condition information is met, the execution module immediately executes the operation corresponding to the interface operable element in the target operation interface, thereby completing the non-immediate type voice control process.

[0095] The voice control method provided in the embodiment of the present application identifies the trigger condition information and target execution information corresponding to the non-instantaneous voice command input by the user, and then matches the target execution information with the interface operable elements, so as to quickly and accurately determine the interface operation required by the user, and then execute the corresponding interface operation according to the trigger condition information, so as to better complete the non-instantaneous voice command, thereby improving the user experience.

[0096] See also Figure 9 , Figure 9 A flow chart of a voice control method according to another embodiment of the present application is shown. The voice control method is applied to the above electronic device. Figure 9 The process shown in FIG. 1 is described in detail. The voice control method may specifically include the following steps:

[0097] Step S410: Acquire voice instructions.

[0098] Step S420: When the instruction type of the voice instruction is a non-instant type, obtain the instruction text corresponding to the voice instruction.

[0099] Step S430: Input the instruction text corresponding to the voice instruction into a pre-trained instruction recognition model to obtain the trigger condition information and target execution information contained in the instruction text. The instruction recognition model is trained based on hierarchical reinforcement learning.

[0100] In an embodiment of the present application, when the electronic device determines that the command type of a voice command is a non-immediate type, when identifying the trigger condition information and target execution information corresponding to the voice command, the electronic device can input the command text corresponding to the voice command into a command recognition model pre-trained based on hierarchical reinforcement learning to obtain the trigger condition information and target execution information contained in the command text. Since multiple tasks need to identify trigger condition information and target execution information, the command recognition model is trained using hierarchical reinforcement learning to improve recognition accuracy.

[0101] In some embodiments, the instruction recognition model includes a first submodule corresponding to trigger condition information, a second submodule corresponding to target execution information, and a collaborative control module. The first submodule corresponding to trigger condition information is used to identify the recognition task of the trigger condition information, and the second submodule corresponding to the target execution information is used to identify the recognition task of the target execution information. The collaborative control module is used to determine the action of the corresponding recognition task, specifically determining the execution order of the submodule's work tasks and assigning work tasks to the submodules. In this approach, the hierarchical reinforcement learning model can more accurately identify trigger condition information and target execution information.

[0102] In this way, the training process of the instruction recognition model includes: creating a first submodule corresponding to the recognition task for identifying trigger condition information, a second submodule corresponding to the recognition task for identifying target execution information, and a collaborative control module for coordinating the recognition task, the first submodule and the second submodule are used to decide the actions of their corresponding recognition tasks, and the decision priority of the collaborative control module is higher than the decision priority of the first submodule and the second submodule; based on text samples marked with trigger condition information, target execution information and the recognition order of the recognition task, the collaborative control module, the first submodule and the second submodule are deep reinforcement learning trained to obtain the trained instruction recognition model.

[0103] As a possible implementation manner, based on a text sample annotated with trigger condition information, target execution information, and the recognition order of the recognition task, deep reinforcement learning training is performed on the collaborative control module, the first submodule, and the second submodule to obtain the trained instruction recognition model, which may include: inputting the text sample into the collaborative control module, the first submodule, and the second submodule to obtain the output results of the first submodule and the second submodule, and the execution order coordinated by the collaborative control module; determining a first reward corresponding to the first submodule based on the output result of the first submodule and the trigger condition information annotated with the text sample; determining a second reward corresponding to the second submodule based on the output result of the second submodule and the target execution information annotated with the text sample; determining a third reward corresponding to the collaborative control module based on the execution order coordinated by the collaborative control module and the recognition order annotated with the text sample; performing deep reinforcement learning training on the first submodule based on the first reward, performing deep reinforcement learning training on the second submodule based on the second reward, and performing deep reinforcement learning training on the collaborative control module based on the third reward, until a preset termination condition is met to obtain the trained instruction recognition model.

[0104] In this embodiment, deep reinforcement learning training is performed on the first submodule based on the first reward, deep reinforcement learning training is performed on the second submodule based on the second reward, and deep reinforcement learning training is performed on the collaborative control module based on the third reward, that is, model training. The algorithm of reinforcement learning training may not be limited, for example, it may be an Advantage Actor Critic (A2C) algorithm, an asynchronous Advantage Actor-Critic (A3C) algorithm, or a Deep Q-Network (DQN) algorithm, etc.

[0105] The following describes the instruction recognition model in the embodiment of the present application, taking as an example the trigger condition information including the condition field and condition parameters, and the target execution information including the execution field and execution parameters. The condition field, condition parameters, execution field, and execution parameters are defined as four slots. The definitions of the condition field, condition parameters, execution field, and execution parameters can be found in the preceding embodiments and will not be repeated here.

[0106] This instruction recognizes the model as Figure 10As shown in the figure, "Agent" refers to an intelligent agent in the reinforcement learning concept, which is trained using training data. The command recognition model consists of one top-level agent and four bottom-level agents. Each bottom-level agent is responsible for identifying information about four specific slots to be identified. The top-level agent determines which bottom-level agent should be selected for identification based on the state of the top-level agent, thereby controlling the working order of the bottom-level agents. The top-level agent can be understood as the coordination control module mentioned above, and the bottom-level agents can be understood as the first and second submodules mentioned above, with the first and second submodules corresponding to two bottom-level agents, respectively.

[0107] The state of the top-level agent refers to the current completion status of the top-level agent's work. Since the top-level agent is responsible for assigning work to the bottom-level agents, the state of the top-level agent can be expressed using the following formula: SH represents the state of the top-level agent, and st1, st2, st3, and st4 represent the states of the four bottom-level agents, respectively. The states of the bottom-level agents store the current slot recognition confidence, recognition content, and historical actions (actions, execution steps of reinforcement learning agents). That is:

[0108] S H ={st1,st2,st3,st4}

[0109] The training goal of the underlying agent is to identify the information it is responsible for identifying (conditional domain, conditional parameters, execution domain, and execution parameters) as accurately as possible. Therefore, during the training process, the training reward (reward, which refers to the feedback on the agent's execution results in reinforcement learning) of the underlying agent is defined as:

[0110]

[0111] Here, st represents the execution state of the underlying agent, and gt is the annotation information of the sample text. When the execution state is exactly the same as the true state, the training reward is 1. When the execution state is empty, the training reward is -p, where p is the set null penalty. When the execution state is all other than the above states, the training reward is -1. During the training process, the training reward is calculated every time the underlying agent takes an action.

[0112] Furthermore, since the information in the bottom slots is mutually constrained and interdependent, after correctly identifying the information of a certain slot, the information search range of another slot to be identified can be narrowed. For example, if the execution domain is identified as adjusting the switch status of the air conditioner, then the execution function can only be the smart home control application in the electronic device. For another example, if the execution function is identified as the volume, then the execution domain can only be the volume adjustment button on the graphical interface of the electronic device. Therefore, the training purpose of the top-level agent is to allocate the execution order of the bottom-level agents as reasonably and efficiently as possible, and give priority to identifying slots containing rich information. During the training process, the training reward of the top-level agent is defined as:

[0113]

[0114] in, It represents the cumulative training reward when the top-level agent's execution state is st and the ground-truth label is gt. The top-level agent's reward is the sum of the valid accumulated rewards of the bottom-level agents up to step i and backtracking N steps. And, only when the bottom-level agent selected by the top-level agent at a certain step is the same as the ground-truth labeled bottom-level agent (that is, the execution order is the same as the labeled execution order) is it considered valid accumulation, otherwise the reward of that step is recorded as 0.

[0115] In addition, the instruction recognition model can also include an information specification module. Each time the bottom-level Agent is executed, the current slot recognition status will be updated. The information specification module integrates and processes the slot recognition status. If the probability of the recognition result of a certain bottom-level Agent is lower than the set threshold, the information specification module can match the slot information identified in the recognition result with the interface operable elements to verify whether the recognition result is reasonable. For example, if it is identified as "adjust NetEase Cloud Music to vibration mode", the information specification module can match it with the interface operable elements, and through matching, it can be determined that NetEase Cloud Music is a software application rather than a device terminal, and the operation cannot be performed; the information specification module will mark the recognition result of this step as incomplete and requires further processing, and feedback it to the top-level Agent. In this way, the accuracy of recognition can be improved.

[0116] Of course, there is no limitation on the specific method of performing deep reinforcement learning training on the instruction recognition model.

[0117] Step S440: When the trigger condition corresponding to the trigger condition information is satisfied, the target operation is executed on the target operation interface corresponding to the target execution information.

[0118] In the embodiment of the present application, step S440 can refer to the contents of other embodiments and will not be repeated here.

[0119] The voice control method provided in the embodiment of the present application is to identify the trigger condition information and target execution information corresponding to the non-instantaneous voice command input by the user through a command recognition model pre-trained in a hierarchical reinforcement learning manner. Then, when the trigger condition corresponding to the trigger condition information is met, the target operation is performed on the target operation interface corresponding to the target execution information. This can quickly and accurately determine the interface operation required by the user, and then perform the corresponding interface operation according to the trigger condition information, so as to better complete the non-instantaneous voice command, thereby improving the user experience. Moreover, since the command recognition model is trained based on a hierarchical reinforcement learning method, it can ensure the accuracy of identifying the trigger condition information and the target execution information, thereby improving the accuracy of voice control.

[0120] See also Figure 11 , Figure 11 The flow chart of the voice control method provided by another embodiment of the present application is shown. The voice control method is applied to the above electronic device. Figure 11 The process shown in FIG. 1 is described in detail. The voice control method may specifically include the following steps:

[0121] Step S510: Acquire voice instructions.

[0122] In the embodiment of the present application, step S410 can refer to the content of the aforementioned embodiment and will not be repeated here.

[0123] Step S520: Input the instruction text corresponding to the voice instruction into a pre-trained instruction type recognition model to obtain the instruction type of the voice instruction. The instruction type recognition model is used to identify whether the instruction type corresponding to the input instruction text is an immediate type or a non-immediate type.

[0124] In an embodiment of the present application, when the instruction type of a voice instruction is identified, it can be identified based on a pre-trained instruction type recognition model. The text vector of the instruction text corresponding to the voice instruction can be input into the pre-trained instruction type recognition model to obtain the instruction type of the voice instruction. The instruction type recognition model can be a text semantics model based on a BERT (Deep Bidirectional Transformers for Language Understanding) model. Figure 2 Classification (whether it is non-immediate type) deep learning network.

[0125] In some embodiments, see Figure 12The instruction type recognition model consists of an encoding module, a decoding module, and a BERT text classification model. The BERT model can be a publicly pre-trained semantic model. After the instruction text is input into the instruction type recognition model, it is first converted into the input format received by the BERT text classification model through the encoding module. After encoding, the encoded text vector is input into the BERT network for classification. The classification result vector is then decoded by a decoding device to obtain the classification result, which includes both immediate type and non-immediate type.

[0126] In some embodiments, before inputting the instruction text corresponding to the voice instruction into the above-mentioned instruction type recognition model, the electronic device may also perform preset correction processing on the instruction text corresponding to the voice instruction; then, the instruction text after the preset correction processing is input into the above-mentioned instruction type recognition model to obtain the instruction type of the voice instruction.

[0127] Optionally, the preset correction process may include: vocabulary calibration based on edit distance, common vocabulary correction based on Bayesian method, etc. The specific correction process may not be limited.

[0128] Step S530: When the command type of the voice command is a non-immediate type, the trigger condition information and target execution information corresponding to the voice command are obtained.

[0129] Step S540: When the trigger condition corresponding to the trigger condition information is satisfied, the target operation is executed on the target operation interface corresponding to the target execution information.

[0130] In the embodiment of the present application, step S530 and step S540 can refer to the contents of the aforementioned embodiment and will not be repeated here.

[0131] The voice control method provided in the embodiments of the present application uses a pre-trained command type recognition model to identify the command type when identifying the command type of a voice command. For non-immediate voice commands input by the user, the method identifies the trigger condition information and target execution information, and executes the target execution information based on the trigger condition information. This allows for better recognition of the type of voice command and the completion of both immediate and non-immediate voice commands, thereby improving the user experience.

[0132] Next, pass Figure 13 The voice control method involved in the above embodiment is introduced.

[0133] like Figure 13As shown, after the electronic device obtains the voice command, it parses the voice command to obtain the voice text corresponding to the voice command; then it determines whether the command type is a non-immediate type; if it is identified as a non-immediate type, it identifies the trigger condition information and target execution information through command recognition; then it synthesizes the command based on the target execution information and the trigger condition information, and hands it over to the graphical interface for execution; if it is identified as an immediate type, it directly matches the interface operable elements to obtain the corresponding target operable elements, which are then synthesized into commands and handed over to the graphical interface for execution.

[0134] See also Figure 14 , which shows a structural block diagram of a voice control device 400 provided in an embodiment of the present application. The voice control device 400 applies the above-mentioned electronic device, and the voice control device 400 includes: an instruction acquisition module 410, an information acquisition module 420, and an operation execution module 430. The instruction acquisition module 410 is used to acquire a voice instruction; the information acquisition module 420 is used to acquire the trigger condition information and target execution information corresponding to the voice instruction when the instruction type of the voice instruction is a non-immediate type; the operation execution module 430 is used to execute the target operation on the target operation interface corresponding to the target execution information when the trigger condition corresponding to the trigger condition information is met.

[0135] In some embodiments, the operation execution module 430 can be specifically used to: when the trigger condition corresponding to the trigger condition information is met, match the target execution information with the interface operable elements to obtain the target operation interface, and execute the target operation on the target operation interface.

[0136] As a possible implementation, the operation execution module 430 can be specifically used to: match the target execution information with the interface operable elements, obtain the matching interface operable elements in the target operation interface as the target operable elements; and execute the operation corresponding to the target operable elements in the target interface.

[0137] In some embodiments, the voice control device 400 may further include an interface recognition module. The interface recognition module is configured to: when the trigger condition corresponding to the trigger condition information is satisfied, before the target operation is performed on the target operation interface corresponding to the target execution information, match the target execution information with interface operable elements to obtain the target operation interface.

[0138] As a possible implementation, the operation execution module 430 can be specifically used to: generate corresponding control instructions based on the trigger condition information and the target operation interface; execute the control instructions, and the control instructions are used to perform the target operation on the target operation interface corresponding to the target execution information when the trigger condition corresponding to the trigger condition information is met.

[0139] In some implementations, the information acquisition module 420 may be specifically configured to: acquire the instruction text corresponding to the voice instruction; and acquire the trigger condition information and target execution information contained in the instruction text.

[0140] In one possible implementation, the information acquisition module 420 can be specifically used to: input the instruction text corresponding to the voice instruction into a pre-trained instruction recognition model to obtain the trigger condition information and target execution information contained in the instruction text, and the instruction recognition model is trained based on hierarchical reinforcement learning.

[0141] Optionally, the instruction recognition model includes a first submodule corresponding to the trigger condition information, a second submodule corresponding to the target execution information, and a collaborative control module. The voice control device 400 may also include a model training module. The model training module can be used to: create a first submodule corresponding to the recognition task for identifying the trigger condition information, a second submodule corresponding to the recognition task for identifying the target execution information, and a collaborative control module for coordinating the recognition task, the first submodule and the second submodule are used to decide the action of the corresponding recognition task, and the decision priority of the collaborative control module is higher than the decision priority of the first submodule and the second submodule; based on a text sample marked with trigger condition information, target execution information and the recognition order of the recognition task, the collaborative control module, the first submodule and the second submodule are subjected to deep reinforcement learning training to obtain the trained instruction recognition model.

[0142] Furthermore, the model training module can be specifically used to: input the text sample into the collaborative control module, the first submodule and the second submodule to obtain the output results of the first submodule and the second submodule, and the execution order coordinated by the collaborative control module; determine the first reward corresponding to the first submodule based on the output result of the first submodule and the trigger condition information marked with the text sample; determine the second reward corresponding to the second submodule based on the output result of the second submodule and the target execution information marked with the text sample; determine the third reward corresponding to the collaborative control module based on the execution order coordinated by the collaborative control module and the recognition order marked with the text samples; perform deep reinforcement learning training on the first submodule based on the first reward, perform deep reinforcement learning training on the second submodule based on the second reward, and perform deep reinforcement learning training on the collaborative control module based on the third reward, until the preset termination condition is met to obtain the trained instruction recognition model.

[0143] In some embodiments, the speech recognition device 400 may further include a type recognition module. The type recognition module is configured to input the instruction text corresponding to the voice instruction into a pre-trained instruction type recognition model to obtain the instruction type of the voice instruction before obtaining the trigger condition information and target execution information corresponding to the voice instruction when the instruction type of the voice instruction is a non-immediate type, wherein the instruction type recognition model is configured to identify whether the instruction type corresponding to the input instruction text is an immediate type or a non-immediate type.

[0144] In a possible embodiment, the type recognition module can also be specifically used to: perform preset correction processing on the instruction text corresponding to the voice instruction; input the instruction text after the preset correction processing into a pre-trained instruction type recognition model to obtain the instruction type of the voice instruction.

[0145] In some embodiments, the information acquisition module 420 can also be used to identify the target operation information corresponding to the voice instruction after the voice instruction is acquired, if the instruction type of the voice instruction is an immediate type; the operation execution module 440 can also be used to respond to the target operation information and execute the operation corresponding to the target operation information on the interface corresponding to the target operation information.

[0146] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the above-described devices and modules can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0147] In several embodiments provided in this application, the coupling between modules may be electrical, mechanical or other forms of coupling.

[0148] In addition, the functional modules in the various embodiments of the present application may be integrated into a processing module, or each module may exist physically separately, or two or more modules may be integrated into a single module. The above-mentioned integrated modules may be implemented in the form of hardware or software functional modules.

[0149] In summary, the solution provided by this application obtains a voice command. When the command type of the voice command is a non-instantaneous type, the trigger condition information and target execution information corresponding to the voice command are obtained. When the trigger condition corresponding to the trigger condition information is met, the target operation is executed on the target operation interface corresponding to the target execution information. In this way, it is possible to identify the trigger condition information of a non-instantaneous voice command input by the user, and then execute the corresponding interface operation based on the trigger condition information, thereby better completing the non-instantaneous voice command and improving the user experience.

[0150] Please refer to Figure 15 , which shows a structural block diagram of an electronic device provided in an embodiment of the present application. The electronic device 100 can be an electronic device capable of running applications, such as a smartphone, a tablet computer, a smart watch, smart glasses, a laptop computer, etc. The electronic device 100 in the present application may include one or more of the following components: a processor 110, a memory 120, and one or more applications, wherein the one or more applications may be stored in the memory 120 and configured to be executed by one or more processors 110, and the one or more programs are configured to execute the method described in the aforementioned method embodiment.

[0151] The processor 110 may include one or more processing cores. The processor 110 utilizes various interfaces and circuits to connect various components within the electronic device 100. It executes instructions, programs, code sets, or instruction sets stored in the memory 120, and accesses data stored in the memory 120 to perform various functions and process data within the electronic device 100. Optionally, the processor 110 may be implemented using at least one of the following hardware forms: a digital signal processing (DSP), a field-programmable gate array (FPGA), or a programmable logic array (PLA). The processor 110 may integrate one or a combination of a central processing unit (CPU), a graphics processing unit (GPU), and a modem. The CPU primarily processes the operating system, user interface, and application programs; the GPU is responsible for rendering and drawing display content; and the modem handles wireless communications. It is understood that the modem may not be integrated into the processor 110 and may be implemented separately via a communications chip.

[0152] The memory 120 may include a random access memory (RAM) or a read-only memory (ROM). The memory 120 may be used to store instructions, programs, codes, code sets, or instruction sets. The memory 120 may include a program storage area and a data storage area, wherein the program storage area may store instructions for implementing an operating system, instructions for implementing at least one function (such as a touch function, a sound playback function, an image playback function, etc.), instructions for implementing the following various method embodiments, etc. The data storage area may also store data created by the electronic device 100 during use (such as a phone book, audio and video data, chat history data), etc.

[0153] Please refer to Figure 16 , which shows a block diagram of a computer-readable storage medium provided in an embodiment of the present application. The computer-readable medium 800 stores program code, which can be called by a processor to execute the method described in the above method embodiment.

[0154] The computer-readable storage medium 800 can be an electronic memory such as a flash memory, an EEPROM (Electrically Erasable Programmable Read-Only Memory), an EPROM, a hard disk, or a ROM. Alternatively, the computer-readable storage medium 800 includes a non-transitory computer-readable storage medium. The computer-readable storage medium 800 has storage space for program code 810 for executing any of the method steps described above. These program codes can be read from or written to one or more computer program products. The program code 810 can be compressed, for example, in a suitable form.

[0155] An embodiment of the present application also provides a computer program product, including a computer program / instruction, characterized in that when the computer program / instruction is executed by a processor, the voice control method provided in the aforementioned embodiment is implemented.

[0156] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A voice control method, characterized in that: The method comprises: Get voice commands; When the command type of the voice command is a non-immediate type, obtaining a command text corresponding to the voice command; The instruction text corresponding to the voice instruction is input into a pre-trained instruction recognition model to obtain the trigger condition information and target execution information contained in the instruction text, the instruction recognition model includes a first submodule corresponding to the recognition task for identifying the trigger condition information, a second submodule corresponding to the recognition task for identifying the target execution information, and a collaborative control module for coordinating the recognition task, the first submodule and the second submodule are used to decide the actions of their corresponding recognition tasks, and the decision priority of the collaborative control module is higher than the decision priority of the first submodule and the second submodule, and the instruction recognition model is obtained after deep reinforcement learning training of the collaborative control module, the first submodule and the second submodule based on text samples marked with trigger condition information, target execution information and the recognition order of the recognition task; When the trigger condition corresponding to the trigger condition information is met, the target operation is executed on the target operation interface corresponding to the target execution information.

2. The method according to claim 1, characterized in that When the trigger condition corresponding to the trigger condition information is met, executing the target operation on the target interface corresponding to the target execution information includes: When the trigger condition corresponding to the trigger condition information is met, the target execution information is matched with the interface operable elements to obtain the target operation interface, and the target operation is performed on the target operation interface.

3. The method according to claim 2, characterized in that The step of matching the target execution information with the interface operable elements to obtain the target operation interface, and performing the target operation on the target operation interface includes: Matching the target execution information with the interface operable elements to obtain the matched interface operable elements in the target operation interface as the target operable elements; Execute the operation corresponding to the target operable element on the target interface.

4. The method according to claim 1, wherein When the trigger condition corresponding to the trigger condition information is satisfied, before executing the target operation on the target operation interface corresponding to the target execution information, the method further includes: The target execution information is matched with the interface operable elements to obtain the target operation interface.

5. The method according to claim 4, characterized in that When the trigger condition corresponding to the trigger condition information is met, executing the target operation on the target operation interface corresponding to the target execution information includes: Generate corresponding control instructions according to the trigger condition information and the target operation interface; Execute the control instruction, where the control instruction is used to execute a target operation on a target operation interface corresponding to the target execution information when a trigger condition corresponding to the trigger condition information is met.

6. The method according to claim 1, characterized in that Based on text samples annotated with trigger condition information, target execution information, and the recognition order of the recognition task, deep reinforcement learning training is performed on the collaborative control module, the first submodule, and the second submodule to obtain the trained instruction recognition model, including: Inputting the text sample into the collaborative control module, the first submodule, and the second submodule to obtain output results of the first submodule and the second submodule, and an execution order coordinated by the collaborative control module; Determining a first reward corresponding to the first submodule based on an output result of the first submodule and trigger condition information of the text sample being marked; Determining a second reward corresponding to the second submodule based on an output result of the second submodule and target execution information annotated with the text sample; determining a third reward corresponding to the collaborative control module based on the execution order coordinated by the collaborative control module and the recognition order in which the text samples are annotated; The first submodule is subjected to deep reinforcement learning training based on the first reward, the second submodule is subjected to deep reinforcement learning training based on the second reward, and the collaborative control module is subjected to deep reinforcement learning training based on the third reward until a preset termination condition is met, thereby obtaining the trained instruction recognition model.

7. The method according to any one of claims 1 to 6, characterized in that When the instruction type of the voice instruction is a non-immediate type, before obtaining the trigger condition information and target execution information corresponding to the voice instruction, the method further includes: The instruction text corresponding to the voice instruction is input into a pre-trained instruction type recognition model to obtain the instruction type of the voice instruction. The instruction type recognition model is used to identify whether the instruction type corresponding to the input instruction text is an immediate type or a non-immediate type.

8. The method according to claim 7, characterized in that Before inputting the instruction text corresponding to the voice instruction into a pre-trained instruction type recognition model to obtain the instruction type of the voice instruction, the method further includes: Performing preset correction processing on the instruction text corresponding to the voice instruction; The step of inputting the instruction text corresponding to the voice instruction into a pre-trained instruction type recognition model to obtain the instruction type of the voice instruction includes: The command text after the preset correction process is input into a pre-trained command type recognition model to obtain the command type of the voice command.

9. The method according to any one of claims 1 to 6, characterized in that: After acquiring the voice instruction, the method further includes: If the command type of the voice command is an immediate type, identifying target operation information corresponding to the voice command; In response to the target operation information, an operation corresponding to the target operation information is performed on an interface corresponding to the target operation information.

10. A voice control device, characterized in that: The device includes: an instruction acquisition module, an information acquisition module and an operation execution module, wherein: The instruction acquisition module is used to acquire voice instructions; The information acquisition module is used to obtain the instruction text corresponding to the voice instruction when the instruction type of the voice instruction is a non-immediate type; the instruction text corresponding to the voice instruction is input into a pre-trained instruction recognition model to obtain the trigger condition information and target execution information contained in the instruction text. The instruction recognition model includes a first submodule corresponding to the recognition task for identifying the trigger condition information, a second submodule corresponding to the recognition task for identifying the target execution information, and a collaborative control module for coordinating the recognition task. The first submodule and the second submodule are used to decide the actions of their corresponding recognition tasks, and the decision priority of the collaborative control module is higher than the decision priority of the first submodule and the second submodule. The instruction recognition model is based on a text sample marked with trigger condition information, target execution information and the recognition order of the recognition task, and is obtained after deep reinforcement learning training is performed on the collaborative control module, the first submodule and the second submodule; The operation execution module is configured to execute a target operation on a target operation interface corresponding to the target execution information when a trigger condition corresponding to the trigger condition information is satisfied.

11. An electronic device, characterized in that: include: one or more processors; Memory; One or more application programs, wherein the one or more application programs are stored in the memory and configured to be executed by the one or more processors, and the one or more programs are configured to execute the method according to any one of claims 1 to 9.

12. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores program code, which can be called by a processor to execute the method according to any one of claims 1 to 9.

13. A computer program product comprising a computer program / instructions, characterized in that When the computer program / instruction is executed by a processor, the steps of the method according to any one of claims 1 to 9 are implemented.

Citation Information

Patent Citations

  • Voice positioning method and device, smart television and storage medium

    CN109600646A

  • Intelligent equipment control method and device, readable storage medium and computer equipment

    CN112415908A