Control methods, devices, electronic equipment and storage media for voice devices
By recognizing user intents of voice devices and determining control strategies based on intent types, the problem of multi-user control of voice devices in public scenarios is solved, performance loss is reduced, and intelligence and application scenarios are improved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-30
- Publication Date
- 2026-03-13
AI Technical Summary
Voice devices indiscriminately respond to repeated control requests from multiple users in public settings, leading to performance degradation. Furthermore, each response interrupts the currently playing resources, limiting their application scenarios.
By acquiring voice commands, voice recognition is performed to determine the type of user intent, and a target control strategy is determined based on the type of intent. Different control measures are taken for different types of intent, including resource playback, device control, and status query, to avoid indiscriminate responses.
It effectively reduces the performance loss of voice devices, improves intelligence, and enriches their application scenarios, especially the user experience in public scenarios.
Smart Images

Figure CN115171687B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of voice device technology, and in particular to a control method, apparatus, electronic device and storage medium for a voice device. Background Technology
[0002] With the development of science and technology, voice-controlled smart devices (hereinafter referred to as: voice devices) are being used more and more widely in people's daily lives, from smart voice speakers and smartphones to smart air conditioners, etc.
[0003] In existing technologies, in most everyday scenarios, as long as the user inputs a voice command, the voice device will receive the voice command and then respond to the voice command to perform the corresponding follow-up processing.
[0004] Therefore, if the use of voice devices is limited to small family settings, there is not much of a problem. However, if voice devices are introduced into public settings, such as barbershops and restaurants, the indiscriminate response to repeated control from multiple users can easily cause performance degradation of the voice devices. In addition, each response will interrupt the resources currently being played by the voice devices, which to some extent limits the application scenarios of voice devices. Summary of the Invention
[0005] In view of this, in order to solve the technical problems in the prior art where voice devices respond indiscriminately to repeated control from multiple users, which can easily cause performance loss of voice devices, and each response will interrupt the currently playing resources of the voice device, thus limiting the application scenarios of voice devices to a certain extent, this invention provides a control method, device, electronic device and storage medium for a voice device.
[0006] In a first aspect, embodiments of the present invention provide a control method for a voice device, the method comprising:
[0007] Acquire voice commands collected by the voice device;
[0008] The voice command is subjected to speech recognition to obtain the user intent corresponding to the voice command;
[0009] Determine the intent type to which the user intent belongs;
[0010] Determine the target control strategy based on the intent type to which the user intent belongs;
[0011] The voice device is controlled according to the target control strategy.
[0012] In one possible implementation, before determining the intent type to which the user intent belongs, the method further includes:
[0013] Determine the operating mode configured for the voice device;
[0014] If it is determined that the voice device is configured in a first operating mode, the step of determining the intent type to which the user intent belongs is performed.
[0015] In one possible implementation, after determining the configured operating mode of the voice device, the method further includes:
[0016] If it is determined that the voice device is configured in the second working mode, a control command corresponding to the user's intent is generated and sent to the voice device for execution; then the process ends.
[0017] In one possible implementation, the intent type includes a first intent type;
[0018] The step of determining the target control strategy based on the intent type to which the user intent belongs includes:
[0019] If the intent type to which the user intent belongs is determined to be the first intent type, the target control strategy is determined to be to control the voice device to play the resources corresponding to the user intent;
[0020] Controlling the voice device according to the target control strategy includes:
[0021] If the intent type to which the user intent belongs is determined to be the first intent type, then the resource corresponding to the user intent is determined.
[0022] The resources are added to the playback queue corresponding to the voice device, and resources are sequentially retrieved from the playback queue and sent to the voice device for playback.
[0023] In one possible implementation, the method further includes:
[0024] When the playback queue is detected to be non-empty, the state of the voice device is set to resource playback state;
[0025] When the playback queue is detected to be empty, the state of the voice device is set to idle state.
[0026] In one possible implementation, the intent type includes a second intent type;
[0027] The step of determining the target control strategy based on the intent type to which the user intent belongs includes:
[0028] If the intent type to which the user intent belongs is determined to be the second intent type, the target control strategy is determined to control the voice device not to perform the device control operation corresponding to the user intent;
[0029] Controlling the voice device according to the target control strategy includes:
[0030] If it is determined that the user's intent belongs to the second intent type, the voice device is controlled not to perform the device control operation corresponding to the user's intent.
[0031] In one possible implementation, the intent type includes a third intent type;
[0032] The step of determining the target control strategy based on the intent type to which the user intent belongs includes:
[0033] If it is determined that the user intent belongs to the third intent type, the target control strategy is to control the voice device to play the device status parameters corresponding to the user intent according to the current state.
[0034] Controlling the voice device according to the target control strategy includes:
[0035] If the intent type to which the user intent belongs is determined to be a third intent type, the current state of the voice device is determined.
[0036] If the current state is an idle state, then control the voice device to play the device state parameters corresponding to the user's intent;
[0037] The method further includes:
[0038] If the current state is a resource playback state, then the voice device is controlled not to play the device state parameters corresponding to the user's intent.
[0039] In a second aspect, embodiments of the present invention provide a control device for a voice device, the device comprising:
[0040] The instruction acquisition module is used to acquire voice instructions collected by the voice device;
[0041] A speech recognition module is used to perform speech recognition on the speech command to obtain the user intent corresponding to the speech command;
[0042] The type determination module is used to determine the intent type to which the user intent belongs;
[0043] The strategy determination module is used to determine the target control strategy based on the intent type to which the user intent belongs;
[0044] The control module is used to control the voice device according to the target control strategy.
[0045] In one possible implementation, the device further includes:
[0046] The mode determination module is used to determine the operating mode configured for the voice device;
[0047] The first execution module is configured to perform the step of determining the intent type to which the user intent belongs when it is determined that the voice device is configured in a first working mode.
[0048] In one possible implementation, the device further includes:
[0049] The second execution module is used to generate a control command corresponding to the user's intent when it is determined that the voice device is configured in a second working mode, and to send the control command to the voice device so that the voice device can execute the control command; and then end the process.
[0050] In one possible implementation, the intent type includes a first intent type;
[0051] The strategy determination module is specifically used for:
[0052] If the intent type to which the user intent belongs is determined to be the first intent type, the target control strategy is determined to be to control the voice device to play the resources corresponding to the user intent;
[0053] The control module is specifically used for:
[0054] If the intent type to which the user intent belongs is determined to be the first intent type, then the resource corresponding to the user intent is determined.
[0055] The resources are added to the playback queue corresponding to the voice device, and resources are sequentially retrieved from the playback queue and sent to the voice device for playback.
[0056] In one possible implementation, the device further includes:
[0057] The first state module is used to set the state of the voice device to resource playback state when it is detected that the playback queue is not empty.
[0058] The second state module is used to set the state of the voice device to idle state when the playback queue is detected to be empty.
[0059] In one possible implementation, the intent type includes a second intent type;
[0060] The strategy determination module is specifically used for:
[0061] If the intent type to which the user intent belongs is determined to be the second intent type, the target control strategy is determined to control the voice device not to perform the device control operation corresponding to the user intent;
[0062] The control module is specifically used for:
[0063] If it is determined that the user's intent belongs to the second intent type, the voice device is controlled not to perform the device control operation corresponding to the user's intent.
[0064] In one possible implementation, the intent type includes a third intent type;
[0065] The strategy determination module is specifically used for:
[0066] If it is determined that the user intent belongs to the third intent type, the target control strategy is to control the voice device to play the device status parameters corresponding to the user intent according to the current state.
[0067] The control module is specifically used for:
[0068] If the intent type to which the user intent belongs is determined to be a third intent type, the current state of the voice device is determined.
[0069] If the current state is an idle state, then control the voice device to play the device state parameters corresponding to the user's intent;
[0070] The method further includes:
[0071] If the current state is a resource playback state, then the voice device is controlled not to play the device state parameters corresponding to the user's intent.
[0072] Thirdly, embodiments of the present invention provide an electronic device, including: a processor and a memory, wherein the processor is configured to execute a control program for a voice device stored in the memory to implement the control method for the voice device described in any one of the first aspects.
[0073] Fourthly, embodiments of the present invention provide a storage medium storing one or more programs, which can be executed by one or more processors to implement the control method of the voice device described in any one aspect.
[0074] The technical solution provided in this invention acquires voice commands collected by a voice device, performs voice recognition on the voice commands to obtain the user intent corresponding to the voice commands, and then determines the intent type to which the user intent belongs. Finally, based on the intent type to which the user intent belongs, a target control strategy is determined, and the voice device is controlled according to the target control strategy. Compared with the prior art, because the intent type to which the user intent belongs is determined first, and then the target control strategy is determined based on the intent type to which the user intent belongs, it avoids indiscriminately responding to repeated control from multiple users, effectively reducing the performance loss of the voice device, improving the intelligence of the voice device, and enriching the application scenarios of the voice device. Attached Figure Description
[0075] Figure 1 This is a schematic diagram illustrating an application scenario of the voice device control method provided in an embodiment of the present invention;
[0076] Figure 2 A flowchart illustrating an embodiment of a control method for a voice device provided by the present invention;
[0077] Figure 3 A flowchart illustrating an embodiment of another voice device control method provided by the present invention;
[0078] Figure 4 A flowchart illustrating another embodiment of a control method for a voice device provided by the present invention;
[0079] Figure 5 A flowchart illustrating another embodiment of a control method for a voice device provided by an embodiment of the present invention;
[0080] Figure 6 A flowchart illustrating an embodiment of a control method for a voice device provided by the present invention;
[0081] Figure 7 A block diagram illustrating an embodiment of a control device for a voice device provided in this invention;
[0082] Figure 8 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention. Detailed Implementation
[0083] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0084] The following is an exemplary description of the application scenarios of the voice device control method provided in the embodiments of the present invention:
[0085] See Figure 1 This is a schematic diagram illustrating an application scenario of the voice device control method provided in an embodiment of the present invention. Figure 1 The application scenarios shown include: smart speaker 11, smart TV 12, smart air conditioner 13, and cloud server 14.
[0086] Voice devices refer to devices that support voice control technology, including but not limited to: smart speakers 11, smart TVs 12, smart air conditioners 13, etc. Figure 1 In the application scenario shown, taking smart speaker 11 as an example, smart speaker 11 can collect voice commands from the surrounding environment and send the voice commands to cloud server 14. Cloud server 14 parses the voice commands, obtains the corresponding control commands, and sends them to smart speaker 11. Smart speaker 11 executes the control commands to realize voice control of smart speaker 11.
[0087] As an example application scenario, voice devices can be used in home life scenarios, such as living rooms and bedrooms.
[0088] As another exemplary application scenario, voice devices can be used in public settings, such as barbershops and restaurants.
[0089] It is understood that in real-life scenarios, the number and type of voice devices can be set by the user according to actual needs, and this embodiment of the invention does not impose any restrictions on this.
[0090] The cloud server 14 is the cloud server to which the voice device is connected, and the method provided in the embodiments of the present invention can be used to control the voice device.
[0091] In addition, the cloud server 14 can also provide other basic services to voice devices, such as updates and upgrades, and setting reminders.
[0092] The control method for the voice device provided by the present invention will be further explained below with reference to the accompanying drawings and specific embodiments. These embodiments do not constitute a limitation on the embodiments of the present invention.
[0093] See Figure 2 This is a flowchart illustrating an embodiment of a control method for a voice device provided by an embodiment of the present invention. Figure 2 As shown, the method may include the following steps:
[0094] Step 201: Obtain the voice commands collected by the voice device.
[0095] exist Figure 1In the application scenario shown, the smart speaker 11 can collect voice commands from the surrounding environment and send the collected voice commands to the cloud server 14, which can then obtain the voice commands collected by the smart speaker 11.
[0096] Step 202: Perform speech recognition on the voice command to obtain the user intent corresponding to the voice command.
[0097] Speech recognition technology is a technology that converts audio into text, similar to the human ear, which understands what others are saying. Speech recognition can convert voice commands into text.
[0098] The cloud server 14 can perform speech recognition on the above voice commands, parse the speech recognition text, and obtain semantic analysis results. Based on the semantic analysis results, the user intent corresponding to the voice command is obtained.
[0099] Step 203: Determine the intent type to which the user's intent belongs.
[0100] In this embodiment of the invention, the above-mentioned intent types can be classified into the following three types: a first intent type, a second intent type, and a third intent type. Specifically, the first intent type may refer to resource acquisition (such as acquiring audio, acquiring video, etc.), the second intent type may refer to device control (such as turning on the device, turning off the device, adjusting the volume, etc.), and the third intent type may refer to status query (such as link status query, running status query, etc.).
[0101] After receiving the user intent corresponding to the voice command, the cloud server 14 can determine the intent type to which the user intent belongs.
[0102] Optionally, the cloud server can determine the intent type of a user's intent through keyword matching. Table 1 below shows an example of the correspondence between intent types and keywords:
[0103] Table 1
[0104]
[0105]
[0106] The intent type classification method given above and the corresponding relationship in Table 1 are merely examples. Specific settings can be made by professionals, and this invention does not impose any limitations on them.
[0107] Optionally, the scope of each intent type can be adjusted according to actual needs, and a fourth intent type can be added to refer to dialogue (such as request-response, chat, etc.).
[0108] For example, if the user's intent is to play a song, the cloud server 14 can determine that the intent type to which the user's intent belongs is the first intent type.
[0109] Step 204: Determine the target control strategy based on the intent type to which the user's intent belongs.
[0110] Step 205: Control the voice device according to the target control strategy.
[0111] Steps 204 and 205 will be described in a unified manner below:
[0112] In this embodiment of the invention, the cloud server 14 can determine the target control strategy according to the intent type to which the user's intent belongs, and control the smart speaker 11 according to the target control strategy.
[0113] Specifically, how the target control strategy is determined based on the type of user intent, and how the voice device is controlled according to the target control strategy, will be explained below. Figure 4 The process shown will be explained in detail here.
[0114] This concludes the process. Figure 2 The process described is as follows.
[0115] pass Figure 2 As shown in the flowchart, the technical solution provided in this embodiment of the invention acquires voice commands collected by a voice device, performs voice recognition on the voice commands to obtain the user intent corresponding to the voice commands, and then determines the intent type to which the user intent belongs. Finally, based on the intent type to which the user intent belongs, a target control strategy is determined, and the voice device is controlled according to the target control strategy. Compared with the prior art, because the intent type to which the user intent belongs is determined first, and then the target control strategy is determined based on the intent type to which the user intent belongs, it avoids indiscriminately responding to repeated control from multiple users, effectively reducing the performance loss of the voice device, improving the intelligence of the voice device, and enriching the application scenarios of the voice device.
[0116] See Figure 3 This is a flowchart illustrating an embodiment of a control method for a voice device provided by the present invention. Figure 3 As shown, the method may include the following steps:
[0117] Step 301: Obtain the voice commands collected by the voice device.
[0118] Step 302: Perform speech recognition on the voice command to obtain the user intent corresponding to the voice command.
[0119] For a detailed description of steps 301 and 302, please refer to the above. Figure 2The relevant descriptions of steps 201 and 202 in the embodiments will not be repeated here.
[0120] Step 303: Determine the operating mode configured for the voice device. If the voice device is determined to be configured in the first operating mode, proceed to step 304; if the voice device is determined to be configured in the second operating mode, proceed to step 307.
[0121] Step 304: Determine the intent type to which the user's intent belongs.
[0122] Step 305: Determine the target control strategy based on the intent type to which the user's intent belongs.
[0123] Step 306: Control the voice device according to the target control strategy.
[0124] Step 307: Generate control commands corresponding to the user's intent and send the control commands to the voice device so that the voice device can execute the control commands.
[0125] After executing step 307, the process ends.
[0126] Steps 303 to 307 are described in a unified manner below:
[0127] The first working mode described above is a mode that does not respond to user requests for device control and can be applied in public scenarios. The second working mode is a normal working mode that responds to various user requests indiscriminately and can be applied in home scenarios.
[0128] Based on this, in this embodiment of the invention, after the cloud server acquires the voice commands collected by the voice device, performs voice recognition on the voice commands to obtain the user intent corresponding to the voice commands, it first determines the working mode configured for the voice device. If the voice device is determined to be configured in the first working mode, the intent type of the user intent is then determined. Based on the intent type of the user intent, a target control strategy is determined to control the voice device according to the target control strategy.
[0129] If it is determined that the voice device is configured in the second working mode, a control command corresponding to the user's intent is generated and sent to the voice device for execution.
[0130] Specifically, how the target control strategy is determined based on the type of user intent, and how the voice device is controlled according to the target control strategy, will be explained below. Figure 4 The process shown will be explained in detail here.
[0131] This concludes the process. Figure 3 The process described is as follows.
[0132] pass Figure 3 As shown in the flowchart, the technical solution provided by this embodiment of the invention determines the configured working mode of the voice device. When the voice device is configured in a first working mode, it executes steps to determine the intent type of the user's intent and, based on the intent type, determine a target control strategy, and control the voice device according to the target control strategy. When the voice device is configured in a second working mode, it generates a control command corresponding to the user's intent and sends the control command to the voice device for execution. Because different control strategies are adopted for different working modes of the voice device, the voice device can be used in both home and public environments, improving its applicability and enriching its application scenarios.
[0133] See Figure 4 This is a flowchart illustrating another embodiment of a voice device control method provided by the present invention. Figure 4 As shown, the method may include the following steps:
[0134] Step 401: Obtain the voice commands collected by the voice device.
[0135] Step 402: Perform speech recognition on the voice command to obtain the user intent corresponding to the voice command.
[0136] For a detailed description of steps 401 to 402, please refer to the above. Figure 2 The relevant descriptions of steps 201 to 202 in the embodiments will not be repeated here.
[0137] Step 403: Determine the intent type to which the user intent belongs. If the intent type to which the user intent belongs is the first intent type, proceed to step 404; if the intent type to which the user intent belongs is the second intent type, proceed to step 407; if the intent type to which the user intent belongs is the third intent type, proceed to step 409.
[0138] Step 404: Determine the target control strategy as controlling the voice device to play resources corresponding to the user's intent.
[0139] 405. Determine the resources corresponding to the user's intent.
[0140] 406. Add the resources to the playback queue corresponding to the voice device, and sequentially retrieve the resources from the playback queue and send them to the voice device for playback.
[0141] Steps 404 to 406 are described in a unified manner below:
[0142] In this embodiment of the invention, when the cloud server determines that the user's intent belongs to a first intent type, it determines the target control strategy as controlling the voice device to play the resource corresponding to the user's intent based on the first intent type (e.g., resource acquisition). The cloud server then controls the voice device according to the aforementioned target control strategy.
[0143] Specifically, the cloud server can first determine the resource corresponding to the user's intent, then add the resource to the playback queue corresponding to the voice device, and sequentially retrieve the resource from the playback queue and send it to the voice device for playback.
[0144] In this way, by introducing a streaming resource playback control scheme, the resource on-demand function is completed. At the same time, the streaming resources are delivered to the voice device without interrupting the resources currently being played on the voice device, thus optimizing the user experience of the voice device.
[0145] In an optional embodiment, when the cloud server detects that the playback queue is not empty, the state of the locally recorded voice device is set to resource playback state; when the cloud server detects that the playback queue is empty, the state of the locally recorded voice device is set to idle state.
[0146] Therefore, the cloud server can update the status of the voice device in real time, providing a basis for the cloud server to determine different control strategies for the same voice command under different statuses of the voice device.
[0147] Step 407: Determine the target control strategy as controlling the voice device to not perform device control operations corresponding to the user's intent.
[0148] Step 408: Control the voice device to not perform device control operations corresponding to the user's intent.
[0149] Steps 407 and 408 will be described in a unified manner below:
[0150] In this embodiment of the invention, when the cloud server determines that the intent type to which the user intent belongs is a second intent type, it determines the target control strategy as controlling the voice device not to perform the device control operation corresponding to the user intent based on the second intent type (e.g., device control type).
[0151] The cloud server controls the voice device according to the above-mentioned target control strategy, that is, it controls the voice device not to perform device control operations corresponding to the user's intent.
[0152] In this way, voice commands that do not respond to device control intentions can be implemented, effectively avoiding repeated control of the voice device by multiple users and reducing the performance loss of the voice device.
[0153] Step 409: Determine the target control strategy as controlling the voice device to play device status parameters corresponding to the user's intent based on the current state.
[0154] Step 410: Determine the current state of the voice device. If the current state is idle, proceed to step 411; if the current state is in resource playback mode, proceed to step 412.
[0155] Step 411: Control the voice device to play the device status parameters corresponding to the user's intent.
[0156] Step 412: Control the voice device to not play the device status parameters corresponding to the user's intent.
[0157] Steps 409 to 412 are described in a unified manner below:
[0158] In this embodiment of the invention, when the cloud server determines that the user's intent belongs to a third intent type, it determines the target control strategy based on the third intent type (e.g., a status query type) to control the voice device to play the device status parameters corresponding to the user's intent according to the current status. The cloud server then controls the voice device according to the aforementioned target control strategy.
[0159] Specifically, the cloud server can first determine the current state of the voice device; if the current state is idle, it controls the voice device to play the device state parameters corresponding to the user's intent. If the current state is resource playback, it controls the voice device not to play the device state parameters corresponding to the user's intent.
[0160] This allows the system to respond to status query commands only when the voice device is idle, thus avoiding interruption of the resources currently being played by the voice device and improving the intelligence of the voice device.
[0161] Optionally, if other intent types exist, the specific control strategy can be determined based on whether or not it will interrupt the playback of the current resource.
[0162] For example, if a fourth intent type exists (such as dialogue), a control strategy similar to that for the second intent type can be adopted to prevent the voice device from executing a response corresponding to the user's intent, so as to ensure that the current resource playback is not interrupted.
[0163] This concludes the process. Figure 4 The process described is as follows.
[0164] pass Figure 4As shown in the flowchart, the technical solution provided by this embodiment of the invention, when determining that the user's intent belongs to a first intent type, determines the target control strategy as controlling the voice device to play resources corresponding to the user's intent. Then, it determines the resources corresponding to the user's intent, adds the resources to the playback queue corresponding to the voice device, and sequentially retrieves resources from the playback queue and sends them to the voice device for playback. When determining that the user's intent belongs to a second intent type, it determines the target control strategy as controlling the voice device not to execute device control operations corresponding to the user's intent. When determining that the user's intent belongs to a third intent type, it determines the target control strategy as controlling the voice device to play device status parameters corresponding to the user's intent based on the current state. Then, it determines the current state of the voice device; if the current state is idle, it controls the voice device to play the device status parameters corresponding to the user's intent; if the current state is a resource playback state, it controls the voice device not to play the device status parameters corresponding to the user's intent. By responding to voice commands related to resource acquisition but not to voice commands related to device control, and only responding to voice commands related to status queries when the device is idle, the device can effectively avoid wear and tear and potential disputes caused by repeated control by multiple users. It also avoids the device indiscriminately issuing control commands and interrupting the current content every time a resource is played, thus improving the intelligence of the voice device and optimizing the user experience of the voice device in public scenarios.
[0165] See Figure 5 This is a flowchart illustrating another embodiment of a voice device control method provided by the present invention. Figure 5 As shown, the method may include the following steps:
[0166] Figure 5 The illustrated embodiment can be effective in scenarios with voice devices. Regardless of whether the user is in a public setting, they can access the voice device and turn on the switch for the solution shown in this embodiment through an APP (Application) on a mobile phone or other associated terminal.
[0167] Optionally, the solution activation switch shown in this embodiment can also be turned on directly on the voice device.
[0168] When a voice device receives a voice command input by a user, the cloud service obtains the voice command through the voice device and first performs speech recognition and semantic parsing.
[0169] If semantic parsing fails, the input voice command is discarded and determined to be a non-user intent; if semantic parsing succeeds, the corresponding NluData (Natural Language Understanding Data) is obtained.
[0170] The NluData function is used to determine whether the user intent corresponding to the voice command is a device control intent.
[0171] If the intent is a device control intent (e.g., turn on the air conditioner), the voice command is discarded and not executed.
[0172] If the intent is not device control, NluData is used to determine if the user intent corresponding to the voice command is a status query intent. If it is a status query intent (e.g., "What is the current status of the air conditioner?"), then the status of the voice device is checked to see if it is in an idle state.
[0173] If the voice device is in resource playback mode, the voice command is discarded and not executed. If the device is idle, the connection status between the device and the cloud server is queried, a response corresponding to the voice command is generated, and the response is sent to the voice device for playback.
[0174] If the intent is not a status query, and the user intent corresponding to the voice command is determined to be a resource acquisition intent through NluData, then the resource list corresponding to the voice command is obtained and added to the cloud server playback queue so that the voice device can play the resource.
[0175] Furthermore, if it is determined from the above steps that the user intent corresponding to the voice command is not a device control intent or a status query intent, an attempt can be made to obtain resources based on the request content in the voice command. If the resources are successfully obtained, they are added to the cloud server playback queue and then distributed. If the resources are not successfully obtained, the voice command is discarded and not executed.
[0176] Understandably, if it is determined that the user intent corresponding to the voice command is another type of intent, such as casual conversation or checking the weather, the voice command can be discarded, not executed, and no response can be made.
[0177] Specifically, it can be done through Figure 6 The illustrated embodiment provides another method for controlling a voice device to control the playback of voice device resources.
[0178] If the cloud server's playback queue is not empty, or when data is inserted into the cloud server's playback queue, it can be determined whether the voice device and the cloud server are in a state of maintaining a connection.
[0179] Optionally, the inserted data can be resources obtained through voice commands or resources uploaded to the cloud server by the user through the associated terminal of the voice device (such as a mobile APP).
[0180] Existing technologies can be used to connect cloud servers, voice devices, and associated terminals. This allows users to configure voice devices, manage their status, and distribute on-demand resources on associated terminals.
[0181] If the voice device and the cloud server are not in a persistent connection state, establish the connection between them first. When the voice device and the cloud server are in a persistent connection state, retrieve the transmitted resources based on the voice device status and send them to the voice device in sequence, while marking the voice device status as resource playback status.
[0182] In other words, the cloud server controls the voice device to broadcast resources by streaming resources. After each resource is delivered, it will check again whether the playback queue is empty.
[0183] If the playback queue is not empty, continue to execute the above steps of "obtaining the voice device status and sending the resources to the voice device according to the order, and marking the voice device status as the resource playback status".
[0184] If the playback queue is empty, disconnect the connection to the voice device and mark the voice device as idle.
[0185] This concludes the process. Figure 5 The process described is as follows.
[0186] The above embodiments are based on voice control devices and cloud data parsing application services. They only process resource requests and stream them for playback, thus completing the resource on-demand function. This avoids interrupting the resources currently being played by the voice device, thereby avoiding the performance loss and potential disputes caused by multiple users controlling the device repeatedly at the same time, and optimizing the user experience of voice devices in public scenarios.
[0187] Corresponding to the aforementioned embodiments of the control method for voice devices, the present invention also provides embodiments of a control device for voice devices.
[0188] See Figure 7 This is a block diagram illustrating an embodiment of a control device for a voice device provided by an embodiment of the present invention. Figure 7 As shown, the device includes:
[0189] The instruction acquisition module 71 is used to acquire voice instructions collected by the voice device;
[0190] The voice recognition module 72 is used to perform voice recognition on the voice command to obtain the user intent corresponding to the voice command;
[0191] The type determination module 73 is used to determine the intent type to which the user intent belongs;
[0192] The strategy determination module 74 is used to determine the target control strategy based on the intent type to which the user intent belongs;
[0193] The control module 75 is used to control the voice device according to the target control strategy.
[0194] In one possible implementation, the device further includes (not shown in the figure):
[0195] The mode determination module is used to determine the operating mode configured for the voice device;
[0196] The first execution module is configured to perform the step of determining the intent type to which the user intent belongs when it is determined that the voice device is configured in a first working mode.
[0197] In one possible implementation, the device further includes (not shown in the figure):
[0198] The second execution module is configured to, when determining that the voice device is configured in a second working mode, generate a control command corresponding to the user's intent and send the control command to the voice device so that the voice device executes the control command.
[0199] In one possible implementation, the intent type includes a first intent type;
[0200] The strategy determination module 74 is specifically used for:
[0201] If the intent type to which the user intent belongs is determined to be the first intent type, the target control strategy is determined to be to control the voice device to play the resources corresponding to the user intent;
[0202] The control module 75 is specifically used for:
[0203] If the intent type to which the user intent belongs is determined to be the first intent type, then the resource corresponding to the user intent is determined.
[0204] The resources are added to the playback queue corresponding to the voice device, and resources are sequentially retrieved from the playback queue and sent to the voice device for playback.
[0205] In one possible implementation, the device further includes (not shown in the figure):
[0206] The first state module is used to set the state of the voice device to resource playback state when it is detected that the playback queue is not empty.
[0207] The second state module is used to set the state of the voice device to idle state when the playback queue is detected to be empty.
[0208] In one possible implementation, the intent type includes a second intent type;
[0209] The strategy determination module 74 is specifically used for:
[0210] If the intent type to which the user intent belongs is determined to be the second intent type, the target control strategy is determined to control the voice device not to perform the device control operation corresponding to the user intent;
[0211] The control module 75 is specifically used for:
[0212] If it is determined that the user's intent belongs to the second intent type, the voice device is controlled not to perform the device control operation corresponding to the user's intent.
[0213] In one possible implementation, the intent type includes a third intent type;
[0214] The strategy determination module 74 is specifically used for:
[0215] If it is determined that the user intent belongs to the third intent type, the target control strategy is to control the voice device to play the device status parameters corresponding to the user intent according to the current state.
[0216] The control module 75 is specifically used for:
[0217] If the intent type to which the user intent belongs is determined to be a third intent type, the current state of the voice device is determined.
[0218] If the current state is an idle state, then control the voice device to play the device state parameters corresponding to the user's intent;
[0219] The method further includes:
[0220] If the current state is a resource playback state, then the voice device is controlled not to play the device state parameters corresponding to the user's intent.
[0221] Figure 8 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention. Figure 8The illustrated electronic device 800 includes at least one processor 801, a memory 802, at least one network interface 804, and a user interface 803. The various components in the electronic device 800 are coupled together via a bus system 805. It is understood that the bus system 805 is used to implement communication between these components. In addition to a data bus, the bus system 805 also includes a power bus, a control bus, and a status signal bus. However, for clarity, ... Figure 8 The general labeled all buses as Bus System 805.
[0222] The user interface 803 may include a display, keyboard or clicking device (e.g., mouse, trackball), touchpad or touch screen.
[0223] It is understood that the memory 802 in the embodiments of the present invention can be volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. The non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. The volatile memory can be random access memory (RAM), which is used as an external cache. By way of example, but not limitation, many forms of RAM are available, such as Static Random Access Memory (SRAM), Dynamic Random Access Memory (DRAM), Synchronous DRAM (SDRAM), Double Data Rate SDRAM (DDRSDRAM), Enhanced Synchronous DRAM (ESDRAM), Synchlink DRAM (SLDRAM), and Direct Rambus RAM (DRRAM). The memory 802 described herein is intended to include, but is not limited to, these and any other suitable types of memory.
[0224] In some implementations, memory 802 stores elements, executable units or data structures, or subsets thereof, or extended sets thereof: operating system 8021 and application programs 8022.
[0225] The operating system 8021 includes various system programs, such as the framework layer, core library layer, and driver layer, used to implement various basic business functions and handle hardware-based tasks. The application program 8022 includes various applications, such as a media player and a browser, used to implement various application functions. The program implementing the method of this embodiment can be included in the application program 8022.
[0226] In this embodiment of the invention, by calling the program or instructions stored in the memory 802, specifically the program or instructions stored in the application program 8022, the processor 801 executes the method steps provided in each method embodiment, including, for example:
[0227] Acquire voice commands collected by the voice device;
[0228] The voice command is subjected to speech recognition to obtain the user intent corresponding to the voice command;
[0229] Determine the intent type to which the user intent belongs;
[0230] Determine the target control strategy based on the intent type to which the user intent belongs;
[0231] The voice device is controlled according to the target control strategy.
[0232] The methods disclosed in the above embodiments of the present invention can be applied to or implemented by processor 801. Processor 801 may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by the integrated logic circuit of the hardware in processor 801 or by instructions in the form of software. The processor 801 may be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of the present invention. The general-purpose processor may be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of the present invention can be directly embodied in the execution of a hardware decoding processor, or executed by a combination of hardware and software units in the decoding processor. The software units may be located in random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, or other mature storage media in the art. The storage medium is located in memory 802. Processor 801 reads the information in memory 802 and, in conjunction with its hardware, completes the steps of the above method.
[0233] It is understood that the embodiments described herein can be implemented in hardware, software, firmware, middleware, microcode, or a combination thereof. For hardware implementation, the processing unit can be implemented in one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), general-purpose processors, controllers, microcontrollers, microprocessors, other electronic units for performing the functions described herein, or combinations thereof.
[0234] For software implementation, the techniques described herein can be implemented by units that perform the functions described herein. The software code can be stored in memory and executed by a processor. The memory can be implemented in the processor or external to the processor.
[0235] The electronic device provided in this embodiment may be as follows: Figure 8 The electronic device shown can perform the following: Figure 2-6 All steps of the control method for Chinese voice devices, thereby achieving Figure 2-6 For details on the technical effects of the control method for the voice device shown, please refer to [link / reference]. Figure 2-6 The relevant descriptions are presented concisely and will not be elaborated upon here.
[0236] This invention also provides a storage medium (computer-readable storage medium). This storage medium stores one or more programs. The storage medium may include volatile memory, such as random access memory; the memory may also include non-volatile memory, such as read-only memory, flash memory, hard disk, or solid-state drive; the memory may also include combinations of the above types of memory.
[0237] When one or more programs in the storage medium can be executed by one or more processors to implement the above-described control method for a voice device executed on the electronic device side.
[0238] The processor is used to execute a control program for the voice device stored in the memory to implement the following steps of a voice device control method executed on the electronic device side:
[0239] Acquire voice commands collected by the voice device;
[0240] The voice command is subjected to speech recognition to obtain the user intent corresponding to the voice command;
[0241] Determine the intent type to which the user intent belongs;
[0242] Determine the target control strategy based on the intent type to which the user intent belongs;
[0243] The voice device is controlled according to the target control strategy.
[0244] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this invention.
[0245] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein can be implemented in hardware, a software module executed by a processor, or a combination of both. The software module can be located in random access memory (RAM), main memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium known in the art.
[0246] The specific embodiments described above further illustrate the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the above description is only a specific embodiment of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A control method of a voice device, characterized by, The voice device is applied in a public scene, and the method is used to control the voice device resource to play, and the method comprises: In the case of repeated control of multiple users, the voice instruction collected by the voice device is obtained; The voice instruction is subjected to voice recognition to obtain the user intent corresponding to the voice instruction; The intent type to which the user intent belongs is determined; According to the intent type to which the user intent belongs, a target control strategy is determined; wherein the intent type comprises a first intent type, a second intent type and a third intent type; the first intent type is a resource acquisition type; the second intent type is a device control type; the third intent type is a state query type; according to the intent type to which the user intent belongs, the target control strategy is determined, which comprises: in the case that the intent type to which the user intent belongs is determined to be the first intent type, the target control strategy is determined to be to control the voice device to play the resource corresponding to the user intent; in the case that the intent type to which the user intent belongs is determined to be the second intent type, the target control strategy is determined to be to control the voice device not to perform the device control operation corresponding to the user intent; in the case that the intent type to which the user intent belongs is determined to be the third intent type, the target control strategy is determined to be to control the voice device to play the device state parameter corresponding to the user intent according to the current state; The voice device is controlled according to the target control strategy.
2. The method of claim 1, wherein, Before the intent type to which the user intent belongs is determined, it further comprises: The working mode configured for the voice device is determined; In the case that the voice device is determined to be configured as a first working mode, the step of determining the intent type to which the user intent belongs is performed.
3. The method of claim 2, wherein, After the working mode configured for the voice device is determined, it further comprises: In the case that the voice device is determined to be configured as a second working mode, a control instruction corresponding to the user intent is generated, and the control instruction is issued to the voice device to be executed by the voice device; the process is ended.
4. The method of claim 1, wherein, The voice device is controlled according to the target control strategy, which comprises: In the case that the intent type to which the user intent belongs is determined to be the first intent type, the resource corresponding to the user intent is determined; The resource is added to the play queue corresponding to the voice device, and the resource is sequentially obtained from the play queue and issued to the voice device for playing by the voice device.
5. The method of claim 4, wherein, The method further comprises: When it is detected that the play queue is not empty, the state of the voice device is set to a resource playing state; When it is detected that the play queue is empty, the state of the voice device is set to an idle state.
6. The method of claim 1, wherein, The voice device is controlled according to the target control strategy, which comprises: In the case that the intent type to which the user intent belongs is determined to be the second intent type, the voice device is controlled not to perform the device control operation corresponding to the user intent.
7. The method of claim 1, wherein, The voice device is controlled according to the target control strategy, which comprises: determining a current state of the voice device in a case where the intent type to which the user intent belongs is determined to be a third intent type; controlling the voice device to play a device state parameter corresponding to the user intent in a case where the current state is an idle state; The method further comprises: controlling the voice device not to play the device state parameter corresponding to the user intent in a case where the current state is a resource playing state.
8. A control device for a voice device, characterized by The voice device is applied in a public scene, and the device is used to control resource playing of the voice device, and the device comprises: an instruction obtaining module configured to obtain a voice instruction collected by the voice device in a case where multiple users repeatedly control; a voice recognition module configured to perform voice recognition on the voice instruction to obtain a user intent corresponding to the voice instruction; a type determining module configured to determine an intent type to which the user intent belongs; a strategy determining module configured to determine a target control strategy according to the intent type to which the user intent belongs; wherein the intent type comprises a first intent type, a second intent type and a third intent type; the first intent type is a resource obtaining type; the second intent type is a device control type; the third intent type is a state query type; the determination of the target control strategy according to the intent type to which the user intent belongs comprises: in a case where the intent type to which the user intent belongs is determined to be the first intent type, determining that the target control strategy is to control the voice device to play a resource corresponding to the user intent; in a case where the intent type to which the user intent belongs is determined to be the second intent type, determining that the target control strategy is to control the voice device not to perform a device control operation corresponding to the user intent; and in a case where the intent type to which the user intent belongs is determined to be the third intent type, determining that the target control strategy is to control the voice device to play a device state parameter corresponding to the user intent according to a current state; a control module configured to control the voice device according to the target control strategy.
9. An electronic device, comprising: comprise: a processor and a memory, wherein the processor is configured to execute a control program of a voice device stored in the memory, so as to implement the control method of the voice device according to any one of claims 1 to 7.
10. A storage medium, characterized by The storage medium stores one or more programs, and the one or more programs are executable by one or more processors to implement the control method of the voice device according to any one of claims 1 to 7.
Citation Information
Patent Citations
Playing control method, device and intelligent apparatus
CN107146608A
Man-machine interaction method and device, electronic equipment and storage medium
CN110689889A
Voice control method and device, electronic equipment and readable storage medium
CN112581945A