Device control method, device, server and storage medium
By using the NLU model in the first server to parse the voice command and determine the execution device, the problem of high delay in execution of voice commands across devices is solved, and the effect of quickly establishing connections and improving voice control efficiency is achieved.
Patent Information
- Application Number
- CN202111266199.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-10-28
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2041-10-28
AI Technical Summary
In the prior art, the delay of voice commands to be executed across devices is high, resulting in low voice interaction efficiency.
By analyzing voice instructions using a natural language understanding (NLU) model in the first server, the type of execution device is determined, and a third terminal device is determined in the second terminal device with a set relationship with the first terminal device to establish a quick connection and execute the instructions.
The decision-making speed of the cloud-based voice command execution device is improved, and the time for the first terminal device to establish a connection between the third terminal device is shortened, thereby improving the voice control efficiency during cross-device execution.
Smart Images

Figure CN114005443B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of communication technology, and in particular to a device control method, device, server and storage medium. Background Art
[0002] At present, electronic devices equipped with voice assistants can support voice interaction operations with users. In actual applications, there are scenarios where voice interaction operations need to be performed across devices, that is, the wake-up device receives a voice command, and the execution device performs the related operations indicated by the voice command. In related technologies, the delay of executing voice commands across devices is relatively high. Summary of the invention
[0003] In view of this, embodiments of the present application provide a device control method, apparatus, server, and storage medium to at least solve the problem of high delay in executing voice commands across devices in the related art.
[0004] To achieve the above purpose, the technical solution of this application is implemented as follows:
[0005] The present application provides a device control method, which is applied to a first server, and the method includes:
[0006] Receiving a first instruction sent by a first terminal device; the first instruction is used to instruct the corresponding terminal device to perform a first operation;
[0007] Inputting the first instruction into a first natural language understanding (NLU) model to obtain at least one first device type output by the first NLU model; the first NLU model is used to perform intent analysis on the input voice instruction regarding the execution device, and output the device type obtained by the analysis for executing the corresponding voice instruction;
[0008] In a case where the first terminal device does not belong to the first device type, determining a third terminal device from at least one second terminal device having a set relationship with the first terminal device according to at least one first device type;
[0009] A first message is sent to the first terminal device, so that the first terminal device establishes a first connection with the third terminal device based on the first message, and controls the third terminal device to perform the first operation through the first connection.
[0010] The present application also provides a device control apparatus, including:
[0011] A first receiving unit, configured to receive a first instruction sent by a first terminal device; the first instruction is used to instruct the corresponding terminal device to perform a first operation;
[0012] a first input unit, configured to input a first instruction into a first NLU model, and obtain at least one first device type output by the first NLU model; the first NLU model is configured to perform intent analysis on the input voice instruction regarding the execution device, and output the device type obtained by the analysis for executing the corresponding voice instruction;
[0013] a first processing unit, configured to determine, if the first terminal device does not belong to the first device type, a third terminal device from at least one second terminal device having a set relationship with the first terminal device according to at least one first device type;
[0014] The first sending unit is used to send a first message to the first terminal device, so that the first terminal device establishes a first connection with the third terminal device based on the first message, and controls the third terminal device to perform the first operation through the first connection.
[0015] The embodiment of the present application further provides a server, comprising: a processor and a memory for storing a computer program that can be run on the processor,
[0016] Wherein, the processor is used to execute the steps of the above-mentioned device control method when running the computer program.
[0017] An embodiment of the present application further provides a storage medium having a computer program stored thereon, wherein the computer program implements the steps of the above-mentioned device control method when executed by a processor.
[0018] In an embodiment of the present application, the first server parses the intention of the execution device of the first instruction through the first NLU model, obtains at least one first device type corresponding to the execution of the first instruction, and when the first terminal device does not belong to the first device type, determines whether it is a cross-device execution according to at least one first device type. When the judgment result is a cross-device execution, the first server determines the third terminal device that executes the first operation from at least one second terminal device that has a set relationship with the first terminal device according to at least one first device type, and sends the first information to the first terminal device so that the first terminal device establishes a first connection with the third terminal device. In this way, the first server determines the third terminal device that executes the first operation by adopting an NLU model for parsing the intention of the execution device of the instruction, and on the basis of the parsing result of the NLU model and the terminal device information provided by the first terminal device, thereby improving the decision speed of the cloud to determine the execution device of the first instruction, so that the first terminal device can quickly establish a connection with the third terminal device, and then complete the cross-device execution of the instruction, thereby improving the voice control efficiency during cross-device execution. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Figure 1 A schematic diagram of a device control method in the related art;
[0020] Figure 2 A schematic diagram of a device control system provided in an embodiment of the present application;
[0021] Figure 3 A schematic diagram of the implementation flow of the device control method provided in the embodiment of the present application;
[0022] Figure 4 A schematic diagram of the implementation flow of the device control method provided in the application embodiment of the present application;
[0023] Figure 5 A schematic diagram of the structure of the device control structure provided in the embodiment of the present application;
[0024] Figure 6 A schematic diagram of the hardware composition structure of the server provided in an embodiment of the present application. DETAILED DESCRIPTION
[0025] In daily life, users can use their mobile phones equipped with voice assistants to control smart TVs. Figure 1 In the implementation scheme shown, the mobile phone receives the user's voice command and sends the voice command to the cloud server; the dialogue system of the cloud server obtains the classification result according to the voice command, executes the dialogue management process, and sends the returned dialogue management result to the mobile phone; based on the dialogue management result, the mobile phone sends the voice command to the smart TV, and the smart TV sends the voice command and the relevant status information of the smart TV to the server, such as the playback status, playback mode, etc., and the dialogue system of the server returns the corresponding protocol according to the voice command and the relevant status information. In this way, the smart TV can obtain the corresponding resources according to the returned protocol and execute the corresponding playback action. When the cloud server parses the voice command according to the dialogue system, it parses the skill intention of the voice command through multiple skill NLU models in the dialogue system, and judges whether it is necessary to establish a connection between the wake-up device and the execution device based on the parsing result obtained by the skill NLU model. If the skill NLU model recognizes that a connection needs to be established between the wake-up device and the execution device, that is, the voice command is recognized by the skill NLU model as a connection instruction, then the operation of establishing a connection between the wake-up device and the execution device is triggered to achieve control of the execution device through the connection. To identify various skill intents, skill NLU models include but are not limited to: multimedia NLU model, used to identify playback functions; system setting NLU model, used to identify adjustment functions.
[0026] However, in order to parse the skill intent of the voice command and identify whether the voice command is a connection command, the cloud server needs to parse the voice command through the skill NLU model. Since different types of execution devices are usually deployed in the system, a large number of skill NLU models are set up in the dialogue system of the cloud server. Therefore, the cloud server will consume a lot of time in judging cross-device execution. The time required for the first terminal device to establish a connection with the third terminal device is relatively long, resulting in a high delay in the cross-device execution of voice commands.
[0027] Based on this, an embodiment of the present application provides a device control method, in which the first server parses the intention of the execution device of the first instruction through the first NLU model, obtains at least one first device type corresponding to the execution of the first instruction, and in the case where the first terminal device does not belong to the first device type, determines whether it is a cross-device execution according to at least one first device type. When the judgment result is a cross-device execution, the first server determines the third terminal device that executes the first operation from at least one second terminal device that has a set relationship with the first terminal device according to at least one first device type, and sends a first message to the first terminal device so that the first terminal device establishes a first connection with the third terminal device. In this way, the first server determines the third terminal device that executes the first operation by adopting the NLU model for parsing the intention of the execution device of the instruction, and on the basis of the parsing result of the NLU model and the terminal device information provided by the first terminal device, thereby improving the decision speed of the cloud to determine the execution device of the first instruction, so that the first terminal device can quickly establish a connection with the third terminal device, and then complete the cross-device execution of the instruction, thereby improving the voice control efficiency during cross-device execution.
[0028] In order to make the purpose, technical solution and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0029] Figure 2A schematic diagram of a device control system provided for an embodiment of the present application. Among them, a group of terminal devices consisting of at least two terminal devices is deployed in the device control system, and the terminal devices communicate with each other based on short-range wireless communication technology, and the terminal devices in the group have the ability to communicate with each other. Short-range wireless communication technology includes Bluetooth communication technology or wireless fidelity (WiFi, Wireless Fidelity) communication technology. When a group includes multiple terminal devices, one-to-one communication, one-to-many communication or many-to-one communication can be performed between the terminal devices based on short-range wireless communication technology. In actual application, the group can be established based on a third-party communication framework, such as an OAF communication framework (OPlusAccessory Framework) established based on a third-party protocol. A group consisting of at least two terminal devices communicates wirelessly with a server 21. In a cross-device execution scenario, a terminal device 22 that receives a voice command is usually referred to as a wake-up device, and a terminal device 23 that executes the operation indicated by the control command corresponding to the voice command is referred to as an execution device. For example, a user inputs a first voice command "Play ××× (TV series) on TV" to a mobile phone, the mobile phone is a wake-up device, and the smart TV is an execution device.
[0030] It should be noted that Figure 2 The terminal devices in the device control system shown are all installed with a first application, for example, the first application can be a voice assistant. The first application is used to support users to input voice commands, and to support the interaction between the terminal device and the cloud (i.e., server 21), and between terminal devices at the application layer. Based on the first application, the user can operate any terminal device in the group by inputting voice commands.
[0031] In practical application, for example, Figure 2 The terminal device may include at least one of the following:
[0032] Mobile phones, smart speakers, laptops, smart watches, tablets, TVs, refrigerators, air conditioners and other IoT terminals.
[0033] In the IoT scenario, with the family as the unit, the terminal devices in a group include at least various terminal devices and smart appliances used by family members. For example, the terminal devices used by family members and the smart appliances in the home are grouped based on short-range wireless communication technology. As a result, all the above-mentioned terminal devices can communicate based on short-range wireless communication technology, and voice commands can be issued to any terminal device in the group to control another terminal device to perform the operation corresponding to the voice command, wherein the wake-up device and execution device corresponding to the voice command do not have to be terminal devices that rely on the same account. In actual application, family members or guests can use different wake-up devices to voice control the smart TV.
[0034] The following will specifically describe the technical solution of the present application and how the technical solution of the present application solves the above technical problems through embodiments and in conjunction with the accompanying drawings. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments.
[0035] Figure 3 The following is a schematic diagram of the implementation process of the device control method provided in the embodiment of the present application. The execution subject of this process is the first server, which is Figure 2 The first terminal device is Figure 2 In an embodiment of the present application, the first terminal device is located in the first group, and each terminal device in the first group communicates based on the short-range wireless communication technology. The first group is a group; the first terminal device is a wake-up device, and the third terminal device is an execution device.
[0036] like Figure 3 As shown, the device control method includes:
[0037] Step 301: Receive a first instruction sent by a first terminal device; the first instruction is used to instruct the corresponding terminal device to perform a first operation.
[0038] In one embodiment, the first instruction can be represented as a voice instruction. When the user wants to perform voice control through the first terminal device, the user inputs the first instruction to the first terminal device through the first application in the first terminal device. The first server receives a message sent by the first terminal device carrying the first instruction. Among them, the first instruction includes an execution device intention and a skill intention. The execution device intention represents the execution device on which the user wants to execute the operation corresponding to the first instruction. The skill intention represents the goal that the user wants to achieve within the skills supported by the corresponding terminal device, which is used to instruct the corresponding terminal device to perform the first operation. Skills generally refer to the capabilities or achievable functions of a terminal device.
[0039] For example, when the first instruction is to turn up the volume of the TV, the execution device is the TV, and the skill intention is to turn up the volume of the playback.
[0040] Step 302: Input a first instruction into a first NLU model to obtain at least one first device type output by the first NLU model; the first NLU model is used to perform intent analysis on the input voice instruction regarding the execution device, and output the device type obtained by analysis for executing the corresponding voice instruction.
[0041] The first server inputs the first instruction into the first NLU model, and the first NLU model performs intent analysis of the execution device on the input first instruction, and outputs at least one device type obtained by the analysis for executing the first instruction. When the first NLU model performs intent analysis of the execution device on the input first instruction, the information device slot of the first instruction can be extracted. The slot represents important information in the instruction information, and the execution device corresponding to the voice instruction can be determined by extracting the device slot, so that at least one device type corresponding to the first instruction can be determined.
[0042] In one embodiment, before extracting the device slot, the first NLU model can obtain information based on multiple rounds of dialogue with the user to fill the slot. In this way, when the first instruction does not contain the field of the corresponding execution device, the device type used to execute the corresponding voice instruction can still be parsed.
[0043] For example, the first instruction is "Open ××× (TV series) on TV", and the first server extracts the device slot through the first NLU model and obtains the corresponding device type as smart TV. In the embodiment of the present application, the execution device intent of the first instruction is parsed through the first NLU model to determine the execution device on which the user wants to perform the operation corresponding to the first instruction.
[0044] Here, the device type can distinguish the terminal devices, and the device type includes but is not limited to: the type of terminal device, such as smart TV, smart audio, smart air conditioner, etc.; the functional type of terminal device, such as refrigeration appliances, kitchen appliances, cleaning appliances, audio-visual appliances, etc.
[0045] Step 303: When the first terminal device does not belong to the first device type, a third terminal device is determined from at least one second terminal device having a set relationship with the first terminal device according to at least one first device type.
[0046] When the first terminal device receiving the first instruction does not belong to the first device type, in other words, when the first terminal device is not a device that executes the first instruction, the first server determines a third terminal device corresponding to the first instruction from at least one second terminal device that satisfies a set relationship with the first terminal device according to at least one first device type. The third terminal device is the device that executes the first instruction.
[0047] In actual application, the setting relationship can be set according to the device information of the first terminal device.
[0048] Step 304: Send a first message to the first terminal device, so that the first terminal device establishes a first connection with the third terminal device based on the first message, and controls the third terminal device to perform the first operation through the first connection.
[0049] The first message sent by the first server carries an instruction to instruct the first terminal device to perform a set operation, so that the first terminal device establishes a first connection with the third terminal device based on the first message, and the first terminal device controls the third terminal device to perform the first operation through the first connection.
[0050] In this embodiment, the first server adopts an NLU model for parsing the intention of the execution device of the instruction, and determines the third terminal device that executes the first operation based on the parsing result of the NLU model and the terminal device information provided by the first terminal device, thereby improving the decision-making speed of the cloud in determining the execution device of the first instruction, allowing the first terminal device to quickly establish a connection with the third terminal device, and then complete the cross-device execution of the instruction, thereby improving the voice control efficiency during cross-device execution.
[0051] In one embodiment, the first NLU model is further used to output a first confidence level corresponding to each device type in the at least one device type;
[0052] The determining, according to at least one first device type, a third terminal device from at least one second terminal device having a set relationship with the first terminal device, includes:
[0053] According to the at least one device type and the first confidence level corresponding to each device type, a third terminal device is determined from at least one second terminal device that has a set relationship with the first terminal device.
[0054] The first server inputs the first instruction into the first NLU model, and the first NLU model performs intent analysis on the input first instruction to determine the execution device, and outputs at least one device type obtained through analysis for executing the first instruction and a first confidence level corresponding to each device type. In the case where the first terminal device receiving the first instruction does not belong to the device type for executing the first instruction, the first server determines the execution device corresponding to the first instruction from the terminal devices that satisfy a set relationship with the first terminal device based on the first confidence level corresponding to each first device type in at least one first device type. Here, the first confidence level represents the confidence level of each second terminal device obtained by the first NLU model from parsing the input first instruction. The higher the confidence level, the more likely it is that the corresponding first device type is the device type of the execution device corresponding to the voice instruction.
[0055] In this way, the first server determines the execution device of the first instruction according to the first confidence corresponding to each first device type in at least one first device type, thereby improving the accuracy of determining the terminal device that executes the first operation, thereby improving the accuracy of establishing a connection between terminal devices.
[0056] As mentioned above, the set relationship can be set according to the device information of the first terminal device, and can be set to the user account information logged in by the first terminal device, and / or the current connection information of the first terminal device. In one embodiment, the set relationship includes at least one of the following:
[0057] The same as the user account of the first terminal device;
[0058] Short-range wireless communication has been established with the first terminal device.
[0059] When the relationship is set to be characterized as being the same as the user account of the first terminal device, the user account for logging in to the first terminal device and at least one second terminal device is the same.
[0060] When the set relationship is characterized by short-range wireless communication having been established with the first terminal device, the first terminal device has a direct communication capability with at least one second terminal device.
[0061] When the set relationship is characterized as being the same as the user account of the first terminal device and short-range wireless communication has been established with the first terminal device, the first terminal device and at least one second terminal device have the same logged-in user account, and the first terminal device and at least one second terminal device have the ability to communicate directly.
[0062] Taking the case where the relationship is characterized as being the same as the user account of the first terminal device as an example, the user accounts of the two terminal devices are the same, which can be the same account logged in to the two terminal devices in the past and / or currently. Here, the user can set the user accounts of the two terminal devices to the same group, for example, the user sets a family group with the same permissions, and the user accounts under the same group are regarded as the same account.
[0063] Here, the first terminal device has logged in to account A / B, the second terminal device A and the second terminal device B have logged in to account B, and the second terminal device C has logged in to account A. If the relationship is set that the two terminal devices are currently logged in to the same account, the third terminal device is determined from the second terminal device A and the second terminal device B; if the relationship is set that the two terminal devices are currently and historically logged in to the same account, the third terminal device is determined from the second terminal device A, the second terminal device B, and the second terminal device C.
[0064] In actual applications, it is assumed that the relationship is set to that the terminal devices are associated with the same user. User A is a guest at user B's home. User A carries smartphone A and smart watch B. User B has smart TV C and smart air conditioner D at home. When user A inputs "open ××× (TV series)" through smartphone A (the first terminal device), since smart TV C, smart air conditioner D are associated with the first terminal device with user B, and smart TV C and smart air conditioner D are associated with different users, they will not be identified as third terminal devices.
[0065] In this way, the first server combines the terminal device information provided by the first terminal device to determine the third terminal device from at least one second terminal device that has a set relationship with the first terminal device, thereby improving the accuracy of determining the terminal device that executes the first operation, thereby improving the accuracy of establishing a connection between terminal devices. At the same time, the decision speed of the cloud to determine the execution device of the first instruction is improved, so that the first terminal device can quickly establish a connection with the third terminal device, and then complete the cross-device execution of the instruction, thereby improving the voice control efficiency during cross-device execution.
[0066] In actual application scenarios, the setting relationship can also be characterized as the absence of a connection between the first terminal device and the second terminal device for controlling the terminal device. In actual applications, in order to determine whether a first connection needs to be established between the wake-up device and the execution device, the setting relationship can be set to the absence of a connection relationship between the first terminal device and the second terminal device. In this way, the first server can exclude the terminal device that has established a connection when determining the connection instruction based on the connection status of the first terminal device.
[0067] In one embodiment, before determining the third terminal device from at least one second terminal device having a set relationship with the first terminal device according to the at least one first device type, the method further includes:
[0068] receiving first information sent by the first terminal device; wherein,
[0069] The first information represents the at least one second terminal device.
[0070] Here, the first information is used to describe a second terminal device that has a set relationship with the first terminal device, and the first server determines at least one second terminal device that has a set relationship with the first terminal device according to the first information.
[0071] Here, the first terminal device may send the first information while sending the first instruction, or the first server may send a request to the first terminal when determining that the first terminal device does not belong to at least one second terminal device, and the first terminal device may send the first information based on the request. In other words, for the first terminal device, the first information may be actively reported (for example, sent together with the first instruction), or reported in response to a request sent by the first server.
[0072] In this way, the first server can accurately determine the execution device of the first instruction in the terminal device based on the first information sent by the first terminal device, thereby improving the accuracy of determining the terminal device that executes the first operation, thereby improving the accuracy of establishing connections between terminal devices.
[0073] In addition, the first information may also be reported by at least one terminal device to the first server at every set period.
[0074] In one embodiment, the determining, according to at least one first device type, a third terminal device from at least one second terminal device having a set relationship with the first terminal device includes:
[0075] According to the at least one first device type and according to the second information corresponding to each second terminal device in the at least one second terminal device, the third terminal device is determined in the at least one second terminal device; wherein,
[0076] The second information represents the device status of the corresponding terminal device.
[0077] In this embodiment, the third terminal device is determined according to the device status of the terminal device corresponding to at least one first device type and at least one second terminal device output by the first NLU model. Here, the second information, i.e., the information of the device status of the corresponding terminal, includes but is not limited to: application information installed on the terminal device, online status information, and supported control modes.
[0078] Furthermore, in this embodiment, the manner in which the first server obtains the second information is not limited.
[0079] In actual applications, taking the second information including online status information as an example, the two device types output by the first NLU model are playback devices and control devices. The playback device is a smart TV A in at least one second terminal device, and A is online. The control device is a smart air conditioner B, and B is offline. According to the device status of the second terminal device, smart TV A is determined as the third terminal device.
[0080] In this way, the first server can accurately determine the execution device of the first instruction in the terminal device based on the second information corresponding to the second terminal device, thereby improving the accuracy of determining the terminal device that executes the first operation, thereby improving the accuracy of establishing connections between terminal devices.
[0081] As mentioned above, the manner in which the first server obtains the second information is not limited. Therefore, in one embodiment, before determining the third terminal device from at least one second terminal device that has a set relationship with the first terminal device, the method further includes:
[0082] Second information of at least one terminal device is received.
[0083] The first server obtains the second information of at least one terminal device in the form of the terminal device sending the second information to the first server at set intervals, or the first server actively sends a request and receives the second information.
[0084] When the first server receives the second information of at least one terminal device, the terminal device that sends the second information to the first server is not limited. For example, in actual applications, each terminal device may report to the first server at a set period; or a management device may be determined from the terminal devices that have the ability to access the cloud, and other terminal devices may report the first information to the management device at a set period, and the management device may send the first information of at least one terminal device to the first server.
[0085] In this way, the first server determines the execution device of the first instruction based on the real-time device status of the terminal device, thereby improving the accuracy of determining the terminal device that executes the first operation, thereby improving the accuracy of determining whether to establish a connection between the wake-up device and the execution device.
[0086] In one embodiment, inputting the first instruction into the first NLU model includes:
[0087] Preprocessing the received first instruction in a setting manner;
[0088] Inputting the preprocessed first instruction into the first NLU model;
[0089] The preprocessing of the setting method includes normalization processing and / or word segmentation processing.
[0090] In this embodiment, before the first instruction is input into the first NLU model, the first server performs preprocessing by standardization, word segmentation, etc., processes the information of the first instruction into information that satisfies the set rules, and inputs the processed information into the first NLU model.
[0091] Here, normalization is to process the instructions into a standard form. For example, "I want to play XXX (TV series) on TV", "Help me play XXX (TV series) on TV", "I want to play, um, XXX (TV series) on TV", are processed into "Play XXX (TV series) on TV". Word segmentation is to segment the Chinese character sequence of the instruction into individual words. For example, "I want to play XXX (TV series) on TV" is processed into "I want + play XXX (TV series) on TV".
[0092] In actual applications, after the above instructions are preprocessed by normalization and word segmentation, the instructions input into the first NLU model are all "play ××× (TV series) on TV". In other words, through preprocessing, instructions representing the same meaning can be processed into standardized instructions. In this way, the first NLU model performs parsing based on standardized instructions, thereby improving the accuracy of NLU model recognition.
[0093] As mentioned above, when the cloud server parses the voice command according to the dialogue system, it parses the skill intent of the voice command through multiple skill NLU models in the dialogue system. In one embodiment, after receiving the first command sent by the first terminal device, the method further includes:
[0094] The first instruction is input into a second NLU model, and the skill intent corresponding to the first instruction output by the second NLU model is obtained; wherein,
[0095] The second NLU model is used to parse the input voice command regarding the skill intent and output the skill intent corresponding to the voice command.
[0096] In the embodiment of the present application, the first server parses the execution device intent of the first instruction through the first NLU model to determine whether it is necessary to establish a first connection between the wake-up device and the execution device. In addition, the first server also parses the skill intent of the first instruction through at least one second NLU model, so that the wake-up device can control the execution device to perform the corresponding first operation through the first connection.
[0097] Here, the first server can also determine at least one device type through the first NLU model, and then use the corresponding second NLU model according to the determined at least one device type, thereby reducing the number of skill NLU models required when parsing voice commands, thereby reducing the computing work consumed in parsing skill intents.
[0098] In one embodiment, sending the first message to the first terminal device includes:
[0099] Sending a first message carrying a second instruction to the first terminal device; the second instruction is used to instruct the first terminal device to establish the first connection with the third terminal device; or,
[0100] A first message carrying a third instruction is sent to the first terminal device; the third instruction is used to instruct the first terminal device to output fourth information; and the fourth information is used to prompt whether to establish a first connection with the third terminal device.
[0101] In some scenarios, the first message sent by the first server to the first terminal device carries the second instruction, and the first terminal device directly establishes a first connection with the third terminal device based on the second instruction.
[0102] In other scenarios, the first message sent by the first server to the first terminal device carries a third instruction. The first terminal device outputs fourth information based on the third instruction. The fourth information is prompt information, which is used to prompt whether to establish a first connection with the third terminal device. Upon receiving a response confirming the establishment of the first connection, the first terminal device establishes a first connection with the third terminal device.
[0103] The following is a further description of the embodiments of the present application in conjunction with the application examples:
[0104] Below, the cross-device execution scenario is taken as the cloud server as the first server, mobile phone A as the first terminal device, smart TV B and smart tablet C as the second terminal device, and the application is further described in detail in combination with the application embodiment.
[0105] Figure 4 The following is a schematic diagram of the implementation flow of the device control method provided in the application embodiment of the present application. Figure 4 As shown, the device control method includes:
[0106] Step 401: The mobile phone collects the current login account information.
[0107] Mobile phone A, namely the control side device and the wake-up device, collects information about the terminal devices owned by the currently logged-in user account to obtain the device status of the terminal devices.
[0108] Taking an actual application scenario as an example, the current login account of mobile phone A has logged in to smart TV B and smart tablet C, so mobile phone A collects device information of smart TV B and smart tablet C.
[0109] Step 402: The mobile phone sends a voice command, current login account information and device information of the terminal device to the cloud server.
[0110] Mobile phone A sends a voice command to the cloud server, and sends the historical login device information of the current login account of mobile phone A and the device information currently connected to mobile phone A (i.e., the device information collected from smart TV B and smart tablet C).
[0111] Among them, the device information includes application information installed on the terminal device, online status information, and supported control modes.
[0112] Here, the voice command is "I want to watch ××× (TV series) on TV", and the device currently connected to mobile phone A is smart tablet C.
[0113] Step 403: The cloud server pre-processes the voice command.
[0114] The cloud server sets up a cloud command normalization module to perform normalization, standardization, word segmentation and other pre-processing on the voice commands to obtain standardized commands. The model performs recognition based on the standardized commands, which can improve the accuracy of the recognition results.
[0115] Here, the voice command "I want to watch ××× (TV series) on TV" is normalized to "Play ××× (TV series) on TV".
[0116] Step 404: The cloud server inputs the voice command into the cross-device decision NLU model.
[0117] The cloud server models the voice commands by setting up a cross-device decision NLU model, that is, the first NLU model, so as to classify the intent of the voice commands and extract the device slots.
[0118] Here, the recognition accuracy of the NLU model can be improved by setting a rule for at least one voice command, such as "play ×× on TV" which may be generated by the user. The NLU model here is usually a neural network model.
[0119] In practical applications, the corresponding corpus samples can also be determined based on the rules, and the NLU model can be trained based on the corpus samples, thereby improving the accuracy of the NLU model in recognizing the corresponding voice commands.
[0120] Step 405: The cloud server obtains at least one first device type output by the cross-device decision NLU model and the confidence level corresponding to each device type.
[0121] The cross-device decision NLU model has a classification method and confidence score for each classification, and will correspond to a confidence level for each device type output. The higher the confidence level output by the cross-device decision NLU model, the greater the confidence of the NLU model in identifying the corresponding device type.
[0122] Step 406: The cloud server identifies whether the voice command is a connection instruction, and sends the connection instruction to the mobile phone when it is identified as a connection instruction.
[0123] The cloud server sets up a cloud-based connection decision service module, obtains the execution device of the voice command based on the NLU model output results, the device information of the terminal device and other auxiliary information (such as user preferences), and identifies whether the voice command is a connection command.
[0124] The cloud-based connection decision service module makes a decision based on at least one first device type, the confidence level corresponding to each device type, and the terminal device information owned by the current user account, obtains the execution device of the voice command, and identifies whether the voice command is a connection command.
[0125] Here, the cloud connection decision service module decides that the execution device of the voice command is smart TV B, and according to the connection information, short-range wireless communication is not established between smart TV B and mobile phone A, and the voice command is recognized as a connection command.
[0126] Step 407: The mobile phone establishes a first connection with the smart TV.
[0127] In the related art, in order to parse the skill intent of voice commands and identify whether the voice commands are association commands, the cloud server needs to parse the voice commands through the skill NLU model. Since the device control system usually deploys execution devices of different device types, a large number of skill NLU models are set in the dialogue system of the cloud server. Therefore, the cloud server will consume a lot of time in judging cross-device execution, and the delay of voice command cross-device execution is high. In addition, in order to judge whether the voice command is an association command, for each additional type of skill intent that the dialogue system supports, the cloud server needs to train the corpus of multiple skill NLU models, which requires a huge amount of computing work.
[0128] In the application embodiment of the present application, the cloud server normalizes the voice instructions, and the cross-device decision NLU model performs intent analysis of the execution device of the multi-device connection instruction, and makes a decision through the cloud connection decision service module in combination with the model output result and the device information issued by the control side device to determine whether the voice instruction is a connection instruction. In this way, the first server uses the NLU model used to parse the intention of the execution device of the instruction, and on the basis of the analysis result of the NLU model and the terminal device information provided by the first terminal device, determines the third terminal device that performs the first operation, thereby improving the decision speed of the cloud to determine the execution device of the first instruction, so that the first terminal device can quickly establish a connection with the third terminal device, and then complete the cross-device execution of the instruction, thereby improving the voice control efficiency during cross-device execution. At the same time, combined with the connection information between devices on the control side, the accuracy of model recognition is improved, and the corpus training of each skill NLU model on the connection instruction is saved.
[0129] In order to implement the device control method of the embodiment of the present application, the embodiment of the present application also provides a device control device, such as Figure 5 As shown, the device control device includes:
[0130] The first receiving unit 501 is used to receive a first instruction sent by a first terminal device; the first instruction is used to instruct the corresponding terminal device to perform a first operation;
[0131] A first processing unit 502 is configured to input a first instruction into a first natural language understanding NLU model to obtain at least one first device type output by the first NLU model; the first NLU model is configured to perform intent analysis on the input voice instruction regarding the execution device, and output the device type obtained by the analysis for executing the corresponding voice instruction;
[0132] A second processing unit 503 is configured to determine, if the first terminal device does not belong to the first device type, a third terminal device from at least one second terminal device having a set relationship with the first terminal device according to at least one first device type;
[0133] The first sending unit 504 is configured to send a first message to the first terminal device, so that the first terminal device establishes a first connection with the third terminal device based on the first message, and controls the third terminal device to perform the first operation through the first connection.
[0134] In one embodiment, the first NLU model is further used to output a first confidence level corresponding to each device type in the at least one device type; and the second processing unit 503 is used to:
[0135] According to the at least one device type and the first confidence level corresponding to each device type, a third terminal device is determined from at least one second terminal device that has a set relationship with the first terminal device.
[0136] In one embodiment, the device control apparatus further comprises:
[0137] The second receiving unit is used to receive the first information sent by the first terminal device; wherein:
[0138] The first information represents the at least one second terminal device.
[0139] In one embodiment, the second processing unit 503 is configured to:
[0140] According to the at least one first device type and according to the second information corresponding to each second terminal device in the at least one second terminal device, the third terminal device is determined in the at least one second terminal device; wherein,
[0141] The second information represents the device status of the corresponding terminal device.
[0142] In one embodiment, the first processing unit 502 is configured to:
[0143] Preprocessing the received first instruction in a setting manner;
[0144] Inputting the preprocessed first instruction into the first NLU model;
[0145] The preprocessing of the setting method includes normalization processing and / or word segmentation processing.
[0146] In one embodiment, the device control apparatus further comprises:
[0147] The third processing unit is used to input the first instruction into a second NLU model to obtain the skill intent corresponding to the first instruction output by the second NLU model; wherein:
[0148] The second NLU model is used to parse the input voice command regarding the skill intent and output the skill intent corresponding to the voice command.
[0149] In one embodiment, the first sending unit 504 is configured to:
[0150] Sending a first message carrying a second instruction to the first terminal device; the second instruction is used to instruct the first terminal device to establish the first connection with the third terminal device; or,
[0151] A first message carrying a third instruction is sent to the first terminal device; the third instruction is used to instruct the first terminal device to output fourth information; and the fourth information is used to prompt whether to establish a first connection with the third terminal device.
[0152] In actual application, the first receiving unit 501 and the second receiving unit can be implemented by a communication interface based on the device control device, the first processing unit 502 and the second processing unit 503 can be implemented by a processor based on the device control device, and the first sending unit 504 can be implemented by a processor based on the device control device combined with the communication interface.
[0153] It should be noted that: the device control device provided in the above embodiment only uses the division of the above program modules as an example when performing device control. In actual applications, the above processing can be assigned to different program modules as needed, that is, the internal structure of the device is divided into different program modules to complete all or part of the processing described above. In addition, the device control device and the device control method embodiment provided in the above embodiment belong to the same concept, and the specific implementation process is detailed in the method embodiment, which will not be repeated here.
[0154] Based on the hardware implementation of the above program modules, and in order to implement the method of the embodiment of the present application, the embodiment of the present application also provides a server. Figure 6 Schematic diagram of the hardware structure of the server in the embodiment of the present application. Figure 6 As shown, the server includes:
[0155] Communication interface 1, capable of exchanging information with other devices such as network devices;
[0156] The processor 2 is connected to the communication interface 1 to implement information exchange with other devices and is used to execute the device control method provided by one or more technical solutions when running a computer program. The computer program is stored in the memory 3.
[0157] Of course, in actual application, the various components in the server are coupled together through the bus system 4. It can be understood that the bus system 4 is used to realize the connection and communication between these components. In addition to the data bus, the bus system 4 also includes a power bus, a control bus and a status signal bus. However, for the sake of clarity, Figure 6 Various buses are labeled as bus system 4.
[0158] The memory 3 in the embodiment of the present invention is used to store various types of data to support the operation of the server. Examples of such data include: any computer program used to operate on the server.
[0159] It can be understood that the memory 3 can be a volatile memory or a non-volatile memory, and can also include both volatile and non-volatile memories. Among them, the non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), a magnetic random access memory (FRAM), a flash memory, a magnetic surface memory, an optical disc, or a compact disc read-only memory (CD-ROM); the magnetic surface memory can be a disk memory or a tape memory. The volatile memory can be a random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of RAM are available, such as static random access memory (SRAM), synchronous static random access memory (SSRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM, SyncLink Dynamic Random Access Memory), and direct RAM bus random access memory (DRRAM, Direct Rambus Random Access Memory).The memory 2 described in the embodiments of the present invention is intended to include but is not limited to these and any other suitable types of memories.
[0160] The method disclosed in the above embodiment of the present invention can be applied to the processor 2, or implemented by the processor 2. The processor 2 may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by the hardware integrated logic circuit in the processor 2 or the instruction in the form of software. The above processor 2 can be a general-purpose processor, a DSP, or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components, etc. The processor 2 can implement or execute the methods, steps and logic block diagrams disclosed in the embodiment of the present invention. The general-purpose processor can be a microprocessor or any conventional processor, etc. In combination with the steps of the method disclosed in the embodiment of the present invention, it can be directly embodied as a hardware decoding processor to execute, or it can be executed by a combination of hardware and software modules in the decoding processor. The software module can be located in a storage medium, which is located in the memory 3. The processor 2 reads the program in the memory 3 and completes the steps of the above method in combination with its hardware.
[0161] When the processor 2 executes the program, the corresponding processes in the various methods of the embodiments of the present invention are implemented, which will not be described here for the sake of brevity.
[0162] In an exemplary embodiment, the embodiment of the present invention further provides a storage medium, namely a computer storage medium, specifically a computer readable storage medium, for example, including a memory 3 storing a computer program, and the above-mentioned computer program can be executed by a processor 2 to complete the steps of the above-mentioned method. The computer readable storage medium can be a memory such as FRAM, ROM, PROM, EPROM, EEPROM, Flash Memory, magnetic surface storage, optical disk, or CD-ROM.
[0163] In the several embodiments provided in the present application, it should be understood that the disclosed devices, terminals and methods can be implemented in other ways. The device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation, such as: multiple units or components can be combined, or can be integrated into another system, or some features can be ignored or not executed. In addition, the coupling, direct coupling, or communication connection between the components shown or discussed can be through some interfaces, and the indirect coupling or communication connection of the devices or units can be electrical, mechanical or other forms.
[0164] The units described above as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units; some or all of the units may be selected according to actual needs to achieve the purpose of the present embodiment.
[0165] In addition, all functional units in the embodiments of the present invention may be integrated into one processing unit, or each unit may be separately used as a unit, or two or more units may be integrated into one unit; the above-mentioned integrated units may be implemented in the form of hardware or in the form of hardware plus software functional units.
[0166] A person of ordinary skill in the art can understand that: all or part of the steps of implementing the above-mentioned method embodiment can be completed by hardware related to program instructions, and the aforementioned program can be stored in a computer-readable storage medium, which, when executed, executes the steps of the above-mentioned method embodiment; and the aforementioned storage medium includes: various media that can store program codes, such as mobile storage devices, ROM, RAM, disks or optical disks.
[0167] Alternatively, if the above-mentioned integrated unit of the present invention is implemented in the form of a software function module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the embodiment of the present invention can be essentially or partly reflected in the form of a software product that contributes to the prior art. The computer software product is stored in a storage medium and includes several instructions for an electronic device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the methods described in each embodiment of the present invention. The aforementioned storage medium includes: various media that can store program codes, such as mobile storage devices, ROM, RAM, magnetic disks or optical disks.
[0168] The term "and / or" herein is only a description of the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. In addition, the term "at least one" herein represents any combination of at least two of any one or more of a plurality. For example, at least one of A, B, and C can represent any one or more elements selected from the set consisting of A, B, and C.
[0169] The above is only a specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed by the present invention, which should be included in the protection scope of the present invention. Therefore, the protection scope of the present invention should be based on the protection scope of the claims.
Claims
1. A device control method, characterized in that: Applied to a first server, the method comprises: Receiving a first instruction sent by a first terminal device; the first instruction is used to instruct the corresponding terminal device to perform a first operation; Inputting the first instruction into a first natural language understanding NLU model, and obtaining at least one first device type output by the first NLU model; the first NLU model is used to perform intent analysis on the input voice instruction regarding the execution device, and output the device type obtained by the analysis for executing the corresponding voice instruction; In a case where the first terminal device does not belong to the first device type, determining a third terminal device from at least one second terminal device having a set relationship with the first terminal device according to at least one first device type; Sending a first message to the first terminal device, so that the first terminal device establishes a first connection with the third terminal device based on the first message, and controls the third terminal device to perform the first operation through the first connection; The first NLU model is further used to output a first confidence level corresponding to each device type in the at least one device type; The determining, according to at least one first device type, a third terminal device from at least one second terminal device having a set relationship with the first terminal device comprises: According to the at least one device type and the first confidence level corresponding to each device type, a third terminal device is determined from at least one second terminal device that has a set relationship with the first terminal device.
2. The method according to claim 1, characterized in that Before determining the third terminal device from at least one second terminal device having a set relationship with the first terminal device according to the at least one first device type, the method further includes: receiving first information sent by the first terminal device; wherein, The first information represents the at least one second terminal device.
3. The method according to claim 1, characterized in that The determining, according to at least one first device type, a third terminal device from at least one second terminal device having a set relationship with the first terminal device comprises: According to the at least one first device type and according to the second information corresponding to each second terminal device in the at least one second terminal device, the third terminal device is determined in the at least one second terminal device; wherein, The second information represents the device status of the corresponding terminal device.
4. The method according to claim 1, characterized in that The step of inputting the first instruction into the first NLU model includes: Preprocessing the received first instruction in a set manner; Inputting the preprocessed first instruction into the first NLU model; The preprocessing of the setting method includes normalization processing and / or word segmentation processing.
5. The method according to claim 1, characterized in that After receiving the first instruction sent by the first terminal device, the method further includes: The first instruction is input into a second NLU model, and the skill intent corresponding to the first instruction output by the second NLU model is obtained; wherein, The second NLU model is used to parse the input voice command regarding the skill intent and output the skill intent corresponding to the voice command.
6. The method according to claim 1, characterized in that The sending a first message to the first terminal device includes: Sending a first message carrying a second instruction to the first terminal device; the second instruction is used to instruct the first terminal device to establish the first connection with the third terminal device; or, A first message carrying a third instruction is sent to the first terminal device; the third instruction is used to instruct the first terminal device to output fourth information; and the fourth information is used to prompt whether to establish a first connection with the third terminal device.
7. A device control device, characterized in that: include: A first receiving unit, configured to receive a first instruction sent by a first terminal device; The first instruction is used to instruct the corresponding terminal device to perform a first operation; a first input unit, configured to input a first instruction into a first NLU model, and obtain at least one first device type output by the first NLU model; the first NLU model is configured to perform intent analysis on the input voice instruction regarding the execution device, and output the device type obtained by the analysis for executing the corresponding voice instruction; a first processing unit, configured to determine, if the first terminal device does not belong to the first device type, a third terminal device from at least one second terminal device having a set relationship with the first terminal device according to at least one first device type; A first sending unit, configured to send a first message to the first terminal device, so that the first terminal device establishes a first connection with the third terminal device based on the first message, and controls the third terminal device to perform the first operation through the first connection; The first NLU model is further used to output a first confidence level corresponding to each device type in the at least one device type; The second processing unit is used to determine a third terminal device from at least one second terminal device that has a set relationship with the first terminal device according to the at least one device type and the first confidence level corresponding to each device type.
8. A server, characterized in that: include: a first processor and a first memory for storing a computer program executable on the processor, Wherein, when the first processor is used to run the computer program, it executes the steps of the device control method described in any one of claims 1 to 6.
9. A storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the device control method according to any one of claims 1 to 6 are implemented.
Citation Information
Patent Citations
Voice control method and device, server, terminal equipment and storage medium
CN113127609A