Device Control Method, Device, Electronic Device and Computer Readable Storage Medium

By acquiring and analyzing user information, environmental information and equipment information, combined with equipment control instructions, safer and more flexible intelligent device control is achieved, and the problem of insufficient security and flexibility of equipment control in the prior art is solved.

CN111352348BActive Publication Date: 2025-06-10BEIJING SAMSUNG TELECOM R&D CENT +1
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN201910267629.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2018-12-24
Filing Date
2019-04-03
Publication Date
2025-06-10
Estimated Expiration
2039-04-03

AI Technical Summary

Technical Problem

The prior art fails to fully consider user information, device information and environmental information in intelligent device control, resulting in low security and flexibility of the device control process.

Method used

By obtaining the device control instructions input by the user, as well as user information, environmental information and equipment information, comprehensive analysis and processing are carried out to determine the operations that the device should perform to achieve safer and more flexible device control.

Benefits of technology

It improves the security and flexibility of smart device control, avoids potential dangers and inconvenience caused by misoperation of equipment or incorrect parameters, and improves the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN111352348B_ABST
    Figure CN111352348B_ABST
Patent Text Reader

Abstract

An embodiment of the present application provides a device control method, apparatus, electronic device, and computer-readable storage medium. The method includes: obtaining a device control instruction input by a user, obtaining at least one of the following information: user information, environmental information, and device information, and controlling at least one target device to perform corresponding operations based on the obtained information and the device control instruction. The embodiment of the present application realizes safe and convenient control of intelligent devices to perform corresponding operations.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence technology. Specifically, the present application relates to a device control method, device, electronic device, and computer-readable storage medium. Background Art

[0002] With the development of information technology, various devices have entered people's daily lives. For example, air conditioners, washing machines, refrigerators, etc. Users can control these devices to perform corresponding operations by manually adjusting the buttons on the devices or remote controlling the buttons on the devices.

[0003] With the further development of artificial intelligence technology, intelligent devices have gradually entered people's daily lives. For example, smart home devices such as smart speakers, smart air conditioners, smart TVs, and smart ovens. Users can control smart home devices to perform corresponding operations without manually adjusting the buttons on the home devices or remote controlling. For example, users can control these smart home devices to perform corresponding operations through application programs installed on terminal devices such as mobile phones. For example, users can control the air conditioner to turn on through smart terminal devices such as mobile phones.

[0004] However, in the current device control methods, generally only the control instructions input by the user are used for control, and other factors that may affect the operation of the device are not considered, which may lead to lower security and flexibility in the device control process. Therefore, how to control smart devices to perform corresponding operations more safely and flexibly has become a key issue. Summary of the Invention

[0005] The present application provides a device control method, device, electronic device, and computer-readable storage medium, which can solve the problem of how to control smart devices to perform corresponding operations more safely and flexibly. The technical solutions are as follows:

[0006] In a first aspect, a device control method is provided. The method includes:

[0007] Obtain a device control instruction input by a user;

[0008] Obtain at least one of the following information: user information; environmental information; device information;

[0009] Based on the obtained information and the device control instruction, control at least one target device to perform corresponding operations.

[0010] In a second aspect, a device control device is provided. The device includes:

[0011] A first obtaining module, configured to obtain a device control instruction input by a user and at least one of the following information: user information; environmental information; device information;

[0012] A control module, configured to control at least one target device to perform corresponding operations based on the information obtained by the first acquisition module and the device control instruction.

[0013] In a third aspect, an electronic device is provided, which includes:

[0014] One or more processors;

[0015] A memory;

[0016] One or more applications, where one or more applications are stored in the memory and configured to be executed by one or more processors, and one or more programs are configured to: execute the device control method shown in the first aspect.

[0017] In a fourth aspect, a computer-readable storage medium is provided, and the storage medium stores at least one instruction, at least one segment of program, a code set or an instruction set, and at least one instruction, at least one segment of program, the code set or the instruction set is loaded and executed by a processor to implement the device control method shown in the first aspect.

[0018] The beneficial effects brought by the technical solution provided in this application are:

[0019] This application provides a device control method, device, electronic device and computer-readable storage medium. By acquiring at least one of user information, environmental information and device information and the device control instruction input by the user, it is possible to control at least one target device to perform corresponding operations based on the acquired information and the device control instruction. As can be seen, compared with the prior art method of only controlling the device according to the control instruction input by the user, when this application controls the device, in addition to considering the control instruction input by the user, it also takes into account at least one of user information, device information, environmental information and other factors that may affect the device operation, so that the device operation can be controlled more safely and flexibly. For example, acquire device control instructions in the form of voice, text, key presses, gestures, etc. input by the user, and considering at least one of user information, device information and environmental information, directly control the air conditioner device to turn on, turn off or adjust the temperature, etc., so as to achieve safe and convenient control of the intelligent device to perform corresponding operations. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments of the present application.

[0021] Figure 1 It is a schematic flowchart of a device control method provided by an embodiment of the present application;

[0022] Figure 2a It is a schematic diagram of the composition of multi-modal information provided by an embodiment of the present application;

[0023] Figure 2b Schematic diagram of obtaining domain classification results through the domain classifier in the embodiments of the present application;

[0024] Figure 3a Schematic diagram of obtaining an information representation vector (multimodal information representation vector) corresponding to the obtained information (multimodal information) through the obtained information in Embodiment 1 of the present application;

[0025] Figure 3b Schematic diagram of obtaining intent classification results through the intent classifier in the embodiments of the present application;

[0026] Figure 3c Schematic diagram of the corresponding operations performed by the domain classifier, intent classifier, and sequence tagger in the embodiments of the present application;

[0027] Figure 4 Schematic diagram of determining that the intent in the device control instruction in the embodiments of the present application is an allowable intent;

[0028] Figure 5 Schematic diagram of sequence tagging processing through the sequence tagger in Embodiment 1 of the present application;

[0029] Figure 6a Schematic diagram of the device control system corresponding to Embodiment 1 of the present application;

[0030] Figure 6b Schematic diagram of the multimodal information processing module in Embodiment 1 of the present application;

[0031] Figure 7 Schematic diagram of the device control system in the prior art;

[0032] Figure 8 Schematic diagram of obtaining domain classification results through the domain classifier in Embodiment 2 of the present application;

[0033] Figure 9 Schematic diagram of obtaining intent classification results through the intent classifier in Embodiment 2 of the present application;

[0034] Figure 10 Schematic diagram of obtaining environmental information and device information in the multimodal information in the embodiments of the present application;

[0035] Figure 11 Schematic diagram of sequence tagging processing by the sequence tagger in Embodiments 2 and 3 of the present application;

[0036] Figure 12 Schematic diagram of the architecture of the device control system in Embodiment 2 of the present application;

[0037] Figure 13Schematic diagram for obtaining multi-modal information in Embodiment 3 of this application;

[0038] Figure 14 Schematic diagram for determining multi-modal information based on user portrait database, environment database, and user permission database in Embodiment 3 of this application;

[0039] Figure 15 Schematic diagram for obtaining an information representation vector (multi-modal information representation vector) corresponding to the obtained information (multi-modal information) through the obtained information in Embodiment 3 of this application;

[0040] Figure 16 Schematic diagram of the architecture of the device control system in Embodiment 3 of this application;

[0041] Figure 17 Schematic diagram of the architecture of the device control device in this embodiment of the application;

[0042] Figure 18 Schematic diagram of the architecture of the electronic device in this embodiment of the application;

[0043] Figure 19 Schematic diagram of the architecture of the computer system in this embodiment of the application;

[0044] Figure 20a Schematic diagram for obtaining multi-modal information in this embodiment of the application;

[0045] Figure 20b Schematic diagram of the training process of the neural network of Emotional TTS in this embodiment of the application;

[0046] Figure 20c Schematic diagram of the online processing process of the neural network of Emotional TTS in this embodiment of the application. Detailed implementation manners

[0047] The embodiments of the present application will be described in detail below. The examples of the embodiments are shown in the drawings, where the same or similar reference numerals represent the same or similar elements or elements with the same or similar functions from beginning to end. The embodiments described below with reference to the drawings are exemplary and are only used to explain the present application, and cannot be construed as a limitation of the present invention.

[0048] Those skilled in the art can understand that, unless specifically stated otherwise, the singular forms "a", "an", "the" and "said" used herein may also include the plural forms. It should be further understood that the term "comprising" used in the specification of this application means the presence of the stated features, integers, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or their groups. It should be understood that when we say that an element is "connected" or "coupled" to another element, it can be directly connected or coupled to other elements, or there may also be intermediate elements. In addition, the "connection" or "coupling" used herein may include wireless connection or wireless coupling. The phrase "and / or" used herein includes all or any unit and all combinations of one or more associated listed items.

[0049] With the development of technology, smart home has gradually entered people's lives. Smart speakers, smart TVs and other devices that use voice control provide the hardware foundation for smart home. In future life, after devices such as TVs, refrigerators, ovens, washing machines, etc. are intelligentized, smart home will realize full voice control of various household appliances in the home to perform corresponding operations.

[0050] However, when users control smart devices (including: smart home devices, mobile phones, tablet computers PADs and other terminal devices) to perform corresponding operations through instructions (such as instructions in the forms of voice, text, buttons, gestures, etc.), various problems will occur:

[0051] Problem 1:

[0052] There are often children and the elderly in the family. For example, children need to be prohibited from using devices such as ovens; some functions of computers and TVs also need to be prohibited for children; in addition, existing smart home devices are not user-friendly for the elderly, and some operations are relatively complex. Mistakes made by the elderly may even lead to danger. Moreover, there is no appropriate feedback on the operating parameters of smart home appliances in the voice control system. For example, if the oven temperature is too high and other situations are not feedback, and the cooking task is still executed, it will cause danger.

[0053] The existing smart device control system does not distinguish and control users, and performs the same operations on all user requests. For example, it cannot give corresponding protection or suggestions to user groups such as the elderly and children mentioned above, which may bring danger; for another example, it will not give corresponding protection measures according to the device or environmental conditions, which may bring danger.

[0054] For example, the intelligent devices in the prior art execute corresponding operations for device control instructions (such as voice instructions) for children without considering whether children are suitable for operating a certain device or performing a certain function. Even for a device like an oven, which is relatively dangerous for children, children are still allowed to operate it, posing a significant potential danger to this group. Another example is that for instructions for the elderly, no protective measures are provided, making it very cumbersome for the elderly to operate the device and potentially dangerous due to operational errors. That is to say, the intelligent device control system in the prior art does not consider relevant user information when the user controls the device, resulting in a relatively low level of security in the device control process.

[0055] Another example is that when a user operates an intelligent device such as an oven, if the current temperature of the oven is high and it has been working for a long time, it is not suitable for continued high-temperature baking for a long time, which is likely to damage the device or pose a danger to the user. However, the prior art does not provide corresponding protective measures based on the relevant information of the device, resulting in a relatively low level of security in the device control process.

[0056] Another example is that when a user operates an intelligent device such as an air conditioner, if the current environmental temperature is very low, it is not suitable for further cooling, which is likely to pose a health risk to the user. However, the prior art does not provide corresponding protective measures based on the relevant information of the environment, resulting in a relatively low level of security in the device control process.

[0057] Problem 2:

[0058] When a user controls an intelligent device through instructions (such as instructions in the form of voice, text, buttons, gestures, etc.), the operation parameters corresponding to the control instructions may have the following problems:

[0059] 1) It may pose a danger to the device or the user. For example, the control instruction input by the user is: set the oven temperature to 240 degrees. However, when the oven has been working for a long time, the excessive temperature may pose a danger to the user or the device. The prior art does not consider the above-mentioned dangerous situation for the above instruction and still executes the corresponding operation according to the user's instruction.

[0060] 2) The parameters corresponding to the control instruction have exceeded the executable range of the device. For example, the control instruction input by the user is: set the air conditioner temperature to 100 degrees. However, the upper limit of the air conditioner temperature is 32 degrees, and the corresponding 100 degrees in the user's instruction has exceeded the executable range of the air conditioner. Another example is that the control instruction input by the user for the mobile phone is: set the alarm for 8 am on February 30th. However, February 30th does not exist, which exceeds the alarm setting range. In the prior art, no corresponding operation is performed after receiving the above instructions, resulting in a poor user experience.

[0061] 3) The operation corresponding to the instruction may not be applicable to the user who issues the instruction. For example, the control instruction input by a child is: turn on channel 10 of the TV, but channel 10 of the TV is a channel not suitable for children. In the prior art, when the above control instruction input by the child is received, the corresponding instruction is still executed based on the control instruction input by the child, without considering the above inappropriate situation, and the corresponding operation is still performed according to the user instruction.

[0062] 4) When the control instruction input by the user is unclear. For example, the instruction input by the user is: adjust the air conditioner temperature to a suitable temperature. This instruction is unclear for the air conditioner, and the air conditioner cannot determine the final adjusted temperature. Therefore, the prior art does not perform any operation after receiving the above instruction, resulting in a poor user experience.

[0063] As can be seen from the above, in the prior art, after the device receives the control instruction sent by the user through voice or other means, it does not adjust the instruction. Even if there are dangerous or inapplicable situations, the corresponding operation is still performed. Or when the instruction is unclear or beyond the execution range, the device may not perform any operation, resulting in low safety when the user controls these devices through control instructions, or there are inconvenient problems, thereby resulting in a poor user experience.

[0064] Of course, for problem one, there are also some solutions in the prior art to prevent children from operating certain devices. For example, the device is locked based on buttons, passwords, or fingerprints to prevent children from using these devices. Button-based unlocking often requires unlocking based on specific buttons or button combinations, such as induction cookers or washing machines. Password-based unlocking is based on specific passwords, such as TVs and computers. Fingerprint-based unlocking is based on fingerprints, such as mobile phones. However, in the smart home scenario, when controlling various household appliances through voice, all the above protection technologies require the unlocker to walk near the appliance to be unlocked to perform the unlocking, which will increase a lot of inconvenience and may cause inconvenience to the device user. At the same time, the button-based unlocking method is relatively simple. If a child learns the fixed buttons or button combinations, they can use the appliance, which is likely to cause danger. Therefore, the traditional protection technologies are not suitable for the full-voice control scenario in the smart home.

[0065] Therefore, in view of the above problems, the embodiments of the present application propose a method for controlling intelligent devices, which can be applied to the control system of smart homes. First, obtain the instructions input by the user (including voice instructions), and at the same time, the image information of the user can also be obtained. According to the image information and / or voice information of the user, the user information of the user is obtained through database comparison, including user portraits (such as age, gender, etc.), user permissions (such as the device control permissions of the user), etc. In addition, device information (such as the working status information of the device) or environmental information (the working environment information of the device) can also be obtained. Then, corresponding semantic understanding operations are performed according to the user information, and / or device information, and / or environmental information, and corresponding operations are executed according to the semantic understanding results, so as to be able to execute corresponding operations according to the relevant information of the user, device, and environment. Specifically:

[0066] Perform different operations on different users to achieve special protection for groups such as children and the elderly. When the user does not have the control permission for the device, the operation result of refusing to execute the device control instruction can be output, or when the user does not have the control permission for the target function corresponding to the device control instruction, the operation result of refusing to execute the device control instruction can be output. For example, when the user is a child, the device can be controlled accordingly according to the permissions corresponding to the child group, such as not allowing the child to operate a certain device (such as an oven), or restricting a certain function of the device operated by the child, such as not allowing the saved TV channels to be deleted.

[0067] Perform corresponding operations according to the working status of the device to achieve the safe operation of the device and protect the safety of the device and the user. When the device does not meet the execution conditions corresponding to the device control instruction, the operation result of refusing to execute the device control instruction can be output. For example, when the device is an oven, if the oven has been working for a long time, the excessive temperature may pose a danger to the user or the device. At this time, the operation result of refusing to execute the user's operation instruction to increase the temperature of the oven can be output to protect the safety of the device and the user.

[0068] Perform corresponding operations according to the working environment of the device to achieve the safe operation of the device and protect the safety of the device and the user. When the working environment of the device does not meet the execution conditions corresponding to the device control instruction, the operation result of refusing to execute the device control instruction can be output. For example, when the device is an air conditioner, if the current environmental temperature is very low and it is not suitable to continue cooling, which is likely to bring health hazards to the user. At this time, the operation result of refusing to execute the user's operation instruction to cool the air conditioner can be output to protect the safety of the user.

[0069] Furthermore, an embodiment of the present application also proposes a parameter rewriting method to solve Technical Problem 2. When the user's control instruction (including voice instruction) is dangerous or inapplicable, or is unclear or beyond the execution scope, the operation parameters corresponding to the control instruction can be automatically modified to improve the convenience and safety of the user using the device.

[0070] Specifically, to solve the above problems, an embodiment of the present application provides a device control method, as Figure 1 shown, the method includes:

[0071] Step S101, obtain the device control instruction input by the user.

[0072] For the embodiment of the present application, the user can input the device control instruction in text mode, can also input the device control instruction in voice mode, and can also input the device control instruction in other ways such as keys or gestures. There is no limitation in the embodiment of the present application.

[0073] Step S102 (not shown in the figure), based on the obtained device control instruction, control at least one target device to perform corresponding operations.

[0074] Specifically, controlling at least one target device to perform corresponding operations in step S102 includes: step S1021 and step S1022, where

[0075] Step S1021, obtain at least one of the following information: user information; environmental information; device information.

[0076] Step S1022, based on the obtained information and the device control instruction, control at least one target device to perform corresponding operations.

[0077] In a possible implementation manner of the embodiment of the present application, the user information may include the user information of the user who inputs the above device control instruction; the device information may include the device information of the target device corresponding to the user's device control instruction; the environmental information may include the environmental information corresponding to the target device.

[0078] In a possible implementation manner of the embodiment of the present application, the user information includes: user portrait information and / or the user's device control permission information; and / or the device information includes: the working state information of the device; and / or the environmental information includes: the working environment information of the device.

[0079] For the embodiments of the present application, the user portrait information includes user information such as user identity information, age, gender, user preferences, etc., and may also include the historical information of the user controlling the device (for example, when the user previously controlled the air conditioner, the air conditioner temperature was generally set to 28 degrees, or when the user controlled the TV, the TV channels were generally set to Channel 1 and Channel 5, etc.); the device control permission information of the user includes: the control permission of the user for the device and / or the control permission of the user for the target function. Among them, in the embodiments of the present application, the target device refers to the device that the user wants to control, and the target function refers to the function that the user wants to control for the target device. For example, if the user wants to increase the temperature of the air conditioner, then the air conditioner is the target device, and increasing the temperature of the air conditioner is the target function.

[0080] For the embodiments of the present application, the working state information of the device includes at least one of the following: the current working state information of the device (such as temperature, humidity, channel, power, storage condition, the duration of continuous operation, etc.), the target working state (such as the best working state of the device, etc.), the functions executable by the device, the executable parameters (such as the adjustable temperature range of the air conditioner is 16 - 32 degrees), etc.

[0081] For the embodiments of the present application, the working environment information of the device includes: the current working environment information of the device and / or the set target working environment information (such as the best working environment information, etc.); among them, the working environment information includes temperature, humidity, pressure, etc.

[0082] Another possible implementation manner of the embodiments of the present application, step S1022 may specifically include step S10221 (not shown in the figure), where

[0083] Step S10221: Based on the acquired information and the device control instruction, output an operation result of refusing to execute the device control instruction.

[0084] Specifically, the operation result of outputting a refusal to execute the device control instruction in step S10221 includes: when it is determined according to the acquired information that at least one of the following conditions is met, output an operation result of refusing to execute the device control instruction:

[0085] The user does not have the control permission for at least one target device; the user does not have the control permission for the target function corresponding to the device control instruction; at least one target device does not meet the execution conditions corresponding to the device control instruction; the working environment of at least one target device does not meet the execution conditions corresponding to the device control instruction.

[0086] Another possible implementation of the embodiment of the present application for controlling at least one target device to perform corresponding operations includes: determining at least one target device corresponding to the device control instruction and / or the target function corresponding to the device control instruction based on the acquired information and the device control instruction; and controlling at least one target device to perform corresponding operations based on the at least one target device and / or the target function.

[0087] Another possible implementation of the embodiment of the present application for determining at least one target device corresponding to the device control instruction includes: performing domain classification processing based on the acquired information and the device control instruction to obtain the execution probability of each device; if the execution probability of each device is less than the first preset threshold, then output an operation result of rejecting the execution of the device control instruction, otherwise determine at least one target device corresponding to the device control instruction based on the execution probability of each device.

[0088] For the embodiment of the present application, perform domain classification processing based on the acquired information and the device control instruction to obtain the execution probability of each device; if the execution probability of each device is less than the first preset threshold, which indicates that the user does not have the control authority for at least one target device, then output an operation result of rejecting the execution of the device control instruction, otherwise determine at least one target device corresponding to the device control instruction based on the execution probability of each device.

[0089] Another possible implementation of the embodiment of the present application for determining the target function corresponding to the device control instruction includes: performing intention classification processing based on the acquired information and the device control instruction to determine the execution probability of each control function; if the execution probability of each control function is less than the second preset threshold, then output an operation result of rejecting the execution of the device control instruction, otherwise determine the target function corresponding to the device control instruction based on the execution probability of each control function.

[0090] For the embodiment of the present application, if the execution probability of each control function is less than the second preset threshold, which indicates that the user does not have the control authority for the target function corresponding to the device control instruction, then output an operation result of rejecting the execution of the device control instruction.

[0091] For the embodiment of the present application, after obtaining the device control instruction input by the user, when performing corresponding operations, by combining at least one of the user information, environmental information, and device information that may affect the device operation and / or user safety, potential dangers can be avoided, and thus the execution of the device control instruction can be controlled more intelligently, improving the user experience. For example, when the device control instruction input by the user is input by a child and the target device to be controlled is an oven, the operation result of rejecting the execution of the device control instruction may be combined with the user information, etc. to avoid the danger caused by the child operating the oven.

[0092] Another possible implementation of the embodiment of the present application, step S1022 may specifically include: step S10222 (not shown in the figure), where

[0093] Step S10222, based on the obtained information, control at least one target device to perform corresponding operations according to the target parameter information.

[0094] Wherein, the target parameter is the parameter information after changing the parameter information in the device control instruction.

[0095] Specifically, in step S10222, controlling at least one target to perform corresponding operations according to the target parameter information includes: when at least one of the following conditions is met, controlling at least one target device to perform corresponding operations according to the target parameter information:

[0096] The device control instruction does not contain a parameter value;

[0097] The parameter value included in the device control instruction does not belong to the parameter value range determined by the obtained information.

[0098] For example, the device control instruction input by the user is "Adjust the air conditioner to a suitable temperature.", which means that the device control instruction does not contain a parameter value. Then, the parameter information in the device control instruction can be changed according to at least one of the user information, device information, and environmental information to obtain the target parameter information, and control the air conditioner to perform corresponding operations according to the target parameter information. For example, the current environmental temperature is 32 degrees and the season is summer. When the user controls the air conditioner temperature in summer, it is generally set to 25 degrees. Therefore, the parameter value of the target parameter information can be set to 25 degrees, and the air conditioner will operate at 25 degrees.

[0099] Another example, the device control instruction input by the user is "Adjust the temperature of the air conditioner to 100 degrees". The parameter value (100 degrees) included in the device control instruction does not belong to the parameter value range (18 degrees - 32 degrees) determined by the obtained information. Then, the parameter information in the device control instruction is changed to obtain the target parameter information, and the air conditioner is controlled to perform corresponding operations according to the target parameter information. For example, the parameter value of the target parameter information is set to 25 degrees according to the user information and environmental information, and the air conditioner will operate at 25 degrees.

[0100] Another possible implementation of the embodiment of the present application, controlling at least one target to perform corresponding operations according to the target parameter information includes: performing sequence annotation processing on the device control instruction to obtain the parameter information in the device control instruction; based on the parameter information in the device control instruction and the obtained information, determine whether to change the parameter information in the device control instruction; if it is changed, based on the parameter information in the device control instruction and the obtained information, determine the changed target parameter information.

[0101] Another possible implementation of the embodiment of the present application is to determine whether to change the parameter information in the device control instruction based on the parameter information in the device control instruction and the acquired information, including: based on the parameter information in the device control instruction and the acquired information, through logistic regression processing, obtaining a logistic regression result; determining whether to change the parameter information in the device control instruction based on the logistic regression result; and / or, based on the parameter information in the device control instruction and the acquired information, determining the target parameter information after the change, including: based on the parameter information in the device control instruction and the acquired information, through linear regression processing, obtaining a linear regression result; determining the parameter information after the change based on the linear regression result.

[0102] Another possible implementation of the embodiment of the present application is that the method may further include: step Sa (not shown in the figure) and step Sb (not shown in the figure), where

[0103] Step Sa: Obtain a plurality of training data.

[0104] Step Sb: Based on the acquired training data and through a target loss function, train a processing model for changing the parameter information in the device control instruction.

[0105] Wherein, any training data includes the following information:

[0106] Device control instruction; parameter information in the device control instruction; indication information on whether the parameter in the device control instruction is changed; parameter information after the change; user information; environmental information; device information.

[0107] Another possible implementation of the embodiment of the present application is that before step Sb, it may further include: step Sc (not shown in the figure), where

[0108] Step Sc: Determine the target loss function.

[0109] Specifically, step Sc may specifically include: step Sc1 (not shown in the figure), step Sc2 (not shown in the figure), step Sc3 (not shown in the figure), and step Sc4 (not shown in the figure), where

[0110] Step Sc1: Based on the parameter information in the device control instruction in each training data and the predicted parameter information in the device control instruction of the model, determine the first loss function.

[0111] Step Sc2: Based on the indication information on whether the parameter in the device instruction in each training data is changed and the predicted indication information on whether it is changed by the model, determine the second loss function.

[0112] Step Sc3: Determine a third loss function based on the changed parameter information in each training data and the changed parameter information predicted by the model.

[0113] Step Sc4: Determine a target loss function based on the first loss function, the second loss function, and the third loss function.

[0114] For the embodiments of the present application, when controlling a target device to perform corresponding operations based on device control information input by a user, when determining target parameter information corresponding to controlling the target device to perform corresponding operations, by combining at least one of user information, device information, and environmental information, it is possible to determine whether to change the parameter information in the device control instruction and determine the changed parameter information, so as to control the target device to perform corresponding operations based on the modified parameter information, thereby improving the intelligence of performing corresponding operations based on the control instruction and enhancing the user experience.

[0115] For example, when the device control instruction is "raise the oven temperature to 240 degrees", based on the device information, it can be known that the current oven temperature is relatively high and the running time is relatively long, then the oven temperature can be adjusted to a lower temperature, thereby avoiding device damage and personal safety of the user during the operation process; for another example, when the user device instruction is "adjust the air conditioner temperature to 100 degrees", based on the device information, it can be known that the air conditioner temperature cannot be adjusted to 100 degrees, then the target parameter information can be adjusted to 32 degrees, etc., based on at least one of user information, device information, and environmental information, to avoid the situation where the device does not execute due to parameter problems in the device control instruction and enhance the user experience; for another example, the user device instruction input by a child is "turn on TV channel 10", combined with the user information, it can be known that children are not allowed to watch TV channel 10, then the TV channel 10 can be adjusted to a channel suitable for children to watch, thereby improving the intelligence of controlling the device to perform corresponding operations and further enhancing the user experience; for another example, the device control instruction input by the user is "adjust the air conditioner to a suitable temperature", it can be known that the device control instruction does not contain a parameter value, then the suitable temperature can be adjusted to 25 degrees, based on at least one of user information, environmental information, and device information, to avoid the situation where the device does not execute when receiving a vague instruction and improve the user experience.

[0116] Another possible implementation manner of the embodiments of the present application may further include, after step S1021: step Sd (not shown in the figure) and step Se (not shown in the figure), where

[0117] Step Sd: Convert the discrete information in the obtained information into a continuous dense vector.

[0118] Step Se. Determine an information representation vector corresponding to the obtained information based on the converted continuous dense vector and the continuous information in the obtained information.

[0119] Specifically, step S1022 may specifically include: controlling at least one target device to perform an operation based on the information representation vector corresponding to the obtained information and the device control instruction.

[0120] The embodiment of the present application provides a device control method. By obtaining at least one piece of information among user information, environmental information, and device information, and the device control instruction input by the user, it is possible to control at least one target device to perform corresponding operations based on the obtained information and the device control instruction. As can be seen from the above, compared with the prior art method of only controlling the device according to the control instruction input by the user, when the present application controls the device, in addition to considering the control instruction input by the user, it also takes into account at least one factor that may affect the device operation, such as user information, device information, and environmental information. Therefore, the device operation can be controlled more safely and flexibly. For example, obtain the device control instruction in the form of voice, text, key, gesture, etc. input by the user, and consider at least one of the user information, device information, and environmental information, and directly control the air conditioner device to turn on, turn off, or adjust the temperature, etc., so as to realize the safe and convenient control of the intelligent device to perform corresponding operations.

[0121] The following introduces the device control method in combination with specific embodiments, which may include three embodiments, namely Embodiment 1, Embodiment 2, and Embodiment 3. Among them, Embodiment 1 is mainly used to solve the problem in the prior art Problem 1 that when performing corresponding operations based on the device control instruction input by the user, the inputter of the device control instruction is not recognized, which may cause danger when some groups (such as children or the elderly) operate certain devices (such as ovens), or it is impossible to restrict some groups from operating a certain device or a certain function of a certain device. For example, it is impossible to restrict children from turning on the smart TV or impossible to restrict children from adjusting the channel; Embodiment 2 is mainly used to solve the technical problems existing in the prior art Problem 2, including: the parameter value in the device control instruction input by the user may cause damage to the user or the device; the parameter value in the device control instruction input by the user exceeds the range that the device can execute; the parameter value in the device control instruction input by the user is the parameter value restricted for the user to operate; the parameter value in the device control instruction input by the user is unclear, or there is no parameter value at all; Embodiment 3 is a combination of Embodiment 1 and Embodiment 2, which can be used to solve the technical problems existing in the prior art Problem 1 and the prior art Problem 2, as shown below:

[0122] Embodiment 1

[0123] Based on the obtained user information, and / or environmental information, and / or device information, as well as the device control instruction input by the user, at least one target device corresponding to the device control instruction and / or the target function corresponding to the device control instruction can be determined, and corresponding operations can be performed based on the determined at least one target device and / or target function.

[0124] In this embodiment, the user information of the user who inputs the device control instruction is mainly considered to determine whether the user has the permission to operate the target device and / or target function involved in the device control instruction. In addition to the user information considered in the first embodiment, the device information and / or environmental information can also be used to determine whether the user has the permission to operate the target device and / or target function. In the embodiments of the present application, the device control instruction input by the user can be input in ways such as voice, text, button, gesture, etc. This embodiment takes the device control instruction input by the user in the form of voice as an example for introduction, which is specifically as follows:

[0125] Step S201 (not shown in the figure), obtain the device control instruction input by the user.

[0126] For the embodiments of the present application, the device control instruction input by the user can be input by the user in ways such as voice, text, button, gesture, etc. There is no limitation in the embodiments of the present application. This embodiment takes the user inputting the device control instruction in the form of voice as an example for introduction.

[0127] Step S202 (not shown in the figure), obtain at least one of the following information: user information; environmental information; device information.

[0128] For the embodiments of the present application, the user information includes: user portrait information and / or the device control permission information of the user.

[0129] For the embodiments of the present application, the device information includes: the working state information of the device.

[0130] For the embodiments of the present application, the environmental information includes: the working environment information of the device.

[0131] For the embodiments of the present application, the user portrait information may include: user identity information and / or gender and / or age and / or user preference information, etc.

[0132] For the embodiments of the present application, the device control permission information of the user includes: the control permission of the user for the device and / or the control permission of the user for the target function.

[0133] For the embodiments of the present application, the working state information of the device includes at least one of the following: the current working state information of the device (such as temperature, humidity, channel, power, storage condition, continuously operating duration, etc.), the functions that the device can execute, the executable parameters, etc.

[0134] For the embodiments of the present application, the working environment information of the device includes: the current working environment of the device and / or the set target working environment (such as the optimal working environment, etc.).

[0135] Among them, the working environment includes temperature, humidity, pressure, and so on.

[0136] For the embodiments of the present application, step S201 may be executed before step S202, may also be executed after step S202, or may be executed simultaneously with step S202. It is not limited in the embodiments of the present application.

[0137] For the embodiments of the present application, a user portrait database and a user permission database may be set in advance, such as Figure 2a As shown, the user portrait database stores user portrait information, including: gender, age, user level (which can also be called user group), nickname, voice recognition template, and face recognition template, etc. The user group can be divided into four user groups, namely the master user group, the child user group, the elderly user group, and the guest user group. Among them, the user portrait data of the master user group, the child user group, and the elderly user group can be written during registration, and the user portrait data of the guest user group can be written during registration or during use; the user permission database records the device control permissions of each user to use each device. If the permissions are set separately according to user groups, then the user permission database can record the list of device categories that the master user group, the child user group, the elderly user group, and the guest user group can use (i.e., the control permissions of the user for the device) and the functions that each device can use (i.e., the control permissions of the user for the target function). Among them, this list can also be called an intent list, which contains the executable or non-executable intents of each user or user group. An intent includes the target device and / or target function that the user wants to control. For example, controlling the air conditioner belongs to an intent, and raising the temperature of the air conditioner also belongs to an intent.

[0138] Such as Figure 2a As shown, in the permission database, each user or each user group can set an intent list respectively. For example, for the child user group, intents ABCDE may not be allowed, and intent F is allowed to be executed; for the elderly user group, intents A and B may not be allowed to be executed, and intents C, D, E, and F are allowed to be executed; for the guest user group, intents B, D, and E may not be allowed to be executed, and intents A, C, and F are allowed to be executed; for the master user group, intents A, B, C, D, and E, and F are all allowed to be executed. This function list has default settings and can also be set manually, such as being set manually by the user of the master user group.

[0139] For the embodiments of the present application, user information is obtained through voiceprint recognition and / or image recognition. In the embodiments of the present application, after obtaining the device control instruction input by the user in a voice manner, based on voiceprint recognition, the user portrait information of the user who inputs the device control instruction is determined, such as at least one of identity information, gender information, age information, and user group information; if an image acquisition device is provided on some devices, the face image information of the user who inputs the device control instruction can be collected based on the image acquisition device, and based on face image detection technology, the user portrait information of the user who inputs the device control instruction is determined, such as at least one of identity information, gender information, age information, and user group information. Specifically, when the corresponding voice signal and face image signal are collected, identity authentication can be performed by comparing the face authentication information with the user face recognition template of each user in the user portrait database to determine the user identity. When only the collected voice signal is available and the face image signal is not collected, identity authentication is performed by comparing the voiceprint authentication information with the user voiceprint recognition template of each user in the user portrait database to determine the user identity (considering that in the smart home scenario, cameras are often installed on computers and TVs, and when the speaker is in the kitchen, bedroom, etc., the image signal may not exist).

[0140] When the authentication is passed (that is, when the speaker feature has a high similarity with the voiceprint recognition template (or face recognition template) of a certain user in the existing user portrait database), the user portrait of the user in the user database is output, including gender, age, user group, etc.; if the identity authentication fails, it indicates a new user, and a new user portrait data is established and written, and the obtained gender data, age data, etc. are written. The user group can be set to guest, and the newly created user portrait data, including gender, age, user group, etc., is output. Then, according to the user group in the output user portrait data, the user permissions (also called the device control permissions of the user) corresponding to the user group are queried in the permission database and output. The user permissions and user portrait information are integrated into multimodal information, and then based on the integrated multimodal information and the device control instruction input by the user, at least one target device corresponding to the device control instruction and / or the target function corresponding to the device control instruction are determined, such as in step S203.

[0141] For the embodiments of the present application, methods such as Markov random field and convolutional neural network can be used to perform voiceprint recognition on the device control instructions input by voice, so as to determine at least one of the identity information, gender information, and age information of the user who inputs the device control instructions. Taking the neural network method as an example, after training a voiceprint classification network with a large amount of data, use this network to extract feature vectors from the user's voiceprint and save them as templates; during authentication, compare the cosine distance between the features of the voiceprint to be authenticated and each feature template in the database. If it exceeds the threshold, the authentication is considered successful, otherwise it fails; for recognizing age information and / or gender information from the voice, a convolutional neural network can also be used, which will not be elaborated in the embodiments of the present application.

[0142] For the embodiments of the present application, after obtaining the device control instructions input by the user by voice, voice noise reduction processing can also be performed on the device control instructions input by the user. In the embodiments of the present application, the voice noise reduction technology can include: multi-microphone collaborative noise reduction technology, convolutional neural network noise reduction technology. This will not be elaborated in the embodiments of the present application.

[0143] Step S203 (not shown in the figure), based on the obtained information and the device control instructions, determine at least one target device corresponding to the device control instructions and / or the target function corresponding to the device control instructions.

[0144] Specifically, based on the obtained information and the device control instructions, control at least one target device to perform corresponding operations, including: based on the information representation vector corresponding to the obtained information and the device control instructions, control at least one target device to perform operations.

[0145] Furthermore, introduce the method of converting multi-modal information (obtained information) into an information representation vector corresponding to the multi-modal information. Specifically, obtain at least one of user information, environmental information, and device information. After that, it further includes: converting the discrete information in the obtained information into a continuous dense vector; according to the converted continuous dense vector and the continuous information in the obtained information, determine the information representation vector (multi-modal information representation vector) corresponding to the obtained information.

[0146] For the embodiments of the present application, the discrete information in the obtained information can be converted into a continuous dense vector through a transformation matrix. In the embodiments of the present application, through the transformation matrix, it is converted into a continuous dense vector; connect the converted continuous dense vector and the information in the obtained information that does not belong to discrete values to obtain a joint vector, and then perform preset processing on the joint vector to obtain the information representation vector corresponding to the obtained information.

[0147] Specifically, such as Figure 3aAs shown, when encoding the acquired information (multi-modal information) (multi-modal information encoding), for example, gender, permissions, and favorite channels, since they are discrete values, they need to be converted into continuous dense vectors through an encoding matrix. Age, favorite temperature, etc. can be directly input. The encoded multi-modal information is concatenated to obtain a joint vector, and after passing through a fully connected layer and a sigmoid activation function, an information representation vector (multi-modal information representation vector) corresponding to the acquired information is obtained. For example, the information corresponding to gender is processed through a gender encoding matrix to obtain a continuous dense vector corresponding to gender information; the device control permission information of the user is processed through a permission encoding matrix to obtain a continuous dense vector corresponding to permission information; the favorite channel is processed through an emotion encoding matrix to obtain a continuous dense vector corresponding to the favorite channel.

[0148] The following details the specific implementation methods for determining at least one target device corresponding to the device control instruction and / or the target function corresponding to the device control instruction based on the acquired information and the device control instruction:

[0149] For the embodiments of the present application, at least one target device corresponding to the device control instruction and / or the target function corresponding to the device control instruction can be determined based on the device control instruction by inputting the user group information to which the user belongs; or at least one target device corresponding to the device control instruction and / or the target function corresponding to the device control instruction can be determined according to the device control instruction by inputting the age information and / or gender information of the user; or at least one target device corresponding to the device control instruction and / or the target function corresponding to the device control instruction can be determined based on the device control instruction by inputting the user group information to which the user belongs and the device control instruction by inputting the age information and / or gender information of the user.

[0150] Of course, at least one target device corresponding to the device control instruction and / or the target function corresponding to the device control instruction can be determined based on the acquired information and the device control instruction and through a trained model. For example, at least one target device corresponding to the device control instruction and / or the target function corresponding to the device control instruction can be determined based on the acquired information (multi-modal information) and the device control instruction and through a trained Domain Classifier (DC); and / or at least one target device corresponding to the device control instruction and / or the target function corresponding to the device control instruction can be determined based on the acquired information (multi-modal information) and the device control instruction through a trained Intent Classifier (IC).

[0151] Specifically, determining at least one target device corresponding to the device control instruction in step S203 includes: performing domain classification processing based on the acquired information and the device control instruction to obtain the execution probability of each device; if the execution probability of each device is less than the first preset threshold, output an operation result of rejecting the execution of the device control instruction, otherwise determine at least one target device corresponding to the device control instruction based on the execution probability of each device.

[0152] Specifically, determining at least one target device corresponding to the device control instruction based on the execution probability of each device includes: determining the device corresponding to the maximum probability among the execution probabilities of each device as the device corresponding to the device control instruction.

[0153] For example, the domain classification results (execution probabilities of each device) are 0.91, 0.01, and 0.08 corresponding to device 1, device 2, and device 3 respectively. Among them, the first preset threshold is 0.5, then the device corresponding to the device control instruction is determined to be device 1.

[0154] For another example, the domain classification results (execution probabilities of each device) are 0.49, 0.48, and 0.03. Since the execution probability of each device is not greater than 0.5, an operation result of rejecting the execution of the device control instruction is output.

[0155] For the embodiments of the present application, since determining at least one target device corresponding to the device control instruction can be determined through a model (such as a domain classifier), before determining at least one target device corresponding to the device control instruction, it may further include: training the model (domain classifier), which will be introduced below taking the domain classifier as an example.

[0156] Specifically, the training data is (s i , m i , d i ), where s i represents the input data (device control instruction) statement text, m i represents multimodal information, including the gender, permission, age, etc. of the user who inputs the device control instruction, d i represents the label of this statement, that is, the domain it belongs to (that is, which device it belongs to), i represents the index of a piece of training data in the training dataset, j represents the device index, d ij represents the probability that the domain of statement i is device j (which can be called the execution probability of the device), d i is in one-hot encoding form, that is, when this statement belongs to the jth device, d ij is 1, and d ik (k≠j) is 0; if this statement is an over - permission statement, that is, the user does not have the control permission for the target device (for example, the command statement of a 4 - year - old user is "Bake a sweet potato for me", and the oven is not allowed for children to use), then d iAll elements are 0, and the loss function for training is as follows:

[0157]

[0158] where M is the total number of devices, N is the total number of input data statements, is the predicted output of the model, that is, the probability that the domain of statement i predicted by the model is device j. When the predicted output is exactly the same as d ij the loss is 0.

[0159] After training the domain classifier based on the above training method, based on the obtained information and device control instructions, and through the trained domain classifier, the domain classification result is obtained. Specifically, the input data is (s, m), where s is the text information corresponding to the device control instruction, m is the multimodal information. Using the above trained DC model, the predicted output (domain classification result), the largest element in is If (c is the first preset threshold, which can be selected as 0.5), then the classification result of this statement is the kth device. If it means that the execution probabilities of all devices are less than the first preset threshold, indicating that this statement belongs to the protected situation and the user does not have the control authority for the target device. Therefore, the device control instruction can be refused to execute (for example, the command statement of a 4-year-old user is "Bake me a sweet potato", and the oven is not allowed for children to use. During training, all elements of the label of this statement are 0. Therefore, when the trained model predicts, the predicted output of this statement is close to 0 and will also be less than the threshold c, so DC will not give a classification and will refuse to execute).

[0160] Specifically, such as Figure 2bAs shown, the input text can be a device control instruction in the text format input by the user, or a device control instruction obtained by converting the voice format device control instruction input by the user through voice-text conversion. After the text undergoes word encoding (word vector conversion and position encoding), an encoded vector is obtained. Then, through a convolutional neural network and a self-attention model, a text representation vector can be obtained. Among them, after the text undergoes word vector conversion, it becomes (w1, w2, w3...), where w1, w2, and w3 represent the word vectors corresponding to each word in the sentence (device control instruction); the position encoding is a function (f(1), f(2), f(3)), which is a function of the position index of the word vector (specifically, a vector (f(1), f(2), f(3)) is concatenated after each word vector, and these vectors are calculated using a function related to the position. This function can be implemented using various methods, and the more commonly used methods include using the sin or cos function for calculation). Then, these two parts are added together to obtain the encoded vector (w1 + f(1), w2 + f(2), w3 + f(3)..). Then, the obtained encoded vector passes through a convolutional neural network and a self-attention model to obtain a text representation vector. After the multi-modal information processing module processes the multi-modal information, a multi-modal information representation vector is obtained, that is, the information representation vector corresponding to the obtained information. After the multi-modal information representation vector and the text representation vector are concatenated, a joint vector is obtained (for example, vector (a1, a2, a3...) and vector (b1, b2, b3...), and after concatenation, it is (a1, a2, a3..., b1, b2, b3,...)). Then, the joint vector is input into the fully connected layer and the domain classification results are output: the execution probability of domain A (device A), the execution probability of domain B (device B), and the execution probability of domain C (device C).

[0161] Therefore, the difference between the DC model in this application and the existing DC model is that multi-modal information is added as input, enabling the model to perform domain classification (determine at least one target device) by referring to the current user profile information and / or environmental information and / or device information. For example, when the oven temperature is too high and the user inputs "bake a cake for one hour", the domain classifier will not classify this statement into the oven domain but will refuse to execute it to ensure the safety of the oven and the user.

[0162] In the embodiments of the present application, an independent DC model can be deployed for each device, that is, domain classification processing is performed for each device separately. Or a DC model can be shared by multiple devices, and this model can be deployed in the cloud. For example, after a device receives a device control instruction input by a user, it uploads the instruction to the cloud, and the cloud performs domain classification based on at least one of device information, user information, and environmental information and the received device control instruction. If the execution probability of each device in the domain classification result is less than a first preset threshold, the device that receives the user input can be instructed to output an operation result of rejecting the execution of the device control instruction. If the maximum execution probability in the domain classification result is not less than the first preset threshold, the device with the maximum execution probability can be determined as the target device, and the instruction is transmitted to the target device for subsequent operations (such as intent classification processing or sequence labeling processing, etc.), or the cloud continues to perform operations such as intent classification processing and sends the determined final operation instruction to the target device for execution. The shared DC model can also be deployed on the terminal. For example, a DC model is deployed in device A. After device A receives a device control instruction input by a user, it performs domain classification based on at least one of device information, user information, and environmental information and the received device control instruction. If the execution probability of each device is less than the first preset threshold, an operation result of rejecting the execution of the device control instruction can be output. If the maximum execution probability is not less than the first preset threshold, the device with the maximum execution probability can be determined as the target device, and the instruction is transmitted to the target device for subsequent operations (such as intent classification processing or sequence labeling processing, etc.), or device A continues to perform operations such as intent classification processing and sends the determined final operation instruction to the target device for execution.

[0163] In the embodiments of the present application, when at least one target device is determined based on the acquired information and the device control instruction, the target function corresponding to the device control instruction can be determined based on the acquired information and the device control instruction. In the embodiments of the present application, the target device can be determined first, and then the target function can be determined. For example, each intelligent device can deploy an IC model separately and share a DC model. At this time, the DC model can first determine the target device, and then the IC model of the target device can determine the target function. Or the target device and the target function can be determined simultaneously, and the specific execution sequence is not specifically limited here. The method for determining the target function corresponding to the device control instruction is specifically described as follows:

[0164] Determining the target function corresponding to the device control instruction in step S203 includes: performing intent classification processing based on the acquired information and the device control instruction to determine the execution probability of each control function; if the execution probability of each control function is less than a second preset threshold, output an operation result of rejecting the execution of the device control instruction, otherwise determine the target function corresponding to the device control instruction based on the execution probability of each control function.

[0165] Specifically, based on the acquired information and device control instructions, intent classification processing is performed through a model (intent classifier). In the embodiments of the present application, when multiple target devices are determined, intent classification can be performed through only one model (shared IC model), or through the models in each target device.

[0166] In the embodiments of the present application, an independent IC model can be deployed for each device, that is, domain classification processing is performed for each device separately. It is also possible to share one IC model among multiple devices. This model can be deployed in the cloud. For example, after a certain device receives a device control instruction input by a user, it uploads it to the cloud. The cloud performs intent classification based on at least one of device information, user information, and environmental information and the received device control instruction. If the execution probability of each control function in the intent classification result is less than a second preset threshold, the device that received the user input can be instructed to output an operation result of rejecting the execution of the device control instruction. If the maximum execution probability in the intent classification result is not less than the second preset threshold, the control function with the maximum execution probability can be confirmed as the target function, and the instruction is transmitted to the target device for subsequent operations (such as sequence labeling processing, etc.), or the cloud continues to perform operations such as sequence labeling processing and sends the determined final operation instruction to the target device for execution. The shared IC model can also be deployed at the terminal. For example, an IC model is deployed in device A. After device A receives a device control instruction input by a user, it performs intent classification based on at least one of device information, user information, and environmental information and the received device control instruction. If the execution probability of each control function is less than the second preset threshold, an operation result of rejecting the execution of the device control instruction can be output. If the maximum execution probability is not less than the second preset threshold, the control function with the maximum execution probability can be confirmed as the target function, and the instruction is transmitted to the target device for subsequent operations (such as sequence labeling processing, etc.), or device A continues to perform operations such as sequence labeling processing and sends the determined final operation instruction to the target device for execution.

[0167] Further, since intent classification processing can be performed through a model (intent classifier), before performing intent classification processing based on the acquired information and device control instructions through the model (intent classifier), it also includes: training the model (intent classifier), and the specific method is as follows:

[0168] Train the intent classifier through the following loss function:

[0169]

[0170] where M is the total number of control functions, N is the total number of input data statements, j represents the function index, I ijIndicates that the intention of statement i is the probability of function j (which can be called the execution probability of the control function). Is the predicted output of the model, that is, the probability that the intention of statement i predicted by the model is function j. i Is the label of the i-th training data. i Is in one-hot encoding form, that is, when this statement belongs to the j-th function. ij Is 1. ik (k≠j) is 0. If this training statement is an over-authorization statement, that is, the user does not have the control permission for the target function (for example, the command statement of a 4-year-old user is "uninstall the XX APP of the TV". Although the TV is open to children, the uninstall app function of the TV is not allowed for children), then i All elements of are 0. When the predicted output And ij Are exactly the same, the loss is 0.

[0171] Among them, when training the intention classifier, the training samples still contain the acquired information (multi-modal information). This multi-modal information can be initialized with the multi-modal information in the trained domain classifier. For example, the weights of the multi-modal intention classifier can be initialized by some methods (such as: using the vectors corresponding to some multi-modal information as the weights of the classifier) to speed up the training speed.

[0172] For the embodiments of the present application, after training the intention classifier in the above manner, based on the device control instruction input by the user and the acquired information, and through the trained intention classifier, the target function corresponding to the device control instruction can be determined. The specific method for determining the target function corresponding to the device control instruction is as follows:

[0173] The input of the intention classifier is (s, m), where s is the text information corresponding to the device control instruction, and m is the multi-modal information. First, the domain (target device) is obtained through the DC model, and then the trained IC model in this domain can be used to obtain the predicted output The largest element in is If (c is a set threshold, and 0.5 can be selected), then the classification result of this device control instruction is the k-th function (target function). If It means that this device control instruction belongs to the protected situation and is refused to be executed. If the owner sets the permissions of children, the elderly, guests, etc. in the user permission database, and the k-th function is just in the shielding list, then it is refused to be executed. For example Figure 4As shown, the device control instructions input by the user include: intention A, intention D, or intention F. If the owner has not set user permissions, the intention classifier directly outputs intention A, intention D, or intention F. If the owner sets user permissions (intention A, intention C, and intention F are allowed to operate, while intention B, intention D, and intention E are not allowed to operate), the intention classifier directly outputs intention A and intention F and rejects the execution of intention D.

[0174] For example, when a child says "Delete the channel list of the TV", this device control instruction will be assigned to the TV domain by the domain classifier. However, the function of deleting the channel list on the TV is not available to children. During training, if the user who inputs this sentence is a child, all elements of the label of this sentence are 0. Therefore, when the trained intention classifier makes a prediction based on user information, the prediction output of this device control instruction is close to 0 and also less than the threshold c. As a result, the intention classifier will not give an intention classification and will reject the execution of this target function.

[0175] Specifically, as Figure 3b shown, the input text can be the device control instruction in the text format input by the user, or the device control instruction obtained by converting the voice format device control instruction input by the user through voice-text conversion. After the text undergoes word encoding (word vector conversion and position encoding), an encoded vector is obtained. Then, through a convolutional neural network and a self-attention model, a text representation vector can be obtained. Among them, after the text undergoes word vector conversion, it becomes (w1, w2, w3...), where w1, w2, and w3 represent the word vectors corresponding to each word in the sentence (device control instruction); the position encoding is a function (f(1), f(2), f(3)), which is a function of the position index of the word vector. Then, these two parts are added together to obtain the encoded vector (w1 + f(1), w2 + f(2), w3 + f(3)..). Then, the obtained encoded vector passes through a convolutional neural network and a self-attention model to obtain a text representation vector. After the multi-modal information processing module processes the multi-modal information, a multi-modal information representation vector is obtained, that is, the information representation vector corresponding to the obtained information. After the multi-modal information representation vector and the text representation vector are concatenated, a joint vector is obtained (for example, vector (a1, a2, a3...) and vector (b1, b2, b3...), and after concatenation, it is (a1, a2, a3..., b1, b2, b3,...)). Then, the joint vector is input into the fully connected layer and the intention classification results are output: the execution probability of function A (intention A), the execution probability of function B (intention B), and the execution probability of function C (intention C).

[0176] Further, based on the acquired information and the device control instruction, and through an intent classifier, when determining the target function, the acquired information (multi-modal information) input to the intent classifier is actually its corresponding representation vector. In the embodiments of the present application, the manner of obtaining the information representation vector corresponding to the acquired information from the acquired information is as described in the above embodiments and will not be elaborated here.

[0177] It should be noted that: if the target device corresponding to the device control instruction is not determined based on the acquired information and the device control instruction (directly output an instruction to reject the corresponding operation), then the target function corresponding to the device control instruction may not be determined, and the device control instruction may not be marked.

[0178] For step S203, the embodiments of the present application provide a specific example: a child's device control instruction "turn on the oven" input by voice, where the oven is a device prohibited for the child to operate, and the operation is rejected through a domain classifier; for another example, the child's device control instruction input by voice is "delete XXX APP on the mobile phone". Here, the mobile phone is a device allowed for the child to operate, but the intention of deleting the APP is not allowed for the child to operate. Therefore, the target device of this device control instruction is determined as the mobile phone through the domain classifier, and the operation is rejected through the intent classifier; for the prior art, multi-modal information is not considered. When the child inputs the device control instruction "adjust the air conditioner to 30 degrees" by voice, the control instruction "adjust the air conditioner to 30 degrees" is passed through the domain classifier to obtain the domain classification result (oven: 0.01; washing machine: 0.02; air conditioner: 0.37), that is, the target device is the air conditioner. The intent classification result is obtained through the intent classifier (intent A, turn on the air conditioner: 0.01; intent B, turn off the air conditioner: 0.02; intent C, set the temperature: 0.97), that is, the target function is to set the temperature. Then, through the sequence tagger, the parameter information in the device control instruction (temperature: "30") is obtained, as Figure 3c shown. Therefore, directly executing the device control instruction input by the user in the prior art may pose a safety risk to the device or the user, and the control of the device is not flexible. In the present application, the target device and / or the target function and / or the target parameter information are determined based on at least one of the device information, the user information, and the environmental information, fully considering various factors that may affect the safe operation of the device, and enabling the user to conveniently perform permission control on the device, which can greatly improve the safety and flexibility when the user controls the device.

[0179] Step S204 (not shown in the figure): Mark the device control instruction to obtain the target parameter information.

[0180] For the embodiments of the present application, the device control instruction is input into a Sequence Tagger (ST) model for annotating the device control instruction to obtain target parameter information. In the embodiments of the present application, step S204 in this embodiment may not annotate the device control instruction based on the obtained information to determine the target parameter information.

[0181] Among them, in step 203, based on the obtained information and the device control instruction, determining the target function corresponding to the device control instruction may be executed in parallel or serially with step S204, which is not limited in the embodiments of the present application. Of course, when in step S203, based on the obtained information and the device control instruction, an instruction to reject executing the corresponding operation is output, then step S204 may not be executed.

[0182] Since the sequence annotation process is processed by a sequence tagger, the model structure of the sequence tagger will be introduced first:

[0183] As Figure 5 shown, the ST model includes: an encoding layer and a decoding layer. The encoding layer includes: word encoding, a Long Short-Term Memory (LSTM) layer, and an attention layer; the decoding layer includes: an LSTM layer and a Multilayer Perceptron (MLP) layer. Among them, x1, x2....xm are the device control instructions of the user. The encoding layer also adopts an encoding form that combines word vector conversion and position encoding. After encoding, each word is represented as a vector with a fixed dimension; the LSTM layer is used for encoding to extract the features h1, h2.....hm of each word; y1, y2....yk are the tags corresponding to x1, x2....xm (the BMO tagging method can be adopted. B indicates that the word is the starting position of the parameter, M indicates that the word is the middle position or the ending position of the parameter, and O indicates that the word is not a parameter). After passing y1, y2...yk through the LSTM layer, it is represented as a hidden state C. Using C and h1, h2.....hm, a vector d is calculated through the attention layer. After passing through multiple MLPs, the vector d becomes the vector f, and after passing through the multilayer perceptron, the label yk+1 at the next moment is output (i.e., the target parameter information). Among them, Figure 5 EOS (English full name: End of sentence) in

[0184] Therefore, the device control instruction is annotated through the ST model (Viterbi decoding method) to obtain the target parameter information.

[0185] Further, since the ST model is a trained ST model in the process of annotating device control instructions through the ST model, before annotating the device control instructions through the ST model, it further includes: training the ST model through training samples and a loss function, and the specific method is as follows:

[0186] The training sample set is (s i , y i , c i , v i , m i ), where s i represents the text information corresponding to the input device control instruction, y i represents the BMO label of this instruction (for example, S i is "set the air conditioner to 30 degrees", and yi is "O O O O B M"), and i is the index of each piece of data in the training sample set. The loss function for training is:

[0187]

[0188] where y ij represents the BMO annotation of the j-th word of the i-th training sample, is the BMO result of the j-th word of the i-th training sample predicted by the model.

[0189] Step S205 (not shown in the figure), based on at least one target device and / or target function and / or target parameter information, control at least one target device to perform corresponding operations.

[0190] For the embodiments of the present application, after determining at least one target device and / or target function and / or target parameter information through step S203, step S204, and step S205, control at least one target device to perform corresponding operations.

[0191] For example, if the device control information input by the user is "adjust the air conditioner temperature to 30 degrees", then the target device is the air conditioner, the target function is to adjust the temperature, and the target parameter information is 30 degrees. Then, according to the determined information, control the air conditioner to adjust the temperature to 30 degrees.

[0192] Further, since there are cases where the domain classifier and the intent classifier directly output operation results of rejecting the execution of device control instructions, this embodiment may further include step S206, where,

[0193] Step S206 (not shown in the figure), based on the obtained information and the device control instruction, output an operation result of rejecting the execution of the device control instruction.

[0194] Specifically, the operation result of rejecting the execution of the device control instruction is output, including: when it is determined based on the acquired information that at least one of the following conditions is met, the operation result of rejecting the execution of the device control instruction is output:

[0195] The user does not have the control authority for at least one target device; the user does not have the control authority for the target function corresponding to the device control instruction; at least one target device does not meet the execution conditions corresponding to the device control instruction; the working environment of at least one target device does not meet the execution conditions corresponding to the device control instruction. Among them, the above execution conditions can be set in advance. For example, the execution condition for adjusting the air conditioner temperature to 30 degrees can be: the ambient temperature is lower than 30 degrees, or the execution condition for adjusting the oven temperature to 260 degrees can be: the continuous operation time of the oven is less than 3 hours.

[0196] For example, the device control instruction input by a child is "increase the oven temperature to 240 degrees". Based on the acquired user information, it is known that the user who inputs the device control instruction is a child, and through the permission database, it is known that the child cannot operate the oven, that is, the child does not have the control authority for the target device. Then, the operation corresponding to the device control instruction is directly rejected.

[0197] Another example, the device control instruction input by a child is "delete the XX application on the TV". Based on the acquired user information, it is known that the user of the device control control instruction is a child, and through the permission database, it is known that the child can operate the TV, but cannot "delete the application" (target function), that is, the child does not have the control authority for the target function. Then, the operation corresponding to the device control instruction is directly rejected.

[0198] Another example, a certain air conditioner does not have the "dehumidification" function. The device control instruction input by the user is "turn on the dehumidification function of the air conditioner", that is, the air conditioner does not meet the execution condition "dehumidification" corresponding to the device control instruction. Then, the operation corresponding to the device control instruction is directly rejected.

[0199] Another example, the current indoor temperature is 30 degrees or it is currently summer. The device control instruction input by the user is "adjust the air conditioner temperature to 32 degrees", which means that the working environment of the air conditioner does not meet the execution conditions corresponding to the device control instruction. Then, the operation corresponding to the device control instruction is directly rejected.

[0200] For the embodiments of the present application, when it is determined based on the acquired information and / or the device control instruction that the operation result of not executing the device control instruction is obtained, the operation can be not executed, and a notification message can be output to inform that the control instruction is currently rejected, or only the corresponding operation can be not executed.

[0201] Furthermore, for Embodiment 1, a device control system is also introduced (taking the example that the user inputs the device control instruction by voice), such asFigure 6a As shown, the system is divided into a sound processing module, an image processing module, a multi-modal information processing module, a speech conversion module, a semantic understanding module, a dialogue management (DM) module, a speech synthesis module, and an execution module. The speech conversion module can also be called an automatic speech recognition (ASR) module, the semantic understanding module can also be called a natural language understanding (NLU) module, and the speech synthesis module can also be called a text-to-speech (TTS) module. Further, the DM module can further include a natural language generation (NLG) module. Among them, after the audio acquisition device (microphone) acquires the sound signal, the sound processing module performs noise reduction and identity recognition, and outputs the sound signal after noise reduction processing and identity authentication information. After the camera acquires the image information, the image processing module performs face extraction and face recognition, and outputs identity authentication information. The identity authentication information output by the above image processing module and sound processing module will be integrated into multi-modal information by the multi-modal information processing module; the sound signal output by the sound processing module will be converted into text information by the speech conversion module; the text information and the multi-modal information are jointly input to the semantic understanding module, and the semantic understanding module outputs the domain (target device), intention (target function), and label (target parameter information) of the statement to the dialogue management module and the execution module; the dialogue management module generates reply text, which is synthesized by the speech synthesis module and then replied; the execution module performs corresponding operations. The prior art does not consider image signals and identity authentication, and when performing semantic understanding, the prior art does not consider multi-modal information ( Figure 6a only an example showing that the multi-modal information includes user information is shown, and the multi-modal information in the present application can also include device information and / or environmental information, which is not shown in Figure 6a ). Specifically, see the introduction to the multi-modal information processing module:

[0202] The function of this module is to process the information obtained by the sound processing module and the image processing module, and summarize and process to obtain multi-modal information. The composition of this module is as Figure 6bAs described above, the multimodal information includes user image information (including age, gender, etc.) obtained from the user portrait database and user permission data (device control permissions of the user) obtained from the permission database. Specifically, when both the sound processing module and the image processing module have collected the corresponding signals, the identity authentication compares the face authentication information output by the image processing module with the user face recognition templates of each user in the user portrait database to determine the user identity. If the authentication is passed, that is, if it is determined as an existing user according to the identity authentication result, the user portrait of this user in the user portrait database is obtained and output, including gender, age, user group, etc. If the identity authentication fails, that is, if it is determined as a new user according to the identity authentication result, a new user portrait data is created and written in the user portrait database, the gender and age data obtained from the sound processing module and the image processing module are written, and the newly created user portrait data, including gender, age, user group, etc., is output. The user permissions corresponding to the user group are queried in the permission database according to the user group in the output user portrait data and output. The user permissions output and the user portrait data are integrated into multimodal information. This multimodal information is output to the semantic understanding module, that is, the Natural Language Understanding (NLU) module, which can also be called the multimodal NLU module.

[0203] Among them, the system architecture diagram of the prior art is as Figure 7 shown. After the sound signal is denoised by the sound processing module, it is converted into text by the speech conversion module. After being processed by the semantic understanding module, the domain, intention, and label of the statement are obtained and output to the dialogue management module and the execution module. The dialogue management module generates reply text, which is synthesized by the speech synthesis module and then replied; executed by the execution module. Therefore, directly executing the device control instructions input by the user in the prior art may pose security risks to the device or the user, and the control of the device is not flexible. However, the present application determines the target device and / or target function and / or target parameter information according to at least one of the device information, user information, and environmental information, fully considering various factors that may affect the safe operation of the device, and also enabling the user to conveniently control the device permissions, which can greatly improve the security and flexibility when the user controls the device.

[0204] Embodiment 2

[0205] This embodiment mainly introduces determining whether the parameter information in the device control instruction input by the user needs to be changed based on the obtained user information and / or environmental information and / or device information. If it needs to be changed, the changed parameter information is output to solve the technical problem in the above-mentioned technical problem two (wherein, in this embodiment, when determining at least one target device and / or target function corresponding to the device control instruction, the obtained information (including: user information, environmental information, and device information) can be not considered), specifically as follows:

[0206] Step S301 (not shown in the figure), obtain the device control instruction input by the user.

[0207] For the embodiment of the present application, the method of obtaining the device control instruction input by the user in step S301 can be seen in the above-mentioned step S201. It will not be elaborated in this embodiment.

[0208] Step S302 (not shown in the figure), based on the obtained device control instruction, determine at least one target device corresponding to the device control instruction and / or the target function corresponding to the device control instruction.

[0209] For the embodiment of the present application, determining at least one target device corresponding to the device control instruction and / or the target function corresponding to the device control instruction based on the obtained device control instruction includes: based on the obtained device control instruction, and based on the domain classifier, determine at least one target device corresponding to the device control instruction; based on the obtained device control instruction, and based on the intent classifier, determine the target function corresponding to the device control instruction.

[0210] For the embodiment of the present application, the model structure of the domain classifier is as Figure 8 shown. Its input text is the device control instruction input by the user in text form, or the text format device control instruction obtained by converting the device control instruction input by the user in voice format through voice-text conversion. The text format device control instruction is encoded (word vector conversion and position encoding) to obtain an encoded vector, and then through a convolutional neural network and a self-attention model, a text expression vector can be obtained. Among them, after word vector conversion, it is (w1, w2, w3...), where w1, w2, w3... represent the word vectors corresponding to each word in the device control instruction; the position encoding is a function (f(1), f(2), f(3)), which is a function of the position index of the word vector; adding these two parts together to obtain the encoded vector (w1 + f(1), w2 + f(2), w3 + f(3)..), and then inputting it into the fully connected layer to output the classification result: the execution probability of domain A (device A), the execution probability of domain B (device B), the execution probability of domain C (device C). Select the device with the highest execution probability as the target device.

[0211] Among them, the model structure of the domain classifier can also adopt Figure 2b the structure shown in the figure, that is, after the multi-modal information processing module processes the multi-modal information, a multi-modal information representation vector is obtained. After being connected with the text expression vector, a joint vector is obtained. The joint vector is input into the fully connected layer and then the domain classification result is output.

[0212] For the embodiments of the present application, the structure of the intent classifier IC is as Figure 9 shown in the figure. It is mainly used to determine the target function corresponding to the device control instruction based on the device control instruction. The structure of the IC is the same as that of the DC. Its input text is the device control instruction input by the user in text form, or the device control instruction in text format obtained by converting the device control instruction input by the user in voice format through voice-to-text conversion. The device control instruction in text format undergoes word encoding (word vector conversion and position encoding) to obtain an encoded vector, and then through a convolutional neural network and a self-attention model, a text expression vector can be obtained. Among them, after word vector conversion, it is (w1, w2, w3...), where w1, w2, w3... represent the word vectors corresponding to each word in the device control instruction; the position encoding is a function (f(1), f(2), f(3)), which is a function of the position index of the word vector; adding these two parts together to obtain the encoded vector (w1 + f(1), w2 + f(2), w3 + f(3)..), and then inputting it into the fully connected layer and outputting the intent classification result: the execution probability of function A (intent A), the execution probability of function B (intent B), and the execution probability of function C (intent C).

[0213] Among them, the model structure of the intent classifier can also adopt Figure 3b the structure shown in the figure, that is, after the multi-modal information processing module processes the multi-modal information, a multi-modal information representation vector is obtained. After being connected with the text expression vector, a joint vector is obtained. The joint vector is input into the fully connected layer and then the intent classification result is output.

[0214] Furthermore, there can be one domain classifier, that is, a shared domain classifier, and multiple intent classifiers (that is, each device corresponds to an intent classifier). It is also possible that both the domain classifier and the intent classifier are one, such as performing domain classification and intent classification in the cloud. This is not limited in the embodiments of the present application.

[0215] In the embodiments of the present application, when determining at least one target device corresponding to a device control instruction based on a domain classifier, the used domain classifier is a pre-trained domain classifier; when determining the target function corresponding to the device control instruction based on an intent classifier, the used intent classifier is also a pre-trained one. The specific training method is as follows: training the domain classifier based on multiple pieces of first training data; training the intent classifier based on multiple pieces of second training data. Any piece of first training data includes: a device control instruction and a label of the domain (target device) corresponding to the device control instruction; any piece of second training data includes: a device control instruction and a label of the target function corresponding to the device control instruction; the more specific training method is not elaborated in the embodiments of the present application.

[0216] Step S303 (not shown in the figure), obtain at least one of the following information: user information; environmental information; device information.

[0217] The embodiments of the present application do not impose any limitation on the execution order among step S301, step S302, and step S303.

[0218] In the embodiments of the present application, the method for obtaining at least one of user information, environmental information, and device information is described in detail in Embodiment 1. Among them, Embodiment 1 mainly introduces the method for obtaining user information. In the embodiments of the present application, the method for obtaining environmental information is mainly introduced. Specifically, as Figure 10 shown, environmental information (including: the temperature of the current environment, the air pressure of the current environment, etc.) is collected by sensors. Appropriate environmental parameters (which can also be called optimal working environment information) are stored in the environmental database. In addition, device information can be obtained in the same way, such as collecting device information (including: the current working temperature of the device, the working humidity, etc.) and the appropriate working parameters (which can also be called optimal working state information) stored in the device database. Multimodal information can be obtained through the obtained device information and / or environmental information.

[0219] Step S304 (not shown in the figure), perform annotation processing on the device control instruction based on the obtained information to obtain target parameter information.

[0220] For the embodiments of the present application, step S304 may specifically include step S3041 (not shown in the figure), where

[0221] Step S3041, based on the obtained information and the device control instruction, and through a sequence tagger, obtain target parameter information.

[0222] Among them, the target parameter information includes any one of the following: the parameter information after the sequence tagger changes the parameter information in the device control instruction; the parameter information in the device control instruction.

[0223] For the embodiments of the present application, if the device control instruction meets the preset conditions, the target parameter information is the parameter information after the parameter information in the device control instruction is changed by the sequence tagger; if the device control instruction does not meet the preset conditions, the target parameter information is the parameter information in the device control instruction.

[0224] Among them, the preset conditions include at least one of the following:

[0225] The device control instruction does not contain a parameter value;

[0226] The parameter value contained in the device control instruction does not belong to the parameter value range determined by the obtained information.

[0227] Further, step S3041 may specifically include: step S30411 (not shown in the figure), step S30412 (not shown in the figure), and step S30413 (not shown in the figure), where

[0228] Step S30411: Perform sequence tagging processing on the device control instruction to obtain the parameter information in the device instruction.

[0229] Step S30412: Based on the device control instruction and the obtained information, determine whether to change the parameter information in the device control instruction.

[0230] Specifically, step S30412 may include: based on the parameter information in the device control instruction and the obtained information, through logistic regression processing, obtain a logistic regression result; based on the logistic regression result, determine whether to change the parameter information in the device control instruction.

[0231] Step S30413: If it is to be changed, based on the parameter information in the device control instruction and the obtained information, determine the changed target parameter information.

[0232] Specifically, determining the changed target parameter information based on the parameter information in the device control instruction and the acquired information in step S30413 may include: obtaining a linear regression result through linear regression processing based on the parameter information in the device control instruction and the acquired information; and determining the changed parameter information based on the linear regression result. In an embodiment of the present application, a prediction result is obtained by fitting a prediction function based on the parameter information in the device control instruction and the acquired information. The preset result may include at least one of whether to change the parameter information in the device control instruction and the changed parameter information. The prediction function may take various forms. Specifically, the fitting prediction function may be a linear function. When the fitting prediction function is a linear function, linear regression processing is performed to obtain a linear regression result. The fitting prediction function may also be an exponential function. When the fitting prediction function is an exponential function, logistic regression is performed to obtain a logistic regression result. Further, the fitting prediction function may also be a polynomial function. When the fitting prediction function is a polynomial function, similar linear regression processing is performed to obtain a similar linear regression result. In an embodiment of the present application, the prediction function may also include other functions, which are not limited herein.

[0233] For an embodiment of the present application, if it is necessary to change the parameter information in the device control instruction, the output result of the sequence tagger includes the changed parameter information, and may further include: an indication information corresponding to the modified parameter information, and the parameter information in the device control instruction;

[0234] For an embodiment of the present application, if not, it is determined that the target parameter information is the parameter information in the device control instruction. Further, if it is not necessary to change the parameter information in the device control instruction, the output result of the sequence tagger includes: an indication information corresponding to the unmodified parameter information, and the parameter information in the device control instruction.

[0235] Regarding the manner of determining whether to change the parameter information in the device control instruction and the changed parameter information through logistic regression and linear regression in steps S30411 - S30413, the specific process of the sequence tagger performing sequence tagging processing is introduced, such as Figure 11As shown, it is the structure of the encoder and decoder (encoder and decoder). x1, x2....xm are the device control instructions of the user. The encoding layer also adopts the encoding form combining word vector conversion and position encoding. After encoding, each word is represented as a vector with a fixed dimension, and then the LSTM layer is used for encoding to extract the features h1, h2.....hm of each word; y1, y2....yk are the tags corresponding to x1, x2....xm (using the BMO tagging method, B indicates the starting position of the parameter of the vocabulary, M indicates the middle position or the ending position of the parameter of the vocabulary, and O indicates that the vocabulary is not a parameter). y1, y2...yk are represented as the hidden state C through the LSTM layer. The vector d is calculated through the attention layer using C and h1, h2.....hm. After passing through the Multilayer Perceptron (MLP), the vector f is obtained. After passing through the multilayer perceptron, the label yk+1 (the parameter information in the device control instruction) at the next moment is output. At the same time, the vector f and the information representation vector corresponding to the obtained information (multimodal information representation vector) pass through logistic regression and linear regression respectively to obtain the logistic regression result and the linear regression result. Among them, Figure 11 EOS (English full name: End of Sentence) in

[0236] Among them, the result of logistic regression determines whether to change the parameter at the k+1 moment (that is, the output result is to change or not to change), and the result of linear regression determines the value after the change (that is, the filling value).

[0237] For the embodiments of the present application, the result of logistic regression determines whether to change the parameter, and the result of linear regression determines the value after the change, so that the network has the ability to rewrite the parameter.

[0238] For example, the device control instruction input by the user is "Set the air conditioner to 100 degrees". This instruction is classified into the air conditioner field by the domain classifier, and the intention classifier assigns the intention of "setting the air conditioner temperature". Since 100 degrees is a temperature that the air conditioner cannot set and the air conditioner cannot execute, after the parameter is marked by the sequence tagger model in the embodiments of the present application, through logistic regression and linear regression, 100 degrees will be rewritten as the upper limit temperature of 30 degrees of the air conditioner in the environmental database (or modified to the temperature of 26 degrees that the user likes in the user portrait database), and this parameter is passed to the air conditioner for execution, thereby increasing the indoor temperature and being more in line with the semantics of "Set the air conditioner to 100 degrees". Another example is the user's statement "Turn up the oven by 240 degrees". The device monitoring module monitors that the current working temperature of the oven is relatively high and the working time is relatively long, and transmits this information to the multimodal information representation vector. After the ST tagging model tags the parameter as 240 degrees, it will rewrite the parameter in combination with the multimodal information representation vector and output 200 degrees to the oven for execution.

[0239] When the MLP in the existing sequence labeling model obtains the label yk+1 at the next moment, it directly serves as the output result without passing through logistic regression and linear regression. This may cause the device to fail to accurately execute the device control instructions input by the user, or the execution result may pose a danger to the device or the user. Through the present application based on linear regression and logistic regression, unreasonable or unclear parameter information in the device instruction information can be modified into reasonable and clear parameter information, improving the safety of device operation, reducing the risk of device failure, enhancing the flexibility of controlling the device, and thus improving the user experience.

[0240] Further, based on the obtained information and the device control instruction, and through the sequence labeler, the target parameter information is obtained. It also includes: obtaining a plurality of training data; training the sequence labeler based on the obtained training data and through the target loss function.

[0241] Wherein, any training data includes the following information:

[0242] Device control instruction; sequence labeling result corresponding to the device control instruction; indication information on whether the parameters in the device control instruction are changed; changed parameter information; obtained information.

[0243] Further, before training the sequence labeler based on the obtained training data and through the target loss function, it also includes: determining the target loss function.

[0244] Wherein, determining the target loss function includes: determining the first loss function based on the sequence labeling result corresponding to the device control instruction in each training data and the predicted labeling result of the sequence labeler; determining the second loss function based on the indication information on whether the parameters in the device instruction in each training data are changed and the predicted indication information on whether to change by the sequence labeler; determining the third loss function based on the changed parameter information in each training data and the changed parameter information output by the sequence labeler; determining the target loss function based on the first loss function, the second loss function and the third loss function.

[0245] Specifically, the training data set used for training the sequence labeler can be (s i , y i , c i , v i , m i ). s i represents the text information corresponding to the input device control instruction, and y i represents the BMO label corresponding to this device control instruction (such as S iFor "set the air conditioner to 30 degrees", yi is "O O O O B M"), c i is 0 or 1 (0 means no parameter change is required, 1 means parameter modification is required), v i represents the filled value after the change (the target parameter information after the change), m i represents multi-modal information (the information obtained) (including the current sensor measurement value, the appropriate value, and the executable range of the device, etc.), and i is the index of each data in the training data set.

[0246] Furthermore, the loss function for its training is:

[0247]

[0248] Among them, M is the total number of words in the training data, N is the total number of input data statements, the first item in Loss represents the annotation error, y ij represents the BMO annotation of the j-th word in the i-th training data, is the BMO result of the j-th word in the i-th training data predicted by the model; the second item in Loss is the parameter correction error, c i represents whether the parameter needs to be modified, c i = 0 means the parameter does not need to be modified, c i = 1 means the parameter needs to be modified; the third item in Loss is the modified value output by the model and the squared difference between the label modified value v i , where α and β are coefficients.

[0249] Step S305 (not shown in the figure), based on at least one target device and / or target function and / or target parameter information, control at least one target device to perform corresponding operations.

[0250] For the embodiments of the present application, based on at least the target device and / or target function determined in step S302, and the target parameter information obtained from step S304 (steps S30411 - S30413), control at least one target device to perform corresponding operations.

[0251] The following introduces a specific example for Embodiment 2:

[0252] The text information corresponding to the device control instruction input by the user is "set the air conditioner to 100 degrees". This is directly classified into the air conditioner field by the domain classifier, and the intention classifier directly assigns this control instruction to the intention of "setting the air conditioner temperature". The sequence tagger tags "100" as a parameter. At the same time, among the device information included in the multi-modal information, the maximum temperature of the air conditioner is 32 degrees and the appropriate temperature is 29 degrees. The output c of the parameter rewriting network's logistic regression iGreater than 0.5, while the linear regression output v i If it is 32 degrees, the parameters are rewritten, and the output result of 32 degrees of the linear regression is passed to the air conditioner for execution.

[0253] Furthermore, for the second embodiment, a device control system is also introduced (taking the example that the user inputs the device control instruction by voice). By collecting the user's voice signal and environmental signals (such as indoor temperature, indoor air quality, etc.), the semantic understanding of the user is finally formed and a voice feedback is given, and the corresponding command is executed. Figure 12 In it, the system is divided into a voice processing module, an environmental monitoring module, a multimodal information processing module, a voice conversion module, a semantic understanding module, a dialogue management module, a voice synthesis module, and an execution module. After the audio acquisition device (microphone) collects the voice signal, the voice processing module performs noise reduction and outputs the voice signal after noise reduction processing; the sensor collects environmental information including temperature, humidity, etc.; the information output by the above voice processing module and environmental monitoring module is integrated into multimodal information by the multimodal information processing module. Figure 12 In it, only an example where the multimodal information includes environmental information is shown. The multimodal information in this application may also include device information and / or user information, which is not shown in Figure 12 In it; the voice signal output by the voice processing module will be converted into text information by the voice conversion module, and this text information and the multimodal information are jointly input to the semantic understanding module. The semantic understanding module outputs the domain, intention, and label of the statement to the dialogue management module and the execution module. The dialogue management module will generate a reply text, which is synthesized by the voice synthesis module and then replied, and the execution module performs the corresponding operation.

[0254] Specifically, the dialogue management module: The dialogue management module is responsible for generating a reply according to the results of the semantic understanding module (including the domain (target device), target function, and annotation result (target parameter information) to which the device control instruction belongs). This part can use manually designed replies or replies through trained models, which will not be elaborated here.

[0255] The voice synthesis module: The language synthesis module converts the results of the dialogue management module into audio output, which will not be elaborated here.

[0256] The execution module: The execution module is responsible for the hardware device that executes the user's device control instruction. The execution module is deployed in intelligent terminals (including smart home devices, mobile phones, etc.).

[0257] Among them, the system architecture diagram of the prior art is as Figure 7As shown, the voice signal is denoised by the voice processing module and then converted into text by the speech conversion module. After being processed by the semantic understanding module, the domain, intention, and label of the sentence are obtained and given to the dialogue management module and the execution module. The dialogue management module generates reply text, which is synthesized by the speech synthesis module and then replied; the execution module performs corresponding operations. Therefore, in the prior art, directly executing according to the parameter information in the device control instruction input by the user may pose a safety risk to the device or the user, or when the parameters are unclear, the device may not be able to accurately execute the corresponding operations. In this application, the target parameter information is adjusted according to at least one of the device information, user information, and environmental information, fully considering various factors that may affect the safe operation of the device. When the parameters in the instruction are unclear or unsafe, the corresponding operations can be executed according to the changed parameters, greatly improving the safety and flexibility when the user controls the device.

[0258] Embodiment III

[0259] This embodiment mainly introduces determining at least one target device and / or target function corresponding to the device control instruction in combination with the acquired information (multi-modal information), and determining whether to modify the parameter information in the device control instruction in combination with the acquired information (multi-modal information), outputting the modified target parameter information (if not modified, outputting the parameter information in the device control instruction), and executing the device control instruction input by the user based on at least one target device and / or target function and / or target parameter information (modified, unmodified), as specifically shown below:

[0260] Step S401 (not shown in the figure), obtain the device control instruction input by the user.

[0261] For the embodiments of this application, the device control instruction input by the user can be input by the user in text form, or can be input by the user in ways such as voice, keys, gestures, etc. There is no limitation in the embodiments of this application. The embodiments of this application are described by taking the user inputting the device control instruction in voice form as an example.

[0262] Step S402 (not shown in the figure), obtain at least one of the following information: user information; environmental information; device information.

[0263] For the embodiments of this application, a user portrait database and a user permission database can be preset, such as Figure 2aAs shown, the user profile database stores user profile data, including: gender, age, user group, nickname, voice recognition template, and face recognition template data, etc. The user group can be divided into four user groups, namely the owner user group, the child user group, the elderly user group, and the guest user group. Among them, the users in the user profile data of the owner user group, the child user group, and the elderly user group can be written during registration, and the users in the user profile data of the guest user group can be written during registration or during use; the user permission database records the categories of devices that the owner user group, the child user group, the elderly user group, and the guest user group can use and the function list that each device can use. For the child user group, intentions ABCDE may all be not allowed, and intention F is allowed to be executed; for the elderly user group, intentions A and B may not be allowed to be executed, and intentions C, D, E, and F are allowed to be executed; for the guest user group, intentions B, D, and E are not allowed to be executed, and intentions A, C, and F are allowed to be executed; for the owner user group, intentions A, B, C, D, and E, F are all allowed to be executed. This function list has default setting values and can also be manually set by the users in the owner user group.

[0264] For the embodiments of the present application, user information is obtained through voiceprint recognition and / or image recognition. In the embodiments of the present application, when a device control instruction input by the user through voice is obtained, the sound processing module determines at least one of the identity information, gender information, and age information of the user who inputs the device control instruction based on voiceprint recognition; if an image acquisition device is provided on some devices, the face image information of the user who inputs the device control instruction can be collected based on the image acquisition device, and the image processing module determines at least one of the identity information, gender information, age information, and user group information of the user who inputs the device control instruction based on face image detection technology. Specifically, when the corresponding sound signal and face image signal are collected, identity authentication is performed by comparing the face authentication information with the user face recognition template of each user in the user profile database to determine the user identity. When only the collected sound signal is available and the face image signal is not collected, identity authentication is performed by comparing the voiceprint authentication information with the user voiceprint recognition template of each user in the user profile database to determine the user identity (considering that in the smart home scenario, the camera is often installed on the computer and TV, and when the speaker is in the kitchen and bedroom, etc., the image signal may not exist).

[0265] When the authentication is passed (i.e., when the speaker feature has a high similarity with the voiceprint recognition template (or face recognition template) of a certain user in the existing user portrait database), the user portrait of this user in the user database is output, including gender, age, user group, etc.; if the identity authentication fails, it is indicated as a new user, a new user portrait data is established and written, the obtained gender data and age data are written, and the user group is guest. And the newly created user portrait data is output, including gender, age, user group, etc., and then the user permissions corresponding to the user group are queried in the user permission database according to the user group in the output user portrait data and output. The user permissions output and the user portrait data are integrated into multimodal information, and then based on the integrated multimodal information and the device control instruction input by the user, at least one target device corresponding to the device control instruction and / or the target function corresponding to the device control instruction are determined, such as in step S403.

[0266] For the embodiments of the present application, methods such as Markov random field and convolutional neural network are used to perform voiceprint recognition on the device control instruction input by voice to determine at least one of the identity information, gender information, and age information of the user who inputs the device control instruction. Taking the neural network method as an example, after training a voiceprint classification network with a large amount of data, the network is used to extract feature vectors from the user's voiceprint and save them as templates; during authentication, the cosine distance between the features of the voiceprint to be authenticated and each feature template in the database is compared, and if it exceeds the threshold, the authentication is considered successful, otherwise it fails; for voice recognition of age information and / or gender information, a convolutional neural network can also be used, which will not be elaborated in the embodiments of the present application.

[0267] For the embodiments of the present application, after obtaining the device control instruction input by the user by voice, the sound processing module can also first perform sound noise reduction processing on the device control instruction input by the user. In the embodiments of the present application, the sound noise reduction technologies can include: multi-microphone collaborative noise reduction technology, convolutional neural network noise reduction technology. This will not be elaborated in the embodiments of the present application.

[0268] Further, after step S402, that is, after obtaining the device control instruction and face image information input by the user in a voice manner, identity authentication is performed. If the authentication is passed, that is, it is determined that there is an existing user according to the identity authentication result, the user portrait information (including: age, gender, user group, etc.) is obtained from the pre-created user portrait database, and based on the user group information, the user permissions are obtained from the pre-created user permission database; if the verification fails, that is, it is determined that it is a new user according to the identity authentication result, the user portrait is obtained based on the device control instruction and face image information input by the user, and is stored in the user portrait database, that is, the new user portrait data is written; and the environmental information (current temperature, current air pressure, etc.) is obtained through the environmental monitoring module, and the appropriate environmental information (appropriate temperature, appropriate air pressure, etc.) is obtained from the preset environmental data. In addition, the device information can also be obtained, and according to the above information, multi-modal information is formed, such as Figure 13 as shown

[0269] In the embodiment of the present application, as Figure 14 shown, the multi-modal information includes: environmental information (including: current temperature, current air pressure, etc.), information obtained from the user portrait database (including: gender, user level (user group), age, etc.), information obtained from the environmental database (including: appropriate temperature, appropriate humidity, appropriate air pressure, etc.), and information obtained from the user permission database. For example, the user permission database records the categories of devices that can be used by the master user group, child user group, elderly user group, and guest user group, and the function list that each device can use. For the child user group, intentions ABCDE may not be allowed, and intention F is allowed to be executed; for the elderly user group, intentions A and B may not be allowed to be executed, and intentions C, D, E, and F are allowed to be executed; for the guest user group, intentions B, D, and E may not be allowed to be executed, and intentions A, C, and F are allowed to be executed; for the master user group, intentions A, B, C, D, and E are all allowed to be executed.

[0270] For the embodiment of the present application, step S401 can be executed before step S402, can also be executed after step S402, or can be executed simultaneously with step S402. It is not limited in the embodiment of the present application.

[0271] Step S403 (not shown in the figure), based on the obtained information and the device control instruction, determines at least one target device corresponding to the device control instruction and / or the target function corresponding to the device control instruction.

[0272] For the embodiments of the present application, the user group information to which the user belongs may be input based on the device control instruction to determine at least one target device corresponding to the device control instruction and / or the target function corresponding to the device control instruction; alternatively, the age information and / or gender information of the user may be input according to the device control instruction to determine at least one target device corresponding to the device control instruction and / or the target function corresponding to the device control instruction; or the user group information to which the user belongs may be input based on the device control instruction and the age information and / or gender information of the user may be input according to the device control instruction to determine at least one target device corresponding to the device control instruction and / or the target function corresponding to the device control instruction.

[0273] Of course, based on the acquired information and the device control instruction, and through the trained model, at least one target device corresponding to the device control instruction and / or the target function corresponding to the device control instruction may be determined. For example, based on the acquired information (multimodal information) and the device control instruction, and through the trained domain classifier (DC), at least one target device corresponding to the device control instruction and / or the target function corresponding to the device control instruction may be determined; based on the acquired information (multimodal information) and the device control instruction, and through the trained intent classifier (IC), the target function corresponding to the device control instruction may be determined.

[0274] Specifically, determining at least one target device corresponding to the device control instruction in step S403 includes: obtaining a domain classification result based on the acquired information and the device control instruction through a domain classifier; if the maximum element value in the domain classification result is not less than the first preset threshold, at least one target device corresponding to the device control instruction is determined based on the domain classification result.

[0275] For the embodiments of the present application, before obtaining a domain classification result based on the acquired information and the device control instruction through a domain classifier, the domain classifier may also be trained.

[0276] Specifically, the training sample is (s i , m i , d i ), where s i represents the input data statement text, m i represents multimodal information, including the gender, permissions, age, etc. of the user input by the device control instruction, and d i represents the label of this statement, that is, the domain to which it belongs (that is, which device it belongs to), i represents the index of a piece of training data in the training dataset, and d i is in one-hot encoding form, that is, when this statement belongs to the jth device, d ij is 1, and d ik(k≠j) is 0; if the statement is an overstepping authority statement (for example, the command statement of a 4-year-old user is "Bake a sweet potato for me", but the oven is not allowed for children to use), then all elements of d i are 0, and the loss function based on which the training is carried out is as follows:

[0277]

[0278] wherein, is the predicted output of the model. When the predicted output and d ij are exactly the same, loss is 0.

[0279] After training the domain classifier based on the above training method, based on the obtained information and device control instructions, and through the trained domain classifier, the domain classification result is obtained. Specifically, the input data is (s, m), where s is the text information corresponding to the device control instruction, and m is the multimodal information. Using the above-trained DC model, the predicted output (domain classification result), the largest element in is If (c is the first preset threshold, and 0.5 can be selected), then the classification result of this statement is the k-th device. If then it means that this statement belongs to the protected situation, and the device control instruction is refused to be executed (for example, the command statement of a 4-year-old user is "Bake a sweet potato for me", but the oven is not allowed for children to use. During training, all elements of the label of this statement are 0. Therefore, when the trained model makes a prediction, the predicted output of this statement is close to 0 and will also be less than the threshold c. Thus, DC will not give a classification and will refuse to execute.).

[0280] Specifically, such as Figure 2bAs shown, the input text can be a device control instruction in the text format input by the user, or a device control instruction obtained by converting the device control instruction in the voice format input by the user through voice-text conversion. After the text undergoes word encoding (word vector conversion and position encoding), an encoded vector is obtained. Then, through a convolutional neural network and a self-attention model, a text representation vector can be obtained. Among them, after the text undergoes word vector conversion, it becomes (w1, w2, w3...), where w1, w2, w3 represent the word vectors corresponding to each word in the sentence (device control instruction); the position encoding is a function (f(1), f(2), f(3)), which is a function of the position index of the word vector. Then, these two parts are added to obtain the encoded vector (w1 + f(1), w2 + f(2), w3 + f(3)..). Then, the obtained encoded vector passes through a convolutional neural network and a self-attention model to obtain a text representation vector. Then, after the multi-modal information representation vector is connected to the text representation vector, a joint vector is obtained (for example, vector (a1, a2, a3...) and vector (b1, b2, b3...), and after connection, it is (a1, a2, a3..., b1, b2, b3,...)). Then, the joint vector is input into a fully connected layer and the classification result is output (the execution probability of device A, the execution probability of device B, the execution probability of device C).

[0281] Furthermore, a method for converting multi-modal information (acquired information) into an information representation vector corresponding to the multi-modal information is introduced. Specifically, at least one of user information, environmental information, and device information is acquired. After that, it further includes: converting the discrete information in the acquired information into a continuous dense vector; and determining an information representation vector (multi-modal information representation vector) corresponding to the acquired information according to the converted continuous dense vector and the continuous information in the acquired information.

[0282] For the embodiments of the present application, the discrete information in the acquired information can be converted into a continuous dense vector through a transformation matrix. In the embodiments of the present application, through the transformation matrix, it is converted into a continuous dense vector; the converted continuous dense vector and the information in the acquired information that does not belong to discrete values are linked to obtain a joint vector; then, the joint vector is subjected to a preset process to obtain an information representation vector corresponding to the acquired information.

[0283] Specifically, such as Figure 15As shown, when encoding the obtained information (multi-modal information), since gender, permission, and favorite channels are discrete values, they need to be converted into continuous dense vectors through an encoding matrix. Age, favorite temperature, device information, current temperature, etc. can be directly input. The encoded multi-modal information is concatenated to obtain a joint vector, and after passing through a fully connected layer and a sigmoid activation function, an information representation vector (multi-modal information representation vector) corresponding to the obtained information is obtained. For example, the information corresponding to gender is processed through a gender encoding matrix to obtain a continuous dense vector corresponding to gender information; the device control permission information of the user is processed through a permission encoding matrix to obtain a continuous dense vector corresponding to permission information; the favorite channel is processed through an emotion encoding matrix to obtain a continuous dense vector corresponding to the favorite channel.

[0284] Based on the obtained information and the device control instruction, controlling at least one target device to perform corresponding operations, including:

[0285] Based on the information representation vector corresponding to the obtained information and the device control instruction, controlling at least one target device to perform an operation.

[0286] Therefore, the difference between the DC model and the existing DC model is that multi-modal information is added as an input, enabling the model to perform domain classification (determine at least one target device) by referring to user information and / or device information and / or environmental information. For example, when the oven temperature is too high and the command is "bake a cake for one hour", the domain classifier will not classify this statement into the oven domain but will reject the execution.

[0287] For the embodiments of the present application, when at least one target device is determined based on the obtained information and the device control instruction, the target function corresponding to the device control instruction can be determined based on the obtained information and the device control instruction and through an intent classifier; it can also be when at least one target device is determined based on the obtained information and the device control instruction, based on the obtained information and the device control instruction, and through the intent classifiers respectively corresponding to each target device in the at least one target device, the target functions respectively corresponding to the device control instructions are determined respectively. Specifically as follows:

[0288] Determining the target function corresponding to the device control instruction in step S403 includes: performing intent classification processing based on the obtained information and the device control instruction to determine the execution probability of each control function; if the execution probabilities of all control functions are less than a second preset threshold, then output an operation result of rejecting the execution of the device control instruction, otherwise determine the target function corresponding to the device control instruction based on the execution probabilities of all control functions.

[0289] Specifically, based on the acquired information and device control instructions, intent classification processing is performed through a model (intent classifier). In the embodiments of the present application, when multiple target devices are determined, intent classification can be performed through only one model, or through the models in each target device.

[0290] Further, since intent classification processing can be performed through a model (intent classifier), before performing intent classification processing based on the acquired information and device control instructions through the model (intent classifier), it further includes: training the model (intent classifier), and the specific method is as follows: training the model (intent classifier) through the following loss function:

[0291]

[0292] Where is the predicted output of the model, I i is the label of the i-th training data, I i is in one-hot encoding form, that is, when the statement belongs to the j-th function (target function), I ij is 1, I ik (k≠j) is 0. If the training statement is an over-authority statement (for example, the command statement of a 4-year-old user is "uninstall the XX APP of the TV set". Although the TV set is open to children, the function of uninstalling the app on the TV set is not allowed for children), then all elements of I i are 0.

[0293] Among them, when training the intent classifier, the training samples still contain the acquired information (multi-modal information), and the multi-modal information can be initialized with the multi-modal information in the pre-trained domain classifier to accelerate the training speed.

[0294] For the embodiments of the present application, after training the intent classifier through the above method, based on the device control instructions input by the user and the acquired information, and through the trained intent classifier, the target function corresponding to the device control instructions can be determined. The specific method for determining the target function corresponding to the device control instructions is as follows:

[0295] The input of the intent classifier is (s, m), where s is the text information corresponding to the device control instructions, and m is the multi-modal information. First, the domain (device) to which it belongs is obtained through the DC model, and then the predicted output The maximum element in is If (c is a set threshold, and 0.5 can be selected), then the classification result of this device control instruction is the k-th function (target function). If It indicates that the device control instruction belongs to the protected situation and the execution is refused. At the same time, if the owner sets the permissions of children, the elderly, guests, etc. in the user permission database and the k-th function is just in the shielding list, then the execution is refused. For example, Figure 4 As shown, the device control instructions input by the user include: intention A, intention D or intention F. If the owner does not set user permissions, the intention classifier directly outputs intention A, intention D or intention F; if the owner sets user permissions (intention A, intention C, intention F are allowed to operate, intention B, intention D, intention E are not allowed to operate), then the output intention classifier directly outputs intention A or intention F, and the execution of intention D is refused.

[0296] For example, when a child says "Delete the channel list of the TV", the device control instruction will be assigned to the TV field by the domain classifier. However, the function of deleting the channel list of the TV is not open to children. Since all elements of the label of this sentence are 0 during training, when the trained intention classifier makes a prediction, the prediction output of the device control instruction is close to 0 and is also less than the threshold c. Therefore, the intention classifier will not give an intention classification and will refuse to execute the target function.

[0297] Furthermore, based on the obtained information and the device control instruction, and through the intention classifier, the multi-modal information input to the intention classifier when determining the target function is its corresponding representation vector. In the embodiments of the present application, the manner of obtaining the information representation vector corresponding to the obtained information from the obtained information is as described in the above embodiments and will not be elaborated here.

[0298] It should be noted that: if the target device corresponding to the device control instruction is not determined based on the obtained information and the device control instruction (directly output an instruction to refuse to execute the corresponding operation), then the target function corresponding to the device control instruction may not be determined, and the device control instruction may not be marked.

[0299] Step S404 (not shown in the figure), based on the obtained information, perform a marking process on the device control instruction to obtain target parameter information.

[0300] For the embodiments of the present application, step S404 may specifically include step S4041 (not shown in the figure), where

[0301] Step S4041, based on the obtained information and the device control instruction, and through a sequence tagger, obtain target parameter information.

[0302] Among them, the target parameter information includes any one of the following: the parameter information after the sequence tagger changes the parameter information in the device control instruction; the parameter information in the device control instruction.

[0303] For the embodiments of the present application, if the device control instruction meets the preset conditions, the target parameter information is the parameter information after the parameter information in the device control instruction is changed by the sequence tagger; if the device control instruction does not meet the preset conditions, the target parameter information is the parameter information in the device control instruction.

[0304] Among them, the preset conditions include at least one of the following:

[0305] The device control instruction does not contain parameter values;

[0306] The parameter values included in the device control instruction do not belong to the parameter values within the parameter value range determined by the acquired information.

[0307] Further, step S4041 may specifically include: step S40411 (not shown in the figure), step 430412 (not shown in the figure), step 430413 (not shown in the figure), step S40414 (not shown in the figure), and step S40415 (not shown in the figure), where

[0308] Step S40411: Perform sequence tagging processing on the device control instruction to obtain the parameter information in the device instruction.

[0309] Step S40412: Based on the device control instruction and the acquired information, determine whether to change the parameter information in the device control instruction.

[0310] Specifically, step S40412 may include: Based on the parameter information in the device control instruction and the acquired information, through logistic regression processing, obtain a logistic regression result; based on the logistic regression result, determine whether to change the parameter information in the device control instruction.

[0311] Step S40413: If it is to be changed, based on the parameter information in the device control instruction and the acquired information, determine the changed target parameter information.

[0312] Specifically, determining the changed target parameter information based on the parameter information in the device control instruction and the acquired information in step S40413 may include: Based on the parameter information in the device control instruction and the acquired information, through linear regression processing, obtain a linear regression result; based on the linear regression result, determine the changed parameter information.

[0313] For the embodiments of the present application, if it is necessary to change the parameter information in the device control instruction, the output result of the sequence tagger includes the changed parameter information, and may also include: the indication information corresponding to the modified parameter information, the parameter information in the device control instruction.

[0314] For the embodiments of the present application, if there is no need to change the parameter information in the device control instruction, the output result of the sequence tagger includes: indication information corresponding to the unmodified parameter information, and the parameter information in the device control instruction.

[0315] Regarding steps S40411 - S40413, the specific process of the sequence tagger for sequence tagging processing is introduced. For example, Figure 11 As shown, it is the encoder and decoder structure (encoder and decoder). x1, x2....xm are the device control instructions of the user. The encoding layer also adopts the encoding form combining word vector conversion and position encoding. After encoding, each word is represented as a vector with a fixed dimension, and then LSTM is used for encoding to extract the features h1, h2.....hm of each word; y1, y2....yk are the tags corresponding to x1, x2....xm (using the BMO tagging method, B indicates that the word is the starting position of the parameter, M indicates that the word is the middle position or the ending position of the parameter, and O indicates that the word is not a parameter). y1, y2...yk are represented as the hidden state C through the LSTM layer. The vector d is calculated through the attention layer using C and h1, h2.....hm. After passing through the Multilayer Perceptron (MLP), the vector f is obtained. After passing through the MLP, the tag yk + 1 (parameter information in the device control instruction) at the next moment is output. At the same time, the vector f and the multi-modal information representation vector respectively pass through logistic regression and linear regression to obtain the logistic regression result and the linear regression result.

[0316] Among them, the result of the logistic regression determines whether to change the parameter at the k + 1 moment, and the result of the linear regression determines the changed value.

[0317] For the embodiments of the present application, the result of the logistic regression determines whether to change the parameter, and the result of the linear regression determines the changed value, so that the network has the ability to rewrite the parameter.

[0318] For example, the device control instruction input by the user is "set the air conditioner to 100 degrees". This instruction is classified into the air conditioner field by the field classifier, the intention classifier assigns the intention of "setting the air conditioner temperature", and the ST model marks "100 degrees" as the parameter to be passed to the air conditioner. Obviously, the air conditioner cannot set this temperature and thus cannot execute. However, after the parameter is marked by the sequence tagger model in the embodiment of the present application, through logistic regression and linear regression, 100 degrees will be rewritten as the upper limit temperature of the air conditioner in the environmental database (or modified to the temperature liked by this user in the user portrait database), and this parameter is passed to the air conditioner for execution, thereby increasing the indoor temperature and being more in line with the semantics of "set the air conditioner to 100 degrees". Another example is the user statement "turn up the oven by 240 degrees". It is monitored that the current working temperature of the oven is relatively high and the working time is relatively long, and this information is passed into the multi-modal information representation vector. After the ST annotation model marks the parameter as 240 degrees, it will rewrite this parameter in combination with the multi-modal information representation vector and output 200 degrees for the oven to execute.

[0319] Further, since the target parameter information is obtained based on the acquired information and the device control instruction and through the sequence tagger, it further includes: obtaining a plurality of training data; training the sequence tagger based on the acquired training data and through the target loss function.

[0320] Wherein, any training data includes the following information:

[0321] Device control instruction; sequence annotation result corresponding to the device control instruction; indication information on whether the parameter in the device control instruction is changed; changed parameter information; acquired information.

[0322] Further, before training the sequence tagger based on the acquired training data and through the target loss function, it further includes: determining the target loss function.

[0323] Wherein, determining the target loss function includes: determining the first loss function based on the sequence annotation result corresponding to the device control instruction in each training data and the predicted annotation result of the sequence tagger; determining the second loss function based on the indication information on whether the parameter in the device instruction in each training data is changed and the indication information on whether it is changed predicted by the sequence tagger; determining the third loss function based on the changed parameter information in each training data and the changed parameter information output by the sequence tagger; determining the target loss function based on the first loss function, the second loss function and the third loss function.

[0324] Specifically, the training data set used for training the sequence tagger can be (s i , y i , c i , v i , m i),s i Indicates the text information corresponding to the input device control instruction, y i Indicates the BMO label corresponding to this device control instruction (for example, S i is "Set the air conditioner to 30 degrees", and yi is "O O O O B M"), c i is 0 or 1 (0 indicates that no parameter needs to be changed, and 1 indicates that the parameter needs to be modified), v i Indicates the filled value after the change (the parameter information after the change), m i Indicates the multimodal information (the information obtained) (including the current sensor measurement value, the appropriate value, and the executable range of the device, etc.), and i is the index of each data in the training dataset.

[0325] Furthermore, the loss function for its training is:

[0326]

[0327] Among them, the first item in Loss represents the annotation error, y ij Indicates the BMO annotation of the j-th word in the i-th training data, is the BMO result of the j-th word in the i-th training data predicted by the model; the second item in Loss is the parameter correction error, c i Indicates whether the parameter needs to be modified, c i = 0 indicates that the parameter does not need to be modified, c i = 1 indicates that the parameter needs to be modified; the third item in Loss is the modified value output by the model and the squared difference between the label modified value v i

[0328] Step S405a (not shown in the figure), based on at least one target device and / or target function and / or target parameter information, control at least one target device to perform corresponding operations.

[0329] For the embodiments of the present application, if step S403 determines at least one target device and / or the target function corresponding to the device control instruction, and the target parameter information output by step S404, control at least one device to perform corresponding operations.

[0330] Step S405b (not shown in the figure), based on the obtained information and the device control instruction, output the operation result of rejecting the execution of the device control instruction.

[0331] For the embodiments of the present application, when it is determined through step S403 that the device control instruction input by the user cannot be executed, do not execute the device control instruction input by the user, and an instruction for rejecting the execution of the corresponding operation can be output.

[0332] ​Specifically, the operation result of rejecting the execution of the device control instruction is output, including: when it is determined according to the acquired information that at least one of the following conditions is met, the operation result of rejecting the execution of the device control instruction is output:

[0333] The user does not have the control authority for at least one target device; the user does not have the control authority for the target function corresponding to the device control instruction; at least one target device does not have the execution condition corresponding to the device control instruction; the working environment of at least one target device does not have the execution condition corresponding to the device control instruction.

[0334] For the above conditions, the examples of the operation result of rejecting the execution of the device control instruction are shown in detail in Embodiment 1 and will not be elaborated here.

[0335] Furthermore, for Embodiment 3, a device control system is also introduced (taking the example that the user inputs the device control instruction by voice). By collecting the user's voice signal, image signal and environmental signal (such as indoor air quality, etc.), the semantic understanding of the user is finally formed and voice feedback is given, and the corresponding command is executed. Figure 16 In this system, it is divided into a voice processing module, an image processing module, an environmental monitoring module, a multimodal information processing module, a voice conversion module, a semantic understanding module, a dialogue management module, a voice synthesis module, and an execution module. The main improvements in the embodiments of the present application lie in the multimodal information processing module and the semantic understanding module. Among them, after the audio acquisition device (microphone) collects the voice signal, the voice processing module performs noise reduction and identity recognition, and outputs the noise-reduced voice signal and identity authentication information; after the image acquisition device (camera) collects the face image information, the image processing module performs face extraction and face recognition, and outputs the identity authentication information; the sensor collects environmental information including temperature and humidity, etc.; the identity authentication information, environmental information, etc. output by the above image processing module, voice processing module, and environmental monitoring module will be integrated into multimodal information by the multimodal information processing module; the voice signal output by the voice processing module will be converted into text by the voice conversion module; this text and the multimodal information are jointly input to the semantic understanding module, and the semantic understanding module outputs the domain, intention, and label of this statement to the dialogue management module and the execution module; the dialogue management module will generate reply text, which is synthesized by the voice synthesis module and then replied; it is executed by the execution module.

[0336] Based on the above device control method, there can be the following multiple implementation manners for the hardware device in the embodiments of the present application:

[0337] A. Monolithic type: That is, relying on intelligent devices such as smart speakers and smart TVs as the hardware. The image processing module, voice processing module, voice conversion module, semantic understanding module, dialogue understanding module, voice synthesis module, etc. are all implemented on this intelligent hardware, and the user needs to issue instructions to this intelligent hardware.

[0338] B, Distributed: That is, all intelligent devices separately store their own IC models and the common DC model. According to whether the intelligent device is equipped with a microphone and a camera, it separately stores an image processing module, a sound processing module, and a voice conversion module. Devices can communicate with each other, and users can issue instructions to any device.

[0339] C, Communication-based: The image processing module, the sound processing module, the voice conversion module, the semantic understanding module, the dialogue understanding module, the voice synthesis module, etc. are all stored in a remote server (i.e., the cloud). After the user's smart home device acquires sound and images through an audio acquisition device (microphone) and an image acquisition device (camera), the server end understands them and returns the results.

[0340] The process of the TTS module is introduced in detail below.

[0341] The TTS module in the embodiments of the present application can generate emotional voices for the replied text. For example, voices with different intonations, different speech rates, and / or different volumes can be obtained. This processing process can be called Emotional TTS.

[0342] The embodiments of the present application can use the obtained multi-modal information to perform Emotional TTS. Further, the user information in the multi-modal information (such as user age, gender, identity information, etc.) can be used to perform Emotional TTS. In addition, the embodiments of the present application also propose that the user information can include not only user portraits (such as age, gender, etc.), user permissions (such as the device control permissions of the user), but also the emotional information of the user. The following specifically explains how to obtain the emotional information of the user:

[0343] Perform emotional recognition processing on the processing results of the sound processing module and the processing results of the image processing module to obtain the user emotional information corresponding to the input user. The obtained user emotional information can be used as the user information in the multi-modal. When performing multi-modal fusion processing, the user emotional information can also be fused, such as Figure 20aAs shown in the figure, specifically, the user emotion information is obtained after performing emotion recognition processing on the processing results of the sound processing module and the image processing module; for identity authentication, the face authentication information output by the image processing module is compared with the user face recognition templates of each user in the user portrait database to determine the user identity; if the authentication is passed, that is, if it is determined as an existing user according to the identity authentication result, then the user portrait of this user in the user portrait database is obtained and output, including gender, age, user group, etc.; if the identity authentication fails, that is, if it is determined as a new user according to the identity authentication result, then a new user portrait data is created and written in the user portrait database, the gender and age data obtained from the sound processing module and the image processing module are written, and the newly created user portrait data, including gender, age, user group, etc., is output; the user permissions of the corresponding user group are queried in the permission database according to the user group in the output user portrait data and output; the environmental information (current temperature, current air pressure, etc.) is obtained through the environmental monitoring module, and the suitable environmental information (suitable temperature, suitable air pressure, etc.) is obtained from the preset environmental data; the obtained environmental information, user emotion information, user permission information, and user portrait data are integrated into multimodal information, and the multimodal information is output to the multimodal NLU module.

[0344] In the embodiment of the present application, the user information (such as user emotion information, user identity information, user age information, user gender information, etc.) in the multimodal information can be used to perform Emotional TTS, so as to output voices with different emotions for different users, or output voices with different emotions for different states of the same user.

[0345] The embodiment of the present application proposes that the neural network of Emotional TTS can be pre-trained, and then the neural network is used to perform online Emotional TTS processing to obtain voices with different emotions for different users, or voices with different emotions for different states of the same user.

[0346] As Figure 20b shown, it is a schematic diagram of the training process of the neural network of Emotional TTS. First, the combination of all text samples is used as the input of network training, and the combinations of each emotional voice sample corresponding to the text samples are used as the output of network training to train the initial neural network of Emotional TTS. This initial neural network can obtain a relatively neutral emotional representation, and the emotional information is not rich enough. Then, the text samples and emotional encodings are used as the input of network training, and the emotional voice samples corresponding to the text samples are used as the output of network training to train the above neural network again, so as to obtain a neural network with better performance.

[0347] More specifically, the database stores text samples and corresponding emotional speech samples, and also stores one-hot encodings of emotional categories. When training the neural network, first preprocess the text samples to extract features of the text samples, such as full label features, etc., and extract corresponding audio features (which can also be called acoustic features) from the emotional speech samples. The one-hot encodings of emotional categories in the database can be processed through an encoding matrix to obtain embedded emotional encodings, that is, the "emotional encodings" in the figure. Using the emotional encodings and the features of the text samples, input features can be obtained. For example, the input features can be directly concatenated. According to the obtained input features and the acoustic features obtained by preprocessing, train a Bi-directional Long Short-Term Memory (Bi-LSTM) network so that the output acoustic features are close to the acoustic features corresponding to the emotional speech samples.

[0348] As Figure 20c shown, it is a schematic diagram of the online processing process of the neural network of Emotional TTS. The emotional category of the user's expected response (which can be called the expected emotional category, that is, the emotional category corresponding to the response speech of the device that the user expects to receive) can be determined from the obtained multimodal information. Among them, the expected emotional category can be related to user information, such as related to user age, gender, user emotional information, etc. The user's expected emotional category can be obtained from the multimodal information containing user information. After the device's DM module generates the response text (that is, the response text, which can also be simply referred to as text), extract the features of the text. According to the features of the text and the emotional encoding corresponding to the user's expected emotional category, input features can be obtained (for example, the features of the text and the emotional encoding corresponding to the user's expected emotional category can be directly concatenated to obtain the input features). Input the input features into the trained Bi-LSTM network. The Bi-LSTM network outputs the acoustic features corresponding to the response text, and then the vocoder generates the speech corresponding to the response text according to the acoustic features (corresponding Figure 20c to the generated speech in

[0349] The above TTS module of the embodiment of the present application can generate emotional speech for the response text. For example, speech with different intonations, different speech rates, and / or different volumes can be obtained, thereby improving the user experience.

[0350] The above are some specific implementation manners of the device control method provided by the embodiment of the present application. Based on this, the embodiment of the present application also provides a device control device. Next, the device control device provided by the embodiment of the present application will be introduced from the perspective of functional modularization in combination with the accompanying drawings.

[0351] The embodiment of the present application provides a device control device, as Figure 17As shown, the device 1700 may include: a first acquisition module 1701 and a control module 1702, where,

[0352] The first acquisition module 1701 is configured to acquire a device control instruction input by a user and at least one of the following pieces of information: user information; environmental information; device information.

[0353] The control module 1702 is configured to control at least one target device to perform corresponding operations based on the information acquired by the first acquisition module 1701 and the device control instruction.

[0354] In another possible implementation manner of the embodiments of the present application, the user information includes: user portrait information and / or user's device control permission information; and / or, the device information includes: the working state information of the device; and / or, the environmental information includes: the working environment information of the device.

[0355] In another possible implementation manner of the embodiments of the present application, the control module 1702 is specifically configured to output an operation result of refusing to execute the device control instruction based on the acquired information and the device control instruction.

[0356] In another possible implementation manner of the embodiments of the present application, the control module 1702 is specifically configured to output an operation result of refusing to execute the device control instruction when it is determined according to the acquired information that at least one of the following is satisfied:

[0357] The user does not have the control permission for at least one target device; the user does not have the control permission for the target function corresponding to the device control instruction; at least one target device does not meet the execution conditions corresponding to the device control instruction; the working environment of at least one target device does not meet the execution conditions corresponding to the device control instruction.

[0358] In another possible implementation manner of the embodiments of the present application, the control module includes: a first determination unit and a control unit, where,

[0359] The first determination unit is configured to determine at least one target device corresponding to the device control instruction and / or the target function corresponding to the device control instruction based on the acquired information and the device control instruction;

[0360] The control unit is specifically configured to control at least one target device to perform corresponding operations based on at least one target device and / or target function determined by the first determination unit.

[0361] In another possible implementation manner of the embodiment of the present application, the first determination unit is specifically configured to perform domain classification processing based on the acquired information and the device control instruction to obtain the execution probability of each device. When the execution probability of each device is less than the first preset threshold, an operation result of rejecting the execution of the device control instruction is output; otherwise, at least one target device corresponding to the device control instruction is determined based on the execution probability of each device.

[0362] In another possible implementation manner of the embodiment of the present application, the first determination unit is specifically configured to perform intention classification processing based on the acquired information and the device control instruction to determine the execution probability of each control function. When the execution probability of each control function is less than the second preset threshold, an operation result of rejecting the execution of the device control instruction is output; otherwise, the target function corresponding to the device control instruction is determined based on the execution probability of each control function.

[0363] In another possible implementation manner of the embodiment of the present application, the control module 1702 is specifically configured to control at least one target device to perform corresponding operations according to the target parameter information based on the acquired information.

[0364] Wherein, the target parameter is the parameter information obtained by changing the parameter information in the device control instruction.

[0365] In another possible implementation manner of the embodiment of the present application, the control module 1702 is further specifically configured to control at least one target device to perform corresponding operations according to the target parameter information when at least one of the following conditions is satisfied:

[0366] The device control instruction does not include a parameter value;

[0367] The parameter value included in the device control instruction does not belong to the parameter value range determined by the acquired information.

[0368] In another possible implementation manner of the embodiment of the present application, the control module 1702 includes: a sequence annotation processing unit, a second determination unit, and a third determination unit, wherein

[0369] The sequence annotation processing unit is configured to perform sequence annotation processing on the device control instruction to obtain the parameter information in the device control instruction;

[0370] The second determination unit is configured to determine whether to change the parameter information in the device control instruction based on the parameter information in the device control instruction and the acquired information;

[0371] The third determination unit is configured to, when the second determination unit determines to change the parameter information in the device control instruction, determine the changed target parameter information based on the parameter information in the device control instruction and the acquired information.

[0372] For the embodiments of the present application, the first determination unit, the second determination unit, and the third determination unit may all be the same unit, may all be different units, or any two of them may be the same unit. There is no limitation in the embodiments of the present application.

[0373] In another possible implementation manner of the embodiments of the present application, the second determination unit is specifically configured to, based on the parameter information in the device control instruction and the acquired information, perform logistic regression processing to obtain a logistic regression result, and determine whether to change the parameter information in the device control instruction based on the logistic regression result; and / or,

[0374] The third determination unit is specifically configured to, based on the parameter information in the device control instruction and the acquired information, perform linear regression processing to obtain a linear regression result, and determine the changed parameter information based on the linear regression result.

[0375] In another possible implementation manner of the embodiments of the present application, the device 1700 further includes: a second acquisition module and a training module, where,

[0376] The second acquisition module is configured to acquire a plurality of training data;

[0377] For the embodiments of the present application, the first acquisition module and the second acquisition module may be the same acquisition module or different acquisition modules. There is no limitation in the embodiments of the present application.

[0378] The training module is configured to, based on the training data acquired by the second acquisition module, and through a target loss function, train a processing model for changing the parameter information in the device control instruction.

[0379] Wherein, any one of the training data includes the following information:

[0380] Device control instruction; parameter information in the device control instruction; indication information on whether the parameter in the device control instruction is changed; changed parameter information; user information; environment information; device information.

[0381] In another possible implementation manner of the embodiments of the present application, the device 1700 further includes: a first determination module;

[0382] The first determination module is configured to determine a target loss function;

[0383] Wherein, the first determination module includes: a fourth determination unit, a fifth determination unit, a sixth determination unit, and a seventh determination unit, where,

[0384] The fourth determination unit is configured to determine a first loss function based on the parameter information in the device control instruction in each training data and the parameter information in the device control instruction predicted by the model;

[0385] A fifth determination unit, configured to determine a second loss function based on indication information on whether parameters in device instructions in each piece of training data are changed, and indication information on whether changes are predicted by the model;

[0386] A sixth determination unit, configured to determine a third loss function based on the changed parameter information in each piece of training data and the changed parameter information predicted by the model;

[0387] A seventh determination unit, configured to determine a target loss function based on the first loss function determined by the fourth determination unit, the second loss function determined by the fifth determination unit, and the third loss function determined by the sixth determination unit.

[0388] For the embodiments of the present application, the fourth determination unit, the fifth determination unit, the sixth determination unit, and the seventh determination unit may all be the same determination unit, may all be different determination units, or any two of them may be the same determination unit, or any three of them may be the same determination unit, etc. There is no limitation in the embodiments of the present application.

[0389] Another possible implementation manner of the embodiments of the present application is that the apparatus 1700 further includes: a conversion module and a second determination module, where

[0390] The conversion module is configured to convert discrete information in the acquired information into a continuous dense vector;

[0391] The second determination module is configured to determine an information representation vector corresponding to the acquired information according to the continuous dense vector converted by the conversion module and the continuous information in the acquired information;

[0392] For the embodiments of the present application, the first determination module and the second determination module may be the same determination module or different determination modules. There is no limitation in the embodiments of the present application.

[0393] The control module 1702 is specifically configured to control at least one target device to perform an operation based on the information representation vector corresponding to the acquired information determined by the second determination module and the device control instruction.

[0394] An embodiment of the present application provides a device control device. By obtaining at least one of user information, environmental information, and device information, as well as a device control instruction input by a user, it can control at least one target device to perform corresponding operations based on the obtained information and the device control instruction. As can be seen from the above, compared with the prior art method of only controlling a device according to a control instruction input by a user, when the embodiment of the present application controls a device, in addition to considering the control instruction input by the user, it also takes into account at least one factor that may affect the operation of the device, such as user information, device information, and environmental information. Therefore, the device operation can be controlled more safely and flexibly. For example, obtain device control instructions in the form of voice, text, keys, gestures, etc. input by the user, and consider at least one of user information, device information, and environmental information, and directly control the air conditioning device to turn on, turn off, or adjust the temperature, etc., so as to realize safe and convenient control of the intelligent device to perform corresponding operations.

[0395] The device control device provided by the embodiment of the present application is applicable to the above method embodiment, and will not be elaborated here.

[0396] The device control device provided by the embodiment of the present application is introduced from the perspective of functional modularization above. Next, the electronic device provided by the embodiment of the present application will be introduced from the perspective of hardware implementation, and the computing system of the electronic device will be introduced at the same time.

[0397] An embodiment of the present application provides an electronic device, which is applicable to the above method embodiment, as Figure 18 shown, including: a processor 1801; and a memory 1802, configured to store machine-readable instructions, and when the above instructions are executed by the above processor 1801, the above processor 1801 executes the above device control method.

[0398] Figure 19 Schematically shows a block diagram of a computing system of an electronic device that can be used to implement the present application according to an embodiment of the present application. As Figure 19 shown, the computing system 1900 includes a processor 1910, a computer-readable storage medium 1920, an output interface 1930, and an input interface 1940. The computing system 1900 can execute the method described above with reference to Figure 1 to control at least one target device to perform corresponding operations based on a device control instruction input by a user. Specifically, the processor 1910 may include, for example, a general microprocessor, an instruction set processor, and / or a related chipset, and / or a dedicated microprocessor (for example, an application specific integrated circuit (ASIC)), etc. The processor 1910 may also include on-board memory for caching purposes. The processor 1910 may be a single processing unit or multiple processing units for performing different actions of the method flow described with reference to Figure 1 above.

[0399] A computer-readable storage medium 1920 can be, for example, any medium capable of containing, storing, transmitting, propagating, or transporting instructions. For example, the readable storage medium can include, but is not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, components, or propagation media. Specific examples of the readable storage medium include: magnetic storage devices such as magnetic tapes or hard disk drives (HDDs); optical storage devices such as compact discs (CD-ROMs); memories such as random access memories (RAMs) or flash memories; and / or wired / wireless communication links.

[0400] The computer-readable storage medium 1920 can include a computer program 1921, and the computer program 1921 can include code / computer-executable instructions that, when executed by the processor 1910, cause the processor 1910 to execute, for example, the method processes and any variations thereof described above in conjunction with Figure 1 The computer program 1921 can be configured to have computer program code that includes, for example, computer program modules. For example, in an exemplary embodiment, the code in the computer program 1921 can include one or more program modules, such as module 1921A, module 1921B,.... It should be noted that the way of dividing the modules and the number of modules are not fixed, and those skilled in the art can use appropriate program modules or combinations of program modules according to the actual situation. When these combinations of program modules are executed by the processor 1910, the processor 1910 can execute, for example, the method processes and any variations thereof described above in conjunction with Figure 1 The computer program 1921 can be configured to have computer program code that includes, for example, computer program modules. For example, in an exemplary embodiment, the code in the computer program 1921 can include one or more program modules, such as module 1921A, module 1921B,.... It should be noted that the way of dividing the modules and the number of modules are not fixed, and those skilled in the art can use appropriate program modules or combinations of program modules according to the actual situation. When these combinations of program modules are executed by the processor 1910, the processor 1910 can execute, for example, the method processes and any variations thereof described above in conjunction with

[0401] According to an embodiment of the present disclosure, the processor 1910 can use the output interface 1930 and the input interface 1940 to execute the method processes and any variations thereof described above in conjunction with Figure 1 The computer program 1921 can be configured to have computer program code that includes, for example, computer program modules. For example, in an exemplary embodiment, the code in the computer program 1921 can include one or more program modules, such as module 1921A, module 1921B,.... It should be noted that the way of dividing the modules and the number of modules are not fixed, and those skilled in the art can use appropriate program modules or combinations of program modules according to the actual situation. When these combinations of program modules are executed by the processor 1910, the processor 1910 can execute, for example, the method processes and any variations thereof described above in conjunction with

[0402] An embodiment of the present application provides an electronic device. By obtaining at least one piece of information from user information, environmental information, and device information, as well as a device control instruction input by the user, it can control at least one target device to perform corresponding operations based on the obtained information and the device control instruction. As can be seen from the above, compared with the prior art method of only controlling the device according to the control instruction input by the user, when the embodiment of the present application controls the device, in addition to considering the control instruction input by the user, it also takes into account at least one factor that may affect the device operation, such as user information, device information, environmental information, etc., so that the device operation can be controlled more safely and flexibly. For example, by obtaining device control instructions in the form of voice, text, key presses, gestures, etc. input by the user, and taking into account at least one of the user information, device information, and environmental information, directly controlling operations such as turning on, turning off, or adjusting the temperature of the air conditioning device, so that the intelligent device can be safely and conveniently controlled to perform corresponding operations.

[0403] The electronic device and the computing system of the electronic device provided in the embodiments of the present application are applicable to the above method embodiments, and will not be elaborated here.

[0404] For the embodiments of the present application, the explanations of the same or similar terms in each embodiment can be used for reference to each other. For example, the manner of determining whether the parameter information in the user equipment instruction is modified and / or the modified parameter information in Embodiment 3 can refer to "Based on the parameter information in the device control instruction and the acquired information, a prediction result is obtained through a fitting prediction function. The preset result may include at least one of whether to change the parameter information in the device control instruction and the changed parameter information. The prediction function can take various forms. Specifically, the fitting prediction function can be a linear function. When the fitting prediction function is a linear function, linear regression processing is performed to obtain a linear regression result; the fitting prediction function can also be an exponential function. When the fitting prediction function is an exponential function, logistic regression is performed to obtain a logistic regression result; further, the fitting prediction function can also be a polynomial function. When the fitting prediction function is a polynomial function, similar linear regression processing is performed to obtain a similar linear regression result. In the embodiments of the present application, the prediction function can also include other functions, which are not limited herein." in Embodiment 2.

[0405] It should be understood that although the steps in the flowchart of the accompanying drawings are shown in sequence according to the indication of the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise clearly stated in this article, the execution of these steps has no strict order limitation, and they can be executed in other orders. Moreover, at least a part of the steps in the flowchart of the accompanying drawings may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily executed at the same moment, but can be executed at different moments, and their execution order is not necessarily sequential, but can be executed alternately or alternately with at least a part of other steps or sub-steps or stages of other steps.

[0406] The above are only some embodiments of the present invention. It should be pointed out that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.

Claims

1. A method executed by an electronic device, characterized in that, comprising: obtaining a device control instruction of a user; obtaining at least one of the following information: user information; environmental information; device information; performing domain classification processing based on the obtained information and the device control instruction to obtain the execution probability of each device; if the execution probability of at least one device is not less than a first preset threshold, performing intent classification processing on the at least one device based on the obtained information and the device control instruction to determine the execution probability of the control function; in the case where the execution probability of the control function is less than a second preset threshold, outputting an operation result of rejecting to execute the device control instruction, otherwise determining a target function corresponding to the device control instruction according to the execution probability of the control function, and controlling at least one target device to perform corresponding operations according to the target function.

2. The method according to claim 1, characterized in that, controlling at least one target device to perform corresponding operations according to the target function includes: determining the operations of at least one target device based on the obtained information and the device control instruction; controlling at least one target device to perform corresponding operations according to target parameter information based on a processing model and the obtained information; wherein the target parameter information is parameter information obtained by changing the parameter information in the device control instruction; the processing model is obtained by the following method: obtaining a plurality of training data; determining a target loss function according to the parameter information in the training data and the predicted parameter information of the processing model; training a processing model for changing the parameter information in the device control instruction based on the obtained training data and through the target loss function.

3. The method according to claim 1, characterized in that, further comprising: determining the identity of the user by recognizing the user's voice and / or face; obtaining the user information from a database based on the identity of the user.

4. The method according to claim 2, characterized in that, determining the operations of at least one target device based on the obtained information and the device control instruction includes: determining the intent of at least one target device and the device control instruction based on the obtained information; determining the operations of at least one target device based on the obtained information and the intent of the device control instruction.

5. The method according to claim 1, characterized in that, the user information includes: user profile information and / or the device control permission information of the user; and / or the device information includes: the working status information of the device; and / or the environmental information includes: the working environment information of the device.

6. The method according to claim 1, characterized in that, outputting an operation result of rejecting to execute the device control instruction includes: when it is determined according to the obtained information that at least one of the following is satisfied, outputting an operation result of rejecting to execute the device control instruction: The user does not have the control authority for the at least one target device; the user does not have the control authority for the target function corresponding to the device control instruction; the at least one target device does not meet the execution conditions corresponding to the device control instruction; the working environment of the at least one target device does not meet the execution conditions corresponding to the device control instruction.

7. The method according to any one of claims 1-6, wherein, controlling at least one target device to perform corresponding operations according to a target function includes: based on the acquired information and the device control instruction, determining at least one target device corresponding to the device control instruction and / or the target function corresponding to the device control instruction; based on the at least one target device and / or the target function, controlling at least one target device to perform corresponding operations.

8. The method according to claim 1, wherein, further includes: if the execution probability of each device is less than the first preset threshold, outputting an operation result of rejecting to execute the device control instruction.

9. The method according to claim 2, wherein, controlling at least one target device to perform corresponding operations according to target parameter information includes: when at least one of the following is satisfied, controlling the at least one target device to perform corresponding operations according to target parameter information: the device control instruction does not contain a parameter value; the parameter value included in the device control instruction does not belong to the parameter value range determined by the acquired information.

10. The method according to claim 1, wherein, controlling at least one target device to perform corresponding operations according to a target function includes: performing sequence annotation processing on the device control instruction to obtain parameter information in the device control instruction; based on the acquired information, determining whether to change the parameter information in the device control instruction; if it is changed, determining the changed parameter information as target parameter information.

11. The method according to claim 10, wherein, based on the parameter information in the device control instruction and the acquired information, determining whether to change the parameter information in the device control instruction includes: based on the parameter information in the device control instruction and the acquired information, performing logistic regression processing to obtain a logistic regression result; determining whether to change the parameter information in the device control instruction based on the logistic regression result; and / or based on the parameter information in the device control instruction and the acquired information, determining the changed target parameter information includes: based on the parameter information in the device control instruction and the acquired information, performing linear regression processing to obtain a linear regression result; determining the changed parameter information based on the linear regression result.

12. The method according to any one of claims 2-11, wherein, any one of the multiple training data includes at least one of the following information: device control instruction; parameter information in the device control instruction; indication information on whether the parameter in the device control instruction is changed; changed parameter information; user information; environment information; device information.

13. The method according to claim 12, wherein, Determine a target loss function according to the parameter information in the training data and the predicted parameter information of the processing model, including: Determine a first loss function based on the parameter information in the device control instructions in each piece of training data and the predicted parameter information of the device control instructions of the model; Determine a second loss function based on the indication information of whether the parameters in the device instructions in each piece of training data are changed and the indication information of whether the change is predicted by the model; Determine a third loss function based on the changed parameter information in each piece of training data and the changed parameter information predicted by the model; Determine the target loss function based on the first loss function, the second loss function, and the third loss function.

14. The method according to any one of claims 1-13, wherein, obtain at least one of user information, environmental information, and device information, and then further include: Convert the discrete information in the obtained information into a continuous dense vector; Determine an information representation vector corresponding to the obtained information according to the converted continuous dense vector and the continuous information in the obtained information; Based on the obtained information and the device control instruction, control at least one target device to perform corresponding operations according to the target function, including: Control at least one target device to perform corresponding operations according to the target function based on the information representation vector corresponding to the obtained information and the device control instruction.

15. The method according to claim 14, wherein, Convert the discrete information in the obtained information into a continuous dense vector, including: Convert the discrete information in the obtained information into a continuous dense vector through an encoding matrix for converting discrete values into continuous dense vectors.

16. An electronic device, wherein, it includes: One or more processors; A memory; One or more applications, wherein the one or more applications are stored in the memory and configured to be executed by the one or more processors, and the one or more programs are configured to: execute the device control method according to any one of claims 1-15.

17. A computer-readable storage medium, wherein, The storage medium stores at least one instruction, at least one segment of program, a code set, or an instruction set, and the at least one instruction, the at least one segment of program, the code set, or the instruction set is loaded and executed by the processor to implement the device control method according to any one of claims 1-15.

Citation Information

Patent Citations

  • Intelligent household management device and system

    CN105182786A

  • Medium-range communication connection system and implementation method

    CN107065583A

  • Information processing method, terminal and compute readable medium

    CN108227565A

  • System and method for sharing record linkage information

    US20140358829A1

  • Device control method, device control system, and server device

    US9774608B2