Control method, device and equipment of intelligent equipment and computer readable storage medium

Through centralized control and reinforced learning technology, intelligent devices automatically select devices that perform target operations, solving the problem of simultaneous feedback from multiple intelligent devices, improving response timeliness and accuracy, and reducing user interference and device power consumption.

CN119937360AActive Publication Date: 2025-05-06CHINA MOBILEHANGZHOUINFORMATION TECH CO LTD +1
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202311445098.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-01
Publication Date
2025-05-06
Estimated Expiration
2043-11-01

AI Technical Summary

Technical Problem

In a smart home environment, when multiple smart devices receive external information and feedback at the same time, it is easy to cause trouble for users, and multiple accounts or smart devices without accounts cannot wake up together.

Method used

Adopting centralized control ideas and reinforcement learning technology, through pre-established learning models and environmental information collected by smart devices, the target intelligent device that performs target operations is automatically selected to reduce user operations and signaling interactions.

Benefits of technology

It improves the timeliness and accuracy of multi-intelligent devices, reduces user interference and device power consumption, and solves the problem that multiple or non-account devices cannot wake up together.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119937360A_ABST
    Figure CN119937360A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a control method, device and equipment of intelligent equipment and a computer readable storage medium. The method comprises the steps that target input is received; in response to the target input, obtaining first environment information collected by the plurality of intelligent devices; determining a reward value of each intelligent device according to a pre-established learning model and the first environment information collected by the plurality of intelligent devices; determining at least one intelligent device as target intelligent preparation based on the reward value of each intelligent device; and controlling the target intelligent device to execute the target operation. The target intelligent device executing the target operation can be automatically selected based on the environment information, active operation of a user is not needed, and interference of simultaneous response of multiple intelligent devices to the user is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application belongs to the technical field of intelligent devices, and in particular, relates to a control method, device, equipment and computer-readable storage medium of an intelligent device. Background Art

[0002] With the development of technology, multiple smart devices are becoming more and more popular, especially driven by smart device manufacturers, smart home solutions are becoming increasingly popular. Multiple smart devices such as mobile phones, smart speakers, TVs, air conditioners, refrigerators, etc. have achieved a unified ecosystem, and devices and people can interact smoothly and intelligently. Sometimes, multiple smart devices will receive external information at the same time. If these devices give feedback at the same time, it will cause trouble to people. Summary of the invention

[0003] The embodiments of the present application provide a control method, apparatus, device and computer-readable storage medium for a smart device, which can automatically select a target smart device to perform a target operation, reduce user operations, and reduce interference to the user.

[0004] In a first aspect, an embodiment of the present application provides a method for controlling an intelligent device, the method for controlling an intelligent device comprising: receiving a target input; in response to the target input, obtaining first environmental information collected by each of multiple intelligent devices; determining a reward value for each intelligent device based on a pre-established learning model and the first environmental information collected by each of the multiple intelligent devices; determining at least one intelligent device as a target intelligent device based on the reward value of each intelligent device; and controlling the target intelligent device to perform a target operation.

[0005] According to the implementation scheme of the first aspect of the present application, the reward value of each smart device is determined based on a pre-established learning model and the first environmental information collected by each of the multiple smart devices, including: for any smart device, the first environmental information collected by the smart device is calculated based on the reward function in the learning model to obtain the reward value of the smart device.

[0006] According to any of the aforementioned implementations of the first aspect of the present application, the first environmental information includes the sound and / or image of the target object.

[0007] According to any of the aforementioned embodiments of the first aspect of the present application, the reward function includes a first reward value, a second reward value, an incentive function, a first weight, a second weight and a third weight; for any smart device, the first environmental information collected by the smart device is calculated based on the reward function in the learning model to obtain the reward value of the smart device, including: for any smart device, the first reward value is calculated based on the size of the sound of the target object collected by the smart device at the target time, and / or, the second reward value is calculated based on the detection result of the target object in the image collected by the smart device at the target time; the sum of the first product, the second product and the third product is calculated to obtain the reward value of the smart device, the first product includes the product of the first weight and the first reward value, the second product includes the product of the second weight and the second reward value, and the third product includes the product of the third weight and the incentive function.

[0008] According to any of the aforementioned embodiments of the first aspect of the present application, before receiving the target input, the control method of the intelligent device also includes: setting a first intelligent device and a second intelligent device from multiple intelligent devices, and setting a reward function; selecting a target action; in response to the target action, switching the first intelligent device to a third intelligent device; obtaining second environmental information collected by the third intelligent device; calculating the second environmental information based on the reward function to obtain a reward value for the third intelligent device; determining whether the third intelligent device is the second intelligent device; when the third intelligent device is not the second intelligent device, updating the reward function and updating the third intelligent device to the first intelligent device, and returning to the step of selecting the target action until the third intelligent device is the second intelligent device and the number of training times reaches a first preset threshold, thereby obtaining a trained learning model.

[0009] According to any of the aforementioned implementations of the first aspect of the present application, after controlling the target smart device to perform the target operation, the smart device control method further includes: when it is detected that the user switches from the target smart device to other smart devices to perform the target operation, retraining the learning model.

[0010] According to any of the aforementioned implementations of the first aspect of the present application, the method for controlling a smart device further includes: constructing a smart device library including multiple smart devices; and retraining the learning model when smart devices are added or reduced in the smart device library.

[0011] According to any of the aforementioned embodiments of the first aspect of the present application, based on the reward values ​​of each smart device, at least one smart device is determined as a target smart device, including: determining a smart device whose reward value is greater than or equal to a first preset threshold as a target smart device.

[0012] According to any of the aforementioned embodiments of the first aspect of the present application, at least one smart device is determined as a target smart device based on the reward value of each smart device, and also includes: obtaining the working status of each smart device; determining a working smart device whose reward value is less than a first preset threshold and whose reward value is greater than or equal to a second preset threshold as a target smart device, and the second preset threshold is less than the first preset threshold.

[0013] According to any of the aforementioned implementations of the first aspect of the present application, controlling the target smart device to perform the target operation includes: controlling the target smart device to issue a prompt message.

[0014] In the second aspect, an embodiment of the present application provides a control device for an intelligent device, the control device for the intelligent device comprising: a receiving module for receiving a target input; a data acquisition module for acquiring first environmental information collected by each of multiple intelligent devices in response to the target input; a learning module for determining a reward value for each intelligent device based on a pre-established learning model and the first environmental information collected by each of the multiple intelligent devices, and determining at least one intelligent device as a target intelligent preparation based on the reward value of each intelligent device; and an instruction issuing module for controlling the target intelligent device to perform a target operation.

[0015] In a third aspect, an embodiment of the present application provides an electronic device, comprising: a processor, a memory, and a computer program stored in the memory and executable on the processor, wherein when the computer program is executed by the processor, the steps of the control method of the smart device provided in the first aspect are implemented.

[0016] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the control method of the smart device provided in the first aspect are implemented.

[0017] The control method, device, equipment and computer-readable storage medium of the smart device of the embodiment of the present application include: receiving target input; in response to the target input, obtaining the first environmental information collected by each of the multiple smart devices; determining the reward value of each smart device according to the pre-established learning model and the first environmental information collected by each of the multiple smart devices; based on the reward value of each smart device, determining at least one smart device as the target smart preparation; controlling the target smart device to perform the target operation. On the one hand, the embodiment of the present application adopts a centralized control concept to uniformly control multiple smart devices, and determines which target smart preparation performs the target operation, thereby reducing the number of signaling interactions between different smart devices and improving the timeliness of the response of multiple smart devices; on the other hand, by constructing a learning model, the target smart device that performs the target operation is automatically selected based on the environmental information, without the need for active operation by the user, thereby improving the response speed of the smart device and reducing the interference of multiple smart devices responding at the same time to the user; on the other hand, the judgment of the target smart device by the embodiment of the present application does not need to rely on the user account, which can solve the problem that smart devices with multiple accounts or no account cannot be awakened collaboratively. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] In order to more clearly illustrate the technical solution of the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0019] Figure 1 A schematic diagram of a flow chart of a control method for an intelligent device provided in an embodiment of the present application;

[0020] Figure 2 A schematic diagram of a flow chart of S103 in the control method of the smart device provided in an embodiment of the present application;

[0021] Figure 3 A flowchart of a training process in a control method for a smart device provided in an embodiment of the present application;

[0022] Figure 4 Another schematic diagram of a flow chart of a control method for a smart device provided in an embodiment of the present application;

[0023] Figure 5 A schematic diagram of another flow chart of the control method of the smart device provided in the embodiment of the present application;

[0024] Figure 6 A schematic diagram of a flow chart of S104 in the control method of the smart device provided in an embodiment of the present application;

[0025] Figure 7A schematic diagram of a structure of a control device for a smart device provided in an embodiment of the present application;

[0026] Figure 8 A schematic diagram of the hardware structure of an electronic device provided in an embodiment of the present application is shown. DETAILED DESCRIPTION

[0027] The features and exemplary embodiments of various aspects of the present application will be described in detail below. In order to make the purpose, technical solutions and advantages of the present application clearer, the present application will be further described in detail below in conjunction with the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are only intended to explain the present application, rather than to limit the present application. For those skilled in the art, the present application can be implemented without the need for some of these specific details. The following description of the embodiments is only to provide a better understanding of the present application by illustrating the examples of the present application.

[0028] It should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the statement "include..." do not exclude the presence of other identical elements in the process, method, article or device including the elements.

[0029] It should be understood that the term "and / or" used in this article is only a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. In addition, the character " / " in this article generally indicates that the associated objects before and after are in an "or" relationship.

[0030] It is obvious to those skilled in the art that various modifications and changes can be made in the present application without departing from the spirit or scope of the present application. Therefore, the present application is intended to cover modifications and changes of the present application that fall within the scope of the corresponding claims (technical solutions for protection) and their equivalents. It should be noted that the implementation methods provided in the embodiments of the present application can be combined with each other without contradiction.

[0031] Before describing the technical solutions provided by the embodiments of the present application, in order to facilitate the understanding of the embodiments of the present application, the present application first specifically describes the problems existing in the related art:

[0032] With the development of science and technology, multiple smart devices are becoming more and more popular, especially driven by smart device manufacturers, smart home solutions are becoming increasingly popular. For example, multiple smart devices such as mobile phones, smart speakers, TVs, air conditioners, refrigerators, etc. have achieved a unified ecosystem, and devices and people can interact smoothly and intelligently. Sometimes, multiple smart devices will receive external information at the same time. If these devices give feedback at the same time, it will cause trouble to people.

[0033] For example, when a user's mobile phone receives an incoming call, the watch, tablet, and computer of the same account may all ring, and the user needs to manually select one of the devices to answer the call. Of course, the user can also set the response priority in each smart device in advance, and when these smart devices receive the event, they will respond according to the user's settings.

[0034] The collaborative response solution among multiple smart devices is upgraded from the previous proximity strategy to an intelligent decision-making of a suitable device response based on information such as distance and active status. Before the user operates the smart devices of the same account, the collaborative response will aggregate and network multiple devices. When the user specifies a device to respond, such as waking up the device with a wake-up word, the collaborative wake-up will extract the direction and distance feature data of the wake-up word and interact between devices in the group. When the global feature data is collected or the decision waiting timeout is reached, a distributed decision is made based on the status of each device and other information, and finally a suitable device response is selected.

[0035] There are many low-power devices in home scenarios, and the power consumption requirements are relatively high. The collaborative response solution between multiple intelligent devices requires each intelligent device to perform more signaling interaction and data processing. According to the principle of collaborative wake-up, when the transmission delay of the wake-up feature message is large or the wake-up delay of different devices is very different, it will lead to the inability to receive all global information within the judgment waiting time, causing inconsistent judgment results of different devices and simultaneous wake-up. The setting of the judgment waiting time will affect the simultaneous wake-up rate and the device response speed. When the judgment waiting time is too long, the device response speed slows down. When the judgment waiting time is too short, the probability of simultaneous wake-up increases.

[0036] In addition, the collaborative response solution between multiple smart devices requires that multiple smart devices are under the same account system. When multiple accounts or no account is logged in (such as a TV), multiple accounts or no account devices cannot be woken up collaboratively.

[0037] In order to solve at least one of the above-mentioned problems in the related art, the embodiments of the present application provide a control method, apparatus, device and computer-readable storage medium of a smart device.

[0038] The technical concept of the embodiments of the present application is that: the embodiments of the present application adopt a centralized control concept to uniformly control multiple smart devices, determine which target smart device is to perform the target operation, reduce the number of signaling interactions between different smart devices, and improve the timeliness of the response of multiple smart devices; on the other hand, by constructing a learning model, the target smart device to perform the target operation is automatically selected based on environmental information without the need for active user operation, thereby improving the response speed of the smart device and reducing the interference of multiple smart devices responding at the same time to the user; on another hand, the embodiments of the present application do not need to rely on user accounts to determine the target smart device, which can solve the problem that smart devices with multiple accounts or no accounts cannot be woken up collaboratively.

[0039] The following first introduces the control method of the smart device provided in the embodiment of the present application.

[0040] Figure 1 A flow chart of a control method for a smart device provided in an embodiment of the present application. Figure 1 As shown, the method may include the following steps S101 to S105.

[0041] S101: Receive target input.

[0042] Exemplarily, the target input includes but is not limited to monitoring a target event, such as receiving a wake-up instruction, an external incoming call signal, and the like.

[0043] S102: In response to a target input, obtain first environmental information collected by each of the plurality of smart devices.

[0044] The smart device has an environment perception capability and can obtain the environment information (or environment data) around the smart device. For the sake of distinction, the environment information collected by the smart device in S102 is referred to as the first environment information.

[0045] S103: Determine a reward value for each smart device according to a pre-established learning model and first environmental information collected by each of the plurality of smart devices.

[0046] A learning model may be established in advance. Exemplarily, the learning model includes but is not limited to a reinforcement learning model established based on a reinforcement learning (RL) algorithm. The learning model may process the first environment information collected by each of the multiple smart devices to determine a reward value for each smart device. The reward value, i.e., the score, may be used to determine whether the smart device can be used as a smart device for executing a target operation.

[0047] S104. Based on the reward values ​​of the various smart devices, determine at least one smart device as a target smart preparation.

[0048] At least one smart device can be selected from multiple smart devices as the target smart device based on the size of the reward value of each smart device. The number of target smart devices can be flexibly adjusted according to actual conditions, and this embodiment of the application does not limit this.

[0049] S105: Control the target smart device to perform the target operation.

[0050] The control method, device, equipment and computer-readable storage medium of the smart device of the embodiment of the present application include: receiving target input; in response to the target input, obtaining the first environmental information collected by each of the multiple smart devices; determining the reward value of each smart device according to the pre-established learning model and the first environmental information collected by each of the multiple smart devices; based on the reward value of each smart device, determining at least one smart device as the target smart preparation; controlling the target smart device to perform the target operation. On the one hand, the embodiment of the present application adopts a centralized control concept to uniformly control multiple smart devices, and determines which target smart preparation performs the target operation, thereby reducing the number of signaling interactions between different smart devices and improving the timeliness of the response of multiple smart devices; on the other hand, by constructing a learning model, the target smart device that performs the target operation is automatically selected based on the environmental information, without the need for active operation by the user, thereby improving the response speed of the smart device and reducing the interference of multiple smart devices responding at the same time to the user; on the other hand, the judgment of the target smart device by the embodiment of the present application does not need to rely on the user account, which can solve the problem that smart devices with multiple accounts or no account cannot be awakened collaboratively.

[0051] According to some embodiments of the present application, optionally, the control method of the smart device of the embodiment of the present application can be applied to a central controller of multiple smart devices (referred to as the central controller), and the central controller executes the control method of the smart device of the embodiment of the present application, uniformly controls multiple smart devices, and determines which target intelligent preparation executes the target operation, thereby reducing the number of signaling interactions between different smart devices and improving the timeliness of the response of multiple smart devices.

[0052] The embodiments of the present application are based on the idea of ​​centralized control and decouple information processing and decision-making from multiple devices. Smart devices only need to collect sound or images, and there is no need for signaling data to interact between devices, thereby reducing signaling overhead and device data processing consumption.

[0053] The specific implementation methods of the above steps are introduced below.

[0054] According to some embodiments of the present application, optionally, S103, determining the reward value of each smart device according to the pre-established learning model and the first environment information collected by each of the multiple smart devices, may include the following steps:

[0055] For any smart device, the first environment information collected by the smart device is calculated based on the reward function in the learning model to obtain a reward value of the smart device.

[0056] Specifically, the first environment information collected by the smart device can be parameterized, that is, the first environment information can be converted into a numerical value that is easy to calculate. A reward function is set in the learning model, and the reward function can be used to calculate the reward value of the smart device based on the first environment information collected by the smart device.

[0057] According to some embodiments of the present application, the first environmental information may optionally include a sound and / or an image of a target object. Exemplarily, the target object includes but is not limited to a user. For example, the smart device may collect sound through a microphone, and / or the smart device may capture an image of the surrounding environment through a camera.

[0058] Based on the sound and / or image of the target object, the smart devices near the target object, i.e., the smart devices that are closer to the target object, can be determined. In some examples, for example, the smart devices near the target object can be used as target smart devices to perform target operations, such as sending a prompt message when a call comes in, so as to remind the user of the incoming call while avoiding waking up smart devices that are far away from the user, thereby reducing power consumption and avoiding the user walking a long distance to turn off the smart devices that are far away from the user.

[0059] The embodiments of the present application are based on the idea of ​​centralized control and decouple information processing and decision-making from multiple devices. Smart devices only need to collect sound or images, and there is no need for signaling data to interact between devices, thereby reducing signaling overhead and device data processing consumption.

[0060] According to some embodiments of the present application, optionally, the reward function may include a first reward value, a second reward value, an incentive function, a first weight, a second weight, and a third weight.

[0061] Figure 2 This is a flow chart of S103 in the control method of the smart device provided in the embodiment of the present application. Figure 2 As shown, according to some embodiments of the present application, optionally, S103, for any smart device, the first environmental information collected by the smart device is calculated based on the reward function in the learning model to obtain the reward value of the smart device, which may include the following steps S201 and S202.

[0062] S201. For any smart device, a first reward value is calculated based on the volume of the sound of the target object collected by the smart device at the target moment, and / or a second reward value is calculated based on the detection result of the target object in the image collected by the smart device at the target moment.

[0063] S202, calculating the sum of the first product, the second product and the third product to obtain a reward value of the smart device, wherein the first product includes the product of the first weight and the first reward value, the second product includes the product of the second weight and the second reward value, and the third product includes the product of the third weight and the incentive function.

[0064] In some examples, the reward function is expressed as follows:

[0065] R t (0,a) = α·R V + β·R H +γ·δ(s i -s target ) (1)

[0066]

[0067] R H = sgn(H t )if has human H > 0 ; else H < 0 (3)

[0068] Among them, s i represents any smart device, R represents receiving the target input at time t, and selecting smart device s i The reward obtained when α, β, and γ are three positive parameters, α represents the first weight, β represents the second weight, and γ represents the third weight. V Indicates smart devices i The reward value obtained based on the sound of the target object at time t. When the sound of the target object is detected, the louder the sound of the target object, the greater the reward value. H Indicates smart devices i The reward value obtained based on the image of the target object at time t. When the image of the target object (i.e., a portrait) is detected, R H is 1, otherwise R H = -1. i -s target ) represents the activation function. sgn(x) represents the step function, x is the input value, such as x = V t or H t When the input value x>0, the output value is 1; when the input value x=0, the output value is 0. V t Indicates smart devices i The volume of the target object's sound collected at time t, V max Indicates the preset maximum volume, V min Indicates the preset minimum volume value.

[0069] According to some embodiments of the present application, optionally, before S101, receiving the target input, the learning model may be trained to obtain a trained learning model.

[0070] Figure 3 A flow chart of the training process in the control method of the smart device provided in the embodiment of the present application. Figure 3 As shown, according to some embodiments of the present application, optionally, before S101, receiving the target input, the control method of the smart device may further include the following steps S301 to S307.

[0071] S301. Set a first smart device and a second smart device from a plurality of smart devices, and set a reward function. The first smart device can be understood as an initial smart device, and the second smart device can be understood as a smart device that is ultimately expected to be switched to. In S301, an initial reward function can be set. In some examples, in S301, an initial action value table (Q value table) can also be set. The Q value table can include the maximum expected reward for each state, which is used to guide the best action to be taken in each state, such as guiding the switch from smart device A to smart device B.

[0072] S302: Select a target action.

[0073] Exemplarily, an action ε-greedy strategy is adopted to select the target action.

[0074] S303: In response to the target action, switch the first smart device to a third smart device.

[0075] S304: Acquire second environment information collected by the third smart device.

[0076] The third smart device may be any smart device other than the first smart device. Exemplarily, the second environment information may include the sound and / or image of the target object. Exemplarily, the target object includes but is not limited to the user. For example, the third smart device may collect sound through a microphone, and / or the third smart device may capture the surrounding environment image through a camera.

[0077] S305: Calculate the second environment information based on the reward function to obtain a reward value of the third smart device.

[0078] For example, the second environment information collected by the third smart device can be calculated based on the reward function in a similar manner to the above-mentioned first environment information to calculate the reward value of the smart device, such as the above-mentioned expressions (1) to (3), to obtain the reward value of the third smart device. In S304, the Q value table can also be updated based on the reward value of the third smart device.

[0079] S306: Determine whether the third smart device is the second smart device.

[0080] S307, when the third smart device is not the second smart device, update the reward function, update the third smart device to the first smart device, and return to the step of selecting the target action, until the third smart device is the second smart device and the number of training times reaches the first preset number threshold, and a trained learning model is obtained. If the number of training times does not reach the first preset number threshold, the number of training times is increased by 1 and returns to S302.

[0081] Among them, the size of the first preset number threshold can be flexibly adjusted according to actual conditions, and the embodiment of the present application does not limit this.

[0082] According to some embodiments of the present application, optionally, a maximum number of steps for a single training session may also be set. When the third smart device is not the second smart device, it is determined whether the number of steps for a single training session reaches the maximum number of steps for a single training session. If the number of steps for a single training session does not reach the maximum number of steps for a single training session, the reward function is updated, the third smart device is updated to the first smart device, and S302 is returned. If the number of steps for a single training session reaches the maximum number of steps for a single training session, it is continued to be determined whether the number of training sessions reaches a first preset number threshold. When the number of training sessions reaches the first preset number threshold, the training ends.

[0083] According to some embodiments of the present application, optionally, after S307, it may be determined whether the first smart device and the second smart device have traversed all smart devices. If not, S301 to S307 are continued.

[0084] Figure 4 Another flow chart of the control method of the smart device provided in the embodiment of the present application. Figure 4 As shown, according to some embodiments of the present application, optionally, after controlling the target smart device to perform the target operation, the smart device control method may further include the following steps:

[0085] S401: When it is detected that the user switches from the target smart device to another smart device to perform a target operation, the learning model is retrained.

[0086] When it is detected that the user switches from the target smart device to another smart device to perform the target operation, it can be said that the user does not expect the target smart device to perform the target operation, or the user expects another smart device to perform the target operation. In this case, the learning model can be retrained to improve the accuracy of the learning model prediction.

[0087] That is, when it is detected that the user switches from the target smart device to another smart device to perform the target operation, it is considered that the learning model is no longer applicable to the current environment. The central controller can re-issue instructions, and the smart device re-acquires the environment information and uploads it to re-train the learning model.

[0088] Figure 5 This is another flow chart of the control method of the smart device provided in the embodiment of the present application. Figure 5 As shown, according to some embodiments of the present application, optionally, the control method of the smart device may further include the following steps S501 and S502.

[0089] S501: Build a smart device library including multiple smart devices.

[0090] In S501, a smart device library including multiple smart devices can be constructed based on multiple existing smart devices. As shown in expression (4), where represents the smart device library, s i Represents a smart device, 0≤i≤n, i and n are integers.

[0091] S = {s0,s1,...,s n} (4)

[0092] S502: When smart devices are added or reduced in the smart device library, the learning model is retrained.

[0093] When smart devices are added or removed from the smart device library, the learning model is no longer applicable to the current environment. The central controller can re-issue instructions, and the smart devices can re-acquire environmental information and upload it to re-train the learning model.

[0094] In the embodiment of the present application, the learning model is updated based on user behavior or device status. When external conditions such as the spatial distribution of multiple devices, the number of devices, or user behavior habits change, the device portrait information contained in the original learning model is no longer available, and the dynamically responding device does not match the user's preference, causing the user to actively change the device status. At this time, the smart device is re-controlled to obtain environmental information and the learning model is retrained to ensure the accuracy of the model.

[0095] According to some embodiments of the present application, optionally, S104, based on the reward values ​​of the respective smart devices, determining at least one smart device as a target smart preparation, may include the following steps:

[0096] A smart device having a reward value greater than or equal to a first preset threshold is determined as a target smart device.

[0097] Among them, the first preset threshold can be flexibly adjusted according to actual conditions, and the embodiments of the present application do not limit this. Smart devices with larger reward values ​​are generally closer to the target object (i.e., the user). Selecting smart devices with larger reward values ​​as target smart devices usually meets the user's expectations and improves the user experience.

[0098] Figure 6 This is a flow chart of S104 in the control method of the smart device provided in the embodiment of the present application. Figure 6 As shown, according to some embodiments of the present application, optionally, S104, based on the reward value of each smart device, determining at least one smart device as the target smart preparation, may also include the following steps S601 and S602.

[0099] S601: Obtain the working status of each smart device.

[0100] S602: Determine a working smart device whose reward value is less than a first preset threshold and whose reward value is greater than or equal to a second preset threshold as a target smart device, where the second preset threshold is less than the first preset threshold.

[0101] When the smart device is working, it means that the user may be using the smart device, for example, the user is watching TV. At this time, if the reward value of the smart device is less than the first preset threshold, but the reward value of the smart device is greater than or equal to the second preset threshold, the working smart device can still be determined as the target smart device to perform the target operation. For example, the target smart device can send a prompt message to ensure that the user can receive the prompt message to a greater extent.

[0102] According to some embodiments of the present application, optionally, S105, controlling the target smart device to perform the target operation, may include the following steps:

[0103] Control the target smart device to send out prompt information.

[0104] For example, the target smart device may be controlled to ring, vibrate, and / or issue a prompt, etc., which is not limited in this embodiment of the present application.

[0105] The embodiments of the present application are based on the idea of ​​centralized control and decouple information processing and decision-making from multiple devices. Smart devices only need to collect sound or images, and there is no need for signaling data to interact between devices, thereby reducing signaling overhead and device data processing consumption.

[0106] The embodiment of the present application uses reinforcement learning technology to build a learning model, automatically selects a response device based on the environment, and does not require active user operation, thereby improving the device response speed. Incorporating human voice, human image, and user behavior into the decision-making basis for multiple intelligent devices to respond to events improves the accuracy and intelligence of the response device, and realizes the intelligent response function of the intelligent device changing with the changes in user behavior.

[0107] The embodiment of the present application combines the centralized control concept and reinforcement learning technology to collect environmental information and train the learning model, and make decisions based on the learning model. Multiple intelligent devices only need to be in the controller network, without logging into a unified account, so that they can be compatible with more models of different devices and have good scalability.

[0108] Based on the control method of the smart device provided in the above embodiment, the present application also provides a specific implementation of the control device of the smart device. Please refer to the following embodiment.

[0109] Figure 7 A schematic diagram of a structure of a control device for an intelligent device provided in an embodiment of the present application. Figure 7 As shown, the control device 70 of the smart device provided in the embodiment of the present application may include the following modules:

[0110] Receiving module 701, used for receiving target input;

[0111] The data collection module 702 is used to obtain first environment information collected by each of the plurality of smart devices in response to a target input;

[0112] The learning module 703 is used to determine the reward value of each smart device according to the pre-established learning model and the first environment information collected by each of the multiple smart devices, and determine at least one smart device as a target smart device based on the reward value of each smart device;

[0113] The instruction issuing module 704 is used to control the target smart device to perform the target operation.

[0114] The smart device provided in the embodiment of the present application adopts a centralized control concept to uniformly control multiple smart devices, and determines which target smart device is to perform the target operation, thereby reducing the number of signaling interactions between different smart devices and improving the timeliness of the response of multiple smart devices; on the other hand, by constructing a learning model, the target smart device to perform the target operation is automatically selected based on environmental information without the need for active user operation, thereby improving the response speed of the smart device and reducing the interference of multiple smart devices responding at the same time to the user; on another hand, the embodiment of the present application does not need to rely on user accounts to determine the target smart device, and can solve the problem that smart devices with multiple accounts or no accounts cannot be woken up collaboratively.

[0115] According to some embodiments of the present application, optionally, the control device 70 of the smart device may further include a data processing module for performing parameterized processing on the first environmental information collected by the smart device, such as enabling the first environmental information to be converted into a numerical value that is easy to calculate.

[0116] According to some embodiments of the present application, optionally, the learning module 703 is specifically used to calculate the first environmental information collected by the smart device based on the reward function in the learning model for any smart device to obtain a reward value of the smart device.

[0117] According to some embodiments of the present application, optionally, the first environmental information includes sound and / or image of the target object.

[0118] According to some embodiments of the present application, optionally, the reward function includes a first reward value, a second reward value, an incentive function, a first weight, a second weight, and a third weight. The learning module 703 is specifically used to calculate the first reward value for any smart device based on the size of the sound of the target object collected by the smart device at the target time, and / or, based on the detection result of the target object in the image collected by the smart device at the target time, calculate the second reward value; calculate the sum of the first product, the second product and the third product to obtain the reward value of the smart device, the first product includes the product of the first weight and the first reward value, the second product includes the product of the second weight and the second reward value, and the third product includes the product of the third weight and the incentive function.

[0119] According to some embodiments of the present application, optionally, the control device 70 of the smart device may also include a training module, which is used to set a first smart device and a second smart device from multiple smart devices, and set a reward function; select a target action; in response to the target action, switch the first smart device to a third smart device; obtain the second environmental information collected by the third smart device; calculate the second environmental information based on the reward function to obtain a reward value for the third smart device; determine whether the third smart device is the second smart device; when the third smart device is not the second smart device, update the reward function and update the third smart device to the first smart device, and return to the step of selecting the target action, until the third smart device is the second smart device and the number of training times reaches a first preset number threshold, and a trained learning model is obtained.

[0120] According to some embodiments of the present application, optionally, the training module can also be used to retrain the learning model when it is detected that the user switches from the target smart device to other smart devices to perform the target operation.

[0121] According to some embodiments of the present application, optionally, the training module can also be used to build a smart device library containing multiple smart devices; when smart devices are added or reduced in the smart device library, the learning model is retrained.

[0122] According to some embodiments of the present application, optionally, the learning module 703 is specifically configured to determine a smart device having a reward value greater than or equal to a first preset threshold as a target smart device.

[0123] According to some embodiments of the present application, optionally, the learning module 703 is specifically used to obtain the working status of each smart device; a working smart device whose reward value is less than a first preset threshold and whose reward value is greater than or equal to a second preset threshold is determined as a target smart device, and the second preset threshold is less than the first preset threshold.

[0124] According to some embodiments of the present application, optionally, the instruction issuing module 704 is specifically used to control the target smart device to issue a prompt message.

[0125] Figure 7 Each module / unit in the device shown has the function of implementing each step in the control method of the smart device provided in the above method embodiment and can achieve its corresponding technical effect. For the sake of concise description, it will not be repeated here.

[0126] Based on the control method of the smart device provided in the above embodiment, the present application also provides a specific implementation method of the electronic device accordingly. Please refer to the following embodiment.

[0127] Figure 8 A schematic diagram of the hardware structure of an electronic device provided in an embodiment of the present application is shown.

[0128] The electronic device may include a processor 801 and a memory 802 storing computer program instructions.

[0129] Specifically, the processor 801 may include a central processing unit (CPU), or an application specific integrated circuit (ASIC), or may be configured to implement one or more integrated circuits of the embodiments of the present application.

[0130] The memory 802 may include a large capacity memory for data or instructions. By way of example and not limitation, the memory 802 may include a hard disk drive (HDD), a floppy disk drive, a flash memory, an optical disk, a magneto-optical disk, a magnetic tape, or a universal serial bus (USB) drive, or a combination of two or more of these. In one example, the memory 802 may include a removable or non-removable (or fixed) medium, or the memory 802 is a non-volatile solid-state memory. The memory 802 may be inside or outside the electronic device.

[0131] In one example, the memory 802 may be a read-only memory (ROM). In one example, the ROM may be a mask-programmed ROM, a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (EEPROM), an electrically rewritable ROM (EAROM), or a flash memory, or a combination of two or more of these.

[0132] The memory 802 may include a read-only memory (ROM), a random access memory (RAM), a magnetic disk storage medium device, an optical storage medium device, a flash memory device, an electrical, optical or other physical / tangible memory storage device. Thus, typically, the memory includes one or more tangible (non-transitory) computer-readable storage media (e.g., a memory device) encoded with software including computer-executable instructions, and when the software is executed (e.g., by one or more processors), it is operable to perform the operations described with reference to the method according to an aspect of the present application.

[0133] The processor 801 implements the method / steps in the above method embodiment by reading and executing the computer program instructions stored in the memory 802, and achieves the corresponding technical effects achieved by the method embodiment executing its method / steps, which will not be repeated here for the sake of brevity.

[0134] In one example, the electronic device may further include a communication interface 803 and a bus 810. Figure 8 As shown, the processor 801, the memory 802, and the communication interface 803 are connected via a bus 810 and communicate with each other.

[0135] The communication interface 803 is mainly used to implement communication between various modules, devices, units and / or equipment in the embodiments of the present application.

[0136] Bus 810 includes hardware, software or both, and the parts of electronic equipment are coupled to each other. For example, but not limitation, bus may include accelerated graphics port (Accelerated Graphics Port, AGP) or other graphics bus, enhanced industry standard architecture (Extended Industry Standard Architecture, EISA) bus, front side bus (Front Side Bus, FSB), Hyper Transport (Hyper Transport, HT) interconnection, industry standard architecture (Industry Standard Architecture, ISA) bus, infinite bandwidth interconnection, low pin count (LPC) bus, memory bus, micro channel architecture (MCA) bus, peripheral component interconnection (PCI) bus, PCI-Express (PCI-X) bus, serial advanced technology attachment (SATA) bus, video electronics standard association local (VLB) bus or other suitable bus or two or more of these combinations. In appropriate cases, bus 810 may include one or more buses. Although the present application embodiment describes and shows a specific bus, the application considers any suitable bus or interconnection.

[0137] In addition, in combination with the control method of the smart device in the above embodiment, the embodiment of the present application may provide a computer-readable storage medium for implementation. The computer-readable storage medium stores computer program instructions; when the computer program instructions are executed by the processor, any one of the control methods of the smart device in the above embodiment is implemented. Examples of computer-readable storage media include non-transitory computer-readable storage media, such as electronic circuits, semiconductor memory devices, ROM, random access memory, flash memory, erasable ROM (EROM), floppy disks, CD-ROMs, optical disks, and hard disks.

[0138] It should be clear that the present application is not limited to the specific configuration and processing described above and shown in the figures. For the sake of simplicity, a detailed description of the known method is omitted here. In the above embodiments, several specific steps are described and shown as examples. However, the method process of the present application is not limited to the specific steps described and shown, and those skilled in the art can make various changes, modifications and additions, or change the order between the steps after understanding the spirit of the present application.

[0139] The functional blocks shown in the above-described block diagram can be implemented as hardware, software, firmware or a combination thereof. When implemented in hardware, it can be, for example, an electronic circuit, an application-specific integrated circuit (Application Specific Integrated Circuit, ASIC), appropriate firmware, plug-in, function card, etc. When implemented in software, the elements of the present application are programs or code segments used to perform the required tasks. The program or code segment can be stored in a machine-readable medium, or transmitted on a transmission medium or communication link by a data signal carried in a carrier. "Machine-readable medium" may include any medium capable of storing or transmitting information. Examples of machine-readable media include electronic circuits, semiconductor memory devices, ROM, flash memory, erasable ROM (EROM), floppy disks, CD-ROMs, optical disks, hard disks, optical fiber media, radio frequency (Radio Frequency, RF) links, etc. The code segment can be downloaded via a computer network such as the Internet, an intranet, etc.

[0140] It should also be noted that the exemplary embodiments mentioned in this application describe some methods or systems based on a series of steps or devices. However, this application is not limited to the order of the above steps, that is, the steps can be performed in the order mentioned in the embodiment, or in a different order from the embodiment, or several steps can be performed simultaneously.

[0141] The above reference is according to the method of the embodiment of the present application, the flow chart of the device (system) and the computer program product and / or the block diagram described various aspects of the present application.It should be understood that each square box in the flow chart and / or the block diagram and the combination of each square box in the flow chart and / or the block diagram can be realized by computer program instructions.These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer or other programmable data processing device to produce a machine so that these instructions executed by the processor of the computer or other programmable data processing device enable the realization of the function / action specified in one or more square boxes of the flow chart and / or the block diagram.Such a processor can be but is not limited to a general-purpose processor, a special-purpose processor, a special application processor or a field programmable logic circuit.It can also be understood that each square box in the block diagram and / or the flow chart and the combination of the square boxes in the block diagram and / or the flow chart can also be realized by the dedicated hardware that performs the specified function or action, or can be realized by the combination of dedicated hardware and computer instructions.

[0142] The above is only a specific implementation of the present application. Those skilled in the art can clearly understand that for the convenience and simplicity of description, the specific working processes of the systems, modules and units described above can refer to the corresponding processes in the aforementioned method embodiments, and will not be repeated here. It should be understood that the protection scope of the present application is not limited to this. Any technician familiar with the technical field can easily think of various equivalent modifications or replacements within the technical scope disclosed in this application, and these modifications or replacements should be included in the protection scope of this application.

Claims

1. A control method for an intelligent device, characterized in that: include: Receive target input; In response to the target input, obtaining first environmental information collected by each of the plurality of smart devices; Determine a reward value for each smart device according to a pre-established learning model and first environmental information collected by each of the plurality of smart devices; Based on the reward values ​​of the respective smart devices, determining at least one of the smart devices as a target smart device; Control the target smart device to perform a target operation.

2. The method according to claim 1, characterized in that The step of determining the reward value of each smart device according to the pre-established learning model and the first environment information collected by each of the plurality of smart devices includes: For any one of the smart devices, the first environmental information collected by the smart device is calculated based on the reward function in the learning model to obtain a reward value of the smart device.

3. The method according to claim 1 or 2, characterized in that: The first environmental information includes the sound and / or image of the target object.

4. The method according to claim 3, characterized in that The reward function includes a first reward value, a second reward value, an incentive function, a first weight, a second weight and a third weight; The step of calculating, for any one of the smart devices, the first environmental information collected by the smart device based on the reward function in the learning model to obtain a reward value of the smart device includes: For any one of the smart devices, the first reward value is calculated based on the volume of the sound of the target object collected by the smart device at the target time, and / or the second reward value is calculated based on the detection result of the target object in the image collected by the smart device at the target time; Calculate the sum of a first product, a second product, and a third product to obtain a reward value of the smart device, wherein the first product includes the product of the first weight and the first reward value, the second product includes the product of the second weight and the second reward value, and the third product includes the product of the third weight and the incentive function.

5. The method according to claim 2, characterized in that: Before receiving the target input, the method further includes: Setting a first smart device and a second smart device from the plurality of smart devices, and setting a reward function; Select the target action; In response to the target action, switching the first smart device to a third smart device; Acquire second environment information collected by the third smart device; Calculating the second environment information based on the reward function to obtain a reward value of the third smart device; Determining whether the third smart device is the second smart device; When the third smart device is not the second smart device, update the reward function, update the third smart device to the first smart device, and return to the step of selecting the target action until the third smart device is the second smart device and the number of training times reaches a first preset threshold, thereby obtaining a trained learning model.

6. The method according to claim 1, characterized in that After controlling the target smart device to perform the target operation, the method further includes: When it is detected that the user switches from the target smart device to another smart device to perform the target operation, the learning model is retrained.

7. The method according to claim 1, characterized in that Also includes: Building a smart device library including the plurality of smart devices; When smart devices are added or reduced in the smart device library, the learning model is retrained.

8. The method according to claim 1, characterized in that The step of determining at least one of the smart devices as a target smart device based on the reward values ​​of the smart devices includes: The smart device whose reward value is greater than or equal to a first preset threshold is determined as the target smart device.

9. The method according to claim 1, characterized in that: The step of determining at least one of the smart devices as a target smart device based on the reward values ​​of the smart devices further includes: Obtaining the working status of each of the smart devices; The working smart device whose reward value is less than the first preset threshold and whose reward value is greater than or equal to a second preset threshold is determined as the target smart device, and the second preset threshold is less than the first preset threshold.

10. The method according to claim 1, characterized in that The controlling the target smart device to perform a target operation includes: Control the target smart device to send a prompt message.

11. A control device for an intelligent device, characterized in that: include: A receiving module, used for receiving target input; A data collection module, configured to obtain first environmental information collected by each of the plurality of smart devices in response to the target input; A learning module, configured to determine a reward value of each smart device according to a pre-established learning model and first environmental information collected by each of the plurality of smart devices, and determine at least one of the smart devices as a target smart device based on the reward value of each smart device; The instruction issuing module is used to control the target smart device to perform the target operation.

12. An electronic device, characterized in that: The electronic device comprises: a processor, a memory, and a computer program stored in the memory and executable on the processor, wherein the computer program implements the steps of the control method of the smart device according to any one of claims 1 to 10 when executed by the processor.

13. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the control method of the intelligent device according to any one of claims 1 to 10 are implemented.

Citation Information

Patent Citations

  • Wake-up method and device for voice smart equipment, equipment and storage medium

    CN109391528A

  • Response method, device and system for multiple intelligent devices, and storage medium

    CN110211580A

  • Smart device control method, apparatus and system based on IoT operating system

    CN110361978A

  • Method and device for managing intelligent equipment, computer readable medium and equipment

    CN112415901A

  • Terminal device, awakening method and device thereof and computer readable storage medium

    CN113096658A