Control method, device and apparatus of intelligent device, and computer readable storage medium
By using centralized control and learning models, the system automatically selects target intelligent devices, solving the problems of untimely response from multiple intelligent devices and user interference, and achieving efficient and low-power collaborative control of intelligent devices.
Patent Information
- Application Number
- CN202311445098.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-01
- Publication Date
- 2026-02-24
- Estimated Expiration
- 2043-11-01
AI Technical Summary
When multiple smart devices receive external information simultaneously, it can easily cause user confusion. Existing collaborative response solutions suffer from signaling interaction delays, high power consumption, and account dependency issues, resulting in untimely responses and increased interference.
By adopting a centralized control approach, a learning model is built to automatically select target smart devices based on environmental information, reducing signaling interactions, improving response speed, and solving the problem of collaborative wake-up of devices without accounts.
It improves the responsiveness of multiple intelligent devices, reduces user interference, lowers signaling overhead and device data processing consumption, and is compatible with different device types.
Smart Images

Figure CN119937360B_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of intelligent device technology, and in particular relates to a control method, apparatus, device and computer-readable storage medium for an intelligent device. Background Technology
[0002] With the development of technology, smart devices are becoming increasingly common, especially driven by smart device manufacturers, leading to a surge in demand for smart home solutions. Smartphones, smart speakers, televisions, air conditioners, refrigerators, and other smart devices are forming a unified ecosystem, enabling seamless smart interaction between devices and people. However, sometimes multiple smart devices receive external information simultaneously, and if these devices respond at the same time, it can actually cause confusion and inconvenience. Summary of the Invention
[0003] This application provides a control method, apparatus, device, and computer-readable storage medium for a smart device, which can automatically select the target smart device to perform the target operation, reducing user operations and interference to the user.
[0004] In a first aspect, embodiments of this application provide a control method for an intelligent device. The control method includes: receiving a target input; in response to the target input, acquiring first environmental information collected by each of a plurality of intelligent devices; determining a reward value for each intelligent device based on a pre-established learning model and the first environmental information collected by each of the plurality of intelligent devices; identifying at least one intelligent device as a target intelligent device based on the reward values of each intelligent device; and controlling the target intelligent device to perform a target operation.
[0005] According to the implementation method of the first aspect of this application, the reward value of each smart device is determined based on a pre-established learning model and first environmental information collected by each of the multiple smart devices, including: for any smart device, calculating the first environmental information collected by the smart device based on the reward function in the learning model to obtain the reward value of the smart device.
[0006] According to any of the foregoing embodiments of the first aspect of this application, the first environmental information includes the sound and / or image of the target object.
[0007] According to any of the foregoing embodiments of the first aspect of this application, the reward function includes a first reward value, a second reward value, an activation function, a first weight, a second weight, and a third weight; for any smart device, the reward value of the smart device is obtained by calculating the first environmental information collected by the smart device based on the reward function in the learning model, including: for any smart device, calculating the first reward value based on the volume of the sound of the target object collected by the smart device at the target time, and / or calculating the second reward value based on the detection result of the target object in the image collected by the smart device at the target time; calculating the sum of the first product, the second product, and the third product to obtain the reward value of the smart device, wherein the first product includes the product of the first weight and the first reward value, the second product includes the product of the second weight and the second reward value, and the third product includes the product of the third weight and the activation function.
[0008] According to any of the foregoing embodiments of the first aspect of this application, before receiving the target input, the control method of the intelligent device further includes: setting a first intelligent device and a second intelligent device from a plurality of intelligent devices, and setting a reward function; selecting a target action; in response to the target action, switching the first intelligent device to a third intelligent device; acquiring second environmental information collected by the third intelligent device; calculating the reward value of the third intelligent device based on the reward function; determining whether the third intelligent device is the second intelligent device; when the third intelligent device is not the second intelligent device, updating the reward function, updating the third intelligent device to the first intelligent device, and returning to the step of selecting the target action, until the third intelligent device is the second intelligent device and the number of training times reaches a first preset threshold, thereby obtaining a trained learning model.
[0009] According to any of the foregoing embodiments of the first aspect of this application, after controlling the target smart device to perform the target operation, the control method of the smart device further includes: when it is detected that the user switches from the target smart device to another smart device to perform the target operation, retraining the learning model.
[0010] According to any of the foregoing embodiments of the first aspect of this application, the control method for a smart device further includes: constructing a smart device library containing multiple smart devices; and retraining the learning model when smart devices are added to or removed from the smart device library.
[0011] According to any of the foregoing embodiments of the first aspect of this application, determining at least one smart device as a target smart device based on the reward value of each smart device includes: determining a smart device whose reward value is greater than or equal to a first preset threshold as a target smart device.
[0012] According to any of the foregoing embodiments of the first aspect of this application, determining at least one smart device as a target smart device based on the reward value of each smart device further includes: acquiring the working status of each smart device; determining a working smart device whose reward value is less than a first preset threshold and whose reward value is greater than or equal to a second preset threshold as a target smart device, wherein the second preset threshold is less than the first preset threshold.
[0013] According to any of the foregoing embodiments of the first aspect of this application, controlling a target smart device to perform a target operation includes: controlling the target smart device to issue a prompt message.
[0014] Secondly, embodiments of this application provide a control device for an intelligent device, comprising: a receiving module for receiving target input; a data acquisition module for acquiring first environmental information collected by multiple intelligent devices in response to the target input; a learning module for determining the reward value of each intelligent device based on a pre-established learning model and the first environmental information collected by the multiple intelligent devices, and identifying at least one intelligent device as a target intelligent device based on the reward value of each intelligent device; and an instruction issuing module for controlling the target intelligent device to perform a target operation.
[0015] Thirdly, embodiments of this application provide an electronic device, which includes: a processor, a memory, and a computer program stored in the memory and executable on the processor. When the computer program is executed by the processor, it implements the steps of the control method for the intelligent device provided in the first aspect.
[0016] Fourthly, embodiments of this application provide a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps of the control method for the intelligent device provided in the first aspect.
[0017] This application discloses a control method, apparatus, device, and computer-readable storage medium for intelligent devices. The method includes: receiving a target input; in response to the target input, acquiring first environmental information collected by multiple intelligent devices; determining a reward value for each intelligent device based on a pre-established learning model and the first environmental information collected by each intelligent device; identifying at least one intelligent device as a target intelligent device based on the reward values of each intelligent device; and controlling the target intelligent device to perform a target operation. On one hand, this application adopts a centralized control approach to uniformly control multiple intelligent devices and determine which target intelligent device will perform the target operation, reducing the number of signaling interactions between different intelligent devices and improving the timeliness of multi-intelligent device responses. On the other hand, by constructing a learning model, the target intelligent device for performing the target operation is automatically selected based on environmental information, eliminating the need for user intervention, improving the response speed of intelligent devices, and reducing interference to the user from simultaneous responses by multiple intelligent devices. Furthermore, the determination of the target intelligent device in this application does not rely on user accounts, solving the problem of multiple accounts or unaccounted intelligent devices being unable to coordinate wake-up. Attached Figure Description
[0018] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the embodiments of this application will be briefly introduced below. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0019] Figure 1 A flowchart illustrating a control method for an intelligent device provided in an embodiment of this application;
[0020] Figure 2 A flowchart illustrating step S103 of the control method for a smart device provided in an embodiment of this application;
[0021] Figure 3 A flowchart illustrating the training process in the control method for an intelligent device provided in this application embodiment;
[0022] Figure 4 Another schematic flowchart illustrating the control method for a smart device provided in an embodiment of this application;
[0023] Figure 5 This is another schematic flowchart illustrating the control method for an intelligent device provided in an embodiment of this application;
[0024] Figure 6 A flowchart illustrating step S104 of the control method for a smart device provided in an embodiment of this application;
[0025] Figure 7A schematic diagram of the structure of a control device for an intelligent device provided in an embodiment of this application;
[0026] Figure 8 A schematic diagram of the hardware structure of the electronic device provided in an embodiment of this application is shown. Detailed Implementation
[0027] The features and exemplary embodiments of various aspects of this application will be described in detail below. To make the objectives, technical solutions, and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are only intended to explain this application and not to limit it. For those skilled in the art, this application can be implemented without some of these specific details. The following description of the embodiments is merely to provide a better understanding of this application by illustrating examples.
[0028] It should be noted that, in this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising..." does not exclude the presence of additional identical elements in the process, method, article, or apparatus that includes said element.
[0029] It should be understood that the term "and / or" used in this article is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, and B existing alone. Additionally, the character " / " in this article generally indicates that the preceding and following related objects have an "or" relationship.
[0030] Various modifications and variations can be made to this application without departing from its spirit or scope, which will be apparent to those skilled in the art. Therefore, this application is intended to cover modifications and variations falling within the scope of the corresponding claims (the claimed technical solutions) and their equivalents. It should be noted that the embodiments provided in this application can be combined with each other without contradiction.
[0031] Before describing the technical solutions provided in the embodiments of this application, in order to facilitate understanding of the embodiments of this application, this application first specifically explains the problems existing in the related technologies:
[0032] With the development of technology, smart devices are becoming increasingly common, especially driven by smart device manufacturers, leading to a surge in demand for smart home solutions. Multiple smart devices, such as smartphones, smart speakers, televisions, air conditioners, and refrigerators, are forming a unified ecosystem, enabling seamless smart interaction between devices and people. However, sometimes multiple smart devices receive external information simultaneously, and if these devices respond at the same time, it can actually cause confusion and inconvenience.
[0033] For example, when a user's phone receives an incoming call, watches, tablets, and computers under the same account may also ring, requiring the user to manually select one device to answer the call. Alternatively, users can pre-set response priorities for each smart device, so that when these devices receive an event, they will respond according to the user's settings.
[0034] The multi-device collaborative response solution upgrades the previous proximity strategy to intelligently decide which device should respond based on information such as distance and activity status. Before user interaction, collaborative response aggregates and networks smart devices under the same account. When a user specifies a device to respond, such as waking the device with a wake word, collaborative wake-up extracts the direction and distance features of the wake word and interacts with devices within the group. After collecting all global feature data or when the decision-making timeout occurs, distributed decision-making is performed by combining information such as the status of each device, ultimately selecting a suitable device to respond.
[0035] In home environments, many devices are low-power, requiring high power efficiency. Collaborative response solutions for multiple smart devices necessitate extensive signaling interaction and data processing among them. According to the principle of collaborative wake-up, when the transmission delay of wake-up feature messages is large, or when there are significant differences in wake-up delays between different devices, it can lead to incomplete global information collection during the decision waiting period, causing inconsistent decision results and simultaneous wake-ups. The decision waiting time affects the simultaneous wake-up rate and device response speed; if the decision waiting time is too long, the device response speed slows down, while if the decision waiting time is too short, the probability of simultaneous wake-ups increases.
[0036] Furthermore, collaborative response solutions between multiple smart devices require all devices to be under the same account system. When multiple accounts are involved, or when no account is logged in (such as a TV), multiple accounts or devices without accounts cannot be woken up collaboratively.
[0037] To address at least one of the aforementioned problems in the related technologies, embodiments of this application provide a control method, apparatus, device, and computer-readable storage medium for a smart device.
[0038] The technical concept of this application embodiment is as follows: This application embodiment adopts a centralized control concept to uniformly control multiple smart devices, determine which target smart device should perform the target operation, reduce the number of signaling interactions between different smart devices, and improve the timeliness of multi-smart device response; on the other hand, by constructing a learning model, the target smart device for performing the target operation is automatically selected based on environmental information, without the need for active user operation, which improves the response speed of smart devices and reduces the interference to the user caused by multiple smart devices responding simultaneously; furthermore, the determination of the target smart device in this application embodiment does not rely on user accounts, which can solve the problem that multiple smart devices with different accounts or without accounts cannot be woken up in a coordinated manner.
[0039] The control method for the intelligent device provided in the embodiments of this application will be introduced first below.
[0040] Figure 1 This is a schematic flowchart illustrating a control method for an intelligent device provided in an embodiment of this application. Figure 1 As shown, the method may include the following steps S101 to S105.
[0041] S101, Receive target input.
[0042] For example, target inputs include, but are not limited to, listening to target events, such as receiving a wake-up command or an external incoming call.
[0043] S102. In response to the target input, acquire the first environmental information collected by each of the multiple smart devices.
[0044] Intelligent devices have environmental sensing capabilities and can acquire environmental information (or environmental data) around them. For ease of distinction, the environmental information collected by the intelligent device in S102 is referred to as the first environmental information.
[0045] S103. Based on the pre-established learning model and the first environmental information collected by each of the multiple smart devices, determine the reward value of each smart device.
[0046] A learning model can be pre-built. For example, the learning model includes, but is not limited to, a reinforcement learning model based on a reinforcement learning (RL) algorithm. The learning model can process the initial environmental information collected by multiple smart devices to determine the reward value for each smart device. The reward value, or score, can be used to determine whether a smart device can be used to perform the target operation.
[0047] S104. Based on the reward value of each smart device, identify at least one smart device as the target smart device.
[0048] At least one smart device can be selected as the target smart device based on the reward value of each smart device. The number of target smart devices can be flexibly adjusted according to the actual situation, and this application embodiment does not limit this.
[0049] S105. Control the target intelligent device to perform the target operation.
[0050] This application discloses a control method, apparatus, device, and computer-readable storage medium for intelligent devices. The method includes: receiving a target input; in response to the target input, acquiring first environmental information collected by multiple intelligent devices; determining a reward value for each intelligent device based on a pre-established learning model and the first environmental information collected by each intelligent device; identifying at least one intelligent device as a target intelligent device based on the reward values of each intelligent device; and controlling the target intelligent device to perform a target operation. On one hand, this application adopts a centralized control approach to uniformly control multiple intelligent devices and determine which target intelligent device will perform the target operation, reducing the number of signaling interactions between different intelligent devices and improving the timeliness of multi-intelligent device responses. On the other hand, by constructing a learning model, the target intelligent device for performing the target operation is automatically selected based on environmental information, eliminating the need for user intervention, improving the response speed of intelligent devices, and reducing interference to the user from simultaneous responses by multiple intelligent devices. Furthermore, the determination of the target intelligent device in this application does not rely on user accounts, solving the problem of multiple accounts or unaccounted intelligent devices being unable to coordinate wake-up.
[0051] Optionally, according to some embodiments of this application, the control method of the intelligent device in the embodiments of this application can be applied to a multi-intelligent device central controller (hereinafter referred to as the central controller). The central controller executes the control method of the intelligent device in the embodiments of this application to uniformly control multiple intelligent devices, determine which target intelligent device will perform the target operation, thereby reducing the number of signaling interactions between different intelligent devices and improving the timeliness of multi-intelligent device response.
[0052] Based on the concept of centralized control, this application decouples information processing and decision-making from multiple devices. The intelligent device only needs to collect sound or images, without the need for signaling data to be exchanged between devices, thus reducing signaling overhead and device data processing consumption.
[0053] The specific implementation methods for each of the above steps are described below.
[0054] According to some embodiments of this application, optionally, S103, determining the reward value of each intelligent device based on a pre-established learning model and the first environmental information collected by each of the multiple intelligent devices, may include the following steps:
[0055] For any smart device, the reward value of the smart device is obtained by calculating the first environmental information collected by the smart device based on the reward function in the learning model.
[0056] Specifically, the initial environmental information collected by the smart device can be parameterized, transforming it into easily calculable numerical values. The learning model includes a reward function, which can be used to calculate the reward value for the smart device based on the initial environmental information collected.
[0057] According to some embodiments of this application, optionally, the first environmental information may include the sound and / or image of the target object. Exemplarily, the target object includes, but is not limited to, a user. For example, a smart device may collect sound via a microphone, and / or, a smart device may capture images of the surrounding environment via a camera.
[0058] Based on the sound and / or image of the target object, smart devices near the target object can be identified, i.e., smart devices that are relatively close to the target object. In some examples, smart devices near the target object can be used as target smart devices to perform targeted operations, such as issuing a notification message when a call comes in. This not only alerts the user to the incoming call but also prevents smart devices that are farther away from the user from being woken up, thereby reducing power consumption and preventing the user from having to walk a long distance to turn off smart devices that are far away from the user.
[0059] Based on the concept of centralized control, this application decouples information processing and decision-making from multiple devices. The intelligent device only needs to collect sound or images, without the need for signaling data to be exchanged between devices, thus reducing signaling overhead and device data processing consumption.
[0060] According to some embodiments of this application, optionally, the reward function may include a first reward value, a second reward value, an incentive function, a first weight, a second weight, and a third weight.
[0061] Figure 2 This is a schematic flowchart of step S103 in the control method for an intelligent device provided in an embodiment of this application. Figure 2 As shown, according to some embodiments of this application, optionally, S103, for any smart device, calculating the reward value of the smart device based on the reward function in the learning model on the first environmental information collected by the smart device, may include the following steps S201 and S202.
[0062] S201. For any smart device, calculate a first reward value based on the volume of the target object's sound collected by the smart device at the target time, and / or calculate a second reward value based on the detection result of the target object in the image collected by the smart device at the target time.
[0063] S202. Calculate the sum of the first product, the second product, and the third product to obtain the reward value of the intelligent device. The first product includes the product of the first weight and the first reward value, the second product includes the product of the second weight and the second reward value, and the third product includes the product of the third weight and the activation function.
[0064] In some examples, the reward function is expressed as follows:
[0065] R t (0,a) = α·R V + β·R H +γ·δ(s i -s target (1)
[0066]
[0067] R H = sgn(H t )if has human H > 0 ; else H < 0 (3)
[0068] Among them, s i Let R represent any smart device, and let R represent the smart device selected when the target input is received at time t. i The reward obtained at that time. α, β, and γ are three positive parameters, where α represents the first weight, β represents the second weight, and γ represents the third weight. R V Indicates smart device s i The reward value is obtained based on the sound of the target object at time t. The higher the volume of the target object's sound when it is detected, the greater the reward value. R H Indicates smart device s i The reward value obtained based on the image of the target object at time t. When the image of the target object (i.e., a human figure) is detected, R... H If it is 1, otherwise R H =-1. δ(s) i -s target ) represents the excitation function. sgn(x) represents the step function, where x is the input value, such as x = V. t Or H t When the input value x > 0, the output value is 1; when the input value x = 0, the output value is 0. t Indicates smart device s i The volume of the target object's sound, V, is collected at time t. max V represents the preset maximum volume. min This indicates the preset minimum volume value.
[0069] According to some embodiments of this application, optionally, before receiving the target input in S101, the learning model can be trained to obtain a trained learning model.
[0070] Figure 3 This is a schematic flowchart illustrating the training process in the control method for an intelligent device provided in an embodiment of this application. Figure 3 As shown, according to some embodiments of this application, optionally, before receiving the target input in S101, the control method of the smart device may further include the following steps S301 to S307.
[0071] S301. From multiple smart devices, a first smart device and a second smart device are selected, and a reward function is defined. The first smart device can be understood as the initial smart device, and the second smart device can be understood as the smart device to which the final switch is expected to occur. In S301, an initial reward function can be defined. In some examples, an initial action value table (Q-value table) can also be defined in S301. The Q-value table can include the maximum expected reward for each state, used to guide the best action to be taken in each state, such as guiding the switch from smart device A to smart device B.
[0072] S302, Select the target action.
[0073] For example, an action ε-greedy strategy is used to select the target action.
[0074] S303, In response to the target action, switch the first smart device to the third smart device.
[0075] S304. Obtain the second environmental information collected by the third intelligent device.
[0076] The third smart device can be any smart device other than the first smart device. For example, the second environmental information may include the sound and / or image of the target object. For example, the target object includes, but is not limited to, a user. For instance, the third smart device may collect sound via a microphone, and / or capture images of the surrounding environment via a camera.
[0077] S305. Calculate the reward value of the third intelligent device based on the reward function for the second environmental information.
[0078] For example, the reward value of the intelligent device can be calculated based on the reward function of the second environmental information collected by the third intelligent device, similar to the calculation of the reward value of the first environmental information mentioned above, as shown in expressions (1) to (3) above. In S304, the Q-value table can also be updated based on the reward value of the third intelligent device.
[0079] S306. Determine whether the third smart device is the second smart device.
[0080] S307. When the third intelligent device is not the second intelligent device, update the reward function, update the third intelligent device to the first intelligent device, and return to the step of selecting the target action, until the third intelligent device is the second intelligent device and the number of training iterations reaches the first preset threshold, thus obtaining the trained learning model. If the number of training iterations does not reach the first preset threshold, increment the number of training iterations by 1 and return to S302.
[0081] The size of the first preset number of times threshold can be flexibly adjusted according to the actual situation, and this application embodiment does not limit it.
[0082] According to some embodiments of this application, optionally, a maximum number of steps per training session can also be set. When the third intelligent device is not the second intelligent device, it is determined whether the number of steps in a single training session has reached the maximum number of steps in a single training session. If the number of steps in a single training session has not reached the maximum number of steps in a single training session, the reward function is updated, the third intelligent device is updated to the first intelligent device, and the process returns to S302. If the number of steps in a single training session has reached the maximum number of steps in a single training session, it is further determined whether the number of training sessions has reached a first preset threshold. When the number of training sessions reaches the first preset threshold, the training ends.
[0083] Optionally, according to some embodiments of this application, after S307, it can be determined whether the first smart device and the second smart device have traversed all smart devices. If not, S301 to S307 are executed.
[0084] Figure 4 This is another schematic flowchart illustrating the control method for a smart device provided in an embodiment of this application. For example... Figure 4 As shown, according to some embodiments of this application, optionally, after controlling the target smart device to perform the target operation, the control method of the smart device may further include the following steps:
[0085] S401. When it is detected that the user switches from the target smart device to another smart device to perform the target operation, the learning model is retrained.
[0086] When it is detected that a user switches from the target smart device to another smart device to perform the target operation, it could indicate that the user did not expect the target smart device to perform the target operation, or that the user expected another smart device to perform the target operation. In this case, the learning model can be retrained to improve the accuracy of its predictions.
[0087] In other words, when it is detected that a user switches from the target smart device to another smart device to perform the target operation, it is considered that the learning model is no longer applicable to the current environment. The central controller can reissue instructions, and the smart device can reacquire and upload environmental information to retrain the learning model.
[0088] Figure 5 This is another schematic flowchart illustrating the control method for an intelligent device provided in an embodiment of this application. For example... Figure 5 As shown, according to some embodiments of this application, optionally, the control method of the smart device may further include the following steps S501 and S502.
[0089] S501. Construct a smart device library containing multiple smart devices.
[0090] In S501, a smart device library containing multiple smart devices can be constructed based on existing smart devices. This is shown in expression (4); where represents the smart device library, s i Let i represent a smart device, where 0 ≤ i ≤ n, and i and n are integers.
[0091] S = {s0,s1,...,s} n} (4)
[0092] S502. When smart devices are added to or removed from the smart device library, the learning model is retrained.
[0093] When smart devices are added to or removed from the smart device library, it is considered that the learning model is no longer applicable to the current environment. The central controller can reissue instructions, and the smart devices can reacquire environmental information, upload it again, and retrain the learning model.
[0094] In this embodiment, the learning model is updated based on user behavior or device status. When external conditions such as the spatial distribution of multiple devices, the number of devices, or user behavior habits change, the device profile information contained in the original learning model becomes unavailable, and the dynamically responding devices no longer match the user's preferences, causing the user to actively change the device status. In this case, the smart device is re-controlled to acquire environmental information, and the learning model is retrained to ensure the model's accuracy.
[0095] According to some embodiments of this application, optionally, S104, determining at least one smart device as the target smart device based on the reward value of each smart device, may include the following steps:
[0096] Smart devices with reward values greater than or equal to a first preset threshold are identified as target smart devices.
[0097] The first preset threshold can be flexibly adjusted according to actual conditions, and this application embodiment does not limit it. Smart devices with higher reward values are generally closer to the target object (i.e., the user). Selecting a smart device with a higher reward value as the target smart device usually better meets the user's expectations and improves the user experience.
[0098] Figure 6 This is a schematic flowchart of step S104 in the control method for a smart device provided in an embodiment of this application. Figure 6 As shown, according to some embodiments of this application, optionally, step S104, determining at least one smart device as the target smart device based on the reward value of each smart device, may also include the steps S601 and S602.
[0099] S601, Obtain the working status of each smart device.
[0100] S602. The working smart devices whose reward value is less than the first preset threshold and whose reward value is greater than or equal to the second preset threshold are identified as target smart devices, where the second preset threshold is less than the first preset threshold.
[0101] When a smart device is active, it indicates that the user may be using it, such as watching television. In this case, if the smart device's reward value is less than a first preset threshold, but greater than or equal to a second preset threshold, the active smart device can still be identified as the target smart device to perform the target operation. For example, the target smart device can issue a notification message to maximize the chances of the user receiving it.
[0102] According to some embodiments of this application, optionally, step S105, controlling the target smart device to perform the target operation, may include the following steps:
[0103] Control the target smart device to send a prompt message.
[0104] For example, it is possible to control the target smart device to ring, vibrate, and / or issue prompts, etc., but this application embodiment does not limit this.
[0105] Based on the concept of centralized control, this application decouples information processing and decision-making from multiple devices. The intelligent device only needs to collect sound or images, without the need for signaling data to be exchanged between devices, thus reducing signaling overhead and device data processing consumption.
[0106] This application utilizes reinforcement learning technology to construct a learning model that automatically selects response devices based on the environment, eliminating the need for active user operation and improving device response speed. By incorporating human voice, image, and user behavior into the decision-making process for multi-device response events, the accuracy and intelligence of the response devices are enhanced, enabling intelligent devices to respond intelligently in accordance with changes in user behavior.
[0107] This application combines centralized control principles with reinforcement learning techniques to collect environmental information and train a learning model, making decisions based on the learned model. Multiple intelligent devices only need to be within the controller network and do not require logging into a unified account, enabling compatibility with a wider range of different device models and providing excellent scalability.
[0108] Based on the control method for intelligent devices provided in the above embodiments, this application also provides specific implementations of the control device for intelligent devices. Please refer to the following embodiments.
[0109] Figure 7 This is a schematic diagram of a control device for a smart device provided in an embodiment of this application. Figure 7 As shown, the control device 70 for the smart device provided in this application embodiment may include the following modules:
[0110] Receiver module 701 is used to receive target input;
[0111] The data acquisition module 702 is used to acquire the first environmental information collected by multiple smart devices in response to the target input;
[0112] The learning module 703 is used to determine the reward value of each intelligent device based on the pre-established learning model and the first environmental information collected by each of the multiple intelligent devices, and to identify at least one intelligent device as the target intelligent device based on the reward value of each intelligent device.
[0113] The instruction issuing module 704 is used to control the target smart device to perform the target operation.
[0114] The intelligent device provided in this application adopts a centralized control approach to uniformly control multiple intelligent devices, determining which target intelligent device should execute the target operation. This reduces the amount of signaling interaction between different intelligent devices and improves the timeliness of multi-intelligent device response. On the other hand, by constructing a learning model, the target intelligent device for executing the target operation is automatically selected based on environmental information, without requiring active user operation, thus improving the response speed of the intelligent device and reducing the interference to the user caused by multiple intelligent devices responding simultaneously. Furthermore, the determination of the target intelligent device in this application does not rely on user accounts, which can solve the problem of multiple accounts or unaccounted intelligent devices being unable to coordinate wake-up.
[0115] According to some embodiments of this application, optionally, the control device 70 of the smart device may further include a data processing module for parameterizing the first environmental information collected by the smart device, such as converting the first environmental information into numerical values that are easy to calculate.
[0116] According to some embodiments of this application, optionally, the learning module 703 is specifically used to calculate the reward value of the smart device based on the reward function in the learning model for any smart device, using the first environmental information collected by the smart device.
[0117] According to some embodiments of this application, optionally, the first environmental information includes the sound and / or image of the target object.
[0118] According to some embodiments of this application, optionally, the reward function includes a first reward value, a second reward value, an activation function, a first weight, a second weight, and a third weight. Specifically, the learning module 703 is used to calculate, for any given smart device, a first reward value based on the volume of the sound of the target object collected by the smart device at a target time, and / or, a second reward value based on the detection result of the target object in the image collected by the smart device at the target time; and to calculate the sum of the first product, the second product, and the third product to obtain the reward value of the smart device, wherein the first product includes the product of the first weight and the first reward value, the second product includes the product of the second weight and the second reward value, and the third product includes the product of the third weight and the activation function.
[0119] According to some embodiments of this application, optionally, the control device 70 of the smart device may further include a training module, used to set a first smart device and a second smart device from multiple smart devices, and set a reward function; select a target action; in response to the target action, switch the first smart device to a third smart device; acquire second environmental information collected by the third smart device; calculate the reward value of the third smart device based on the reward function; determine whether the third smart device is the second smart device; when the third smart device is not the second smart device, update the reward function, update the third smart device to the first smart device, and return to the step of selecting the target action, until the third smart device is the second smart device and the number of training times reaches a first preset threshold, thereby obtaining a trained learning model.
[0120] According to some embodiments of this application, optionally, the training module can also be used to retrain the learning model when it is detected that the user switches from the target smart device to another smart device to perform the target operation.
[0121] According to some embodiments of this application, optionally, the training module can also be used to construct a smart device library containing multiple smart devices; when smart devices are added or removed from the smart device library, the learning model is retrained.
[0122] According to some embodiments of this application, optionally, the learning module 703 is specifically used to identify smart devices whose reward value is greater than or equal to a first preset threshold as target smart devices.
[0123] According to some embodiments of this application, optionally, the learning module 703 is specifically used to obtain the working status of each smart device; and to determine the working smart device whose reward value is less than a first preset threshold and whose reward value is greater than or equal to a second preset threshold as the target smart device, wherein the second preset threshold is less than the first preset threshold.
[0124] According to some embodiments of this application, optionally, the instruction issuing module 704 is specifically used to control the target smart device to issue prompt information.
[0125] Figure 7 Each module / unit in the device shown has the function of implementing each step in the control method of the intelligent device provided in the above method embodiment, and can achieve its corresponding technical effect. For the sake of brevity, it will not be described in detail here.
[0126] Based on the control method for intelligent devices provided in the above embodiments, this application also provides specific implementation methods for electronic devices. Please refer to the following embodiments.
[0127] Figure 8 A schematic diagram of the hardware structure of the electronic device provided in an embodiment of this application is shown.
[0128] Electronic devices may include a processor 801 and a memory 802 storing computer program instructions.
[0129] Specifically, the processor 801 may include a central processing unit (CPU), an application specific integrated circuit (ASIC), or one or more integrated circuits that can be configured to implement the embodiments of this application.
[0130] Memory 802 may include mass storage for data or instructions. For example, and not limitingly, memory 802 may include a hard disk drive (HDD), floppy disk drive, flash memory, optical disk, magneto-optical disk, magnetic tape, or Universal Serial Bus (USB) drive, or a combination of two or more of these. In one example, memory 802 may include removable or non-removable (or fixed) media, or memory 802 may be a non-volatile solid-state memory. Memory 802 may be internal or external to an electronic device.
[0131] In one example, memory 802 may be read-only memory (ROM). In one example, the ROM may be a mask-programmed ROM, a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (EEPROM), an electrically rewritable ROM (EAROM), or flash memory, or a combination of two or more of these.
[0132] Memory 802 may include read-only memory (ROM), random access memory (RAM), disk storage media device, optical storage media device, flash memory device, electrical, optical, or other physical / tangible memory storage device. Therefore, typically, memory includes one or more tangible (non-transitory) computer-readable storage media (e.g., memory devices) encoded with software including computer-executable instructions, and when the software is executed (e.g., by one or more processors), it is operable to perform the operations described with reference to the method according to one aspect of this application.
[0133] The processor 801 reads and executes the computer program instructions stored in the memory 802 to implement the methods / steps in the above method embodiments and achieve the corresponding technical effects achieved by the method embodiments in executing their methods / steps. For the sake of brevity, these details will not be repeated here.
[0134] In one example, the electronic device may also include a communication interface 803 and a bus 810. For example, Figure 8 As shown, the processor 801, memory 802, and communication interface 803 are connected through bus 810 and complete communication with each other.
[0135] The communication interface 803 is mainly used to realize communication between various modules, devices, units and / or equipment in the embodiments of this application.
[0136] Bus 810 includes hardware, software, or both, that couples components of an electronic device together. For example, and not limitingly, the bus may include an Accelerated Graphics Port (AGP) or other graphics bus, an Extended Industry Standard Architecture (EISA) bus, a Front Side Bus (FSB), a Hyper Transport (HT) interconnect, an Industry Standard Architecture (ISA) bus, an Infinite Bandwidth Interconnect, a Low Pin Count (LPC) bus, a memory bus, a Microchannel Architecture (MCA) bus, a Peripheral Component Interconnect (PCI) bus, a PCI-Express (PCI-X) bus, a Serial Advanced Technology Attachment (SATA) bus, a Video Electronics Standards Association Local (VLB) bus, or other suitable buses, or combinations of two or more of these. Where appropriate, bus 810 may include one or more buses. Although specific buses are described and illustrated in embodiments of this application, this application contemplates any suitable bus or interconnect.
[0137] Furthermore, in conjunction with the control methods for intelligent devices in the above embodiments, this application embodiment can provide a computer-readable storage medium for implementation. This computer-readable storage medium stores computer program instructions; when these computer program instructions are executed by a processor, they implement any of the control methods for intelligent devices in the above embodiments. Examples of computer-readable storage media include non-transitory computer-readable storage media, such as electronic circuits, semiconductor memory devices, ROM, random access memory, flash memory, erasable ROM (EROM), floppy disks, CD-ROMs, optical disks, and hard disks.
[0138] It should be clarified that this application is not limited to the specific configurations and processes described above and shown in the figures. For the sake of brevity, detailed descriptions of known methods are omitted here. In the above embodiments, several specific steps are described and shown as examples. However, the method process of this application is not limited to the specific steps described and shown. Those skilled in the art can make various changes, modifications, and additions, or change the order of steps, after understanding the spirit of this application.
[0139] The functional blocks shown in the above-described block diagram can be implemented as hardware, software, firmware, or a combination thereof. When implemented in hardware, they can be, for example, electronic circuits, application-specific integrated circuits (ASICs), appropriate firmware, plug-ins, function cards, etc. When implemented in software, the elements of this application are programs or code segments used to perform the required tasks. Programs or code segments can be stored on a machine-readable medium or transmitted over a transmission medium or communication link via data signals carried on a carrier wave. "Machine-readable medium" can include any medium capable of storing or transmitting information. Examples of machine-readable media include electronic circuits, semiconductor memory devices, ROM, flash memory, erasable ROM (EROM), floppy disks, CD-ROMs, optical disks, hard disks, fiber optic media, radio frequency (RF) links, etc. Code segments can be downloaded via computer networks such as the Internet, intranets, etc.
[0140] It should also be noted that the exemplary embodiments mentioned in this application describe methods or systems based on a series of steps or apparatus. However, this application is not limited to the order of the above steps; that is, the steps can be performed in the order mentioned in the embodiments, or in a different order, or several steps can be performed simultaneously.
[0141] The aspects of this application have been described above with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It should be understood that each block in the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that these instructions, executable via the processor of the computer or other programmable data processing apparatus, enable the implementation of the functions / actions specified in one or more blocks of the flowchart illustrations and / or block diagrams. Such a processor can be, but is not limited to, a general-purpose processor, a special-purpose processor, a special application processor, or a field-programmable logic circuit. It is also understood that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can also be implemented by dedicated hardware performing the specified functions or actions, or can be implemented by a combination of dedicated hardware and computer instructions.
[0142] The above description is merely a specific implementation of this application. Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, modules, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here. It should be understood that the protection scope of this application is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in this application, and these modifications or substitutions should all be covered within the protection scope of this application.
Claims
1. A control method for an intelligent device, characterized in that, include: Receive target input; In response to the target input, first environmental information collected by each of the multiple smart devices is acquired; Based on the pre-established learning model and the first environmental information collected by multiple smart devices, the reward value of each smart device is determined. Based on the reward value of each smart device, at least one of the smart devices is identified as the target smart device; Control the target intelligent device to perform the target operation; The step of determining the reward value for each intelligent device based on a pre-established learning model and the first environmental information collected by each of the multiple intelligent devices includes: For any one of the smart devices, the reward value of the smart device is obtained by calculating the first environmental information collected by the smart device based on the reward function in the learning model; Before receiving the target input, the method further includes: From the plurality of smart devices, a first smart device and a second smart device are selected, and a reward function is defined; Select the target action; In response to the target action, the first smart device is switched to a third smart device; Obtain the second environmental information collected by the third intelligent device; The reward value of the third intelligent device is obtained by calculating the second environmental information based on the reward function; Determine whether the third smart device is the second smart device; When the third intelligent device is not the second intelligent device, update the reward function, update the third intelligent device to the first intelligent device, and return to the step of selecting the target action, until the third intelligent device is the second intelligent device and the number of training times reaches the first preset threshold, and obtain the trained learning model.
2. The method according to claim 1, characterized in that, The first environmental information includes the sound and / or image of the target object.
3. The method according to claim 2, characterized in that, The reward function includes a first reward value, a second reward value, an incentive function, a first weight, a second weight, and a third weight; For any one of the smart devices, the reward value of the smart device is calculated based on the reward function in the learning model using the first environmental information collected by the smart device, including: For any one of the smart devices, the first reward value is calculated based on the volume of the sound of the target object collected by the smart device at the target time, and / or the second reward value is calculated based on the detection result of the target object in the image collected by the smart device at the target time; The sum of the first product, the second product, and the third product is calculated to obtain the reward value of the intelligent device. The first product includes the product of the first weight and the first reward value, the second product includes the product of the second weight and the second reward value, and the third product includes the product of the third weight and the activation function.
4. The method according to claim 1, characterized in that, After controlling the target smart device to perform the target operation, the method further includes: When it is detected that the user switches from the target smart device to another smart device to perform the target operation, the learning model is retrained.
5. The method according to claim 1, characterized in that, Also includes: Construct a smart device library that includes the aforementioned smart devices; When smart devices are added to or removed from the smart device library, the learning model is retrained.
6. The method according to claim 1, characterized in that, The step of identifying at least one of the smart devices as the target smart device based on the reward value of each smart device includes: The smart devices whose reward value is greater than or equal to a first preset threshold are identified as the target smart devices.
7. The method according to claim 1, characterized in that, The step of identifying at least one of the smart devices as the target smart device based on the reward value of each smart device also includes: Obtain the working status of each of the aforementioned smart devices; The smart devices that are currently in operation and whose reward value is less than the first preset threshold and whose reward value is greater than or equal to the second preset threshold are identified as the target smart devices, where the second preset threshold is less than the first preset threshold.
8. The method according to claim 1, characterized in that, The control of the target smart device to perform the target operation includes: Control the target smart device to issue a prompt message.
9. A control device for an intelligent device, characterized in that, include: The receiving module is used to receive target input; The data acquisition module is used to respond to the target input and acquire the first environmental information collected by each of the multiple smart devices. The learning module is used to determine the reward value of each intelligent device based on a pre-established learning model and the first environmental information collected by each of the multiple intelligent devices, and to identify at least one of the intelligent devices as the target intelligent device based on the reward value of each intelligent device. The instruction issuing module is used to control the target smart device to perform the target operation; The step of determining the reward value for each intelligent device based on a pre-established learning model and the first environmental information collected by each of the multiple intelligent devices includes: For any one of the smart devices, the reward value of the smart device is obtained by calculating the first environmental information collected by the smart device based on the reward function in the learning model; Before receiving the target input, the method further includes: From the plurality of smart devices, a first smart device and a second smart device are selected, and a reward function is defined; Select the target action; In response to the target action, the first smart device is switched to a third smart device; Obtain the second environmental information collected by the third intelligent device; The reward value of the third intelligent device is obtained by calculating the second environmental information based on the reward function; Determine whether the third smart device is the second smart device; When the third intelligent device is not the second intelligent device, update the reward function, update the third intelligent device to the first intelligent device, and return to the step of selecting the target action, until the third intelligent device is the second intelligent device and the number of training times reaches the first preset threshold, and obtain the trained learning model.
10. An electronic device, characterized in that, The electronic device includes: a processor, a memory, and a computer program stored in the memory and executable on the processor, wherein the computer program, when executed by the processor, implements the steps of the control method for the intelligent device as described in any one of claims 1 to 8.
11. A computer-readable storage medium, characterized in that, A computer program is stored on the computer-readable storage medium, which, when executed by a processor, implements the steps of the control method for the intelligent device as described in any one of claims 1 to 8.
Citation Information
Patent Citations
Terminal device, awakening method and device thereof and computer readable storage medium
CN113096658A