Smart home reinforcement learning interaction method and system, terminal equipment and storage medium

Through the smart home reinforcement learning interaction method, the AI model is used to automatically control the device, combined with environmental and behavioral data, the problem of traditional smart home interaction is solved, the device's autonomous adjustment and user feedback optimization are realized, and the degree of intelligence and interactive experience are improved.

CN120353144APending Publication Date: 2025-07-22ULTIMATE IOT (HENAN) TECHNOLOGY LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510710845.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-29
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

The traditional smart home interaction method is not smart enough, and requires users to actively set and adjust, which cannot improve the interactive experience.

Method used

The sensors obtain environmental data and behavioral data of smart home devices, input pre-trained AI models, generate behavioral decisions, and automatically control device operations based on the decisions, and iterative optimization is performed through user feedback.

Benefits of technology

Reduce user operation frequency, improve the intelligence of smart home devices, adapt to family life habits, and enhance interactive experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120353144A_ABST
    Figure CN120353144A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of artificial intelligence, in particular to a smart home reinforcement learning interaction method, and the method comprises the steps: obtaining environment data through a sensor; acquiring behavior data of each smart home device; inputting the environment data and the behavior data into a pre-trained AI model to obtain a behavior decision output by the AI model; and according to the behavior decision, controlling the corresponding smart home device to execute a corresponding task. Through the automatic identification and control operation of the AI model, the interaction demand of the user is reduced, and the intelligent degree of the smart home equipment is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and particularly to a smart home reinforcement learning interaction method, system, terminal device, and storage medium. Background Art

[0002] In the traditional smart home field, interactions mainly rely on physical buttons, sensor triggers, voice, gestures, etc. However, these interaction devices are still not convenient enough and lack intelligence. In some scenarios, users still need to actively set and adjust, and cannot improve the interaction experience from the aspects of users and the environmental ecosystem, making users unable to feel the sense of intelligence during the use of smart homes. Summary of the Invention

[0003] In view of this, embodiments of this application provide a smart home reinforcement learning interaction method, which can effectively solve problems such as lack of intelligence.

[0004] In a first aspect, embodiments of this application provide a smart home reinforcement learning interaction method, including: Obtaining environmental data through sensors; Obtaining the behavior data of each smart home device; Inputting the environmental data and the behavior data into a pre-trained AI model to obtain a behavior decision output by the AI model; Controlling the corresponding smart home device to execute corresponding tasks according to the behavior decision.

[0005] In some embodiments, the method further includes: When receiving an active interaction event from the user, obtaining the environmental data at the current moment and the status data of the device corresponding to the interaction event; Using the active interaction event, the environmental data at the previous moment, and the status data as reward factors, and feeding the reward factors back into the AI model to iteratively optimize the AI model.

[0006] In some embodiments, the step of controlling the corresponding smart home device to execute corresponding tasks according to the behavior decision includes: Prompting the user about the behavior decision through voice broadcast, and when receiving the consent instruction issued by the user, controlling the corresponding smart home device to execute the behavior decision.

[0007] In some embodiments, when receiving the consent instruction issued by the user, it further includes: Using the consent instruction as positive feedback and feeding it back into the AI model to iteratively optimize the AI model.

[0008] In some embodiments, after prompting the user of the behavior decision by voice broadcast, it further includes: When receiving the disagreement instruction issued by the user, do not execute the behavior decision, and use the disagreement instruction as negative feedback and feedback it into the AI model to iteratively optimize the AI model.

[0009] In a second aspect, the present application further provides a smart home reinforcement learning interaction system, including a plurality of smart home devices, sensors, and a control terminal; The control terminal is used to obtain environmental data through the sensors and obtain corresponding behavior data through each smart home device; Input the environmental data and the behavior data into a pre-trained AI model in the control terminal to obtain a behavior decision output by the AI model; The control terminal sends a control command to the corresponding smart home device according to the behavior decision to execute the behavior decision.

[0010] In some embodiments, it includes: When the control terminal receives an active interaction event from the user, obtain the environmental data at the current moment and the status data of the device corresponding to the interaction event; Use the active interaction event, the environmental data at the previous moment, and the status data as reward factors, and feedback the reward factors into the AI model to iteratively optimize the AI model.

[0011] In some embodiments, the sending a control command to the corresponding smart home device according to the behavior decision to execute the behavior decision includes: The control terminal prompts the user of the behavior decision by voice broadcast. When receiving the consent instruction issued by the user, send a control instruction to the corresponding smart home device to execute the behavior decision.

[0012] In a third aspect, the present application further provides a terminal device, the terminal device includes a processor and a memory, the memory stores a computer program, and the processor is used to execute the computer program to implement the smart home reinforcement learning interaction method described above.

[0013] In a fourth aspect, the present application further provides a readable storage medium, which stores a computer program, and when the computer program is executed on a processor, it implements the smart home reinforcement learning interaction method described above.

[0014] The embodiments of the present application have the following beneficial effects: This application uses an AI model to identify the acquired environmental data and the behavior data of smart home devices, automatically generate behavior decisions, and control the smart home devices to automatically execute corresponding tasks, thereby reducing the user's manipulation frequency. Through the automatic identification and control operations of the AI model, the user's interaction requirements are reduced, and the intelligence level of the smart home devices is improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] To more clearly illustrate the technical solutions of the embodiments of this application, the drawings required for use in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of this application, and thus should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can also be obtained based on these drawings without creative efforts.

[0016] Figure 1 Shows a schematic flowchart of a smart home reinforcement learning interaction method according to an embodiment of this application; Figure 2 Shows a schematic architecture diagram of a smart home reinforcement learning interaction system according to an embodiment of this application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0017] The technical solutions in the embodiments of this application will be clearly and completely described below with reference to the drawings in the embodiments of this application. Obviously, the described embodiments are only some of the embodiments of this application, rather than all of the embodiments.

[0018] The components of the embodiments of this application usually described and shown in the drawings here can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of this application provided in the drawings is not intended to limit the scope of the claimed application, but only represents the selected embodiments of this application. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of this application.

[0019] In the following text, the terms "including", "having" and their cognates that can be used in various embodiments of this application are only intended to indicate specific features, numbers, steps, operations, elements, components or combinations of the foregoing items, and should not be understood as first excluding the existence or adding the possibility of one or more other features, numbers, steps, operations, elements, components or combinations of the foregoing items. In addition, the terms "first", "second", "third", etc. are only used for differentiating descriptions and cannot be understood as indicating or implying relative importance.

[0020] Unless otherwise defined, all terms used herein (including technical and scientific terms) have the same meaning as commonly understood by one of ordinary skill in the art to which the various embodiments of this application belong. The terms (such as those defined in a commonly used dictionary) will be interpreted as having the same meaning as the contextual meaning in the relevant technical field and will not be interpreted as having an idealized meaning or an overly formal meaning, unless clearly defined in the various embodiments of this application.

[0021] The following will, with reference to the accompanying drawings, elaborate on some embodiments of this application. Without conflict, the following embodiments and the features in the embodiments can be combined with each other.

[0022] The current smart home still requires users to perform interactive control through physical buttons, sensor triggers, voice, gestures, etc. However, there is a sense of fragmentation among these interactive devices, and it cannot improve the interactive experience from the aspects of users and the environmental ecosystem. This application provides a smart home reinforcement learning interaction method, which obtains environmental data and behavior data through smart home devices and sensors, and then performs intelligent automatic control through an AI model, reducing user operations and improving the intelligence level.

[0023] The following will illustrate this smart home reinforcement learning interaction method with some specific embodiments.

[0024] Figure 1 A flowchart showing an embodiment of the smart home reinforcement learning interaction method of this application is presented. Exemplarily, this smart home reinforcement learning interaction method includes the following steps: Step S100, obtaining environmental data through sensors.

[0025] Environmental data refers to data such as temperature, humidity, time, light, etc. that can represent the indoor environment. These data can be obtained by relying on various sensors, such as thermometers, hygrometers, electronic clocks, etc. Some smart home devices are equipped with these sensors, so these data can be obtained by relying on the sensors of smart home devices.

[0026] Sensors can also be monitoring devices such as cameras and microphones. Currently, cameras are installed in most homes to monitor the interior, and the captured images and recorded sounds can also reflect the current environment. For example, if a user only installs a camera in the living room, the obtained image data can intuitively feedback the environmental state of the living room, such as whether there is anyone and the lighting conditions.

[0027] Environmental data can mainly be used to judge the differences in indoor environmental areas. For example, there will be the sounds of ignition and range hoods in the kitchen, and the temperature will be relatively high. In the bedroom, there are often only voices or the sounds of opening and closing doors, etc. Different environmental data can help locate the layout of the room and the current state to assist the subsequent intelligent operation process.

[0028] Step S200: Obtain the behavior data of each smart home device.

[0029] There are many types of smart home devices, such as air conditioners, curtains, door locks, etc. The variety of smart home devices is related to the user's usage situation. However, according to the actual family situation, the number and types of smart home devices equipped in each family are different. In this embodiment, it is not required that all devices in the home are smart home devices. There can be some smart home devices and some non-smart home devices. In this embodiment, the control operations related to smart home devices are limited to the existing smart home devices, and there is no limit on their number and types.

[0030] For smart home devices, there will be corresponding behavior data. These behavior data can be automatically triggered by the system. For example, for an air conditioner, when it determines that the current temperature of 30 degrees is too high, it starts the cooling mode and sets the temperature to 20 degrees. This is a piece of behavior data.

[0031] Another example is that at 8:30 in the morning, the bedroom curtains automatically open. This is also a piece of behavior data. Obtaining behavior data can monitor the status information of each smart home device.

[0032] At the same time, as mentioned above, different home environments have different types and numbers of smart home devices. Therefore, in this embodiment, there is no fixed and rigid requirement for the collected behavior data and environmental data. That is, it is not necessary to collect data such as temperature and humidity. If the corresponding sensors or smart home devices do not exist, only the data that can be collected can be collected, and the default processing can be performed on other uncollectable data. It can be understood that if the air conditioner is a smart home device, environmental data such as temperature and humidity useful for adjusting the air conditioner can be obtained according to the sensors in the air conditioner. If the curtain is a smart home device, there are sensors such as a clock or a photosensitive sensor to obtain environmental data such as brightness or time. The corresponding smart home devices are basically equipped with corresponding sensors to obtain the required environmental data, and the currently available data is sufficient for controlling the operations of various smart home devices.

[0033] It should be noted that not only the operations performed after the smart home device receives control are considered behavior data. For most smart home devices, they will always be in a certain behavior mode. For example, the refrigerator is in the freezing or refrigerating operation at a certain set temperature, and the smart air conditioner is in the cooling operation at a certain temperature. These operations remain unchanged for a long time and seem static, but these are also behavior data.

[0034] Step S300: Input the environmental data and the behavior data into a pre-trained AI model to obtain the behavior decision output by the AI model.

[0035] The above environmental data and behavior data are input into a pre-trained AI model. The AI model will output behavior decisions based on these two types of data. For example, when the current temperature is 30 degrees and the air conditioner is set to cool at 20 degrees, but 20 degrees is too cold, then adjust 20 degrees to 24 degrees.

[0036] Among them, the AI model can be a neural network model, and its training data can be the execution data of smart home devices, environmental data, and interaction data of human-computer interaction. During training, these data are labeled, and the corresponding home scenarios for each data are determined, and then used as training data to be input into the AI model for inference training, so that the AI model can output the interaction operations desired by the user according to the execution data and environmental data, so as to automatically adjust in advance and liberate the user's operations.

[0037] For example, when labeling a certain set of environmental data, device execution data, and interaction data, it can be specified which venue in the home space this set of data is in. This venue can be environments such as the living room, master bedroom, secondary bedroom, kitchen, etc. For different venues, the operation decisions that the AI model should output will definitely be different. Therefore, the specific venue also needs to be labeled into each set of training data. In addition to the environment, the execution data and interaction data also need to be labeled. For example, label the corresponding execution weights for the execution data and interaction data, so that the AI model can learn which environmental data corresponds to which scenarios and what the corresponding better processing measures are.

[0038] For example, a set of training data includes environmental data, interaction data, and device behavior data. At the same time, the weight of the interaction data is greater than the device behavior data, indicating that under the environmental data of this set of data, the interaction data is a better choice, and the device behavior data is not a recommended choice. Another example is that a set of training data only includes environmental data and interaction data, which means that in this environmental data, it is recommended to process it with this interaction data. Another example is that a set of training data includes environmental data, interaction data, and device behavior data. At the same time, the weight of the interaction data is less than the device behavior data, which means that in this case of environmental data, the device behavior data is superior to the interaction data.

[0039] It should be noted that there are many different types of data in the above-mentioned execution data, environmental data, and interaction data of manual interaction. For different smart home devices, the types of data they care about are actually different, and the execution data will also be different due to different devices. In different scenarios, the types of environmental data are also different. Therefore, in this embodiment, when training the AI model, it can be trained separately based on different scenarios and different devices. By gradually adding new training data, the scenarios that the AI model can recognize and process are gradually expanded, and finally the AI model can handle behavioral decision operations in multiple scenarios.

[0040] Step S400: According to the behavior decision, control the corresponding smart home device to perform the corresponding task.

[0041] Behavioral decisions can be controlling a smart home device to make corresponding adjustments, such as controlling the temperature of the air conditioner, or controlling the opening and closing of curtains, or controlling the opening and closing of lights. For specific behavioral decisions, control commands need to be sent to the corresponding smart home device for control.

[0042] For example, if the current time is 8:30 in the morning and the user needs to be woken up, a curtain opening control command will be sent to the curtains to automatically open the curtains. If an alarm is set, the alarm will be turned on. If the user's voice command "stop" is received, the scene the user is in, such as the master bedroom or the second bedroom, can be determined based on the sound source. At this time, the alarm in the corresponding scene can be turned off and the curtains in the corresponding scene can be turned off.

[0043] Among them, before making a behavioral decision, the user can be prompted with voice or text about the current behavioral decision to confirm with the user whether it is okay to do so. When a consent instruction from the user is received, the corresponding smart home device is controlled to execute the behavioral decision, and the consent instruction is used as positive feedback and fed back to the AI model to iteratively optimize the AI model.

[0044] Similarly, if a disagreement instruction is received from the user, the behavioral decision is not executed, and the disagreement instruction is fed back to the AI model as negative feedback to iteratively optimize the AI model.

[0045] During iterative optimization, the environmental data, device behavior data, and user interaction data of this interaction process can be sent to the AI model as training data, as in the aforementioned training process, and the weight of the interaction data is set to be greater than the device behavior data. The specific weight value can be determined based on the time when the interaction behavior occurs.

[0046] The above two feedback methods can enhance the accuracy of the behavioral decisions made by the AI model. It can be understood that different families have different living schedules and habits. Therefore, it is difficult for smart home device manufacturers to set a set of parameters that everyone is used to and agrees with. By receiving the user's feedback on the automatically generated behavioral decisions, the AI model can also perform self-learning and reinforcement after deployment, so that the decisions made by the AI model become more and more in line with the user's habits.

[0047] In addition, the AI model of this embodiment will also obtain the active interaction events performed by the user. For example, in some scenarios, the AI model does not make a behavioral decision, but the user actively controls the working state of the smart home device through some control operations. This kind of operation is outside the AI model mode. When the user performs such an operation, the AI model needs to learn this operation. Therefore, at this time, the environmental data at the current moment and the state data of the device corresponding to the interaction event will be obtained; the active interaction event, the environmental data at the previous moment, and the state data will be used as reward factors, and the reward factors will be fed back into the AI model to perform iterative optimization on the AI model.

[0048] For example, when the room temperature is 20 degrees and the air conditioner's cooling temperature is 20 degrees, the user actively raises the air conditioner temperature to 25 degrees. At this time, this active interaction event will be obtained by the AI model, and environmental data such as the current room temperature, humidity, and even data such as season and time can also be obtained. After obtaining these data, they can be used as training data or iterative data to further perform iterative learning on the AI model in accordance with the model training method.

[0049] When performing iterative optimization, iterative optimization can be carried out through methods such as gradient descent, attention enhancement, and value function correction. In a possible scenario, the user is currently in the living room and issues a voice prompt to inform the user to lower the air conditioner temperature. When the voice prompt is issued, the user moves from the living room to the master bedroom and issues an approval instruction. At this time, the user's location can be determined according to the sound source. Although the voice prompt was issued in the living room, the voice command was received in the master bedroom. Therefore, at this time, the AI model will output an operation decision to lower the air conditioner temperature in the master bedroom.

[0050] However, in this scenario, the AI may also output a decision to lower the air conditioner in the living room. In this case, the user is very likely to actively lower the temperature of the air conditioner in the master bedroom again through the remote control or voice command. At this time, this interaction instruction will also be obtained and fed back to the AI model. The AI model will then know that in the previous situation, it should adjust the air conditioner in the master bedroom instead of the one in the living room.

[0051] Among them, the result of iterative learning can be operations such as enhancing the execution weight of the operation of raising the room temperature under such environmental data. By adjusting the parameters of the AI model, the AI model can automatically execute the operations that the user wants to perform one step ahead of the user at an appropriate time, thus becoming more and more intelligent.

[0052] The smart home reinforcement learning interaction method of this embodiment enables smart home devices to be uniformly controlled by an AI model. The AI model can continuously iterate itself based on training data and learning data in daily life, so that it is more suitable for the living habits of the current family. As the AI model continues to iterate, for users, they increasingly do not need to actively control the operation of smart home devices, because the AI model will automatically adjust at an appropriate time, allowing users to be in a habitual environment.

[0053] Figure 2 FIG. shows a schematic structural diagram of a smart home reinforcement learning interaction system according to an embodiment of the present application. Exemplarily, the smart home reinforcement learning interaction system includes: A plurality of smart home devices 100, sensors 200, and a control terminal 300.

[0054] Among them, the control terminal 300 is used to obtain environmental data through the sensor 200 and obtain corresponding behavior data through each smart home device 100. The control terminal 300 can be a server located in the cloud or a gateway device deployed on-site, such as a smart speaker or other devices that can be connected to other smart home devices 100.

[0055] The sensor 200 can be some dedicated sensors separately set at home or sensors built into the smart home device 100. The smart home device 100 is a device that connects various devices (such as lighting, electrical appliances, security, etc.) in the home through Internet of Things technology, network communication technology, and automation control technology.

[0056] Each smart home device 100 can be controlled by the user according to its own unique control method or by the control terminal 300.

[0057] The control terminal 300 is internally equipped with an AI model, and can input the environmental data and the behavior data into the AI model, and obtain the behavior decision output by the AI model.

[0058] When these behavior decisions need to be executed, the control terminal 300 will send control commands to the corresponding smart home device 100 to execute the behavior decisions.

[0059] It can be understood that the control terminal 300 has the ability to send control commands to all smart home devices 100 and also has the ability to collect data generated by all smart home devices 100.

[0060] Among them, when the control terminal 300 receives an active interaction event from the user, it obtains the environmental data at the current moment and the status data of the device corresponding to the interaction event.

[0061] Taking the active interaction event, the environmental data at the previous moment, and the status data as reward factors, and feeding the reward factors back into the AI model to iteratively optimize the AI model.

[0062] The iterative optimization operation in this part is similar to that in the foregoing embodiments and will not be elaborated here.

[0063] The control terminal 300 can also prompt the user about the behavior decision by voice broadcast. When receiving the consent instruction issued by the user, it sends a control instruction to the corresponding smart home device to execute the behavior decision.

[0064] It can be understood that the system in this embodiment corresponds to the smart home reinforcement learning interaction method in the above embodiment. The optional items in the above embodiment also apply to this embodiment, so they will not be repeated here.

[0065] It can be understood that different families have different types and quantities of smart home devices. When the control terminal 300 connects to the network of smart home devices, it can detect and determine how many smart home devices there are in the current family. At this time, it can automatically adjust the weights to ensure the accuracy of its own recognition.

[0066] For example, during training, the types of input data are 100, and these 100 types of data come from 20 types of smart home devices and sensors. However, there are only 5 types of smart home devices in a certain family. No matter how data is obtained, these 100 types of data cannot be obtained. That is, there will be a deviation between the input in the actual scenario and the input during training. At this time, the control terminal 300 can determine which types of input data there are according to the 5 types of smart home devices currently present. For example, if there are 20 types of input, it can adjust the weights of these 20 types of input data so that when the AI model performs inference, it only considers these 20 types of data and does not consider the other 80 types of data. In this way, the AI model can correctly identify the current scenario and output a reasonable decision.

[0067] The smart home reinforcement learning interaction system of this embodiment collects various data of sensors and smart home devices in the system through a control terminal, and performs intelligent analysis through a built-in AI model to obtain an analyzed behavior decision, and then performs corresponding control operations according to the behavior decision to improve the intelligence of the operations of all smart home devices in the entire system, making the entire system more and more intelligent and more adaptable to the current home, thereby reducing the number of user interactions.

[0068] This application also provides a terminal device. Exemplarily, the terminal device includes a processor and a memory. Among them, the memory stores a computer program, and the processor runs the computer program, so that the terminal device executes the above-mentioned smart home reinforcement learning interaction method or the functions of each module in the above-mentioned smart home reinforcement learning interaction system.

[0069] Among them, the processor can be an integrated circuit chip with signal processing capabilities. The processor can be a general-purpose processor, including at least one of a central processing unit (CPU), a graphics processing unit (GPU), a network processor (NP), a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, and discrete hardware components. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc., and can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present application.

[0070] The memory can be, but is not limited to, a random access memory (RAM), a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), etc. Among them, the memory is used to store a computer program, and after receiving an execution instruction, the processor can execute the computer program accordingly.

[0071] This application also provides a readable storage medium for storing the computer program used in the above terminal device.

[0072] In several embodiments provided in this application, it should be understood that the disclosed devices and methods can also be implemented in other ways. The device embodiments described above are merely illustrative. For example, the flowcharts and structure diagrams in the accompanying drawings show the possible architectures, functions, and operations of the devices, methods, and computer program products according to multiple embodiments of this application. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code, and the module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in an alternative implementation, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks can actually be executed substantially in parallel, and they can sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the structure diagram and / or flowchart, as well as the combination of blocks in the structure diagram and / or flowchart, can be implemented by a dedicated hardware-based system that performs the specified functions or actions, or can be implemented by a combination of dedicated hardware and computer instructions.

[0073] In addition, in each embodiment of this application, the various functional modules or units can be integrated together to form an independent part, or each module can exist alone, or two or more modules can be integrated to form an independent part.

[0074] If the above functions are implemented in the form of software function modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a smart phone, a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of this application. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.

[0075] The above is only the specific implementation manner of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art within the technical scope disclosed in this application can easily think of changes or substitutions, which should all be covered by the protection scope of this application.

Claims

1. A smart home reinforcement learning interaction method, characterized in that, include: Obtain environmental data through sensors; Get the behavior data of each smart home device; Inputting the environmental data and the behavioral data into a pre-trained AI model to obtain a behavioral decision output by the AI model; According to the behavior decision, the corresponding smart home device is controlled to perform the corresponding task.

2. The smart home reinforcement learning interaction method according to claim 1, wherein Also includes: When receiving an active interaction event from the user, obtaining the current environment data and the status data of the device corresponding to the interaction event; The active interaction event, the environmental data at the current moment and the state data are used as reward factors, and the reward factors are fed back to the AI model to iteratively optimize the AI model.

3. The smart home reinforcement learning interaction method according to claim 1, wherein, According to the behavior decision, controlling the corresponding smart home device to perform the corresponding task includes: The behavior decision is prompted to the user by voice broadcast, and when a consent instruction issued by the user is received, the corresponding smart home device is controlled to execute the behavior decision.

4. The smart home reinforcement learning interaction method according to claim 3, wherein, When receiving the consent instruction issued by the user, the method further includes: The consent instruction is fed back into the AI model as positive feedback to iteratively optimize the AI model.

5. The smart home reinforcement learning interaction method according to claim 3, wherein After the behavior decision is prompted to the user by voice broadcast, the method further includes: When a disagreement instruction is received from the user, the behavioral decision is not executed, and the disagreement instruction is fed back into the AI model as negative feedback to iteratively optimize the AI model.

6. A smart home reinforcement learning interaction system, characterized in that, Includes multiple smart home devices, sensors and control terminals; The control terminal is used to obtain environmental data through the sensor, and obtain behavior data corresponding to each of the smart home devices; Inputting the environmental data and the behavior data into a pre-trained AI model in the control terminal to obtain a behavior decision output by the AI model; The control terminal sends a control command to the corresponding smart home device according to the behavior decision to execute the behavior decision.

7. The smart home reinforcement learning interaction system according to claim 6, wherein include: When the control terminal receives an active interaction event from the user, it obtains the current environment data and the status data of the device corresponding to the interaction event; The active interaction event, the environmental data at the previous moment and the state data are used as reward factors, the reward factors are fed back to the AI model, and the AI model is iteratively optimized.

8. The intelligent home reinforcement learning interaction system according to claim 6, characterized in that, The step of sending a control command to a corresponding smart home device according to the behavior decision to execute the behavior decision includes: The control terminal prompts the user of the behavior decision by voice broadcast, and when receiving the consent instruction issued by the user, sends a control instruction to the corresponding smart home device to execute the behavior decision.

9. A terminal device, characterized in that, The terminal device includes a processor and a memory, the memory stores a computer program, and the processor is used to execute the computer program to implement the smart home reinforcement learning interaction method according to any one of claims 1 to 5.

10. A readable storage medium, characterized in that, It stores a computer program, which, when executed on a processor, implements the smart home reinforcement learning interaction method according to any one of claims 1-5.