Control method of air conditioner and related product

By dynamically adjusting the optimization target weights of the air conditioner through a multi-objective reinforcement learning model and combining user input information, the problem of data distortion caused by sensor noise or sudden interference in traditional air conditioning control methods is solved, realizing a control strategy that better meets user needs and improving user experience.

CN121655103APending Publication Date: 2026-03-13QINGDAO HAIER AIR CONDITIONING ELECTRONICS CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-30
Publication Date
2026-03-13

AI Technical Summary

Technical Problem

Traditional air conditioning control methods rely on fixed weights or single sensor data, which are prone to data distortion due to sensor noise or sudden interference. They are difficult to effectively distinguish between noise or sudden interference and real signals, leading to deviations in control strategies.

Method used

A multi-objective reinforcement learning model is adopted to dynamically adjust the weights of the optimization objectives by acquiring the current and historical operating parameters of the air conditioner. Combined with user input information, the optimal control command is output to avoid misjudgment due to sensor noise or sudden interference.

Benefits of technology

It improves the matching degree between the air conditioner's control strategy and user needs, enhances the user experience, solves the problem that traditional air conditioners cannot adapt to personalized needs, and enhances the stability and flexibility of control.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121655103A_ABST
    Figure CN121655103A_ABST
Patent Text Reader

Abstract

The invention provides a control method of an air conditioner and a related product. The method comprises the steps that current operation parameters and historical operation parameters of the air conditioner are obtained; and according to the current operation parameter and the historical operation parameter, determining a historical target parameter of the user. And determining the weight of each optimization target in the multi-target reinforcement learning model according to the current operation parameters and / or the historical operation parameters. And determining state data in the multi-target reinforcement learning model according to the current operation parameters and the historical target parameters. And the multi-target reinforcement learning model outputs a control instruction for controlling the air conditioner based on the weight and the state data. And under the condition that the weight is not equal to the initial weight in the multi-target reinforcement learning model, determining whether an emergency occurs or not according to the current operation parameter and / or the historical operation parameter. And if the emergency is not generated, replacing the initial weight with the weight. And judging whether an emergency situation occurs or not is added, so that misjudgment caused by sudden interference when the air conditioner generates the control instruction can be avoided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of air conditioning technology, and in particular to a control method for an air conditioner and related products. Background Technology

[0002] Traditional air conditioning control methods rely on fixed weights or single sensor data. In the environmental perception stage, data distortion often occurs due to sensor noise or sudden interference. For example, a sudden increase in indoor heat sources or local temperature anomalies caused by opening doors and windows may be misjudged as global demand. Traditional air conditioning control methods have difficulty effectively distinguishing noise or sudden interference from real signals, leading to deviations in control strategies. Summary of the Invention

[0003] In view of the above problems, the present invention is proposed to provide a control method and related products for an air conditioner that overcomes or at least partially solves the above problems, effectively avoiding data distortion caused by sensor noise or sudden interference, and improving the user experience.

[0004] Specifically, the present invention provides a control method for an air conditioner, comprising: Obtain the current and historical operating parameters of the air conditioner; Based on the current operating parameters and the historical operating parameters, determine the user's historical target parameters; The weights of each optimization objective in the multi-objective reinforcement learning model are determined based on the current operating parameters and / or the historical operating parameters. The state data in the multi-objective reinforcement learning model is determined based on the current operating parameters and the historical target parameters. The multi-objective reinforcement learning model outputs control commands to control the air conditioner based on the weights and the state data. If the weights are not equal to the initial weights in the multi-objective reinforcement learning model, determine whether a sudden event occurs based on the current operating parameters and / or the historical operating parameters. If no sudden event occurs, the initial weights are replaced with the stated weights.

[0005] Optionally, determining whether a sudden event has occurred based on the current operating parameters and / or the historical operating parameters includes: The current operating parameters and the historical operating parameters are compared to obtain the comparison result; The comparison results will determine whether an emergency has occurred.

[0006] Optionally, the comparison result is the probability of the current operating parameters being generated; Determining whether a sudden event has occurred based on the comparison results specifically involves: determining that a sudden event has occurred when the probability is lower than a preset probability value; Specifically, after receiving user input information, the current operating parameters include the parameters input by the user, and when the probability is higher than the preset probability value, it is determined that the user's behavior has changed.

[0007] Optionally, determining the state data in the multi-objective reinforcement learning model based on the current operating parameters and the historical target parameters includes: Determine whether the current running parameters contain the target parameters input by the user; If so, the current operating parameters shall be used as the status data; If not, the current operating parameters and the historical target parameters shall be used as the status data.

[0008] Optionally, the current operating parameters include environmental parameters or time parameters of the air conditioner; determining the user's historical target parameters based on the current operating parameters and the historical operating parameters includes: Historical data corresponding to the time parameter or the environmental parameter are selected from the historical operating parameters. The historical target parameters are determined based on the historical data.

[0009] Optionally, the optimization objectives include user comfort, energy saving, and air conditioner health; The operating parameters include indoor temperature, indoor humidity, and / or indoor concentration of toxic gases; Determining the weights of each optimization objective in the multi-objective reinforcement learning model based on the current operating parameters and / or the historical operating parameters includes: Obtain the initial weights for each of the optimization objectives; When the air conditioner is cooling, it is determined whether the rate of increase of the temperature exceeds a first threshold, or whether the temperature exceeds a second threshold, or whether the rate of change of the humidity exceeds a third threshold, or whether the humidity exceeds a fourth threshold; or whether the rate of increase of the concentration of toxic gas exceeds a fifth threshold, or whether the concentration of toxic gas exceeds a sixth threshold. If so, increase the initial weight corresponding to the user comfort level.

[0010] Optionally, when the rate of increase of the toxic gas concentration exceeds a fifth threshold, or when the toxic gas concentration exceeds a sixth threshold, the weight corresponding to the user comfort is increased to obtain a first weight. When the rate of increase in temperature exceeds a first threshold, or the temperature exceeds a second threshold, or the rate of change in humidity exceeds a third threshold, or the humidity exceeds a fourth threshold, the weight corresponding to the user comfort is increased to obtain a second weight; the first weight is greater than the second weight.

[0011] Optionally, the optimization objectives include user comfort, energy saving, and air conditioner health; The operating parameters include peak electricity consumption periods; Determining the weights of each optimization objective in the multi-objective reinforcement learning model based on the current operating parameters and / or the historical operating parameters includes: Determine whether the air conditioner is operating during peak electricity consumption hours; If so, increase the weight corresponding to the energy saving; and / or, The operating parameters include the continuous operating time of the compressor and the number of times the compressor is protected. Determine whether the continuous running time of the air conditioner's compressor exceeds a preset time, or determine whether the number of times the compressor has been protected exceeds a preset number; If so, increase the weight corresponding to the health status of the air conditioner; and / or The operating parameters include target parameters input by the user, and the target parameters include target temperature; Determining the weights of each optimization objective in the multi-objective reinforcement learning model based on the current operating parameters and / or the historical operating parameters includes: Based on the current operating parameters and the historical operating parameters, the user's historical target parameters are determined; the historical target parameters include historical target temperature. When the air conditioner is cooling and the target temperature is lower than the historical target temperature, the weight corresponding to the energy saving is reduced; and / or The operating parameters include the outdoor ambient temperature; Determining the weights of each optimization objective in the multi-objective reinforcement learning model based on the current operating parameters and / or the historical operating parameters includes: When the air conditioner is cooling, it is determined whether the outdoor ambient temperature is higher than the outdoor ambient temperature in the historical operating parameters; If so, reduce the weight corresponding to energy saving and increase the weight corresponding to user comfort.

[0012] Optionally, the multi-objective reinforcement learning model includes multiple action spaces, and the execution object corresponding to each action space is different from the execution object corresponding to each of the other action spaces; the multi-objective reinforcement learning model outputs control commands for controlling the air conditioner based on the weights and the state data, including: Based on the weights, the state data, and the multiple action spaces, multiple control strategies and a comprehensive score corresponding to each control strategy are obtained; each target control strategy contains the control instructions. Output the control strategy with the highest overall score.

[0013] The present invention also provides a computer-readable storage medium, including a memory, a processor, and a computer program stored on the memory, wherein the processor executes the computer program to implement the steps of the control method for an air conditioner described in any of the preceding claims.

[0014] In the air conditioner control method and related products of this invention, a multi-objective reinforcement learning model continuously learns user habits and environmental patterns through iterative weight updates. The longer the running time, the more closely the control strategy matches user needs, solving the problems of limited user preference prediction and inability to adapt to personalized needs in traditional air conditioners. Adding a judgment on whether an emergency occurs avoids misjudgments caused by sensor noise or sudden interference when the air conditioner generates control commands, effectively preventing data distortion caused by sensor noise or sudden interference. This makes the air conditioner's operation better meet user needs and improves the user experience.

[0015] The above and other objects, advantages and features of the present invention will become more apparent to those skilled in the art from the following detailed description of specific embodiments of the invention in conjunction with the accompanying drawings. Attached Figure Description

[0016] The following sections will describe some specific embodiments of the invention in detail by way of example and not limitation, with reference to the accompanying drawings. The same reference numerals in the drawings denote the same or similar parts or portions. Those skilled in the art should understand that these drawings are not necessarily drawn to scale. In the drawings: Figure 1 This is a schematic flowchart of a control method for an air conditioner according to an embodiment of the present invention; Figure 2 This is a schematic flowchart of a control method for an air conditioner according to an embodiment of the present invention; Figure 3 This is a schematic flowchart of a control method for an air conditioner according to an embodiment of the present invention; Figure 4 This is a schematic flowchart of a control method for an air conditioner according to an embodiment of the present invention; Figure 5 This is a schematic flowchart of a control method for an air conditioner according to an embodiment of the present invention; Figure 6 This is a schematic diagram of a computer program product according to an embodiment of the present invention; Figure 7This is a schematic diagram of a computer-readable storage medium according to an embodiment of the present invention; and Figure 8 This is a schematic diagram of a computer device according to an embodiment of the present invention. Detailed Implementation

[0017] The following reference Figures 1 to 8 This invention describes the control method for an air conditioner and related products according to embodiments of the present invention. In this description, it should be understood that the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Therefore, a feature defined with "first" and "second" may explicitly or implicitly include at least one of that feature, that is, include one or more of that feature. In the description of the present invention, "a plurality of" means at least two, such as two, three, etc., unless otherwise explicitly specified. When a feature "includes or contains" one or more of the features it encompasses, unless otherwise specifically described, this indicates that other features are not excluded and may be further included.

[0018] Unless otherwise expressly specified and limited, the terms "set up," "install," "connect," "link," "fix," and "couple" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral part; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; they can refer to the internal communication of two components or the interaction between two components, unless otherwise expressly limited. Those skilled in the art should be able to understand the specific meaning of the above terms in this invention according to the specific circumstances.

[0019] Furthermore, in the description of this embodiment, "above" or "below" the second feature can include direct contact between the first and second features, or it can include contact between the first and second features through another feature between them. That is, in the description of this embodiment, "above," "over," and "on top" of the second feature includes the first feature being directly above or diagonally above the second feature, or simply indicates that the first feature is at a higher horizontal level than the second feature. "Below," "below," or "below" of the second feature can mean the first feature is directly below or diagonally below the second feature, or simply indicates that the first feature is at a lower horizontal level than the second feature.

[0020] In the description of this embodiment, the terms "one embodiment," "some embodiments," "illustrative embodiment," "example," "specific example," or "some examples," etc., refer to specific features, structures, materials, or characteristics described in connection with that embodiment or example, which are included in at least one embodiment or example of the present invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.

[0021] Figure 1 This is a schematic flowchart of an air conditioner control method according to an embodiment of the present invention, as shown below. Figure 1 As shown, and refer to Figures 2 to 5 This invention provides a control method for an air conditioner, which generally includes: Step S100: Obtain the current operating parameters and historical operating parameters of the air conditioner; Step S200: Determine the user's historical target parameters based on the current operating parameters and historical operating parameters; Step S300: Determine the weights of each optimization objective in the multi-objective reinforcement learning model based on the current and historical operating parameters; Step S400: Determine the state data in the multi-objective reinforcement learning model based on the current running parameters and historical target parameters; Step S500: The multi-objective reinforcement learning model outputs control commands for the air conditioner based on weight and state data; Step S600: Determine whether the weights are equal to the initial weights in the multi-objective reinforcement learning model; If not, proceed to step S700 to determine whether a sudden event has occurred based on the current operating parameters and historical operating parameters; If not, proceed to step S800 to replace the initial weights with the new weights.

[0022] First, obtain the current and historical operating parameters of the air conditioner. Combine these parameters to analyze and derive the user's historical target parameters, which represent the user's long-term needs, such as desired temperature, humidity, and fan speed. Based on these parameters, assign different weights to the various optimization objectives of the multi-objective reinforcement learning model. Integrate the current and historical target parameters as input state data for the model, allowing it to understand the current environment and user needs. The model calculates and outputs the optimal control command based on the weights and state data. If the current weights differ from the initial weights, compare the current and historical operating parameters to determine if a sudden event has occurred. If no sudden event occurs, the weight adjustment is considered reasonable, and the new weights replace the initial weights, allowing the model to directly use the optimized weights for future decisions, achieving continuous self-optimization. If a sudden event occurs, the new weights do not replace the initial weights to avoid interference signals misleading the control strategy; the model still uses the initial weights to issue control commands.

[0023] This solution uses iterative weight updates to allow the multi-objective reinforcement learning model to continuously learn user habits and environmental patterns. The longer the model runs, the more closely the control strategy aligns with user needs, solving the problem of traditional air conditioners having overly limited user preference prediction capabilities and being unable to adapt to personalized requirements. Adding a check for unforeseen circumstances prevents misjudgments caused by sensor noise or sudden interference when generating control commands, effectively avoiding data distortion due to sensor noise or sudden interference. This ensures that the air conditioner's operation better meets user needs and improves the user experience.

[0024] In other embodiments of the present invention, the weights of each optimization objective in the multi-objective reinforcement learning model are determined based on the current operating parameters.

[0025] In other embodiments of the present invention, the weights of each optimization objective in the multi-objective reinforcement learning model are determined based on historical operating parameters.

[0026] In some embodiments of the present invention, such as Figure 2 As shown, step S700, determining whether a sudden event has occurred based on current operating parameters and / or historical operating parameters, includes: Step S710: Compare the current operating parameters with the historical operating parameters to obtain the comparison result; Step S720: Determine whether an emergency has occurred based on the comparison results.

[0027] By comparing the current operating parameters with historical operating parameters, it can be determined whether a sudden event has occurred. For example, if the temperature rises at the same time in historical operating data, then the temperature rise at this time cannot be identified as a sudden event.

[0028] In some embodiments of the present invention, the comparison result is the probability of the current operating parameters being generated; Determining whether a sudden event has occurred based on the comparison results specifically means: if the probability is lower than the preset probability value, a sudden event is determined to have occurred.

[0029] Specifically, after receiving user input, the current operating parameters include the parameters input by the user, and when the probability is higher than the preset probability value, it is determined that the user's behavior has changed.

[0030] By calculating the probability of the current operating parameters and comparing it with preset probability values, if the probability is below a threshold, it is determined to be an unexpected event, such as a user temporarily opening a window or turning on a new heater. In this case, the model weights are not updated to avoid interference signals misleading the control strategy. If the probability is above the threshold, it is determined to be a change in user habits, and the weights are updated to adapt the model to the new habits. Simultaneously, by incorporating user input parameters, the accuracy of the judgment is further improved, addressing the pain point of traditional air conditioners' inability to distinguish between noise and actual needs, thus balancing control stability and adaptability.

[0031] In some embodiments of the present invention, such as Figure 3 As shown, step S400 determines the state data in the multi-objective reinforcement learning model based on the current running parameters and historical target parameters, including: Step S410: Determine whether the current running parameters have the target parameters input by the user; If so, proceed to step S420 and use the current running parameters as status data; If not, proceed to step S430, using the current running parameters and historical target parameters as status data.

[0032] When the operating parameters include user-inputted target parameters, the current operating parameters are used directly as the status data to prioritize responding to the user's proactive needs. When there are no user-inputted target parameters, the current operating parameters are combined with historical target parameters, and the decision-making basis is supplemented by the user's past habits, ensuring that the control strategy aligns with long-term preferences. This method ensures both immediate response to user-initiated actions and utilizes historical data to optimize for better meeting user needs in the absence of input parameters.

[0033] In some embodiments of the present invention, such as Figure 5 As shown, the current operating parameters include the environmental parameters or time parameters of the air conditioner. Step S200: Based on the current operating parameters and historical operating parameters, determine the user's historical target parameters, including: Step S210: Select historical data corresponding to time parameters or environmental parameters from historical operating parameters; Step S220: Determine the historical target parameters based on historical data.

[0034] By filtering historical operating parameters that better match the current environmental and time parameters, interference from irrelevant data is avoided, making the state data of the multi-objective reinforcement learning model more targeted and relevant to the current usage scenario.

[0035] In some embodiments of the present invention, the optimization objectives include user comfort, energy saving, and air conditioner health.

[0036] Operating parameters include indoor temperature, indoor humidity, and / or indoor concentration of toxic gases.

[0037] The weights of each optimization objective in the multi-objective reinforcement learning model are determined based on current and / or historical operating parameters, including: Obtain the initial weights for each optimization objective; When the air conditioner is cooling, it determines whether the rate of increase in temperature exceeds the first threshold, or whether the temperature exceeds the second threshold, or whether the rate of change in humidity exceeds the third threshold, or whether the humidity exceeds the fourth threshold; or whether the rate of increase in toxic gas concentration exceeds the fifth threshold, or whether the toxic gas concentration exceeds the sixth threshold. If so, increase the initial weight corresponding to user comfort.

[0038] During cooling, by monitoring the absolute values ​​and rates of change of temperature, humidity, and toxic gas concentration, once a threshold is triggered, such as a sudden rise in temperature, excessive humidity, or a sudden increase in toxic gas concentration, the weight of user comfort is immediately increased. Parameters are adjusted first to improve user comfort, prioritizing core user needs while also considering energy conservation and air conditioning health. The toxic gas concentration here can be the concentration of carbon dioxide. There can be only one set of initial weights or multiple sets; the air conditioner has a built-in lookup table for time and initial weights, and the initial weights are called based on time.

[0039] In some embodiments of the present invention, the air conditioner includes a temperature sensor, a humidity sensor, and a carbon dioxide sensor to collect environmental data.

[0040] In some embodiments of the present invention, when the rate of increase of toxic gas concentration exceeds a fifth threshold, or the toxic gas concentration exceeds a sixth threshold, the weight corresponding to user comfort is increased to obtain a first weight. When the rate of increase of temperature exceeds the first threshold, or the temperature exceeds a second threshold, or the rate of change of humidity exceeds a third threshold, or the humidity exceeds a fourth threshold, the weight corresponding to user comfort is increased to obtain a second weight; the first weight is greater than the second weight.

[0041] To ensure users' health, prioritizing user comfort is more important than prioritizing toxic gas levels when they exceed acceptable limits, in order to reduce toxic gas concentrations as quickly as possible and avoid harming users' health.

[0042] In some embodiments of the present invention, the optimization objectives include user comfort, energy saving, and air conditioner health. Operating parameters include peak electricity consumption periods. The weights of each optimization objective in the multi-objective reinforcement learning model are determined based on current and / or historical operating parameters, including: Determine if the air conditioner is operating during peak electricity consumption hours; If so, increase the weight corresponding to energy saving.

[0043] When the system determines that it is currently in a peak electricity consumption period, it directly increases the weight of the energy-saving target, causing the air conditioner to operate in a low-power mode and reduce the electricity load.

[0044] In some embodiments of the present invention, the operating parameters include the continuous operating time of the compressor and the number of times the compressor is protected. The weights of each optimization objective in the multi-objective reinforcement learning model are determined based on the current operating parameters and / or historical operating parameters, including: determining whether the continuous operating time of the air conditioner's compressor exceeds a preset duration, or determining whether the number of times the compressor is protected exceeds a preset number; If so, increase the weighting of the health-related aspects of the air conditioner.

[0045] When the continuous running time of the compressor exceeds the preset value, or the number of protection triggers exceeds the preset value, the weight of the air conditioner's health target is increased to control and reduce the compressor's operating power and shorten the continuous running time, so as to avoid equipment overload damage.

[0046] In some embodiments of the present invention, the operating parameters include target parameters input by the user, and the target parameters include target temperature.

[0047] The weights of each optimization objective in the multi-objective reinforcement learning model are determined based on current and / or historical operating parameters, including: Based on the current operating parameters and historical operating parameters, determine the user's historical target parameters; the historical target parameters include the historical target temperature. When the air conditioner is cooling and the target temperature is lower than the historical target temperature, the weight corresponding to energy saving is reduced.

[0048] In cooling mode, if the target temperature set by the user is lower than the target temperature in the historical parameters, for example, if the user usually sets it to 26℃ and now sets it to 24℃, it means that the user is more concerned about rapid cooling. The system will reduce the weight of the energy-saving target and allow the air conditioner to run at a higher power.

[0049] In some embodiments of the present invention, the operating parameters include outdoor ambient temperature. Determining the weights of each optimization objective in the multi-objective reinforcement learning model based on current and / or historical operating parameters includes: When the air conditioner is cooling, determine whether the outdoor ambient temperature is higher than the outdoor ambient temperature in the historical operating parameters; If so, reduce the weight of energy saving and increase the weight of user comfort.

[0050] In cooling mode, if the current outdoor temperature is higher than the historical average for the same period, it indicates that the cooling difficulty has increased. The system will reduce the energy saving weight and increase the user comfort weight, and ensure the cooling effect by increasing power to avoid the room temperature not reaching the target due to energy saving limitations.

[0051] In some embodiments of the present invention, such as Figure 4 As shown, the multi-objective reinforcement learning model includes multiple action spaces, each with a different execution object than the others. These action spaces include a first action space, a second action space, and a third action space. The first action space controls the compressor speed of the air conditioner. The second action space controls both the compressor speed and the fresh air module of the air conditioner. The third action space controls both the compressor speed and either the humidification or heating module of the air conditioner. Step S500: The multi-objective reinforcement learning model outputs control commands for the air conditioner based on weight and state data, including: Step S510: Based on weight and state data, and multiple action spaces, obtain multiple control strategies and a comprehensive score corresponding to each control strategy; each target control strategy contains control instructions; Step S520: Output the control strategy with the highest overall score.

[0052] The control strategy and score are obtained based on weights, state data and multiple action spaces. The control strategy with the highest comprehensive score is output, which can avoid the decision bias of single goal orientation and ensure that the output strategy achieves the optimal balance between comfort, energy saving and equipment health.

[0053] In some embodiments of the present invention, the current operating parameters and historical operating parameters of the air conditioner are obtained. The current operating parameters include environmental parameters or time parameters of the air conditioner. Historical data corresponding to the time or environmental parameters are selected from the historical operating parameters, and then historical target parameters are determined based on the historical data. Initial weights for each optimization objective are obtained; these optimization objectives include user comfort, energy saving, and air conditioner health. Operating parameters also include indoor temperature, indoor humidity, and / or indoor toxic gas concentration. When the air conditioner is cooling, it is determined whether the rate of temperature increase exceeds a first threshold, or whether the temperature exceeds a second threshold, or whether the rate of humidity change exceeds a third threshold, or whether the humidity exceeds a fourth threshold. Alternatively, it is determined whether the rate of increase in toxic gas concentration exceeds a fifth threshold, or whether the toxic gas concentration exceeds a sixth threshold. If any of these determinations is true, the initial weight corresponding to user comfort is increased. That is, when making the above determinations, priority is given to the impact on user comfort. In particular, when the toxic gas concentration exceeds the standard, the increase in the weight corresponding to user comfort is greater than the increase in the weight when the temperature and humidity exceed the standard. Operating parameters also include peak electricity consumption periods. When the air conditioner is in peak electricity consumption periods, the weight corresponding to energy saving is increased. Operating parameters also include the continuous running time of the compressor and the number of compressor protection cycles. When the continuous running time of the air conditioner's compressor exceeds the preset duration, or the number of compressor protection cycles exceeds the preset number, the weight corresponding to air conditioner health is increased. Operating parameters include the user-input target temperature. Based on the current user-input target temperature and historical user-input target temperatures, the user's historical target temperature is determined. When the air conditioner is cooling and the target temperature is lower than the historical target temperature, the weight corresponding to energy saving is reduced. Operating parameters include the outdoor ambient temperature. When the air conditioner is cooling and the outdoor ambient temperature is higher than the outdoor ambient temperature in the historical operating parameters, the weight corresponding to energy saving is reduced, and the weight corresponding to user comfort is increased. Based on the weights and state data, as well as multiple action spaces, multiple control strategies and a comprehensive score corresponding to each control strategy are obtained. Each target control strategy contains control instructions. The control strategy with the highest comprehensive score is output. If the weights are not equal to the initial weights, the current operating parameters and historical operating parameters are compared to obtain the comparison result. When the probability is lower than the preset probability value, a sudden event is determined, and the new weights do not replace the initial weights to avoid interference signals misleading the control strategy. If the probability is not lower than the preset probability value, it is determined that it is not a sudden event, and the initial weights are replaced with the new weights. The multi-objective reinforcement learning model can then directly use the optimized weights to make decisions in the next iteration, achieving continuous self-optimization.

[0054] This embodiment also provides a computer program product 10, a computer-readable storage medium 20, and a computer device 30. Figure 6 This is a schematic diagram of a computer program product 10 according to an embodiment of the present invention. Figure 7 This is a schematic diagram of a computer-readable storage medium 20 according to an embodiment of the present invention. Figure 8 This is a schematic diagram of a computer device 30 according to an embodiment of the present invention. The computer program product 10 includes a computer program 11, which, when executed by the processor 32, implements the steps of the control method for an air conditioner described above. A computer-readable storage medium 20 stores the computer program 11 thereon, which, when executed by the processor 32, implements the steps of the control method for an air conditioner according to any of the above embodiments. The computer device 30 may include a memory 31, a processor 32, and the computer program 11 stored in the memory 31 and running on the processor 32.

[0055] The computer program 11 used to perform the operations of this invention may be assembly instructions, Instruction Set Architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, status setting data, integrated circuit configuration data, or source code or object code written in any combination of one or more programming languages ​​and procedural programming languages. The computer program 11 may execute entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the latter case, the remote computer may be connected to the user's computer via any type of network, including a Local Area Network (LAN) or Wide Area Network (WAN), or may be connected to an external computer. In some embodiments, to perform aspects of this invention, electronic circuits, including, for example, programmable logic circuits, Field-Programmable Gate Arrays (FPGAs), or Programmable Logic Arrays (PLAs), may execute computer-readable program instructions to personalize the electronic circuits by utilizing status information of the computer-readable program instructions.

[0056] For the purposes of this embodiment, computer program product 10 is a related product that includes computer program 11.

[0057] For the purposes of this embodiment, computer-readable storage medium 20 is a tangible device capable of holding and storing a computer program 11. It can be any device capable of containing, storing, communicating, propagating, or transmitting the program 11 for use by or in conjunction with an instruction execution system, apparatus, or device. More specific examples (a non-exhaustive list) of computer-readable storage medium 20 include: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable optical disc read-only memory (CD-ROM), digital versatile disc (DVD), memory stick, floppy disk, mechanical encoding device, and any suitable combination thereof.

[0058] Computer device 30 can be, for example, a server, desktop computer, laptop computer, tablet computer, or smartphone. In some examples, computer device 30 can be a cloud computing node. Computer device 30 can be described in the general context of computer system executable instructions (such as program modules) executed by a computer system. Typically, program modules can include routines, programs, object programs, components, logic, data structures, etc., that perform specific tasks or implement specific abstract data types. Computer device 30 can be implemented in a distributed cloud computing environment where tasks are performed by remote processing devices linked through a communication network. In a distributed cloud computing environment, program modules can reside on local or remote computing system storage media, including storage devices.

[0059] Computer device 30 may include a processor 32 adapted to execute stored instructions and a memory 31 that provides temporary storage space for the operation of said instructions during operation. The processor 32 may be a single-core processor, a multi-core processor, a computing cluster, or any other configuration. The memory 31 may include random access memory (RAM), read-only memory, flash memory, or any other suitable storage system.

[0060] Computer device 30 may also include a network adapter / interface and an input / output (I / O) interface. The I / O interface allows external devices that can be connected to the computer device to input and output data. The network adapter / interface provides communication between the computer device and a network, typically represented as a communication network.

[0061] Therefore, those skilled in the art should recognize that although numerous exemplary embodiments of the present invention have been shown and described in detail herein, many other variations or modifications conforming to the principles of the present invention can be directly determined or derived from the disclosure of the present invention without departing from the spirit and scope of the invention. Thus, the scope of the present invention should be understood and construed as covering all such other variations or modifications.

Claims

1. A control method for an air conditioner, characterized in that, include: Obtain the current and historical operating parameters of the air conditioner; Based on the current operating parameters and the historical operating parameters, determine the user's historical target parameters; The weights of each optimization objective in the multi-objective reinforcement learning model are determined based on the current operating parameters and / or the historical operating parameters. The state data in the multi-objective reinforcement learning model is determined based on the current operating parameters and the historical target parameters. The multi-objective reinforcement learning model outputs control commands to control the air conditioner based on the weights and the state data. If the weights are not equal to the initial weights in the multi-objective reinforcement learning model, determine whether a sudden event occurs based on the current operating parameters and / or the historical operating parameters. If no sudden event occurs, the initial weights are replaced with the stated weights.

2. The control method according to claim 1, characterized in that, The step of determining whether a sudden event has occurred based on the current operating parameters and / or the historical operating parameters includes: The current operating parameters and the historical operating parameters are compared to obtain the comparison result; The comparison results will determine whether an emergency has occurred.

3. The control method according to claim 2, characterized in that, The comparison result represents the probability of the current operating parameters being generated. Determining whether a sudden event has occurred based on the comparison results specifically involves: determining that a sudden event has occurred when the probability is lower than a preset probability value; Specifically, after receiving user input information, the current operating parameters include the parameters input by the user, and when the probability is higher than the preset probability value, it is determined that the user's behavior has changed.

4. The control method according to claim 1, characterized in that, The step of determining the state data in the multi-objective reinforcement learning model based on the current operating parameters and the historical target parameters includes: Determine whether the current running parameters contain the target parameters input by the user; If so, the current operating parameters shall be used as the status data; If not, the current operating parameters and the historical target parameters shall be used as the status data.

5. The control method according to claim 1, characterized in that, The current operating parameters include environmental parameters or time parameters of the air conditioner; determining the user's historical target parameters based on the current operating parameters and the historical operating parameters includes: Historical data corresponding to the time parameter or the environmental parameter are selected from the historical operating parameters. The historical target parameters are determined based on the historical data.

6. The control method according to claim 1, characterized in that, The optimization objectives include user comfort, energy saving, and air conditioner health. The operating parameters include indoor temperature, indoor humidity, and / or indoor concentration of toxic gases; Determining the weights of each optimization objective in the multi-objective reinforcement learning model based on the current operating parameters and / or the historical operating parameters includes: Obtain the initial weights for each of the optimization objectives; When the air conditioner is cooling, it is determined whether the rate of increase of the temperature exceeds a first threshold, or whether the temperature exceeds a second threshold, or whether the rate of change of the humidity exceeds a third threshold, or whether the humidity exceeds a fourth threshold; or whether the rate of increase of the concentration of toxic gas exceeds a fifth threshold, or whether the concentration of toxic gas exceeds a sixth threshold. If so, increase the initial weight corresponding to the user comfort level.

7. The control method according to claim 6, characterized in that, When the rate of increase of the toxic gas concentration exceeds the fifth threshold, or when the toxic gas concentration exceeds the sixth threshold, the weight corresponding to the user comfort is increased to obtain the first weight. When the rate of increase in temperature exceeds a first threshold, or the temperature exceeds a second threshold, or the rate of change in humidity exceeds a third threshold, or the humidity exceeds a fourth threshold, the weight corresponding to the user comfort is increased to obtain a second weight; the first weight is greater than the second weight.

8. The control method according to claim 1, characterized in that, The optimization objectives include user comfort, energy saving, and air conditioner health. The operating parameters include peak electricity consumption periods; Determining the weights of each optimization objective in the multi-objective reinforcement learning model based on the current operating parameters and / or the historical operating parameters includes: Determine whether the air conditioner is operating during peak electricity consumption hours; If so, increase the weight corresponding to the energy saving; and / or, The operating parameters include the continuous operating time of the compressor and the number of times the compressor is protected. Determine whether the continuous running time of the air conditioner's compressor exceeds a preset duration, or determine whether the number of times the compressor has been protected exceeds a preset number; If so, increase the weight corresponding to the health status of the air conditioner; and / or The operating parameters include target parameters input by the user, and the target parameters include target temperature; Determining the weights of each optimization objective in the multi-objective reinforcement learning model based on the current operating parameters and / or the historical operating parameters includes: Based on the current operating parameters and the historical operating parameters, the user's historical target parameters are determined; the historical target parameters include historical target temperature. When the air conditioner is cooling and the target temperature is lower than the historical target temperature, the weight corresponding to the energy saving is reduced; and / or The operating parameters include the outdoor ambient temperature; Determining the weights of each optimization objective in the multi-objective reinforcement learning model based on the current operating parameters and / or the historical operating parameters includes: When the air conditioner is cooling, it is determined whether the outdoor ambient temperature is higher than the outdoor ambient temperature in the historical operating parameters; If so, reduce the weight corresponding to energy saving and increase the weight corresponding to user comfort.

9. The control method according to claim 1, characterized in that, The multi-objective reinforcement learning model includes multiple action spaces, and the execution object corresponding to each action space is different from the execution object corresponding to each of the other action spaces. The multi-objective reinforcement learning model outputs control commands for the air conditioner based on the weights and the state data, including: Based on the weights, the state data, and the multiple action spaces, multiple control strategies and a comprehensive score corresponding to each control strategy are obtained; each target control strategy contains the control instructions. Output the control strategy with the highest overall score.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the control method for the air conditioner according to any one of claims 1 to 9.