An air conditioning control method, device, system, electronic equipment, and air conditioner.

By generating personalized air conditioner noise control strategies through deep reinforcement learning models and state prediction technology, the problem that existing air conditioner noise control methods cannot meet user needs is solved, thus improving the comfort experience of air conditioning.

CN119737669BActive Publication Date: 2026-05-26GREE ELECTRIC APPLIANCE INC OF ZHUHAI

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
GREE ELECTRIC APPLIANCE INC OF ZHUHAI
Filing Date
2024-11-12
Publication Date
2026-05-26

AI Technical Summary

Technical Problem

Existing air conditioning noise control methods cannot meet the personalized needs of different users, and preset noise reduction strategies are difficult to meet the actual needs of users.

Method used

A deep reinforcement learning model is used to generate noise control strategies. Combining the current environmental state of the air conditioner with user needs, the environmental state is predicted and updated through a state prediction model, and the noise control strategies are selectively executed or updated based on user feedback.

Benefits of technology

It enables flexible noise control based on users' individual needs, improving the comfort experience of air conditioning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119737669B_ABST
    Figure CN119737669B_ABST
Patent Text Reader

Abstract

This application relates to the field of air conditioning technology, providing an air conditioning control method, device, system, electronic device, and air conditioner. The air conditioning control method includes: acquiring input parameters, the input parameters including a current environmental state, or the input parameters including the current environmental state and a user's air conditioning usage needs, wherein the environmental state includes values ​​of environmental factors related to noise during air conditioning operation; generating a noise control strategy based on the input parameters and a strategy generation model; predicting an updated environmental state after adopting the noise control strategy based on the noise control strategy and the current environmental state; and selectively executing or updating the noise control strategy in response to user feedback regarding the updated environmental state and / or the noise control strategy. Using this application is beneficial for obtaining noise control strategies that meet the needs of different users.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of air conditioning technology, and in particular to an air conditioning control method, device, system, electronic device and air conditioner. Background Technology

[0002] With the widespread use of air conditioning, it has become a necessity in people's lives. People's needs have shifted from the most basic cooling and heating to pursuing a higher level of comfort.

[0003] Air conditioner noise has always been a major concern for comfort. Currently, the most common method for detecting air conditioner noise is to collect noise levels in decibels and then apply noise reduction strategies based on these levels and preset noise thresholds. However, preset noise reduction levels and strategies cannot meet the personalized needs of different users and are therefore difficult to adapt to their actual requirements. Summary of the Invention

[0004] To address the problem that existing technologies cannot meet the personalized needs of different users, embodiments of this application provide an air conditioning control method, device, system, electronic device, and air conditioner.

[0005] A first aspect of this application provides an air conditioning control method, the method comprising:

[0006] Obtain input parameters, which include the current environmental state, or the input parameters also include at least one of the user's air conditioning usage requirements and air conditioning settings, wherein the environmental state includes the values ​​of environmental factors related to noise during air conditioning operation;

[0007] A noise control strategy is generated based on the input parameters and the strategy generation model.

[0008] Based on the noise control strategy and the current environmental state, predict the updated environmental state after adopting the noise control strategy;

[0009] In response to user feedback regarding the updated environmental state and / or the noise control strategy, the noise control strategy may be selectively executed or updated.

[0010] In conjunction with the first aspect, in one implementation of the first aspect, the environmental factors include at least one of noise decibels, indoor temperature, fan speed, compressor frequency, and air conditioning usage time.

[0011] In conjunction with the first aspect, in one implementation of the first aspect, the policy generation model is a deep reinforcement learning model, and the step of generating a noise control policy based on the input parameters and the policy generation model includes:

[0012] The input parameters are input into the deep reinforcement learning model to obtain the noise control strategy output by the deep reinforcement learning model, wherein the deep reinforcement learning model takes the input parameters as input and the noise control strategy as output.

[0013] In conjunction with the first aspect, in one implementation of the first aspect, the deep reinforcement learning model is trained in the following manner:

[0014] Obtain the first value estimate of executing the first noise control strategy under the environmental conditions at the first moment;

[0015] Obtain the second value estimate of executing the second noise control strategy under the environmental conditions at the second time point;

[0016] Based on the first value estimate and the second value estimate, update the model parameters of the deep reinforcement learning model;

[0017] Wherein, the first noise control strategy is a noise control strategy obtained by the deep reinforcement learning model with the first environmental state at the first moment, the user's first air conditioning usage needs and the first air conditioning setting value as input, the environmental state at the second moment is the environmental state predicted based on the first environmental state at the first moment and the first action, and the second noise control strategy is a noise control strategy obtained by the deep reinforcement learning model with the second environmental state, the user's second air conditioning usage needs and the second air conditioning setting value as input.

[0018] In conjunction with the first aspect, in one implementation of the first aspect, predicting the updated environmental state after adopting the noise control strategy based on the noise control strategy and the current environmental state includes:

[0019] The noise control strategy and the current environmental state are input into the state prediction model to obtain the updated environmental state output by the state prediction model, wherein the state prediction model takes the noise control strategy and the current environmental state as input and the updated environmental state as output.

[0020] In conjunction with the first aspect, in one implementation of the first aspect, the training data used by the state prediction model during the training process includes:

[0021] The first air conditioning setting value and the first environmental state corresponding to the first air conditioning setting value within the target time period;

[0022] The first air conditioning setting includes at least one of the following: target temperature, target fan speed, and target operating mode;

[0023] The first environmental condition includes at least one of the following: actual indoor temperature, actual fan speed, actual noise level in decibels, actual air conditioning usage time, and actual compressor frequency.

[0024] In conjunction with the first aspect, in one implementation of the first aspect, selectively executing or updating the noise control policy in response to user feedback regarding the updated environmental state and / or the noise control policy includes:

[0025] If the user is satisfied with the updated environmental state and / or the noise control strategy, the noise control strategy is executed; and / or,

[0026] If the user is dissatisfied with the updated environmental state and / or the noise control strategy, the noise control strategy will be regenerated based on the user's improved air conditioning usage needs and the current environmental state.

[0027] In conjunction with the first aspect, in one implementation of the first aspect, the method further includes:

[0028] The air conditioner usage requirements are extracted from user input based on a language model, and the form of the air conditioner usage requirements conforms to the format requirements of the input parameters of the strategy generation model.

[0029] In conjunction with the first aspect, in one implementation of the first aspect, the method further includes:

[0030] Send the noise control strategy and the updated environment status corresponding to the noise control strategy to the user, and obtain user feedback on the updated environment status and / or the noise control strategy;

[0031] The number of strategy generation models is at least one, and the number of noise control strategies is at least one.

[0032] A second aspect of this application provides an air conditioning control device, the air conditioning control device comprising:

[0033] An input parameter acquisition module is used to acquire input parameters, which include the current environmental state, or the input parameters also include at least one of the user's air conditioning usage requirements and air conditioning settings, wherein the environmental state includes the values ​​of environmental factors related to noise during air conditioning operation;

[0034] A strategy generation model is used to generate a noise control strategy based on the input parameters;

[0035] A state prediction model is used to predict the updated environmental state after adopting the noise control strategy, based on the noise control strategy and the current environmental state.

[0036] The strategy processing module is used to selectively execute the noise control strategy or update the noise control strategy in response to user feedback regarding the updated environmental state and / or the noise control strategy.

[0037] In conjunction with the second aspect, one implementation of the second aspect also includes:

[0038] The strategy display module is used to send or present the noise control strategy and the updated environmental status to the user.

[0039] The user feedback module is used to obtain user feedback on the updated environmental status and / or the noise control strategy.

[0040] A third aspect of this application provides an electronic device, the electronic device comprising:

[0041] Memory, used to store computer instructions;

[0042] A processor is used to invoke the computer instructions to implement the air conditioning control method described in the first aspect of the embodiments of this application.

[0043] The fourth aspect of this application provides an air conditioner, which employs the air conditioner control method described in the first aspect of the application, or includes the air conditioner control device of the second aspect of the application, or includes the electronic device of the third aspect of the application.

[0044] A fifth aspect of this application provides an air conditioning control system, the air conditioning control system comprising:

[0045] The server employs the air conditioning control method described in the first aspect of the embodiments of this application.

[0046] User equipment is configured to receive the updated environmental status and / or the noise control policy sent by the server, present the updated environmental status and / or the noise control policy to the user for selection, and,

[0047] If the user feedback regarding the updated environmental status and / or the noise control strategy is satisfactory, a control command is sent to the air conditioner according to the noise control strategy, or...

[0048] If a user is dissatisfied with the updated environmental state and / or the noise control strategy, the updated user request will be sent to the server so that the server can regenerate the noise control strategy.

[0049] The method provided in this embodiment has several advantages. First, it generates a noise control strategy based on input parameters and a strategy generation model, which offers greater flexibility compared to the fixed use of preset noise control schemes in existing technologies. Second, it predicts the updated environmental state after adopting the noise control strategy based on the noise control strategy and the current environmental state, which helps to present the possible environmental state corresponding to the noise control strategy to the user, thereby providing the user with an objective expectation. Third, in response to user feedback on the updated environmental state and / or noise control strategy, it selectively executes the noise control strategy or updates the noise control, thereby enabling the noise control strategy to meet the user's personalized needs and improve the comfort experience of air conditioning control. Attached Figure Description

[0050] Figure 1 This is a flowchart illustrating an air conditioning control method according to an embodiment of this application;

[0051] Figure 2 This is a flowchart illustrating an air conditioning control method according to an embodiment of this application;

[0052] Figure 3 This is a schematic diagram illustrating the training process of a deep reinforcement learning model according to an embodiment of this application;

[0053] Figure 4 This is a block diagram of an air conditioning control device according to an embodiment of this application. Detailed Implementation

[0054] Embodiments of the present invention are described in detail below. Examples of these embodiments are illustrated in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and intended to explain the present invention, and should not be construed as limiting the present invention.

[0055] Throughout the specification and claims, the following terms will have at least the meaning explicitly associated herein, unless the context otherwise requires. The meanings defined below are not intended to limit the terms, but are merely illustrative examples.

[0056] In the description of this invention, the phrase "in one embodiment" does not necessarily refer to the same embodiment, although it may refer to the same embodiment. Similarly, the phrase "in some embodiments," as used herein, does not necessarily refer to the same embodiment when used multiple times, although it may refer to the same embodiment. As used herein, the term "or" is an inclusive "or" operator and is equivalent to the term "and / or," unless the context clearly specifies otherwise. The term "based on" is not exclusive and allows for reliance on additional factors not described, unless the context clearly specifies otherwise. The word "exemplary" herein means "used as an example, instance, or illustration." Any embodiment described herein as "exemplary" is not necessarily to be construed as superior to or better than other embodiments. The scope of this invention is limited only by the scope of the appended claims, and any examples set forth in this specification are not intended to be limiting, but merely illustrate some of the many possible embodiments of the claimed invention. The various embodiments provided in this invention should not be construed as limiting the scope of protection of this invention.

[0057] In the description of this invention, it should be understood that the terms "center," "longitudinal," "lateral," "length," "width," "thickness," "upper," "lower," "front," "rear," "left," "right," "vertical," "horizontal," "top," "bottom," "inner," "outer," "clockwise," and "counterclockwise," etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. They are only for the convenience of describing this invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on this invention.

[0058] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of this invention, "a plurality of" means two or more, unless otherwise explicitly specified.

[0059] In this invention, unless otherwise explicitly specified and limited, the terms "installation," "connection," "linking," and "fixing," etc., should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral part; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; they can refer to the internal communication of two components or the interaction between two components. Those skilled in the art can understand the specific meaning of the above terms in this invention according to the specific circumstances.

[0060] In this invention, unless otherwise explicitly specified and limited, "above" or "below" the second feature can include direct contact between the first and second features, or contact between the first and second features through another feature between them. Furthermore, "above," "over," and "on top" of the second feature includes the first feature directly above or diagonally above the second feature, or simply indicates that the first feature is at a higher horizontal level than the second feature. "Below," "below," and "under" the second feature includes the first feature directly below or diagonally below the second feature, or simply indicates that the first feature is at a lower horizontal level than the second feature.

[0061] The terms and names used in this application are explained below. These explanations are intended to help those skilled in the art understand the embodiments of this application, and are not intended to limit the understanding of the embodiments of this application.

[0062] Deep Reinforcement Learning (DRL) is a technique that combines Reinforcement Learning (RL) with Deep Learning (DL), aiming to enable agents to learn optimal behavioral strategies through interaction with the environment.

[0063] The working principle of a deep reinforcement learning model can be summarized in the following steps: Environment perception: The agent acquires state information about the environment through sensors or interfaces; State representation: The state information usually needs to be preprocessed (e.g., image scaling, feature extraction) before being input into the deep neural network; Decision making: Based on the input state information, the neural network outputs the probability distribution of one or more actions or directly outputs an action; Action execution: The agent executes the corresponding behavior according to the action or action probability distribution output by the network; Receiving feedback: After executing the action, the agent receives an immediate reward from the environment and enters a new state; Model update: Based on the agent's behavioral results (reward and new state), the parameters of the neural network are updated using a learning algorithm (e.g., Q-learning, policy gradient) to make better decisions in the future.

[0064] The main components of deep reinforcement learning include: Agent: the subject that learns and executes the policy; Environment: the external world in which the agent resides, providing state information and rewards; State: describing the current state of the environment; Action: the behavior that the agent can perform; Reward: the feedback the agent receives from the environment after performing an action, used to evaluate the quality of the action; Policy: the principle by which the agent chooses actions, which can be deterministic or random; Value Function: evaluating the value of a state or state-action pair, such as the Q function or V function; Deep Neural Network: the model used to learn the policy or value function.

[0065] LSTM (Long Short-Term Memory) is a special type of Recurrent Neural Network (RNN) designed to address the vanishing or exploding gradient problems encountered by traditional RNNs when processing long sequences of data. LSTM introduces a gate mechanism, enabling the network to learn long-term dependencies and effectively retain important information.

[0066] Air conditioner noise has always been a major concern for comfort. Currently, the most common method for detecting air conditioner noise is to collect noise levels in decibels and then apply noise reduction strategies based on these levels and preset noise thresholds. However, preset noise reduction levels and strategies cannot meet the personalized needs of different users and are therefore difficult to adapt to their actual requirements.

[0067] For example, some noise reduction control methods acquire air conditioner noise and ambient noise, compare the difference between the two noise levels with a preset difference, and then adjust the indoor fan speed and compressor frequency to control the air conditioner noise to approximate the ambient noise. This control strategy, based on the noise difference, only considers the noise factor and does not take into account the impact of reducing fan speed and compressor frequency on the different noise experiences experienced by different users.

[0068] For example, in noise control methods for fresh air conditioners, a pre-configured fan speed control strategy is used based on the current ambient noise level to match the fresh air speed with the ambient noise, thereby reducing noise. This also controls the air conditioner's operation based on objective environmental noise factors, without considering people's subjective perceptions of noise, and the use of preset control strategies cannot meet the different usage needs of different groups of people.

[0069] In response, this application provides an air conditioning control method suitable for meeting users' personalized needs, thereby improving user comfort. For example... Figure 1 The diagram shown is a schematic flowchart of an air conditioning control method according to an embodiment of this application. (Refer to...) Figure 1 The method includes the following processing steps.

[0070] S100: Get input parameters.

[0071] The input parameters include the current environmental status, or the input parameters may also include at least one of the user's air conditioning usage requirements and air conditioning settings.

[0072] For example, the environmental state refers to the actual operating state of the air conditioner, including the values ​​of environmental factors related to noise during air conditioner operation. For instance, environmental factors include at least one of the following: noise level in decibels, indoor temperature, fan speed, compressor frequency, and air conditioner usage time. The noise level in decibels includes the decibels at the air outlet of the indoor unit and / or the decibels at the outdoor unit.

[0073] For example, a user's air conditioning usage needs may include target control parameters of the air conditioner (i.e., which parameters of the air conditioner they want to adjust), such as at least one of compressor frequency, indoor fan speed, indoor unit, outdoor unit, etc.

[0074] For example, the air conditioner setpoint reflects the intervention in the future (e.g., the next moment after the current moment) state of the air conditioner. For example, it includes at least one of the following: user-set target temperature (which affects the outdoor unit compressor), target fan speed (which affects the fan speed), and target operating mode (which affects the operation of the whole unit, such as cooling mode, heating mode; or, for example, cooling sleep mode, cooling energy saving mode, etc.).

[0075] S102: Generate a noise control strategy based on the input parameters and the strategy generation model.

[0076] The policy generation model takes at least the input parameters as input and the noise control policy as output.

[0077] S104: Based on the noise control strategy and the current environmental state, predict the updated environmental state after the noise control strategy is adopted.

[0078] For example, a state prediction model can be used to predict the environmental state at the next moment based on the noise control strategy at a previous moment and the current environmental state (i.e., update the environmental state).

[0079] S106: In response to user feedback regarding updated environmental conditions and / or noise control strategies, selectively implement or update the noise control strategy.

[0080] For example, if the user feedback is satisfactory, the noise control strategy is implemented; if the user is dissatisfied, the noise control strategy is updated.

[0081] The method provided in this embodiment has several advantages. First, it generates a noise control strategy based on input parameters and a strategy generation model, which offers greater flexibility compared to the fixed use of preset noise control schemes in existing technologies. Second, it predicts the updated environmental state after adopting the noise control strategy based on the noise control strategy and the current environmental state, which helps to present the possible environmental state corresponding to the noise control strategy to the user, thereby providing the user with an objective expectation. Third, in response to user feedback on the updated environmental state and / or noise control strategy, it selectively executes the noise control strategy or updates the noise control, thereby enabling the noise control strategy to meet the user's personalized needs and improve the comfort experience of air conditioning control.

[0082] Furthermore, when the input parameters include the user's air conditioning usage requirements, it is more beneficial for the embodiments of this application to provide a noise control strategy that meets the user's needs, thereby satisfying the needs of different users and improving comfort and intelligent experience.

[0083] Furthermore, when the input parameters include the air conditioner's setpoint, it is beneficial to generate a noise control strategy that better reflects the actual operating characteristics of the air conditioner by comprehensively considering both the actual operating state of the air conditioner and its ideal operating state (i.e., the state reflected by the user's setpoint).

[0084] Optionally, in one implementation of this embodiment, in S102, the policy generation model is a deep reinforcement learning model, and the noise control policy is generated according to the input parameters and the policy generation model, including: inputting the input parameters into the deep reinforcement learning model to obtain the noise control policy output by the deep reinforcement learning model, wherein the deep reinforcement learning model takes the input parameters as input and the noise control policy as output.

[0085] Using this implementation method, a noise control strategy corresponding to the current environmental state (or the current environmental state and user needs) can be obtained based on a deep reinforcement learning model.

[0086] It should be noted that deep reinforcement learning models are part of Markov Decision Processes (MDPs). Therefore, other MDPs can be used instead of deep reinforcement learning models in other embodiments.

[0087] For example, a deep reinforcement learning model is trained in the following manner.

[0088] First, a first value estimate is obtained for executing the first noise control strategy under the environmental state at the first time step. Then, a second value estimate is obtained for executing the second noise control strategy under the environmental state at the second time step. Finally, the model parameters of the deep reinforcement learning model are updated based on the first and second value estimates. The first noise control strategy is derived from the first environmental state at the first time step, the user's first air conditioning usage requirement, and the first air conditioning setting (here, "first time step" refers to a historical time step corresponding to the training data; the first environmental state, the user's first air conditioning usage requirement, and the first air conditioning setting are all historical data) as inputs. The environmental state at the second time step is the environmental state predicted based on the environmental state at the first time step and the first action. The second noise control strategy is derived from the deep reinforcement learning model using the second environmental state at the second time step, the user's second air conditioning usage requirement, and the second air conditioning setting (here, "second time step" refers to a historical time step corresponding to the training data; the second environmental state, the user's second air conditioning usage requirement, and the second air conditioning setting are all historical data) as inputs. The second time step is the time step following the first time step.

[0089] The first environmental state / second environmental state includes at least one of the following values: noise decibels, indoor temperature, fan speed, compressor frequency, and air conditioning usage time; the user's first air conditioning usage requirement / user's second air conditioning usage requirement includes at least one of the following: compressor frequency, indoor fan speed, indoor unit, outdoor unit, etc.; the first air conditioning setting value / second air conditioning setting value includes at least one of the following: target temperature, target fan speed, and target operating mode.

[0090] In this way, the user's air conditioning needs and air conditioning settings are taken into account during the model training process, so that the trained deep reinforcement learning model can output noise control strategies that meet user needs and the characteristics of the air conditioner itself with a high probability.

[0091] For example, the environmental state at the second time step can be predicted using the LSTM model mentioned below.

[0092] Optionally, in one implementation of this application embodiment, in S104, predicting the updated environmental state after adopting the noise control strategy based on the noise control strategy and the current environmental state includes: inputting the noise control strategy and the current environmental state into the state prediction model to obtain the updated environmental state output by the state prediction model, wherein the state prediction model takes the noise control strategy and the current environmental state as input and the updated environmental state as output.

[0093] For example, considering that the air conditioner's operating state is a time-continuous metastable state, the inventors can use an LSTM (Long Short-Term Memory Recurrent Neural Network) model to predict noise levels based on the air conditioner's operating parameters. This helps avoid the negative user experience caused by repeated parameter adjustments and excessive parameter changes. That is, the state prediction model can be an LSTM model. It should be noted that LSTM is an optimization algorithm of Recurrent Neural Networks (RNNs), and alternatively, other RNNs can be used in other embodiments. Examples include: basic RNNs, RNN variants, LSTM variants, etc.

[0094] This implementation method can relatively accurately predict the environmental state after implementing the noise control strategy under the current environmental conditions, based on the noise control strategy and the current environmental state, which is beneficial for providing users with an intuitive reference.

[0095] For example, the training data used by the state prediction model during training includes: a first air conditioning setpoint and a first environmental state within the target time period; wherein, the first air conditioning setpoint includes at least one value of target temperature, target fan speed, and target operating mode, and the first environmental state includes at least one value of actual indoor temperature, actual fan speed, actual noise level in decibels, actual air conditioning usage time, and actual compressor frequency. Of course, the air conditioning setpoint is not limited to this example; any setpoint that affects the operating state of the air conditioner can be included.

[0096] By using the training data provided in this embodiment and the conventional LSTM model training method, a state prediction model that can accurately predict the environmental state at the next moment can be obtained.

[0097] Optionally, in one implementation of this embodiment, in S106, the noise control strategy can be executed if the user is satisfied with the updated environmental state and / or noise control strategy. This allows the air conditioning to meet the user's actual needs.

[0098] Optionally, in one implementation of this embodiment, in S106, if the user is dissatisfied with the updated environmental state and / or noise control strategy, the noise control strategy is regenerated based on the user's improved air conditioning usage needs and the current environmental state. This iterative process helps generate a noise control strategy that meets the user's needs.

[0099] Optionally, in one implementation of this embodiment, the air conditioning usage requirements can be extracted from the user's input requirements based on a language model (e.g., a large language model, GPT, Wenxin Yiyan, Doubao, Tongyi Qianwen, or a specially trained language model), and the form of the air conditioning usage requirements conforms to the format requirements of the input parameters of the strategy generation model.

[0100] For example, user requests can be voice or text, and the content can be natural language, such as: "I feel very hot, but I also find the room very noisy." The corresponding air conditioning usage requests could be: cooling / heating - temperature adjustment - compressor frequency; noise - indoor fan speed, compressor power; indoor - indoor unit.

[0101] Optionally, in one implementation of this application embodiment, before S106, the method may further include: sending a noise control strategy and an updated environmental state corresponding to the noise control strategy to the user, and obtaining user feedback on the updated environmental state and / or the noise control strategy. Here, the number of strategy generation models is at least one, and the number of noise control strategies is at least one. In other words, input parameters can be input into multiple different strategy generation models to obtain different noise control strategies, and then the different noise control strategies can be input into the same state prediction model to obtain environmental states corresponding to different noise control strategies, which are then simultaneously provided to the user for selection.

[0102] Figure 2 This is a flowchart illustrating an air conditioning control method according to an embodiment of this application. More specifically, it is a flowchart illustrating a method for controlling air conditioning noise based on deep reinforcement learning. (Refer to...) Figure 2 The method includes the following processing steps.

[0103] 1) Information collection.

[0104] Specifically, since air conditioner noise is directly affected by the frequency of the air conditioner's fan and compressor, noise sensors are placed at the air outlet of the indoor unit and the outdoor unit (e.g., at the outdoor unit casing or the air outlet of the outdoor unit's fan) to collect noise decibels. Simultaneously, the air conditioner's main control chip uses the air conditioner's setpoint, actual temperature, fan speed, compressor power, air conditioner usage time, and the noise decibel values ​​collected by the noise sensors as current vector information.

[0105] The air conditioner settings include: temperature, which affects the outdoor unit compressor; fan speed, which affects the fan speed; and operating mode, which affects the overall operation of the unit. Operating mode can be cooling mode, heating mode, or further subdivided into cooling sleep mode, cooling energy-saving mode, etc. The overall state of the air conditioner is a dynamic process, and the air conditioner settings can be considered as interventions into the future state of the air conditioner. Air conditioner settings should not be limited to the three examples mentioned above; any setting that affects the air conditioner's operating state should be considered a variable. LSTM will automatically assign weights to these variables during training based on their impact on noise levels.

[0106] The vector information above represents the parameters that have a significant impact on air conditioner noise. In practice, it is not limited to the examples above; other information can be added as vectors. The LSTM model will automatically assign weights to each parameter based on the calculation results.

[0107] In a practical application, the collected information can be transmitted to the cloud platform via the main control chip-gateway-cloud platform, and the cloud platform records the collected information as a set of vectors st.

[0108] 2) LSTM model creation environment status.

[0109] Specifically, an LSTM model can be trained based on data collected over a period of time T, including air conditioner setpoints, actual temperatures, fan speeds, compressor power, and noise levels. The environmental state *st* at a given time *t* is then input into the LSTM model, predicting the environmental state *st+1* at time *t+1*, based on the noise level, actual temperature, and fan speed. This information at time *t+1* is virtual information predicted by the LSTM and does not affect the actual environment.

[0110] In this embodiment, environmental conditions, including noise levels in decibels, fan speed, and indoor temperature, are used as examples for illustration. In other embodiments, compressor frequency, air conditioning usage time (indicating how long the air conditioner has been used continuously), and operating mode may also be included.

[0111] The LSTM model is explained below:

[0112] ① Calling the LSTM model:

[0113] When calling the model, the system checks whether the user has updated the LSTM model: if not, the pre-trained model is used, which uses noise judgment standards from a laboratory environment. If the model has been updated, the user's personal LSTM model is used.

[0114] Using the current vector st and the sequence of running states over a past period of time ht-1 as input (ht-1=0 when t=0), predict the current environment state st+1.

[0115] it = σ(Wi[ht-1,st] + bi)

[0116] ft=σ(Wf[ht-1,st]+bf)

[0117] ct=ftct-1+ittanh(Wc[ht-1,st]+bc)

[0118] ot=σ(Wo[ht-1,st]+bo)

[0119] st+1 = ottanh(ct)

[0120] Where σ is the standard sigmoid function; i, f, o, and c are the input gate, forget gate, output gate, and memory unit, respectively; bi, bf, bo, and bc are the bias vectors of the input gate, forget gate, output gate, and memory unit, respectively; and W is the weight matrix between each unit and the gate vector.

[0121] ② Model training and update:

[0122] When the predicted noise level and indoor temperature decibels deviate too much from the actual values, the dataset is constructed using n running information vectors s = (s1, s2, ..., sn) from the last model training. The LSTM model is then trained and updated to adapt to the air conditioner's own state and environmental factors. The impact of the air conditioner's noise and indoor temperature is divided based on the length of the time series.

[0123] Short-time series: Changes in values ​​within a short period of time can have a significant impact on dust accumulation in air conditioners. It is necessary to remember or forget these changes in a short period of time, such as fan speed and compressor frequency.

[0124] Long-term time series: Changes in values ​​over a long period of time will have an impact on air conditioner noise. There is no need to remember or forget changes in a short period of time, such as the duration of air conditioner use.

[0125] Under normal circumstances, short-time and long-time time series do not need to be manually divided and labeled. The LSTM model can automatically remember and forget based on the influence of these factors on air conditioning noise and indoor temperature.

[0126] The output of the information vector s after passing through the LSTM model is y:

[0127] {y1,y2,...,yn}=LSTM{s1,s2,..,sn}

[0128] Loss function:

[0129] The smaller the loss function, the higher the accuracy of the LSTM model, and the more accurately it can predict the noise generated by the air conditioner and the changes in indoor temperature after the state changes.

[0130] 3) Different noise control strategies can be obtained from multiple strategy generation models. That is, multiple different noise control strategies can be obtained based on multiple trained strategy generation models. Of course, in other embodiments, multiple different noise control strategies can also be obtained from a single strategy generation model.

[0131] 4) Predict the operating results of the noise control strategy based on the LSTM model (i.e., predict the environmental state at the next moment).

[0132] 5) Obtain user feedback. If the user feedback is satisfactory, save the satisfied deep reinforcement learning model and update the LSTM model. If the user feedback is unsatisfactory and suggestions for improvement are made, return to step 3 based on the user's suggestions.

[0133] Specifically, the system can present users with multiple optimal value function action plans (an) obtained through various deep reinforcement learning models, along with the corresponding expected reductions in noise levels (decibels), indoor temperature, and fan speed. Users can then select the appropriate plan based on their needs. If a user is not satisfied with the proposed plan, they can request improvements. The deep reinforcement learning model will then adjust its parameters and continue training until the user is satisfied. At this point, the deep reinforcement learning model for that plan will be saved as the preferred model, and the parameters in the LSTM model will be updated to increase prediction accuracy.

[0134] 6) Air conditioning control based on local applications. Specifically, the cloud platform provides calculated noise control strategies. Users select appropriate noise control strategies on the APP (i.e., local application). The APP issues control commands to the air conditioner based on the selected noise control strategy, and the air conditioner's main control chip adjusts the air conditioner load to reduce noise.

[0135] For example, after the air conditioner has been running for a period of time, the information acquisition module collects vector information over that period and transmits it to the cloud platform to create an LSTM environment. The deep reinforcement learning module calculates several different noise control strategies and imports the air conditioner parameters corresponding to these strategies into the previously created LSTM model to obtain the expected air conditioner operation results (i.e., the predicted environmental state) under the proposed noise control strategies. The cloud platform then distributes the noise control strategies and the predicted air conditioner operation results to the user's app. For example, reducing the fan speed reduces noise by 30 decibels; or reducing the compressor frequency reduces noise by 40 decibels, and the indoor temperature is expected to increase by 2 degrees Celsius. When the user feedback is unsatisfactory (e.g., the user feels the noise could be slightly higher for a more comfortable indoor temperature), the deep reinforcement learning can adjust the noise reduction strategy according to the user's needs and recalculate. This process continues until the user is satisfied. The app then issues control commands to the air conditioner based on the user's selected noise control strategy. The air conditioner's main control chip receives the control commands and adjusts the air conditioner load accordingly. The deep reinforcement learning model saves this preferred scheme and updates the LSTM model to increase prediction accuracy.

[0136] Figure 3 This is a schematic diagram illustrating the training process of a deep reinforcement learning model according to an embodiment of this application. The following is in conjunction with... Figure 3 The training process of a deep reinforcement learning model is explained.

[0137] In this embodiment, branches can be formed by different deep reinforcement learning models with different strategies, such as: hetero-strategy models that use noise reduction as the behavioral strategy and user comfort as the target strategy; homo-strategy models that use noise reduction as both the behavioral and target strategies; DQN networks that encourage exploratory behavior, etc.

[0138] In this embodiment, refer to Figure 2 The training process is as follows.

[0139] ① Based on user needs, a deep reinforcement learning model is used to simulate the air conditioner setpoint and the current environmental state st (e.g., fan speed, compressor frequency, and air conditioner usage time) to obtain an action a.

[0140] For example, user needs can be converted into user air conditioning usage needs through GPT models or direct function settings. For example, the conversion can yield: cooling / heating - temperature adjustment - compressor frequency; noise - indoor fan speed and compressor power; indoor / outdoor - indoor unit, outdoor unit, and so on.

[0141] Regarding action a, for example, if my requirement is "I feel very hot, but I also feel that the room is very noisy," then action a could be: lowering the air conditioner temperature setting, increasing the compressor frequency, and lowering the indoor fan speed. In the embodiments of this application, any parameters related to noise and user experience that are controllable by the air conditioner can be included in action a, including but not limited to indoor and outdoor fan speeds, compressor frequency, and air conditioner set temperature.

[0142] ② The action 'a' is combined with the current environment 'st' as input to the LSTM model. The LSTM predicts the noise, indoor temperature, and fan speed at a future time interval 't+1' as the state 'st+1'. The length of each time interval can be customized as needed, for example, 10 minutes.

[0143] ③ The deep reinforcement learning model scores the returned function reward R(st,a,st+1) γ (γ∈[0,1]) based on the target policy π. The action value function can be expressed as:

[0144] Qπ(st,at)=E[Rt+γ·Rt+1+γ2·Rt+2+...+γn·Rt+n|St=st,At=at]

[0145] Where Rt+n is the reward at time t+n predicted by the LSTM model based on the current state st and action a.

[0146] The target strategy π can be one or a weighted combination of factors such as action a (noise), indoor temperature value, or user habit conformity (comfort).

[0147] ④ Next, input st, the user's air conditioning usage needs, and the air conditioning setpoint into the deep reinforcement learning model to obtain the action at. The action at is then predicted by the LSTM model to obtain a value estimate qt and a new environment st+1. Then, input the new environment st+1, the user's air conditioning usage needs, and the air conditioning setpoint into the deep reinforcement learning model to obtain a new action at+1 and a new value qt+1. It can be considered that the new environment st+1 is the target obtained based on the target policy π, thus generating an error between the target and the current value.

[0148] L(w)=1 / 2[Q(st,at;w)-(rt+γ·Q(st+1,at+1;w))]2

[0149] Update the current model parameter w:

[0150]

[0151] ⑤ Deep reinforcement learning models are trained based on the value score of the target policy π and the behavioral policy (which can be the same as or different from the target policy; different deep reinforcement learning models have different target policies and behavioral rewards) to obtain multiple optimal value function action schemes an.

[0152] In this embodiment, the training of LSTM and deep reinforcement learning is carried out on the cloud platform. After the training is completed, the model is saved on the cloud platform, and the energy management solution obtained by deep reinforcement learning training is sent back to the user's smart device APP through the cloud platform-gateway-smart device APP.

[0153] Figure 4 This is a block diagram of an air conditioning control device according to an embodiment of this application. (Refer to...) Figure 4 The air conditioning control device includes the following modules.

[0154] The input parameter acquisition module is used to acquire input parameters, which include at least one of the user's air conditioning usage requirements and air conditioning settings. The environmental status includes the values ​​of environmental factors related to noise during air conditioning operation.

[0155] A policy generation model is used to generate noise control policies based on input parameters.

[0156] A state prediction model is used to predict the updated environmental state after the noise control strategy is adopted, based on the noise control strategy and the current environmental state.

[0157] The policy processing module is used to selectively execute or update the noise control policy in response to user feedback regarding updated environmental conditions and / or noise control policies.

[0158] The air conditioning control device provided in this embodiment has several advantages. First, it generates a noise control strategy based on input parameters and a strategy generation model, which offers greater flexibility compared to the fixed use of preset noise control schemes in existing technologies. Second, it predicts the updated environmental state after adopting the noise control strategy based on the noise control strategy and the current environmental state, which helps to present the possible environmental state corresponding to the noise control strategy to the user, thereby providing the user with an objective expectation. Third, in response to user feedback on the updated environmental state and / or noise control strategy, it selectively executes the noise control strategy or updates the noise control, thereby enabling the noise control strategy to meet the user's personalized needs and improve the comfort experience of air conditioning control.

[0159] Optionally, in one implementation of this embodiment, the air conditioning control module further includes: a strategy display module, used to send or present the noise control strategy and the updated environmental status to the user; and a user feedback module, used to obtain user feedback on the updated environmental status and / or the noise control strategy.

[0160] For a detailed description of each module in this embodiment, please refer to the description in the previous method embodiment; it will not be repeated here.

[0161] This application also provides an electronic device, which includes a memory and a processor. The memory is used to store computer instructions, and the processor is used to invoke the computer instructions to implement the methods provided in this application (e.g., air conditioning control method, model training method).

[0162] For example, an electronic device may be a chip, a circuit board with integrated chips, or hardware containing a circuit board. For instance, an electronic device may be a user terminal device.

[0163] This application also provides an air conditioner, which uses the air conditioner control method or model training method provided in this application, or the air conditioner integrates the electronic equipment provided in this application.

[0164] This application embodiment also provides an air conditioning control system, the air conditioning control system including:

[0165] The server uses the air conditioning control method provided in the embodiments of this application.

[0166] The user equipment is configured to receive updated environmental conditions and / or noise control policies from the server, present the updated environmental conditions and / or noise control policies to the user for selection, and, if the user feedback on the updated environmental conditions and / or noise control policies is satisfactory, send control commands to the air conditioner according to the noise control policies; or, if the user feedback on the updated environmental conditions and / or noise control policies is unsatisfactory, send updated user requirements to the server so that the server can regenerate the noise control policies.

[0167] In addition, the air conditioning control system may also include an air conditioner, whose main control chip is used to receive control commands sent by the user equipment and adjust the corresponding load according to the control commands. For example, when the control command includes the load to be controlled (e.g., compressor frequency, fan speed) and the adjustment range, the corresponding adjustment is performed according to the control command.

[0168] In the above embodiments of this application, the descriptions of each embodiment have their own emphasis. Parts not described in detail in a certain embodiment can be referred to in the relevant descriptions of other embodiments. The steps illustrated in the related flowcharts can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowcharts, in some cases, the steps shown or described may be performed in a different order than that shown here. In other words, the order of steps described in the foregoing embodiments is merely an example. Reasonable adjustments to the order of steps based on the content of the embodiments of this application are also within the protection scope of the embodiments of this application.

[0169] The sequence numbers or order of description of the embodiments in this application are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.

[0170] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. In addition, those skilled in the art can combine and integrate the different embodiments or examples described in this specification.

[0171] Although embodiments of the present invention have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those skilled in the art can make changes, modifications, substitutions and variations to the above embodiments within the scope of the present invention.

Claims

1. An air conditioning control method, characterized in that, The method includes: The system acquires input parameters, including the current environmental state, the user's air conditioning usage requirements, and air conditioning settings. The environmental state includes values ​​of environmental factors related to noise during air conditioning operation. The air conditioning usage requirements include target control parameters for the air conditioning, including at least one of compressor frequency, indoor fan speed, indoor unit speed, and outdoor unit speed. The air conditioning settings reflect intervention in the future state of the air conditioning, including at least one of the user-set target temperature, target fan speed, and target operating mode. A noise control strategy is generated based on the input parameters and the strategy generation model. The noise control strategy and the current environmental state are input into the state prediction model to obtain the updated environmental state output by the state prediction model. The state prediction model takes the noise control strategy and the current environmental state as input and the updated environmental state as output. If the user is satisfied with the updated environmental state and / or the noise control strategy, the noise control strategy is executed; and / or, if the user is not satisfied with the updated environmental state and / or the noise control strategy, the noise control strategy is regenerated based on the user's improved air conditioning usage needs and the current environmental state. Wherein, the policy generation model is a deep reinforcement learning model, and the step of generating a noise control policy based on the input parameters and the policy generation model includes: The input parameters are input into the deep reinforcement learning model to obtain the noise control strategy output by the deep reinforcement learning model, wherein the deep reinforcement learning model takes the current environmental state, the air conditioning usage requirements and the air conditioning setting as input, and the noise control strategy as output.

2. The air conditioning control method according to claim 1, characterized in that, The environmental factors include at least one of the following: noise level (decibels), indoor temperature, fan speed, compressor frequency, and air conditioning usage time.

3. The air conditioning control method according to claim 1, characterized in that, The deep reinforcement learning model was trained in the following manner: Obtain the first value estimate of executing the first noise control strategy under the environmental conditions at the first moment; Obtain the second value estimate of executing the second noise control strategy under the environmental conditions at the second time point; Based on the first value estimate and the second value estimate, update the model parameters of the deep reinforcement learning model; Wherein, the first noise control strategy is a noise control strategy obtained by the deep reinforcement learning model with the first environmental state at the first moment, the user's first air conditioning usage needs and the first air conditioning setting value as input, the environmental state at the second moment is the environmental state predicted based on the first environmental state at the first moment and the first action, and the second noise control strategy is a noise control strategy obtained by the deep reinforcement learning model with the environmental state at the second moment, the user's second air conditioning usage needs and the second air conditioning setting value as input.

4. The air conditioning control method according to claim 1, characterized in that, The training data used in the training process of the state prediction model includes: The first air conditioning setting and the first environmental state within the target time period; The first air conditioning setting value includes at least one of the following: target temperature, target fan speed, and target operating mode. The first environmental condition includes at least one of the following: actual indoor temperature, actual fan speed, actual noise level in decibels, actual air conditioning usage time, and actual compressor frequency.

5. The air conditioning control method according to claim 1, characterized in that, The method further includes: The air conditioner usage requirements are extracted from user input based on a language model, and the form of the air conditioner usage requirements conforms to the format requirements of the input parameters of the strategy generation model.

6. The air conditioning control method according to claim 1, characterized in that, The method further includes: Send the noise control strategy and the updated environment status corresponding to the noise control strategy to the user, and obtain user feedback on the updated environment status and / or the noise control strategy; The number of strategy generation models is at least one, and the number of noise control strategies is at least one.

7. An air conditioning control device, characterized in that, The air conditioning control device includes: The input parameter acquisition module is used to acquire input parameters, including the current environmental state, the user's air conditioning usage requirements, and air conditioning settings. The environmental state includes values ​​of environmental factors related to noise during air conditioning operation. The air conditioning usage requirements include target control parameters for the air conditioning, including at least one of compressor frequency, indoor fan speed, indoor unit speed, and outdoor unit speed. The air conditioning settings reflect intervention in the future state of the air conditioning, including at least one of the user-set target temperature, target fan speed, and target operating mode. A strategy generation model is used to generate a noise control strategy based on the input parameters; A state prediction model is used as input to the noise control strategy and the current environmental state, and outputs an updated environmental state. The strategy processing module is used to execute the noise control strategy when the user is satisfied with the updated environmental state and / or the noise control strategy; and / or, when the user is not satisfied with the updated environmental state and / or the noise control strategy, to regenerate the noise control strategy based on the user's improved air conditioning usage needs and the current environmental state. The strategy generation model is a deep reinforcement learning model, which takes the current environmental state, the air conditioning usage requirements, and the air conditioning settings as inputs, and the noise control strategy as output.

8. The air conditioning control device according to claim 7, characterized in that, Also includes: The strategy display module is used to send or present the noise control strategy and the updated environmental status to the user. The user feedback module is used to obtain user feedback on the updated environmental status and / or the noise control strategy.

9. An electronic device, characterized in that, The electronic device includes: Memory, used to store computer instructions; A processor is configured to invoke the computer instructions to implement the air conditioning control method as described in any one of claims 1-6.

10. An air conditioner, characterized in that, The air conditioner employs the air conditioner control method as described in any one of claims 1-6, or includes the air conditioner control device as described in claim 7 or 8, or includes the electronic device as described in claim 9.

11. An air conditioning control system, characterized in that, The air conditioning control system includes: The server employs the air conditioning control method as described in any one of claims 1-6; User equipment is configured to receive the updated environment status and / or the noise control policy sent by the server, present the updated environment status and / or the noise control policy to the user for selection, and, If the user feedback regarding the updated environmental status and / or the noise control strategy is satisfactory, a control command is sent to the air conditioner according to the noise control strategy, or... If a user is dissatisfied with the updated environmental state and / or the noise control strategy, the updated user request will be sent to the server so that the server can regenerate the noise control strategy.