Air conditioner and control method and device thereof, storage medium and computer program product
By combining an ambient sound detection and low-power monitoring module with a preset whisper recognition mechanism, the problem of high wake-up noise, inaccurate recognition, and high power consumption in traditional air conditioning voice control systems in low-noise scenarios has been solved, achieving accurate interaction and energy-saving operation in low-noise environments.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-18
- Publication Date
- 2026-04-03
AI Technical Summary
Traditional air conditioner voice control systems suffer from high wake-up noise in low-noise scenarios, insufficient sensitivity in recognizing soft voices, poor environmental adaptability, and high power consumption during monitoring, resulting in a degraded user experience.
An ambient sound detection module is used to acquire noise parameters and scene information, determine low-noise triggering conditions, enable a low-power monitoring module to monitor voice signals, use a preset whisper recognition mechanism for recognition and analysis, and provide control commands in a non-voice manner.
It reduces wake-up noise, improves the recognition sensitivity of whispered and hushed commands, enhances the system's adaptability to low-noise environments, and reduces operating power consumption.
Smart Images

Figure CN121782701A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of air conditioning technology, and specifically relates to an air conditioning control method, device, air conditioner, storage medium, and computer program product. Background Technology
[0002] Voice control functionality has become a mainstream feature in air conditioners. Traditional air conditioner voice control solutions generally employ a uniform high-volume wake-up and recognition logic. On the one hand, voice wake-up requires users to issue wake-up commands at high sound pressure levels. In low-noise environments such as bedrooms and conference rooms at night, this wake-up method is prone to noise interference, disrupting the tranquility of the environment and affecting others' rest or meeting progress. On the other hand, to ensure recognition success rates, the system often uses microphone arrays with fixed sensitivity. This not only fails to accurately capture whispered commands from users in low-noise environments, leading to recognition failures when controlling from a distance or in a soft voice, but also results in unnecessary energy waste because the voice recognition module needs to maintain a high-power, full-volume listening state around the clock. Furthermore, it cannot dynamically adjust the recognition strategy based on the intensity of ambient noise.
[0003] In addition, traditional systems often provide feedback via voice broadcast after executing commands. This feedback method can further exacerbate noise pollution in low-noise scenarios and cannot meet users' needs for silent control and privacy protection in private scenarios.
[0004] In summary, existing technologies cannot simultaneously achieve undisturbed wake-up in low-noise scenarios, accurate whisper recognition, low-power operation, and silent feedback, resulting in a significant decline in the user experience of air conditioner voice control in sensitive scenarios such as bedrooms and conference rooms.
[0005] The above content is only used to help understand the technical solution of the present invention and does not represent an admission that the above content is prior art. Summary of the Invention
[0006] The purpose of this invention is to provide an air conditioner control method, device, air conditioner, storage medium, and computer program product to solve the problems of high wake-up noise, insufficient sensitivity in recognizing whispered commands, poor environmental adaptability, and high power consumption in related air conditioner voice control systems. This invention aims to reduce wake-up noise in air conditioner voice control, improve the sensitivity in recognizing whispered and hushed commands, enhance the system's environmental adaptability to low-noise scenarios, and reduce power consumption during voice monitoring.
[0007] This invention provides a control method for an air conditioner, the air conditioner including an ambient sound detection module, a low-power monitoring module, and a voice recognition module; the low-power monitoring module is capable of monitoring voice features in a low-power state; the control method includes: acquiring ambient noise parameters through the ambient sound detection module, determining whether a low-noise trigger condition is met based on the ambient noise parameters and current scene information; if the low-noise trigger condition is met, activating the low-power monitoring module to perform feature monitoring on the user's voice signal; when the detected voice signal matches preset whisper features, the voice recognition module uses a preset whisper recognition mechanism to recognize and parse the voice signal to obtain control commands for the air conditioner; and controlling the air conditioner to execute the control commands.
[0008] In some implementations, determining whether the low-noise triggering condition is met based on the environmental noise parameters and the current scene information includes: if the environmental noise parameters are less than a preset noise threshold, or if the scene information is a preset low-noise sensitive scene, then it is determined that the low-noise triggering condition is met.
[0009] In some implementations, the speech recognition module uses a preset whisper recognition mechanism to recognize and analyze the speech signal to obtain the control commands of the air conditioner, including: extracting the cepstral coefficient features, spectral slope features, and low-frequency energy ratio features of the speech signal, and generating a feature vector; inputting the feature vector into a preset neural network model for joint acoustic and language recognition to obtain the control commands of the air conditioner.
[0010] In some embodiments, the speech recognition module includes a main processing unit and a whisper recognition unit; the method further includes: when the low-noise triggering condition is met, controlling the main processing unit to enter a sleep state and waking up the whisper recognition unit, the whisper recognition unit performing an operation to recognize and analyze the speech signal after being woken up; when the low-noise triggering condition is not met, controlling the whisper recognition unit to enter a sleep state and waking up the main processing unit.
[0011] In some embodiments, the method further includes: after controlling the air conditioner to execute the control command, providing feedback to the user on the command execution status in a non-voice manner; the non-voice manner includes flashing lights and changes in airflow.
[0012] In some embodiments, the method further includes: when the low-noise triggering condition is continuously met and no new voice signal is detected within a preset time period, turning off the voice recognition module and keeping the low-power monitoring module in operation.
[0013] In conjunction with the above method, another aspect of the present invention provides a control device for an air conditioner, the air conditioner including an ambient sound detection module, a low-power monitoring module, and a voice recognition module; the low-power monitoring module is capable of monitoring voice features in a low-power state; the control device includes: a judgment unit configured to acquire ambient noise parameters through the ambient sound detection module, and determine whether a low-noise triggering condition is met based on the ambient noise parameters and current scene information; a monitoring unit configured to activate the low-power monitoring module if the low-noise triggering condition is met, and perform feature monitoring on the user's voice signal through the low-power monitoring module; a recognition unit configured to, when the voice signal is detected to match preset whisper features, enable a preset whisper recognition mechanism to recognize and parse the voice signal to obtain control commands for the air conditioner; and a control unit configured to control the air conditioner to execute the control commands.
[0014] In some implementations, the determination unit determines whether the low-noise triggering condition is met based on the environmental noise parameters and the current scene information, including: if the environmental noise parameters are less than a preset noise threshold, or the scene information is a preset low-noise sensitive scene, then it is determined that the low-noise triggering condition is met.
[0015] In some implementations, the recognition unit, wherein the speech recognition module enables a preset whisper recognition mechanism to recognize and analyze the speech signal to obtain the control command of the air conditioner, including: extracting the cepstral coefficient features, spectral slope features, and low-frequency energy ratio features of the speech signal, and generating a feature vector; inputting the feature vector into a preset neural network model for joint acoustic and speech recognition to obtain the control command of the air conditioner.
[0016] In some embodiments, the speech recognition module includes a main processing unit and a whisper recognition unit; the recognition unit is further configured to: when the low-noise triggering condition is met, control the main processing unit to enter a sleep state and wake up the whisper recognition unit, and the whisper recognition unit performs the operation of recognizing and parsing the speech signal after being woken up; when the low-noise triggering condition is not met, control the whisper recognition unit to enter a sleep state and wake up the main processing unit.
[0017] In some implementations, the control unit is also configured to: after controlling the air conditioner to execute the control command, provide feedback to the user on the command execution status in a non-voice manner; the non-voice manner includes flashing lights and changes in airflow.
[0018] In some implementations, the recognition unit is further configured to: shut down the voice recognition module and keep the low-power monitoring module running when the low-noise triggering condition is continuously met and no new voice signal is detected within a preset time period.
[0019] In conjunction with the above-described device, the present invention further provides an air conditioner, comprising: the control device for the air conditioner described above.
[0020] In conjunction with the above method, the present invention further provides a storage medium comprising a stored program, wherein, when the program is executed, the device on which the storage medium is located controls the air conditioner control method described above to be performed.
[0021] In conjunction with the above method, the present invention further provides a computer program product comprising a computer program that, when processed and executed, implements the steps of the above-described air conditioner control method.
[0022] The present invention provides an air conditioner comprising an ambient sound detection module, a low-power monitoring module, and a voice recognition module. The low-power monitoring module can monitor voice features in a low-power state. The ambient sound detection module acquires environmental noise parameters and, combined with current scene information, determines whether a low-noise trigger condition is met. If met, the low-power monitoring module is activated to monitor the user's voice signal. When the detected voice signal matches preset whispered characteristics, the voice recognition module uses a preset whispered recognition mechanism to recognize and parse the voice signal to obtain air conditioner control commands, thereby controlling the air conditioner to execute these commands. This effectively reduces wake-up noise from air conditioner voice control, improves the recognition sensitivity of soft and whispered commands, enhances the system's environmental adaptability to low-noise scenarios such as sleep and meetings, and reduces power consumption during voice monitoring.
[0023] Other features and advantages of the invention will be set forth in the description which follows, and will be apparent in part from the description, or may be learned by practicing the invention.
[0024] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Attached Figure Description
[0025] Figure 1 This is a flowchart illustrating an embodiment of the air conditioner control method of the present invention; Figure 2 This is a schematic diagram of the structure of an embodiment of the air conditioner control device of the present invention; Figure 3 This is a flowchart illustrating the low-noise voice control mode for air conditioners.
[0026] Referring to the accompanying drawings, the reference numerals in the embodiments of the present invention are as follows: 101-Decision unit; 102-Monitoring unit; 103-Identification unit; 104-Control unit. Detailed Implementation
[0027] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below in conjunction with specific embodiments and corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of this invention, and not all of them. Based on the embodiments of this invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this invention.
[0028] According to an embodiment of the present invention, a control method for an air conditioner is provided. The air conditioner includes an ambient sound detection module, a low-power monitoring module, and a voice recognition module. The low-power monitoring module is capable of monitoring voice features in a low-power state.
[0029] The ambient sound detection module is equipped with ambient sound signal acquisition, processing, and noise parameter calculation functions. It is used to capture ambient sound signals in real time and convert them into quantifiable ambient noise parameters. The low-power monitoring module minimizes energy consumption while maintaining basic monitoring functions. Its core function is to selectively monitor speech signals that meet specific characteristics, avoiding the high energy consumption problem of full-volume speech monitoring. The speech recognition module receives speech signals and converts them into executable control commands. It has built-in recognition logic specifically adapted to quiet interaction scenarios, accurately parsing speech signals with specific characteristics and generating corresponding control commands.
[0030] like Figure 1 The flowchart of an embodiment of the method of the present invention is shown. The air conditioner control method may include steps S110 to S140.
[0031] In step S110, the ambient noise parameters are obtained through the ambient sound detection module, and it is determined whether the low noise triggering condition is met based on the ambient noise parameters and the current scene information.
[0032] Environmental noise parameters are quantitative indicators obtained by analyzing and calculating environmental sound signals, used to characterize the noise intensity of the current environment. Scene information describes the scenario in which the user is currently using the air conditioner, including but not limited to sleep scenarios, meeting scenarios, and daily scenarios. This information can be obtained through various input methods.
[0033] Traditional air conditioning voice control systems use a fixed high-volume wake-up and recognition mode, which cannot dynamically adjust according to environmental conditions. This makes them prone to interference in low-noise environments and consumes a lot of energy. Environmental noise parameters directly reflect the quietness of the current environment, while scene information reflects the user's usage scenario needs. Combining these two factors to determine low-noise trigger conditions ensures that the system only activates the corresponding control logic in scenarios requiring low-noise interaction. This avoids unnecessary mode switching and accurately adapts to user needs in different scenarios, fundamentally solving the problems of poor environmental adaptability and disturbance to others inherent in traditional systems.
[0034] Specifically, the ambient sound detection module continuously collects sound signals from the surrounding environment in real time. It analyzes the collected sound signals using a built-in signal processing algorithm to calculate the corresponding ambient noise parameters. Simultaneously, the system acquires current scene information, which can be obtained in various ways (including but not limited to user-defined settings, automatic system judgment based on time periods, and analysis based on auxiliary environmental parameters). Subsequently, the system compares the acquired ambient noise parameters with preset noise judgment standards and performs a comprehensive evaluation based on the current scene information to determine whether the low-noise trigger conditions for activating the low-noise voice control mode are met.
[0035] In some implementations, step S110, determining whether the low-noise triggering condition is met based on the environmental noise parameters and the current scene information, includes: if the environmental noise parameters are less than a preset noise threshold, or the scene information is a preset low-noise sensitive scene, then it is determined that the low-noise triggering condition is met.
[0036] The preset noise threshold is a noise intensity benchmark value pre-set in the air conditioning system to determine whether the environment is in a low-noise state. This value is determined based on the typical noise level of a low-noise scenario. The preset low-noise sensitive scenarios are a pre-defined set of scenarios that have specific requirements for environmental quietness and where users cannot easily issue high-volume voice commands. Their core characteristics are "the need to maintain a low-noise environment and low-volume interaction constraints." Specifically, they include sleep scenarios, meeting scenarios, library scenarios, hospital ward scenarios, etc., covering all sensitive scenarios that require non-intrusive voice control.
[0037] Specifically, the system first collects the sound signal of the current environment through the ambient sound detection module, calculates the ambient noise parameter through the signal processing algorithm, and then compares the ambient noise parameter with the preset noise threshold to determine whether the ambient noise parameter is less than the preset noise threshold. At the same time, the system obtains the current scene information through the scene recognition module. The scene recognition module performs comprehensive analysis on multimodal input information (including but not limited to time period, ambient light parameters, user terminal APP settings, detection results of the number of people in the environment, user voice content, etc.) to determine whether the current scene is a low-noise sensitive scene such as a sleep scene. The system performs a logical "OR" operation on the above two judgment results. That is, as long as either of the two conditions "ambient noise parameter is less than the preset noise threshold" or "scene information is a low-noise sensitive scene" is met, it is determined that the current low-noise trigger condition is met, and the system is triggered to start the low-noise voice control mode. If neither condition is met, it is determined that the low-noise trigger condition is not met, and the system maintains the normal voice control mode or standby state.
[0038] The preset noise threshold can be adjusted according to the actual application scenario requirements, with a typical setting range of 35-40dB. This range can effectively cover the noise intensity characteristics of low-noise scenarios such as sleep and meetings. When determining a sleep scenario, the scene recognition module can take into account information such as a specific time period at night (e.g., 22:00-07:00), ambient light parameters being lower than the preset light threshold, and the user setting a "sleep mode" through the APP. When determining a meeting scenario, it can take into account information such as the user setting a "meeting mode", the detection of multiple voice signals in the environment with overall low noise levels, and the absence of frequent large-amplitude movements for a long period of time, to ensure the accuracy of scene determination.
[0039] In step S120, if the low-noise triggering condition is met, the low-power monitoring module is activated, and the user's voice signal is monitored for features through the low-power monitoring module.
[0040] When the system determines that the low-noise trigger condition is met, it means that the user is in a scenario requiring quiet interaction (such as sleeping at night or during a meeting). In this case, using the traditional high-power full-volume voice monitoring mode would result in unnecessary energy waste and could lead to false triggers due to an overly broad monitoring range. The low-power monitoring module, however, monitors only voice signals with specific characteristics. This allows for accurate capture of the user's interaction intent in low-noise scenarios while minimizing system energy consumption, achieving a balance between energy saving and accurate monitoring, and avoiding the false triggering interference that can occur with full-volume monitoring.
[0041] Specifically, when the system determines that the low-noise triggering condition is met, it automatically starts the low-power monitoring module. This module operates in a low-power state to avoid the impact of high-energy-consumption operation on the overall energy consumption of the air conditioner. After the low-power monitoring module is started, it focuses on monitoring the characteristics of the surrounding voice signals. It does not perform full-scale voice signal analysis and processing, but only filters voice signals that meet the preset characteristics, filtering out ambient noise (such as the sound of curtains rubbing, the sound of writing on paper and pen), reducing the interference of invalid signals on the system, and ensuring the targeting and accuracy of the monitoring.
[0042] In step S130, when the voice signal is detected to match the preset whisper characteristics, the voice recognition module activates the preset whisper recognition mechanism to recognize and analyze the voice signal, thereby obtaining the control command for the air conditioner.
[0043] In step S140, the air conditioner is controlled to execute the control command.
[0044] The preset whisper features are a pre-defined set of features used to distinguish whisper signals from other sound signals. This feature set is built based on the low-energy, high-airflow characteristics of whispers, enabling accurate identification of the unique attributes of whisper signals. The whisper recognition mechanism is a built-in recognition logic within the speech recognition module specifically designed to analyze whisper signals. Through targeted feature extraction and signal parsing algorithms, it achieves accurate identification and command conversion of whisper signals. Control commands are operational instructions generated by the speech recognition module after recognizing and parsing the whisper signals, which can be executed by the air conditioner, such as adjusting the temperature, controlling the fan speed, and closing the air vents.
[0045] In low-noise environments, users are reluctant to issue loud voice commands and often interact via whisper. Whisper signals possess unique properties of low energy and high airflow, making them difficult for traditional speech recognition mechanisms to accurately identify. The preset whisper features are specifically designed for the properties of whisper signals, effectively distinguishing whispers from other sounds. The whisper recognition mechanism employs dedicated logic adapted to whisper signals. Through targeted algorithm design, it addresses the insufficient sensitivity of traditional recognition mechanisms for whisper signals, ensuring that the user's interactive intentions in low-noise environments can be accurately captured and translated into control commands.
[0046] Specifically, during continuous monitoring, the low-power monitoring module compares the captured voice signals with preset whisper features in real time. When a voice signal is detected to perfectly match the preset whisper features, the voice recognition module is triggered to activate the preset whisper recognition mechanism. After the whisper recognition mechanism is activated, it first extracts targeted features from the voice signal, capturing its key features related to low energy and high airflow. Then, it analyzes and processes the extracted features using a built-in parsing algorithm, converting the voice signal into corresponding text information. Combined with the logic rules related to air conditioning control, the text information is converted into control commands that the air conditioner can execute, completing the conversion process from whisper signal to control command. The control command is then transmitted to the air conditioner's main control unit through the system's internal signal transmission channel. After receiving the control command, the main control unit verifies and parses the command to determine the corresponding air conditioning operation type (such as temperature adjustment, fan adjustment, shutdown, etc.). Subsequently, the main control unit sends operation signals to the corresponding actuators of the air conditioner (such as the compressor, fan, and air outlet adjustment mechanism), controlling the actuators to operate according to the command requirements and completing the user's expected air conditioning control operation.
[0047] In some implementations, in step S130, the speech recognition module uses a preset whisper recognition mechanism to recognize and analyze the speech signal to obtain the control command of the air conditioner, including: extracting the cepstral coefficient features, spectral slope features, and low-frequency energy ratio features of the speech signal, and generating a feature vector; inputting the feature vector into a preset neural network model for joint acoustic and language recognition to obtain the control command of the air conditioner.
[0048] Cepstral coefficient features are characteristic parameters obtained after performing Fourier transform, Mel filtering, logarithmic operations, and discrete cosine transform on the speech signal. They can effectively capture the spectral envelope information of the speech signal and can employ 39 dimensions (including 13-dimensional basic cepstral coefficients and first- and second-order differences). They can accurately characterize the spectral distortion characteristics of whispers caused by airflow, distinguishing the spectral differences between whispers and ordinary speech. Spectral slope features are parameters obtained by calculating the "frequency-spectral amplitude weighted ratio," used to describe the degree of sloping of the speech signal spectrum with frequency changes. They can effectively identify the high airflow characteristics of whispers and are one of the key features for distinguishing whispers from environmental noise. Low-frequency energy ratio features refer to the ratio of energy in the 0-500Hz frequency band to the energy of the entire frequency band in the speech signal. They can highlight the low-energy characteristics of whispers and avoid misinterpreting low-frequency noise in the environment as valid speech commands. The feature vector is a high-dimensional data vector formed by concatenating the extracted cepstral coefficient features, spectral slope features, and low-frequency energy ratio features after normalization according to preset rules. It integrates the core feature information of the whisper signal. The pre-trained neural network model is an artificial intelligence model specifically adapted for whisper recognition. It has the ability to analyze and process feature vectors and convert them into control commands. It can be a lightweight hybrid neural network model combining CNN and BiLSTM, a DNN model, or a Transformer model.
[0049] Whisper signals differ significantly from ordinary speech signals in terms of spectral distribution and energy intensity. Traditional speech recognition mechanisms have not adapted to these characteristics, resulting in insufficient sensitivity and a high false positive rate in whisper signal recognition. By extracting three core features—cep spectral coefficients, spectral slope, and low-frequency energy ratio—the unique attributes of whisper signals can be comprehensively captured from different dimensions. The resulting feature vector can effectively distinguish whispers from environmental noise. Furthermore, the joint acoustic and speech recognition approach combines signal-level features with semantic-level logic, resolving issues such as incomplete pronunciation and discontinuous signals in whispered speech. This improves the accuracy and completeness of command parsing while reducing recognition latency, ensuring rapid response to user control needs in low-noise environments.
[0050] Specifically, after the speech recognition module activates the whisper recognition mechanism, it first receives a speech signal that conforms to preset whisper characteristics transmitted by the low-power monitoring module. This signal is then preprocessed (including noise reduction, pre-emphasis, framing, and windowing) to eliminate environmental noise and signal interference, improving signal quality. Three types of core features are extracted from the preprocessed speech signal: 39-dimensional cepstral coefficients are calculated using a dedicated algorithm to capture spectral distortion information; spectral slope is calculated using the "frequency-spectral amplitude weighted ratio" to identify high airflow characteristics (the spectral slope of whispers is typically ≥-10dB / Hz); and the ratio of energy in the 0-500Hz band to the total energy in the entire frequency band is obtained using an energy calculation algorithm, i.e., the low-frequency energy ratio feature (the low-frequency energy ratio of whispers is typically ≥0.3). The extracted three types of features are then normalized to eliminate dimensional and numerical differences between different feature dimensions, ensuring that each feature is consistent with its intended function. The influence of features on the recognition results is balanced; the three types of normalized features are concatenated in a preset order to form a feature vector of a unified dimension, which serves as the input data for the preset neural network model; after receiving the feature vector, the preset neural network model performs feature depth extraction and analysis through the built-in network layers: if a hybrid model combining lightweight CNN and BiLSTM is used, the CNN layer will capture the local spectral peaks of the whisper signal (such as the spectral bulge of airflow sound), and the BiLSTM layer will capture the temporal relationship of the command (such as the continuous speech logic of "adjust-lower-temperature-degree"), avoiding command breaks caused by intermittent whispers; the model combines the extracted acoustic features with the language model of the air conditioning control scenario through the joint recognition logic of acoustics and language, and achieves alignment between the two through an attention mechanism to directly parse out the complete control command; the model outputs the parsed control command, completing the entire recognition and parsing process.
[0051] The training process of the preset neural network model needs to be based on the whispered command data of air conditioners in low-noise scenarios. The CTC loss function is used to solve the problem of mismatch between whispered speech and text length, ensuring that the model can achieve a recognition accuracy of ≥90% in low-noise environments below 35dB, and the command recognition delay is ≤300ms.
[0052] In some implementations, the speech recognition module includes a main processing unit and a whisper recognition unit. The main processing unit is the functional unit within the speech recognition module responsible for recognizing ordinary speech commands. It employs a deep neural network (DNN) or Transformer acoustic model, adapting to the recognition and parsing of normal-volume speech signals (such as everyday conversational volume). It possesses full-volume speech signal processing, feature extraction, and command generation capabilities, suitable for routine voice control needs in non-low-noise scenarios. The whisper recognition unit is a functional unit within the speech recognition module specifically adapted for whisper signal recognition. Designed based on the characteristics of whispers, it achieves accurate recognition and parsing of whisper signals through targeted feature extraction algorithms and a lightweight recognition model. Its core application is in disturbance-free control in low-noise scenarios.
[0053] The method further includes: when the low-noise triggering condition is met, controlling the main processing unit to enter a sleep state and waking up the whisper recognition unit, the whisper recognition unit performing the operation of recognizing and parsing the speech signal after being woken up; when the low-noise triggering condition is not met, controlling the whisper recognition unit to enter a sleep state and waking up the main processing unit.
[0054] Traditional speech recognition modules typically operate with a single recognition unit continuously. Maintaining high power consumption is necessary to ensure good recognition performance in normal scenarios, while prioritizing low power consumption sacrifices accuracy in low-noise environments. Dividing the speech recognition module into a main processing unit and a whispering recognition unit, and using a dynamic switching logic adapted to different scenarios, achieves the goals of "uninterrupted accuracy in low-noise scenarios, efficient recognition in normal scenarios, and low power consumption across all scenarios." In low-noise scenarios, the main processing unit is put into sleep mode, and only the whispering recognition unit is activated. This avoids energy waste caused by the main processing unit's high power consumption and improves whispering recognition accuracy through a dedicated unit. In normal scenarios, switching to the main processing unit ensures efficient recognition of normal volume commands.
[0055] Specifically, the system receives environmental noise parameters from the ambient sound detection module and scene information from the scene recognition module in real time, continuously determining whether the low-noise triggering conditions are met. When the system determines that the low-noise triggering conditions are met, it automatically sends a mode switching command to the speech recognition module, controlling the main processing unit to enter a sleep state and simultaneously waking up the whisper recognition unit. After the main processing unit enters a sleep state, it disables the full-volume speech signal parsing function, retaining only basic signal monitoring capabilities to maintain standby with minimal power consumption and avoid unnecessary energy consumption. After being woken up, the whisper recognition unit immediately enters the working state, receives the speech signal transmitted by the low-power monitoring module, and prepares to perform whisper recognition and parsing operations to ensure that user whisper commands can be captured in a timely manner. When the system determines that the low-noise triggering conditions are not met (such as ambient noise exceeding a preset threshold or the scene being a daily leisure scene), it sends a reverse switching command to the speech recognition module, controlling the whisper recognition unit to enter a sleep state and simultaneously waking up the main processing unit. After the whisper recognition unit enters a sleep state, it stops whisper feature parsing related functions to reduce the system's operating load. After the main processing unit is woken up, it starts the normal speech recognition mode, continuously monitoring and recognizing normal-volume speech commands to ensure interactive response speed and recognition accuracy in daily scenarios. During unit switching, the system ensures a smooth transition through an internal signal synchronization mechanism, avoiding situations where commands are missed or delayed. The switching response time can be controlled within a short time to ensure that the user is unaware of the process.
[0056] The control signals for unit sleep and wake-up are issued by the air conditioner main control unit. The main control unit generates a switching command based on the comprehensive judgment result of environmental noise parameters and scene information, and transmits it to the voice recognition module through the internal bus to achieve precise control of the unit's operating status. After the whisper recognition unit is woken up, its operating power consumption is only a portion of that of a traditional single recognition unit, which can significantly reduce the overall energy consumption of the air conditioner.
[0057] In some embodiments, the method further includes: after controlling the air conditioner to execute the control command, providing feedback to the user on the command execution status in a non-voice manner; the non-voice manner includes flashing lights and changes in airflow.
[0058] Non-voice feedback refers to feedback methods that convey information to users silently, such as through visual, tactile, or airflow perception, rather than through voice broadcasts or prompts. The core characteristic is the absence of noise, avoiding interference in low-noise environments. Flashing lights convey the status of commands by controlling LED indicator lights, soft light panels, or other light-emitting components on the air conditioner, using preset flashing times, frequencies, or brightness changes. The light intensity is soft and does not cause visual interference to the user. Changes in airflow convey the status of commands by controlling the airflow speed, direction, or duration at the air vents, using brief, slight changes in airflow. The airflow intensity is gentle and does not affect the comfort of the current environment. Command execution status refers to the air conditioner's response to control commands, including successful command execution, command failure, and command in progress.
[0059] The core requirement for low-noise scenarios is maintaining a quiet and undisturbed environment. Traditional air conditioning voice control systems typically use voice broadcasting as feedback after executing commands. This method generates additional noise, disrupting the tranquility of low-noise scenarios and even disturbing others' rest or meetings. Non-voice methods, on the other hand, convey information through silent feedback. This allows users to clearly understand the command's execution status without generating any noise interference, perfectly meeting the undisturbed requirements of low-noise environments. Furthermore, flashing lights and changes in airflow are intuitive and easily perceptible forms of feedback, allowing users to obtain information without focusing their attention, thus improving ease of use.
[0060] Specifically, after receiving the control command generated by the voice recognition module, the air conditioning main control unit sends an operation signal to the corresponding execution component, controlling the execution component to operate according to the command requirements, performing operations such as temperature adjustment, fan adjustment, and shutdown. After the execution component completes the operation, it returns an execution result signal (such as a success signal or a failure signal) to the main control unit. The main control unit confirms the command execution status based on this signal. The main control unit transmits the command execution status signal to the feedback control module. The feedback control module selects the corresponding non-voice feedback method according to preset feedback rules: if the current scenario is a sleep scenario, the feedback control module controls the air conditioner's LED soft light indicator to flash once at a preset frequency (such as completing a "on-off" cycle within 1 second) to avoid strong light or frequent flashing affecting the user's sleep; if the current scenario is a meeting scenario, the feedback control module controls the air conditioner's air outlet to output a brief, gentle airflow (such as a breeze lasting 0.5 seconds), informing the user that the command has been executed successfully through the tactile sensation of the airflow. After the feedback is completed, the feedback control module returns a feedback end signal to the main control unit, and the system returns to the low-noise voice control standby state, waiting for the next command input.
[0061] Among them, the feedback control module is the core hardware component that realizes non-voice feedback. It can realize corresponding feedback actions by controlling the air conditioner's light drive circuit, fan control circuit, etc. The preset feedback rules can be automatically matched based on the scene type, or they can be customized by the user through the mobile APP or air conditioner panel to meet the usage habits of different users.
[0062] In some embodiments, the method further includes: when the low-noise triggering condition is continuously met and no new voice signal is detected within a preset time period, turning off the voice recognition module and keeping the low-power monitoring module in operation.
[0063] In low-noise scenarios, user voice interactions are typically intermittent. Keeping the voice recognition module running continuously without new interactions would result in unnecessary energy waste, violating the energy-saving requirements of low-noise environments. Temporarily shutting down the voice recognition module can promptly reduce system energy consumption when there are no further user actions, achieving an organic balance between accurate monitoring and energy-saving operation in low-noise scenarios.
[0064] Specifically, after the low-noise trigger condition is met, the system activates the low-noise voice control mode. At this time, the voice recognition module is in working condition, and the low-power monitoring module simultaneously monitors the voice signal. The system's built-in timing module starts timing, and in real time, it counts the non-interaction duration since the last valid voice signal was detected (or after the low-noise mode was activated). During the timing process, the low-power monitoring module continuously monitors the surrounding voice signals. If a new voice signal is detected, the timing module immediately resets and restarts counting the non-interaction duration. The voice recognition module remains in working condition to meet new recognition needs. If the non-interaction duration counted by the timing module reaches the preset duration, and no new voice signal is detected during this period, the system determines that the user has no intention of further interaction and sends a shutdown command to the voice recognition module, controlling the voice recognition module to stop the recognition and parsing function, retaining only basic state monitoring. After the voice recognition module is shut down, the low-power monitoring module continues to maintain the basic voice feature monitoring function, operating in a low-power state, and continuously capturing the surrounding voice signals. When the low-power monitoring module detects a new voice signal that matches the preset characteristics again, it immediately sends a wake-up signal to the system. The system quickly activates the voice recognition module, restoring the full functionality of the low-noise voice control mode, ensuring that new user commands can be recognized and parsed in a timely manner.
[0065] The preset duration can be flexibly adjusted according to the actual application scenario, with a typical setting range of 1-5 minutes. This avoids frequent start-stop cycles caused by closing the module after a short period of no interaction, and effectively controls energy consumption during long periods of no interaction. When maintaining basic voice feature monitoring, the low-power monitoring module consumes only a very small percentage of the power of the voice recognition module, which can significantly reduce the overall energy consumption of the air conditioner.
[0066] Figure 3 The flowchart for the low-noise voice control mode of the air conditioner is shown, specifically including steps 1 to 10.
[0067] Step 1: After the air conditioner is powered on, it automatically completes the hardware initialization and software parameter loading of each functional module.
[0068] Step 2: The scene recognition module analyzes information such as time period, ambient light, and user settings to determine whether the current scene belongs to a low-noise scene such as sleep, meeting, or silent mode.
[0069] Step 3: If the scene is determined to be low noise, the system starts the low power monitoring module. This module only maintains the basic sound energy detection function and does not perform full speech recognition, and operates with low power consumption.
[0070] Step 4: The ambient sound detection module collects ambient sound signals in real time, calculates the noise intensity and compares it with a preset threshold to determine whether the current environment is in a low-noise state.
[0071] Step 5: If the ambient noise is below the threshold, the system switches to the silent trigger state, and the low-power monitoring module begins to specifically monitor low-energy voice signals that match the characteristics of whispers.
[0072] Step 6: The low-power monitoring module captures a voice signal that meets the characteristics of low energy and high airflow, and determines it to be a valid whisper command signal.
[0073] Step 7: The system sends a wake-up signal to the main recognition chip of the voice recognition module to switch it from sleep mode to working mode.
[0074] Step 8: The main recognition chip starts the whisper recognition algorithm, and sequentially completes the feature extraction of the voice signal, text conversion and parsing of the air conditioning control commands.
[0075] Step 9: The air conditioning main control unit receives the parsed control command, drives the corresponding components to perform the operation, and provides feedback on the command execution status to the user in a non-voice manner.
[0076] Step 10: If no new voice interaction signal is detected within the preset time period, the system shuts down the main recognition chip and only retains the low-power monitoring module to maintain basic monitoring, and resumes low-power operation.
[0077] This solution achieves precise triggering of low-noise voice control mode through joint judgment of environmental noise parameters and scene information, effectively solving the problem of loud wake-up noise and easy disturbance to others in low-noise scenarios of traditional air conditioning voice control systems; with the targeted monitoring of low-power monitoring modules, the system significantly reduces operating energy consumption while ensuring the capture of effective voice signals, solving the drawback of traditional systems constantly resident in high-power monitoring state; through preset whisper features and a dedicated whisper recognition mechanism, the accuracy of whisper signal recognition is significantly improved, making up for the lack of sensitivity of traditional voice recognition modules to soft speech or whispers.
[0078] For example, when a user uses the air conditioner in a quiet bedroom at night, the ambient noise level detected by the ambient sound detection module is below the preset noise threshold. Simultaneously, the system determines the current time (e.g., 2 AM) to be a sleep environment. Combining these two pieces of information, the system determines that the low-noise trigger condition is met. At this point, the system automatically activates the low-power monitoring module, which operates in a low-power state, continuously monitoring surrounding voice signals. When the user, feeling the temperature is too high, whispers "lower the temperature," the low-power monitoring module captures this voice signal and detects that it matches preset whisper characteristics. This triggers the voice recognition module to activate the whisper recognition mechanism. The voice recognition module extracts and analyzes the features of the "lower the temperature" whisper signal, generating a control command to "lower the air conditioner temperature by 1°C." The air conditioner's main control unit then receives this command and controls the compressor and other components to lower the air conditioner temperature by 1°C, fulfilling the user's control request. Throughout this process, the user does not need to issue a loud command, and the system does not need to continuously monitor with high power consumption, avoiding disturbance to others while achieving precise control.
[0079] For example, if a user is using the air conditioner in a meeting room while a meeting is in progress, the system, based on the "meeting mode" set by the user via a mobile app and combined with low ambient noise parameters collected by the ambient sound detection module, determines that the low-noise trigger condition is met and activates the low-power monitoring module. When the user needs to turn off the air conditioner vents and whispers "turn off the airflow," the low-power monitoring module detects the voice signal that matches the preset whisper characteristics. The voice recognition module then uses a whisper recognition mechanism to parse and generate a control command to "turn off the air conditioner vent speed." The air conditioner executes this command, turning off the vent speed. The entire interaction process is quiet and undisturbed, without affecting the normal progress of the meeting.
[0080] The technical solution adopted in this embodiment includes an air conditioner comprising an ambient sound detection module, a low-power monitoring module, and a voice recognition module. The low-power monitoring module can monitor voice features in a low-power state. It acquires ambient noise parameters through the ambient sound detection module and, combined with current scene information, determines whether the low-noise trigger condition is met. If met, the low-power monitoring module is activated to monitor the user's voice signal. When the detected voice signal matches preset whispered characteristics, the voice recognition module uses a preset whispered recognition mechanism to recognize and parse the voice signal to obtain the air conditioner control command, thereby controlling the air conditioner to execute the command. This effectively reduces wake-up noise from air conditioner voice control, improves the recognition sensitivity of soft and whispered commands, enhances the system's environmental adaptability to low-noise scenarios such as sleep and meetings, and reduces operating power consumption in voice monitoring mode.
[0081] According to an embodiment of the present invention, a control device for an air conditioner corresponding to the control method for an air conditioner is also provided. The air conditioner includes an ambient sound detection module, a low-power monitoring module, and a voice recognition module; the low-power monitoring module is capable of monitoring voice features in a low-power state.
[0082] The ambient sound detection module is equipped with ambient sound signal acquisition, processing, and noise parameter calculation functions. It is used to capture ambient sound signals in real time and convert them into quantifiable ambient noise parameters. The low-power monitoring module minimizes energy consumption while maintaining basic monitoring functions. Its core function is to selectively monitor speech signals that meet specific characteristics, avoiding the high energy consumption problem of full-volume speech monitoring. The speech recognition module receives speech signals and converts them into executable control commands. It has built-in recognition logic specifically adapted to quiet interaction scenarios, accurately parsing speech signals with specific characteristics and generating corresponding control commands.
[0083] See Figure 2 The schematic diagram shown is a structural diagram of an embodiment of the device of the present invention. The control device of the air conditioner may include: a determination unit 101, a listening unit 102, an identification unit 103, and a control unit 104.
[0084] The determination unit 101 is configured to acquire environmental noise parameters through the environmental sound detection module, and determine whether the low-noise triggering condition is met based on the environmental noise parameters and the current scene information. For the specific functions and processing of this unit, please refer to step S110.
[0085] Environmental noise parameters are quantitative indicators obtained by analyzing and calculating environmental sound signals, used to characterize the noise intensity of the current environment. Scene information describes the scenario in which the user is currently using the air conditioner, including but not limited to sleep scenarios, meeting scenarios, and daily scenarios. This information can be obtained through various input methods.
[0086] Traditional air conditioning voice control systems use a fixed high-volume wake-up and recognition mode, which cannot dynamically adjust according to environmental conditions. This makes them prone to interference in low-noise environments and consumes a lot of energy. Environmental noise parameters directly reflect the quietness of the current environment, while scene information reflects the user's usage scenario needs. Combining these two factors to determine low-noise trigger conditions ensures that the system only activates the corresponding control logic in scenarios requiring low-noise interaction. This avoids unnecessary mode switching and accurately adapts to user needs in different scenarios, fundamentally solving the problems of poor environmental adaptability and disturbance to others inherent in traditional systems.
[0087] Specifically, the ambient sound detection module continuously collects sound signals from the surrounding environment in real time. It analyzes the collected sound signals using a built-in signal processing algorithm to calculate the corresponding ambient noise parameters. Simultaneously, the system acquires current scene information, which can be obtained in various ways (including but not limited to user-defined settings, automatic system judgment based on time periods, and analysis based on auxiliary environmental parameters). Subsequently, the system compares the acquired ambient noise parameters with preset noise judgment standards and performs a comprehensive evaluation based on the current scene information to determine whether the low-noise trigger conditions for activating the low-noise voice control mode are met.
[0088] In some implementations, the determination unit 101 determines whether the low-noise triggering condition is met based on the environmental noise parameters and the current scene information, including: if the environmental noise parameters are less than a preset noise threshold, or the scene information is a preset low-noise sensitive scene, then it is determined that the low-noise triggering condition is met.
[0089] The preset noise threshold is a noise intensity benchmark value pre-set in the air conditioning system to determine whether the environment is in a low-noise state. This value is determined based on the typical noise level of a low-noise scenario. The preset low-noise sensitive scenarios are a pre-defined set of scenarios that have specific requirements for environmental quietness and where users cannot easily issue high-volume voice commands. Their core characteristics are "the need to maintain a low-noise environment and low-volume interaction constraints." Specifically, they include sleep scenarios, meeting scenarios, library scenarios, hospital ward scenarios, etc., covering all sensitive scenarios that require non-intrusive voice control.
[0090] Specifically, the system first collects the sound signal of the current environment through the ambient sound detection module, calculates the ambient noise parameter through the signal processing algorithm, and then compares the ambient noise parameter with the preset noise threshold to determine whether the ambient noise parameter is less than the preset noise threshold. At the same time, the system obtains the current scene information through the scene recognition module. The scene recognition module performs comprehensive analysis on multimodal input information (including but not limited to time period, ambient light parameters, user terminal APP settings, detection results of the number of people in the environment, user voice content, etc.) to determine whether the current scene is a low-noise sensitive scene such as a sleep scene. The system performs a logical "OR" operation on the above two judgment results. That is, as long as either of the two conditions "ambient noise parameter is less than the preset noise threshold" or "scene information is a low-noise sensitive scene" is met, it is determined that the current low-noise trigger condition is met, and the system is triggered to start the low-noise voice control mode. If neither condition is met, it is determined that the low-noise trigger condition is not met, and the system maintains the normal voice control mode or standby state.
[0091] The preset noise threshold can be adjusted according to the actual application scenario requirements, with a typical setting range of 35-40dB. This range can effectively cover the noise intensity characteristics of low-noise scenarios such as sleep and meetings. When determining a sleep scenario, the scene recognition module can take into account information such as a specific time period at night (e.g., 22:00-07:00), ambient light parameters being lower than the preset light threshold, and the user setting a "sleep mode" through the APP. When determining a meeting scenario, it can take into account information such as the user setting a "meeting mode", the detection of multiple voice signals in the environment with overall low noise levels, and the absence of frequent large-amplitude movements for a long period of time, to ensure the accuracy of scene determination.
[0092] The monitoring unit 102 is configured to activate the low-power monitoring module if the low-noise trigger condition is met, and to perform feature monitoring on the user's voice signal through the low-power monitoring module. For the specific functions and processing of this unit, please refer to step S120.
[0093] When the system determines that the low-noise trigger condition is met, it means that the user is in a scenario requiring quiet interaction (such as sleeping at night or during a meeting). In this case, using the traditional high-power full-volume voice monitoring mode would result in unnecessary energy waste and could lead to false triggers due to an overly broad monitoring range. The low-power monitoring module, however, monitors only voice signals with specific characteristics. This allows for accurate capture of the user's interaction intent in low-noise scenarios while minimizing system energy consumption, achieving a balance between energy saving and accurate monitoring, and avoiding the false triggering interference that can occur with full-volume monitoring.
[0094] Specifically, when the system determines that the low-noise triggering condition is met, it automatically starts the low-power monitoring module. This module operates in a low-power state to avoid the impact of high-energy-consumption operation on the overall energy consumption of the air conditioner. After the low-power monitoring module is started, it focuses on monitoring the characteristics of the surrounding voice signals. It does not perform full-scale voice signal analysis and processing, but only filters voice signals that meet the preset characteristics, filtering out ambient noise (such as the sound of curtains rubbing, the sound of writing on paper and pen), reducing the interference of invalid signals on the system, and ensuring the targeting and accuracy of the monitoring.
[0095] The recognition unit 103 is configured to, when the voice signal is detected to match a preset whisper feature, activate a preset whisper recognition mechanism to recognize and analyze the voice signal, thereby obtaining the control command for the air conditioner. For the specific functions and processing of this unit, please refer to step S130.
[0096] The control unit 104 is configured to control the air conditioner to execute the control commands. The specific functions and processing of this unit are described in step S140.
[0097] The preset whisper features are a pre-defined set of features used to distinguish whisper signals from other sound signals. This feature set is built based on the low-energy, high-airflow characteristics of whispers, enabling accurate identification of the unique attributes of whisper signals. The whisper recognition mechanism is a built-in recognition logic within the speech recognition module specifically designed to analyze whisper signals. Through targeted feature extraction and signal parsing algorithms, it achieves accurate identification and command conversion of whisper signals. Control commands are operational instructions generated by the speech recognition module after recognizing and parsing the whisper signals, which can be executed by the air conditioner, such as adjusting the temperature, controlling the fan speed, and closing the air vents.
[0098] In low-noise environments, users are reluctant to issue loud voice commands and often interact via whisper. Whisper signals possess unique properties of low energy and high airflow, making them difficult for traditional speech recognition mechanisms to accurately identify. The preset whisper features are specifically designed for the properties of whisper signals, effectively distinguishing whispers from other sounds. The whisper recognition mechanism employs dedicated logic adapted to whisper signals. Through targeted algorithm design, it addresses the insufficient sensitivity of traditional recognition mechanisms for whisper signals, ensuring that the user's interactive intentions in low-noise environments can be accurately captured and translated into control commands.
[0099] Specifically, during continuous monitoring, the low-power monitoring module compares the captured voice signals with preset whisper features in real time. When a voice signal is detected to perfectly match the preset whisper features, the voice recognition module is triggered to activate the preset whisper recognition mechanism. After the whisper recognition mechanism is activated, it first extracts targeted features from the voice signal, capturing its key features related to low energy and high airflow. Then, it analyzes and processes the extracted features using a built-in parsing algorithm, converting the voice signal into corresponding text information. Combined with the logic rules related to air conditioning control, the text information is converted into control commands that the air conditioner can execute, completing the conversion process from whisper signal to control command. The control command is then transmitted to the air conditioner's main control unit through the system's internal signal transmission channel. After receiving the control command, the main control unit verifies and parses the command to determine the corresponding air conditioning operation type (such as temperature adjustment, fan adjustment, shutdown, etc.). Subsequently, the main control unit sends operation signals to the corresponding actuators of the air conditioner (such as the compressor, fan, and air outlet adjustment mechanism), controlling the actuators to operate according to the command requirements and completing the user's expected air conditioning control operation.
[0100] In some embodiments, the recognition unit 103, the speech recognition module, enables a preset whisper recognition mechanism to recognize and analyze the speech signal to obtain the control command of the air conditioner, including: extracting the cepstral coefficient features, spectral slope features, and low-frequency energy ratio features of the speech signal, and generating a feature vector; inputting the feature vector into a preset neural network model for joint acoustic and language recognition to obtain the control command of the air conditioner.
[0101] Cepstral coefficient features are characteristic parameters obtained after performing Fourier transform, Mel filtering, logarithmic operations, and discrete cosine transform on the speech signal. They can effectively capture the spectral envelope information of the speech signal and can employ 39 dimensions (including 13-dimensional basic cepstral coefficients and first- and second-order differences). They can accurately characterize the spectral distortion characteristics of whispers caused by airflow, distinguishing the spectral differences between whispers and ordinary speech. Spectral slope features are parameters obtained by calculating the "frequency-spectral amplitude weighted ratio," used to describe the degree of sloping of the speech signal spectrum with frequency changes. They can effectively identify the high airflow characteristics of whispers and are one of the key features for distinguishing whispers from environmental noise. Low-frequency energy ratio features refer to the ratio of energy in the 0-500Hz frequency band to the energy of the entire frequency band in the speech signal. They can highlight the low-energy characteristics of whispers and avoid misinterpreting low-frequency noise in the environment as valid speech commands. The feature vector is a high-dimensional data vector formed by concatenating the extracted cepstral coefficient features, spectral slope features, and low-frequency energy ratio features after normalization according to preset rules. It integrates the core feature information of the whisper signal. The pre-trained neural network model is an artificial intelligence model specifically adapted for whisper recognition. It has the ability to analyze and process feature vectors and convert them into control commands. It can be a lightweight hybrid neural network model combining CNN and BiLSTM, a DNN model, or a Transformer model.
[0102] Whisper signals differ significantly from ordinary speech signals in terms of spectral distribution and energy intensity. Traditional speech recognition mechanisms have not adapted to these characteristics, resulting in insufficient sensitivity and a high false positive rate in whisper signal recognition. By extracting three core features—cep spectral coefficients, spectral slope, and low-frequency energy ratio—the unique attributes of whisper signals can be comprehensively captured from different dimensions. The resulting feature vector can effectively distinguish whispers from environmental noise. Furthermore, the joint acoustic and speech recognition approach combines signal-level features with semantic-level logic, resolving issues such as incomplete pronunciation and discontinuous signals in whispered speech. This improves the accuracy and completeness of command parsing while reducing recognition latency, ensuring rapid response to user control needs in low-noise environments.
[0103] Specifically, after the speech recognition module activates the whisper recognition mechanism, it first receives a speech signal that conforms to preset whisper characteristics transmitted by the low-power monitoring module. This signal is then preprocessed (including noise reduction, pre-emphasis, framing, and windowing) to eliminate environmental noise and signal interference, improving signal quality. Three types of core features are extracted from the preprocessed speech signal: 39-dimensional cepstral coefficients are calculated using a dedicated algorithm to capture spectral distortion information; spectral slope is calculated using the "frequency-spectral amplitude weighted ratio" to identify high airflow characteristics (the spectral slope of whispers is typically ≥-10dB / Hz); and the ratio of energy in the 0-500Hz band to the total energy in the entire frequency band is obtained using an energy calculation algorithm, i.e., the low-frequency energy ratio feature (the low-frequency energy ratio of whispers is typically ≥0.3). The extracted three types of features are then normalized to eliminate dimensional and numerical differences between different feature dimensions, ensuring that each feature is consistent with its intended function. The influence of features on the recognition results is balanced; the three types of normalized features are concatenated in a preset order to form a feature vector of a unified dimension, which serves as the input data for the preset neural network model; after receiving the feature vector, the preset neural network model performs feature depth extraction and analysis through the built-in network layers: if a hybrid model combining lightweight CNN and BiLSTM is used, the CNN layer will capture the local spectral peaks of the whisper signal (such as the spectral bulge of airflow sound), and the BiLSTM layer will capture the temporal relationship of the command (such as the continuous speech logic of "adjust-lower-temperature-degree"), avoiding command breaks caused by intermittent whispers; the model combines the extracted acoustic features with the language model of the air conditioning control scenario through the joint recognition logic of acoustics and language, and achieves alignment between the two through an attention mechanism to directly parse out the complete control command; the model outputs the parsed control command, completing the entire recognition and parsing process.
[0104] The training process of the preset neural network model needs to be based on the whispered command data of air conditioners in low-noise scenarios. The CTC loss function is used to solve the problem of mismatch between whispered speech and text length, ensuring that the model can achieve a recognition accuracy of ≥90% in low-noise environments below 35dB, and the command recognition delay is ≤300ms.
[0105] In some implementations, the speech recognition module includes a main processing unit and a whisper recognition unit. The main processing unit is the functional unit within the speech recognition module responsible for recognizing ordinary speech commands. It employs a deep neural network (DNN) or Transformer acoustic model, adapting to the recognition and parsing of normal-volume speech signals (such as everyday conversational volume). It possesses full-volume speech signal processing, feature extraction, and command generation capabilities, suitable for routine voice control needs in non-low-noise scenarios. The whisper recognition unit is a functional unit within the speech recognition module specifically adapted for whisper signal recognition. Designed based on the characteristics of whispers, it achieves accurate recognition and parsing of whisper signals through targeted feature extraction algorithms and a lightweight recognition model. Its core application is in disturbance-free control in low-noise scenarios.
[0106] The recognition unit 103 is further configured to: when the low-noise triggering condition is met, control the main processing unit to enter a sleep state and wake up the whisper recognition unit, and the whisper recognition unit performs the operation of recognizing and parsing the speech signal after being woken up; when the low-noise triggering condition is not met, control the whisper recognition unit to enter a sleep state and wake up the main processing unit.
[0107] Traditional speech recognition modules typically operate with a single recognition unit continuously. Maintaining high power consumption is necessary to ensure good recognition performance in normal scenarios, while prioritizing low power consumption sacrifices accuracy in low-noise environments. Dividing the speech recognition module into a main processing unit and a whispering recognition unit, and using a dynamic switching logic adapted to different scenarios, achieves the goals of "uninterrupted accuracy in low-noise scenarios, efficient recognition in normal scenarios, and low power consumption across all scenarios." In low-noise scenarios, the main processing unit is put into sleep mode, and only the whispering recognition unit is activated. This avoids energy waste caused by the main processing unit's high power consumption and improves whispering recognition accuracy through a dedicated unit. In normal scenarios, switching to the main processing unit ensures efficient recognition of normal volume commands.
[0108] Specifically, the system receives environmental noise parameters from the ambient sound detection module and scene information from the scene recognition module in real time, continuously determining whether the low-noise triggering conditions are met. When the system determines that the low-noise triggering conditions are met, it automatically sends a mode switching command to the speech recognition module, controlling the main processing unit to enter a sleep state and simultaneously waking up the whisper recognition unit. After the main processing unit enters a sleep state, it disables the full-volume speech signal parsing function, retaining only basic signal monitoring capabilities to maintain standby with minimal power consumption and avoid unnecessary energy consumption. After being woken up, the whisper recognition unit immediately enters the working state, receives the speech signal transmitted by the low-power monitoring module, and prepares to perform whisper recognition and parsing operations to ensure that user whisper commands can be captured in a timely manner. When the system determines that the low-noise triggering conditions are not met (such as ambient noise exceeding a preset threshold or the scene being a daily leisure scene), it sends a reverse switching command to the speech recognition module, controlling the whisper recognition unit to enter a sleep state and simultaneously waking up the main processing unit. After the whisper recognition unit enters a sleep state, it stops whisper feature parsing related functions to reduce the system's operating load. After the main processing unit is woken up, it starts the normal speech recognition mode, continuously monitoring and recognizing normal-volume speech commands to ensure interactive response speed and recognition accuracy in daily scenarios. During unit switching, the system ensures a smooth transition through an internal signal synchronization mechanism, avoiding situations where commands are missed or delayed. The switching response time can be controlled within a short time to ensure that the user is unaware of the process.
[0109] The control signals for unit sleep and wake-up are issued by the air conditioner main control unit. The main control unit generates a switching command based on the comprehensive judgment result of environmental noise parameters and scene information, and transmits it to the voice recognition module through the internal bus to achieve precise control of the unit's operating status. After the whisper recognition unit is woken up, its operating power consumption is only a portion of that of a traditional single recognition unit, which can significantly reduce the overall energy consumption of the air conditioner.
[0110] In some embodiments, the control unit 104 is further configured to: after controlling the air conditioner to execute the control command, provide feedback to the user on the command execution status in a non-voice manner; the non-voice manner includes flashing lights and changes in airflow.
[0111] Non-voice feedback refers to feedback methods that convey information to users silently, such as through visual, tactile, or airflow perception, rather than through voice broadcasts or prompts. The core characteristic is the absence of noise, avoiding interference in low-noise environments. Flashing lights convey the status of commands by controlling LED indicator lights, soft light panels, or other light-emitting components on the air conditioner, using preset flashing times, frequencies, or brightness changes. The light intensity is soft and does not cause visual interference to the user. Changes in airflow convey the status of commands by controlling the airflow speed, direction, or duration at the air vents, using brief, slight changes in airflow. The airflow intensity is gentle and does not affect the comfort of the current environment. Command execution status refers to the air conditioner's response to control commands, including successful command execution, command failure, and command in progress.
[0112] The core requirement for low-noise scenarios is maintaining a quiet and undisturbed environment. Traditional air conditioning voice control systems typically use voice broadcasting as feedback after executing commands. This method generates additional noise, disrupting the tranquility of low-noise scenarios and even disturbing others' rest or meetings. Non-voice methods, on the other hand, convey information through silent feedback. This allows users to clearly understand the command's execution status without generating any noise interference, perfectly meeting the undisturbed requirements of low-noise environments. Furthermore, flashing lights and changes in airflow are intuitive and easily perceptible forms of feedback, allowing users to obtain information without focusing their attention, thus improving ease of use.
[0113] Specifically, after receiving the control command generated by the voice recognition module, the air conditioning main control unit sends an operation signal to the corresponding execution component, controlling the execution component to operate according to the command requirements, performing operations such as temperature adjustment, fan adjustment, and shutdown. After the execution component completes the operation, it returns an execution result signal (such as a success signal or a failure signal) to the main control unit. The main control unit confirms the command execution status based on this signal. The main control unit transmits the command execution status signal to the feedback control module. The feedback control module selects the corresponding non-voice feedback method according to preset feedback rules: if the current scenario is a sleep scenario, the feedback control module controls the air conditioner's LED soft light indicator to flash once at a preset frequency (such as completing a "on-off" cycle within 1 second) to avoid strong light or frequent flashing affecting the user's sleep; if the current scenario is a meeting scenario, the feedback control module controls the air conditioner's air outlet to output a brief, gentle airflow (such as a breeze lasting 0.5 seconds), informing the user that the command has been executed successfully through the tactile sensation of the airflow. After the feedback is completed, the feedback control module returns a feedback end signal to the main control unit, and the system returns to the low-noise voice control standby state, waiting for the next command input.
[0114] Among them, the feedback control module is the core hardware component that realizes non-voice feedback. It can realize corresponding feedback actions by controlling the air conditioner's light drive circuit, fan control circuit, etc. The preset feedback rules can be automatically matched based on the scene type, or they can be customized by the user through the mobile APP or air conditioner panel to meet the usage habits of different users.
[0115] In some embodiments, the recognition unit 103 is further configured to: shut down the voice recognition module and keep the low-power monitoring module running when the low-noise triggering condition is continuously met and no new voice signal is detected within a preset time period.
[0116] In low-noise scenarios, user voice interactions are typically intermittent. Keeping the voice recognition module running continuously without new interactions would result in unnecessary energy waste, violating the energy-saving requirements of low-noise environments. Temporarily shutting down the voice recognition module can promptly reduce system energy consumption when there are no further user actions, achieving an organic balance between accurate monitoring and energy-saving operation in low-noise scenarios.
[0117] Specifically, after the low-noise trigger condition is met, the system activates the low-noise voice control mode. At this time, the voice recognition module is in working condition, and the low-power monitoring module simultaneously monitors the voice signal. The system's built-in timing module starts timing, and in real time, it counts the non-interaction duration since the last valid voice signal was detected (or after the low-noise mode was activated). During the timing process, the low-power monitoring module continuously monitors the surrounding voice signals. If a new voice signal is detected, the timing module immediately resets and restarts counting the non-interaction duration. The voice recognition module remains in working condition to meet new recognition needs. If the non-interaction duration counted by the timing module reaches the preset duration, and no new voice signal is detected during this period, the system determines that the user has no intention of further interaction and sends a shutdown command to the voice recognition module, controlling the voice recognition module to stop the recognition and parsing function, retaining only basic state monitoring. After the voice recognition module is shut down, the low-power monitoring module continues to maintain the basic voice feature monitoring function, operating in a low-power state, and continuously capturing the surrounding voice signals. When the low-power monitoring module detects a new voice signal that matches the preset characteristics again, it immediately sends a wake-up signal to the system. The system quickly activates the voice recognition module, restoring the full functionality of the low-noise voice control mode, ensuring that new user commands can be recognized and parsed in a timely manner.
[0118] The preset duration can be flexibly adjusted according to the actual application scenario, with a typical setting range of 1-5 minutes. This avoids frequent start-stop cycles caused by closing the module after a short period of no interaction, and effectively controls energy consumption during long periods of no interaction. When maintaining basic voice feature monitoring, the low-power monitoring module consumes only a very small percentage of the power of the voice recognition module, which can significantly reduce the overall energy consumption of the air conditioner.
[0119] This solution achieves precise triggering of low-noise voice control mode through joint judgment of environmental noise parameters and scene information, effectively solving the problem of loud wake-up noise and easy disturbance to others in low-noise scenarios of traditional air conditioning voice control systems; with the targeted monitoring of low-power monitoring modules, the system significantly reduces operating energy consumption while ensuring the capture of effective voice signals, solving the drawback of traditional systems constantly resident in high-power monitoring state; through preset whisper features and a dedicated whisper recognition mechanism, the accuracy of whisper signal recognition is significantly improved, making up for the lack of sensitivity of traditional voice recognition modules to soft speech or whispers.
[0120] Since the processing and functions implemented by the device in this embodiment are basically the same as the embodiments, principles and examples of the aforementioned methods, any details not covered in the description of this embodiment can be found in the relevant descriptions in the aforementioned embodiments, and will not be repeated here.
[0121] The air conditioner using the technical solution of this invention includes an ambient sound detection module, a low-power monitoring module, and a voice recognition module. The low-power monitoring module can monitor voice features in a low-power state. The ambient sound detection module acquires environmental noise parameters and, combined with current scene information, determines whether the low-noise trigger condition is met. If met, the low-power monitoring module is activated to monitor the user's voice signal. When the detected voice signal matches preset whispered characteristics, the voice recognition module uses a preset whispered recognition mechanism to recognize and parse the voice signal to obtain the air conditioner control command, thereby controlling the air conditioner to execute the command. This effectively reduces wake-up noise from air conditioner voice control, improves the recognition sensitivity of soft and whispered commands, enhances the system's environmental adaptability to low-noise scenarios such as sleep and meetings, and reduces operating power consumption in voice monitoring mode.
[0122] According to an embodiment of the present invention, an air conditioner corresponding to an air conditioner control device is also provided. This air conditioner may include the air conditioner control device described above.
[0123] Since the processing and functions implemented by the air conditioner in this embodiment are basically the same as the embodiments, principles and examples of the aforementioned device, any details not covered in the description of this embodiment can be found in the relevant descriptions in the aforementioned embodiments, and will not be repeated here.
[0124] The air conditioner using the technical solution of this invention includes an ambient sound detection module, a low-power monitoring module, and a voice recognition module. The low-power monitoring module can monitor voice features in a low-power state. The ambient sound detection module acquires environmental noise parameters and, combined with current scene information, determines whether the low-noise trigger condition is met. If met, the low-power monitoring module is activated to monitor the user's voice signal. When the detected voice signal matches preset whispered characteristics, the voice recognition module uses a preset whispered recognition mechanism to recognize and parse the voice signal to obtain the air conditioner control command, thereby controlling the air conditioner to execute the command. This effectively reduces wake-up noise from air conditioner voice control, improves the recognition sensitivity of soft and whispered commands, enhances the system's environmental adaptability to low-noise scenarios such as sleep and meetings, and reduces operating power consumption in voice monitoring mode.
[0125] According to an embodiment of the present invention, a storage medium corresponding to an air conditioner control method is also provided, the storage medium including a stored program, wherein the program controls the device where the storage medium is located to execute the air conditioner control method described above when it is executed.
[0126] Since the processing and functions implemented by the storage medium in this embodiment are basically the same as the embodiments, principles and examples of the aforementioned methods, any details not covered in this embodiment can be found in the relevant descriptions in the aforementioned embodiments, and will not be repeated here.
[0127] The air conditioner using the technical solution of this invention includes an ambient sound detection module, a low-power monitoring module, and a voice recognition module. The low-power monitoring module can monitor voice features in a low-power state. The ambient sound detection module acquires environmental noise parameters and, combined with current scene information, determines whether the low-noise trigger condition is met. If met, the low-power monitoring module is activated to monitor the user's voice signal. When the detected voice signal matches preset whispered characteristics, the voice recognition module uses a preset whispered recognition mechanism to recognize and parse the voice signal to obtain the air conditioner control command, thereby controlling the air conditioner to execute the command. This effectively reduces wake-up noise from air conditioner voice control, improves the recognition sensitivity of soft and whispered commands, enhances the system's environmental adaptability to low-noise scenarios such as sleep and meetings, and reduces operating power consumption in voice monitoring mode.
[0128] According to an embodiment of the present invention, a computer program product corresponding to the control method for an air conditioner is also provided. The computer program product includes a computer program that, when processed and executed, implements the steps of the control method for the air conditioner described above.
[0129] Since the processing and functions implemented by the computer program product in this embodiment are basically corresponding to the embodiments, principles and examples of the aforementioned methods, any details not covered in the description of this embodiment can be found in the relevant descriptions in the aforementioned embodiments, and will not be repeated here.
[0130] The air conditioner using the technical solution of this invention includes an ambient sound detection module, a low-power monitoring module, and a voice recognition module. The low-power monitoring module can monitor voice features in a low-power state. The ambient sound detection module acquires environmental noise parameters and, combined with current scene information, determines whether the low-noise trigger condition is met. If met, the low-power monitoring module is activated to monitor the user's voice signal. When the detected voice signal matches preset whispered characteristics, the voice recognition module uses a preset whispered recognition mechanism to recognize and parse the voice signal to obtain the air conditioner control command, thereby controlling the air conditioner to execute the command. This effectively reduces wake-up noise from air conditioner voice control, improves the recognition sensitivity of soft and whispered commands, enhances the system's environmental adaptability to low-noise scenarios such as sleep and meetings, and reduces operating power consumption in voice monitoring mode.
[0131] In summary, it is readily understood by those skilled in the art that, without conflict, the aforementioned advantageous methods can be freely combined and superimposed.
[0132] The above description is merely an embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of the claims of the present invention.
Claims
1. A method for controlling an air conditioner, characterized in that, The air conditioner includes an ambient sound detection module, a low-power monitoring module, and a voice recognition module. The low-power monitoring module is capable of monitoring voice features in a low-power state. The control method includes: The ambient noise parameters are obtained through the ambient sound detection module, and it is determined whether the low noise triggering condition is met based on the ambient noise parameters and the current scene information. If the low-noise triggering condition is met, the low-power monitoring module is activated, and the user's voice signal is monitored for features through the low-power monitoring module. When the voice signal is detected to match the preset whisper features, the voice recognition module activates the preset whisper recognition mechanism to recognize and analyze the voice signal, thereby obtaining the control command for the air conditioner. Control the air conditioner to execute the control command.
2. The air conditioning control method according to claim 1, characterized in that, Determine whether the low-noise triggering condition is met based on the environmental noise parameters and the current scene information, including: If the environmental noise parameter is less than the preset noise threshold, or if the scene information is a preset low-noise sensitive scene, then the low-noise triggering condition is determined to be met.
3. The air conditioning control method according to claim 1, characterized in that, The voice recognition module uses a preset whisper recognition mechanism to recognize and parse the voice signal to obtain the control commands for the air conditioner, including: Extract the cepstral coefficient features, spectral slope features, and low-frequency energy ratio features of the speech signal, and generate a feature vector; The feature vector is input into a preset neural network model for joint acoustic and language recognition to obtain the control commands for the air conditioner.
4. The air conditioning control method according to claim 1 or 2, characterized in that, The speech recognition module includes a main processing unit and a whisper recognition unit; The method further includes: When the low-noise triggering condition is met, the main processing unit is controlled to enter a sleep state and the whisper recognition unit is woken up. After being woken up, the whisper recognition unit performs the operation of recognizing and parsing the speech signal. When the low-noise triggering condition is not met, the whisper recognition unit is controlled to enter a sleep state and the main processing unit is woken up.
5. The air conditioning control method according to claim 1, characterized in that, The method further includes: After the air conditioner executes the control command, the system provides feedback on the command execution status to the user in a non-voice manner; the non-voice manner includes flashing lights and changes in airflow.
6. The air conditioning control method according to claim 1, characterized in that, The method further includes: When the low-noise triggering condition is continuously met and no new voice signal is detected within a preset time period, the voice recognition module is turned off, while the low-power monitoring module remains in operation.
7. A control device for an air conditioner, characterized in that, The air conditioner includes an ambient sound detection module, a low-power monitoring module, and a voice recognition module. The low-power monitoring module is capable of monitoring voice features in a low-power state. The control device includes: The determination unit is configured to acquire environmental noise parameters through the environmental sound detection module, and determine whether the low noise triggering condition is met based on the environmental noise parameters and the current scene information. The monitoring unit is configured to activate the low-power monitoring module if the low-noise triggering condition is met, and to perform feature monitoring on the user's voice signal through the low-power monitoring module. The recognition unit is configured such that when the voice signal is detected to match the preset whisper features, the voice recognition module activates the preset whisper recognition mechanism to recognize and analyze the voice signal, thereby obtaining the control command of the air conditioner; The control unit is configured to control the air conditioner to execute the control commands.
8. An air conditioner, characterized in that, include: The air conditioning control device as described in claim 7.
9. A storage medium, characterized in that, The storage medium includes a stored program, wherein, when the program is executed, it controls the device containing the storage medium to perform the air conditioning control method according to any one of claims 1 to 6.
10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the air conditioning control method according to any one of claims 1 to 6.