Air conditioner and control method and device thereof, storage medium and computer program product

By combining camera equipment and millimeter-wave radar, the air conditioner can automatically identify the location of furniture and the user's trajectory, build an accurate indoor space model, solve the problem of manual annotation by the user in the existing technology, improve the intelligence and comfort of air conditioner control, and ensure privacy and security.

CN120868565APending Publication Date: 2025-10-31GREE ELECTRIC APPLIANCE INC OF ZHUHAI
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202511090284.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-05
Publication Date
2025-10-31

AI Technical Summary

Technical Problem

Existing automatic air conditioning control systems require users to manually confirm the location of furniture. Image recognition and radar data are not effectively integrated, making it difficult to form a stable and efficient closed loop for home space perception and control.

Method used

Panoramic images are acquired by camera equipment for image recognition. The user's position is detected by millimeter-wave radar. The furniture position is automatically identified using target detection and segmentation models. A spatial structure model is constructed using depth estimation algorithms and extended Kalman filtering algorithms to achieve fusion of the three-dimensional coordinates of the furniture and the user's position. The operation of the air conditioner is controlled by combining the user's activity trajectory.

Benefits of technology

It enables automatic modeling of indoor spaces without manual annotation by the user, accurately reconstructs the indoor space structure, precisely senses the user area, improves the intelligence and comfort of air conditioning control, and has energy-saving effects while ensuring privacy and security.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120868565A_ABST
    Figure CN120868565A_ABST
Patent Text Reader

Abstract

The invention provides an air conditioner and a control method and device thereof, a storage medium and a computer program product. The method comprises the steps that a panoramic image, shot by camera equipment, of a space where the air conditioner is located is obtained; performing image recognition on the acquired panoramic image to obtain the types of the home devices in the space and the positions of the home devices in the panoramic image; detecting the position of a user in the space and the human body skeleton information of the user through a millimeter wave radar, and outputting the three-dimensional coordinates of the space containing the position of the user and the movement track of the user; fusing the obtained position of each home device in the space in the panoramic image with the three-dimensional coordinate of the space output by the millimeter wave radar, and constructing a space structure model of the space; and controlling the operation of the air conditioner by combining the constructed space structure model and the user movement track detected by the millimeter wave radar. According to the scheme, indoor space structure modeling can be completed without user operation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of control, and more particularly to an air conditioner and its control method, apparatus, storage medium and computer program product. Background Technology

[0002] With the development of IoT technology, smart home control systems have been widely used in home environments. In the field of air conditioning control, traditional methods relying on timers, remote controls, or mobile apps are gradually evolving towards higher levels of automation. Modern systems are beginning to introduce sensors (such as infrared sensors, temperature and humidity sensors, and Wi-Fi modules) to sense environmental changes or user behavior, thereby achieving a certain degree of automatic operation.

[0003] Related technologies are beginning to combine image recognition with sensor data to build indoor maps or assist in control. However, these systems are still in the early stages of exploration, with applications mainly concentrated in security patrols or service robot navigation, and a stable and efficient closed-loop system for home space perception and control has not yet been formed. Air conditioning automatic control systems in related technologies generally use millimeter-wave radar to monitor user location, combined with time-based air conditioning strategies. For example, settings might include "automatically turning on cooling when someone approaches the air-conditioned area" or "entering energy-saving mode 15 minutes after the user leaves." This type of technology achieves a degree of automation by detecting user location, but furniture arrangement information is usually manually entered by the user or marked via an app during the initial installation phase. More advanced systems attempt to capture images using indoor cameras and identify the locations of devices such as air conditioners and sofas, then compare them with millimeter-wave data to assist in spatial reasoning. However, furniture positions still require manual confirmation by the user, lacking automatic modeling capabilities, and image recognition and radar data are not effectively integrated, serving only as complementary references and making it difficult to reconstruct the complete spatial structure. Summary of the Invention

[0004] The main objective of this invention is to overcome the deficiencies of the aforementioned related technologies and provide an air conditioner and its control method, device, storage medium, and computer program product to solve the problem that the location of home appliances in the related technologies needs to be manually confirmed by the user.

[0005] This invention provides a method for controlling an air conditioner, comprising: acquiring a panoramic image of a space captured by a camera; the panoramic image containing various home appliances in the space; performing image recognition on the acquired panoramic image to obtain the types and positions of the various home appliances in the space; detecting the user's location and skeletal information within the space using millimeter-wave radar, and outputting the three-dimensional coordinates of the space containing the user's location and the user's activity trajectory; fusing the obtained positions of the various home appliances in the space in the panoramic image with the three-dimensional coordinates of the space output by the millimeter-wave radar to construct a spatial structure model of the space; and controlling the operation of the air conditioner by combining the constructed spatial structure model and the user's activity trajectory detected by the millimeter-wave radar.

[0006] Optionally, image recognition is performed on the acquired panoramic image to obtain the types and positions of each home furnishing device in the space, including: inputting the acquired panoramic image into a pre-trained target detection model for target proximity detection to obtain the type and position detection boxes of each home furnishing device in the space; inputting the identified position detection boxes of each home furnishing device in the panoramic image into a pre-trained target segmentation model for edge segmentation to obtain the outline of each home furnishing device.

[0007] Optionally, the positions of each home appliance in the space obtained in the panoramic image are fused with the three-dimensional coordinates of the space output by the millimeter-wave radar to construct a spatial structure model of the space. This includes: estimating the estimated three-dimensional positions of each home appliance in the space in the coordinate system of the camera device using a depth estimation algorithm; converting the estimated three-dimensional positions of each home appliance in the coordinate system of the camera device into three-dimensional coordinates in the three-dimensional coordinate system of the millimeter-wave radar; and fusing the converted three-dimensional coordinates of each home appliance in the three-dimensional coordinate system of the millimeter-wave radar with the three-dimensional coordinates of the space output by the millimeter-wave radar to construct a spatial structure model of the space.

[0008] Optionally, controlling the operation of the air conditioner by combining the constructed spatial structure model and the user activity trajectory detected by the millimeter-wave radar includes: dividing the spatial structure model into two or more regions using a spatial partitioning algorithm; superimposing the user activity trajectory detected by the millimeter-wave radar onto the spatial structure model to determine whether the user has entered any of the two or more regions; if it is determined that the user has entered any of the two or more regions and the triggering condition of the preset scene mode corresponding to that region is met, then controlling the air conditioner to execute the preset scene mode corresponding to that region.

[0009] Optionally, the triggering conditions for the preset scene mode corresponding to any one of the two or more areas include: the user entering the area and receiving a voice command to activate the preset scene mode corresponding to the area.

[0010] Optionally, controlling the air conditioner to execute a preset scene mode corresponding to the area includes: controlling the air conditioner to execute air conditioner control parameters corresponding to the preset scene mode corresponding to the area, wherein different air conditioner control parameters correspond to different scene modes; the different air conditioner control parameters corresponding to different scene modes are set by the user, or set according to the user's operation behavior data of the air conditioner.

[0011] Optionally, it further includes: after controlling the air conditioner to execute the preset scene mode corresponding to the area, detecting whether the user has left the area; if the user is detected to have left the area and the departure time exceeds a first preset time, controlling the air conditioner to switch to a preset energy-saving mode.

[0012] Optionally, the method further includes: acquiring historical activity data of the user in the two or more areas, the historical activity data including: time information and control information of the user appearing in each of the two or more areas each day; predicting the area the user will enter in the future and the control intention based on the acquired historical activity data using a time-series prediction algorithm; the control intention being whether the user will issue a voice command to activate the preset scene mode corresponding to that area; and controlling the air conditioner in advance based on the predicted area the user will enter and the control intention.

[0013] Optionally, it further includes: acquiring the user's historical activity trajectory detected by the millimeter-wave radar; superimposing the user's historical activity trajectory onto the spatial structure model to form a heat map, so as to determine the user's active area in the two or more regions; and controlling the air conditioner according to the determined user's active area in the two or more regions.

[0014] Another aspect of the present invention provides a control device for an air conditioner, comprising: a first acquisition unit, configured to acquire a panoramic image of the space captured by a camera device; the panoramic image including various home appliances in the space; an identification unit, configured to perform image recognition on the panoramic image acquired by the first acquisition unit to obtain the types and positions of the various home appliances in the space within the panoramic image; a detection unit, configured to detect the location of a user and the user's skeletal information within the space using millimeter-wave radar, and output the three-dimensional coordinates of the space including the user's location and the user's activity trajectory; a construction unit, configured to fuse the obtained positions of the various home appliances in the space within the panoramic image with the three-dimensional coordinates of the space output by the millimeter-wave radar to construct a spatial structure model of the space; and a control unit, configured to control the operation of the air conditioner by combining the spatial structure model constructed by the construction unit and the user activity trajectory detected by the detection unit using the millimeter-wave radar.

[0015] Optionally, the recognition unit performs image recognition on the panoramic image acquired by the first acquisition unit to obtain the types and positions of each home furnishing device in the space, including: inputting the acquired panoramic image into a pre-trained target detection model for target proximity detection to obtain the type and position detection boxes of each home furnishing device in the space; inputting the identified position detection boxes of each home furnishing device in the panoramic image into a pre-trained target segmentation model for edge segmentation to obtain the outline of each home furnishing device.

[0016] Optionally, the construction unit fuses the obtained positions of each home appliance in the space in the panoramic image with the three-dimensional coordinates of the space output by the millimeter-wave radar to construct a spatial structure model of the space, including: estimating the estimated three-dimensional positions of each home appliance in the space in the coordinate system of the camera device using a depth estimation algorithm; converting the estimated three-dimensional positions of each home appliance in the coordinate system of the camera device into three-dimensional coordinates in the three-dimensional coordinate system of the millimeter-wave radar; and fusing the converted three-dimensional coordinates of each home appliance in the three-dimensional coordinate system of the millimeter-wave radar with the three-dimensional coordinates of the space output by the millimeter-wave radar to construct a spatial structure model of the space.

[0017] Optionally, the control unit, combining the constructed spatial structure model and the user activity trajectory detected by the millimeter-wave radar, controls the operation of the air conditioner, including: dividing the spatial structure model into two or more regions using a spatial partitioning algorithm; superimposing the user activity trajectory detected by the millimeter-wave radar onto the spatial structure model to determine whether the user has entered any of the two or more regions; if it is determined that the user has entered any of the two or more regions and the preset scene mode triggering condition corresponding to that region is met, then controlling the air conditioner to execute the preset scene mode corresponding to that region.

[0018] Optionally, the triggering conditions for the preset scene mode corresponding to any one of the two or more areas include: the user entering the area and receiving a voice command to activate the preset scene mode corresponding to the area.

[0019] Optionally, the control unit controls the air conditioner to execute a preset scene mode corresponding to the area, including: controlling the air conditioner to execute air conditioner control parameters corresponding to the preset scene mode corresponding to the area, wherein different air conditioner control parameters correspond to different scene modes; the different air conditioner control parameters corresponding to different scene modes are set by the user, or set according to the user's operation behavior data of the air conditioner.

[0020] Optionally, the control unit is further configured to: after controlling the air conditioner to execute the preset scene mode corresponding to the area, detect whether the user has left the area; if the user is detected to have left the area and the departure time exceeds a first preset time, control the air conditioner to switch to a preset energy-saving mode.

[0021] Optionally, it further includes: a second acquisition unit, configured to acquire historical activity data of the user in the two or more areas, the historical activity data including: time information and control information of the user appearing in each of the two or more areas each day; a prediction unit, configured to predict the area the user will enter in the future and the control intention based on the historical activity data acquired by the second acquisition unit using a time-series prediction algorithm; the control intention being whether the user will issue a voice command to activate the preset scene mode corresponding to that area; the control unit is further configured to: control the air conditioner in advance based on the area the user will enter in advance and the control intention predicted by the prediction unit.

[0022] Optionally, it further includes: a third acquisition unit and a determination unit; the third acquisition unit is used to acquire the user's historical activity trajectory detected by the millimeter-wave radar; the determination unit is used to superimpose the user's historical activity trajectory acquired by the third acquisition unit onto the spatial structure model to form a heat map, so as to determine the user's active area in the two or more regions; the control unit is further used to: control the air conditioner according to the determined user's active area in the two or more regions.

[0023] In another aspect, the present invention provides a storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of any of the methods described above.

[0024] In another aspect, the present invention provides an air conditioner, including a processor, a memory, and a computer program stored in the memory that can run on the processor, wherein the processor executes the program to implement the steps of any of the methods described above.

[0025] In another aspect, the present invention provides an air conditioner including any of the control devices described above.

[0026] In another aspect, the present invention provides a computer program product, including a computer program that, when executed by a processor, implements the steps of any of the methods described above.

[0027] According to the technical solution of the present invention, the following beneficial effects are achieved:

[0028] 1. Achieve true "zero manual annotation" automatic modeling capability for interior spaces:

[0029] Related technologies require users to manually mark the locations of devices such as air conditioners and sofas in an app. This invention, however, automatically identifies uploaded images using a target detection model (YOLOv5) and a target segmentation model (Mask R-CNN), and then uses millimeter-wave radar to achieve 3D coordinate positioning. Furniture layout recognition and functional area division can be completed without any user intervention. This technology significantly lowers the user barrier and improves the system's automation level.

[0030] 2. By fusing imagery and radar data, the indoor spatial structure can be accurately reconstructed:

[0031] This invention uses a coordinate transformation and extended Kalman filter fusion algorithm based on image recognition and millimeter-wave radar perception to accurately estimate the actual positional relationship between each piece of furniture and the wall. It establishes a consistent three-dimensional spatial coordinate system among different sensor data, achieving high-precision spatial modeling and providing a solid foundation for subsequent behavior recognition and air conditioning control.

[0032] 3. Accurately detects the user's current location, enabling more targeted control:

[0033] This invention uses millimeter-wave skeleton tracking technology to identify the user's functional area in real time and generates an "activity heat map" by combining historical trajectory data. This accurately determines the user's active area, enabling the air conditioner to implement precise control strategies such as "directional air delivery" and "air delivery avoiding people," thereby improving user comfort.

[0034] 4. The voice + location dual-trigger mechanism enhances the naturalness of interaction and reduces the tolerance for accidental touches:

[0035] Voice control in related technologies relies solely on command recognition, which is prone to false triggering. This invention introduces "current spatial location" as a prerequisite for scene determination, ensuring that control logic is executed only when the user is in a specific area and issues a corresponding voice command, significantly improving the system's intelligence and reliability.

[0036] 5. Automatically generate air conditioning operation strategies and support personalized scene matching:

[0037] This invention can automatically generate air conditioning parameter settings (temperature, wind speed, air swing direction, etc.) based on the identified scene mode (such as sofa area + movie mode), without requiring users to preset operations, greatly improving the level of intelligence and ease of use.

[0038] 6. Significant energy-saving effect; the system can dynamically switch operating states:

[0039] This invention automatically switches the air conditioner to energy-saving mode (such as increasing the temperature and reducing the fan speed) after detecting that the user has been away from the functional area for a long time, effectively reducing unnecessary energy consumption.

[0040] 7. Robust privacy and security mechanisms ensure all data processing is completed locally:

[0041] All image analysis and behavior judgment in this invention can be completed locally in the edge computing gateway without uploading images to the cloud, ensuring that the user's home space privacy is not leaked and enhancing user trust. Attached Figure Description

[0042] The accompanying drawings, which are included to provide a further understanding of the invention and form part of this invention, illustrate exemplary embodiments of the invention and are used to explain the invention, but do not constitute an undue limitation of the invention. In the drawings:

[0043] Figure 1 This is a schematic diagram of an embodiment of the air conditioner control method provided by the present invention;

[0044] Figure 2 The diagram illustrates a specific implementation of the present invention's steps for performing image recognition on the acquired panoramic image to obtain the types of home appliances in the space and their positions in the panoramic image.

[0045] Figure 3 The diagram illustrates a specific implementation of the steps for fusing the positions of the identified home appliances in the space in the panoramic image with the three-dimensional coordinates of the area containing the human body output by the millimeter-wave radar to obtain a spatial structure model of the space.

[0046] Figure 4The diagram illustrates a specific implementation of the steps for controlling the operation of the air conditioner by combining the constructed spatial structure model and the user activity trajectory detected by the millimeter-wave radar.

[0047] Figure 5 This is a schematic diagram of another embodiment of the air conditioner control method provided by the present invention;

[0048] Figure 6 This is a schematic diagram of a specific embodiment of the air conditioner control method provided by the present invention;

[0049] Figure 7 This is a structural block diagram of an embodiment of the air conditioner control device provided by the present invention;

[0050] Figure 8 This is a structural block diagram of another embodiment of the air conditioner control device provided by the present invention. Detailed Implementation

[0051] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below in conjunction with specific embodiments and corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of this invention, and not all of them. Based on the embodiments of this invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this invention.

[0052] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0053] The relevant technologies fail to provide a complete solution for automatically identifying indoor structures, determining human activity areas, and dynamically triggering air conditioning operation scenarios without user intervention.

[0054] This invention provides a method for controlling an air conditioner. This control method can be implemented in a local gateway (e.g., an edge gateway) or a server. Specifically, the edge gateway can be a small smart host deployed in a user's home, such as an embedded box, a smart home center, a router built-in module, or a high-performance IoT controller (such as a Raspberry Pi, Qualcomm IoT chip module, etc.).

[0055] Figure 1 This is a schematic diagram of an embodiment of the air conditioner control method provided by the present invention.

[0056] like Figure 1 As shown, according to an embodiment of the present invention, the air conditioner control method includes at least steps S110, S120, S130, S140 and S150.

[0057] Step S110: Obtain a panoramic image of the space captured by a camera device.

[0058] Specifically, the panoramic image includes all the home appliances in the space. In one specific embodiment, a panoramic image of the space taken by the user through the camera device is acquired. For example, when a user controls the air conditioner for the first time using an app installed on a mobile terminal, the user takes a panoramic image of the room through the mobile terminal. The image includes all the home appliances in the space, such as doors, windows, furniture, and appliances. The user can follow the guidance to take photos of the room in good daylight conditions, preferably from a top-down diagonal perspective, covering major furniture such as the air conditioner, sofa, television, and dining table. After taking the photo, it is uploaded to a local gateway (e.g., a smart air conditioner with certain computing and communication capabilities) or a server through the mobile terminal.

[0059] Step S120: Perform image recognition on the acquired panoramic image to obtain the types of home appliances in the space and their positions in the panoramic image.

[0060] Figure 2 The diagram illustrates a specific implementation of the present invention's steps for performing image recognition on the acquired panoramic image to obtain the types of home appliances in the space and their positions in the panoramic image.

[0061] like Figure 2 As shown, in one specific embodiment, step S120 may specifically include steps S121 and S122.

[0062] Step S121: Input the acquired panoramic image into a pre-trained target detection model for target detection to obtain the types of home appliances in the space and their location detection boxes in the panoramic image.

[0063] Specifically, after acquiring a panoramic image of the space, the panoramic image can be input into a pre-trained target detection model for target detection to obtain the types (e.g., air conditioners, sofas, doors and windows) and location detection boxes of various home appliances in the space.

[0064] The object detection model can specifically be the YOLOv5 (You Only Look Once, version 5) object detection model. The YOLOv5 object detection model is pre-trained to detect objects in home furnishings (e.g., appliances, furniture, doors, windows, etc.). The output of the object detection model is the type and location bounding boxes (e.g., rectangles) for each home furnishing item. Using the YOLOv5 object detection algorithm for object detection offers fast processing speed, wide coverage, and rapid filtering of candidate target regions.

[0065] Step S122: Input the identified location detection boxes of each home appliance in the panoramic image into a pre-trained target segmentation model for edge segmentation to obtain the outline of each home appliance.

[0066] To improve recognition accuracy, the detected location bounding boxes of each home appliance in the identified space are further input into a pre-trained target segmentation model for edge segmentation, obtaining the precise contours of each home appliance in the panoramic image, making subsequent spatial mapping more accurate. Specifically, the target segmentation model can be a Mask R-CNN target segmentation model. That is, using the Mask R-CNN (Region Convolutional Neural Network) algorithm, pixel-level edge segmentation is performed on the detected location bounding boxes of each home appliance to obtain the precise contour of each appliance. For example, for each rectangular region detected by the YOLOv5 model, this rectangular region is fed into the Mask R-CNN model. The Mask R-CNN model further refines this defined area, determining which pixels belong to the target home appliance and which belong to the background, outputting a mask image, i.e., the precise contour boundary of each home appliance in the image.

[0067] This invention fuses the "target category and coarse location" provided by YOLOv5 with the "pixel-level boundary" provided by Mask R-CNN. YOLOv5 is responsible for "finding" the target, while Mask R-CNN is responsible for "depicting" it clearly, ultimately providing a more detailed geometric description of each target home appliance in the panoramic image, including its position, category, area, and shape. This combination of "detection box + segmentation mask" combines the advantages of both algorithms: YOLOv5 provides high speed and high recall, ensuring no missed detections; Mask R-CNN provides fine contours, which is beneficial for subsequent spatial localization and volume estimation. Using rectangular boxes as input regions also significantly reduces the computational burden of Mask R-CNN, improving segmentation efficiency. Therefore, this invention employs a two-stage recognition process of "YOLOv5 detection first, Mask R-CNN refinement later" to achieve both fast and accurate target recognition and contour extraction, providing high-quality input for subsequent 3D spatial reconstruction and functional area division.

[0068] Step S130: Detect the user's location and skeletal information within the space using millimeter-wave radar, and output the three-dimensional coordinates of the space containing the user's location and the user's activity trajectory.

[0069] Specifically, by deploying millimeter-wave radar indoors, and continuously outputting human point cloud information at preset intervals (e.g., 20 milliseconds), it is possible to accurately sense "whether there is someone in the room", "where the person is", and "what their movement trajectory is", without affecting privacy or depending on lighting conditions.

[0070] Millimeter-wave radar detects the signal characteristics reflected by objects (especially the human body) in space by transmitting and receiving high-frequency electromagnetic waves. This allows it to identify the human body's position in three-dimensional space, its posture and skeletal structure (i.e., skeletal information), and whether the body is moving, stationary, or performing a certain action (such as sitting or walking). Specifically, the radar senses the spatial location and energy distribution of "multiple reflection points" in each time frame, clusters these points using algorithms, and fits the human body shape.

[0071] The millimeter-wave radar outputs the following core data at preset intervals (e.g., 20ms):

[0072] The three-dimensional coordinates (x, y, z) of the human body's position (e.g., the center point): represent the current position of the human body relative to the radar coordinate system;

[0073] Key point coordinates of the human skeleton: such as the three-dimensional coordinates of the head, shoulders, knees, and soles of the feet (depending on radar capabilities);

[0074] Human motion trajectory: Records the changes in the center position of the human body over a continuous period of time, forming a trajectory sequence;

[0075] The core function of millimeter-wave radar is to provide the three-dimensional position and continuous movement trajectory of the human body in real space, providing crucial support for subsequent steps.

[0076] Step S140: The positions of each home appliance in the space identified in the panoramic image are fused with the three-dimensional coordinates of the area containing the human body output by the millimeter-wave radar to construct a spatial structure model of the space.

[0077] Figure 3 The diagram illustrates a specific implementation of the steps for fusing the positions of the identified home appliances in the space within the panoramic image with the three-dimensional coordinates of the area containing the human body output by the millimeter-wave radar to obtain a spatial structure model of the space.

[0078] like Figure 3 As shown, in one specific embodiment, step S140 may specifically include: step S141, step S142 and step S143.

[0079] Step S141: The estimated three-dimensional position of each home appliance in the space in the coordinate system of the camera device is obtained by using a depth estimation algorithm.

[0080] In one specific implementation, the panoramic image is input into a pre-trained depth estimation model to obtain the depth value of each pixel in the panoramic image; the depth value of each pixel is then correlated with the identified positions of each home appliance in the space within the panoramic image to obtain the estimated three-dimensional position of each home appliance in the coordinate system of the camera device.

[0081] The depth estimation model can specifically be a MiDaS monocular depth estimation model. That is, the MiDaS monocular depth estimation algorithm estimates the depth value of each pixel in the panoramic image, thereby obtaining the depth information of the positions of each home appliance in the panoramic image, realizing the reconstruction of the actual spatial distance of the home appliances from the planar image. Since the panoramic image of the space captured by the camera only contains two-dimensional image information and does not contain depth information, the MiDaS monocular image depth estimation algorithm is introduced. The MiDaS monocular image depth estimation algorithm can determine the distance of each pixel in the image from the camera device. MiDaS (Monocular Depth Sensing) is a monocular image-based depth estimation algorithm that can predict the relative depth map of a scene from a regular RGB image through deep learning. Inputting the original panoramic image into the MiDaS monocular image depth estimation model outputs a depth map of the same size as the image, with each pixel having a relative depth value. Combining this with the pixel position of each home appliance in the image, these depth values ​​are mapped to each identified home appliance to obtain an estimated three-dimensional position of each home appliance from the camera device's perspective. The original panoramic image is input into the MiDaS monocular depth estimation model to obtain the depth value of each pixel and output a depth map. The larger the value, the farther the pixel is from the camera device. The depth value of each pixel is matched with the position (position detection box) of each identified home appliance in the panoramic image. The depth of the corresponding region is extracted according to the position of the detection box in the panoramic image to obtain the average / center depth value of each home appliance and the estimated three-dimensional position of each home appliance in the coordinate system of the camera device.

[0082] For example, the bounding box of the "sofa" in the image (i.e., the location detection box output by the location object detection model, such as the rectangle output by the YOLOv5 model) is located at (x1, y1, x2, y2), and the average depth of the corresponding area in the depth map is 2 meters. Based on this, it can be determined that the sofa is 2 meters in front of the camera, at a 30-degree angle to the right.

[0083] The original image coordinate system of a camera device is a two-dimensional pixel coordinate system (i.e., the x / y axis of the image). However, when the image incorporates depth estimation results, each pixel is no longer just "positional," but has a "depth value." That is, each pixel can be represented as: (x... 像素 ,y 像素 ,d 深度At this point, the image coordinates plus the depth value constitute a three-dimensional estimated coordinate system under the camera coordinate system, which is related to the actual position of the object in real space relative to the camera. Although this coordinate system is still "estimated," it can already reflect the spatial layout relationship of the furniture relative to the camera. For example: - The sofa is located in the lower right corner of the image, with a depth of 1.5 meters; - The air conditioner is located in the upper left corner of the image, with a depth of 3.2 meters; their "relative spatial positions" under the camera's view are clear. The obtained estimated three-dimensional position of the furniture refers to its estimated spatial position in the coordinate system of the camera device, no longer a pure two-dimensional pixel coordinate. Through the restoration process of "pixel coordinates + depth value," the two-dimensional image information has actually been "upgraded" to a three-dimensional coordinate space. This three-dimensional coordinate is still under the "camera coordinate system," that is, with the camera as the origin (0,0,0), the front as the Z-axis, the left and right as the X-axis, and the up and down as the Y-axis, and the unit can be meters or centimeters (according to the depth model). Therefore, the final position of each piece of furniture is an "estimated three-dimensional position," and its coordinate form is (X,Y,Z), which is used for subsequent transformation to the radar coordinate system or spatial model.

[0084] The MiDaS monocular depth estimation algorithm can obtain roughly reliable object distance relationships from ordinary photographs without the need for lasers, binocular cameras, or dedicated sensors, providing crucial support for automatic indoor space modeling. This algorithm boasts high execution efficiency, making it suitable for deployment in edge computing gateways and meeting the real-time and resource consumption requirements of smart home systems. Therefore, the technical solution proposed in this invention is fully compatible with ordinary consumer-grade cameras or smartphones in the image acquisition stage, requiring no replacement or configuration of special equipment, and possesses significant practicality and promotional value.

[0085] Step S142: The estimated three-dimensional position of each home appliance in the coordinate system of the camera device is obtained and converted into three-dimensional coordinates in the three-dimensional coordinate system of the millimeter-wave radar.

[0086] In one specific embodiment, based on the pre-calibrated spatial mapping relationship between the coordinate system of the camera device and the three-dimensional coordinate system of the millimeter-wave radar, the estimated three-dimensional position of each home appliance in the coordinate system of the camera device is projected into the three-dimensional coordinate system of the millimeter-wave radar to obtain the three-dimensional coordinates of each home appliance in the three-dimensional coordinate system of the millimeter-wave radar.

[0087] Specifically, the spatial mapping relationship between the two-dimensional coordinate system of the camera device and the three-dimensional coordinate system of the millimeter-wave radar is pre-calibrated. This spatial mapping relationship can be, for example, a camera-radar extrinsic parameter matrix. The positions of each household appliance identified in the image can be converted into three-dimensional coordinates in the three-dimensional coordinate system of the millimeter-wave radar using this camera-radar extrinsic parameter matrix.

[0088] The camera-radar extrinsic parameter matrix refers to the spatial geometric relationship between the camera and the millimeter-wave radar, specifically: their relative positions in actual physical installation (e.g., the camera is 20 cm higher than the radar and 15 cm to the left), and the difference in their orientation angles (e.g., the angle between the camera's front and the radar's front is 10 degrees). This information, combined, constitutes the "extrinsic parameters," which can be understood as how the "space seen by the camera" corresponds to the "space perceived by the radar."

[0089] The extrinsic parameters include two parts: 1. Relative position: the straight-line distance between the camera and the radar, including offsets in the front-back, left-right, and vertical directions; 2. Relative angle: the orientation difference between the camera and the radar in three-dimensional space, including pitch, yaw, and rotation angles. Once the installation position relationship between the two is known, the data coordinate systems of the two devices can be mathematically "aligned" based on these extrinsic parameters. This process is called "coordinate transformation" or "coordinate alignment." First, the relative position and angle between the camera and the radar are measured in advance through a calibration program; then, when the camera identifies the position of a household object (based on the camera), the spatial position of this object in the radar's view can be calculated; in this way, the data sensed by the two devices are in the same unified space, allowing for fusion analysis. For example, the camera identifies "the sofa is 1.5 meters in front and 30 degrees to the right"; after conversion based on the extrinsic parameters, the radar can also know the direction and distance of the "sofa" in its own view; subsequently, the trajectory of a person can be accurately correlated with spatial areas such as the sofa and air conditioner. This external parameter conversion is a key step in realizing the fusion of image recognition information and radar data, and it is also the foundation for establishing an indoor spatial structure model.

[0090] Step S143: The converted three-dimensional coordinates of each home appliance in the three-dimensional coordinate system of the millimeter-wave radar are fused with the three-dimensional coordinates of the space output by the millimeter-wave radar to construct a spatial structure model of the space.

[0091] The spatial structure model may specifically include the three-dimensional coordinates and boundary contours of various home appliances (home appliances, furniture), walls, doors, and windows. In one specific implementation, the extended Kalman filter (EKF) algorithm is used to fuse the transformed three-dimensional coordinates of each home appliance in the three-dimensional coordinate system of the millimeter-wave radar with the three-dimensional coordinates of the space output by the millimeter-wave radar to construct a spatial structure model of the space.

[0092] Specifically, the converted three-dimensional coordinates of each home appliance in the three-dimensional coordinate system of the millimeter-wave radar are combined with the three-dimensional coordinates of the area containing the human body output by the millimeter-wave radar, and the extended Kalman filter algorithm is used to determine which is more reliable.

[0093] Extended Kalman Filter (EKF) is an algorithm that can fuse sensor data from different sources, and is particularly suitable for handling location information containing uncertainty. In this invention, EKF is used to fuse the location of home appliances obtained from image recognition and depth estimation (which is an estimate and contains errors) with the actual user location provided by millimeter-wave radar to construct a more accurate and stable spatial structure model. The specific fusion process is as follows:

[0094] 1. Define the state: Treat the position of each home appliance in three-dimensional space as a dynamic state variable, and continuously optimize its accuracy.

[0095] 2. Image recognition results as initial values: Initially, image recognition combined with depth estimation provides an estimated 3D position for each piece of furniture. These data are estimates, and therefore contain some systematic error.

[0096] 3. Millimeter-wave radar trajectory as a correction reference: As the user moves indoors, the millimeter-wave radar continuously outputs their position. Based on the fact that "the user is near the furniture when using it," the actual range of the furniture can be deduced.

[0097] For example, the system initially estimates that "the sofa is at position X", but if the user repeatedly sits or stays near X, the system will consider this estimate to be reliable; if the user frequently appears in places that are inconsistent with the image estimate, the system will dynamically adjust the estimated position of the sofa based on radar feedback.

[0098] 4. Extended Kalman Filter (EKF) at Work: With each new radar input, the EKF predicts and updates the positions of the furniture, compares the prediction with the actual measurement, and automatically adjusts the furniture coordinates to gradually approach the true values. Simultaneously, it controls error convergence, ensuring the consistency of the entire spatial structure model. Image recognition provides the initial model structure; radar trajectories provide feedback for dynamic correction; the EKF, as an information fusion unit, combines the two to continuously optimize the spatial model, making the spatial position of each piece of furniture closer to the actual arrangement.

[0099] The advantages of this mechanism are: it can handle the uncertainty of data from different sources; it optimizes the model accuracy over time; and it combines the advantages of "clear structure" in images with "dynamic accuracy" in radar.

[0100] Step S150: Combining the constructed spatial structure model and the user activity trajectory detected by the millimeter-wave radar, control the operation of the air conditioner.

[0101] Specifically, the user activity trajectory detected by the millimeter-wave radar is superimposed onto the spatial structure model to determine the user's location; the operation of the air conditioner is controlled based on the determined user location.

[0102] Figure 4 The diagram illustrates a specific implementation of the steps for controlling the operation of the air conditioner by combining the constructed spatial structure model and the user activity trajectory detected by the millimeter-wave radar.

[0103] like Figure 4 As shown, in a preferred embodiment, step S150 includes: step S151, step S152 and step S153.

[0104] Step S151: Divide the spatial structure model into two or more regions using a spatial partitioning algorithm.

[0105] The two or more areas have different functions. Specifically, the Voronoi spatial partitioning algorithm is applied to the spatial structure model to divide the room into several "influence areas" using furniture as reference points, and these areas are labeled. Each area is assigned a functional area label; for example, the area near the sofa is automatically labeled as the "sofa area," and the dining table and its surrounding area are labeled as the "dining area." The system executes the Voronoi spatial partitioning algorithm to automatically generate labels such as "sofa area," "dining table area," and "children's activity area" based on the relative positions and distances between furniture. Each area is associated with one or more home appliances and their spatial boundary information.

[0106] Specifically, using the position of each piece of furniture in the spatial structure model as the center point, the entire room is divided into several zones, each zone belonging to the piece of furniture closest to it. The specific execution process is as follows:

[0107] 1. Use the center point of the furniture as input:

[0108] Through image recognition, depth estimation, coordinate transformation, and filtering fusion, the 3D coordinates of all furniture in the room have been obtained. The horizontal position of the furniture on the ground plane (ignoring height) is then used as the seed point for Voronoi partitioning.

[0109] 2. Construct a two-dimensional room projection diagram:

[0110] The three-dimensional spatial structure model is compressed to a horizontal plane to form an interior layout diagram similar to "viewing from above," which is used for division operations.

[0111] 3. Perform Voronoi partitioning:

[0112] Using the standard Voronoi algorithm: Draw the center point of all furniture in the interior layout diagram; calculate the distance between any point in the diagram and each piece of furniture, and determine the furniture closest to that point, i.e., "which piece of furniture is closest"; based on this, assign each point to the "influence area" of the closest furniture.

[0113] In this way, all region boundaries will automatically form a set of non-overlapping polygons, called Voronoi units.

[0114] For example: the sofa is at point A, and the dining table is at point B; the perpendicular bisector of the midpoint between A and B is their boundary; the room is naturally divided into several areas such as the "sofa area", "dining table area", and "bed area".

[0115] 4. Output labeled spatial regions:

[0116] Each area is assigned a corresponding furniture name (such as "Sofa Area"), which serves as the basic spatial unit for subsequent "User Activity Attribution", "Hot Zone Generation", and "Scene Triggering".

[0117] This method of partitioning eliminates the need for manual area drawing and does not rely on CAD drawings of the building. Instead, it is entirely based on the furniture positions and is automatically generated, making it a key technology for achieving "spatial self-awareness." Its advantages include: fully automatic generation, adapting to any interior layout; each area directly corresponding to a specific piece of furniture, facilitating subsequent behavior judgment; and dynamic updates, automatically reconstructing the partitioning results when furniture positions change.

[0118] Step S152: Superimpose the user activity trajectory detected by the millimeter-wave radar onto the spatial structure model to determine whether the user has entered any of the two or more regions.

[0119] Specifically, the system continuously receives the user's human skeleton trajectory output by millimeter-wave radar, and superimposes the user's activity trajectory output by the millimeter-wave radar onto a two-dimensional plane parallel to the ground in the spatial structure model to determine whether the user has entered any of the two or more regions.

[0120] The user activity trajectory output by millimeter-wave radar is compressed and superimposed onto the horizontal plane (i.e., the ground plane) of the indoor spatial structure model during processing, rather than being directly superimposed onto the complete three-dimensional spatial structure model. In other words, the height information (i.e., the y-axis direction) from the three-dimensional coordinate data detected by the millimeter-wave radar is discarded, retaining only the horizontal position coordinates (x-axis and z-axis). The processed user activity trajectory data is then projected onto a two-dimensional plane parallel to the ground. The technical consideration behind this processing method is that, for the air conditioning control application this invention focuses on, the user's horizontal position is the primary basis for determining whether their activity area overlaps with functional areas, while the user's standing, sitting, or vertical position information at a certain height in space has minimal impact on control decisions. Simplifying the user activity trajectory to two dimensions not only makes data processing more efficient but also facilitates rapid matching with defined functional areas (such as sofa areas and dining areas). In contrast, if the user trajectory were fully preserved in three-dimensional space, richer behavioral analysis could be achieved, such as distinguishing whether the user is at a high position, lying down, or moving vertically. However, this would significantly increase the complexity of data processing and is not substantially necessary in typical applications such as air conditioning and energy-saving control. Therefore, this invention prioritizes the use of a trajectory planarization method.

[0121] Step S153: If it is determined that the user has entered any of the two or more areas and the triggering condition of the preset scene mode corresponding to that area is met, then the air conditioner is controlled to execute the preset scene mode corresponding to that area.

[0122] In one specific implementation, the triggering condition for a preset scene mode corresponding to any one of the two or more areas may specifically include: the user entering the area and receiving a voice command to activate the preset scene mode corresponding to that area. Different areas with different functions within the two or more areas correspond to different scene modes. For example, the sofa area corresponds to a movie / TV show mode.

[0123] Specifically, when a user enters any area, the system receives a voice command from the user. When the system receives a voice command from the user to activate a preset scene mode corresponding to that area, it controls the air conditioner to execute that preset scene mode. More specifically, the system receives and recognizes the user's voice command via a voice recognition module. After receiving the user's voice command, it identifies whether it is a voice command to activate a preset scene mode corresponding to that area. If so, it controls the air conditioner to execute that preset scene mode. For example, a user can use a mobile app, smart speaker, or voice remote control to say natural language commands such as "turn on movie mode" or "I want to watch a movie." The voice recognition module converts these commands into structured commands and matches them with existing scene configurations.

[0124] In the "Scene Triggering Rule Engine" of this invention, "Entering a Region + Voice 'Movie Mode'" is a typical composite triggering condition, specifically including two independent sources of input information: 1. "Entering a Region" refers to the user entering a certain region, which is determined based on whether the user's current spatial coordinates fall within the boundary of that functional area; 2. "Voice 'Movie Mode'" refers to the control command actively issued by the user via voice. Only when both conditions are met simultaneously, i.e., the user is in a specific region and issues a voice command matching the preset scene mode corresponding to that region, does the system recognize that the "triggering condition" of the preset scene mode corresponding to that region is met, and thus controls the air conditioner according to the corresponding air conditioning control parameters (such as temperature 23℃, medium fan speed, horizontal sweep, etc.). This composite triggering method of "spatial location + active voice" not only ensures the accuracy of the user's intention but also reduces the possibility of false triggering, which is one of the important mechanisms of this invention to improve control accuracy and interaction naturalness.

[0125] Different preset scene modes correspond to different air conditioning control parameters. Controlling the air conditioner to execute the preset scene mode corresponding to that area means controlling the air conditioner to execute the air conditioning control parameters corresponding to the preset scene mode corresponding to that area.

[0126] In one specific implementation, different air conditioning control parameters corresponding to different scene modes can be set by the user. Specifically, parameter templates for different air conditioning control parameters corresponding to different preset scene modes can be preset. These parameter templates pre-set default air conditioning control parameters for the corresponding preset scene modes. The user can modify the default air conditioning control parameters in the parameter templates to generate the air conditioning control parameters corresponding to the preset scene mode required by the user.

[0127] For example, the air conditioning system has several built-in preset scene modes and corresponding air conditioning control parameter combinations. Taking "Movie Mode" as an example, its preset parameter template is: Temperature: 23℃, Fan Speed: Medium, Airflow Direction: Horizontal Swing, Mode: Cooling. When the user uses it for the first time or has not modified their preferences, the system defaults to using the air conditioning control parameters in this template as the control command output. If the user wants to modify them, they can modify the values ​​of the control parameters in this parameter template.

[0128] In another specific implementation, different air conditioning control parameters corresponding to different preset scene modes are set based on user operation data of the air conditioner. Specifically, initial air conditioning control parameters corresponding to different preset scene modes are preset; based on user adjustments to the air conditioning control parameters during the operation of the corresponding preset scene mode, the air conditioning control parameters corresponding to that preset scene mode are modified. That is, in different preset scene modes, the air conditioner is controlled according to the initial air conditioning control parameters corresponding to the different preset scene modes. When running in any preset scene mode, if any control parameter is adjusted, the value of that control parameter in that preset scene mode is modified to the adjusted value. For example, if the user has manually adjusted the control parameters in a certain scene (e.g., set the temperature to 22℃), the system will record this setting and prioritize its use the next time the same scene is triggered. This "learning mechanism" allows the control parameters to gradually conform to user habits.

[0129] Optionally, the same preset scene mode can have different air conditioning control parameters for different time periods. For example, if the "movie mode" is triggered at night, the temperature will be set to a comfortable 23°C; during the day when the outdoor temperature is higher, the temperature will automatically drop to 22°C under the same scene.

[0130] Optionally, after controlling the air conditioner to execute the preset scene mode corresponding to the area, it detects whether the user has left the area; if the user is detected to have left the area and the departure time exceeds a first preset time, the air conditioner is controlled to switch to a preset energy-saving mode. If the user is detected to have re-entered the area, the preset scene mode corresponding to the area continues to be executed.

[0131] For example, the system continuously monitors whether the user is active. If the user leaves the area for 10 minutes, it automatically determines that "no one is present"; the system switches the air conditioner to energy-saving mode (such as increasing the set temperature and / or reducing the indoor fan speed) to reduce power consumption during idling. When the user re-enters the area, the system quickly restores the previous settings, achieving seamless control that "restores settings as soon as someone enters".

[0132] Figure 5 This is a schematic diagram of another embodiment of the air conditioner control method provided by the present invention.

[0133] like Figure 5 As shown, according to another embodiment of the present invention, the air conditioner control method further includes steps S160, S170 and S180.

[0134] Step S160: Obtain the user's historical activity data in the two or more regions.

[0135] The historical activity data includes: the time information and control information of the user's appearance in each of the two or more areas each day. The control information refers to whether the user issues a voice command to activate the preset scene mode corresponding to the currently located area.

[0136] For example, during long-term operation, continuous recording of users' daily behaviors mainly includes the following information: 1. Time information: the specific time the user appeared in a certain area (e.g., 18:30); 2. Spatial information: which area the user was in (e.g., sofa area, dining area); 3. Control intent information: whether the user issued a voice command and its content (e.g., "turn on movie mode"). Using "time period + spatial location + control command" as a joint tag, data records are continuously generated to build a user behavior database. For example, recording the user's activity trajectory and voice commands issued during each time period of the day forms a combination of "daily time period - activity area - command" to establish a behavior database.

[0137] Step S170: Based on the acquired historical activity data, predict the area the user will enter in the future and the control intent using a time-series prediction algorithm.

[0138] The control intent refers to whether the user will issue a voice command to activate the preset scene mode corresponding to that area. In one specific implementation, the acquired historical activity data can be modeled using an LSTM (Long Short-Term Memory) model or a time-series-based regression algorithm to predict the area the user is about to enter and their control intent. For example, predicting "when the user might enter the sofa area and activate the movie / TV mode." The historical activity data with time characteristics (a combination of "daily time period - activity area - command") serves as the model input, and the model output is the predicted time window for the user to next enter any area and trigger the corresponding preset scene mode.

[0139] Step S180: Based on the predicted area the user is about to enter and the control intention, control the air conditioner in advance.

[0140] For example, if it is predicted that "there is a greater than 90% probability that users will enter the sofa area and start the movie mode around 18:30", the following operation can be performed 5 minutes in advance: switch the air conditioner from energy-saving standby mode to "movie mode" parameters (temperature 23℃, medium fan speed, horizontal sweep); this achieves the proactive response effect that the air conditioner is already preset before the user enters the area, which can significantly improve the user experience, especially avoiding the lag of waiting for the device to adjust the temperature in hot or cold weather.

[0141] According to the above embodiments of the present invention, by long-term recording of the three elements of "time-location-command" and the predictive capability of the LSTM model, behavior pattern recognition and proactive response control are realized, enabling the air conditioning control system to evolve from "passively waiting for commands" to an intelligent state of "actively anticipating user needs".

[0142] Optionally, the historical activity trajectory of the user detected by the millimeter-wave radar is acquired; the historical activity trajectory of the user is superimposed on the spatial structure model to form a heat map, thereby determining the user active areas in the two or more regions; and the air conditioner is controlled according to the determined user active areas in the two or more regions. The region where the user occurrence frequency (times / hour) reaches a preset value is determined as the user active area.

[0143] Specifically, by continuously recording users' historical activity trajectories using millimeter-wave radar, these trajectories are overlaid onto the spatial structure model. The frequency and / or duration of user stays in each of the two or more areas are then statistically analyzed to create a spatial thermal layer. This allows for the derivation of usage intensity and habits in a specific area from "users' actual daily behavior," providing a supporting basis for personalized air conditioning control and behavior prediction.

[0144] Precise air conditioning control can be achieved by controlling the air conditioning to perform directional airflow, airflow avoidance of people, and / or differential airflow based on the identified user activity areas. Specifically, based on the thermal map, it can be determined which areas are the most frequently used by users. When controlling the air conditioning, the comfort of these areas can be prioritized. For example: directional airflow: concentrate airflow to the user's active areas; airflow avoidance of people: if someone is detected approaching the cold air outlet in a low-frequency area, the system can temporarily close the airflow direction to avoid direct airflow; differential temperature control: in a multi-outlet system, different air supply temperatures or air speeds can be set for each area according to the distribution of user activity areas.

[0145] The heat map in this invention can also be used for user behavior modeling and prediction. The heat map can be used as a "behavioral label" input to a prediction model (such as LSTM) to train "when and in which areas users are likely to appear"; thereby entering the corresponding control state in advance, such as: users often appear in the sofa area in the evening → start the "movie mode" air conditioning parameters in advance; users concentrate in the dining table area during breakfast time → temporarily accelerate air purification or ventilation.

[0146] The heat map in this invention can also be used for dynamic scene matching (rule triggering assistance): different areas can be dynamically prioritized according to their popularity weight. For example, if someone is currently in the "sofa area (popularity 80)" and the "desk area (popularity 30)", the system is more inclined to think that the user is in a "leisure scene" rather than a "work scene". This can avoid false triggers or help the system make a more appropriate scene judgment.

[0147] In this invention, the heat map is not only the core basis for air conditioning air supply decisions, but also an important semantic layer for behavior prediction and scene judgment, and is one of the key components for realizing personalized space control.

[0148] The air conditioning control of this invention can be completed via serial port protocol, infrared transmitter (general purpose), or Wi-Fi protocol (smart air conditioner), with high compatibility; all image and trajectory data processing can be completed locally, ensuring privacy is not leaked, and cloud-based model updates are also supported.

[0149] To clearly illustrate the technical solution of the present invention, the execution flow of the air conditioner control method provided by the present invention will be described below with reference to a specific embodiment.

[0150] Figure 6 This is a schematic diagram of a specific embodiment of the air conditioner control method provided by the present invention. Figure 6 As shown, the user takes a photo and uploads an indoor image, which is then processed for image recognition and home appliance detection. Millimeter-wave radar collects the user's location information in real time. The image recognition results are fused with the millimeter-wave radar detection data to reconstruct a three-dimensional spatial structure, automatically dividing and labeling functional areas. The user's trajectory is overlaid onto the spatial structure model to form activity hotspots. Scene rules are triggered via voice and location to generate air conditioning operation strategies and send corresponding commands. If the user leaves a functional area, the system enters energy-saving mode.

[0151] The present invention also provides a control device for an air conditioner. This control method can be implemented in a local gateway (e.g., an edge computing gateway) or a server. Specifically, the edge gateway can be a small smart host deployed in a user's home, such as an embedded box, a smart home center, a router built-in module, or a high-performance IoT controller (e.g., a Raspberry Pi, a Qualcomm IoT chip module, etc.).

[0152] Figure 7 This is a structural block diagram of an embodiment of the air conditioner control device provided by the present invention. Figure 7 As shown, the control device 100 includes: a first acquisition unit 110, an identification unit 120, a detection unit 130, a construction unit 140, and a control unit 150.

[0153] The first acquisition unit 110 is used to acquire a panoramic image of the space captured by a camera device; the panoramic image includes various home appliances in the space.

[0154] Specifically, the panoramic image includes all the home appliances in the space. In one specific embodiment, a panoramic image of the space taken by the user through the camera device is acquired. For example, when a user controls the air conditioner for the first time using an app installed on a mobile terminal, the user takes a panoramic image of the room through the mobile terminal. The image includes all the home appliances in the space, such as doors, windows, furniture, and appliances. The user can follow the guidance to take photos of the room in good daylight conditions, preferably from a diagonal downward view, covering major furniture such as the air conditioner, sofa, television, and dining table. After taking the photo, it is uploaded to a local gateway (e.g., a smart air conditioner with certain computing and communication capabilities) or a server through the mobile terminal.

[0155] The identification unit 120 is used to perform image recognition on the panoramic image acquired by the first acquisition unit to obtain the types of home appliances in the space and their positions in the panoramic image.

[0156] In one specific embodiment, the identification unit performs image recognition on the acquired panoramic image to obtain the types of home appliances in the space and their positions in the panoramic image. This may specifically include the following steps:

[0157] (1) Input the acquired panoramic image into a pre-trained target detection model to perform target detection, and obtain the types of home appliances in the space and their location detection boxes in the panoramic image.

[0158] Specifically, after acquiring a panoramic image of the space, the panoramic image can be input into a pre-trained target detection model for target detection to obtain the types (e.g., air conditioners, sofas, doors and windows) and location detection boxes of various home appliances in the space.

[0159] The object detection model can specifically be the YOLOv5 (You Only Look Once, version 5) object detection model. The YOLOv5 object detection model is pre-trained to detect objects in home furnishings (e.g., appliances, furniture, doors, windows, etc.). The output of the object detection model is the type and location bounding boxes (e.g., rectangles) for each home furnishing item. Using the YOLOv5 object detection algorithm for object detection offers fast processing speed, wide coverage, and rapid filtering of candidate target regions.

[0160] (2) Input the identified location detection boxes of each home appliance in the panoramic image into a pre-trained target segmentation model for edge segmentation to obtain the outline of each home appliance.

[0161] To improve recognition accuracy, the detected location bounding boxes of each home appliance in the identified space are further input into a pre-trained target segmentation model for edge segmentation, obtaining the precise contours of each home appliance in the panoramic image, making subsequent spatial mapping more accurate. Specifically, the target segmentation model can be a Mask R-CNN target segmentation model. That is, using the Mask R-CNN (Region Convolutional Neural Network) algorithm, pixel-level edge segmentation is performed on the detected location bounding boxes of each home appliance to obtain the precise contour of each appliance. For example, for each rectangular region detected by the YOLOv5 model, this rectangular region is fed into the Mask R-CNN model. The Mask R-CNN model further refines this defined area, determining which pixels belong to the target home appliance and which belong to the background, outputting a mask image, i.e., the precise contour boundary of each home appliance in the image.

[0162] This invention fuses the "target category and coarse location" provided by YOLOv5 with the "pixel-level boundary" provided by Mask R-CNN. YOLOv5 is responsible for "finding" the target, while Mask R-CNN is responsible for "depicting" it clearly, ultimately providing a more detailed geometric description of each target home appliance in the panoramic image, including its position, category, area, and shape. This combination of "detection box + segmentation mask" combines the advantages of both algorithms: YOLOv5 provides high speed and high recall, ensuring no missed detections; Mask R-CNN provides fine contours, which is beneficial for subsequent spatial localization and volume estimation. Using rectangular boxes as input regions also significantly reduces the computational burden of Mask R-CNN, improving segmentation efficiency. Therefore, this invention employs a two-stage recognition process of "YOLOv5 detection first, Mask R-CNN refinement later" to achieve both fast and accurate target recognition and contour extraction, providing high-quality input for subsequent 3D spatial reconstruction and functional area division.

[0163] The detection unit 130 is used to detect the user's location and human skeleton information in the space using millimeter-wave radar, and outputs the three-dimensional coordinates of the user's location and the user's activity trajectory.

[0164] Specifically, deploying millimeter-wave radar indoors and continuously outputting human point cloud information at preset intervals (e.g., 20 milliseconds) can accurately detect "whether there is someone in the room", "where the person is", and "how their movements are", without affecting privacy or depending on lighting conditions.

[0165] Millimeter-wave radar detects the signal characteristics reflected by objects (especially the human body) in space by transmitting and receiving high-frequency electromagnetic waves. This allows it to identify the human body's position in three-dimensional space, its posture and skeletal structure (i.e., skeletal information), and whether the body is moving, stationary, or performing a certain action (such as sitting or walking). Specifically, the radar senses the spatial location and energy distribution of "multiple reflection points" in each time frame, clusters these points using algorithms, and fits the human body shape.

[0166] The millimeter-wave radar outputs the following core data at preset intervals (e.g., 20ms):

[0167] The three-dimensional coordinates (x, y, z) of the human body's position (e.g., the center point): represent the current position of the human body relative to the radar coordinate system;

[0168] Key point coordinates of the human skeleton: such as the three-dimensional coordinates of the head, shoulders, knees, and soles of the feet (depending on radar capabilities);

[0169] Human motion trajectory: Records the changes in the center position of the human body over a continuous period of time, forming a trajectory sequence;

[0170] The core function of millimeter-wave radar is to provide the three-dimensional position and continuous movement trajectory of the human body in real space, providing crucial support for subsequent steps.

[0171] The construction unit 140 is used to fuse the positions of each home appliance in the identified space in the panoramic image with the three-dimensional coordinates of the space output by the millimeter-wave radar to construct a spatial structure model of the space.

[0172] In one specific implementation, the construction unit fuses the positions of each home appliance in the identified space in the panoramic image with the three-dimensional coordinates of the area containing the human body output by the millimeter-wave radar to obtain a spatial structure model of the space, which may specifically include the following steps:

[0173] (1) The estimated three-dimensional position of each home appliance in the space under the coordinate system of the camera device is obtained by using a depth estimation algorithm.

[0174] In one specific implementation, the panoramic image is input into a pre-trained depth estimation model to obtain the depth value of each pixel in the panoramic image; the depth value of each pixel is then correlated with the identified positions of each home appliance in the space within the panoramic image to obtain the estimated three-dimensional position of each home appliance in the coordinate system of the camera device.

[0175] The depth estimation model can specifically be a MiDaS monocular depth estimation model. That is, the MiDaS monocular depth estimation algorithm estimates the depth value of each pixel in the panoramic image, thereby obtaining the depth information of the positions of each home appliance in the panoramic image, realizing the reconstruction of the actual spatial distance of the home appliances from the planar image. Since the panoramic image of the space captured by the camera only contains two-dimensional image information and does not contain depth information, the MiDaS monocular image depth estimation algorithm is introduced. The MiDaS monocular image depth estimation algorithm can determine the distance of each pixel in the image from the camera device. MiDaS (Monocular Depth Sensing) is a monocular image-based depth estimation algorithm that can predict the relative depth map of a scene from a regular RGB image through deep learning. Inputting the original panoramic image into the MiDaS monocular image depth estimation model outputs a depth map of the same size as the image, with each pixel having a relative depth value. Combining this with the pixel position of each home appliance in the image, these depth values ​​are mapped to each identified home appliance to obtain an estimated three-dimensional position of each home appliance from the camera device's perspective. The original panoramic image is input into the MiDaS monocular depth estimation model to obtain the depth value of each pixel and output a depth map. The larger the value, the farther the pixel is from the camera device. The depth value of each pixel is matched with the position (position detection box) of each identified home appliance in the panoramic image. The depth of the corresponding region is extracted according to the position of the detection box in the panoramic image to obtain the average / center depth value of each home appliance and the estimated three-dimensional position of each home appliance in the coordinate system of the camera device.

[0176] For example, the bounding box of the "sofa" in the image (i.e., the location detection box output by the location object detection model, such as the rectangle output by the YOLOv5 model) is located at (x1, y1, x2, y2), and the average depth of the corresponding area in the depth map is 2 meters. Based on this, it can be determined that the sofa is 2 meters in front of the camera, at a 30-degree angle to the right.

[0177] The original image coordinate system of a camera device is a two-dimensional pixel coordinate system (i.e., the x / y axis of the image). However, when the image incorporates depth estimation results, each pixel is no longer just "positional," but has a "depth value." That is, each pixel can be represented as: (x... 像素 ,y 像素 ,d 深度At this point, the image coordinates plus the depth value constitute a three-dimensional estimated coordinate system under the camera coordinate system, which is related to the actual position of the object in real space relative to the camera. Although this coordinate system is still "estimated," it can already reflect the spatial layout relationship of the furniture relative to the camera. For example: - The sofa is located in the lower right corner of the image, with a depth of 1.5 meters; - The air conditioner is located in the upper left corner of the image, with a depth of 3.2 meters; their "relative spatial positions" under the camera's view are clear. The obtained estimated three-dimensional position of the furniture refers to its estimated spatial position in the coordinate system of the camera device, no longer a pure two-dimensional pixel coordinate. Through the restoration process of "pixel coordinates + depth value," the two-dimensional image information has actually been "upgraded" to a three-dimensional coordinate space. This three-dimensional coordinate is still under the "camera coordinate system," that is, with the camera as the origin (0,0,0), the front as the Z-axis, the left and right as the X-axis, and the up and down as the Y-axis, and the unit can be meters or centimeters (according to the depth model). Therefore, the final position of each piece of furniture is an "estimated three-dimensional position," and its coordinate form is (X,Y,Z), which is used for subsequent transformation to the radar coordinate system or spatial model.

[0178] The MiDaS monocular depth estimation algorithm can obtain roughly reliable object distance relationships from ordinary photographs without the need for lasers, binocular cameras, or dedicated sensors, providing crucial support for automatic indoor space modeling. This algorithm boasts high execution efficiency, making it suitable for deployment in edge computing gateways and meeting the real-time and resource consumption requirements of smart home systems. Therefore, the technical solution proposed in this invention is fully compatible with ordinary consumer-grade cameras or smartphones in the image acquisition stage, requiring no replacement or configuration of special equipment, and possesses significant practicality and promotional value.

[0179] (2) The estimated three-dimensional position of each home appliance in the coordinate system of the camera device is obtained and converted into three-dimensional coordinates in the three-dimensional coordinate system of the millimeter-wave radar.

[0180] In one specific embodiment, based on the pre-calibrated spatial mapping relationship between the coordinate system of the camera device and the three-dimensional coordinate system of the millimeter-wave radar, the estimated three-dimensional position of each home appliance in the coordinate system of the camera device is projected into the three-dimensional coordinate system of the millimeter-wave radar to obtain the three-dimensional coordinates of each home appliance in the three-dimensional coordinate system of the millimeter-wave radar.

[0181] Specifically, the spatial mapping relationship between the two-dimensional coordinate system of the camera device and the three-dimensional coordinate system of the millimeter-wave radar is pre-calibrated. The spatial mapping management, for example, is a camera-radar extrinsic parameter matrix; the positions (outlines) of each home appliance identified in the image can be converted into three-dimensional coordinates in the three-dimensional coordinate system of the millimeter-wave radar through this camera-radar extrinsic parameter matrix.

[0182] The camera-radar extrinsic parameter matrix refers to the spatial geometric relationship between the camera and the millimeter-wave radar, specifically: their relative positions in actual physical installation (e.g., the camera is 20 cm higher than the radar and 15 cm to the left), and the difference in their orientation angles (e.g., the angle between the camera's front and the radar's front is 10 degrees). This information, combined, constitutes the "extrinsic parameters," which can be understood as how the "space seen by the camera" corresponds to the "space perceived by the radar."

[0183] The extrinsic parameters include two parts: 1. Relative position: the straight-line distance between the camera and the radar, including offsets in the front-back, left-right, and vertical directions; 2. Relative angle: the orientation difference between the camera and the radar in three-dimensional space, including pitch, yaw, and rotation angles. Once the installation position relationship between the two is known, the data coordinate systems of the two devices can be mathematically "aligned" based on these extrinsic parameters. This process is called "coordinate transformation" or "coordinate alignment." First, the relative position and angle between the camera and the radar are measured in advance through a calibration program; then, when the camera identifies the position of a household object (based on the camera), the spatial position of this object in the radar's view can be calculated; in this way, the data sensed by the two devices are in the same unified space, allowing for fusion analysis. For example, the camera identifies "the sofa is 1.5 meters in front and 30 degrees to the right"; after conversion based on the extrinsic parameters, the radar can also know the direction and distance of the "sofa" in its own view; subsequently, the trajectory of a person can be accurately correlated with spatial areas such as the sofa and air conditioner. This external parameter conversion is a key step in realizing the fusion of image recognition information and radar data, and it is also the foundation for establishing an indoor spatial structure model.

[0184] (3) The three-dimensional coordinates of each home appliance in the three-dimensional coordinate system of the millimeter-wave radar are fused with the three-dimensional coordinates of the space output by the millimeter-wave radar to construct a spatial structure model of the space.

[0185] The spatial structure model may specifically include the three-dimensional coordinates and boundary contours of various home appliances (home appliances, furniture), walls, doors, and windows. In one specific implementation, the extended Kalman filter (EKF) algorithm is used to fuse the transformed three-dimensional coordinates of each home appliance in the three-dimensional coordinate system of the millimeter-wave radar with the three-dimensional coordinates of the space output by the millimeter-wave radar to construct a spatial structure model of the space.

[0186] Specifically, the converted three-dimensional coordinates of each home appliance in the three-dimensional coordinate system of the millimeter-wave radar are combined with the three-dimensional coordinates of the area containing the human body output by the millimeter-wave radar, and the extended Kalman filter algorithm is used to determine which is more reliable.

[0187] Extended Kalman Filter (EKF) is an algorithm that can fuse sensor data from different sources, and is particularly suitable for handling location information containing uncertainty. In this invention, EKF is used to fuse the location of home appliances obtained from image recognition and depth estimation (which is an estimate and contains errors) with the actual user location provided by millimeter-wave radar to construct a more accurate and stable spatial structure model. The specific fusion process is as follows:

[0188] 1. Define the state: Treat the position of each home appliance in three-dimensional space as a dynamic state variable, and continuously optimize its accuracy.

[0189] 2. Image recognition results as initial values: Initially, image recognition combined with depth estimation provides an estimated 3D position for each piece of furniture. These data are estimates, and therefore contain some systematic error.

[0190] 3. Millimeter-wave radar trajectory as a correction reference: As the user moves indoors, the millimeter-wave radar continuously outputs their position. Based on the fact that "the user is near the furniture when using it," the actual range of the furniture can be deduced.

[0191] For example, the system initially estimates that "the sofa is at position X", but if the user repeatedly sits or stays near X, the system will consider this estimate to be reliable; if the user frequently appears in places that are inconsistent with the image estimate, the system will dynamically adjust the estimated position of the sofa based on radar feedback.

[0192] 4. Extended Kalman Filter (EKF) at Work: With each new radar input, the EKF predicts and updates the positions of the furniture, compares the prediction with the actual measurement, and automatically adjusts the furniture coordinates to gradually approach the true values. Simultaneously, it controls error convergence, ensuring the consistency of the entire spatial structure model. Image recognition provides the initial model structure; radar trajectories provide feedback for dynamic correction; the EKF, as an information fusion unit, combines the two to continuously optimize the spatial model, making the spatial position of each piece of furniture closer to the actual arrangement.

[0193] The advantages of this mechanism are: it can handle the uncertainty of data from different sources; it optimizes the model accuracy over time; and it combines the advantages of "clear structure" in images with "dynamic accuracy" in radar.

[0194] The control unit 150 controls the operation of the air conditioner by combining the spatial structure model constructed by the building unit and the user activity trajectory detected by the millimeter-wave radar by the detection unit.

[0195] Specifically, the user activity trajectory detected by the millimeter-wave radar is superimposed onto the spatial structure model to determine the user's location; the operation of the air conditioner is controlled based on the determined user location.

[0196] In a preferred embodiment, controlling the operation of the air conditioner by combining the spatial structure model constructed by the building unit and the user activity trajectory detected by the millimeter-wave radar by the detection unit may specifically include the following steps:

[0197] (1) Divide the spatial structure model into two or more regions using a spatial partitioning algorithm.

[0198] The two or more areas have different functions. Specifically, the Voronoi spatial partitioning algorithm is applied to the spatial structure model to divide the room into several "influence areas" using furniture as reference points, and these areas are labeled. Each area is assigned a functional area label; for example, the area near the sofa is automatically labeled as the "sofa area," and the dining table and its surrounding area are labeled as the "dining area." The system executes the Voronoi spatial partitioning algorithm to automatically generate labels such as "sofa area," "dining table area," and "children's activity area" based on the relative positions and distances between furniture. Each area is associated with one or more home appliances and their spatial boundary information.

[0199] Specifically, using the position of each piece of furniture in the spatial structure model as the center point, the entire room is divided into several zones, each zone belonging to the piece of furniture closest to it. The specific execution process is as follows:

[0200] 1. Use the center point of the furniture as input:

[0201] Through image recognition, depth estimation, coordinate transformation, and filtering fusion, the 3D coordinates of all furniture in the room have been obtained. The horizontal position of the furniture on the ground plane (ignoring height) is then used as the seed point for Voronoi partitioning.

[0202] 2. Construct a two-dimensional room projection diagram:

[0203] The three-dimensional spatial structure model is compressed to a horizontal plane to form an interior layout diagram similar to "viewing from above," which is used for division operations.

[0204] 3. Perform Voronoi partitioning:

[0205] Using the standard Voronoi algorithm: Draw the center point of all furniture in the interior layout diagram; calculate the distance between any point in the diagram and each piece of furniture, and determine the furniture closest to that point, i.e., "which piece of furniture is closest"; based on this, assign each point to the "influence area" of the closest furniture.

[0206] In this way, all region boundaries will automatically form a set of non-overlapping polygons, called Voronoi units.

[0207] For example: the sofa is at point A, and the dining table is at point B; the perpendicular bisector of the midpoint between A and B is their boundary; the room is naturally divided into several areas such as the "sofa area", "dining table area", and "bed area".

[0208] 4. Output labeled spatial regions:

[0209] Each area is assigned a corresponding furniture name (such as "Sofa Area"), which serves as the basic spatial unit for subsequent "User Activity Attribution", "Hot Zone Generation", and "Scene Triggering".

[0210] This method of partitioning eliminates the need for manual area drawing and does not rely on CAD drawings of the building. Instead, it is entirely based on the furniture positions and is automatically generated, making it a key technology for achieving "spatial self-awareness." Its advantages include: fully automatic generation, adapting to any interior layout; each area directly corresponding to a specific piece of furniture, facilitating subsequent behavior judgment; and dynamic updates, automatically reconstructing the partitioning results when furniture positions change.

[0211] (2) The user activity trajectory detected by the millimeter-wave radar is superimposed on the spatial structure model to determine whether the user has entered any of the two or more regions.

[0212] Specifically, the system continuously receives the user's human skeleton trajectory output by millimeter-wave radar, and superimposes the user's activity trajectory output by the millimeter-wave radar onto a two-dimensional plane parallel to the ground in the spatial structure model to determine whether the user has entered any of the two or more regions.

[0213] The user activity trajectory output by millimeter-wave radar is compressed and superimposed onto the horizontal plane (i.e., the ground plane) of the indoor spatial structure model during processing, rather than being directly superimposed onto the complete three-dimensional spatial structure model. In other words, the height information (i.e., the y-axis direction) from the three-dimensional coordinate data detected by the millimeter-wave radar is discarded, retaining only the horizontal position coordinates (x-axis and z-axis). The processed user activity trajectory data is then projected onto a two-dimensional plane parallel to the ground. The technical consideration behind this processing method is that, for the air conditioning control application this invention focuses on, the user's horizontal position is the primary basis for determining whether their activity area overlaps with functional areas, while the user's standing, sitting, or vertical position information at a certain height in space has minimal impact on control decisions. Simplifying the user activity trajectory to two dimensions not only makes data processing more efficient but also facilitates rapid matching with defined functional areas (such as sofa areas and dining areas). In contrast, if the user trajectory were fully preserved in three-dimensional space, richer behavioral analysis could be achieved, such as distinguishing whether the user is at a high position, lying down, or moving vertically. However, this would significantly increase the complexity of data processing and is not substantially necessary in typical applications such as air conditioning and energy-saving control. Therefore, this invention prioritizes the use of a trajectory planarization method.

[0214] (3) If it is determined that the user enters any of the two or more areas and the triggering condition of the preset scene mode corresponding to that area is met, then the air conditioner is controlled to execute the preset scene mode corresponding to that area.

[0215] In one specific implementation, the triggering condition for a preset scene mode corresponding to any one of the two or more areas may specifically include: the user entering the area and receiving a voice command to activate the preset scene mode corresponding to that area. Different areas with different functions within the two or more areas correspond to different scene modes. For example, the sofa area corresponds to a movie / TV show mode.

[0216] Specifically, when a user enters any area, the system receives a voice command from the user. When the system receives a voice command from the user to activate a preset scene mode corresponding to that area, it controls the air conditioner to execute that preset scene mode. More specifically, the system receives and recognizes the user's voice command via a voice recognition module. After receiving the user's voice command, it identifies whether it is a voice command to activate a preset scene mode corresponding to that area. If so, it controls the air conditioner to execute that preset scene mode. For example, a user can use a mobile app, smart speaker, or voice remote control to say natural language commands such as "turn on movie mode" or "I want to watch a movie." The voice recognition module converts these commands into structured commands and matches them with existing scene configurations.

[0217] In the "Scene Triggering Rule Engine" of this invention, "Entering a Region + Voice 'Movie Mode'" is a typical composite triggering condition, specifically including two independent sources of input information: 1. "Entering a Region" refers to the user entering a certain region, which is determined based on whether the user's current spatial coordinates fall within the boundary of that functional area; 2. "Voice 'Movie Mode'" refers to the control command actively issued by the user via voice. Only when both conditions are met simultaneously, i.e., the user is in a specific region and issues a voice command matching the preset scene mode corresponding to that region, does the system recognize that the "triggering condition" of the preset scene mode corresponding to that region is met, and thus controls the air conditioner according to the corresponding air conditioning control parameters (such as temperature 23℃, medium fan speed, horizontal sweep, etc.). This composite triggering method of "spatial location + active voice" not only ensures the accuracy of the user's intention but also reduces the possibility of false triggering, which is one of the important mechanisms of this invention to improve control accuracy and interaction naturalness.

[0218] Different preset scene modes correspond to different air conditioning control parameters. Controlling the air conditioner to execute the preset scene mode corresponding to that area means controlling the air conditioner to execute the air conditioning control parameters corresponding to the preset scene mode corresponding to that area.

[0219] In one specific implementation, different air conditioning control parameters corresponding to different scene modes can be set by the user. Specifically, parameter templates for different air conditioning control parameters corresponding to different preset scene modes can be preset. These parameter templates pre-set default air conditioning control parameters for the corresponding preset scene modes. The user can modify the default air conditioning control parameters in the parameter templates to generate the air conditioning control parameters corresponding to the preset scene mode required by the user.

[0220] For example, the air conditioning system has several built-in preset scene modes and corresponding air conditioning control parameter combinations. Taking "Movie Mode" as an example, its preset parameter template is: Temperature: 23℃, Fan Speed: Medium, Airflow Direction: Horizontal Swing, Mode: Cooling. When the user uses it for the first time or has not modified their preferences, the system defaults to using the air conditioning control parameters in this template as the control command output. If the user wants to modify them, they can modify the values ​​of the control parameters in this parameter template.

[0221] In another specific implementation, different air conditioning control parameters corresponding to different preset scene modes are set based on user operation data of the air conditioner. Specifically, initial air conditioning control parameters corresponding to different preset scene modes are preset; based on user adjustments to the air conditioning control parameters during the operation of the corresponding preset scene mode, the air conditioning control parameters corresponding to that preset scene mode are modified. That is, in different preset scene modes, the air conditioner is controlled according to the initial air conditioning control parameters corresponding to the different preset scene modes. When running in any preset scene mode, if any control parameter is adjusted, the value of that control parameter in that preset scene mode is modified to the adjusted value. For example, if the user has manually adjusted the control parameters in a certain scene (e.g., set the temperature to 22℃), the system will record this setting and prioritize its use the next time the same scene is triggered. This "learning mechanism" allows the control parameters to gradually conform to user habits.

[0222] Optionally, it also includes different air conditioning control parameters for the same preset scene mode at different times. For example, if the "movie mode" is triggered at night, the temperature is set to a comfortable 23°C; during the day when the outdoor temperature is high, the temperature automatically drops to 22°C under the same scene.

[0223] Optionally, the control unit is further configured to: after controlling the air conditioner to execute the preset scene mode corresponding to the area, detect whether the user has left the area; if the user is detected to have left the area and the departure time exceeds a first preset time, control the air conditioner to switch to a preset energy-saving mode.

[0224] For example, the system continuously monitors whether the user is active. If the user leaves the area for 10 minutes, it automatically determines that "no one is present"; the system switches the air conditioner to energy-saving mode (such as increasing the set temperature and / or reducing the indoor fan speed) to reduce power consumption during idling. When the user re-enters the area, the system quickly restores the previous settings, achieving seamless control that "restores settings as soon as someone enters".

[0225] Figure 8 This is a structural block diagram of an embodiment of the air conditioner control device provided by the present invention. Figure 8 As shown, the control device 100 further includes: a second acquisition unit 160 and a prediction unit 170.

[0226] The second acquisition unit 160 is used to acquire historical activity data of the user in the two or more areas.

[0227] The historical activity data includes: the time information and control information of the user's appearance in each of the two or more areas each day. The control information refers to whether the user issues a voice command to activate the preset scene mode corresponding to the currently located area.

[0228] For example, during long-term operation, continuous recording of users' daily behaviors mainly includes the following information: 1. Time information: the specific time the user appeared in a certain area (e.g., 18:30); 2. Spatial information: which area the user was in (e.g., sofa area, dining area); 3. Control intent information: whether the user issued a voice command and its content (e.g., "turn on movie mode"). Using "time period + spatial location + control command" as a joint tag, data records are continuously generated to build a user behavior database. For example, recording the user's activity trajectory and voice commands issued during each time period of the day forms a combination of "daily time period - activity area - command" to establish a behavior database.

[0229] The prediction unit 170 is used to predict the area that the user will enter in the future and the control intention based on the historical activity data obtained by the second acquisition unit through a time-series prediction algorithm.

[0230] The control intent refers to whether the user will issue a voice command to activate the preset scene mode corresponding to that area. In one specific implementation, the acquired historical activity data can be modeled using an LSTM (Long Short-Term Memory) model or a time-series-based regression algorithm to predict the area the user is about to enter and their control intent. For example, predicting "when the user might enter the sofa area and activate the movie / TV mode." The historical activity data with time characteristics (a combination of "daily time period - activity area - command") serves as the model input, and the model output is the predicted time window for the user to next enter any area and trigger the corresponding preset scene mode.

[0231] The control unit 150 is further configured to: control the air conditioner in advance based on the area the user is about to enter, as predicted by the prediction unit, and the user's control intention.

[0232] For example, if it is predicted that "there is a greater than 90% probability that users will enter the sofa area and start the movie mode around 18:30", the following operation can be performed 5 minutes in advance: switch the air conditioner from energy-saving standby mode to "movie mode" parameters (temperature 23℃, medium fan speed, horizontal sweep); this achieves the proactive response effect that the air conditioner is already preset before the user enters the area, which can significantly improve the user experience, especially avoiding the lag of waiting for the device to adjust the temperature in hot or cold weather.

[0233] According to the above embodiments of the present invention, by long-term recording of the three elements of "time-location-command" and the predictive capability of the LSTM model, behavior pattern recognition and proactive response control are realized, enabling the air conditioning control system to evolve from "passively waiting for commands" to an intelligent state of "actively anticipating user needs".

[0234] Optionally, the control device 100 further includes a third acquisition unit and a determination unit (not shown).

[0235] The third acquisition unit is used to acquire the user's historical activity trajectory detected by the millimeter-wave radar; the determination unit is used to overlay the user's historical activity trajectory onto the spatial structure model to form a heat map, so as to determine the user's active area in the two or more regions; the control unit is also used to control the air conditioner according to the determined user's active area in the two or more regions. The region where the user's occurrence frequency (times / hour) reaches a preset value is determined as the user's active area.

[0236] Specifically, by continuously recording users' historical activity trajectories using millimeter-wave radar, these trajectories are overlaid onto the spatial structure model. The frequency and / or duration of user stays in each of the two or more areas are then statistically analyzed to create a spatial thermal layer. This allows for the derivation of usage intensity and habits in a specific area from "users' actual daily behavior," providing a supporting basis for personalized air conditioning control and behavior prediction.

[0237] Precise air conditioning control can be achieved by controlling the air conditioning to perform directional airflow, airflow avoidance of people, and / or differential airflow based on the identified user activity areas. Specifically, based on the thermal map, it can be determined which areas are the most frequently used by users. When controlling the air conditioning, the comfort of these areas can be prioritized. For example: directional airflow: concentrate airflow to the user's active areas; airflow avoidance of people: if someone is detected approaching the cold air outlet in a low-frequency area, the system can temporarily close the airflow direction to avoid direct airflow; differential temperature control: in a multi-outlet system, different air supply temperatures or air speeds can be set for each area according to the distribution of user activity areas.

[0238] The heat map in this invention can also be used for user behavior modeling and prediction. The heat map can be used as a "behavioral label" input to a prediction model (such as LSTM) to train "when and in which areas users are likely to appear"; thereby entering the corresponding control state in advance, such as: users often appear in the sofa area in the evening → start the "movie mode" air conditioning parameters in advance; users concentrate in the dining table area during breakfast time → temporarily accelerate air purification or ventilation.

[0239] The heat map in this invention can also be used for dynamic scene matching (rule triggering assistance): different areas can be dynamically prioritized according to their popularity weight. For example, if someone is currently in the "sofa area (popularity 80)" and the "desk area (popularity 30)", the system is more inclined to think that the user is in a "leisure scene" rather than a "work scene". This can avoid false triggers or help the system make a more appropriate scene judgment.

[0240] In this invention, the heat map is not only the core basis for air conditioning air supply decisions, but also an important semantic layer for behavior prediction and scene judgment, and is one of the key components for realizing personalized space control.

[0241] The present invention also provides a storage medium corresponding to the control method of the air conditioner, wherein a computer program is stored thereon, and the program, when executed by a processor, implements the steps of any of the aforementioned methods.

[0242] The present invention also provides an air conditioner corresponding to the control method of the air conditioner, comprising a processor, a memory, and a computer program stored in the memory that can run on the processor, wherein the processor executes the program to implement the steps of any of the aforementioned methods.

[0243] The present invention also provides an air conditioner corresponding to the control device of the air conditioner, including the control device of any of the aforementioned air conditioners.

[0244] The present invention also provides a computer program product corresponding to the control method of the air conditioner, including a computer program that, when executed by a processor, implements the steps of any of the aforementioned methods.

[0245] This invention optimizes the following key aspects, solving the complete technical process of automatically identifying spatial structure, sensing user location in real time, and intelligently controlling air conditioning;

[0246] 1. Image recognition and millimeter-wave data fusion enables automatic indoor 3D spatial modeling and functional area labeling:

[0247] By using YOLOv5 and Mask R-CNN algorithms to identify key targets such as air conditioners, sofas, doors and windows in indoor photos, their positions and outlines in the images are obtained. At the same time, human trajectory data output by millimeter-wave radar is used to convert the two-dimensional position information in the image into real three-dimensional spatial coordinates through camera-radar coordinate transformation and extended Kalman filter algorithms. This enables the complete presentation of the indoor space structure and furniture distribution without user intervention, and automatically labels functional areas (such as "sofa area", "dining table area" etc.).

[0248] 2. By combining users' historical activity trajectories to construct dynamic activity heatmaps, the accuracy of behavior recognition and the personalization of control are improved:

[0249] The system records human skeleton data continuously output by millimeter-wave radar, combines it with spatial partitioning, and generates an "activity heat map" through overlay analysis. This map is used to depict the frequency and habits of users in different time periods and areas, and can dynamically reflect users' daily behavior patterns, making air conditioning control more in line with real-life scenarios, such as switching to sleep mode based on nighttime activity heat zones.

[0250] 3. Construct a voice + location combined triggering mechanism to achieve high-precision intelligent control:

[0251] A combined triggering mechanism of "area awareness + voice command" was designed, which not only supports user voice invocation scenarios (such as "start movie mode"), but also requires the user to be in a matching area (such as the sofa area) to trigger the operation, avoiding accidental operation and improving control accuracy. It can combine three types of conditions—time, location, and user command—to achieve highly personalized control logic.

[0252] 4. Intelligent matching of air conditioning control strategies and automatic generation mechanism of operating parameters:

[0253] After identifying the scene, the system automatically calculates the air conditioner's operating status (such as setting the temperature to 23℃, the fan speed to medium, and the airflow direction to horizontal sweep) based on user habits or preset parameter templates, without requiring manual intervention, achieving "what you see is what you get" intelligent operation.

[0254] 5. Intelligent detection of users leaving their seats and energy-saving standby strategies achieve significant energy consumption optimization:

[0255] The system continuously monitors whether a user has left the current functional area and automatically switches the air conditioner to energy-saving mode (e.g., increasing temperature and decreasing fan speed) after a period of inactivity (e.g., 10 minutes). Compared to constant temperature control or timed shutdown modes in related technologies, this invention achieves more efficient energy-saving control, reducing air conditioner energy consumption by 10%-15% in tests.

[0256] 6. Supports edge deployment and local data processing, enhancing user privacy and system stability:

[0257] All image analysis, data fusion, and control logic can be completed on the local gateway, eliminating the need to upload images to a cloud server and ensuring user privacy. The system is also designed with strong fault tolerance and scalability, making it easy to deploy and use in different home environments.

[0258] Accordingly, the solution provided by this invention addresses several technical shortcomings in current human-sensing-based air conditioning control systems by proposing an automatic modeling and control strategy that integrates image recognition and millimeter-wave data, specifically solving the following key technical problems:

[0259] 1. The relevant technology relies on users manually marking furniture positions, which is complex and prone to errors:

[0260] In the initial deployment phase, intelligent air conditioning control systems often require users to manually mark the locations of furniture such as sofas and air conditioners in an app for zone-based control. This method not only has a high learning curve but is also prone to mislabeling, leading to inaccurate system responses.

[0261] This invention utilizes YOLOv5 and Mask R-CNN image recognition technology, combined with millimeter-wave radar spatial perception capabilities, to achieve automatic modeling capabilities that enable "instant recognition upon power-on," completing spatial structure recognition and functional area division without user intervention.

[0262] 2. In related technologies, image and sensor data are not fused, making it impossible to reconstruct the true three-dimensional spatial structure:

[0263] Many air conditioning systems in related technologies only use image recognition and radar detection separately, lacking an effective coordinate alignment and spatial fusion mechanism. As a result, the system can only roughly determine whether "the user is near the air conditioner" and cannot determine the specific functional area where the user is located (such as whether he is sitting in the sofa area).

[0264] This invention achieves the unification of coordinate systems for data collected by different devices through radar-camera calibration and extended Kalman filtering algorithm, truly realizing "image + radar" dual-mode fusion modeling and accurately constructing spatial maps.

[0265] 3. Air conditioning control in related technologies relies on fixed timing or single-dimensional sensing, lacking a multi-condition intelligent triggering mechanism:

[0266] The relevant technologies mainly rely on simple rules such as "manned / unmanned" and "day / night" to execute preset control strategies, which cannot perceive specific user behaviors or flexibly adjust the operating status according to the needs of the scenario.

[0267] This invention introduces a dual-dimensional trigger engine of "voice command + functional area" to realize composite trigger logic such as "user enters the sofa area and says 'turn on movie mode'", making control more precise, intelligent and in line with human behavior habits.

[0268] 4. The control strategy lacks adaptability and struggles to cover complex home scenarios:

[0269] Most air conditioning systems in related technologies operate in a "constant temperature operation" mode or use user-preset scenario parameters. They lack the ability to learn from and adjust to users' historical behavior, and therefore cannot provide a personalized comfort experience.

[0270] This invention generates an "activity heat map" and a "behavioral habit model" by continuously recording human activity trajectories and voice operation commands, thereby achieving dynamic optimization and automatic adjustment of control strategies.

[0271] 5. Energy-saving strategies rely on time control or manual settings, lacking proactive sensing capabilities:

[0272] Most of the energy-saving modes in related technologies are based on "automatic shutdown after a set time" or timed temperature adjustment, which cannot detect whether anyone is actually present, and may result in the air conditioner being turned off accidentally or resources being wasted.

[0273] This invention monitors in real time whether the user is still in a certain functional area and automatically switches to energy-saving standby mode when no one is around for a long time, while retaining the ability to restore the original settings, effectively improving energy efficiency.

[0274] 6. The related technologies lack sufficient system privacy protection capabilities, and there is a risk of leakage when uploading user data:

[0275] Most camera image recognition solutions rely on uploading images to the cloud for analysis, which poses a risk of image leakage, especially in private home spaces where users generally have security concerns.

[0276] The system of this invention supports image recognition and spatial reconstruction on the local gateway, with data remaining within the home network environment, thus protecting user privacy and improving response speed.

[0277] The functions described herein can be implemented in hardware, software executed by a processor, firmware, or any combination thereof. If implemented in software executed by a processor, the functions can be stored as one or more instructions or codes on or transmitted via a computer-readable medium. Other examples and embodiments are within the scope and spirit of this invention and the appended claims. For example, due to the nature of software, the functions described above can be implemented using software executed by a processor, hardware, firmware, hardwired, or any combination thereof. Furthermore, the functional units can be integrated into a single processing unit, or each unit can exist physically separately, or two or more units can be integrated into a single unit.

[0278] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. The device embodiments described above are merely illustrative; for example, the division of units can be a logical functional division, and in actual implementation, there may be other division methods. For instance, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the displayed or discussed mutual coupling, direct coupling, or communication connection may be through some interfaces; the indirect coupling or communication connection between units or modules may be electrical or other forms.

[0279] The units described as separate components may or may not be physically separate. Similarly, the components of the control device may or may not be physical units; they may be located in one place or distributed across multiple units. Some or all of the units can be selected to achieve the purpose of this embodiment, depending on actual needs.

[0280] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to related technologies, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, read-only memory (ROM), random access memory (RAM), portable hard drives, magnetic disks, or optical disks.

[0281] The above description is merely an embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of the claims of the present invention.

Claims

1. A method for controlling an air conditioner, characterized in that, include: Acquire a panoramic image of the space captured by a camera device; the panoramic image includes all the home appliances in the space; Image recognition is performed on the acquired panoramic image to obtain the types of home appliances in the space and their positions in the panoramic image; The system uses millimeter-wave radar to detect the user's location and skeletal information within the space, and outputs the three-dimensional coordinates of the user's location and the user's activity trajectory. The positions of each home appliance in the space obtained in the panoramic image are fused with the three-dimensional coordinates of the space output by the millimeter-wave radar to construct a spatial structure model of the space. The operation of the air conditioner is controlled by combining the constructed spatial structure model and the user activity trajectory detected by the millimeter-wave radar.

2. The method according to claim 1, characterized in that, Image recognition is performed on the acquired panoramic image to obtain the types of home appliances in the space and their positions in the panoramic image, including: The acquired panoramic image is input into a pre-trained target detection model for target proximity detection, thereby obtaining the type and location detection boxes of each home appliance in the space. The identified location detection boxes of each home appliance in the panoramic image are input into a pre-trained target segmentation model for edge segmentation to obtain the outline of each home appliance.

3. The method according to claim 1, characterized in that, The positions of each home appliance in the space obtained in the panoramic image are fused with the three-dimensional coordinates of the space output by the millimeter-wave radar to construct a spatial structure model of the space, including: The estimated three-dimensional position of each home appliance in the space is obtained in the coordinate system of the camera device by using a depth estimation algorithm; The estimated three-dimensional positions of each home appliance in the coordinate system of the camera device are obtained and converted into three-dimensional coordinates in the three-dimensional coordinate system of the millimeter-wave radar. The converted three-dimensional coordinates of each home appliance in the three-dimensional coordinate system of the millimeter-wave radar are fused with the three-dimensional coordinates of the space output by the millimeter-wave radar to construct a spatial structure model of the space.

4. The method according to claim 1, characterized in that, Combining the constructed spatial structure model and the user activity trajectory detected by the millimeter-wave radar, the operation of the air conditioner is controlled, including: The spatial structure model is divided into two or more regions using a spatial partitioning algorithm; The user activity trajectory detected by the millimeter-wave radar is superimposed on the spatial structure model to determine whether the user has entered any of the two or more regions. If it is determined that the user has entered any of the two or more areas and the triggering conditions of the preset scene mode corresponding to that area are met, then the air conditioner is controlled to execute the preset scene mode corresponding to that area.

5. The method according to claim 4, characterized in that, The triggering conditions for the preset scene mode corresponding to any one of the two or more areas include: the user entering the area and receiving a voice command to activate the preset scene mode corresponding to the area.

6. The method according to claim 4 or 5, characterized in that, Controlling the air conditioner to execute the preset scene mode corresponding to this area includes: The air conditioner is controlled to execute the air conditioner control parameters corresponding to the preset scene mode of the area, wherein different scene modes correspond to different air conditioner control parameters; Different air conditioning control parameters corresponding to different scenario modes can be set by the user, or set based on the user's operation behavior data of the air conditioner.

7. The method according to claim 4 or 5, characterized in that, Also includes: After controlling the air conditioner to execute the preset scene mode corresponding to the area, detect whether the user has left the area; If the system detects that a user has left the area and the time spent away exceeds a first preset time, the system controls the air conditioner to switch to a preset energy-saving mode.

8. The method according to claim 4 or 5, characterized in that, Also includes: Obtain historical activity data of the user in the two or more regions, wherein the historical activity data includes: time information and control information of the user appearing in each of the two or more regions each day; Based on the acquired historical activity data, a time-series prediction algorithm is used to predict the area the user will enter in the future and their control intentions; the control intentions refer to whether the user will issue a voice command to activate the preset scene mode corresponding to that area. The air conditioner is controlled in advance based on the predicted area the user will enter and their control intentions.

9. The method according to claim 4 or 5, characterized in that, Also includes: Obtain the user's historical activity trajectory detected by the millimeter-wave radar; The user's historical activity trajectory is overlaid onto the spatial structure model to form a heat map, thereby determining the user's active areas in the two or more regions; The air conditioner is controlled based on the user activity areas identified in the two or more regions.

10. A control device for an air conditioner, characterized in that, include: The first acquisition unit is used to acquire a panoramic image of the space captured by a camera device; the panoramic image includes various home appliances in the space. The identification unit is used to perform image recognition on the panoramic image acquired by the first acquisition unit to obtain the types of home appliances in the space and their positions in the panoramic image; The detection unit is used to detect the user's location and human skeleton information in the space using millimeter-wave radar, and outputs the three-dimensional coordinates of the user's location and the user's activity trajectory. The construction unit is used to fuse the positions of each home appliance in the space in the panoramic image with the three-dimensional coordinates of the space output by the millimeter-wave radar to construct a spatial structure model of the space. The control unit, combining the spatial structure model constructed by the building unit and the user activity trajectory detected by the millimeter-wave radar by the detection unit, controls the operation of the air conditioner.

11. A storage medium, characterized in that, It stores a computer program that, when executed by a processor, implements the steps of the method according to any one of claims 1-9.

12. An air conditioner, characterized in that, It includes a processor, a memory, and a computer program stored in the memory that can run on the processor, wherein the processor executes the program to implement the steps of any of the methods of claims 1-9, or includes the control device as described in claim 10.

13. A computer program product, characterized in that, Includes a computer program that, when executed by a processor, implements the steps of the method according to any one of claims 1-9.

Citation Information

Cited By

  • Smart home control method, system and device and storage medium

    CN121300112A