A pointing device recognition and control method and device based on vision and spatial perception, and a storage medium

By combining visual and spatial perception to create a directional device recognition method, and utilizing laser pointers and camera recognition combined with a lightweight neural network, the problem of complex operation and misidentification of smart home devices is solved. This achieves low-cost, fast, and accurate device control, making it suitable for the elderly and protecting their privacy.

CN122284404APending Publication Date: 2026-06-26BEIJING HAOWANG TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIJING HAOWANG TECHNOLOGY CO LTD
Filing Date
2026-03-05
Publication Date
2026-06-26

AI Technical Summary

Technical Problem

In existing technologies, the control methods for smart home devices suffer from problems such as complex operation, high misidentification rate, high cost, complex deployment, and inability to achieve scene linkage between devices, making them particularly unsuitable for the elderly and users with limited resources.

Method used

A directional device recognition method based on vision and spatial perception is adopted. It combines laser pointer spot and miniature camera with lightweight neural network for device recognition and utilizes existing smart home devices for collaborative positioning to achieve accurate device recognition and control.

Benefits of technology

It enables device control without screen operation, at low cost, and with fast response. It has high recognition accuracy and precise control, making it suitable for the elderly. It also does not rely on the network and protects user privacy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122284404A_ABST
    Figure CN122284404A_ABST
Patent Text Reader

Abstract

This invention discloses a method, device, and storage medium for directional device identification and control based on vision and spatial perception, belonging to the field of intelligent control and human-computer interaction technology. The method is applied to portable terminals, achieving "point-and-shoot" by using a coaxially integrated laser and camera module to acquire target images containing laser pointer spots. To address the identification challenges caused by differences in target object size and distance, the method proposes multi-scale adaptive region extraction, generating and matching multiple candidate regions centered on the laser pointer spot. Simultaneously, to eliminate control ambiguities for similar devices in different rooms, the method innovatively utilizes smart home device nodes in known locations within the environment as Bluetooth beacons, achieving zero-deployment-cost collaborative spatial perception and obtaining the terminal's precise spatial location. Finally, by jointly querying and adjudicating the visual recognition results and spatial location information, the target device is uniquely identified and control commands are generated. This invention achieves intuitive interaction with zero learning costs, solving the problems of complexity and confusion in traditional control methods. It has advantages such as accurate identification, fast response, simple deployment, and privacy security, making it particularly suitable for smart home and smart elderly care scenarios.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent control and human-computer interaction technology, specifically to a method, device and storage medium for directional device identification and control based on the fusion of vision and spatial perception, applicable to scenarios such as smart homes, smart elderly care, smart exhibition halls, and smart spaces. Background Technology

[0002] With the increasing prevalence of the Internet of Things (IoT) and smart homes, the number of controllable devices in homes is growing rapidly. The mainstream methods for users to control these devices include mobile apps, smart speaker voice control, and traditional multiple remote controls. All of these methods have significant limitations: Mobile App: Users need to find and turn on their phones, navigate through complex app interfaces to locate corresponding devices and controls, and navigate long operation paths, making it extremely unfriendly to the elderly and users unfamiliar with smart devices.

[0003] Smart speakers rely on voice commands, which are not effective in noisy environments, private scenarios, or when precise control of specific objects is required (such as "turn on the left light"), and they also have problems with false wake-up and misrecognition.

[0004] Multiple traditional remote controls: Users need to manage and remember multiple remote controls, which are easy to lose or get confused, and cannot achieve scene linkage between devices.

[0005] Existing technologies have also seen some attempts to simplify interaction. For example, patent application CN113XXX discloses a "home appliance control method based on image recognition," which pops up a virtual control panel after recognizing the home appliance through a camera. However, this method still requires users to make secondary selections on a small mobile phone screen, failing to achieve true "what you point to is what you get." Another example is the use of UWB or Bluetooth beacons for high-precision indoor positioning, but this requires users to deploy dedicated base stations or beacon networks, increasing cost and deployment complexity.

[0006] Therefore, the industry urgently needs an intelligent control solution that is highly intuitive, requires no secondary screen operation, is easy to deploy, and can accurately distinguish between similar devices. Summary of the Invention

[0007] (a) Purpose of the invention To address the shortcomings of the existing technology, this invention aims to provide a method, apparatus, and storage medium for directional device identification and control based on vision and spatial perception. The objective of this invention is: It provides a zero-learning-cost interaction method that allows users to accurately select the target device through natural pointing gestures.

[0008] This addresses the problem of inaccurate positioning of the recognition box on portable terminals with no screen or limited screen, due to changes in the size and distance of the target object.

[0009] It enables low-cost, highly reliable room-level positioning without the need for additional dedicated positioning hardware, and eliminates control ambiguities for similar devices in different spaces by combining this positioning information.

[0010] Build a complete control loop that enables localized processing on resource-constrained embedded terminals, ensuring fast response, good privacy, and no dependence on the network.

[0011] (II) Technical Solution To achieve the above objectives, the present invention adopts the following technical solution: In a first aspect, the present invention provides a directional device identification and control method based on vision and spatial perception, which is applied to a portable terminal with image acquisition, laser emission and wireless communication functions.

[0012] Figure 1 A flowchart illustrating a control method provided in an embodiment of the present invention. Figure 1 As shown, the method includes the following steps: S110: Point-triggered operation and image acquisition. In response to a user's point-triggered operation (such as a specific gesture detected by the IMU, or a physical button being pressed), the laser emission module at the top of the control terminal emits a visible laser beam (e.g., 650nm) towards the target object, forming a conspicuous indicator light spot on the surface of the target object. At the same time, a scene image containing the light spot and its surrounding environment is simultaneously captured by a miniature camera module integrated coaxially or precision excimer with the laser, serving as the "target image".

[0013] S120: Multi-scale Adaptive Region Extraction. This step aims to address the problem of ineffective recognition using a fixed-size bounding box due to the variable distance between the user and the target object, and the varying physical dimensions of the target object itself. The processor first identifies the bright laser spot in the target image and calculates its pixel coordinates and range. Then, using the center of the spot as a reference, the algorithm automatically generates multiple candidate cropping boxes of different sizes and / or aspect ratios. For example, three boxes can be generated: one slightly smaller than the spot (for recognizing small objects such as switches), one similar in size to the spot, and one much larger than the spot (for recognizing large objects such as televisions). These candidate boxes cover different scales that the target object may present.

[0014] S130: On-device parallel feature matching. Multiple candidate image regions obtained in step S120 are simultaneously input into a lightweight neural network model (such as a variant of MobileNetV2) integrated locally on the terminal. This model acts as a feature extractor, outputting a high-dimensional feature vector for each candidate region. Subsequently, these feature vectors are compared with a device feature vector library pre-stored in the terminal's local flash memory using similarity calculations (such as cosine similarity). The device type corresponding to the candidate region with the highest matching degree is determined as the "preliminary device type identifier".

[0015] S140: Collaborative Spatial Awareness. Simultaneously or before / after visual recognition, the terminal's wireless communication module (such as Bluetooth) continuously scans the surrounding environment. Crucially, the scanned objects are not ordinary Bluetooth devices, but rather other smart home device nodes whose physical installation locations are known within the system (such as a smart light Bluetooth module installed in the living room, or a smart curtain motor Bluetooth module installed in the bedroom). These nodes are configured to periodically broadcast their own device ID and pre-bound "spatial location identifiers" (such as "living room" or "master bedroom"). Based on the strength of these received signals (RSSI), the terminal uses a specific positioning algorithm (such as nearest neighbor or fingerprint matching) to determine its precise room or area of ​​current location.

[0016] S150: Unique Target Decision. This step is central to resolving control ambiguity. The terminal maintains a local device database where each record contains not only the device's feature vector and name but also an explicit association with its associated "spatial location identifier." The processor performs a joint query on the "preliminary device type identifier" obtained in step S130 (e.g., "ceiling light") and the "precise spatial location" obtained in step S140 (e.g., "living room"). Specifically, it first filters the database for all devices located in the "living room," and then finds the record within this subset that has the highest match score with the "ceiling light" type. The device corresponding to this record is the unique target device that the user ultimately wants to control. This fundamentally eliminates the possibility of accidentally controlling the same device in other rooms.

[0017] S160: Control Command Generation and Execution. After uniquely identifying the target device, the terminal generates a corresponding control command based on the user's subsequent operation (such as pressing the "power" button). This command can be an infrared code, a radio frequency signal, or a network protocol command sent via Bluetooth / Wi-Fi, and is directly issued by the terminal to control the target device to perform the corresponding action.

[0018] (III) Beneficial Effects Compared with the prior art, the present invention has the following significant advantages: Extremely intuitive interaction: It achieves true "what you point to is what you get", and users do not need to learn complex commands or operate multi-layered menus. It is especially suitable for the elderly, children and all users who pursue convenience, effectively bridging the digital divide.

[0019] High recognition accuracy and strong robustness: The "multi-scale adaptive region extraction" technology effectively overcomes the recognition challenges caused by changes in target scale and distance, thus improving the success rate of first-shot recognition.

[0020] By integrating "visual recognition" with "collaborative spatial perception based on existing equipment" for decision-making, the problem of identification ambiguity of similar equipment in different rooms is perfectly solved, and the control accuracy is close to 100%.

[0021] The system is low-cost and easy to deploy: Users do not need to purchase and install dedicated indoor positioning base stations or beacons. It innovatively reuses existing smart home devices as positioning reference points, achieving "zero-cost" room-level positioning capabilities, greatly reducing the deployment threshold and complexity of the entire system. Swift response and secure privacy: All image processing, feature matching, and decision-making are completed locally on the terminal, forming an end-to-end control loop. This not only results in extremely low operation latency (typically <300ms) but also avoids uploading images of the user's home environment and habit data to the cloud, fully protecting user privacy.

[0022] High degree of functional integration: This solution highly integrates advanced human-computer interaction, machine vision, embedded AI and IoT technologies into a portable terminal device, representing the future development direction of smart home control terminals. Attached Figure Description

[0023] Figure 1 A flowchart of a control method provided in an embodiment of the present invention.

[0024] Figure 2 This is a hardware structure block diagram of a portable intelligent control terminal provided in an embodiment of the present invention.

[0025] Figure 3 This is a schematic diagram of "multi-scale adaptive region extraction" in one embodiment of the present invention.

[0026] Figure 4 This is a schematic diagram illustrating the working principle of "cooperative spatial perception" and "unique target adjudication" in one embodiment of the present invention. Detailed Implementation

[0027] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be noted that the following description is merely illustrative and does not constitute a limitation on the scope of protection of this invention.

[0028] Example 1: Reference Figures 1-4 This embodiment describes the working process of a complete "TianShu Intelligent Remote Control".

[0029] The hardware structure of the portable intelligent control terminal (i.e., remote control) is as follows: Figure 2 As shown, it includes: a main control MCU, an NPU (for running feature extraction models), memory, a coaxial integrated camera and 650nm laser module, a six-axis IMU, a Bluetooth / Wi-Fi / infrared / RF multi-mode communication module, buttons, a microphone, a speaker, and a battery.

[0030] The user picks up the remote control, the IMU detects the movement, and automatically illuminates a laser pointer. The user points the pointer at the ceiling light in the living room and presses the "Recognize" button.

[0031] S110: The camera takes a picture that includes the light spot and the ceiling light.

[0032] S120: The processor uses the light spot as the center to crop out three sub-images: large, medium, and small regions (e.g., ...). Figure 3 (As shown).

[0033] S130: The NPU processes three sub-graphs in parallel, extracts features, and matches them against the local database. The result points to the "ceiling light" category.

[0034] S140: The Bluetooth module simultaneously scanned strong signals from both the "Living Room Smart Main Light Switch" and the "Living Room Curtain Motor". Both broadcast the space ID "Living Room", so it is determined that the terminal is located in the living room.

[0035] S150: Query the local database. There is only one record that meets the criteria of "Device Type = Ceiling Light" and "Spatial Location = Living Room", which is the user-preset "Living Room Main Light".

[0036] S160: The remote control announces "Living room main light" via voice. The user then presses the "On" button, and the remote control transmits the corresponding 2.4G radio frequency switch command, turning on the "Living room main light".

[0037] Example 2: New device self-learning process.

[0038] A user purchases a new smart fan and places it in the bedroom. The remote control then enters "self-learning" mode.

[0039] The user points to the fan, presses the learning button, completes S110-S130, and obtains the fan's feature vector.

[0040] The system matches the vector with the built-in "General Home Appliance Feature Library" and prompts "Identified as: Fan, is that correct?" The user confirmed via voice: "Yes."

[0041] The system prompted, "Please name it." The user replied, "Bedroom fan."

[0042] At this point, step S140 has determined the current location to be "Bedroom". The system automatically binds the name "Bedroom Fan", the fan's feature vector, and the spatial ID "Bedroom" and stores them in the local database.

[0043] At this point, the new device has completed its learning process and can then be pointed at and controlled like any other device.

[0044] The above description is merely a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.

Claims

1. A pointing device identification and control method based on vision and spatial perception, applied to a portable terminal with image acquisition, laser emission and wireless communication functions, characterized in that, The method includes: Step S110: In response to the pointing trigger operation, control the terminal to emit an indicator laser to form an indicator spot on the surface of the target object, and synchronously acquire a target image containing the spot through an image acquisition module associated with the laser optical path; Step S120: Multi-scale adaptive region extraction step: Based on the visual features of the indicator spot in the target image, with the spot as the center, automatically determine and extract at least two candidate image regions with different scales and / or aspect ratios to adapt to the problem of single recognition box matching failure caused by the unknown actual physical size of the target object or the uncertain distance between the terminal and the target object. Step S130: On-device parallel feature matching step: Using a lightweight neural network model deployed locally on the terminal, feature vectors of each candidate image region are extracted and matched in parallel with the device feature vector library pre-stored locally on the terminal. A preliminary device type identifier is obtained based on the matching results. Step S140: Cooperative spatial perception step: During or before / after the execution of steps S110 to S130, the terminal's wireless communication module scans the wireless beacon signals broadcast by one or more controllable device nodes with known spatial location information in its current environment; based on the strength of each received beacon signal and the spatial location identifier bound to it, the precise spatial location of the terminal is determined. Step S150: Unique Target Determination Step: Perform a joint query between the preliminary device type identifier and the precise spatial location, and determine a unique target device identifier from the local database that stores the device identifier, device feature vector and spatial location in association, so as to resolve the identification ambiguity caused by similar devices in different physical spaces; Step S160: Control command generation and execution step: Based on the unique target device identifier, generate the corresponding device control command and execute it.

2. The method of claim 1, wherein, The "pointer trigger operation" in step S110 is determined by any one or a combination of the following methods: The inertial measurement unit within the terminal detects movements that conform to preset gesture characteristics; The system detects user actions on specific physical buttons or touch areas on the terminal.

3. The method of claim 1, wherein, The step S120, "automatically identifying and extracting at least two candidate image regions with different scales and / or aspect ratios," specifically includes: Identify the pixel range of the indicated light spot in the image and calculate its equivalent diameter or circumscribed rectangle size; Based on the center coordinates of the light spot, multiple candidate boxes are generated, wherein the size of at least one candidate box is smaller than the range of the light spot, and the size of at least one candidate box is larger than the range of the light spot.

4. The method according to claim 1, characterized in that, In step S140: The wireless communication module is a Bluetooth module; The controllable device node is a smart home execution device with Bluetooth broadcasting function, and the beacon signal broadcast by the device includes at least its unique device identifier and a pre-configured spatial location identifier.

5. The method according to claim 1, characterized in that, The step S150, "determining a unique target device identifier from a local database that associates and stores device identifiers, device feature vectors, and spatial locations," includes: First, all device records that match the precise spatial location are filtered out from the local database to form a first candidate set; Then, in the first candidate set, the device record with the highest matching degree with the preliminary device type identifier is selected, and its device identifier is the unique target device identifier.

6. The method according to claim 1, characterized in that, After step S110 and before step S120, an image acquisition stabilization step is also included**: Acquire multiple frames of preview images continuously output by the image acquisition module after triggering; Analyze the pixel-level displacement of the indicator spot in the multi-frame preview images; The current frame or the next frame image is determined as the target image for subsequent steps only when the displacement is continuously below a preset threshold.

7. The method according to claim 1, characterized in that, The method also includes a new device self-learning step: In response to the self-learning instruction, steps S110 to S130 are executed to obtain sample images of the new device and their feature vectors. The sample feature vector of the new device is matched with a pre-stored general device feature library; If the match is successful, the system will output the suggested device category information to the user and receive confirmation or correction. The confirmed equipment category information, user-defined names, sample feature vectors, and the currently determined precise spatial location are associated and stored in the local database.

8. A portable intelligent control terminal, characterized in that, include: case; The main controller, memory, image sensor, laser emitter, inertial measurement unit, and wireless communication module are housed within the casing. The image sensor and the laser emitter are integrated into the pointing end of the housing in a coaxial or fixed paraaxial optical structure. The memory stores computer programs and device feature vector libraries; The main controller is connected to the image sensor, laser emitter, inertial measurement unit, and wireless communication module, and is configured to execute the computer program to implement the method as described in any one of claims 1 to 7.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the method as described in any one of claims 1 to 7.