A multi-modal voice interaction intelligent hardware privacy protection system
By using multimodal voice interaction and millimeter-wave radar technology, the system can perceive and predict the location of targets within the camera's field of view in real time, generate and overlay occlusion masks, solve the problems of occlusion lag and robustness in intelligent camera systems, and achieve efficient privacy protection.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- MINAMI ACOUSTICS LTD
- Filing Date
- 2026-03-02
- Publication Date
- 2026-06-12
Smart Images

Figure CN122197059A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of smart hardware, and more specifically, to a privacy protection system for smart hardware with multimodal voice interaction. Background Technology
[0002] With the development of smart hardware technology, cameras have been widely deployed in private settings such as homes and offices for various purposes, including security monitoring, remote care, and smart connectivity. However, while providing convenience, smart camera systems have also brought increasingly serious privacy protection issues. How to promptly obscure sensitive areas or specific individuals in the frame without interrupting monitoring functions when visitors enter the home, or when users change clothes or engage in other private activities in front of the camera has become a key research focus in the current smart hardware field.
[0003] Most existing privacy occlusion technologies rely on image processing algorithms, such as object detection and facial recognition, to occlude the displayed target after the camera captures the image. These methods suffer from a significant lag issue: the object must first be captured and identified by the camera before the system can respond, inevitably leading to the risk of brief exposure of private information. Furthermore, in dynamic scenes, if someone quickly enters the camera's field of view, the generation and rendering of the occlusion mask often cannot synchronize with the target object's movement path, easily resulting in misaligned occlusion, missed occlusion, or jitter, affecting the occlusion effect and user trust.
[0004] Furthermore, traditional occlusion systems mostly involve fixed-area occlusion, lacking semantic understanding capabilities. They cannot flexibly adjust the occlusion target, area, and strategy based on user natural language commands, and struggle to handle complex privacy requirements involving multiple targets, semantics, and orientations. Additionally, existing systems generally rely on image information for occlusion calculations, often losing their occlusion capabilities and robustness when faced with camera malfunctions such as strong light, backlighting, occlusion, or image blur.
[0005] Therefore, there is an urgent need for an intelligent system that integrates multimodal perception capabilities, supports voice interaction, has predictive occlusion capabilities, and can continuously perform occlusion tasks even when images are unavailable. This system would be able to predict the location of a target and complete occlusion preparations before the target enters the camera's field of view, thereby effectively improving the response speed, occlusion accuracy, and overall system stability for privacy protection. Summary of the Invention
[0006] The technical problem to be solved by the present invention is to provide a smart hardware privacy protection system for multimodal voice interaction, so as to solve the problems mentioned in the background art.
[0007] To achieve the above objectives, the present invention adopts the following technical solution: A privacy protection system for smart hardware with multimodal voice interaction, comprising: The voice interaction module is used to acquire user voice commands and perform semantic parsing to determine the target object to be occluded and the occlusion strategy; The millimeter-wave radar module is used to perceive target objects approaching the camera in real time within a spatial buffer area in front of the camera's shooting direction but not yet within the camera's field of view, and to obtain their three-dimensional spatial position and motion trajectory. The occlusion prediction module is used to predict the image position of the target object after it enters the camera's field of view based on the three-dimensional position information sensed by millimeter-wave radar and the camera imaging parameters, and to generate a pre-occlusion command at the predicted position. The occlusion generation module is used to generate a corresponding occlusion mask at a corresponding position in the camera image coordinate system according to the pre-occlusion instruction and the occlusion strategy. The video rendering module is used to overlay the occlusion mask onto the camera's captured image in real time before the target object enters the camera's field of view; The occlusion dynamic update module is used to dynamically adjust the position and range of the occlusion mask based on continuous tracking data from millimeter-wave sensing after the target enters the camera's field of view, so that it always covers the position of the target object in the picture.
[0008] In some embodiments, the occlusion prediction module is further used to predict the target's movement trajectory within a preset time range based on the target's current position, speed, and direction of movement obtained by millimeter-wave radar, and to pre-generate a continuous occlusion mask strip to be used in the image area corresponding to the trajectory, and to call the continuous occlusion mask strip after the target actually enters the camera's field of view.
[0009] In some embodiments, the occlusion dynamic update module is used to adjust the position, shape, or extension range of the occlusion mask strip in real time after the target object actually enters the camera's field of view, based on the continuous tracking results of the millimeter-wave radar and the comparison and analysis of the camera image data, so as to maintain the continuous matching between the occlusion and the actual movement path of the target.
[0010] In some embodiments, the occlusion prediction module obtains the three-dimensional spatial position of the target object at consecutive time points, combines its velocity vector and acceleration information, and uses a motion trajectory prediction algorithm to calculate the target's movement path within a future preset time window.
[0011] In some embodiments, the motion trajectory prediction algorithm includes a multinomial prediction model based on Kalman filtering, extended Kalman filtering, or historical trajectory fitting to determine the expected movement area of the target in the image coordinate system.
[0012] In some embodiments, the video rendering module employs layer compositing technology and retains the original image data when performing occlusion mask overlay to support controlled occlusion recovery under subsequent authorized conditions.
[0013] In some embodiments, the occlusion strategy includes full occlusion, blurred occlusion, semi-transparent occlusion, edge fading occlusion, or time-controlled occlusion, and the user switches the occlusion type in real time via voice.
[0014] In some embodiments, the occlusion generation module generates an occlusion mask in the camera image coordinate system through the following steps: Based on the three-dimensional spatial coordinate data sensed by millimeter-wave radar, combined with the extrinsic and intrinsic parameters of the camera, the position of the target object is transformed from the world coordinate system to the camera image coordinate system using coordinate projection transformation. The center point of the occlusion mask is determined based on the position of the projected 2D image, and the shape, size, and boundary blurring parameters of the mask are set according to the target size, depth distance, and occlusion strategy parameters.
[0015] In some embodiments, the extrinsic parameters include the position and orientation of the camera.
[0016] In some embodiments, the intrinsic parameters include the camera's focal length, principal point coordinates, and distortion coefficient.
[0017] The advantages of this invention over existing technologies are that it introduces millimeter-wave radar to achieve early detection and occlusion prediction before the target enters the camera's field of view. Combined with voice semantic interaction and a dynamic occlusion update mechanism, it not only effectively avoids the privacy leakage problem of "exposure before occlusion" in traditional systems, but also improves the accuracy and continuity of occlusion. At the same time, it can still rely on millimeter waves to maintain the occlusion function in image interference scenarios, which has stronger robustness and real-time performance, and significantly enhances the privacy protection capabilities in smart hardware scenarios. Attached Figure Description
[0018] Figure 1 This is the overall processing flowchart of the intelligent hardware privacy protection system for multimodal voice interaction of the present invention; Figure 2 This is a flowchart of the voice interaction module and semantic parsing module of the present invention; Figure 3 This is a flowchart of the millimeter-wave radar module and its predictive processing according to the present invention. Detailed Implementation
[0019] The specific embodiments of the present invention will now be described with reference to the accompanying drawings.
[0020] The system of this invention aims to achieve privacy protection in a smart hardware environment through multimodal voice interaction and millimeter-wave radar technology.
[0021] like Figure 1 As shown, this invention combines technologies such as voice command parsing, millimeter-wave radar sensing, occlusion prediction, dynamic mask generation, and video rendering to ensure that the target object is effectively occluded before entering the camera's field of view, while also supporting dynamic adjustment of the occlusion area to adapt to the target's movement trajectory. More specifically, it includes: The voice interaction module is used to acquire user voice commands and perform semantic parsing to determine the target object to be occluded and the occlusion strategy; The millimeter-wave radar module is used to perceive target objects approaching the camera in real time within a spatial buffer area in front of the camera's shooting direction but not yet within the camera's field of view, and to obtain their three-dimensional spatial position and motion trajectory. The occlusion prediction module is used to predict the image position of the target object after it enters the camera's field of view based on the three-dimensional position information sensed by millimeter-wave radar and the camera imaging parameters, and to generate a pre-occlusion command at the predicted position. The occlusion generation module is used to generate a corresponding occlusion mask at a corresponding position in the camera image coordinate system according to the pre-occlusion instruction and the occlusion strategy. The video rendering module is used to overlay the occlusion mask onto the camera's captured image in real time before the target object enters the camera's field of view; The occlusion dynamic update module is used to dynamically adjust the position and range of the occlusion mask based on continuous tracking data from millimeter-wave sensing after the target enters the camera's field of view, so that it always covers the position of the target object in the picture.
[0022] Among them, such as Figure 2 As shown, the voice interaction module of this invention is responsible for receiving voice commands from the user, recognizing the command content through semantic parsing technology, determining the target object to be occluded, and the corresponding occlusion strategy. Users can express their privacy protection needs through natural language, such as "occlude a moving human figure in the living room" or "blur the kitchen area."
[0023] This module combines deep learning-based speech recognition models (such as end-to-end speech recognition models based on the Transformer architecture) with natural language processing models (such as BERT or its variants). The speech signal is first acquired through a microphone array, preprocessed (e.g., noise reduction and echo cancellation), and then input into the speech recognition model to generate text commands. Subsequently, the text commands are sent to the semantic parsing module, which extracts key information through intent recognition and slot filling techniques, such as the target object (person, pet, or other object), the occlusion region (specific room or area), and the type of occlusion (complete occlusion, blurred occlusion, etc.).
[0024] To ensure the accuracy of semantic parsing, the system uses a dedicated speech dataset containing smart hardware scenarios when training the speech recognition model, covering various language expressions and accents. The semantic parsing model is optimized for common commands in the home environment through supervised learning and few-shot learning techniques. For example, when a user says "blurry occlusion of the person in the living room," the system can resolve the target object to "person," the area to "living room," and the occlusion strategy to "blurry occlusion." Furthermore, the module supports real-time command updates, allowing users to switch occlusion strategies at any time via voice, such as from complete occlusion to semi-transparent occlusion, to meet privacy needs in different scenarios.
[0025] Millimeter-wave radar modules are used to sense target objects within a spatial buffer zone in front of the camera's field of view, acquiring their three-dimensional spatial position and trajectory. The spatial buffer zone is defined as an area within a certain range (e.g., 2 to 5 meters) of the camera, not yet within the camera's field of view. Millimeter-wave radar operates in the 77 GHz to 81 GHz frequency band, possessing high resolution and penetration capabilities, and can detect the target's distance, velocity, and angle information. The radar system employs a multi-input multi-output antenna array and uses beamforming technology to achieve high-precision three-dimensional positioning.
[0026] like Figure 3 As shown, in the specific implementation, the millimeter-wave radar scans the spatial buffer area every fixed time interval (e.g., 10 milliseconds) to generate point cloud data of the target object. The point cloud data contains the target's x, y, and z coordinates and velocity vector. The target's motion information is extracted from the radar echo using signal processing algorithms (such as Fast Fourier Transform (FFT) and Constant False Alarm Rate (CFAR) detection). The system performs cluster analysis on the point cloud data to distinguish different target objects (such as people and pets) and assigns a unique tracking ID to each target. The tracking algorithm is based on a Kalman filter and combines historical data to smooth the target's trajectory and reduce noise interference. For example, when a person is detected approaching the camera at a speed of 0.5 m / s, the system records their current position as (2m, 1m, 1m) and updates their velocity vector and direction of motion in real time.
[0027] The occlusion prediction module, based on the 3D position information acquired by millimeter-wave radar and camera imaging parameters, predicts the image position of a target object as it enters the camera's field of view and generates a pre-occlusion command. This module first transforms the target's 3D world coordinates into a 2D position in the camera's image coordinate system through coordinate projection transformation. This transformation involves the camera's intrinsic and extrinsic parameters. Extrinsic parameters include the camera's position (e.g., 3D coordinates relative to the room's origin) and orientation (expressed in Euler angles), while intrinsic parameters include focal length (in pixels), principal point coordinates (image center point), and distortion coefficients (used to correct lens distortion). Assuming the target object's world coordinates are (x... w ,y w ,zw The image coordinates (u,v) are calculated using the following projection formula: u=f x ×(x w / z w )+c x v=f y ×(y w / z w )+c y Among them, f x and f y c is the focal length of the camera. x and c y The coordinates of the main point are used. Distortion correction further adjusts (u,v) using distortion coefficients k1, k2, etc., to ensure the accuracy of the projected position.
[0028] To predict the target's trajectory, the system employs Kalman filtering and extended Kalman filtering algorithms, combining the target's current position, velocity, and acceleration information to calculate its expected position within a preset time window (e.g., 0.5 to 2 seconds). Kalman filtering assumes the target's motion is a linear model, with the state vector including position (x, y, z) and velocity (v). x ,v y ,v z The state is updated using the following state transition equation: X t =F×X t-1 +w; Where F is the state transition matrix, w is the process noise, t represents time, and X represents the state vector. Extended Kalman filtering handles nonlinear motion scenarios, improving prediction accuracy by linearizing the motion equations. For complex motion patterns, the system can also use a multinomial fitting model based on historical trajectories to fit the target's position data at continuous time points, generating a smooth trajectory curve. For example, if the target moves at a constant velocity in a straight line, the system predicts its position after 0.5 seconds as (x...). t +0.5V x ,y t +0.5V y ,z t +0.5V z ).
[0029] After prediction, the system generates a continuous occlusion mask strip in the image area where the target is expected to enter. The mask strip consists of a series of two-dimensional occlusion masks, covering all possible positions of the target within a preset time window. The shape and size of the mask are determined based on the target's dimensions (estimated from the radar point cloud) and occlusion strategy parameters. For example, for a humanoid target, the mask might be elliptical, 0.5 meters wide and 1.8 meters high, with a boundary blur parameter set to 10 pixels for a smooth transition.
[0030] The occlusion generation module generates an occlusion mask in the camera image coordinate system based on the pre-occlusion instructions and occlusion strategy. The specific steps include: First, based on the three-dimensional coordinates provided by the millimeter-wave radar and the camera's intrinsic and extrinsic parameters, the target position is converted into image coordinates (u,v) using the aforementioned projection formula. Then, the shape and size of the mask are determined according to the target's dimensions (e.g., the width and height of a person) and depth distance (which affects the mask's pixel size). The occlusion strategy parameters further define the mask type; for example, complete occlusion uses a pure black rectangle (RGB value 0,0,0), blurred occlusion uses a Gaussian blur kernel (standard deviation range of 5 to 15 pixels), semi-transparent occlusion sets transparency parameters (α value range of 0.3 to 0.7), and edge fading occlusion uses linear interpolation to achieve a gradual change in boundary pixel transparency.
[0031] For example, when the target is a human-shaped object located 3 meters away from the camera, the system calculates its image coordinates as (u=640, v=360), and the mask size is 200x400 pixels (corresponding to a width of 0.5 meters and a height of 1.8 meters). If the user selects blur occlusion, the system generates a Gaussian blur mask with a standard deviation of 10 pixels, covering the target area. After the mask is generated, it is stored as a temporary image layer for use by the video rendering module.
[0032] The video rendering module is responsible for overlaying an occlusion mask onto the camera's field of view in real time before the target object enters the camera's field of view. The system employs layer compositing technology, merging the mask as a foreground layer with the original camera image (background layer). The compositing process uses OpenCV or a similar image processing library, calculating the composite pixel values using the following formula: I out =(1-α)×I bg +α×I mask ; Among them, I out To output the image, I bg For background image, I mask For masking, α represents the transparency of the mask. For complete masking, α = 1; for semi-transparent masking, α ranges from 0.3 to 0.7; for blurred masking, α = 1. maskThis represents the blurred image area. To support subsequent occlusion recovery under authorized conditions, the system stores the original image data in a separate buffer. The original, unoccluded image can be accessed after authorization verification.
[0033] For example, when a target enters the field of view, the system overlays a 200x400 pixel blur mask at image coordinates (640, 360) to generate a privacy-preserving video stream. The rendering process remains real-time, with a frame rate of no less than 30fps to ensure smooth playback.
[0034] The occlusion dynamic update module dynamically adjusts the position, shape, and extent of the occlusion mask based on continuous tracking data from the millimeter-wave radar and camera image data after the target enters the camera's field of view. The system achieves dynamic updates through the following steps: First, the millimeter-wave radar updates the target's 3D position and velocity information every 10 milliseconds, generating new point cloud data. The occlusion prediction module recalculates the target's image coordinates and expected trajectory based on this data. Subsequently, the system compares and analyzes the millimeter-wave radar's predicted position with the target position in the camera image, using target detection algorithms (such as YOLO or SSD) to identify the target contour in the image and correct prediction errors.
[0035] For example, if the radar predicts the target is located at image coordinates (650, 370), but image detection shows the actual location is (645, 365), the system adjusts the mask center point to (645, 365) and updates the mask size according to the actual size of the target. The mask shape can be dynamically adjusted according to the target's attitude, for example, changing from a rectangle to an ellipse to accommodate the target's tilt. The update frequency is consistent with the radar scanning frequency (100Hz), ensuring continuous matching between the masked area and the target's movement path.
[0036] The motion trajectory prediction algorithm is the core of the occlusion prediction module, combining Kalman filtering, extended Kalman filtering, and a multinomial fitting model to achieve high-precision prediction. Kalman filtering is suitable for linear motion scenarios, assuming the target moves at a constant speed, and the state transition matrix F is: F=[[1,0,0,dt,0,0], [0,1,0,0,dt,0], [0,0,1,0,0,dt], [0,0,0,1,0,0], [0,0,0,0,1,0], [0,0,0,0,0,1]]; Where dt is the time step (e.g., 0.01 seconds). For nonlinear motion, the extended Kalman filter linearizes the motion equations using the Jacobian matrix to handle changes in the target's acceleration. The polynomial fitting model, based on the target's position data over the past second, fits a quadratic or cubic polynomial curve to predict its position within the next 0.5 seconds. The prediction error is evaluated using the covariance matrix, with the error range controlled within 5 centimeters.
[0037] Multiple occlusion strategies are supported, and users can switch between them in real time via voice commands. Complete occlusion fills the target area with a solid color, suitable for high-privacy scenarios; blurred occlusion uses Gaussian blur to preserve the target outline but hide details; semi-transparent occlusion allows partial background visibility, suitable for low-privacy needs; edge-fading occlusion enhances visual effects through gradual transparency; and time-controlled occlusion allows users to specify the occlusion duration (e.g., automatically canceling after 5 seconds). Parameter ranges are as follows: Blur occlusion: Gaussian kernel standard deviation 5 to 15 pixels.
[0038] Semi-transparent occlusion: transparency α value 0.3 to 0.7.
[0039] Edge fading: Boundary width 10 to 50 pixels.
[0040] Suppose a user in the living room says, "Blurringly occlude the person approaching the camera." The voice interaction module parses the command, identifies the target as a "person," and sets the occlusion strategy to "blurring occlusion." Millimeter-wave radar detects a human-shaped target 3 meters from the camera, approaching at a speed of 0.5 meters per second. The occlusion prediction module uses a Kalman filter to predict that the target will enter the field of view in 0.5 seconds, with image coordinates (640, 360). The occlusion generation module generates a 200x400 pixel blurred mask (standard deviation 10 pixels). The video rendering module overlays the mask onto the image before the target enters the field of view. After the target enters the field of view, the occlusion dynamic update module adjusts the mask position based on radar and image data to ensure the occlusion area always covers the target.
[0041] This invention introduces millimeter-wave radar, enabling the system to predict occlusion before the target enters the camera's field of view, thus avoiding the privacy leaks associated with traditional "exposure before occlusion" methods. A dynamic update mechanism, combining radar and image data, ensures the continuity and accuracy of occlusion prediction. The voice interaction module provides flexible user control, supporting multiple occlusion strategies to adapt to different privacy needs. The system's modular design facilitates expansion; for example, it can integrate a deep learning target detection model to further improve occlusion accuracy.
[0042] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.
Claims
1. A smart hardware privacy protection system for multimodal voice interaction, characterized in that, include: The voice interaction module is used to acquire user voice commands and perform semantic parsing to determine the target object to be occluded and the occlusion strategy; The millimeter-wave radar module is used to perceive target objects approaching the camera in real time within a spatial buffer area in front of the camera's shooting direction but not yet within the camera's field of view, and to obtain their three-dimensional spatial position and motion trajectory. The occlusion prediction module is used to predict the image position of the target object after it enters the camera's field of view based on the three-dimensional position information sensed by millimeter-wave radar and the camera imaging parameters, and to generate a pre-occlusion command at the predicted position. The occlusion generation module is used to generate a corresponding occlusion mask at a corresponding position in the camera image coordinate system according to the pre-occlusion instruction and the occlusion strategy. The video rendering module is used to overlay the occlusion mask onto the camera's captured image in real time before the target object enters the camera's field of view; The occlusion dynamic update module is used to dynamically adjust the position and range of the occlusion mask based on continuous tracking data from millimeter-wave sensing after the target enters the camera's field of view, so that it always covers the position of the target object in the picture.
2. The intelligent hardware privacy protection system for multimodal voice interaction according to claim 1, characterized in that, The occlusion prediction module is further used to predict the target's movement trajectory within a preset time range based on the target's current position, speed, and direction of movement obtained by millimeter-wave radar, and to pre-generate a continuous occlusion mask strip to be used in the image area corresponding to the trajectory, and to call the continuous occlusion mask strip after the target actually enters the camera's field of view.
3. The intelligent hardware privacy protection system for multimodal voice interaction according to claim 2, characterized in that, The occlusion dynamic update module is used to adjust the position, shape, or extension range of the occlusion mask strip in real time after the target object actually enters the camera's field of view, based on the continuous tracking results of the millimeter-wave radar and the comparison and analysis of the camera image data, so as to maintain the continuous matching between the occlusion and the actual movement path of the target.
4. The intelligent hardware privacy protection system for multimodal voice interaction according to claim 2, characterized in that, The occlusion prediction module obtains the three-dimensional spatial position of the target object at continuous time points, combines its velocity vector and acceleration information, and uses a motion trajectory prediction algorithm to calculate the target's movement path within a future preset time window.
5. The intelligent hardware privacy protection system for multimodal voice interaction according to claim 4, characterized in that, The motion trajectory prediction algorithm includes a multinomial prediction model based on Kalman filtering, extended Kalman filtering, or historical trajectory fitting to determine the expected movement area of the target in the image coordinate system.
6. The intelligent hardware privacy protection system for multimodal voice interaction according to claim 1, characterized in that, The video rendering module employs layer compositing technology and retains the original image data when performing occlusion mask overlay to support controllable occlusion recovery under subsequent authorized conditions.
7. The intelligent hardware privacy protection system for multimodal voice interaction according to claim 1, characterized in that, The occlusion strategies include full occlusion, blurred occlusion, semi-transparent occlusion, edge fading occlusion, or time-controlled occlusion, and users can switch the occlusion type in real time via voice.
8. The intelligent hardware privacy protection system for multimodal voice interaction according to claim 1, characterized in that, The occlusion generation module generates an occlusion mask in the camera image coordinate system through the following steps: Based on the three-dimensional spatial coordinate data sensed by millimeter-wave radar, combined with the extrinsic and intrinsic parameters of the camera, the position of the target object is transformed from the world coordinate system to the camera image coordinate system using coordinate projection transformation. The center point of the occlusion mask is determined based on the position of the projected 2D image, and the shape, size, and boundary blurring parameters of the mask are set according to the target size, depth distance, and occlusion strategy parameters.
9. The intelligent hardware privacy protection system for multimodal voice interaction according to claim 8, characterized in that, The external parameters include the position and orientation of the camera.
10. The intelligent hardware privacy protection system for multimodal voice interaction according to claim 8, characterized in that, The intrinsic parameters include the camera's focal length, principal point coordinates, and distortion coefficient.