Privacy protection intelligent visual terminal and visualization method based on physical switch and artificial intelligence
The intelligent vision terminal, which combines physical switches with artificial intelligence, solves the problems of unreliable privacy protection and insufficient behavior recognition in high privacy-sensitive scenarios of existing video surveillance equipment. It enables mode switching for on-site operation control and efficient behavior monitoring, thereby improving user experience and security.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-01-28
- Publication Date
- 2026-04-14
AI Technical Summary
Existing video surveillance equipment, due to its reliance on software control in highly privacy-sensitive scenarios, suffers from unreliable privacy protection, lacks behavioral recognition capabilities in desensitized images, and is disconnected from on-site operations when switching modes, making it difficult to meet the actual needs of high-density care scenarios.
The intelligent vision terminal, which combines physical switches with artificial intelligence, allows mode switching to be triggered directly by physical switching buttons. Combined with artificial intelligence, it generates desensitized videos that are consistent with the environment in real time, ensuring that mode switching is controlled by on-site operation and that video synthesis and analysis are completed on the device.
It ensures the reliability of privacy protection and the accuracy of behavior monitoring in highly privacy-sensitive scenarios, reduces network bandwidth consumption, adapts to the flexible needs of complex care scenarios, and improves user acceptance and security.
Smart Images

Figure FT_1 
Figure FT_2 
Figure FT_3
Abstract
Description
Technical Field
[0001] This invention belongs to the field of video surveillance and artificial intelligence technology, specifically relating to a dual-mode privacy-protecting visual terminal and visualization method based on physical switches and edge intelligence. Background Technology
[0002] In large-scale senior living communities, nursing homes, and hospital wards, caregivers typically conduct manual rounds at fixed intervals, such as every 30 minutes or an hour. Due to the large number of rooms and their wide distribution across floors, a single complete round often takes a considerable amount of time. During this time, if residents experience falls, sudden fainting, unconsciousness while out of bed at night, or prolonged periods of stillness, these abnormalities may go undetected, posing a significant safety risk.
[0003] To improve monitoring efficiency, some institutions have tried installing video surveillance equipment in rooms, hoping to reduce reliance on manual inspections through remote viewing. However, traditional high-definition cameras continuously collect and transmit raw video containing complete human form, facial features, and even exposed body parts. This type of content can easily trigger resistance from residents and their families during daily private activities such as changing clothes, using the toilet, and lying in bed. It also makes caregivers feel uncomfortable when providing close care, affecting the normal conduct of care work.
[0004] Some camera products on the market currently offer a so-called privacy mode. Users can remotely enable image blurring, area occlusion, or switch to a skeleton map display generated by a pose estimation algorithm via a mobile application. However, enabling and disabling these functions relies entirely on software commands, which can be remotely controlled via a network interface. In actual use, there have been multiple cases where privacy mode has been accidentally disabled due to insufficient account password strength, unpatched firmware vulnerabilities, or accidental operation by family members. Because there are no physical traces of the switching process, residents cannot confirm whether the device is currently in a protected state, and administrators also find it difficult to trace whether a mode change was authorized on-site.
[0005] Furthermore, existing desensitization solutions have significant shortcomings in practicality. Some products only blur the overall image, which, while hiding details, fails to identify whether a person is active or in a sitting or lying position, thus losing the basic value of behavior monitoring. Other products can output skeletal images, but because the background image is usually captured during device initialization, even if the furniture position, curtains, or lighting conditions in the room are changed, the system still uses the old background for synthesis, resulting in obvious misalignment between the human figure and the environment. For example, limbs may appear to float in the air or not align with the edge of the bed, affecting the accuracy of the judgment.
[0006] In actual care procedures, the need for video modes is dynamic. Under normal circumstances, staff activity needs to be continuously monitored using an anonymized method; when the system detects abnormal behavior or a family member requests to view the footage, staff may need to temporarily switch to high-definition mode to confirm the situation; after the visit, the privacy mode should be quickly restored. However, existing systems lack a switching mechanism directly linked to on-site operations. Caregivers cannot quickly confirm or change the current mode before entering the room, and remote operation suffers from response delays and complex access control issues. More importantly, if cameras can still be remotely activated for high-definition recording by others via the network while caregivers provide private services such as bathing and changing clothes, it poses a potential threat to the professional dignity of service personnel.
[0007] In real-world applications involving high-density care, high privacy sensitivity, and stringent response times, existing camera technologies fail to ensure privacy protection measures cannot be remotely bypassed, provide stable and reliable behavioral visualization information, and lack the ability to switch modes synchronously with on-site operations. Therefore, it is necessary to develop a novel camera system whose mode switching is directly triggered by a physical button on the device itself. This button's state cannot be modified via software or network commands, and it can generate desensitized video output locally in real time, consistent with the current environment, to meet the practical needs of the aforementioned scenarios. Summary of the Invention
[0008] This invention provides a dual-mode privacy-protecting intelligent visual terminal and visualization method based on physical switches and artificial intelligence, aiming to solve the technical problems of unreliable privacy protection, lack of behavior recognition ability in desensitized images, and disconnect between mode switching and on-site operation caused by the reliance on software control in existing video surveillance equipment in highly privacy-sensitive scenarios.
[0009] The intelligent vision terminal includes an image acquisition module, a physical switching button, an artificial intelligence processing unit, a background capture module, a video synthesis module, and a video output module. The physical switching button is the core hardware component for achieving trusted privacy protection in this invention. This button is a physical mechanical switch or capacitive touch switch located on the terminal casing. Its electrical output is directly connected to a dedicated input pin of the main control chip, and the signal path does not pass through the operating system kernel, application layer, or network service middleware. When the user presses the button, the hardware circuit immediately generates a stable high / low level signal, which is read by the main control chip in real time as a mode status signal. The state of this signal directly determines the selection logic of the video output path and cannot be modified, overwritten, or simulated through any remote network commands, mobile applications, cloud configuration interfaces, or local software commands.
[0010] In high-definition mode, the original high-definition video stream output by the image acquisition module is compressed by the video encoder and then directly transmitted to external devices via PoE, Wi-Fi, or 4G / 5G interfaces by the video output module for remote viewing or temporary diagnosis and treatment.
[0011] In privacy mode, once the physical switch button is triggered, the main control chip immediately initiates the background capture process. Background image acquisition supports two complementary mechanisms. The first is an automatic detection method: the system continuously analyzes the video stream using a human detection model built into the AI processing unit. When no valid human target is detected for several consecutive frames, the system determines the current environment is empty and automatically captures a frame as a static background image. The second is a manual assistance method: suitable for scenarios where operators actively trigger the switch after confirming the room is empty. In this mode, after the user presses the physical switch button, the system does not immediately capture the background but starts a configurable delay timer, defaulting to five seconds, allowing the operator to calmly leave the monitoring area. After the timer expires, the system automatically captures the current image as a new background image. Users can adjust this delay duration through the accompanying application to adapt to different room sizes or mobility requirements. Regardless of the method used, the captured background image is stored in uncompressed RGB or YUV format in local non-volatile memory, preserving complete spatial resolution and color information for subsequent compositing.
[0012] Subsequently, the AI processing unit loads a lightweight human pose estimation model to process each subsequent frame of video in real time, extracting multiple joints of at least one human body in the frame. Each joint contains two-dimensional pixel coordinates and a confidence value. The video compositing module generates corresponding visualizations based on these joints. These visualizations are first constructed in an independent rendering buffer: if the skeleton line graph mode is selected, connecting lines are drawn with preset colors, and circular markers are drawn at the joint positions; if the cartoon humanoid mode is selected, a pre-stored two-dimensional character template is driven according to the joint proportions to generate a simplified humanoid outline with a head, torso, and limbs; in addition, the system also supports various abstract representations such as simplified outlines, 3D mesh projection, motion heatmaps, particle systems, or symbolic icons.
[0013] Crucially, the visualization is not simply overlaid on the original video; instead, it is composited with the aforementioned static background image using pixel-level layer compositing. The system uses the background image as the bottom layer and the visualization as the top layer, employing an alpha blending algorithm for compositing. In the top layer, the alpha channel value for non-graphic areas is set to zero, while the alpha value for graphic areas is set to one, or a gradient value between zero and one depending on edge anti-aliasing requirements. This process is completed in a GPU or dedicated image signal processor, ensuring real-time compositing performance of at least fifteen frames per second.
[0014] Through the aforementioned mechanism, the visualized human body is precisely embedded into the real-world room environment. For example, the foot joints are aligned with the ground, and the buttocks align with the bed surface when seated. This effectively avoids distortions such as the human body appearing to float in the air or pass through furniture, which can occur due to outdated backgrounds or neglecting perspective. The composite result preserves the spatial context of the environment, enabling caregivers to accurately determine the location and posture of individuals.
[0015] The synthesized video frames are fed into a video encoder and compressed into MP4 or other common container formats according to standard H.264 or H.265 encoding protocols. This video stream contains only anonymized visual content and static backgrounds, and does not contain any original human images. The video output module cuts off the output channel of the original high-definition video stream, and only outputs this privacy-preserving video stream to the NVR, cloud platform, or user terminal through the same communication interface. The entire processing is completed on the device side. The original video frames are released from memory after the artificial intelligence inference and synthesis are completed, and are not written to the memory card or transmitted through any network channel.
[0016] Furthermore, the terminal also includes a status indicator, such as a dual-color LED light, whose color is directly driven by the mode status signal generated by the physical buttons, realizing local visual feedback of the mode status without relying on software updates. The artificial intelligence processing unit is also configured to perform fall detection or people counting in privacy mode. All analysis is completed based on keypoint coordinates, and the results are output in the form of text alerts or structured data, without image content, and can be used to link emergency calls or generate care logs.
[0017] In some embodiments, the terminal also includes a physical lens shielding mechanism driven by a micro-motor. Upon receiving a privacy mode signal, the mechanism automatically slides a shielding plate to cover the optical lens, leaving only the infrared window open for motion sensing in low-light environments. This shielding action is achieved by an independent power supply or mechanical linkage, without relying on the main control system software control. Even if the device firmware is tampered with, it can still ensure that optical imaging is physically blocked.
[0018] This invention also provides a corresponding video output control method, the core of which lies in the strict binding of video output content to the hardware state of physical switching buttons. The method includes: generating a mode state signal in response to physical button operation; selecting, based on the signal, to output either the original high-definition video stream or a privacy video stream synthesized from a key visualization graphic and a background image; and outputting the selected video stream via wired or wireless means. This method ensures that regardless of the network connection method, the output content always reflects the user's true intention from the most recent physical operation, eliminating software-level pattern deception.
[0019] This invention offers the following advantages: By directly connecting the mode switching logic to the physical button hardware, changes to the privacy protection status must be actively performed by the on-site user, fundamentally eliminating the possibility of remote attacks, account hijacking, or software misconfiguration bypassing the privacy mode. The background image acquisition mechanism supports both automatic detection and manual assistance. In manual assistance mode, the system delays for several seconds after the physical button is triggered before capturing the background image, providing operators with sufficient time to leave the monitored area, effectively avoiding residual human images in the background and improving the reliability of the initial desensitized image. Video synthesis accurately overlays the visualized graphics generated based on human joint points onto the static background image, naturally aligning the skeleton or cartoon human figure with environmental elements such as the ground and bed. This solves the distortion problems caused by background misalignment in existing technologies, such as human figures appearing to float or penetrate furniture, significantly improving the accuracy of behavior judgment. The artificial intelligence processing unit completes posture estimation, fall detection, and people counting on the device side, without leaving the device, reducing network bandwidth consumption and preventing the leakage of sensitive data. The dual-mode output architecture supports the use of anonymized video in routine monitoring and temporary switching to high-definition mode during emergency response or family visits, flexibly adapting to the actual workflows of complex care scenarios such as senior living communities, nursing homes, and hospital wards. The overall solution ensures effective behavioral monitoring while also considering the privacy needs of residents, families, and caregivers, improving the security, usability, and user acceptance of video surveillance systems in highly sensitive areas. Attached Figure Description
[0020] Figure 1 A system structure block diagram of a privacy-protecting intelligent vision terminal provided in an embodiment of the present invention;
[0021] Figure 2 This is a schematic diagram of the hardware connection between the physical switching button and the main control chip in an embodiment of the present invention;
[0022] Figure 3 This is a flowchart of video processing and synthesis in privacy mode in an embodiment of the present invention;
[0023] Figure 4 This is a schematic diagram of the timing of artificially assisted background capture in an embodiment of the present invention;
[0024] Figure 5 This is a schematic diagram illustrating the effect of Alpha blending and synthesis of a visual graphic and a static background image in an embodiment of the present invention.
[0025] Figure 6 Examples of output screens in different visualization forms in embodiments of the present invention include skeleton line drawings, cartoon human figures, and simplified outlines. Detailed Implementation
[0026] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0027] like Figure 1 The diagram shown is a system structure block diagram of a privacy-protecting smart vision terminal provided in an embodiment of the present invention. The smart camera includes an image acquisition device 101, a physical button 102, an edge AI processing unit 103, a background capture module 104, a video synthesizer 105, and an output module 106. Optionally, it may also include a status indicator device 107 and a physical lens masking mechanism 108.
[0028] Image acquisition device 101 is used to acquire real-time video streams and transmit the video data to edge AI processing unit 103. Physical button 102 is located on the camera body and is a physical mechanical or capacitive switch. Its output is directly connected to the main control logic circuit and is used to receive user trigger signals to switch between high-definition mode and privacy mode, and generate corresponding mode status signals. The mode status signals are not processed by the operating system or network protocol stack, and their status cannot be directly modified by remote software commands.
[0029] The edge AI processing unit 103 is integrated inside the camera and is used to perform real-time analysis of the video stream in privacy mode, identify human bodies in the image, and output multiple key points representing human body structure based on a human pose estimation model, thereby generating a visual graphic to replace the real human body; the visual graphic at least partially presents the key points as visual elements in the privacy video stream, and can be represented as a skeleton line drawing, a cartoon human figure, a simplified human body outline, a 3D mesh model, or a motion heat map; in addition, the edge AI processing unit 103 is also configured to perform fall detection or people counting on the device side and generate alarms or structured data, without relying on cloud computing.
[0030] Background capture module 104 is used to capture an environmental image as a static background base map when the device is first started, each time it is switched to privacy mode, or when a significant change in the scene is detected, and to provide the base map to video compositor 105.
[0031] The video synthesizer 105 is used to overlay the visual graphics generated by the edge AI processing unit 103 onto the static background map provided by the background capture module 104 in privacy mode to generate a desensitized privacy video stream; in high-definition mode, the video synthesizer 105 transmits the original high-definition video stream.
[0032] The output module 106 is used to selectively transmit the original high-definition video stream or the private video stream to an external device via PoE, Wi-Fi or 4G / 5G network according to the mode status signal generated by the physical button 102, and the output content is always consistent with the current state of the physical button 102.
[0033] A status indicator 107 (such as a dual-color LED) displays different visual states based on a mode status signal to indicate whether the current mode is HD or privacy mode. A physical lens shielding mechanism 108, in response to the mode status signal, automatically shields the optical lens when privacy mode is activated, retaining only non-visible light sensing capabilities.
[0034] In one possible implementation, the smart camera of the present invention is implemented using an embedded hardware platform. The platform includes a main control chip with an integrated NPU (e.g., Rockchip RK3588), a 2-megapixel CMOS image sensor as an image acquisition device 101, a mechanical push switch located on the top of the camera housing as a physical button 102, and video processing-related edge AI processing unit 103, background capture module 104, video synthesizer 105, and output module 106.
[0035] When the user presses physical button 102, its output pin is directly connected to the GPIO of the main control chip, generating a hardware level signal. This signal is not processed by the operating system or network protocol stack; it is immediately recognized by the firmware as a mode switching instruction, and a corresponding mode status signal is generated. Because this path completely bypasses the software layer, remote attackers cannot modify the current mode through App, cloud, or firmware vulnerabilities, ensuring the authenticity and unbypassability of privacy protection.
[0036] In privacy mode, the background capture module 104 decides whether to update the static background image according to a preset strategy: if it is the first time the device is started, the user is configured to "update every time it switches", or the edge AI processing unit 103 detects that the structural similarity between the current scene and the existing background is lower than the threshold, then a frame of no people is captured from the image acquisition device 101 as a new background image and stored in the local flash memory in YUV format.
[0037] Subsequently, the edge AI processing unit 103 loads a lightweight pose estimation model (such as MoveNet) to analyze the real-time video stream frame by frame, extracting the coordinates of multiple joints of the human body. The video synthesizer 105 generates visual graphics based on these joints—for example, rendering the joints as circular markers and drawing lines according to the skeletal topology relationships such as shoulder-elbow-wrist, hip-knee-ankle, etc., to form a skeletal line drawing; it can also be configured as a cartoon human figure or a simplified outline. This graphic is superimposed onto a static background image using an alpha blending algorithm, with an alpha value of 1 for the foreground area and 0 for the background area, ensuring that the human figure naturally conforms to the ground or furniture, avoiding floating distortion.
[0038] After the synthesized desensitized video stream is encoded with H.264, it is transmitted to an external device by the output module 106 via PoE, Wi-Fi or 4G network. The output content is always strictly consistent with the current state of the physical button 102.
[0039] In addition, the camera may include a status indicator, such as a dual-color LED, whose color is directly driven by the mode status signal: blue indicates high-definition mode and green indicates privacy mode, making it easy for on-site personnel to quickly confirm the current working status.
[0040] The edge AI processing unit 103 can also perform fall detection or people counting on the device: when an abnormal torso tilt angle is detected and the center of gravity height drops suddenly for several seconds, it is determined to be a fall event and a local alarm is generated; the people counting outputs structured data based on the number of valid human detection boxes or head joints, without the need to upload the original video.
[0041] In some embodiments, the camera may also integrate a physical lens shielding mechanism that automatically slides a shielding plate to cover the optical lens in response to a privacy mode signal, leaving only the infrared sensing window open, thereby further enhancing both psychological and physical privacy protection.
[0042] In one possible implementation, the smart camera of the present invention is deployed in the single resident room of a nursing home to provide behavioral monitoring and privacy protection functions during the daily care of elderly people with dementia.
[0043] In the actual operation of elderly care facilities, nursing staff typically conduct manual rounds at fixed intervals. Due to the large number of rooms spread across different floors, completing a full round of rounds takes a considerable amount of time. During this time, it is difficult to detect abnormal situations such as residents falling, unknowingly getting out of bed, or remaining motionless for extended periods. Some facilities have attempted to achieve remote monitoring by installing traditional high-definition cameras, but because these devices continuously record raw video containing facial features and body contours, they can easily trigger resistance when residents are engaged in private activities such as changing clothes, using the toilet, or lying in bed. They also make nursing staff uncomfortable when providing close care, thus limiting the deployment of such devices in private rooms.
[0044] Most existing surveillance products that support privacy mode rely on mobile applications for remote switching. However, this switching mechanism is implemented through software commands, which poses a risk of being bypassed remotely. For example, if the device account password is weak, the firmware has unpatched vulnerabilities, or a family member accidentally grants shared control permissions, unauthorized users may disable privacy mode and restore the original video recording. Furthermore, some products only apply a global blur to the image, which, while hiding details, fails to identify whether a person is active or in a sitting or lying position, resulting in missing monitoring information.
[0045] In this embodiment, the camera is installed in a corner of the room ceiling, covering the area around the bed and the bathroom passageway. Upon initial activation, staff press a physical button on the camera itself. The system initiates a configurable delay (e.g., 5 seconds) and then automatically captures a frame of stillness as a static background image. Subsequently, the device operates in privacy mode by default: the edge AI processing unit analyzes the real-time video stream locally, extracts human key points, and generates a skeleton line drawing; the video synthesizer then overlays this drawing onto the static background image using an alpha blending algorithm, creating a desensitized video stream. This video stream is encoded and transmitted via PoE network to the nursing station monitoring terminal, allowing nursing staff to view the activity status of people in multiple rooms.
[0046] When the edge AI processing unit detects that the human torso tilt angle is below a preset threshold and the center of gravity continues to drop for a certain period of time, the system determines it as a fall event and generates a structured alarm message without image content locally, which is then pushed to the nursing station workstation. After the nursing staff goes to the scene to confirm, if they need to show the current scene to the family, they can press the physical button again to switch the device to high-definition mode and output the original video stream. After the visit ends, the staff can press the physical button again to switch the device back to privacy mode and recapture the background image after a delay according to the configured policy.
[0047] Throughout the process, the original high-definition video was not uploaded to an external server; all AI inference and video synthesis were completed on the device itself. Before entering the room to provide service, caregivers can confirm the current privacy mode via the status indicator light on the camera. Because mode switching is triggered directly by a physical button, and its signal path does not pass through the operating system or network protocol stack, remote users cannot modify the current mode via software commands. While family members can view the high-definition image, this requires on-site personnel to operate the physical buttons.
[0048] In hospital wards, this camera is used for post-operative patient monitoring. Patients need to limit their movement but may try to get out of bed due to discomfort. Through skeletal visualization, medical staff can remotely determine whether the patient is in bed, sitting up, or walking towards the door, while the original human image is effectively masked. During remote consultations with doctors, nurses can temporarily switch to high-definition mode and switch back to privacy mode afterward. The edge AI processing unit can also perform people counting based on the number of valid human detection boxes, assisting in determining the composition of people in the room.
[0049] like Figure 2 The diagram shown illustrates the hardware connection between the physical switch button and the main control chip in an embodiment of the present invention. The diagram includes a physical button 102, a GPIO interface 201, and a main control chip 202.
[0050] The physical button 102 is a mechanical or capacitive physical switch located on the camera body. One end of the button is grounded, and the other end is directly connected to the general purpose input / output pin (GPIO interface 201) of the main control chip 202. The GPIO interface 201 is configured in input mode and its internal pull-up resistor is enabled, so that the pin remains high when the physical button 102 is not pressed; when the button is pressed, the pin is pulled low, thereby generating a stable digital signal as a mode status signal.
[0051] This connection path bypasses the operating system kernel, application layer, network protocol stack, or any intermediate software modules. The mode status signal is directly read by the firmware layer of the main control chip 202 and used to control the video output logic. Because the signal path is entirely located between the hardware and the underlying firmware, external devices cannot modify, overwrite, or simulate this signal through network commands, mobile applications, cloud configuration interfaces, or local software commands, thus ensuring that the activation of the privacy mode truly reflects the physical operating intent of the user on site.
[0052] In one possible implementation, the present invention provides a highly integrated camera module that can be embedded as a core vision unit into various smart terminal devices that require local privacy protection and behavior awareness capabilities.
[0053] The camera module includes an image sensor, a physical switch button, a main control chip, non-volatile memory, a communication interface, and a power management unit. The image sensor uses a 2-megapixel global shutter CMOS device, supporting 1080p video output at 30 frames per second. Its data pins are directly connected to the image signal processor input port of the main control chip. The physical switch button is a surface-mount mechanical tactile switch, soldered onto the module's printed circuit board. Its output is directly connected to a dedicated general-purpose input / output pin of the main control chip. This pin is configured for input mode and uses an internal pull-up resistor, allowing the main control chip to read the button's state in real-time via high and low voltage levels.
[0054] The main control chip uses domestically produced system-on-a-chip (SoC) supporting edge AI computing, such as Rockchip, Ankai Microelectronics, Qingke, Allwinner, or HiSilicon series chips. These chips integrate multi-core ARM processors and dedicated neural network acceleration units, enabling efficient local execution of lightweight human pose estimation models and completion of video synthesis and encoding tasks. Non-volatile memory uses eight GB eMMC or SPI NAND Flash to store static background images, AI model weights, and firmware programs. Communication interfaces include a 100Mbps Ethernet PHY, Wi-Fi 5 or 6 modules, and optional 4G or 5G communication daughterboards, supporting output of processed video streams to external systems via PoE, wireless, or cellular networks.
[0055] This module can be flexibly deployed in various devices. In service robots, such as elderly care robots or hospital guidance robots, the module is installed in the robot's head to recognize user posture, determine falls or interaction intentions, and ensures a privacy mode is switched before entering private areas such as bedrooms via a physical button. In rehabilitation massage robotic arms or nursing assistive robotic arms, the module is integrated into the end effector or base of the robotic arm to capture the user's sitting posture, spinal angle, or limb position in real time to adjust massage intensity or assistive movements; when the user is changing clothes, preparing to use the toilet, etc., the operator can press a physical button to immediately stop the module from outputting raw video and only provide visual data of joint points for the control algorithm to use.
[0056] Furthermore, this module can be integrated into smart mirrors for home health monitoring, providing posture feedback during morning washing or dressing while ensuring facial and body details are not recorded. In smart beds or nursing beds, the module can be embedded in the headboard or side rail to continuously monitor lying posture, turning frequency, or getting out of bed, providing data support for long-term care. In remote consultation terminals or family doctor workstations, the module can be triggered by a physical button by the patient or family member before the doctor initiates a high-definition video call, achieving a dynamic balance between privacy and treatment needs. In assistive systems in public restrooms or accessible facilities, the module can be used to detect falls or prolonged periods of inactivity, but only outputs skeletal images to avoid infringing on the user's dignity.
[0057] Regardless of the device on which it is deployed, the module's mode status is generated directly via a hardware connection using physical buttons, without relying on operating system scheduling, application calls, or network commands. This ensures that the privacy protection mechanism is unbreakable remotely on all types of terminals. The module also reserves LED driver pins and motor control signals, allowing for the connection of external status indicator lights or physical lens shielding mechanisms to further enhance privacy protection capabilities.
[0058] In one possible implementation, the present invention provides an intelligent camera module with a dual-mode one-click switching mechanism. The module operates in HD mode by default, outputting a standard HD video stream, such as 1080P or 4K resolution, containing complete information about the actual scene and people, suitable for conventional security monitoring. In this mode, the video stream is transmitted directly from the image sensor to the encoder and output to external devices, such as an NVR or cloud platform, via a network interface.
[0059] When the user presses the physical button on the device, the system immediately switches to privacy mode. At this time, the built-in AI chip activates and begins processing the real-time video stream. First, the system accurately identifies and extracts human joint points based on a lightweight pose estimation algorithm. These joint points are used to generate various visual alternative display formats, which users can configure and select through the accompanying application.
[0060] Visualization formats include simple skeletal line drawings, where joints are marked with circles and lines are drawn according to a preset human skeletal topology; cartoonish humanoid silhouettes, which simulate real human movements by driving simplified 2D character models; low-poly style 3D mesh modeling of the human body, retaining basic shape and movement characteristics; abstract geometric shapes, which retain only the outer contour of the human body and remove all details; minimalist heatmaps, which only highlight moving areas without showing the specific human shape; and child-friendly images, such as small robots or animal shapes, whose movements follow the changes in real human postures, suitable for home or childcare scenarios.
[0061] The system automatically captures a single frame of the current room with no people in it the moment it switches to Privacy Mode, or when the device is first powered on, as a static background image. All subsequent videos in Privacy Mode will have an AI-generated visualization overlaid on this background image, creating an immersive effect of "people moving in a real environment," while completely concealing real identities, facial features, and body details. The background capture strategy supports multiple configuration options, including capturing only on first power-on, automatically updating each time the device switches to Privacy Mode, or updating weekly and prompting the user to confirm whether to recapture when significant scene changes are detected.
[0062] To enhance user awareness of device status, the device itself is equipped with a dual-color LED status indicator. A solid blue light indicates that it is currently in HD mode and the original video is being output; a solid green light indicates that it is currently in privacy mode and AI-based anonymization processing is in progress. The accompanying application can display the current operating mode in real time and record the historical time and triggering method of all mode switches, improving operational transparency and user trust.
[0063] In addition, the system supports a timed automatic switching function, such as automatically entering privacy mode between 10 PM and 7 AM the next day. It also supports integration with smart home systems; for example, when the system detects that the bedroom door is closed and the lights are off, it automatically triggers privacy mode to achieve proactive, context-aware privacy protection.
[0064] This camera module can be flexibly integrated into various terminal devices. In service robots, such as elderly care companions or hospital guide robots, the module is installed in the head to recognize user posture, assess fall risk, or interaction intent, while ensuring that it has switched to privacy mode before entering a private space. In rehabilitation massage robotic arms or nursing assistive robotic arms, the module is embedded in the end effector or base to capture the user's sitting posture, spinal angle, or limb position in real time, dynamically adjusting massage intensity or assistive movements. In smart mirrors, it can be used for morning health monitoring, providing posture feedback without recording real images. In smart beds or nursing beds, it can be embedded in the headboard structure to continuously monitor lying posture, turning frequency, or getting out of bed. In remote consultation terminals, patients or family members can temporarily activate high-definition mode for doctors to view via a physical button, and switch back to privacy mode with one click afterward. In public restrooms or accessible facilities, it can be used to detect falls or prolonged stays, outputting only skeletal images to avoid infringing on user dignity.
[0065] Regardless of the device on which it is deployed, the module's mode status is generated directly via a hardware connection using physical buttons. The signals are not processed by the operating system, application, or network protocol stack, ensuring that the privacy protection mechanism cannot be bypassed remotely. The module also has reserved LED driver pins and motor control signal outputs, which can be connected to external status indicator lights or physical lens shielding mechanisms to further enhance privacy protection capabilities.
[0066] like Figure 3 The diagram shows a flowchart of video processing and synthesis in privacy mode according to an embodiment of the present invention. The flowchart, starting from the user pressing the physical switch button, fully demonstrates the entire process from background acquisition to desensitized video output.
[0067] When the user presses the physical button on the device, the system immediately enters privacy mode and initiates the background capture process. During this process, if preset conditions are met, such as the device being powered on for the first time, the user configuring the background to update on each switch, or the system detecting a significant difference between the current scene and a stored background, the camera automatically captures a frame of the environment without human activity as a static background image. This background image is stored in uncompressed format in local non-volatile memory for subsequent video compositing.
[0068] Subsequently, the system extracts key human points from the real-time video stream. The edge AI processing unit loads a lightweight human pose estimation model, analyzes each frame of the image, identifies and outputs the coordinates of multiple joints representing human structures, including the head, shoulders, elbows, wrists, hips, knees, and ankles. This process is completed on the device itself, without relying on cloud computing, ensuring low latency and that data remains within the device.
[0069] Based on the extracted key points, the system generates a visual graphic. Depending on the display format pre-selected by the user through the accompanying application, the graphic can be represented as a simple skeleton line drawing, a cartoonish human silhouette, a low-poly 3D mesh human body, an abstract geometric shape, a minimalist heatmap, or a child-friendly image. All graphics are built in an independent render buffer, containing only pose information and no visually identifiable details.
[0070] Next, the system performs an alpha blending operation. The aforementioned visualization is used as the foreground layer, and the static background image as the background layer, and the blending is performed at the pixel level. The opacity of non-graphic areas in the foreground is set to completely transparent, while graphic areas are set to opaque or use edge anti-aliasing gradients. This ensures that the human figure in the composite image naturally integrates into the real environment, avoiding any floating or misaligned appearances.
[0071] Finally, the synthesized, desensitized video frames are fed into a video encoder, compressed into MP4 or other common container formats according to the H.264 or H.265 standard, and then transmitted to an external device via PoE, Wi-Fi, or 4G / 5G networks by the video output module. At this point, the original high-definition video stream has been cut off, and the output content only contains the privacy-preserving visualization results.
[0072] In one possible implementation, the smart camera of the present invention is deployed in a home-based elderly care environment to support a third-party security monitoring platform in providing remote monitoring services to elderly people living alone. This service is executed based on a service agreement signed between the user and the platform, clearly specifying the time periods or scenarios in which high-definition video mode will be enabled, and the circumstances under which only privacy protection mode will be enabled.
[0073] The elderly person lives in the master bedroom and living room of an ordinary urban apartment. Due to limited daily mobility, they face safety risks such as falls, prolonged periods of inactivity, or getting out of bed at night. Their children have signed a service contract with a qualified security monitoring platform, authorizing the platform to receive video data under specific conditions for abnormal behavior identification. According to the agreement, family members can temporarily authorize high-definition mode for remote visits during the day; at other times, the system defaults to privacy protection mode, transmitting only anonymized, visualized video streams to the platform.
[0074] The devices are installed on the ceilings of the living room and bedroom. Upon initial power-on, community staff operate the system: pressing the physical button on the camera triggers a five-second delay before automatically capturing an image of an empty room as a static background. Afterward, the devices operate continuously in privacy mode. An edge AI processing unit analyzes the real-time video stream, extracting the coordinates of seventeen key points on the human body and generating a linear skeleton diagram. This skeleton diagram displays the key points as circular markers and draws lines according to preset skeletal topological relationships such as shoulder-elbow-wrist and hip-knee-ankle. All graphic elements are displayed in a bright green to ensure clear visibility even in low-light conditions. The original human image is completely obscured, containing no information that could identify the individual, such as face, body shape, or clothing.
[0075] When the system detects an abnormal torso tilt angle and a center of gravity height that remains below a threshold for several seconds, it classifies it as a fall event and immediately sends a structured alarm to the safety monitoring platform, including the event type, location, and time-series key coordinates. Platform staff, after confirming the situation using the skeletal video stream, contact the family or community emergency response team according to the protocol. If further verification is needed, the platform can send a request to the family, who can then remotely instruct the elderly person or their caregiver via a mobile application to press a physical button, temporarily switching to high-definition mode. High-definition video is only output after the button is pressed, and its duration is limited by the protocol; it automatically switches back to privacy mode afterward and updates the background image according to the configured policy.
[0076] Throughout the service process, the original high-definition video was not uploaded to the platform server; all AI inference, video synthesis, and encoding were completed locally on the device. The physical button's direct hardware connection design ensures that mode switching must be actively performed by the user on-site; the platform cannot force the activation of high-definition mode via software commands, thus protecting the user's right to control their privacy. The accompanying application records all mode switching times, trigger methods, and service logs for users to review at any time.
[0077] This implementation demonstrates that, in home-based elderly care scenarios, the technical solution of this invention enables third-party monitoring services to continuously acquire posture information that can be used for behavioral judgment without obtaining actual human images. By giving users control over the mode and solidifying it through a physical mechanism, the contradiction between "services needing visual data" and "users refusing to expose private images" in remote monitoring is resolved. Simultaneously, the skeleton visualization reduces the data complexity of the video stream, decreases network bandwidth consumption, and allows stable access to the monitoring service even in low-bandwidth home environments. The security monitoring platform does not need to process highly sensitive raw video, reducing data storage and compliance management costs, and improving service accessibility and operational efficiency.
[0078] like Figure 4The diagram illustrates the specific process of manually assisted background capture. First, after the user presses button 401, the system enters a preparation state. This button serves as the key point for triggering the privacy mode, marking the start of the background capture process.
[0079] Next is the "Delay N seconds 402" stage, a user-defined time interval. This delay mechanism provides the user with sufficient time to leave the camera's field of view, ensuring that no one is accidentally captured in the subsequent background capture steps. By default, this delay is set to five seconds, but it can be adjusted to suit different scenario requirements.
[0080] After the preset delay time expires, the "Capture Background 403" step will automatically start. At this time, the system will perform an image acquisition operation to obtain a static image from the current viewpoint as the background image. This background image will be used in subsequent video stream processing, especially in privacy protection mode, to replace the actual captured footage in order to protect user privacy.
[0081] Throughout the process, each step is closely interconnected, ensuring a high degree of automation and precise control from user input to final background capture. Furthermore, this design considers flexibility and user experience, enabling effective operation even in complex home environments, thereby enhancing the system's applicability and reliability. In this way, this embodiment provides a simple and effective solution to the technical challenges of background capture while ensuring user privacy and security.
[0082] In one possible implementation, the smart camera of the present invention is deployed in a home-based elderly care environment to support collaborative management of video recording and privacy protection when professional caregivers provide care services at home.
[0083] Caregivers enter the elderly person's residence at the scheduled time to provide close care services, including washing hair, trimming nails, assisting with bathing, or treating bedsores. According to the service agreement and operating procedures, some non-sensitive procedures (such as consultations and medication reminders) can be recorded on video in high-definition mode for service quality review or remote medical collaboration; while procedures involving bodily exposure or private operations (such as bathing, wound cleaning, and changing urinary catheters) must be switched to privacy mode to ensure that raw human images are not collected or transmitted.
[0084] Before initiating sensitive services, the caregiver presses a physical toggle button mounted on the wall or the device itself. This button is a physical mechanical switch whose signal is directly connected to a dedicated input pin of the main control chip, bypassing the operating system or network layer. Upon button activation, the system immediately starts a configurable delay timer, defaulting to five seconds, allowing the caregiver to gracefully exit the camera's field of view. During this time, the camera continues to capture images but does not crop the background, preventing the caregiver from being mistakenly recorded as part of the environment.
[0085] After the delay ends, the system automatically captures the current frame with no human activity as a static background image and officially enters privacy mode. Subsequently, all video streams are analyzed in real time by the edge AI processing unit, extracting the elderly person's key points and generating a highlighted green linear skeleton image overlaid on the background. The actual human body, face, wound areas, and details of caregiving actions are completely obscured, retaining only posture and position information for the remote monitoring platform to determine if any abnormal events such as falls or prolonged periods of stillness have occurred.
[0086] After the service is completed, caregivers can press the physical button again to temporarily switch back to HD mode for recording the healing of local wounds (with the elderly person's authorization) or to confirm the service. This HD footage is only briefly output locally and requires on-site operation; it cannot be forcibly activated by a remote platform. All mode switching records are synchronized to the accompanying application, forming an auditable operation log.
[0087] This implementation addresses the long-standing privacy and regulatory conflict in home-based care services: on the one hand, service providers need to retain evidence of the service process to ensure quality and accountability; on the other hand, the elderly and their families strongly object to their private moments being recorded on video. By introducing a delayed background capture mechanism triggered by physical buttons and local AI desensitization processing, this invention ensures that during sensitive services, the system only outputs irreversible, abstracted posture data, meeting basic behavioral monitoring needs while completely avoiding the risk of identity and body detail leakage. Simultaneously, the delayed mechanism effectively prevents service personnel from being mistakenly saved in the background image due to not leaving promptly after the operation, avoiding "ghost figures" or background misalignment in subsequent composite images, thus improving the accuracy and credibility of the desensitized video. The entire process does not rely on network connections or cloud intervention and can still operate normally in offline environments, making it suitable for various urban and rural home care scenarios.
[0088] In one possible implementation, the technical solution of the present invention is integrated into a portable service recorder, which is worn by community caregivers, rehabilitation therapists or domestic service personnel for recording the process and managing privacy compliance during the provision of home care services.
[0089] Service recorders are typically mounted on the caregiver's chest or shoulder, with the camera facing forward, covering the service area. In non-sensitive service scenarios, such as health inquiries, medication guidance, and environmental tidying, the device operates in high-definition mode, outputting a 1080P video stream containing the entire scene and the person, used for service quality assessment, dispute evidence collection, or remote expert collaboration. However, when services involve high-privacy scenarios such as dressing assistance, skin examination, bathing assistance, pressure ulcer care, or toileting assistance, it must immediately switch to privacy protection mode to prevent raw human images from being recorded or uploaded.
[0090] To this end, the service recorder features a physical switching button. Before entering a private operation, the caregiver manually presses this button. The button signal is transmitted directly to the main control chip via hardware connection, bypassing the operating system's intermediate layer, ensuring that the switching command cannot be intercepted by software or remotely simulated. Upon triggering, the device initiates a configurable delay, defaulting to five seconds. During this period, the device continues to capture footage but does not perform background cropping. This delay provides the caregiver with ample time to adjust their position or temporarily exit the camera's field of view, preventing their image from being mistakenly captured as part of the background.
[0091] After the delay ends, the system automatically captures the current frame with no human activity as a static background image. Then, the device switches to privacy mode: the edge AI processing unit performs real-time human pose estimation on each subsequent frame of video, extracts keypoint coordinates, and generates a highlighted green linear skeleton map. This map is then overlaid onto the static background using an alpha blending algorithm, forming a desensitized video stream. Real human contours, facial features, wound details, and clothing conditions are completely obscured, retaining only limb position and movement trends. This is sufficient to support fall detection or action compliance assessment, but it cannot reconstruct identity or physiological details.
[0092] After the service is completed, the caregiver can press the physical button again to temporarily restore the high-definition mode, which can be used to take close-up shots of authorized areas (such as wound healing) and upload them to the institution's platform immediately. All mode switching is initiated by on-site personnel; the platform cannot remotely force high-definition recording, ensuring the user's control over privacy boundaries. The accompanying application synchronously records the time, duration, and geographical location of each switch, forming a complete audit trail.
[0093] This implementation addresses the core contradiction of mobile service recording devices in real-world in-home scenarios: it must meet the operational needs of traceable and monitorable service processes while respecting residents' dignity and privacy rights in private spaces. By combining physical buttons, delayed background capture, and local AI desensitization, this invention ensures compliant recording—"behavior visible, identity invisible"—even in dynamic, unstructured home environments. Furthermore, because all processing is completed on the device itself, requiring no continuous internet connection, it is suitable for typical scenarios with weak network signals or where elderly residents refuse to share their home Wi-Fi, significantly improving the usability and acceptance of service recorders in grassroots elderly care service systems.
[0094] As attached Figure 5 As shown in the figure, this diagram illustrates the effect of the Alpha blending technology used for privacy-preserving video synthesis in this invention. This technology can be applied to various product forms described in this invention, including fixed smart cameras, portable service recorders, service robots, or auxiliary robotic arms—terminal devices with local video processing capabilities.
[0095] The image contains three key image elements: the original background image 501, the skeleton / cartoon human figure image 502, and the final composite result 503. The original background image 501 is an environmental image captured by the system when switching to privacy mode or when the device is first started; its content does not contain any human activity. This image can be an unprocessed color image or a pre-processed image, such as a grayscale image, a sketch-style image, an edge-enhanced image, or a semantic segmentation image. This processing aims to preserve spatial structure and key environmental features (such as beds, chairs, door frames, etc.) while further reducing visually sensitive information in the background, thereby improving the ability to identify scene layouts while ensuring privacy.
[0096] The skeleton / cartoon human figure image 502 is generated by the edge AI processing unit based on a real-time video stream. Joints are extracted using a lightweight human pose estimation algorithm and connected by highlighted green lines to form a linear skeleton, or rendered as an abstract form such as a cartoon human figure or a low-poly model. This layer only contains pose and position information and does not include realistic human appearance, facial or clothing details.
[0097] The final composite result 503 uses alpha blending technology to overlay the skeleton / cartoon human figure 502 onto the original background image 501. This compositing method is based on the principle of pixel-level transparency control: for each pixel location, the system determines its visibility in the final image based on whether the skeleton graphic covers that location. In the skeleton graphic area, the background is completely masked, displaying only the highlighted visual human structure; in non-skeleton areas, the background image is fully presented. This processing method ensures that the composite image retains the spatial sense of the real environment while completely desensitizing human details.
[0098] In one possible implementation, the technical solution of the present invention is applied to a home-based elderly care scenario in urban residences. In this scenario, an elderly person lives alone in a two-bedroom apartment, and their daily activities are mainly concentrated in the bedroom, living room, and bathroom. Their children live in another area of the same city and want to use technology to remotely monitor whether the elderly person has experienced any abnormalities such as falls, prolonged periods of inactivity, or unconsciously getting out of bed at night. However, the elderly person explicitly refuses to have a camera installed that records real human images, and is particularly concerned about having their private moments such as using the toilet, changing clothes, and lying in bed filmed and stored.
[0099] To address this practical problem, the system employs a localized processing architecture, completing all video analysis and compositing operations on the device itself. Upon initial power-on, a family member or community worker presses a physical button on the camera to trigger the background capture process. After a five-second delay, the system automatically captures a single frame of unoccupied footage as the original background image 501. This image can then be processed in different ways depending on the environment. In public areas such as living rooms, the system retains the original color format to maintain the color and position recognition of furniture such as sofas, coffee tables, and TV cabinets. In bedrooms or bathroom entrances, the background is converted to grayscale, and edge detection algorithms are used to enhance structural features such as wall outlines, door frames, and floor seams, while removing unnecessary details such as curtain patterns, bedding textures, and toiletries labels. This processing method ensures that the background neither contains personal information nor obscures spatial layout, providing a reliable reference for subsequent posture localization.
[0100] When an elderly person enters the monitored area, the edge AI processing unit runs a lightweight pose estimation algorithm in real time, extracting the coordinates of multiple joints of the human body from each frame of video, including the head, shoulders, elbows, wrists, hips, knees, and ankles. Based on these coordinates, the system generates a skeleton / cartoon human figure 502. This figure is not a fixed shape but supports multiple visualization forms. For example, in the default settings, the system draws a linear skeleton diagram, marking the joints with highlighted green circles and drawing white lines according to the preset human skeletal connection relationships; if there are children living in the user's home, the system can switch to cartoon human figure mode, driving a simplified two-dimensional character model whose limb movements are synchronized with the real human body; in rehabilitation monitoring scenarios that require judging sitting stability or trunk tilt angle, the system can enable low-polygon 3D mesh modeling, inferring relative depth through monocular vision, and generating a human body structure with a sense of front-to-back hierarchy; for basic monitoring needs that only require confirming the presence of someone, an abstract geometric shape can be selected, simplifying the human body into an elliptical cylinder, displaying only the centroid position and orientation.
[0101] All the above graphics are overlaid onto the processed background using alpha blending technology. During the compositing process, the pixel area containing the skeleton graphics completely covers the background content, while the remaining areas fully expose the background image, ensuring that in the final output, the human body is presented only in an abstract form, while the room environment remains continuous and natural. The original high-definition video frames are immediately released from memory after keypoint extraction is completed, without being written to flash memory or transmitted over a network.
[0102] This implementation effectively addresses two key obstacles of traditional monitoring in home care scenarios. Firstly, because real human images are never recorded or uploaded, elderly people are significantly more accepting of the device and willing to keep the monitoring function running for extended periods. Secondly, caregivers or family members can accurately determine whether the elderly person is sitting, lying down, standing, or in a fall position using anonymized video, avoiding misjudgments due to blurry or black screen images. For example, when the system detects a sudden drop in hip joint height and a trunk tilt angle of less than 30 degrees for several seconds, it can generate a local alarm, indicating a possible fall. At this time, family members can request someone on-site (such as a resident caregiver) to press a physical button to temporarily switch to high-definition mode to verify the situation, and then switch back to privacy mode afterward.
[0103] The entire process runs automatically under the control of the device firmware, without relying on an external network connection. Even if home Wi-Fi is interrupted, the system can still complete posture recognition and video synthesis locally and send structured alarms via the 4G module. Background processing strategies and skeleton representation can be configured through the accompanying application, enabling the same hardware platform to adapt to various application scenarios, from daily activity monitoring and post-operative rehabilitation tracking to accessible bathroom assistance. This design not only enhances the device's versatility but also ensures that users always retain control over the form of video content in environments with varying levels of privacy sensitivity.
[0104] As shown in the attached diagram, Figure 6 Examples of three different visualization formats for the system's output are shown: skeletal line drawing 601, cartoon human figure 602, and simplified outline 603. Each format is accompanied by a corresponding small application scenario illustration to help readers understand its purpose more intuitively.
[0105] In the skeletal line drawing 601, the human body is abstracted as a series of connecting points and lines. The points represent the locations of major joints, such as the head, shoulders, and elbows, while the lines represent the connections between bones. This format is ideal for situations requiring precise analysis of human posture, such as during rehabilitation training, where healthcare professionals can assess a patient's recovery by observing changes in the positions of these key points. A small application scenario illustration shows an elderly person performing a simple arm-raising exercise; the adjacent skeletal line drawing clearly reflects the movement of each joint during this action.
[0106] For the cartoon humanoid 602, we present a more vivid and easily understood human model that can be flexibly adjusted to resemble either a caregiver or a patient, depending on the target audience. When the focus is on a caregiver, a specific cartoon image of the caregiver is displayed. This helps to clarify role identification in interactive scenarios; for example, when guiding a patient through rehabilitation training, the caregiver image can be used to demonstrate the correct movement steps. A corresponding illustration shows a caregiver guiding a patient through rehabilitation training; through the caregiver's cartoon image, users can clearly see how to perform the movements correctly. Conversely, if the target audience is a patient, a corresponding cartoon image of the patient is displayed to help healthcare professionals or family members better monitor the patient's daily activities and ensure their safety and health.
[0107] Finally, there's the simplified outline 603, a more abstract representation that uses simple geometric shapes to depict the human body and its general posture. This method is particularly suitable for situations where it's only necessary to know if someone is in a certain area or to monitor basic activities. For example, in nighttime monitoring mode, it's sufficient to know if the elderly person has gotten out of bed, without needing to know their specific body posture. The corresponding illustration depicts a bedroom scene at night, where the simplified outline effectively conveys the information of whether the elderly person has gotten out of bed.
[0108] It should be noted that the above three forms are merely specific examples of the visualization solutions supported by this invention, and not limitations on the technical scope. The core of this invention lies in: based on extracted human joint point data, driving a configurable graphics rendering engine to generate any abstract human representation that meets privacy protection requirements. The form, style, color, role identity, and level of detail of this graphic representation can be customized through a user interface or configuration file, including but not limited to linear skeletons, cartoon characters, low-poly 3D models, heatmaps, geometric projections, symbolic icons, etc. As long as its generation logic relies on locally extracted joint point coordinates and is used to replace the original human image for desensitized display, it falls within the technical concept scope of this invention. This design ensures the openness and scalability of the technical solution, while preventing patent protection from being circumvented by simply replacing the visual appearance.
[0109] In one possible implementation, the technical solution of the present invention is deployed in a pilot home care residence managed by an urban community elderly care service center. This residence houses a registered elderly person with disabilities who has limited mobility due to stroke sequelae and requires a contracted caregiver to provide bathing assistance, dressing changes, and limb rehabilitation training services three times a week. Traditional high-definition cameras were already installed in the residence for remote monitoring; however, due to repeated instances where the original images were automatically recorded and uploaded to the cloud during care, the elderly person experienced significant anxiety, and the family repeatedly complained and requested the equipment be deactivated, resulting in a long-term lack of security monitoring.
[0110] To rebuild trust and restore effective monitoring, the service center replaced the device with the intelligent visual terminal described in this invention. The device is installed in a corner of the bedroom ceiling, covering the bed, walker, and bathroom entrance area. Upon initial use, community staff operate the device on-site: pressing the physical button on the device initiates a five-second delay. After the person leaves the area, a frame of unoccupied footage is automatically captured as the original background image. This image is then converted to a grayscale edge-enhanced form, preserving spatial structural features such as wall outlines, bed edges, and floor seams, while removing unnecessary details such as curtain patterns and clothing placement, resulting in a desensitized but still recognizable environmental base map.
[0111] During subsequent services, before the caregiver enters the room to assist the elderly person with bathing, they first press a physical button to trigger privacy mode. The system immediately stops the original video stream output and initiates the local AI processing flow. The edge computing unit analyzes the real-time footage, extracting the individual body joints of both the elderly person and the caregiver. Based on the identity configuration information, the system generates two cartoon avatars with different styles: the elderly person is presented in a light blue elderly image, and the caregiver is presented in a white uniform image. Both actions are driven by their respective joints, synchronously reflecting their real posture. The synthesized video stream is transmitted to the service center monitoring platform via home Wi-Fi. The on-duty nurse can clearly see key action sequences such as "the caregiver is helping the elderly person stand up from the bedside" and "the two move towards the bathroom simultaneously," but cannot obtain any facial, body shape, or exposed body information.
[0112] During one service session, the system detected an abnormal drop in the height of the elderly person's right knee joint and a trunk tilt angle that remained below the threshold for six seconds, classifying it as a fall risk event and immediately displaying an alarm on the platform. The on-duty staff confirmed, using a cartoon-style image, that the elderly person had indeed fallen to the ground, then contacted the on-site caregiver by phone to verify the information and coordinated with the nearest emergency response team to be on standby. The entire process did not involve access to any original video footage, ensuring the elderly person's privacy was not violated, and the response efficiency was significantly better than the previous model relying on manual inspections.
[0113] This implementation addresses three core issues in home-based care scenarios. First, a localized processing mechanism triggered by physical buttons ensures that original videos are not recorded or uploaded during sensitive services, eliminating user concerns about privacy leaks and allowing the device to remain operational. Second, visually distinguishable human figures with structure-preserving background processing enables remote monitoring parties to accurately understand multi-person interactions, avoiding misjudgments due to overly abstract visuals. Third, all processing is completed on the device itself; even if the home network is interrupted, the system can still perform fall detection locally and store alarm logs, synchronizing them once the network is restored to ensure service continuity.
[0114] The practical value of this lies in the fact that service providers, without increasing labor costs, achieve full traceability, intervention, and auditability of high-risk service processes, while obtaining long-term authorization from the elderly and their families. This technological solution does not rely on specific role models or fixed graphic styles; its visualization format can be flexibly adjusted according to the service recipients, cultural habits, or cognitive needs. It is applicable to various types of in-home services, from basic daily care to professional rehabilitation training, providing a reliable technological path for the large-scale promotion of smart home care.
[0115] In one possible implementation, the technical solution of the present invention is deployed in the neurosurgery ward of a tertiary hospital for the safe monitoring of postoperative patients. Patients admitted to this ward are mostly those who have undergone craniocerebral or spinal surgery, and are required by doctors to strictly remain in bed for 48 to 72 hours, prohibiting them from getting out of bed on their own to prevent complications such as intracranial pressure fluctuations, wound dehiscence, or falls. However, some patients, due to pain, anxiety, or cognitive impairment, may still attempt to get out of bed at night to use the toilet or walk around, posing a high safety risk.
[0116] Traditional high-definition surveillance cameras are often rejected by patients and their families because they involve continuous recording of patients' private moments, such as when using the toilet, changing dressings, or lying in bed. Even when installed, caregivers feel uncomfortable providing close care, limiting the use of the equipment. While existing desensitization solutions such as blurring or blacking out the screen obscure details, they cannot determine whether the patient is still in bed, whether they are sitting up, or whether they are walking towards the door, thus rendering the monitoring meaningless.
[0117] This embodiment utilizes the intelligent vision terminal described in this invention. The device is fixedly installed on the ceiling directly above the hospital bed, with the lens vertically covering the entire bed and the bedside walking area. Upon initial activation, the responsible nurse presses a physical button on the device, and the system automatically captures a frame of an empty bed as the original background image after a five-second delay. This image is processed by an algorithm into a grayscale edge-enhanced form, preserving key spatial features such as the bed rails, IV stand base, and floor seams, while removing unnecessary information such as bedding patterns and personal belongings.
[0118] During routine monitoring, the device operates in privacy mode by default. The edge AI processing unit analyzes the video stream in real time, extracts the patient's key points, and generates a bright green linear skeleton map, which is overlaid on the processed background. The nurse station monitoring screen can simultaneously display anonymized images from multiple wards, allowing on-duty nurses to clearly identify states such as "patient lying supine," "semi-sitting," and "attempting to get up." When the system detects that the hip joint height exceeds the bed surface threshold and continues to move, it is determined to be an act of getting out of bed, and an alarm notification immediately pops up at the nurse station workstation.
[0119] When the attending physician needs a remote consultation, the responsible nurse enters the ward and presses a physical button. The device temporarily switches to high-definition mode, outputting a raw 1080P video stream for the physician to view wound dressings, limb swelling, and other conditions via an authorized terminal. After the consultation, the nurse presses the physical button again, and the device switches back to privacy mode. Because it is configured to "update background on each switch," it recaptures the current environment after a delay to use as a new background image, ensuring that the subsequent composite image matches the actual layout.
[0120] Throughout the process, the original high-definition video was not written to the device's storage medium, nor uploaded to the hospital's cloud platform or third-party servers. All AI inference, video synthesis, and encoding were completed locally on the device. The signal paths of the physical buttons are directly connected to the GPIO pins of the main control chip, without going through the operating system or network protocol stack, ensuring that remote users cannot force the high-definition mode to be activated via software commands.
[0121] This implementation addresses three key technical challenges in hospital ward scenarios. First, it enables continuous and effective monitoring of high-risk behaviors while protecting patient privacy and dignity. Second, it allows for dynamic balance between temporary medical needs and long-term privacy protection by enabling on-site personnel to control the high-definition mode via physical buttons. Third, the automatic background update mechanism adapts to the frequent changes in the ward environment (such as adjusting bed positions or adding / removing equipment), preventing distortion of the composite image due to background misalignment.
[0122] The resulting practical effects are: nursing teams can intervene in abnormal behaviors promptly without infringing on patient privacy, reducing the incidence of falls and complications; doctors gain the ability to access high-definition images on demand, improving the quality of remote diagnosis and treatment; and patients and their families have significantly increased acceptance of the device because they have more control over mode switching. This technical solution does not rely on a specific visualization format; its skeletal diagrams, cartoon figures, or simplified outlines can all be configured according to departmental needs, making it suitable for various postoperative monitoring scenarios, from neurosurgery and orthopedics to geriatrics.
[0123] In one possible implementation, the technical solution of the present invention is applied to the public activity hall of an urban community day care center. The center receives approximately fifteen to twenty elderly people daily, providing day care, rehabilitation training, recreational activities, and meals. Due to the high density of people, diverse activities, and the fact that some elderly individuals have mild cognitive impairment or mobility issues, the center needs to monitor for abnormal events such as falls, getting lost, sudden discomfort, or interpersonal conflicts in real time. However, because multiple people are involved in the same frame simultaneously, and some elderly individuals are extremely sensitive to continuous video recording, traditional high-definition monitoring systems are difficult to deploy comprehensively.
[0124] To ensure group safety while respecting individual privacy, the center installed a wide-angle intelligent vision terminal on the ceiling of the activity hall. When the device was first activated, staff pressed a physical button, and after a delay, the system captured a frame of empty space as the original background image. This image was then processed by an algorithm into a sketch style, highlighting key facilities such as the layout of tables and chairs, the direction of passageways, the location of entrances and exits, and the emergency call button. Non-structural details such as wall decorations, personal water cups, and clothing were removed, resulting in a clear but anonymized spatial base map.
[0125] During daily operation, the device defaults to privacy mode. The edge AI processing unit performs multi-target human detection and tracking on the wide-angle video stream, extracting a unique sequence of human key points for each elderly person present. Based on pre-recorded identification (such as wristband ID or check-in order), the system assigns a unique visual identifier to each elderly person: for example, elderly person number 1 is represented by a red skeleton line drawing, elderly person number 2 by a blue cartoon human figure, and elderly person number 3 by a green simplified outline. All graphics are superimposed on the same background, independent of each other and without interference.
[0126] When an elderly person exhibits abnormal posture (such as a sudden drop in torso angle or prolonged immobility) or enters a restricted area (such as approaching an exit gate alone), the system immediately generates a structured alarm, pushes it to the tablet terminal of the front desk staff, and highlights the corresponding visual graphic on the monitoring screen. If multiple elderly people gather in the same area for more than a preset time, the system can determine it as a potential conflict or group activity, triggering different levels of alerts. Throughout the process, the original video is not recorded or uploaded; only anonymized multi-target composite footage is output.
[0127] This implementation effectively resolves a core contradiction in day care center scenarios. On one hand, through multi-target identity differentiation and diverse visualization methods, staff can accurately identify which elderly person is exhibiting abnormal behavior, avoiding the blind spot of "seeing the action but not knowing who it is." On the other hand, because all individuals are presented in an abstract form and the background does not contain sensitive information, even if family members temporarily view the footage, there will be no disputes arising from the exposure of others' images. Furthermore, the system supports dynamic addition and removal of personnel without reconfiguration, adapting to the daily staff turnover characteristics of day care centers.
[0128] The practical effects of this are that the center achieves refined monitoring of group activities without increasing the frequency of manual inspections; elderly people are more willing to participate in group activities because they confirm that their image has been desensitized; and staff can quickly locate risk sources through a visual interface, improving emergency response efficiency. This technical solution is not limited to specific graphic styles or identity binding methods. As long as a distinguishable abstract human figure is generated based on locally extracted key points and superimposed on the processed background, it falls within the scope of this invention and is applicable to various care environments with multiple people, such as public areas of nursing homes, community activity stations, and elderly canteens.
[0129] In one possible implementation, the technical solution of the present invention is integrated into a portable service recorder, which is worn by community rehabilitation therapists when providing home rehabilitation services. The therapist visits five to eight different families daily, serving patients with stroke sequelae, elderly people after joint replacement surgery, and Parkinson's disease patients, among others. Services include gait training, joint range of motion assessment, balance testing, and other procedures requiring precise observation of limb movements.
[0130] Because the environments at each service location vary greatly—some homes are dimly lit, some have densely packed furniture, and some have highly reflective walls—it is impossible to pre-set a fixed background. Furthermore, users are highly sensitive to the continuous recording of raw video in private spaces and have repeatedly refused to allow therapists to wear traditional recording devices. To address this issue, this implementation method uses a portable terminal that supports dynamic background capture and optional background compositing.
[0131] The service recorder is fixed to the therapist's chest with the lens facing forward and contains the camera module described in this invention. Before starting service at a new client's home, the therapist presses the physical button on the side of the device to trigger the privacy mode activation process. The system then initiates a configurable delay, which is four seconds by default. During this time, the therapist quickly retreats to a corner of the room or temporarily turns their back to the camera to ensure they are not captured in the background. After the delay ends, the device automatically captures the current frame of an empty room as the original background image 501.
[0132] However, not all service scenarios require a background display. For example, when performing standardized joint range of motion measurements, therapists only need to focus on the flexion and extension angles of the patient's elbow or knee. In this case, the system can be configured to "no background mode": outputting only a bright green linear skeleton diagram 601 overlaid on a pure black or pure gray background, completely omitting environmental information. In this mode, the computational load is lower, and the image focuses more on the movement itself, making it suitable for monitoring localized movements that are independent of spatial location.
[0133] In scenarios requiring assessment of the overall posture and its relationship to the environment (such as evaluating whether a patient can stand up steadily from the bedside or whether they are close to obstacles), a "background blending mode" is activated: the original background image is processed by grayscale and edge enhancement, serving as a static base image, upon which the skeleton graphic is superimposed using alpha blending. To address slight vibrations caused by device movement, the system introduces a background stabilization algorithm based on optical flow to compensate for the skeleton's position in real time, ensuring accurate spatial relationships between the human figure and the ground and furniture.
[0134] The edge AI processing unit dynamically selects the visualization format based on the service type. Gait analysis uses a linear skeleton graph 601 to preserve geometric accuracy; upper limb training can switch to a cartoon humanoid 602 to enhance approachability; and basic presence detection uses a simplified outline 603 to save resources. Regardless of whether a background is enabled, all graphics are generated driven by locally extracted joints, and the original high-definition video frames are immediately released from memory after inference is completed without being written to storage media.
[0135] The use of a background is not mandatory but a configurable option based on task requirements. The system supports both "immersive compositing with a background" and "pure pose output without a background," with both modes sharing the same set of joint extraction, graphics rendering, and physical button control mechanisms. This design not only adapts to diverse needs ranging from fine-grained motion evaluation to coarse-grained behavior monitoring but also ensures the integrity of the technical solution—any implementation that only outputs an anonymized humanoid figure (regardless of whether it contains a background) and relies on local AI and physical switching falls within the protection scope of this invention.
[0136] In one possible implementation, the technical solution of the present invention is deployed in an accessible public restroom managed by a community elderly care service center to monitor the fall risk of elderly or mobility-impaired users. This restroom is designed specifically for disabled or semi-disabled elderly people and is equipped with handrails, non-slip floor tiles, and an emergency call button; however, it remains a high-risk area for falls due to its enclosed space and slippery floor. Furthermore, because it involves highly private activities such as using the toilet and changing clothes, users are extremely sensitive to any form of human-shaped display, and some elderly people even refuse to use restrooms equipped with cameras.
[0137] To achieve basic security monitoring in extremely privacy-sensitive environments, this implementation adopts the minimalist visualization strategy described in this invention. The device is installed in a corner of the bathroom ceiling, with the lens covering the toilet and the standing area in front of it. When the system is first activated, the staff presses a physical button, and after a delay, a frame of empty bathroom footage is captured as the original background image 501. This image is processed by an algorithm into a completely detextured geometric contour map: only key structures such as the ground boundary, the outer edge of the toilet, and the position of the handrail are retained, while all colors, materials, brand logos, and identifiable objects are removed, forming a near-abstract spatial representation.
[0138] During normal operation, the device defaults to privacy mode and supports two output strategies. The first is the "minimalist heatmap" mode: the system does not generate any skeleton or human shape, but only highlights the detected moving areas with low-saturation color blocks (such as light blue), and the size of the color blocks and the height of the center of gravity change dynamically; when the height of the center of gravity is below the threshold and remains still for more than ten seconds, it is judged as a fall event. The second is the "abstract geometric shape" mode: the human body is simplified into a featureless elliptical cylinder, only reflecting the position of the center of gravity and the general orientation, without including joints, limbs or gender hints.
[0139] Regardless of the mode used, the system does not present any semantic graphics that can be interpreted as "human bodies." Users cannot discern "someone using the toilet" from the image; they can only see "low-activity movement in a certain area." The original high-definition video frames are discarded immediately after motion detection and center of gravity estimation are completed; they are neither stored nor uploaded. All processing is done locally on the device. The physical button is located outside the toilet door, and users decide whether to activate the monitoring function—if they do not wish to be monitored, they can leave the button pressed throughout the process, and the device will remain in sleep mode.
[0140] This implementation addresses the core contradiction in highly privacy-sensitive spaces: the need for automatic identification of life-threatening falls while simultaneously eliminating the psychological burden of being "watched." By reducing visualization to a minimalist form—non-human, non-gestural, containing only spatial-temporal-height information—the system technically achieves a state of "monitoring exists, but identity remains unknown."
[0141] In one possible implementation, the technical solution of the present invention is integrated into a home care service robot to provide daily interaction, safety reminders, and basic living assistance for elderly people living alone. The robot has a mobile chassis, a robotic arm, a voice interaction module, and a top-mounted visual perception unit. It needs to understand the user's posture through a camera to determine whether water needs to be provided, medication reminders needed, assistance in getting up, or fall detection required. However, because the robot needs to operate at close range in private spaces such as bedrooms, living rooms, and kitchens, continuously collecting raw human images can easily cause user anxiety and poses a risk of data leakage.
[0142] To achieve synergy between "proactive service" and "privacy protection," the vision module integrated on the top of the robot adopts the localized desensitization architecture described in this invention. Upon initial startup, the user or a family member presses a physical button on the robot, and after a delay, the system captures an empty view of the current room as the original background image 501. This image can retain color information to identify functional areas such as the dining table, sofa, and bed, or it can be converted into an edge-enhanced image to reduce sensitivity; the specific format is configured by the user through the accompanying application.
[0143] During daily operation, the robot defaults to privacy mode. When a user enters its field of vision, the edge AI processing unit extracts their key points in real time and generates a visual humanoid shape—for example, if the user prefers a simple interface, a linear skeleton diagram 601 is displayed; if there is an elderly person with cognitive impairment at home, it switches to a cartoon humanoid shape 602 to improve acceptance with a gentler image. This graphic is overlaid on the background to form a composite image for use by the robot's internal behavior decision-making module.
[0144] The key is that this visualized data is not only used for remote display, but also directly drives service logic. For example, when the system detects that a user has been sitting on the sofa for a long time with their head drooping, combined with the time context (such as having missed the medication time), the robot will proactively approach and remind the user via voice, "It's time to take your blood pressure medication." When a user tries to stand up from the bedside but their knee joints tremble significantly, the robot will extend its robotic arm to provide support and simultaneously send a "help get up" notification to their child's phone. If a sudden drop in torso angle is detected and the user remains still for more than eight seconds, a fall alarm will be triggered immediately, along with a reassuring voice message and contacting emergency contacts.
[0145] All the above judgments were made based on locally generated desensitized pose data. The original high-definition video never left the device's memory, nor was it uploaded to the cloud platform. When a doctor needs to view a user's facial status during a remote consultation, family members can request temporary high-definition authorization via the app. After the user presses a physical button on the robot, the device briefly outputs the original image, and then automatically switches back to privacy mode and updates the background.
[0146] In one possible implementation, the technical solution of the present invention is integrated into a multimodal health monitoring terminal, which is deployed at the bedside of elderly people living alone in the city to achieve 24 / 7 monitoring of vital signs and behavioral safety. Based on the visual perception module described in the present invention, this device further integrates a 60GHz millimeter-wave radar sensor, a near-infrared thermal imaging module, and ambient light and temperature / humidity sensing units, forming a multi-source sensing system that integrates visible light, thermal radiation, and radio frequency sensing. All sensors are housed in the same housing and share the same physical switching button and local edge AI processing platform.
[0147] Upon initial installation, community staff press a physical button on the device to trigger the initialization process. After a five-second delay, the system automatically captures a frame of visible light in an unoccupied state as the raw background image, and simultaneously records the current silent baseline and infrared thermal field distribution of the millimeter-wave radar. Afterward, the device operates in privacy mode by default. In this mode, the APS imaging path of the visible light camera is cut off by the hardware-level power domain, generating no raw image frames; the millimeter-wave radar continuously emits low-power continuous waves, extracting respiratory rate and heart rate data by detecting subtle chest movements; the infrared thermal imaging module outputs a low-resolution thermal distribution map, used only to determine the area where a human body is present and its approximate orientation, without showing facial or body details.
[0148] The edge AI processing unit performs spatiotemporal alignment and feature fusion of physiological signals extracted by millimeter-wave radar, the spatial distribution of infrared thermal fields, and motion pulse streams output by event cameras. Based on the fusion results, the system generates a composite desensitized visualization graphic: for example, limb posture is represented by a green skeleton line drawing (reconstructed from infrared thermal field contours and event streams), with a dynamically jumping circular icon superimposed on the center of the torso, its flashing frequency synchronized with the real-time heart rate; respiratory status is represented by the scaling animation of concentric circles surrounding the icon. This graphic is superimposed on a sketched background to form the final output image. All raw data—including visible light frames, high-resolution infrared images, and raw radar IQ data—are immediately released from memory after feature extraction is completed, without being written to flash memory or uploaded via Wi-Fi or 4G.
[0149] When family members or doctors need to temporarily check an elderly person's facial condition (such as assessing clinical signs like pallor or cyanosis of the lips), the person on site presses a physical button, and the device switches to high-definition mode. At this time, the APS imaging pathway is powered on again, outputting a 1080P visible light video stream; the millimeter-wave radar and infrared module continue to work, but their data is only used for auxiliary analysis and is not superimposed on the video image. The duration of high-definition mode can be set to 30 seconds, 60 seconds, or manually terminated by the user. After the high-definition mode ends, the device automatically switches back to privacy mode and, because it is configured to "update background on each switch," recaptures the current empty scene as a new background image after a delay.
[0150] It's worth noting that the activation status of the millimeter-wave radar and infrared module is also controlled by physical buttons. In privacy mode, the infrared module reduces its frequency to 1 frame per second, limiting the resolution to below 32×32, ensuring that individuals cannot be identified; the millimeter-wave radar only outputs structured physiological parameters (such as BPM values and respiratory waveforms), without transmitting raw point clouds or range-Doppler images. In high-definition mode, the infrared resolution can be increased to 64×64 to assist in low-light imaging, and the radar can output more detailed motion trajectories. However, regardless of the mode, the raw visible light image never leaves the device, while the data output format of other sensors is dynamically adjusted according to the mode status, ensuring that the overall system does not leak any identifiable information in privacy mode.
[0151] This terminal supports multiple visualization style configurations. Users can choose whether to display a heart rate icon, enable breathing animation, include keypoint markers in the skeletal line drawing, and retain or convert the background to grayscale sketch via the accompanying application. The system can also completely omit the background, outputting only a desensitized human figure on a pure black background. All these options are implemented in the local rendering engine and do not affect the processing logic of the original data.
[0152] Furthermore, the device can be deployed on bedside tables, walls, nursing bed frames, or mobile carts, and is powered by a DC adapter, PoE, or a built-in rechargeable battery. Networking capabilities include Ethernet, Wi-Fi 6, 4G Cat.1, and Bluetooth 5.0, and it can also operate completely offline, recording alarm logs only via a local memory card. Output protocols support RTSP, HTTP-FLV, and GB / T28181, allowing seamless integration with existing elderly care service platforms.
[0153] In one possible implementation, the intelligent visual terminal of the present invention is deployed in the elderly residents' rooms and public activity areas of a medium-sized nursing home in an city, serving as an enhancement component of its existing video surveillance system. The nursing home had previously built a conventional high-definition security system based on an IP network, including multiple 1080P fixed cameras, a gigabit switch supporting PoE power supply, a mainstream brand network video recorder (NVR), and two 4K monitoring splicing screens located at the nurses' station and the property management duty room, used for 24-hour rotating display of images from each floor. However, due to strong resistance from residents and their families regarding traditional cameras in high-privacy areas such as bedrooms and bathroom entrances, these critical areas have long been "blind spots" for surveillance, resulting in the inability to detect falls, sudden illnesses, and other incidents in a timely manner, posing significant care risks.
[0154] To resolve this contradiction, the hospital replaced the traditional cameras in the blind spots with the intelligent vision terminal of this invention without replacing the existing NVR, adding new servers, or modifying the network cabling. This terminal connects to the existing PoE switch via a standard RJ45 network cable, simultaneously obtaining power and transmitting data through the same cable, eliminating the need for additional power adapters or high-voltage wiring. The device is pre-installed with the ONVIF 2.6 protocol stack and GB / T 28181 national standard encoding capabilities. Upon power-up, it is automatically recognized by the existing NVR as a regular IP camera, assigned an IP address, and incorporated into unified management. During daily operation, the terminal defaults to privacy mode, continuously outputting an anonymized video stream composed of skeleton line drawings and static backgrounds to the NVR. The stream uses standard H.264 encoding, 1080P resolution, and 30fps frame rate, and is fully compatible with the decoding and storage capabilities of NVRs. It can be displayed normally on the large screen at the nurse station with the same window size, layout logic, and operation method as other cameras. What nursing staff see is an abstract image with a clear indication of human posture, which can accurately determine whether the elderly are lying in bed, getting up, walking, or falling, but cannot identify their face, clothing, or body features, thus effectively alleviating privacy concerns.
[0155] When the on-duty nurse receives a call or notices abnormalities (such as prolonged stillness or distorted posture) on the desensitized screen, and further confirmation of the elderly patient's complexion, level of consciousness, or presence of clinical signs such as vomiting or bleeding is needed, she can go to the door of the room and press the physical button on the terminal casing. The device immediately switches to high-definition mode: the APS image sensor is re-powered and outputs a raw 1080P visible light video stream; simultaneously, the NVR automatically receives the new video stream and refreshes the corresponding window on the nurse station screen in real time, seamlessly switching the image from a skeleton diagram to a real image. High-definition mode lasts for 60 seconds by default, during which the NVR simultaneously records high-definition clips for later review. After the time expires, the device automatically switches back to privacy mode, and because it is configured to "update background on each switch," it recaptures the current empty room image as a new background image after a five-second delay, ensuring the accuracy of subsequent desensitized screens.
[0156] Throughout the entire process, all data transmission was completed through the existing PoE network, without the addition of any dedicated lines, edge servers, or cloud platforms. The terminal's dual-mode output capability enables it to meet both daily privacy monitoring needs and provide clinical-grade visual information at critical moments, all built upon the nursing home's existing NVRs, switches, monitoring screens, and maintenance procedures. Compared to building a new independent privacy monitoring system (which typically requires an additional investment of tens of thousands of yuan for dedicated servers, customized software, dual-link cabling, and personnel training), this solution only increases the cost per device while achieving the dual goals of functional upgrades and compliance implementation.
[0157] This implementation clearly demonstrates that by employing technologies such as protocol compatibility, PoE plug-and-play, and seamless NVR integration, the barrier to deploying intelligent privacy monitoring capabilities within existing security systems is significantly lowered. For community elderly care centers, primary healthcare stations, old residential property management companies, or home-based aging-in-place renovation projects with limited budgets and technical resources, this solution avoids high equipment purchase costs, complex system integration projects, and lengthy installation and commissioning cycles, truly achieving a practical implementation path of "using conventional facilities, following standard protocols, reducing deployment difficulty, and improving care efficiency."
[0158] In one possible implementation, the intelligent vision terminal of the present invention is deployed in a simulated home behavior research platform of a university's Human Factors Engineering and Intelligent Interaction Laboratory to address a long-standing core contradiction in human behavior science research: how to acquire high-precision human motion data in a compliant and continuous manner without interfering with the natural behavior of the subjects.
[0159] This laboratory is dedicated to studying the behavioral patterns of the elderly and people with mobility impairments in their daily activities, including key movements such as walking, sitting-lying transitions, bending over to pick up objects, and sudden falls, in order to support the development of service robots, accessible environment design, and health risk early warning systems. However, traditional data collection methods face significant obstacles. Using optical motion capture systems requires attaching reflective markers to the subjects' bodies or having them wear special clothing, which is not only cumbersome but also significantly alters their natural behavior, placing a psychological and physiological burden, especially on elderly or disabled subjects. If ordinary cameras are used to record raw video, the inclusion of facial, body shape, and clothing information that can identify individuals makes it difficult to pass research ethics reviews, and the pressure of being "recorded throughout" often distorts the subjects' behavior, affecting the validity of the data.
[0160] To overcome this challenge, the laboratory introduced the intelligent vision terminal of this invention as the core sensing device. Multiple terminals were installed in the corners of the ceilings in simulated living rooms, bedrooms, kitchens, and other areas. Each terminal integrated a visible light image sensor, an edge artificial intelligence processing unit, and a physical switching button located on the device's casing. During system initialization, the experimenter pressed the physical button, and the device initiated a short delay (e.g., 5 seconds). After the person left the field of vision, it automatically captured a frame of the unoccupied environment as a static background image.
[0161] During the formal experiment, all terminals operated in privacy mode by default. Subjects moved freely without awareness or with minimal awareness. The terminals used a locally deployed lightweight human pose estimation model to analyze the real-time video stream frame by frame, extracting the two-dimensional coordinates and confidence scores of multiple human joints. Subsequently, the system rendered these structured features into abstract visual graphics—such as skeleton line drawings, simplified human outlines, or low-polygon 3D projections—and overlaid them on a static background to form a desensitized video stream. This stream was used solely by the experimenter to monitor the process status in real time and did not contain any original image information that could reveal the subject's identity.
[0162] The data actually used for scientific analysis is not video, but structured behavioral logs synchronously generated by the edge AI processing unit. These logs identify subjects with anonymous IDs and record timestamps, coordinates of each key point, movement speed, posture category, and abnormal event markers (such as a fall probability exceeding a threshold). The logs are stored in encrypted format in the device's local non-volatile memory and exported via a wired interface after the experiment for subsequent modeling and algorithm training. Crucially, the original visible light video frames are immediately released from memory after feature extraction; they are neither written to the internal storage card nor transmitted out of the device via any network interface, ensuring that sensitive biometric data never leaves the hardware boundaries.
[0163] In one possible implementation, the technical solution of the present invention is implemented as a highly integrated embedded vision module, installed inside the head structure of a service robot designed for home care scenarios. This robot is primarily used to assist elderly or mobility-impaired users in tasks such as daily activity monitoring, fall warning, remote video interaction, and life reminders. In such applications, the robot needs to continuously acquire information about the user's location, posture, and behavioral status to support its decision-making and response logic. However, directly using conventional cameras to collect and transmit raw video streams will inevitably record the user's facial features, body shape, clothing details, and private space environment, easily raising user concerns about privacy leaks, especially when active in highly sensitive areas such as bedrooms and living rooms, significantly reducing product acceptance.
[0164] To address the aforementioned issues, this service robot integrates the intelligent vision module described in this invention. This module includes a visible light image sensor, an edge AI processing unit, a physical switching button, a background capture module, a video synthesis module, and a video output interface. All functional units are packaged on a single circuit board and connected to the robot's main control system via a standard communication interface (such as MIPI CSI-2 or USB 3.0). The physical switching button is located on the front of the robot body as a physical mechanical structure for easy user operation.
[0165] Upon initial deployment, the caregiver presses the physical switch button. After a five-second delay, the module automatically captures a frame of the environment in an unoccupied state, using it as a static background image and storing it in local non-volatile memory. Thereafter, the module operates in privacy mode by default. In this mode, the visible light image sensor continuously acquires real-time video streams, and the edge AI processing unit performs human detection and pose estimation on the device, extracting the two-dimensional coordinates and confidence values of multiple key points. The video synthesis module generates visual graphics based on these structured features—for example, a skeleton diagram connected by green lines, or a cartoon human figure composed of simplified geometric shapes—and overlays this graphic onto the static background image to form a desensitized privacy video stream. This video stream is transmitted to the robot's main control system through the module's output interface to drive screen display or behavior analysis algorithms.
[0166] Crucially, in privacy mode, the raw high-definition video frames exist only briefly in the memory of the edge AI processing unit for feature extraction and graphics rendering, and are then immediately released. They are neither written to any internal storage medium nor transmitted out of the module via communication channels such as Wi-Fi, Bluetooth, or Ethernet. Therefore, the robot's main control system and remote cloud platform cannot obtain any raw image data containing identifiable information.
[0167] When family members or medical staff need to temporarily check the user's actual condition (e.g., assess facial color, level of consciousness, or wound condition), they can press the physical switch button on the robot. The module then switches to high-definition mode: the visible light image sensor outputs a raw 1080P video stream, which is encoded and transmitted to the main control system, and can be displayed through the robot screen or a remote video call interface. The high-definition mode lasts for a preset value (e.g., 60 seconds), and automatically switches back to privacy mode after the session ends. Because it is configured to "update background on each switch," it recaptures the current empty scene as a new background image after a delay.
[0168] Through the aforementioned technical solution, this service robot effectively avoids the privacy risks caused by the continuous collection and transmission of raw video while meeting the requirements for behavioral perception and interaction functions. Users can explicitly control when high-definition video is activated via physical buttons, ensuring their autonomy over their personal image data; while daily monitoring relies on irreversibly desensitized visual content, preserving spatial context and behavioral semantics while avoiding the exposure of identity information. This integrated approach enables the robot to achieve compliant deployment in scenarios with high privacy requirements without relying on cloud processing, increasing additional network bandwidth, or changing the existing main control architecture, significantly improving the product's usability and trustworthiness in elderly homes and nursing homes.
[0169] In one possible implementation, the technical solution of the present invention is realized as a low-profile, embedded intelligent vision module integrated into the headboard support structure of a high-end intelligent nursing bed for non-contact monitoring of the behavior and physiological state of bedridden patients. This nursing bed is mainly used in hospital rehabilitation departments, long-term care facilities, and home-based elderly care scenarios. It needs to continuously identify key events such as changes in patient position (e.g., lying flat, lying on their side, semi-sitting), getting out of bed, frequency of turning over, and signs of impending falls to support pressure ulcer prevention, nighttime safety monitoring, and improved care efficiency.
[0170] Traditional nursing beds rely heavily on pressure-sensing pads, infrared beams, or microwave radar for monitoring. However, pressure pads cannot distinguish details of body position and are easily affected by sheet wrinkles; while infrared or radar solutions can detect presence, they struggle to accurately determine complex behaviors such as limb posture, sitting angle, or whether someone has slipped off the bed. Adding a regular camera directly to the patient's face and body involves the collection of highly sensitive biometric data, posing significant obstacles in terms of ethical review and user acceptance.
[0171] To overcome the aforementioned limitations, this embodiment embeds the intelligent vision module of the present invention as the core sensing unit into the bed frame. This module includes a visible light image sensor, an edge artificial intelligence processing unit, a physical switching button, a background capture module, a video synthesis module, and a communication interface. Its overall thickness is controlled to within 15 millimeters, allowing it to be completely concealed behind the headboard decorative panel, with only the optical window exposed. The physical switching button is a small mechanical button located inside the bedside rail, facilitating operation by caregivers or the patient.
[0172] When the device is first activated, the caregiver presses the physical switch button. After a five-second delay, the system automatically captures a frame of an empty bed as a static background image after the caregiver leaves the area of view, and stores it in the module's local flash memory. Thereafter, the module operates in privacy mode by default. In this mode, the visible light image sensor continuously acquires video streams of the bed surface area, and the edge AI processing unit performs human segmentation and pose estimation in real time on the device, extracting the coordinates of key joints in the torso, limbs, and head. The video synthesis module generates a simplified human contour graphic based on this data—for example, using soft color blocks to represent the torso and limbs, or using a skeletal line drawing to illustrate joint connections—and overlays this graphic onto the static background image to form a desensitized video stream. This stream is output to the nursing bed's main control system via a UART or Ethernet interface to drive the local display screen or upload it to the nurse station monitoring platform.
[0173] Throughout the privacy mode operation, the original high-definition video frames are only used in the module's memory for feature extraction and graphics rendering, and are released immediately after completion. They are neither written to the internal storage card nor transmitted through any network channel. Therefore, the nursing bed control system, remote server, or mobile terminal cannot obtain the original images containing information about the patient's face, body shape, or clothing.
[0174] When clinicians or family members need to temporarily assess a patient's complexion, skin condition, or level of consciousness, they can initiate a request via an authorized terminal. On-site personnel then press a physical button on the bedside to trigger high-definition mode. The module then outputs a raw 1080P video stream for 60 seconds, automatically switching back to privacy mode and recapturing the current empty bed background to update the base image. This mechanism ensures that enabling high-definition video requires on-site physical operation and cannot be bypassed via remote software commands.
[0175] Furthermore, this module can work in conjunction with a millimeter-wave radar module integrated into the nursing bed. In privacy mode, the heart rate and respiratory rate data output by the radar are transmitted to an edge AI processing unit to enhance the expressiveness of the visualization graphics—for example, overlaying a concentric circle that scales with the breathing rhythm on the torso area, or displaying the heart rate value in the center of a human-shaped icon. However, the radar's raw point cloud or IQ data is not directly used as video output, but only participates in local fusion as auxiliary features.
[0176] Through the integration of the aforementioned technologies, the intelligent nursing bed achieves substantial protection of patient privacy without sacrificing the accuracy of behavioral recognition. Caregivers can accurately determine whether a patient has not turned over for an extended period, has attempted to get out of bed without seeking help, or is in an abnormally still state, while avoiding the collection and transmission of personally identifiable visual information. This solution meets the compliance requirements of medical institutions for "minimum necessary data collection" and "de-identification processing," significantly enhancing patient dignity and care experience, and providing a feasible path for the large-scale deployment of intelligent nursing equipment in highly sensitive medical environments.
[0177] In one possible implementation, the technical solution of the present invention is applied to the low-cost, rapid intelligent monitoring system upgrade of an old urban nursing home that has been operating for over ten years. The nursing home's original video surveillance system was an early-built IP high-definition system, including more than ten 1080P network cameras, an eight-channel network hard disk recorder, a gigabit LAN, and two monitors located at the nurses' station and the property management office. The system consistently covered corridors, entrances, and public activity areas, but high-privacy areas such as elderly residents' rooms, private bathrooms, and public shower rooms remained blind spots due to strong objections from residents and their families. Events such as nighttime falls, sudden illnesses, or prolonged periods of inactivity often went undetected, posing significant care risks.
[0178] To enhance security monitoring capabilities within a limited budget while respecting residents' privacy, the hospital decided to adopt a partial replacement strategy, deploying intelligent vision terminals integrating the technology of this invention in key high-risk areas. These terminals support standard PoE power supply, ONVIF 2.4 protocol, H.264 encoding, and 1080P@25fps video output, and are fully compatible with existing NVRs. Installation personnel only need to connect the new devices to the existing network cable; after powering on, they are automatically recognized as ordinary cameras by the NVR. No replacement of the main control equipment, no new server, no rewiring, and no complex training for on-duty personnel is required. The entire renovation was completed within three days, with seven terminals installed, covering six rooms for disabled elderly residents and one public bathroom entrance, at a total cost of less than 30,000 yuan.
[0179] After the upgrade, all newly added terminals run in privacy mode by default. During their daily shifts, nurses can view these locations on existing monitors—the content is not actual video, but a desensitized video stream generated by overlaying skeletal line drawings onto a sketched static background. This stream clearly shows the human body's position, posture, and direction of movement: for example, a lying skeleton is displayed when the person is in bed, a standing human figure is displayed when they get up, and a sudden change in joint angles accompanied by a downward shift when they fall. Nursing staff can use this to determine if any abnormal behavior has occurred, such as not returning to bed at night, remaining still for extended periods in the shower room, or unstable sitting posture. Because the output format is consistent with the original system, the NVR can record, store, and play back the desensitized video normally. On-duty personnel can access historical records using the existing interface, with no learning curve whatsoever.
[0180] In terms of night shift management, the system has significantly improved response efficiency and workload balance. Previously, night shift nurses had to manually make rounds every two hours, which affected their own rest and made it difficult to cover all emergencies. Now, by observing anonymized images in real time on monitors, nurses can detect abnormal behavior immediately from the duty room. For example, if a room's skeletal diagram shows an elderly person slipping off the bed and remaining still for more than a minute, although the system doesn't automatically alarm, the nurse can immediately notice and go to investigate. This shift from "passive monitoring to proactive visibility" optimizes night shift staffing, allowing one nurse to effectively monitor the entire floor without frequently moving around and disturbing other residents' sleep.
[0181] In daily operations and management, the system also brings multiple efficiency improvements. First, during family visits or when handling complaints, managers can demonstrate the physical button function on-site: pressing it immediately switches the monitor screen from a skeleton diagram to a real high-definition image, which automatically returns to normal after 60 seconds. This transparent and controllable operating mechanism greatly enhances family trust and reduces disputes arising from "whether or not one has been secretly filmed." Second, nursing records are more objective and accurate. In the past, it relied on manual recall to fill in "what time I got up" and "whether I went to the toilet," but now the timeline of behavior can be verified by playing back desensitized videos, improving the quality of care logs. Third, newly hired caregivers can learn typical behavioral patterns (such as the gait of Parkinson's patients and the posture of hemiplegic patients after stroke) by watching desensitized videos, completing pre-job training without infringing on privacy.
[0182] Crucially, all original high-definition video frames remain on the terminal device in privacy mode. After the edge AI processing unit completes pose estimation and graphics rendering, it immediately releases the original image memory, neither writing it to the local storage card nor transmitting it over the network. Therefore, even if the NVR is illegally accessed, attackers can only obtain anonymized footage and cannot reconstruct any identity information. Furthermore, enabling high-definition mode requires a physical button press by an on-site personnel; remote software commands cannot simulate or bypass this, ensuring absolute control over the user's personal video feed.
[0183] Through its protocol-compatible, plug-and-play design and zero-new-infrastructure-required infrastructure, it achieves low-cost transformation and extremely fast deployment cycles. Without altering the existing high-definition monitoring system architecture, it provides a practical and feasible technical path for the intelligent upgrading of a large number of existing elderly care facilities nationwide by simply replacing or adding a small number of terminals.
[0184] In one possible implementation, the technical solution of the present invention is applied to the intelligent transformation of the rehabilitation medicine ward of a large general hospital. This ward admits a large number of inpatients with limited mobility or a high risk of falls. The clinical medical team needs to continuously monitor the patients' activity status in the ward, including whether they get out of bed, whether they can walk independently, whether falls have occurred, and the completion status of rehabilitation training exercises. However, due to concerns about patients' personal image and privacy, directly using conventional cameras for video monitoring faces strong privacy concerns. Furthermore, the hospital's information management department explicitly requires that all patient-related visual data must be strictly retained within the hospital's local environment and must not be transmitted to external networks or third-party platforms in any way.
[0185] Previously, the department had evaluated several intelligent monitoring solutions. One was a system that relied on cloud-based artificial intelligence processing, which required uploading raw video to an internet server for analysis. Although it was feature-rich, it was rejected due to data leakage issues. Another solution involved deploying a dedicated edge server within the department to centrally process multiple video streams. However, this solution not only required high initial investment but also specialized network isolation and dedicated maintenance personnel. Furthermore, if the server failed, the intelligent monitoring function of the entire ward would be completely interrupted, making it unreliable.
[0186] Ultimately, the hospital selected the intelligent vision terminal of this invention as its core sensing device. Multiple terminals were installed inside patient rooms and at the entrances to restrooms. Each device is an independently operating complete system, incorporating a visible light image sensor, an edge AI processing unit, physical switching buttons, a local storage module, and a standard network interface. All terminals are connected to a dedicated departmental switch via the hospital's intranet, and are linked to existing monitors and local network video recorders (NVRs) at the nurses' station. They do not connect to the internet or rely on any central computing node.
[0187] During normal operation, the terminal defaults to privacy mode. The edge AI processing unit analyzes the real-time video stream within the device, identifies the patient's body and extracts the coordinates of key points, then generates visual graphics such as skeleton line drawings, which are overlaid on a pre-captured background map of an unoccupied environment to form a desensitized video stream. This video stream is output in standard H.264 format and received and stored by the local NVR via the ONVIF protocol. Nurses can view the real-time activity status of the abstract human figures in each ward on the monitor in the duty room to determine whether there is prolonged stillness, abnormal posture, or risk of falling.
[0188] Crucially, in privacy mode, the original high-definition video frames only exist briefly in the terminal's memory for feature extraction and graphics rendering, and are released immediately afterward. These original images are neither written to the terminal's internal storage chip nor transmitted out of the device via the network interface. Therefore, even if the local NVR is accessed, only anonymized images that cannot be used to reconstruct the patient's identity can be obtained. Sensitive information such as the patient's face, body shape, and clothing never leaves the terminal's hardware boundaries.
[0189] When the attending physician or responsible nurse needs to temporarily check a patient's complexion, skin condition, or level of consciousness, they can enter the ward and press the physical button on the terminal's casing. The device then switches to high-definition mode, recording and encrypting the original 1080P video stream, which is then stored locally on the terminal. This high-definition clip can only be accessed through authorized terminals within the hospital, and the system automatically deletes it after 72 hours by default to prevent long-term storage. The entire process requires no connection to an external system; all operations are directly triggered by hardware signals and cannot be simulated via remote software commands.
[0190] This deployment method demonstrates significant advantages in localized processing. First, the system is completely decentralized, with each terminal independently completing sensing, analysis, and output, eliminating the risk of single points of failure. Second, it eliminates the need for new servers, firewalls, or dedicated software platforms, reusing existing NVRs and monitoring equipment in the hospital, significantly reducing initial investment and subsequent maintenance burden. Third, all data processing and storage are completed within the physical hospital area, ensuring that patient imaging information remains within a controllable hospital environment, effectively addressing the privacy concerns of clinical, ethical, and management stakeholders.
[0191] In actual operation, the system has significantly improved nursing efficiency and response speed. Night shift nurses can monitor patients' dynamics through monitors without frequent ward rounds; rehabilitation therapists can review desensitized videos to assess the quality of patients' training movements; during family visits, medical staff can demonstrate the physical button switching function on-site, intuitively showcasing the privacy protection mechanism of "not showing real people under normal circumstances, but using high-definition only in emergencies," thereby enhancing trust.
[0192] In medical settings where data security and privacy are paramount, this invention achieves an effective balance between behavioral monitoring and privacy protection by embedding artificial intelligence processing capabilities into the terminal, combining physical switch control with local closed-loop output. The system provides hospitals with a secure, reliable, and easily deployable intelligent monitoring solution without relying on external networks or adding complex infrastructure.
[0193] In one possible implementation, the technical solution of the present invention is applied to a community service upgrade project in an old urban residential area with a high degree of aging. This residential area was built in the early 2000s, and over 40% of its permanent residents are over 65 years old, with a large number of elderly people living alone, in empty nests, or with disabilities or partial disabilities. The existing security system is an early-deployed IP video surveillance network, mainly covering the entrances and exits of the community, main roads, and unit lobbies, but lacks effective visual perception capabilities in key locations such as elevator cars, public areas in corridors, and around accessible facilities. Due to the declining mobility of the elderly and the higher risk of sudden health events, the property service center urgently needs a technical means to promptly identify and respond to abnormal behavior without disturbing residents' normal lives and while protecting their privacy.
[0194] To enhance community care response capabilities, the property management deployed intelligent visual terminals integrating the technology of this invention in high-risk areas without altering the existing monitoring architecture. These terminals were installed inside several older residential elevators, at stairwell corners in several high-rise units, and in the outer passageways of the community's accessible public restrooms. The devices utilize standard PoE power supply and connect to the existing network video recorder (NVR) in the property management office via the community's existing local area network. They output 1080P video streams compliant with the ONVIF protocol, which are displayed on the duty monitor along with the existing camera feeds, eliminating the need for additional servers, dedicated software, or operator training.
[0195] All newly added terminals run in privacy mode by default. In this mode, the terminal's built-in edge AI processing unit performs local analysis of the real-time video stream, identifies human bodies in the image, extracts key point coordinates, generates skeletal line drawings or simplified human outlines, and overlays these visualizations onto a pre-captured background image of an unoccupied environment to create an anonymized video stream. During routine patrols, property management staff can clearly observe whether people have fallen, remained still for extended periods, or lingered abnormally through monitors, but cannot obtain any original image information that can identify their identities. For example, in an elevator, the system can determine if a passenger has suddenly collapsed; in a hallway, it can identify whether an elderly person has been sitting or lying near a doorway for an extended period. This information provides property management with a basis for early intervention, helping to contact family members or coordinate community medical resources within the critical timeframe.
[0196] In cases of emergency or necessary verification requiring access to the actual video feed, authorized personnel can go to the site and press the physical button on the terminal casing to temporarily activate HD mode. The device then outputs the original HD video stream, automatically switching back to privacy mode after a preset duration and updating the background image. This switching mechanism is directly driven by hardware signals and cannot be triggered by remote commands, mobile applications, or background software, ensuring residents' control over their own video data.
[0197] The entire system operates entirely locally. Raw high-definition video frames are used solely for feature extraction within the terminal in privacy mode, and memory is immediately released upon completion; they are neither written to local storage nor transmitted over the network. While the anonymized video stream is recorded by the NVR, the content consists only of abstract graphics and does not contain sensitive information such as faces, body shapes, or clothing. All data processing and storage are confined to the community's internal network environment, without internet connection or reliance on external platforms, complying with the community's management requirements for the security of residents' personal information.
[0198] In actual operation, the system has significantly improved the efficiency of service response for the elderly. Night shift staff can keep track of key areas without frequent patrols; community grid workers can use the anonymized images to trace the daily activity patterns of the elderly and optimize home visit plans; in the event of a sudden health incident, property management can quickly locate and assess the situation, shortening emergency response time. Many elderly residents said that seeing only "line figures" on the screen without real images significantly improved their psychological acceptance and made them more willing to cooperate with relevant safety measures.
[0199] In older residential communities with a high concentration of elderly residents, this invention achieves a balance between behavioral awareness and residents' privacy rights by embedding privacy protection mechanisms into the terminal, combining physical switch control with localized processing. While reusing existing infrastructure, the system provides communities with a low-intrusion, highly available, and easy-to-maintain intelligent monitoring approach, effectively supporting the construction of a refined community service system for the elderly.
[0200] In one possible implementation, the technical solution of this invention is systematically applied to various health and care insurance service scenarios targeting elderly people living at home, covering long-term care insurance, comprehensive elderly liability insurance, disability income compensation insurance, home health management insurance, and commercial health insurance products bundled with services such as community-based elderly care and home-based hospital beds. In these insurance models, insurance companies do not directly provide medical services, but rather procure services from third-party professional care institutions to provide policyholders (mostly elderly, disabled, semi-disabled, or chronically ill patients) with regular in-home care, rehabilitation training, daily living assistance, medication reminders, fall prevention, and other non-medical support services. The actual implementation of these services directly affects the effective use of insurance funds, the fairness of claims processing, and users' trust in the insurance products.
[0201] However, this field has long faced a fundamental contradiction: service quality supervision requires objective evidence, but the process of obtaining evidence is highly susceptible to infringing on user privacy. Traditional methods primarily rely on caregivers manually filling out electronic work orders, recording fields such as "service start time," "end time," and "completed items." This method has serious flaws—it cannot verify whether actions actually occurred (e.g., "turning over" only requires checking a box, no proof is needed), cannot determine whether operations were performed correctly (e.g., whether the angle of rehabilitation movements met standards), and cannot identify service interruptions or perfunctory behavior (e.g., leaving after checking in). Some institutions have attempted to require caregivers to take photos or short videos of services using their mobile phones as evidence, but this is not only cumbersome and disruptive to service flow, but also generates numerous complaints because the images include the user's face, exposed body parts, or private home environment, even leading to policyholders refusing to continue receiving services. This is especially true for the elderly, who have cognitive impairments, aphasia, or mobility issues; they are unable to effectively express their dissatisfaction yet are most in need of reliable service guarantees, placing them in a vulnerable position of "being served but having no voice."
[0202] To overcome this predicament, insurance companies, in collaboration with technology service providers, are deploying the intelligent visual terminal of this invention in policyholders' homes as a standardized, non-invasive service recording device. This device is typically configured and encrypted by the insurance company and installed in a corner of the bedroom ceiling, outside the bathroom door, or in the main activity area of the living room after the user signs an informed consent form. It covers typical service scenarios but avoids direct view of the bed or bathing area. The device has a simple appearance with no obvious camera features, reducing psychological resistance.
[0203] Before each service begins, the caregiver initiates a service request via a dedicated mobile terminal. The system automatically sends a confirmation prompt (such as a doorbell tone or voice announcement) to the policyholder or their guardian. After the user presses the physical button on the terminal casing, the device activates and enters privacy mode. In this mode, a visible light image sensor captures a real-time video stream, and the edge AI processing unit performs all the analysis internally: first, it detects the presence of a human body in the image, then estimates the posture of each person, and extracts the coordinates of multiple joint points. Subsequently, the video synthesis module generates abstract visual graphics based on this structured data—for example, using different colors to distinguish caregivers from those being cared for, connecting joints with green lines to form a skeleton diagram, or driving a simplified two-dimensional human silhouette. This graphic is precisely overlaid on a previously captured background image of an unoccupied environment, forming a desensitized video stream. The entire process is completed locally; the original high-definition video frames only exist briefly in memory for feature extraction and rendering, and are immediately released after completion, neither written to a memory card nor transmitted out of the device via Wi-Fi, 4G, or any network interface.
[0204] This anonymized video stream clearly presents the key behavioral semantics of the service process: for example, when a caregiver assists an elderly person to stand up from the bedside, the system can identify hand contact, weight transfer, and standing posture; during passive joint movement training, it can determine whether the limbs complete the required flexion and extension movements; if there is prolonged stillness during the service (such as more than 5 minutes without interaction), the system automatically marks it as "abnormal stagnation." This information far exceeds the text description of traditional work orders, providing tamper-proof visual evidence for subsequent verification.
[0205] After the service is completed, the terminal automatically generates a structured service record package, including anonymized video clips (service period only), total service duration, identified action types and frequencies, abnormal event logs, device status, and other metadata. This data is then transmitted via an encrypted channel to a local server deployed within the insurance company's intranet environment. This server is not connected to the internet and is only accessible to the audit, settlement, and quality management departments. The settlement team verifies the work order based on the actual actions performed in the video, preventing "payment upon check-in." The risk control department analyzes historical data to identify high-frequency abnormal patterns (such as a caregiver repeatedly stopping immediately after the service begins) and intervenes promptly for investigation. The service quality team uses the anonymized video footage to conduct remote supervision, scoring the standardization of nursing actions and incorporating this into performance evaluations.
[0206] In the event of a service dispute—for example, a family member questions whether a caregiver provided bathing assistance, while the caregiver insists that the user refused to cooperate—the insurance company can retrieve desensitized videos from the corresponding time period for objective review. If it is necessary to view the original images to confirm skin condition, wound treatment details, or environmental safety risks, two conditions must be met: first, the policyholder or their legal guardian must be present; second, a physical button must be pressed on-site to temporarily activate the high-definition mode (if the device supports local playback) or to retrieve high-definition clips recorded with prior explicit consent. This mechanism ensures that the acquisition of high-definition images is always based on the user's on-site, proactive, and verifiable authorization, and cannot be triggered remotely in the background or recorded afterward, fundamentally eliminating the possibility of covert surveillance.
[0207] Through the above technical approach, this invention solves several key problems in insurance service scenarios:
[0208] First, it achieves objectification of service authenticity verification. It no longer relies on subjective reporting, but instead uses quantifiable visual behavioral evidence to support settlement and auditing, significantly reducing the risk of fraud and unnecessary expenditures.
[0209] Second, it established a standardized foundation for service quality assessment. Abstract and visualized graphics preserve the structure of actions and spatial relationships, making remote supervision possible and promoting the transformation of care services from "experience-driven" to "data-driven".
[0210] Third, it achieves a balance between privacy and oversight in a highly sensitive and private environment. Users see "line figures" rather than real images, which makes them psychologically more accepting; at the same time, the physical switch gives them absolute control over the high-definition mode, which conforms to the ethical principles of "informed consent and controllability".
[0211] Fourth, it enhances the protection of the rights and interests of vulnerable elderly groups. For disabled elderly people who cannot clearly express themselves, the system automatically records the service process, providing them with "silent witness" and preventing service absence or abuse.
[0212] Fifth, it reduces the operational complexity and compliance risks for insurance companies. All processing is completed locally on the terminal, the raw data does not leave the device, and the anonymized content is stored on the internal network server. There is no need to rely on public cloud or third-party AI platforms, avoiding the risk of cross-border data leakage or other risks, and there is no need to bear additional data governance burdens.
[0213] Sixth, it promotes the healthy development of the care service ecosystem. High-quality caregivers receive higher ratings and more orders due to their standardized operations, while substandard services are accurately identified and eliminated, forming a positive incentive mechanism.
[0214] The scope of protection of the technical solution proposed in this application should not be limited to the specific implementation details, graphic styles, device forms, or application scenarios listed in the specification. The core of this solution lies in constructing a privacy-preserving visual architecture that uses physical-level control as the starting point of trust, local edge intelligence as the processing hub, and irreversible desensitization as the output endpoint. Under this architecture, the original visible light image exists only briefly in the device's memory as intermediate data to extract structured features related to human behavior, and is then immediately released without being written to any persistent storage medium or transmitted beyond the device boundary via wired or wireless channels. The output content consists entirely of locally generated abstract representations and cannot reconstruct any personally identifiable visual information.
[0215] At the level of visualization, this invention covers all non-realistic image forms generated based on human movement or presence characteristics. This includes not only concrete representations such as skeletal line drawings, cartoon human figures, simplified outlines, low-poly 3D models, and heat maps, but also abstract visual encodings that completely lack "human" semantics, such as dynamic point clouds, trajectory arrows, region highlight blocks, color mapping fields, symbol labels, geometric projections, particle flows, spectrogram maps, and time series visualizations. Regardless of whether joints are presented, line segments are connected, left and right limbs are distinguished, or proportional relationships are preserved, as long as the generation logic relies on locally extracted human-related data (such as joint coordinates, centroid position, motion vectors, event pulses, optical flow fields, foreground masks, posture vectors, height estimates, gait period parameters, etc.), and the output results lack identity recognition, it falls within the technical scope of this application.
[0216] The use of a background image is optional, and its presence or absence does not affect the integrity of the technical solution. The system can choose to capture a frame of an unoccupied environment as a static background, and, as needed, retain color, convert to grayscale, perform edge enhancement, perform semantic segmentation, remove texture details, or replace it with a solid color, a virtual scene, or a procedurally generated image. The system can also completely omit the background, outputting only a desensitized human figure on a monochrome background. Crucially, the background content must not contain any real-time human images. If the environmental portion of the output image originates from a video frame containing people, regardless of whether it has been blurred, pixelated, or masked, it is not within the technical scope of this invention; however, if the background is an independently captured result in an unoccupied state, and the human figure is reconstructed from features, then regardless of the background's form, it falls within the protection scope.
[0217] The technical logic of this invention is also applicable to situations where only de-identified features are used for local intelligent decision-making without generating a video stream. For example, the system can directly input the extracted posture data into a fall detection model, a long-term static alarm module, an area intrusion judgment engine, a people counting unit, or a robot control command generator, without generating an external video signal. This kind of "no display but analysis" implementation, because its core still relies on the technical path of the original image not leaving the device and the behavioral information being de-identified locally, also constitutes an implementation of this application.
[0218] In terms of device carriers, this solution can be deployed on any smart terminal with local AI processing capabilities, far beyond traditional cameras. Typical forms include portable service recorders, head-mounted work terminals, vehicle-mounted in-cabin monitoring modules, integrated vision units on nursing beds, smart mirrors, rehabilitation training equipment, service robots, robotic arm vision guidance systems, access control intercom devices, elevator car monitors, public restroom security terminals, wide-angle equipment in community activity halls, ceiling-mounted monitors in hospital wards, smart home sensors, wearable health devices, and drone indoor navigation modules, etc. Regardless of whether the device is fixed, mobile, handheld, embedded, or wearable, as long as it achieves a technical closed loop of physical button triggering, local feature extraction, desensitized output, and original image isolation, it falls under the reasonable extension of this application.
[0219] The application scenarios are also highly scalable. In addition to the scenarios already described, such as home-based elderly care, community care, door-to-door services, hospital monitoring, and barrier-free facilities, this technology can also be applied to: recording home visits by grassroots public service personnel (social workers, housekeepers, repairmen, delivery personnel); passenger cabin safety monitoring in public transportation (buses, subways, ride-hailing services, taxis); supervision of teaching behavior in educational institutions (kindergartens, special schools, universities for the elderly); seamless recording in judicial auxiliary venues (mediation rooms, interrogation rooms, community correction points); customer movement and dwell time analysis in retail stores; attendance and emergency response in office spaces; human-machine collaborative safety monitoring in factory workshops; abnormal behavior early warning in hotel rooms; home safety protection for young people living alone, people with disabilities, and patients with chronic diseases; remote medical consultation assistance; smart home context sensing; and any behavioral monitoring needs involving bedrooms, bathrooms, dressing areas, and the surrounding areas of private living spaces. Any environment where both behavioral visibility and personal dignity must be considered can be considered a natural extension of this application.
[0220] In terms of hardware implementation, the edge AI processing unit can run on various chip platforms, including SoCs or dedicated accelerators supporting neural network inference from Rockchip, Allwinner, HiSilicon, Ankai, Qingke, NVIDIA Jetson, Qualcomm QCS, and Intel Movidius. Image sensors can employ global shutter CMOS, rolling shutter APS, event vision sensors (EVS), or hybrid dual-mode sensors, with interface types covering MIPI CSI-2, SLVS-EC, LVDS, and USB3 Vision. The physical switching mechanism can be implemented using physical structures such as mechanical microswitches, capacitive touch sensors, Hall effect magnetic induction elements, and opto-isolated buttons. The core requirement is that the signal path bypasses the operating system and network protocol stack, ensuring that mode switching cannot be simulated, intercepted, or tampered with by remote software commands.
[0221] Furthermore, the technical essence of this invention can be extended to multimodal fusion scenarios. For example, when the system is equipped with infrared thermal imaging, millimeter-wave radar, ultrasonic, or ToF depth sensors simultaneously, the visualization generation can fuse multi-source sensing data, but it still needs to meet the premise that "the original visible light image does not leave the device." Even if the final output human figure is driven by radar point clouds or heat maps, as long as it works in conjunction with the local visible light feature extraction process and the original visible light frame is not transmitted externally, it is still within the scope of this application.
[0222] Furthermore, this solution supports concurrent multi-user identification and differentiated visualization. The system can assign desensitized graphics with different colors, shapes, labels, or animation styles to different individuals to distinguish between caregivers and patients, family members and visitors, and staff and service recipients. This identity differentiation logic can be implemented based on ID binding after facial desensitization, wristband beacons, check-in order, or spatial location inference. As long as it does not rely on the transmission of the original facial image, it complies with the privacy principles of this invention.
[0223] In the long run, with the improvement of edge computing capabilities and the development of neuromorphic vision technology, the architecture of this invention can also be adapted to new sensors such as event cameras. In this mode, the APS imaging pathway can be powered off by hardware in privacy mode, with only the EVS module operating, outputting a sparse event stream for reconstructing motion contours. Although this implementation does not have traditional "frame images," its output is still irreversibly desensitized behavioral visualization content, and it is controlled by physical buttons, perfectly matching the core technology of this application.
[0224] The scope of protection for this application shall be determined by the technical features defined in its claims, covering all variations that are substantially the same as the aforementioned core mechanism in terms of function, effect, and implementation logic. Any system built upon the technical closed loop of "physical switch triggering + local feature extraction + original image not leaving the device + desensitized visualization output," regardless of any functional adjustments made to its graphical appearance, data dimensions, rendering engine, communication protocol, storage strategy, power supply method, networking capabilities, or human-computer interaction, as long as its essential technical effect still lies in achieving continuous, usable, and irreversible identity visualization through physically controllable local intelligent processing, even if not explicitly listed in the embodiments of the specification, shall be considered an equivalent implementation of this application and shall be fully protected by law.
Claims
1. A smart visual terminal with privacy protection function, characterized in that, include: The image acquisition module is used to acquire real-time video streams; A physical switching button is located on the terminal body and is used to receive user trigger signals to switch between high-definition mode and privacy mode, and generate corresponding mode status signals. An edge AI processing unit, integrated inside the terminal, is used to perform real-time analysis of the video stream in the privacy mode, identify human bodies in the image, and generate visual graphics to replace real human bodies. The background capture module is used to capture a frame of environmental image as a static background image when switching to privacy mode or when the device is first started. The video synthesis module is used to overlay the visualized graphics onto the static background image in the privacy mode to generate a desensitized privacy video stream; The video output module is used to selectively output the original high-definition video stream or the privacy video stream according to the mode status signal; The physical switch button is a physical hardware switch, and its status signal is generated through a direct hardware connection path and cannot be modified by remote software commands, network protocols or operating system kernels.
2. The intelligent vision terminal according to claim 1, characterized in that, The visualization is generated by the edge AI processing unit based on a human pose estimation model, which is configured to output the coordinates of multiple joints representing the human body structure and to present the joints as visual elements in the privacy video stream in at least part.
3. The intelligent vision terminal according to claim 2, characterized in that, The visualization includes a skeleton line drawing, which is formed by connecting the coordinates of the joint points according to a preset or dynamic topological relationship of the human skeleton.
4. The intelligent vision terminal according to claim 1, characterized in that, The visualization graphics also include any one or more of the following combinations: cartoon human figures, simplified human outlines, three-dimensional mesh models, motion heatmaps, particle systems, symbolic icons, or geometric projections.
5. The intelligent vision terminal according to claim 1, characterized in that, It also includes a status indicator device, which displays different visual or auditory states based on the mode status signal to indicate to the user whether the current mode is HD or privacy mode.
6. The intelligent vision terminal according to claim 1, characterized in that, The edge AI processing unit is also configured to perform at least one behavioral analysis task on the device side, including fall detection, long-term static alarm, area intrusion judgment, or people counting, and generate local alarm signals or structured data based on the analysis results, without relying on cloud computing resources.
7. The intelligent vision terminal according to claim 1, characterized in that, The background capture module is configured to automatically recapture the static background image when the device is first started, each time it is switched to privacy mode, or when a significant change in scene lighting, layout, or object distribution is detected.
8. The intelligent vision terminal according to claim 1, characterized in that, The video output module supports transmitting video streams to external devices via wired or wireless communication interfaces, including Ethernet, PoE, Wi-Fi, 4G, 5G, or Bluetooth, and the output content is always strictly consistent with the current hardware state of the physical switch button.
9. The intelligent vision terminal according to claim 1, characterized in that, The terminal also includes at least one auxiliary sensing module, which includes at least one of millimeter-wave radar, infrared thermal imaging sensor, event camera, ToF depth sensor, or ultrasonic sensor; wherein, the image acquisition module serves as the main sensing unit, used to provide a real-time video stream containing visible light information; in the privacy mode, the edge AI processing unit extracts human-related features based on the video stream and combines them with the sensing data output by the auxiliary sensing module to jointly generate fusion features for enhancing the visualization graphics; the final output privacy video stream is generated by the fusion features and does not contain the original high-definition video frames from the image acquisition module; the physical switching button uniformly controls the working state and data output strategy of the image acquisition module and the auxiliary sensing module.
10. A video output control method based on the intelligent visual terminal according to any one of claims 1-9, characterized in that, include: In response to physical button operations, it switches working modes and generates mode status signals; If in HD mode, the original HD video stream will be output. If in privacy mode: Obtain a static background image with no people moving around; The real-time video stream is analyzed using an edge AI model to identify human bodies and generate visual graphics, which include the coordinates of human body joints presented in a visual manner. The visualized graphics are overlaid onto the static background image to form a private video stream; Output the privacy video stream; In this privacy mode, the original high-definition video stream is neither written to the storage medium nor transmitted through any network channel.