Dual-mode intelligent nursing camera supporting privacy protection mode switching and event stream processing architecture thereof

By using a single-chip dual-mode sensor architecture and a physical mode switching switch, the challenge of balancing privacy protection and behavior analysis functions in smart care cameras has been solved. This allows for the provision of visible light video and grayscale human body contour video without changing the existing monitoring network, reducing costs and ensuring the security of privacy data. This is suitable for smart care needs in both fixed and mobile scenarios.

CN121815072APending Publication Date: 2026-04-07YUNNONG (SHENZHEN) TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-01-09
Publication Date
2026-04-07

AI Technical Summary

Technical Problem

Existing smart surveillance cameras, while balancing privacy protection and behavior analysis functions, suffer from problems such as high hardware costs, increased device size, inability to seamlessly integrate with existing monitoring networks, and easy leakage of privacy data. They are particularly difficult to meet the needs of privacy protection and authenticity verification in both fixed and mobile scenarios.

Method used

It adopts a single-chip dual-mode sensor architecture, integrating an active pixel sensor (APS) module and an event vision sensor (EVS) module. It shares a data channel through a MIPI CSI-2 interface. The edge processing unit switches imaging functions in different modes to generate visible light video or grayscale human body contour video. Privacy protection is ensured through a physical mode switch, and it is compatible with existing NVR systems.

Benefits of technology

It reduces hardware costs without changing the existing NVR access architecture and cabling, ensures irreversible anonymization of privacy data, supports real-time push or offline synchronization of structured alarm events, is suitable for both fixed and mobile scenarios, and improves the reliability and usability of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure FT_1
    Figure FT_1
  • Figure FT_2
    Figure FT_2
  • Figure FT_3
    Figure FT_3
Patent Text Reader

Abstract

The invention discloses a dual-mode intelligent visual device supporting privacy protection mode switching. The device integrates an APS module and an EVS module, a first mode outputs a visible light video stream, a second mode hardware is powered off the APS, a gray contour video is generated only based on an EVS event stream, target behavior analysis is executed, and original images or event data are not output. The built-in lightweight AI model can detect the falling of the human body in real time, and the alarm is pushed through the ONVIF. The video adopts a standard coding and streaming media protocol, is compatible with a mainstream NVR, supports network-free local storage and physical mode switching, completely eradicates biological recognition information generation from the source, and is suitable for a high-privacy sensitive scene.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent visual monitoring technology, and in particular to a dual-mode intelligent care camera that supports privacy protection mode switching and its event stream processing architecture. Background Technology

[0002] High-definition network cameras based on IP networks are now widely used. The vast majority of these cameras support the ONVIF Profile S standard and can be directly connected to network video recorders (NVRs) compatible with mainstream manufacturers, forming a monitoring technology system of "front-end acquisition—central storage—remote access." This system boasts advantages such as low deployment cost, convenient operation and maintenance, and strong compatibility, and has become the current infrastructure for home security.

[0003] However, in key application scenarios such as fall detection and nighttime activity monitoring, the continuous recording and transmission characteristics of traditional visible light video can easily raise privacy concerns among users. For private areas such as bedrooms and bathrooms, most elderly people and their families are reluctant to use traditional cameras, even if these devices have AI behavioral analysis capabilities. This issue results in deployed devices often being out of service or only applicable to non-sensitive areas, thus limiting the effectiveness of smart care systems.

[0004] Meanwhile, portable video recording devices are increasingly being used in grassroots public services and mobile work scenarios. For example, family doctors and caregivers at long-term care insurance designated institutions are generally equipped with service recorders for home visits to provide highly privacy-sensitive services such as chronic disease follow-up, pressure ulcer dressing changes, and bathing assistance; community service providers rely on video recorders to preserve video evidence of service activities; and dashcams are also widely used in the transportation sector to record accident events.

[0005] While such devices solve the problem of "service process traceability", they also fall into the dilemma of privacy compliance.

[0006] To alleviate these privacy concerns, some manufacturers have attempted to use non-visible light alternatives for sensing. For example, they use infrared thermal imaging or millimeter-wave radar for human detection. However, infrared thermal imaging devices are expensive and lack sufficient spatial detail, making it difficult to accurately identify actions such as falls. While millimeter-wave radar can penetrate clothing, it cannot provide visual verification, resulting in a high false alarm rate and requiring caregivers to conduct on-site checks, increasing maintenance burden and response time. Another approach uses a dual-camera architecture, with one active pixel sensor (APS) for regular video surveillance and the other an event vision sensor (EVS) for behavior analysis in privacy mode. However, this solution requires two independent optical systems and image signal processors (ISPs), leading to increased hardware costs and device size. It also lacks direct compatibility with existing monocular cameras and NVRs, requiring rewiring, drilling, and channel configuration, making modification difficult and costly, hindering large-scale application in existing surveillance networks. Existing privacy modes also suffer from inherent flaws in their software control mechanisms: the APS module remains powered on, with video stream output only disabled through firmware configuration. This results in the original visible light image data remaining in the device's memory or cache. If security incidents such as firmware vulnerabilities, remote attacks, or log leaks occur, this data may be illegally extracted or recovered, failing to fundamentally meet the core requirement of privacy protection and exhibiting significant technical security shortcomings.

[0007] Therefore, the industry urgently needs a monitoring camera that can achieve the following technical effects in both fixed installation and mobile portable scenarios, without changing the existing ONVIF / NVR access architecture, adding cabling or servers, and with controllable cost increases: providing high-definition visible light video to support authenticity verification in non-sensitive services; completely cutting off the acquisition of original images in high-privacy exposure operations, outputting only grayscale human outline video that cannot reveal the identity; ensuring that the privacy mode cannot be bypassed by software through a physical switch; and being compatible with the existing video management ecosystem, supporting real-time push or offline synchronous structured alarm events. Currently, no technical solution can simultaneously meet all of the above requirements. Summary of the Invention

[0008] This invention aims to solve the technical challenge of balancing privacy protection and behavior analysis functions in existing camera devices within the widely deployed high-definition network surveillance environment. Currently, homes and elderly care facilities commonly use network cameras supporting the ONVIF Profile S standard connected to mainstream NVR systems. However, users are reluctant to activate the monitoring function in areas such as bedrooms and bathrooms due to concerns about visible light video leaking private information, resulting in a large number of idle devices. Existing alternatives either rely on high-cost dual-sensor architectures or only block image output through software, failing to eliminate the risk of raw data leakage at the hardware level and lacking seamless compatibility with existing systems.

[0009] Meanwhile, in mobile operations and temporary service scenarios, portable video recording devices also face a sharp contradiction between authenticity verification and privacy compliance. For example, when family doctors or caregivers use service recorders to perform procedures such as bathing or changing dressings for pressure ulcers at home, recording high-definition videos would expose the patient's physical privacy; if recording is completely turned off, service credentials cannot be provided to medical insurance or regulatory platforms. Similarly, in passenger transportation such as ride-hailing services, taxis, or customized buses, in-vehicle recorders need to monitor driver behavior and passenger safety, but collecting passenger faces, clothing, or abnormal postures can easily cross the line of personal information protection. Existing portable devices mostly use a single APS sensor and lack physical-level privacy isolation mechanisms. Their "privacy mode" usually only hides the preview screen, while the original image is still cached in memory or on a memory card, posing a risk of being recovered or illegally exported.

[0010] To address the technical challenges of coexisting fixed and mobile scenarios, this invention provides a dual-mode smart monitoring camera and its event stream processing architecture that supports privacy protection mode switching. It can achieve strong privacy monitoring by ensuring that the original visible light image does not leave the device and the output content is irreversibly anonymized, without changing the existing NVR access method or adding new wiring or dedicated servers.

[0011] The camera includes an image sensor chip, an edge processing unit, a video rendering module, and a network communication interface. The image sensor chip integrates an Active Pixel Sensor (APS) module and an Event Vision Sensor (EVS) module, both sharing the same pixel array and transmitting data to the edge processing unit via a shared physical channel through an embedded data packet via a MIPI CSI-2 interface. In the first operating mode, the edge processing unit enables the APS module to generate a visible light video stream; in the second operating mode, it disables the APS module's imaging function and only enables the EVS module to perform human behavior analysis. In the second operating mode, the video rendering module converts the event stream data output by the EVS into a grayscale human contour video stream, which does not contain identifiable personal information such as facial features, skin color, or environmental textures. The network communication interface outputs a visible light video stream in the first operating mode and only outputs a grayscale human contour video stream and structured alarm events in the second operating mode, while prohibiting the output of the original APS image or the original EVS event stream.

[0012] It should be emphasized that the "camera" is not only applicable to fixed network cameras, but can also be packaged as a handheld service recorder, a vehicle passenger monitoring terminal, or a head-mounted recorder. Its power supply can be PoE, a DC adapter, or a rechargeable battery, and its network communication interface can be Ethernet, Wi-Fi, 4G / 5G, or offline storage media. All of these variations do not depart from the technical essence of this invention.

[0013] The edge processing unit incorporates a lightweight AI inference model for real-time detection of human fall events based on event streams in the second operating mode. Upon detection, it generates a structured alarm event, which can be pushed to a network video recorder via the ONVIF event protocol or uploaded to a cloud monitoring platform via secure protocols such as HTTPS / CoAP. In the second operating mode, the raw image data from the APS module is prohibited from being output to the network communication interface, preferably by completely stopping its operation through hardware power-off, thereby physically blocking visible light imaging. The camera also features a physical mode switch, allowing users to manually select between the first and second operating modes. This switch is independent of the operating system and network connection, providing firmware-independent operational redundancy.

[0014] Both visible light video streams and grayscale human contour video streams use standard video encoding formats and are provided through standard streaming media protocols, ensuring plug-and-play compatibility with mainstream NVRs such as Hikvision and Dahua. They can directly replace existing monocular cameras without requiring changes to the backend configuration. For portable devices, contour video and alarm events can be temporarily stored in local encrypted storage (such as a MicroSD card or eMMC), and automatically synchronized to a remote monitoring platform after the device is connected to a workstation or network, achieving flexible deployment of "offline acquisition and online archiving."

[0015] The edge processing unit supports offline operation, enabling it to perform event stream processing, fall detection, alarm generation, and local storage even without an internet connection. It also caches alarm-related video clips via a MicroSD card. The event stream data includes pixel coordinates X and Y, a timestamp t, and polarity pol. The video rendering module maps the event stream into continuous grayscale contour frames based on a spatiotemporal aggregation algorithm. The edge processing unit is implemented using a SoC chip with an integrated hardware video encoder and MIPI CSI-2 interface, balancing processing performance and power consumption control.

[0016] This invention achieves several beneficial effects through the aforementioned technical solution. The single-chip dual-mode sensing architecture significantly reduces hardware cost and size, avoiding the redundant design of traditional dual-camera solutions. In privacy protection mode, only grayscale human contour video reconstructed from the event stream is output; the original visible light image is neither transmitted nor stored, fundamentally meeting the minimum necessary requirements for personal information processing and irreversible anonymization. Because it retains standard video encoding and streaming media protocol interfaces, this camera can seamlessly integrate with existing ONVIF-compatible NVR systems, eliminating the need for rewiring, channel configuration, or new server deployment, significantly lowering the upgrade threshold. The physical mode switch and APS hardware power-off mechanism together provide dual security. Even if the firmware is tampered with or the network is attacked, users can still manually force entry into privacy mode, ensuring privacy is not accidentally leaked. Furthermore, offline alarms and local storage capabilities enable reliable system operation in weak or offline environments, improving practicality and robustness in diverse scenarios such as elderly care, in-home services, and vehicle monitoring. The overall solution achieves true synergy between privacy and functionality in both the already saturated high-definition surveillance network and the emerging mobile recording ecosystem with extremely low incremental costs, demonstrating significant industrial application value. Attached Figure Description

[0017] Figure 1 This is a system architecture block diagram of the dual-mode intelligent care camera of the present invention.

[0018] Figure 2 This is a data processing flowchart of the present invention in the second working mode.

[0019] Figure 3 This is a schematic diagram of the structure and control logic of the physical mode switching switch of the present invention. Detailed Implementation

[0020] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0021] like Figure 1As shown, the dual-mode intelligent monitoring camera of the present invention includes an image sensor chip 101, an edge processing unit 102, a video rendering module 103, a network communication interface 104, and a physical mode switching switch 105. The image sensor chip 101 integrates an active pixel sensor (APS) module 106 and an event vision sensor (EVS) module 107. These two modules share the same pixel array and share the same physical channel via an embedded data packet through a MIPI CSI-2 interface, synchronously transmitting visible light image frames and event stream data to the edge processing unit 102.

[0022] In the first operating mode, the edge processing unit 102 activates the APS module 106, receives its output visible light image frames, encodes them, and outputs a standard high-definition video stream through the network communication interface 104 for routine monitoring. In the second operating mode, the edge processing unit 102 controls the APS module 106 to stop participating in imaging or data output. This can be achieved, for example, by cutting off its power supply or by configuring its internal registers to disable the generation and transmission of raw image data, thus ensuring that no visible light images are acquired or output in privacy protection mode. Simultaneously, the EVS module 107 remains operational, continuously outputting event stream data. This event stream includes pixel coordinates X and Y, a timestamp t, and polarity pol, which the edge processing unit 102 uses to call its built-in lightweight AI inference model for real-time human fall detection. Once a fall event is detected, a structured alarm event is generated and pushed to the network video recorder via the ONVIF event protocol through the network communication interface 104.

[0023] Simultaneously, the edge processing unit 102 sends the event stream data to the video rendering module 103, which maps the event stream into continuous grayscale human contour video frames based on a spatiotemporal aggregation algorithm, forming a grayscale human contour video stream. This video stream does not contain facial features, skin color, or environmental textures from the original image and cannot reconstruct personal identification information. In the second operating mode, the network communication interface 104 only outputs this grayscale human contour video stream and structured alarm events, prohibiting the output of the original APS image or the original EVS event stream, ensuring that sensitive data does not leave the device.

[0024] The physical mode switch 105 is a three-position mechanical toggle switch with its contacts hardwired to a dedicated GPIO pin of the edge processing unit 102. Users can manually select the operating mode, such as "VIDEO," "PRIVACY," or "AUTO." Because this switch does not rely on the operating system or network communication, even if the device firmware malfunctions, it can still force the device into a privacy protection state through physical operation, thus improving system reliability.

[0025] Furthermore, the network communication interface 104 supports H.264 / H.265 standard video encoding and RTSP and ONVIF Profile S streaming media protocols, allowing this camera to directly replace existing monocular cameras without requiring changes to the NVR configuration. The device also supports local storage via MicroSD card, enabling the caching of alarm-related video clips even in offline environments, with automatic uploading upon reconnection.

[0026] In one possible implementation, the dual-mode smart monitoring camera of the present invention uses the Rockchip RV1106 SoC as the edge processing unit 102. According to Rockchip RV1106 Datasheet Rev 1.7, this chip adopts a 22nm process, integrates a dual-core ARM Cortex-A7 processor with a maximum clock speed of 1.2GHz, integrates an NPU unit with 0.5 TOPS computing power, supports H.264 / H.265 video encoding, and has a MIPI CSI-2 interface, which meets the requirements of the present invention for low power consumption and high integration edge computing.

[0027] The image sensor chip 101 uses the Ruisizhi APX003CE, a fusion image sensor that integrates an active pixel sensor (APS) module 106 and an event vision sensor (EVS) module 107, both sharing the same pixel array. According to the APX003CE Datasheet V1.6, its optical pixel area is 3296H × 2480V, with a total pixel count of 3408H × 2600V, and it supports multiple operating modes. In this embodiment, the APX003CE is configured in Mode 2: 8MAPS + 0.5M EVS fusion output mode, where the APS outputs a 3264×2448 resolution Bayer image, and the EVS outputs an event stream with a resolution of 640×480.

[0028] The APX003CE connects to the CSI receiver controller of the RV1106 via a single MIPI CSI-2 4-lane interface. According to page 25 of the APX003CE datasheet, APS and EVS data can be output in three ways; this implementation uses mode 2: APS and EVS use different Virtual Channels but share the MIPI interface of the APS. Specifically, APS image frames are transmitted via Virtual Channel 0, and EVS event streams are transmitted via Virtual Channel 1; both are encapsulated in embedded data packets within the same physical channel.

[0029] The physical mode switch 105 is a three-position mechanical toggle switch with its common terminal grounded and its two normally open terminals connected to the GPIO0 and GPIO1 pins of the RV1106, respectively. Users can select "VIDEO," "PRIVACY," or "AUTO" mode by toggling the switch. The RV1106 reads the GPIO level status during system startup and loads the corresponding operating configuration accordingly.

[0030] In "PRIVACY" mode, the Linux system running on the RV1106 sends control commands to the APX003CE via the I²C interface. According to page 16 of the APX003CE datasheet, the I²C control interface is enabled when the CFG_MODE pin is low, supporting a communication rate of 1Mbps. The system writes configuration to the APX003CE internal registers, disabling the imaging function of the APS module 106. Specifically, this includes disabling the APS exposure control register, turning off the horizontal / vertical sync signal output, and placing the APS data path in a high-impedance state. Simultaneously, the EVS module 107 remains active, continuously outputting event data including pixel coordinates X and Y, timestamp t, and polarity pol.

[0031] The RV1106's MIPI CSI-2 controller is configured to receive only packets from Virtual Channel 1, thereby filtering out all APS image frames at the hardware level. The EVS event stream is written to a dedicated memory buffer by the DMA engine and processed by the video rendering module 103 called by the CPU core. The rendering module uses a spatiotemporal aggregation algorithm to accumulate events within the past 16 milliseconds according to their polarity, generating a 640×480 resolution grayscale human body contour frame.

[0032] Meanwhile, the NPU loads a lightweight fall detection model to perform real-time analysis of the contour frame sequence. When a fall event is detected, the system generates an ONVIF-compatible structured alarm and pushes it to the NVR via network communication interface 104. The entire unit supports local storage via MicroSD card to cache video clips before and after the alarm in the event of a network outage.

[0033] During the entire “PRIVACY” mode operation, the raw APS image is neither captured nor entered into the RV1106 memory system. The raw EVS event stream is only processed temporarily in a protected memory area and is not output externally. The network interface only provides grayscale contour video reconstructed from the event stream, ensuring that user privacy is not compromised.

[0034] This solution fully leverages the dual-mode fusion capability of APX003CE and the efficient edge processing performance of RV1106, achieving a unified approach to strong privacy protection and behavioral analysis functions while keeping costs under control.

[0035] In one possible implementation, the dual-mode smart care camera of the present invention uses Rockchip RK3588 as the edge processing unit 102. According to the RK3588 Datasheet v1.4 released by Rockchip, this chip uses an 8nm LP process, integrates a quad-core ARM Cortex-A76 plus a quad-core Cortex-A55 big.LITTLE CPU architecture, with a maximum clock speed of 2.4GHz, a built-in NPU unit with 6 TOPS computing power, and supports 8K video encoding and decoding in multiple formats such as H.264 / H.265 / VP9 / AV1. It has powerful edge AI inference and multi-channel video processing capabilities, which is suitable for the complex behavior analysis needs of the present invention in high-end smart elderly care or rehabilitation center scenarios.

[0036] The image sensor chip 101 still uses the RK3588 APX003CE, whose internal APS module 106 and EVS module 107 share a pixel array. Visible light image frames and event stream data are output via the MIPI CSI-2 interface in different Virtual Channel modes. The RK3588 is equipped with two independent MIPI CSI-2 receiver controllers. In this embodiment, one of them is connected to the APX003CE and configured to simultaneously receive APS data from Virtual Channel 0 and EVS data from Virtual Channel 1.

[0037] The physical mode switch 105 is a three-position mechanical toggle switch connected to the GPIO pin of the RK3588. Users can manually select "VIDEO" (video mode), "PRIVACY" (privacy mode), or "AUTO" (automatic mode). In "PRIVACY" mode, the system sends a command to the APX003CE via I²C to close the imaging path of the APS module 106, retaining only the event stream output from the EVS module 107. At this time, the original visible light image is not acquired and does not enter the RK3588 memory subsystem.

[0038] Thanks to the RK3588's NPU computing power of up to 6 TOPS and large-capacity on-chip cache, the edge processing unit 102 can run multiple lightweight AI inference models in parallel in the second working mode to perform multi-dimensional real-time analysis of event stream data. First, the system loads a fall detection model optimized based on a spatiotemporal graph convolutional network. The input is the voxelized event tensor accumulated over the past 500 milliseconds, and the output is a sequence of human posture key points and the probability of falling. When the probability exceeds a preset threshold, a structured alarm event conforming to the ONVIF standard is immediately generated and pushed to the network video recorder.

[0039] In addition, the system can simultaneously run a crowd density estimation algorithm. By clustering and counting movement trajectories in the event flow, it outputs the number of people in the room and an activity heatmap in real time to determine whether abnormal gatherings or prolonged periods of vacancy have occurred. In rehabilitation training scenarios, the system loads a standardized rehabilitation posture assessment model. This model is trained based on a large number of clinical rehabilitation movement samples and can compare the user's leg raises, knee flexions, standing balance, and other movements in real time, outputting a performance score and deviation prompts, and encrypting and storing the results on a local MicroSD card.

[0040] Furthermore, the RK3588's powerful multi-tasking capabilities also support enabling auxiliary facial recognition in specific authorized areas. For example, in the common activity area of ​​a nursing home, the system can briefly enable APS imaging in "VIDEO" mode after obtaining user authorization, and call the facial recognition model for identity verification, used for check-in or service calls; however, once switched to "PRIVACY" mode or entering the bedroom area, APS is immediately powered off, and the facial recognition function is automatically disabled to ensure absolute anonymity in private spaces.

[0041] All analysis results derived from the event stream, including fall alarms, people flow statistics, and rehabilitation scores, are output in structured JSON or ONVIF event format via the network communication interface 104. The video stream remains only a grayscale human silhouette, containing no identifiable biometric features. The device supports dual gigabit Ethernet ports and PCIe 3.0 expansion, allowing simultaneous connection to an NVR and a local health management system for multi-terminal collaboration of care data.

[0042] like Figure 2As shown, in the second operating mode of the present invention, the EVS module 107 in the image sensor chip 101 continuously outputs event stream data 201. This event stream data 201 contains asynchronous event packets sorted by timestamps. Each event packet includes at least pixel coordinates X and Y, a timestamp t, and a polarity field pol. The event stream data 201 is transmitted to the edge processing unit 102, where a lightweight AI inference model loaded internally performs fall detection 202. This model is based on a structure optimized by pruning a spatiotemporal graph convolutional network. The input is a voxelized aggregated event tensor from the past 500 milliseconds, and the output is a sequence of human pose key points and a fall probability value. When the fall probability value exceeds a preset threshold, the system determines that a fall event 207 has occurred and immediately generates a structured alarm event 205. The structured alarm event 205 conforms to the event message format defined by the ONVIF Core Specification, including an event type identifier, timestamp, confidence level, and a unique device ID, and is pushed to the network video recorder via the network communication interface 104 using HTTP POST or PullPoint subscription.

[0043] Simultaneously, the event stream data 201 is also sent to the video rendering module 103. This module uses a sliding time window mechanism to accumulate and normalize events within the most recent 16 milliseconds according to their polarity, generating a grayscale human contour video stream 203 with a resolution of 640×480. This video stream only retains the human motion contour information and does not include visual elements that can identify an individual, such as facial features, skin color, and clothing texture. After being encoded with H.264, the grayscale human contour video stream 203 is output as an contour video 204 through the same network communication interface 104 in RTSP stream format for remote real-time viewing.

[0044] During the entire second operating mode, the APS module 106 is explicitly prohibited from outputting raw image data 206. This prohibition can be implemented through hardware or software: for example, the edge processing unit 102 sends a register configuration instruction to the image sensor chip 101 to shut down the analog front-end power supply of the APS module 106; or configures its internal frame controller to prevent it from generating valid line and field synchronization signals; or discards all data packets from the APS virtual channel at the MIPI CSI-2 receiver. Regardless of the method used, it ensures that the raw visible light image is neither acquired nor enters the system memory or network transmission path, thereby meeting the strong isolation requirements for privacy protection.

[0045] Both the structured alarm event 205 and the contour video 204 are output uniformly through the network communication interface 104, but they carry different semantics: the former is metadata triggered by discrete events, while the latter is a continuous video stream. This separate output architecture allows existing NVRs to simultaneously receive behavioral alarms and anonymous videos without modification.

[0046] In one possible implementation, the dual-mode intelligent monitoring camera of the present invention uses an APX003CE image sensor chip and an RV1106 SoC to build a hardware platform. The APX003CE is connected to the RV1106 via a 4-lane MIPI CSI-2 interface and is configured to share a physical channel with the APS module 106 and the EVS module 107, respectively using VirtualChannel 0 and Virtual Channel 1 to output data. The EVS module 107 outputs an asynchronous event stream with a resolution of 640×480, each event containing X[9:0], Y[8:0], t[11:0], and pol[1] fields, with a maximum event rate of 10,000 events / s.

[0047] The device features a three-position physical toggle switch on its top, connected to the GPIO3 and GPIO4 pins of the RV1106. When the user selects privacy mode, the system writes control instructions to the APX003CE via I²C: write 0x00 to register 0x301A to disable the APS analog front-end power supply, write 0x02 to register 0x300C to disable the APS frame synchronization signal, and write 0x01 to register 0x3020 to enable EVS output. Simultaneously, the RV1106's MIPI CSI-2 controller is configured to receive only Virtual Channel 1 packets, and all APS-related data is discarded at the hardware level.

[0048] EVS events are written to a 4MB memory region at physical address 0x88000000 via DMA. This region is set to non-cacheable and accessible only to the kernel. The video rendering module 103 reads data from this region at a frequency of 30Hz, performs spatiotemporal aggregation of events within the past 16ms using a voxel raster algorithm, generates a 640×480 grayscale human silhouette frame, and writes it to the frame buffer at address 0x89000000.

[0049] Edge processing unit 102 loads an 820KB INT8 quantized fall detection model, with input being a 128×128×5 voxel tensor, and performs inference every 33ms. When the output probability is ≥ 0.92, the system generates an ONVIF-compatible structured alarm event with the topic "tns1:RuleEngine / FallDetection / IsTrue", and pushes it to the NVR's PullPoint subscription endpoint via HTTPS POST. Simultaneously, 15 seconds of H.264 video clips before and after the alarm are encrypted and written to the / secure / alert / directory on the MicroSD card, with file system-level encryption enabled.

[0050] Network communication interface 104 only provides an external RTSP stream generated by contour frame encoding, with the stream address rtsp: / / [IP] / stream1. Its SDP description declares an H.264 profile-level-id of 42001F. The system uses kernel netfilter rules to prohibit any process from accessing the raw image device node, ensuring that the raw APS data cannot be read or transmitted.

[0051] The device supports dual firmware partitioning (A / B), selecting the appropriate partition to load during startup based on GPIO status and performing RSA-2048 signature verification. Actual power consumption is measured at 2.7W, with a fall detection accuracy of 96.4%. It enables robust privacy protection for monitoring functions without altering existing NVR configurations.

[0052] In one possible implementation, the dual-mode intelligent care camera of the present invention is deployed as a service recorder for long-term care insurance or life insurance home services, applied in scenarios where insurance companies provide non-cash payment health care services to insured individuals. In such services, insurance companies do not directly pay claims but instead procure home services from professional care institutions, such as haircuts, nail trimming, foot washing, bathing assistance, perineal cleaning, and bedsore care, to ensure the quality of the insured's recovery. To ensure the authenticity of the service, prevent false declarations, and meet the requirements of the Personal Information Protection Law and medical and health data compliance, the entire service process must be recorded on video; however, different service types have drastically different requirements for video recording modes.

[0053] For services with low privacy sensitivity, such as haircuts, nail trimming, assisting with eating, or measuring blood pressure, the service recorder operates in its first working mode, using the APS module 106 to collect and transmit high-definition visible light video streams to the insurance company's designated monitoring platform. Back-office underwriters, service quality supervisors, or third-party auditors can use this video to verify the service personnel's identity, service duration, operational compliance, and service completion status, serving as a basis for service settlement and quality assessment.

[0054] However, when the service involves scenarios with high privacy exposure, such as assisting with bathing, perineal cleaning, incontinence care, or pressure sore treatment, continuing to transmit raw video would severely infringe upon the insured's dignity and violate the data minimization principle. In this case, the service personnel manually switch to the second operating mode using a physical toggle switch on the device. In this mode, the APS module 106 is immediately prohibited from outputting raw image data, and only the EVS module 107 continuously outputs event stream data. (See attached...) Figure 2 As shown, the event flows through the edge processing unit 102 for real-time human behavior analysis. On the one hand, the video rendering module 103 generates a grayscale human outline video stream, retaining only non-identification information such as the relative position of the service personnel and the insured, limb movements, and service duration. On the other hand, the edge processing unit 102 runs a lightweight AI model in parallel to detect the existence of key service actions (such as "wiping", "rinsing", "turning over"), and generates a structured service event log after identifying valid service behaviors.

[0055] The grayscale human contour video stream and structured service event logs are encrypted and uploaded to the monitoring platform via network communication interface 104. Although back-end reviewers cannot see the original footage, they can confirm through the contour video that the service has occurred, there is no replacement or vacancy, and combine this with fields such as action type, duration, and completion confidence in the structured logs to determine whether the service conforms to the standard process. For example, in a pressure ulcer care scenario, the system can identify a three-stage action sequence of "wound exposure—cleaning—medication application." If the sequence is fully detected, the service is marked as valid; if only a brief approach is detected without any action, a review alarm is triggered.

[0056] This mechanism effectively resolves a core contradiction in the regulation of long-term care services: it satisfies insurance companies' strong verification requirements for service authenticity while providing substantial protection for the privacy of insured individuals in highly sensitive scenarios. Because mode switching relies on a physical switch and is independent of network or operating system, even if the device is attacked by malware, service personnel can manually force entry into privacy mode, ensuring that compliance standards are not breached. Simultaneously, all original APS images are never stored in memory or on storage media in privacy mode, fundamentally eliminating the risk of privacy leaks.

[0057] Furthermore, this service recorder supports automatic preset mode strategies based on service work orders. For example, when the dispatch system assigns a "bathing assistance" task, it can remotely issue a command to preset the device to privacy mode, requiring manual confirmation before the service begins; while a "hairdressing" task defaults to high-definition mode. This flexible dual-mode mechanism allows the same device to cover all categories of on-site services, reducing procurement and management costs while meeting multiple compliance requirements regarding insurance regulation, service quality, and personal privacy.

[0058] In one possible implementation, the dual-mode intelligent care camera of the present invention is configured as a community service recorder, suitable for community service personnel to perform services such as home visits and conflict mediation. In such tasks, the entire service process needs to be recorded to ensure service traceability and compliance. Moreover, the work scenarios often involve the interior space of residents' homes. If high-definition video is recorded throughout the process, it will collect private information such as portraits of irrelevant persons, home furnishings, and medical records, which may lead to privacy disputes.

[0059] To address this issue, the community service recorder features a built-in physical three-position toggle switch, corresponding to "Service Mode," "Privacy Mode," and "Automatic Mode." When community service personnel enter ordinary public areas to provide on-site services, the device is in service mode, the APS module 106 operates normally, and records and uploads high-definition video in real time to the community service management platform for service process documentation and post-event traceability.

[0060] However, when tasks involve in-home scenarios, such as conducting security checks on elderly people living alone, following up with patients with mental disorders, checking the health of people under home quarantine, or mediating neighborhood disputes in residents' homes, community service personnel can manually switch to privacy mode. In this case, as shown in the attached... Figure 2 As shown, the APS module 106 is immediately prohibited from outputting raw image data, and only the EVS module 107 continues to capture the dynamic event stream. The edge processing unit 102 processes the event stream in real time: on the one hand, the video rendering module 103 generates a grayscale human body contour video stream, which only displays the relative position, body movements and interaction duration of community service personnel and residents; on the other hand, the built-in lightweight AI model simultaneously analyzes behavioral semantics, identifies key service actions such as "showing identification", "measuring body temperature", "delivering supplies" and "signing notices", and generates a structured service event log after detecting valid behavior.

[0061] The grayscale contour video and structured logs are encrypted and transmitted to the community service audio and video management platform via a dedicated community service network. Backend management staff can use the contour video to confirm that community service personnel have arrived on-site, have not left their posts without authorization, and have not engaged in inappropriate physical contact. Simultaneously, they can verify the completeness of the service process based on the structured logs. For example, in providing health care services to elderly people living alone, if the system identifies a five-stage action sequence—"dialogue inquiry (to assess mental state) — synchronized video recording and archiving (for subsequent basic analysis) — health status assessment — fall risk reminder — service record"—the service task is marked as compliant. If there is only a brief stop without effective service interaction, a service review reminder is triggered.

[0062] This mechanism effectively solves a dilemma in community door-to-door services: it satisfies the basic requirement of "traceability of service process" while avoiding excessive collection of residents' private information in unnecessary scenarios. Because the privacy mode is triggered by a physical switch and does not rely on operating system or network commands, even if the device is infected with malware or remotely hijacked, community service personnel can manually force the original image output to be cut off, ensuring that residents' privacy is not compromised. All original APS images never enter the memory buffer or local storage in privacy mode, eliminating the risk of privacy data leakage at the source.

[0063] This community service recorder supports linkage with community service terminals. When community service personnel receive a "home visit care" task order through the community service terminal, the system can automatically preset the device to privacy mode. After arriving at the scene, the community service personnel need to manually confirm the start, realizing intelligent matching of "task type - recording mode".

[0064] In one possible implementation, the dual-mode intelligent care system of the present invention is applied to scenarios involving family doctor contract services and collaborative management of home-based hospital beds, serving patients with chronic diseases who have limited mobility, post-operative recoveries, or elderly people with disabilities. The system comprises two types of terminal devices: one is a portable service recorder carried by the family doctor, and the other is a fixed, secure camera pre-installed in the patient's room. Both are based on the same dual-mode imaging architecture, but their usage logic is clearly defined and their data is complementary.

[0065] When a family doctor team makes a home visit as scheduled, the service recorder worn by the doctor is in standby mode by default. Before entering the patient's residence, the doctor manually selects the working mode based on the services provided that day. If the service involves physical examination, chronic disease follow-up, medication guidance, or traditional Chinese medicine therapy (such as acupuncture or massage), since it does not involve high privacy exposure, the device switches to the first working mode, enabling the APS module to collect high-definition video, which is uploaded to the regional health information platform in real time for the quality control department to verify the standardization and duration of the service. For example, when changing dressings for foot ulcers in diabetic patients, the high-definition video can clearly record details such as debridement and dressing application, serving as a basis for medical quality assessment.

[0066] However, if the service involves sensitive procedures such as perineal care, incontinence management, turning over in bed, or observation of pressure sores, the doctor immediately switches the recorder to the second working mode. In this mode, the APS module is prohibited from outputting raw images, and only the EVS module continuously captures the event stream. The edge processing unit generates grayscale human contour videos in real time and simultaneously runs an AI model to identify key actions, such as the "turning over—cleaning—back patting" sequence. Structured service logs, along with the contour videos, are uploaded encrypted. Back-end administrators can confirm that the service actually occurred and the process is complete, but they cannot obtain any visual information that could identify the patient's identity or physical characteristics, thus fully respecting the patient's dignity.

[0067] Meanwhile, fixed, secure cameras deployed in the patient's bedroom or living room operate in a second working mode. This device monitors the patient's daily activities continuously via the EVS module without human intervention. Grayscale contour videos generated by the video rendering module are remotely viewed by family members or community nurses to confirm whether the patient has gotten up on time or exhibited any abnormal behaviors such as prolonged periods of stillness. Once the fall detection model built into the edge processing unit identifies a fall, it immediately sends a structured alert to the family doctor's workstation, community health service center, and emergency contacts, triggering a rapid response mechanism.

[0068] In home-based hospital bed management, this secure camera also supports quantitative assessment of rehabilitation behaviors. For example, for patients after hip replacement surgery, the system can identify the number of times and range of motion they perform the "stand-raise-sit" rehabilitation exercise daily, and generate weekly reports for doctors to remotely adjust treatment plans. All analyses are based on event streams, and raw visible light images are never captured, ensuring patients enjoy unobtrusive monitoring and strong privacy protection in their private spaces.

[0069] The data from both types of devices are integrated on the regional health information platform: service recorders provide a chain of evidence from the "service side," while secure cameras provide continuous "home-based" data, together forming a closed loop for evaluating the quality of home-based care services. The entire system supports medical insurance's control over the authenticity of home-based care services while avoiding resident resistance due to excessive monitoring.

[0070] like Figure 3 As shown in the schematic diagram of the structure and control logic of the physical mode switching switch 105 of the present invention, the switch has three positions and its signal connection relationship with the edge processing unit 102. The physical mode switching switch 105 is a three-position mechanical toggle switch, and its three positions correspond to the first working mode 301, the second working mode 302 and the automatic switching mode 303, respectively.

[0071] The first operating mode 301 enables the joint output of the APS module 106 and the EVS module 107. In this mode, the system acquires and transmits high-definition visible light video streams, suitable for low privacy-sensitive service scenarios such as hairdressing, physical examinations, and medication guidance. The second operating mode 302 prohibits the APS module 106 from outputting raw image data and only allows the EVS module 107 to output event streams. The edge processing unit 102 generates grayscale human contour video streams and performs behavior analysis, suitable for high privacy-exposure scenarios such as perineal care, assisted bathing, and bedsore cleaning.

[0072] The automatic switching mode 303 does not rely on manual operation. Instead, the edge processing unit 102 autonomously determines the appropriate working mode based on preset rules or real-time environmental information. For example, when the service type is detected as "home-based assessment", the system can automatically enter the second working mode; when a task in a public area is identified, it switches to the first working mode.

[0073] The physical mode switch 105 is hardwired to a dedicated GPIO pin of the edge processing unit 102, ensuring that the mode switching command can be read before the operating system starts and enforced at the firmware level. This connection does not go through application layer software, so even if the device is attacked by malicious programs or hijacked by the network, the user can still immediately cut off the original image output by physically switching the switch, ensuring basic privacy and security.

[0074] The entire control logic is implemented by the state machine inside the edge processing unit 102. It loads the corresponding register configuration, AI model and video output strategy according to the GPIO level state, thereby completing mode isolation at the hardware level and effectively supporting the compliance and practicality requirements of this invention in various scenarios such as family doctor home visits, insurance service supervision and community policing.

[0075] In one possible implementation, the secure camera of this invention is deployed in the home of a patient receiving home-based care under the management of a community health service center. Specifically, it is installed on the bedroom ceiling near the foot of the bed, with a field of view covering the bed surface, the passageway to the bathroom, and the living area, ensuring no blind spots for critical activities. The device uses a fully enclosed metal casing with no visible lens markings to avoid causing psychological resistance from the person being cared for. The entire device is built on the APX003CE image sensor chip and Rockchip RV1106SoC, with 4GB of LPDDR4 memory, 32GB of eMMC storage, and a Gigabit Ethernet interface. It supports PoE++ power supply, and the overall power consumption is kept below 2.9W.

[0076] After the device is powered on, the system first reads the state of the physical toggle switch 105, which is hardwired to GPIO3 (first operating mode), GPIO4 (second operating mode), and common ground of the RV1106. Based on the pin level combination, the system loads the corresponding firmware partition and initializes the image sensor.

[0077] When the switch is set to the first operating mode 301, the driver writes register 0x301A=0x03 to APX003CE to enable the APS analog front-end, and simultaneously configures the MIPI CSI-2 controller to receive dual-stream data from Virtual Channel 0 (APS) and Virtual Channel 1 (EVS). The APS module 106 outputs 1080p raw images at 30fps, which are then encoded with H.264 to form the main video stream; the EVS module 107 synchronously outputs an event stream for auxiliary motion analysis. This mode is specifically designed for non-invasive services performed by family doctors or contracted caregivers at home, including but not limited to: measuring blood pressure and blood oxygen saturation, guiding insulin injection, demonstrating rehabilitation exercises, conducting cognitive assessments, changing dressings, or trimming nails. During this process, high-definition video fully records the service personnel's operational procedures, the authenticity of doctor-patient interactions, and the duration of the service. The video stream is pushed to the regional health information platform via the RTSP protocol, with the URL format rtsp: / / [IP] / main, for review by medical insurance auditors. At the same time, the system packages the service start / end timestamp, service type code, and employee ID card OCR recognition results into metadata, and uploads them synchronously with the video frames to form a structured service certificate.

[0078] When the switch is set to the second operating mode 302, the system immediately executes a privacy protection process: writing 0x00 to register 0x301A of APX003CE to cut off the power supply to the APS analog front end; writing 0x02 to 0x300C to disable the frame synchronization signal; and configuring the MIPICSI-2 controller to only parse data packets with VC=1. After this, only the event stream output by the EVS module 107 enters the processing pipeline. The video rendering module performs voxelization aggregation on the events using a 16ms sliding window, generating 640×480 grayscale contour frames at a frame rate of 30fps. This video stream does not contain any visual elements that could identify personal information, such as skin color, facial contours, clothing patterns, or room furnishings, but it clearly reflects changes in human posture, movement trajectory, and relative positional relationship with the caregiver. This mode is strictly limited to highly privacy-sensitive scenarios, such as assisting bedridden elderly with bed baths, perineal cleaning, post-incontinence skin care, pressure ulcer irrigation and dressing changes, catheter maintenance, or changing clothes assistance. In such services, although back-end supervisors cannot see the original footage, they can confirm that the service has been executed and that there are no proxy signatures or false check-ins through contour video. They can also judge the integrity of the service by combining the action sequence logs output by the AI ​​model (such as "approach - wipe - disinfect - leave"). If a fall, prolonged stillness for more than 30 minutes, or abnormal struggling movements are detected, the system immediately generates an ONVIF-compatible structured alarm event, pushes it to the family doctor workstation, community nurse station, and emergency contact mobile APP via HTTPS POST, and triggers a short beep from the local buzzer.

[0079] When the switch is in automatic switching mode 303, the device defaults to the second working mode, but continuously listens for service ticket messages from the regional health information platform. Each ticket includes a service type code (e.g., SVC_001=blood pressure measurement, SVC_015=bathing assistance), the estimated start time, and the service personnel ID. The system's built-in mapping table maps low-sensitivity services (SVC_001–SVC_010) to the first working mode, and high-sensitivity services (SVC_011–SVC_020) to the second working mode. Five minutes before the service starts, the device automatically preloads the corresponding mode's register configuration and AI model; 10 minutes after the service ends, it automatically switches back to the second working mode. In addition, during non-service periods, the system continuously runs fall detection and activity statistics models. If multiple people are found together in a room for an extended period, violent physical conflict, or abnormal cries for help (input via an external microphone) are detected, the EVS sampling rate can be temporarily increased to 10,000 events / s, and a more refined behavior recognition sub-model can be activated, but APS output is still strictly prohibited. All automatic failover decisions are logged to the security log partition for post-event auditing.

[0080] The device supports dual firmware A / B partitions, and performs RSA-2048 signature verification upon each boot to prevent unauthorized firmware from being loaded. All video and event data are encrypted using the national standard SM4 before transmission, and local cache files use the Linux fscrypt file-level encryption policy. Even if the device is physically disassembled, the original image data cannot be recovered.

[0081] In one possible implementation, the service recorder of the present invention is distributed to contracted caregivers at designated long-term care insurance service institutions in a certain city for the purpose of providing home-based care services. The device is a portable handheld terminal integrating an APX003CE image sensor, an RV1106 edge processing unit, a 4G / 5G communication module, and a three-position physical toggle switch; the entire device conforms to recorder industry standards.

[0082] Caregivers receive service orders daily via mobile devices from the long-term care insurance management platform. Each order clearly indicates the service type, service duration, the disability level of the person being cared for, and privacy-sensitive information. For example, order A is for "assisting with eating and oral hygiene" and is marked as low-sensitivity; order B is for "perineal irrigation and catheter care" and is marked as high-sensitivity. Upon arriving at the client's residence, the caregiver manually activates a physical switch on top of the device based on the order details: if it's a low-sensitivity service, it's set to the first working mode; if it's a high-sensitivity service, it's set to the second working mode; if unsure, the caregiver selects the automatic switching mode, allowing the device to determine the appropriate mode based on preset rules.

[0083] When in the first working mode, the APS module 106 normally acquires 1080p@30fps visible light video, and the EVS module 107 synchronously outputs the event stream. After H.264 encoding, the video is uploaded in real-time to the municipal long-term care insurance audio-visual supervision platform via the 4G network, with the stream address rtsp: / / [IMEI] / live. Simultaneously, the edge processing unit runs a facial recognition model to verify the caregiver's identity and uses OCR to recognize the patient's bedside card information, ensuring consistency between "person-certificate-service". In services such as "assisting with hair washing" or "trimming nails," high-definition video fully records the operation steps, consumables used, and service duration, serving as core evidence for medical insurance reimbursement per service. Platform quality control personnel can randomly check videos to assess service standardization and prevent false service claims.

[0084] When switching to the second working mode, the system immediately disables the raw image output of the APS module 106, retaining only the EVS event stream processing path. The video rendering module generates a grayscale human silhouette video, showing only the relative movement of two white human silhouettes, making it impossible to identify faces, genders, body features, or room environment. This mode is specifically designed for services involving bodily exposure, such as "bed sponging," "post-incontinence skin care," "pressure ulcer dressing changes," and "catheter maintenance." During this process, the AI ​​model continuously analyzes the event stream, identifying key action sequences, such as "approaching—wiping—disinfecting—tidying clothes—leaving." If the sequence is completely detected and lasts for ≥8 minutes, a structured service completion log is generated, including action type, confidence level (e.g., 94.7%), start and end timestamps, and device ID; if no valid operation is detected, it is marked as "suspected non-service," triggering manual review. The silhouette video and structured log are encrypted and uploaded to the monitoring platform for auditors to verify the authenticity of the service, without exposing any privacy details.

[0085] In automatic switching mode, the device has a built-in service type mapping table: 22 services, such as "haircut," "blood pressure measurement," and "rehabilitation training guidance," are classified as Mode 1; 15 services, such as "bathing assistance," "perineal care," and "bedsore treatment," are classified as Mode 2. When a caregiver initiates a service, the device reads the "service_code" field from the work order JSON and automatically loads the corresponding mode configuration. For example, when service_code = "SVC_BATH_02," the system preset register 0x301A = 0x00, forcibly entering privacy protection mode. Furthermore, the device supports emergency coverage: in the event of a fall in Mode 2, the caregiver can double-click the power button to temporarily activate APS to record a 10-second high-definition video for medical emergency use. Afterward, the device automatically reverts to privacy mode and stores the video clip separately in a secure partition, accessible only to medical institutions.

[0086] All videos and logs are encrypted using the SM4 national cryptographic algorithm during transmission, and local cache files are protected by a hardware-based Trusted Execution Environment (TEE). After each day's service ends, the device automatically sends the complete service package back to the platform and clears the local cache. This entire mechanism meets the requirements of Article 19 of the "Supervision and Management Measures for Long-Term Care Insurance Services" regarding "traceability, verifiability, and auditability of the service process," and also complies with Article 28 of the "Personal Information Protection Law" regarding the principle of "separate consent + minimum necessity" for the processing of sensitive personal information.

[0087] Through the above methods, the service recorder not only safeguards the medical insurance fund but also respects and protects the personal dignity of disabled elderly people, resolving the fundamental contradiction of "both supervision and privacy" in long-term care services, and becoming a key technical support for the sustainable operation of the long-term care insurance system.

[0088] In one possible implementation, the dual-mode intelligent monitoring camera of the present invention is deployed in the operating vehicles of a compliant ride-hailing platform in a certain city. It is installed in the central rearview mirror area of ​​the windshield, with the lens facing the rear passenger compartment, to ensure the safety of drivers and passengers while meeting the dual regulatory requirements of the "Interim Measures for the Administration of Online Ride-Hailing Services" and the "Personal Information Protection Law." The device adopts a compact cylindrical housing and integrates an APX003CE image sensor chip, an RV1106 edge processing unit, a 4G communication module, a MicroSD card slot, and a three-position physical toggle switch. The entire unit draws power from the vehicle's OBD-II interface and supports a wide operating temperature range of -20℃ to +70℃.

[0089] After power-on, the device defaults to the second operating mode. In this mode, the edge processing unit sends a command to the APX003CE to shut down the analog front-end power supply of the active pixel sensor (APS) module (register 0x301A=0x00), completely stopping its imaging; only the event vision sensor (EVS) module continues to operate, outputting an asynchronous event stream with microsecond-level time resolution. The video rendering module performs voxel aggregation on the events based on a sliding time window, generating a 640×480 resolution grayscale human silhouette video stream. The image only shows the white silhouettes of the passenger and driver and their relative motion trajectories, making it impossible to identify facial features, skin color, clothing patterns, gender, or interior details of the vehicle, fundamentally achieving the "irreversible anonymization" required by Article 29 of the Personal Information Protection Law.

[0090] When the system detects abnormal behavior—such as a passenger suddenly falling, struggling violently, remaining still for an extended period, or engaging in physical altercations—the built-in lightweight AI model immediately triggers a structured alarm event, including the event type (e.g., "FallDetected"), confidence level (92.4%), timestamp, and device ID. This alarm event, along with 30 seconds of preceding and following contour video footage, is encrypted and written to a MicroSD card, then pushed to the city-level ride-hailing regulatory platform via HTTPS POST over a 4G network. It is noteworthy that the original APS image is never captured, cached, or transmitted; even if the device is physically disassembled or the firmware is reverse-engineered, no visible light image can be recovered.

[0091] In specific low-privacy-risk scenarios, such as when a vehicle collision occurs, a passenger files a complaint, or the platform remotely authorizes evidence collection, the driver can manually switch the device to the first operating mode by toggling the physical switch. At this time, the APS module powers on again and outputs a 1080p@30fps high-definition visible light video stream, which is encoded with H.264 and uploaded to the monitoring platform in real time via the RTSP protocol, serving as crucial evidence for accident liability determination or dispute resolution. This mode requires active driver intervention, and the platform logs the mode switching for auditing purposes to ensure it is not misused.

[0092] All video and event data are encrypted and stored using the national standard SM4 algorithm, with fscrypt transparent encryption enabled on the local file system. The device supports offline operation: in areas without network access, such as tunnels and underground parking garages, it can still complete event detection, contour generation, and local caching; alarm content is automatically synchronized after network recovery. Furthermore, both contour video and visible light video are encapsulated in standard MP4 containers and provided externally through an ONVIF Profile S-compatible streaming media interface, allowing seamless integration with the vehicle video management platform designated by the Transportation Commission or third-party NVRs without requiring modifications to existing monitoring infrastructure.

[0093] Through the aforementioned mechanism, the vehicle dashcam operates at the highest privacy level by default during daily operation, outputting only silhouette videos that cannot reveal the passenger's identity, effectively alleviating passengers' resistance to being "recorded throughout their journey." When necessary, a physical switch can temporarily activate the high-definition evidence collection function, balancing security supervision with personal dignity. The entire solution replaces the traditional dual-camera design with a single-chip dual-mode architecture, resulting in low cost increase and comparable size and power consumption to existing dashcams, making it suitable for large-scale deployment and providing a feasible technical path for privacy compliance in the era of intelligent transportation.

[0094] In one possible implementation, the dual-mode intelligent monitoring camera of the present invention is deployed in the center front position of a third-grade classroom in a pilot primary school of a certain province's "Smart Campus Safety Supervision Platform". This device is used to ensure teaching order and student safety while strictly complying with Article 71 of the "Law of the People's Republic of China on the Protection of Minors," which mandates the prohibition of collecting biometric information from minors, and Article 31 of the "Personal Information Protection Law," which provides special protection for the processing of children's information.

[0095] The device features a rounded, cylindrical casing, 60 mm in diameter and 85 mm in height, with a frosted white surface to prevent glare from interfering with normal teaching activities. Internally, it integrates an APX003CE image sensor chip, a Rockchip RV1126 edge processing unit, 8GB eMMC storage, a Wi-Fi 6 communication module, and a hidden three-position physical toggle switch. This switch, located at the bottom of the device, requires a special tool to operate, preventing accidental activation by students or unauthorized personnel. The device draws power from the campus gigabit switch via PoE++ (IEEE 802.3bt) and connects to the school's existing Hikvision DS-9632NI-I16 network video recorder, eliminating the need for additional cabling or the deployment of a dedicated server.

[0096] The system defaults to the second operating mode. In this mode, the edge processing unit sends hardware-level instructions to the image sensor chip to cut off the analog power domain of the active pixel sensor module, completely stopping its photosensitivity. Only the event vision sensor module continues to operate, outputting an asynchronous event stream with a microsecond-level time resolution. The video rendering module, based on a spatiotemporal voxel grid algorithm, aggregates the event stream into a grayscale human silhouette video stream at 30 frames per second and a resolution of 720p. The image only presents the dynamic silhouettes of teachers and students and their relative positions, making it impossible to identify facial features, gender, skin color, school uniform patterns, hairstyles, or details of items in the classroom, fundamentally achieving irreversible anonymization of student identity information.

[0097] The edge processing unit incorporates a lightweight Transformer behavior analysis model, capable of detecting various high-risk events in real time. These include abnormal physical contact between teachers and students, such as pulling or pushing; students suddenly fainting or remaining motionless for more than 120 seconds; and group conflicts involving multiple students, characterized by overlapping rapid movement trajectories. Upon detecting any of these events, the system immediately generates a structured alarm event, including the event type, confidence level, timestamp, classroom number, and unique device identifier. This alarm is then pushed to the school-level network video recorder and the district education bureau's monitoring platform via the ONVIF Profile T event interface. Simultaneously, a 60-second contour video clip before and after the event is encrypted using the SM4 national cryptographic algorithm and written to eMMC storage, retained for 30 days for subsequent auditing.

[0098] When an education bureau supervisor initiates a remote inspection request, or when a major safety incident occurs on campus requiring high-definition video as evidence, a temporary authorization command can be sent through the management platform after approval by both the principal and the director of moral education, activating the first working mode. At this time, the active pixel sensor module is powered on again, outputting a 1080p@25fps, H.265 encoded visible light video stream, which is transmitted in real-time to the monitoring terminal via the RTSP protocol. This mode lasts for no more than 5 minutes, after which it automatically switches back to the second working mode. All mode switching operations are recorded in the device security log and synchronized to the blockchain evidence storage node, ensuring the process is traceable and tamper-proof.

[0099] The physical mode switch uses a hardwired mechanical design, directly connecting to the SoC's GPIO interrupt pin, independent of the operating system or network connection status. Even if the device firmware is tampered with or subjected to a remote attack, teachers can still manually toggle the switch to force entry into privacy mode, ensuring that student images are not accidentally captured. This mechanism provides firmware-independent operational redundancy, constituting a dual security guarantee.

[0100] All video streams are encapsulated in standard MP4 format and output via an ONVIF Profile S-compatible interface, allowing seamless integration with existing school network video recorders for plug-and-play functionality. The device supports offline operation, enabling event detection, profile generation, and local storage even during campus network outages. Once the network is restored, the system automatically retransmits alarm data, ensuring continuous monitoring.

[0101] In one possible implementation, the dual-mode intelligent care camera of the present invention is configured as a portable family doctor follow-up terminal, used by the nursing team of a designated community health service center under the long-term care insurance scheme when providing home-based medical services to disabled or semi-disabled elderly people. This device aims to resolve the fundamental contradiction between the rigid requirements of medical insurance supervision for traceability of service processes and prevention of insurance fraud, and patients' high sensitivity to their physical privacy. It ensures that, without violating Article 28 of the Personal Information Protection Law of the People's Republic of China, which stipulates that the processing of sensitive personal information should follow the principles of "separate consent and minimum necessity," the accurate recording and intelligent verification of service activities are achieved.

[0102] This terminal features a lightweight, handheld design, measuring 115mm × 60mm × 20mm and weighing approximately 170 grams, making it easy for nurses to wear on their chest or place in a medical bag. Internally, it integrates an APX003CE image sensor chip, a Rockchip RV1106 edge processing unit, 32GB of encrypted eMMC storage, a 4G full-network compatible communication module, and a sliding physical mode switch located on the side of the device. The entire device is powered by a rechargeable lithium battery, supporting over 5 hours of continuous operation, and can be quickly recharged via a USB-C interface.

[0103] After arriving at the client's home, the nurse manually sets the working mode according to the day's medical orders. If the service involves low-privacy exposure operations such as blood pressure measurement, blood glucose testing, or medication guidance, the switch is set to the first working mode; if it involves high-privacy sensitive operations such as perineal care, catheterization, pressure ulcer dressing changes, or assisted bathing, the switch is set to the second working mode. This physical switch is hardwired directly to the GPIO pins of the main control chip, independent of the operating system, application, or network status. Even if the device firmware is tampered with or subjected to remote attacks, the user can still force entry into a privacy-protected state through mechanical operation, creating firmware-independent security redundancy.

[0104] In the second operating mode, the edge processing unit sends hardware-level instructions to the image sensor chip, completely cutting off the analog power domain of the active pixel sensor module, thus completely stopping its photosensitivity. Only the event vision sensor module continues to operate, outputting an asynchronous event stream with a microsecond-level time resolution. The video rendering module, based on a spatiotemporal aggregation algorithm, converts the event stream into a 640×480 resolution grayscale human silhouette video stream. The image only shows the dynamic silhouettes of the caregiver and the elderly person being cared for, along with their relative movement trajectories. Facial features, body shape, skin condition, clothing details, or home environment cannot be identified, fundamentally achieving irreversible anonymization of biometric information.

[0105] The edge processing unit incorporates a lightweight temporal behavior analysis model, enabling real-time verification of the completeness and standardization of critical nursing procedures. For example, in catheterization services, the system checks whether the standard procedure of "hand disinfection—draping—catheter insertion—catheter fixation—bed unit tidying" is completed in sequence; in pressure ulcer care, it determines whether steps such as "wound assessment—cleaning and disinfection—dressing change—recording" are missing or abnormally interrupted. If an interruption, sequence error, or potential risk event (such as a sudden fall in an elderly person) is detected, the system immediately generates a structured service log, including service type, action sequence status, confidence level, timestamp, patient ID, and unique device identifier.

[0106] All contour video clips and structured logs are encrypted using the national cryptographic algorithm SM4 and stored locally in eMMC. When the device connects to home Wi-Fi or returns to the community health service center, the data is automatically synchronized to the municipal medical insurance intelligent supervision platform via HTTPS protocol. The platform can then use this data to verify service authenticity, audit fee settlements, and conduct quality assessments. Throughout the entire process, the original visible light images are never captured, cached, or transmitted. Even if the device is lost or the storage medium is illegally extracted, it is impossible to recover any images that could identify personal information, fully meeting the requirements of the "Specifications for the Protection of Personal Information in Medical and Health Institutions" regarding "not retaining original biometric data."

[0107] In special circumstances, such as service disputes, medical insurance audits, or family complaints, with dual authorization from the head of the community health service center and the medical insurance specialist, the equipment can temporarily activate its first working mode during subsequent services. This mode records 1080p H.264 encoded high-definition visible light video and uploads it in real time via a 4G network as legal evidence. Activating this mode requires recording a complete operation log and synchronizing it to a blockchain evidence storage node to ensure controllable permissions and auditable processes.

[0108] The device supports fully offline operation, enabling event detection, profile generation, and local encrypted storage even in environments without network coverage, such as older residential areas, rural areas, or underground parking garages. Data is automatically retransmitted once communication is restored, ensuring the continuity and integrity of service records.

[0109] Through the above mechanism, the follow-up terminal operates at the highest compliance level by default in daily high-privacy services, effectively alleviating the strong resistance of disabled elderly people and their families to being "recorded throughout the process"; when high-definition evidence collection is indeed required, the visible light mode can be flexibly activated through a physical switch and authorization process, taking into account both the effectiveness of supervision and personal dignity.

[0110] In one possible implementation, the dual-mode intelligent monitoring camera of the present invention is configured as a portable, handheld recording terminal for use by staff at community service outlets. It is worn and used when conducting home visits, behavior checks, or living condition assessments of service recipients. This device aims to resolve the prominent conflict between the current community service management requirements for traceable processes and verifiable behavior and the protection of the privacy of service recipients and their cohabiting family members, ensuring compliant recording of the activity status of service targets without collecting raw biometric information.

[0111] This terminal features a compact handheld design, measuring 110mm × 58mm × 18mm and weighing approximately 160 grams, making it easy for workers to wear on their chest or place in a work bag. The device integrates an APX003CE image sensor chip, a Rockchip RV1106 edge processing unit, 16GB of encrypted eMMC storage, a 4G communication module, and a side-mounted sliding physical mode switch. Powered by a rechargeable lithium battery, it supports over 4 hours of continuous operation and can be quickly charged via a USB-C interface.

[0112] Upon arrival at the client's residence, the equipment defaults to operating mode two. In this mode, the edge processing unit sends hardware commands to the image sensor to completely shut down the power domain of the active pixel sensor module, preventing it from generating any visible light images; only the event vision sensor module continues to operate, outputting an asynchronous event stream. The video rendering module, based on a spatiotemporal aggregation algorithm, converts the event stream into a 640×480 resolution grayscale human silhouette video stream. The image only displays the dynamic silhouette of the managed individual and their movement trajectory within the room, without identifying facial features, clothing details, family member identities, or private items within the room, effectively preventing unintentional collection of personal information from non-managed individuals (such as residing relatives).

[0113] If any unusual situations are discovered during the investigation, such as suspected contraband, heated verbal conflicts, self-harm tendencies, or unauthorized departures from residence, staff can immediately manually switch the device to its first operating mode. At this time, the active pixel sensor module is powered on again, outputting a 1080p H.264 encoded high-definition visible light video stream to preserve evidence at the scene. This operation is purely mechanically triggered, independent of network or operating system, ensuring rapid response in emergency situations.

[0114] All data is encrypted using the national cryptographic algorithm SM4 and transmitted directly to the municipal-level grassroots governance information platform via a dedicated APN channel, connecting to the unified regulatory network. Contour videos, structured behavior logs, and high-definition evidence clips are stored hierarchically according to access permissions. The original visible light images do not leave the device's memory throughout their entire lifecycle—that is, they are not cached, written to the storage card, or transmitted externally through any interface, fundamentally eliminating the risk of information leakage.

[0115] The device supports offline operation. Even in areas with no signal or limited network coverage, it can still perform event detection, profile generation, and local encrypted storage; upon returning to a network-enabled environment, it automatically synchronizes data, ensuring the integrity of management records. All mode switches and critical operations generate immutable logs for subsequent auditing and accountability.

[0116] Through the above design, the terminal operates in a privacy-first mode during routine visits, meeting the institutional requirements for regulating service behavior while fully respecting the privacy rights of family members and avoiding complaints or legal disputes caused by excessive data collection. When high-definition evidence collection is indeed required, a visible light mode can be quickly activated via a physical switch, achieving a balance between security and efficiency.

[0117] In one possible implementation, the dual-mode intelligent care camera of the present invention is integrated into an indoor autonomous service robot for home care of elderly people living alone. This robot, developed by a smart elderly care technology company, features automatic navigation, item delivery, fall response, and remote medical assistance. Because this robot frequently needs to enter highly private spaces such as bedrooms and bathrooms to perform tasks, traditional solutions equipped with high-definition cameras are prone to raising user concerns about privacy leaks, and may even result in the device being blocked from access. Therefore, the present invention provides an embedded visual perception module that, while ensuring the robot's environmental understanding and behavior recognition capabilities, ensures at the hardware level that user biometric information is not collected, transmitted, or stored.

[0118] The vision module is mounted on the robot's top gimbal, with a field of view of 120 degrees horizontally and 80 degrees vertically, and can automatically adjust its pitch according to task requirements. At its core is a neuromorphic image sensor chip integrating an active pixel sensor (APS) and an event vision sensor (EVS) (example model: APX003CE, but can also be replaced with a Sony IMX636, PropheseeGen4, or other hybrid event camera with dual-mode output capabilities). Both share the same pixel array and transmit data to the main control SoC (example: Rockchip RV1126) via a single MIPI CSI-2 interface in the form of embedded data packets. The module has a built-in independent power management unit, which can control the power supply domains of the APS and EVS separately.

[0119] The system defaults to the second operating mode. In this mode, the main control unit sends hardware commands to the image sensor to cut off the analog front-end power supply of the APS module (by disabling the LDO_EN pin), completely stopping its photosensitivity; only the EVS module continues to operate, outputting an asynchronous event stream containing pixel coordinates (X, Y), timestamps (t), and polarity (pol) at microsecond-level time resolution. The edge processing unit, based on a spatiotemporal voxelization algorithm, aggregates the sparse event stream into a 640×480 resolution grayscale human silhouette video stream in real time. The image only shows the dynamic silhouette of the elderly person and furniture and their relative motion relationship, and cannot identify faces, skin colors, clothing patterns, body details, or private items in the room, fundamentally achieving irreversible anonymization of identity information.

[0120] The contour video stream is used for two purposes: firstly, for the robot's SLAM mapping and obstacle avoidance—because the event vision is robust to changes in lighting and has no motion blur, it is particularly suitable for stable operation in dimly lit bedrooms or bathrooms at night; secondly, a lightweight AI model analyzes the event stream in real time to detect high-risk behaviors, such as elderly people losing their balance when getting up from the bedside, slipping while using the toilet, or lying still for a long time. Once a fall alarm is triggered (confidence ≥90%), the system immediately generates a structured event, including the event type, location coordinates, timestamp, and robot ID, and pushes it to the family's mobile app and the community health management center via 4G network. At the same time, the 45-second contour video clips before and after are encrypted with the SM4 national cryptographic algorithm and written into the robot's built-in eMMC, retained for 7 days for review.

[0121] When an elderly person initiates a remote consultation request, or when medical staff need to check the status of wound dressings, the first working mode can be temporarily activated with authorization from a family member. At this time, the APS module powers on again and outputs a 1080p@30fps H.265 encoded visible light video stream, which is transmitted to the doctor's end via the RTSP protocol. This mode lasts for no more than 3 minutes, after which it automatically switches back to the second mode. All switching operations are recorded in the security log and synchronized to the blockchain evidence node.

[0122] To prevent privacy mode from failing due to software malfunctions or remote attacks, this module features a hardware-level mode control mechanism. Although the robot as a whole is scheduled by the operating system, the APS power enable signal is controlled by an independent GPIO hardwired connection and can be forcibly set by a physical toggle switch (located inside the maintenance port on the bottom of the robot) or a preset security policy (such as automatically locking the APS when entering a bathroom geofence). Even if the main control system is compromised, the APS will still be unable to power on for imaging, providing firmware-independent ultimate protection.

[0123] All video outputs are encapsulated in standard MP4 containers and provided via an ONVIF Profile S-compatible interface, allowing seamless integration with existing home NVRs or cloud platforms. Even in environments without network access (such as basements or elevators), the robot can still perform event detection, contour generation, and local storage, automatically synchronizing alarm data upon returning to a Wi-Fi coverage area.

[0124] Furthermore, this invention is also applicable to other autonomous mobile vehicles that need to enter private spaces. For example, in emergency scenarios of fires in high-rise residential buildings, indoor reconnaissance drones deployed by fire departments can be equipped with this module, flying in the second mode by default, only outputting heat maps of the outlines of people in the fire scene, avoiding the photographing of exposed bodies of affected residents; logistics robots delivering medicines in hospital wards can automatically switch to outline mode during nighttime ward rounds to prevent disturbing patients' rest and protect their dignity; in high-end hotel room service robots, the second mode is used throughout the cleaning or delivery process, completely eliminating guests' concerns about being "secretly photographed."

[0125] Through the above design, this invention deeply embeds strong privacy protection capabilities into the perception layer of mobile intelligent agents, enabling them to gain both user trust and policy compliance advantages while performing necessary service tasks. The entire solution replaces the fragile traditional dual-camera + masking software solution with a single-chip dual-mode architecture, increasing costs by less than 100 yuan and controlling power consumption to within 2.5W, fully adapting to the payload and battery life requirements of mainstream service robots and small drones. This implementation not only expands the application boundaries of event vision technology in the civilian service field but also creates a new product paradigm of "trustworthy intelligent agents that can enter private spaces," possessing broad industrial application prospects.

[0126] The core patent claim of this application is to provide an intelligent vision device that avoids the generation of biometric information at the imaging source. To clearly illustrate its technical features, the following provides an objective comparison of different technical approaches through specific scenarios.

[0127] In a home care environment, when an elderly person falls in their bedroom or bathroom at night, various visual perception solutions will output different images. Traditional solutions typically rely on active pixel sensors (APS) to first acquire complete visible light image frames. Based on this, some systems use Gaussian blur or mosaic algorithms to mask the face and body areas. The resulting image appears as localized color blocks or low-resolution areas. While identity details are weakened, the original high-resolution image has already been briefly stored in the image signal processor or memory, posing a risk of unauthorized extraction. Furthermore, excessive blurring may lead to loss of behavioral semantics, such as difficulty distinguishing between a fall and a bending-over movement.

[0128] Another common approach utilizes pose estimation algorithms, such as OpenPose, to extract human keypoints and draw connecting lines to form a skeleton map. This type of output only contains a limited number of keypoint coordinates and their connections, without displaying skin, clothing, or facial features. However, the skeleton structure itself still retains biometric information such as height proportions, limb lengths, and motion trajectories, which may be used for identity inference under certain conditions. Furthermore, keypoint detection is prone to failure when the target is occluded, crouched, or in low-light conditions, causing the system to fail to output valid results.

[0129] Some solutions attempt to convert real images into cartoonish or sketch-style human models using generative adversarial networks (GANs). While such outputs possess visual abstraction, their generation process relies on identity-related features in the original image, essentially still presupposing complete image acquisition. Furthermore, these methods are computationally complex and difficult to run in real-time on resource-constrained edge devices.

[0130] Non-visible light imaging technologies are also used in privacy protection scenarios. Thermal imaging generates temperature distribution maps by detecting infrared radiation, outputting pseudo-color images. This method does not rely on visible light, but its spatial resolution is low, its sensitivity drops significantly when the ambient temperature is close to human body temperature, and it is difficult to accurately identify fine movements. Millimeter-wave radar or lidar can generate three-dimensional point cloud data, reflecting the spatial contours of a target. Such point clouds consist of a large number of spatial coordinate points. Although they lack texture information, their geometric accuracy is sufficient to characterize body shape, posture, and even gait features, and they have been regarded as identifiable personal information in some judicial practices. In addition, point cloud data requires specialized software for rendering, making it difficult for ordinary users to understand intuitively.

[0131] The technical approach adopted in this invention differs fundamentally from such solutions. In the second operating mode, the device cuts off the analog power domain of the active pixel sensor via hardware instructions, completely stopping its photosensor function. At this time, only the Event Vision Sensor (EVS) module is active. Instead of outputting a complete image at a fixed frame rate, the EVS asynchronously outputs a sparse event stream containing coordinates, timestamps, and polarity information only when pixel brightness changes exceed a preset threshold. This data stream inherently does not contain static texture, color, or background details; it only reflects the dynamic edge changes.

[0132] The edge processing unit, based on a spatiotemporal aggregation algorithm, reconstructs the event stream within a certain time window into a continuous grayscale contour video. In the output image, the human body appears as a smooth silhouette without facial features or clothing patterns, while the background is suppressed to an extremely low grayscale or transparent state. This contour changes in real time with the movement, clearly expressing behavioral semantics such as getting up, walking, falling, and remaining still, but does not contain any biometric information that can be used for identification. Since the original visible light image is never generated, the system does not have the problems of original data residue or reversible desensitization.

[0133] The event vision sensor features microsecond-level temporal resolution, high dynamic range, and ultra-low power consumption, making it suitable for complex environments such as low light, strong backlight, or rapid movement. Its output contour video can be transmitted via standard video protocols, ensuring compatibility with existing monitoring platforms without requiring modifications to the backend infrastructure.

[0134] This invention, through hardware-level architecture design, enables the generation of desensitized visual outputs suitable for behavior identification without acquiring raw visible light images. This solution achieves an effective balance between privacy protection, behavioral readability, and engineering feasibility.

[0135] This invention employs a dual-mode switching architecture, not only for privacy compliance but also to fully consider engineering efficiency and user-friendliness in practical deployment. Integrating the Active Pixel Sensor (APS) and Event Vision Sensor (EVS) onto the same image sensor chip reduces the physical requirements of the device—requiring only one optical lens, one image processing path, and a single video output interface—significantly decreasing the number of components, circuit complexity, and overall size. Compared to the traditional dual-camera design that deploys a visible light camera and an infrared / event camera in parallel, this solution avoids the technical challenges of multi-lens calibration, multi-channel data synchronization, and dual-encoding stream management, reducing hardware costs by approximately 15% to 20%. It also allows for a more compact device form factor, facilitating embedding in space-constrained devices such as service robots, portable recorders, or wall-mounted terminals. More importantly, this architecture supports on-demand activation of high-definition imaging capabilities. It operates in low-power contour mode in everyday scenarios where evidence collection is not required, and temporarily switches to high-definition mode when legal evidence is needed. This balances the energy efficiency requirements of long-term monitoring with the clarity requirements of occasional evidence collection, representing an efficient and economical technical implementation path.

[0136] To ensure operational reliability and user versatility, this invention specifically introduces a physical button or toggle switch as the trigger mechanism for mode switching. This design is not redundant but rather a practical consideration for actual users. In typical application scenarios such as elderly care, community follow-up, and home rehabilitation, frontline service personnel are mostly caregivers, social workers, or primary healthcare workers, whose digital literacy varies greatly, and who face significant obstacles in using smartphone applications, menu settings, or multi-step operations. Relying on software interfaces for mode selection not only easily leads to privacy leaks due to misoperation but may also delay service processes due to operational failures. The physical switch provides an intuitive, definite, and learning-free operating method: toggle up for HD mode and toggle down for privacy mode, with the status clearly displayed. Even in dim lighting, while wearing gloves, or in tense emergency situations, switching can be completed quickly and accurately. Furthermore, the switch directly controls the sensor power domain via hardwired wiring, without relying on an operating system or network connection. Even if the device crashes or is attacked, it can still force entry into privacy protection mode. This user-centric interactive design significantly lowers the technical barrier to use and improves the system's usability and credibility in real-world social service environments.

[0137] It should be understood that the above embodiments are only used to clearly illustrate the technical concept, implementation path and typical application scenarios of this application, and are not intended to limit the scope of patent protection of this application. Those skilled in the art, based on a full understanding of the core ideas of this application, can reasonably adjust, replace, delete or combine specific hardware selections, algorithm implementations, communication protocols, power supply methods, device forms or deployment environments. All such modifications should be considered as natural extensions of this application and fall within its reasonable protection boundaries.

[0138] While the dual-mode intelligent vision device described in this invention supports behavior analysis of various moving targets (such as pets, moving objects, vehicles, etc.) in its architecture, its system design, algorithm parameters, and visualization strategies have all been deeply optimized for human behavior perception scenarios. The sensitivity threshold, time window length, spatial resolution of the event vision sensor, and the training dataset of the lightweight AI model in the edge processing unit are all configured based on typical human movements (such as walking, sitting, standing up, falling, and prolonged stillness). The grayscale contour video stream generated by the video rendering module prioritizes accurate representation of the dynamic structure of the human torso and limbs in terms of shape preservation, edge continuity, and posture readability, ensuring that non-professional users (such as family members and caregivers) can intuitively identify key behavioral events. Therefore, although the technical solution possesses general target analysis capabilities, its core application scenarios and performance advantages are concentrated in human behavior monitoring and emergency response in highly privacy-sensitive environments.

[0139] It should be understood that the "physical mode switching mechanism" described in this invention does not refer only to a traditional manual mechanical switch, but rather to any mode selection device or interface that does not rely on the operating system software stack and can directly control the working state of the image sensor through hardware signals. Its core feature is that the switching command acts directly on the power domain, enable pin, or configuration register of the image sensor chip in the form of electrical signals, thereby physically determining whether the active pixel sensor (APS) module is powered on or activated, ensuring that the privacy protection mode can still be reliably enabled even if the main control system crashes, the firmware is tampered with, or it is subjected to a remote attack. Typical implementations include, but are not limited to: mechanical toggle switches, slide switches, or push-button switches mounted on the device housing; hard-wired GPIO control signal input ports led out through terminals or pin headers, which can be triggered by an external controller (such as a robot main control board or access control module); or hardware latching circuits driven by a security coprocessor, automatically set in response to preset conditions such as geofencing and time policies. Although the above-mentioned implementation methods differ in their interaction forms or triggering sources, they all share the essential attributes of "direct hardware connection, software independence, and state determination," and belong to the equivalent technical means covered by this invention.

[0140] It is particularly important to emphasize that the "APX003CE" image sensor chip mentioned in this application is merely a specific implementation example for illustrative purposes and is by no means a limitation on the sensing front-end hardware. The core of this application lies in employing a single-chip architecture that integrates an active pixel sensor (APS) module and an event vision sensor (EVS) module within the same pixel array. Both modules share optical paths and physical data channels (such as MIPI CSI-2), thereby achieving compact, low-cost, and low-power dual-mode sensing capabilities. Therefore, any neuromorphic vision sensor chip with dual-mode APS and EVS output capabilities, regardless of its manufacturer, interface type, event encoding format, or internal architecture, can be considered an equivalent alternative to this application.

[0141] Such alternative chips include, but are not limited to:

[0142] Pure event dynamic vision sensors (eDVS) based on asynchronous address event representation (AER) output, such as iniVation’s DAVIS series and Prophesee’s Metavision sensor;

[0143] Digital event vision sensors (dDVS) that employ digital interface optimization and support embedded event packet transmission, such as the Sony IMX636 and Samsung ISOCELL event-based sensor prototypes;

[0144] A hybrid image sensor with frame-based imaging and event stream synchronous output capability, that is, a device that natively integrates an APS pixel array and an EVS photosensitive unit on the same chip, supporting independent or collaborative working modes.

[0145] Although the aforementioned sensors differ in terms of temporal resolution (microseconds to milliseconds), spatial sensitivity (adjustable contrast threshold), low-light response capability, event density control mechanism, power consumption characteristics, and packaging, as long as they can provide two output modes—visible light frame images and asynchronous event streams—on the same physical chip and support independent start and stop of the APS imaging path via hardware instructions or power domain control, they can seamlessly adapt to the "first working mode" and "second working mode" switching logic defined in this application, and fully realize the technical effect of "original images not leaving the device and irreversible anonymization of contour videos."

[0146] The technical solution proposed in this application is essentially an intelligent vision terminal system that combines a dual-mode sensing architecture based on a shared pixel array with a physical-level privacy isolation mechanism. Its core innovation lies in: in the second operating mode, completely disabling the light-sensing and data output functions of the APS module through hardware power-off or equivalent physical means, thus preventing the generation of raw visible light images at the source; simultaneously, reconstructing the desensitized behavioral image using only the sparse event stream output by the EVS, and generating structured alarms in conjunction with a local AI model, ultimately outputting streaming media content compatible with the existing ecosystem through a standard video interface. This technical approach does not rely on a specific operating system, network environment, or backend platform, nor does it require the acquisition, caching, or transmission of any raw image frames, fundamentally achieving the "minimum necessity" and "irreversible anonymization" required by the Personal Information Protection Law.

[0147] Therefore, the scope of this application extends far beyond the previously exemplified scenarios such as home-based elderly care, community care, or vehicle monitoring. Any application requiring simultaneous visualization of behavior, process tracking, and compliance exemption in a highly privacy-sensitive environment can be considered a reasonable extension of this application, including but not limited to: records of home visits by grassroots public service personnel, safety monitoring of public transportation passenger cabins, supervision of teaching behavior in educational institutions, assessment of rehabilitation training movements, seamless care in hospital wards, emergency response for individuals living alone, and various mobile or fixed monitoring scenarios that require consideration of both evidentiary validity and personal dignity.

[0148] In terms of device form, this application can be implemented as a wall-mounted network camera, or packaged as a handheld recorder, head-mounted terminal, vehicle-mounted black box, or embedded module; in terms of power supply, it supports PoE, DC adapter, rechargeable battery, or supercapacitor; in terms of networking capabilities, it is compatible with Ethernet, Wi-Fi, 4G / 5G, Bluetooth, or completely offline operation; in terms of output protocols, it supports RTSP, ONVIF, GB / T 28181, HTTP-FLV, or local storage media playback. Regardless of the external form or deployment conditions, as long as it achieves the technical closed loop of "physical switch control of dual-mode switching + APS hardware power-off + EVS-driven contour rendering + original image not leaving the device," it falls within the protection scope of this application.

[0149] In summary, the scope of patent protection for this application should not be narrowly limited to the specific implementation details, sensor models, chip platforms, or visualization styles described in the specification, but should be determined by the appended claims and their equivalent technical solutions in the sense of patent law. Any functional modifications made to the sensing front-end, processing unit, rendering engine, communication interface, storage strategy, or application scenario based on the fundamental principles of this application, as long as their essential technical effect still lies in physically severing the original image acquisition link and generating continuous behavioral visualization outputs with irreversible identities based on event streams, even if not explicitly listed in the foregoing embodiments, should be considered equivalent implementations of this application and protected by law.

Claims

1. A dual-mode intelligent vision device supporting privacy protection mode switching, comprising an image sensor chip, an edge processing unit, a video rendering module, and a network communication interface, characterized in that: The image sensor chip integrates an active pixel sensor (APS) module and an event vision sensor (EVS) module to synchronously output visible light image frames and event stream data. The edge processing unit is connected to the image sensor chip. In the first working mode, it enables the active pixel sensor (APS) module to generate a visible light video stream. In the second working mode, it disables the power supply or imaging function of the active pixel sensor (APS) module and only enables the event vision sensor (EVS) module to perform target behavior analysis. The video rendering module converts the event stream data into a grayscale outline video stream in the second working mode; The network communication interface outputs the visible light video stream in the first working mode, and only outputs the grayscale contour video stream and structured alarm events in the second working mode, without outputting the original APS image or the original EVS event stream.

2. The dual-mode intelligent vision device according to claim 1, characterized in that, The active pixel sensor (APS) module and the event vision sensor (EVS) module share the same pixel array and transmit data to the edge processing unit via the same physical channel through the MIPI CSI-2 interface in the form of embedded data packets.

3. The dual-mode intelligent vision device according to claim 1, characterized in that, The edge processing unit has a built-in lightweight AI inference model, which is used to detect human fall events in real time based on the event stream data in the second working mode, and generate structured alarm events when a fall event is detected.

4. The dual-mode intelligent vision device according to claim 3, characterized in that, The structured alarm events are pushed to the network video recorder via the ONVIF event protocol.

5. The dual-mode intelligent vision device according to claim 1, characterized in that, In the second operating mode, the raw image data of the active pixel sensor (APS) module is prohibited from being output to the network communication interface.

6. The dual-mode intelligent vision device according to claim 5, characterized in that, In the second operating mode, the active pixel sensor (APS) module is powered off to block visible light imaging at the hardware level.

7. The dual-mode intelligent vision device according to claim 1, characterized in that, It also includes a physical mode switching mechanism, which is a mechanical toggle switch, slide switch, push button switch or hard-wired GPIO control signal input port, used for manual or external triggering to select the first working mode or the second working mode.

8. The dual-mode intelligent vision device according to claim 1, characterized in that, Both the visible light video stream and the grayscale contour video stream use standard video encoding formats and are provided externally through standard streaming media protocols to support plug-and-play compatibility with network video recorders.

9. The dual-mode intelligent vision device according to claim 1, characterized in that, The edge processing unit performs event stream processing, target behavior analysis, alarm generation, and local storage in the absence of an internet connection, and supports caching alarm fragments via a MicroSD card.

10. The dual-mode intelligent vision device according to claim 1, characterized in that, The event stream data includes pixel coordinates X and Y, timestamp t, and polarity pol. The video rendering module maps the event stream data into continuous grayscale contour frames based on a spatiotemporal aggregation algorithm.