Privacy protection type service behavior monitoring method and system based on event-driven vision

By extracting structured action features through event-driven visual sensors and neural network models, and combining them with multimodal visualization, the problem of balancing privacy and functionality in service behavior monitoring under high privacy scenarios is solved, and non-intrusive, interpretable intelligent recording and automated verification are achieved.

CN121834876APending Publication Date: 2026-04-10YUNNONG (SHENZHEN) TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-07
Publication Date
2026-04-10

AI Technical Summary

Technical Problem

Existing service behavior monitoring solutions in high-privacy scenarios either sacrifice privacy (high-definition video) or functionality (pure event streams without semantics), failing to achieve structured, interpretable, and auditable intelligent recording of service behavior, especially lacking multi-granularity modeling capabilities in understanding complex human service behavior.

Method used

It uses an asynchronous sensor based on event-driven vision to collect dynamic light intensity changes, extracts structured action features through a neural network model, and combines multimodal visualization rendering to generate a geometric representation without identity information. It supports collaborative interpretation by human supervision and artificial intelligence, and realizes non-intrusive recording and intelligent analysis.

Benefits of technology

It enables compliance with legal requirements without collecting identifiable image information, enhances the ability of inspectors to identify service details, reduces storage and transmission overhead, supports automated verification, improves the efficiency of service quality assessment and insurance review, and is applicable to a variety of hardware platforms.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121834876A_ABST
    Figure CN121834876A_ABST
Patent Text Reader

Abstract

The invention discloses a privacy protection type service behavior monitoring method and system based on event-driven vision. The method comprises the following steps: collecting light intensity change in a service scene through an event-driven asynchronous visual sensor, and generating event stream data without texture information; a neural network model is utilized to extract desensitized structured action features from the event stream, and the features comprise a whole body posture, hand operation and a multi-granularity key point skeleton of a face area and do not contain any identifiable identity information; and performing behavior semantic analysis such as fall detection, service action compliance verification, people flow statistics and the like based on the features, and outputting a visual result for service quality supervision, risk early warning or responsibility tracing. On the premise of not obtaining, storing or transmitting any image frame, non-intrusive and high-precision monitoring of service behaviors in high-privacy sensitive scenes such as home door-to-door service, old-age nursing and medical rehabilitation is achieved, and privacy protection and intelligent supervision requirements are effectively balanced.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of intelligent visual perception and privacy computing, and is used for non-intrusive service behavior monitoring and desensitization analysis in a high-privacy scenario. BACKGROUND

[0002] In recent years, the home service industry has rapidly expanded, including not only traditional basic care such as bathing, feeding, and accompanying diagnosis, but also high-contact and high-privacy sensitive service types such as home rehabilitation training, facial beauty, skin care, eyebrow and face repair, head and neck massage, and physiotherapy relaxation. These services are usually carried out in private spaces such as user bedrooms and bathrooms, and service personnel need to come into close contact with the user's body, even involving the face, neck and other highly sensitive areas. The market demand for such services continues to grow, but at the same time, how to protect user privacy while ensuring service quality has become a core bottleneck for industry development.

[0003] Regulators, insurance companies and platform enterprises urgently need an objective, tamper-proof and traceable technical means to verify key facts such as "whether it is a real service", "whether it is operated according to the standard", and "whether it covers the specified parts".

[0004] Currently, the industry generally uses two types of visual recording devices as a solution: one is a fixed high-definition RGB camera system installed in the living room, and the other is a wearable service recorder worn by service personnel, which generally integrates a 1080P or higher resolution high-definition camera module for recording the entire home service process. Such devices have been used by force or recommended in multiple long-term care insurance pilot cities as the main basis for service verification.

[0005] However, these traditional imaging-based solutions face fundamental contradictions in high-privacy scenarios. On the one hand, high-definition video inevitably records sensitive information such as user's exposed body parts, facial features, and home environment, even if the platform promises "authorized review only" or "automatic face blurring", the collection of raw image data itself constitutes a substantial challenge to the privacy clauses of the Personal Information Protection Law and the Civil Code; on the other hand, users (especially the elderly, women or people with mental disorders) generally have strong psychological resistance to being recorded, often refusing service, blocking the camera, and requesting the device to be turned off, resulting in interrupted recording, distorted data, and weakened regulatory effectiveness. In addition, the massive video relies on manual spot checks, which is costly and difficult to scale, and cannot achieve automated behavior compliance judgment.

[0006] Although Event-Driven Vision technology is considered as a new path of privacy-friendly perception due to its output of only sparse event stream without texture image, existing researches and products still focus on industrial detection, autonomous driving and other fields, and there is no effective solution for understanding complex human service behaviors. Specifically, the current event vision system lacks the multi-granularity modeling capability for full-body posture (such as falling), hand fine action (such as holding tools, feeding trajectory), and facial contact area (such as massage path); the visualization effect is poor, which makes it difficult to support the intuitive interpretation of supervisors; and it is not semantically associated with service work orders, insurance verification rules, and cannot answer business questions such as "whether a 10-minute facial massage is completed" or "whether the hand really contacts the lower jaw area". At the same time, the system is mostly closed hardware, which is difficult to embed in wearable recorders, nursing beds, shower robots and other diversified terminals, limiting the actual landing.

[0007] In summary, although the high-value service market such as home rehabilitation, beauty, massage is growing rapidly, and the demand for anti-fraud, traceable, and privacy-compliant monitoring technology is urgent, existing technical solutions either sacrifice privacy (high-definition video) or sacrifice functionality (pure event stream without semantics), and there is no structured, interpretable, and auditable service behavior intelligent recording system that can be implemented without collecting any image frames. This technical gap needs to be filled. SUMMARY

[0008] Based on this, the present application provides a privacy-protected service behavior monitoring method and system based on event-driven vision, which is used to realize non-intrusive recording and intelligent analysis of service behavior in high-privacy-sensitive scenes such as home care, home rehabilitation, beauty and massage, while avoiding the collection of image information that can identify the identity.

[0009] The method first collects the dynamic light intensity changes in the service scene through an event-driven asynchronous vision sensor. The sensor works in a pixel-level asynchronous manner, and only when the local light intensity changes by more than a preset threshold, it outputs events containing timestamps, pixel coordinates and change polarity. All events are arranged in chronological order to form event stream data. This process does not generate, rely on, or output any continuous image frames containing complete scene texture, eliminating the risk of identity information leakage from the sensor source.

[0010] Subsequently, the event stream data is processed in real time using a neural network model to extract structured motion features related to human behavior. The structured motion features are geometric representations that do not contain identifiable information of identity, specifically including at least one of whole body posture key points, hand joint key points and their topological connection relationships, and facial dissection key points. The neural network model reconstructs the human motion state from sparse event streams through spatio-temporal feature encoding, and dynamically enables key point sets of corresponding granularity according to the current service type or spatial context.

[0011] To support the collaborative interpretation of artificial supervision and artificial intelligence models, the system performs multi-modal visual rendering of the perception results. The visualization content includes the point rendering of the original event stream, the optical flow vector field calculated based on the event stream, and the skeleton line segments corresponding to the structured motion features, which can be selectively superimposed or mixed and displayed on the same screen. Among them, the event stream presents the dynamic area in the form of discrete points, each event point is distinguished by different colors according to its polarity, and the brightness of the event point and the background is enhanced by a dynamic contrast enhancement algorithm to improve the distinction between the two, so that weak motion areas can still be clearly identified on low dynamic range display devices; the optical flow is represented by arrows or color-coded vectors to indicate the direction and speed of motion; the skeleton line segments outline the human structure in a geometric line: the whole body posture is composed of line segments between key points of the torso, limbs and head; the hand restores the structure outline of the palm and fingers through the topological connection of the palm and the finger joints; the face forms a desensitization outline by key points such as eyes, nose tip, corner of the mouth, jaw, and cheekbone, and their connecting lines. All visualization elements do not contain original image texture, and specific parts of the skeleton or key points can be dynamically hidden according to the preset privacy policy.

[0012] On this basis, the system performs behavior semantic analysis related to service compliance. The analysis maps the structured motion features to predefined service behavior semantic units, including fall event detection, whether the hand completes the feeding trajectory, whether the facial massage covers the specified area, or the number of individuals in the scene changes, etc. The analysis results generate structured behavior logs, including behavior type, start and end time, duration, involved body parts, and compliance determination results, which are used for service quality evaluation, insurance verification, or accident tracing.

[0013] Correspondingly, the application also provides a monitoring system for implementing the above method. The system comprises an event-driven asynchronous visual sensor, an edge intelligent processing unit, a visual output unit and an SDK encapsulation unit. The edge intelligent processing unit integrates a lightweight neural network inference engine and is deployed locally on the device, responsible for event stream analysis, feature extraction, behavior analysis and multi-modal visual rendering, ensuring that the original event stream and intermediate feature data do not leave the device. The visual output unit fuses the enhanced contrast event stream, optical flow and skeleton line segment and outputs them to a local display screen or encodes them as a video stream for remote access. The SDK encapsulation unit packages the perception, analysis and visualization functions into a standardized software interface, supporting embedding into special cameras, wearable service recorders, mobile service robots, intelligent lamps, nursing beds or bath devices.

[0014] Compared with the prior art, the application has significant technical effects. Since image frame acquisition is completely abandoned and only event stream is relied on for perception, the acquisition of human faces or body textures is fundamentally avoided, meeting the legal compliance requirements of high-privacy places. The multi-modal visualization mechanism displays the fusion of event stream, optical flow and skeleton line segment, combined with dynamic contrast enhancement of event points, significantly improving the recognition ability of monitoring personnel in reviewing records for on-site dynamic details, especially in complex light or weak motion scenes, which can more accurately judge the operation position of service personnel, user posture changes and potential risk events. Structured behavior logs replace raw videos, greatly reducing storage and transmission overheads, and supporting automated verification to improve the efficiency of long-term care insurance service audits. The edge local processing architecture ensures data security and is suitable for medical institutions or rehabilitation sanatoriums that have mandatory requirements for data not leaving the domain. The SDK design enables the system to quickly adapt to various hardware platforms, reducing deployment costs and enhancing the engineering practicality and large-scale landing ability of the solution. BRIEF DESCRIPTION OF DRAWINGS

[0015] In order to more clearly illustrate the technical solutions in the embodiments of the application or the prior art, brief introductions will be given to the drawings needed to be used in the embodiments or prior art descriptions. Obviously, the drawings in the following description are only some embodiments of the application, and for those skilled in the art, other drawings can also be obtained without creative labor on the basis of these drawings.

[0016] Figure 1 The overall architecture schematic diagram of the privacy protection type service behavior monitoring system based on event-driven vision provided by the embodiments of the application;

[0017] Figure 2 The screen example diagram of event stream, optical flow and multi-granularity skeleton superposition visualization in the embodiments of the application;

[0018] Figure 3A visual effect schematic diagram of the embodiment of the present application in a fall detection scene;

[0019] Figure 4 A hand fine operation visualization schematic diagram of the embodiment of the present application. DETAILED DESCRIPTION

[0020] The technical solutions in the embodiments of the present application will be clearly and completely described with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by a person of ordinary skill in the art without creative labor fall within the protection scope of the present application.

[0021] In one specific embodiment of the present application, the privacy protection type service behavior monitoring system is deployed in a home nursing scene, and is used for non-intrusive recording and analysis of high privacy sensitive services such as bathing, feeding, facial massage, etc. The system hardware includes an event camera module integrated with a dynamic visual sensor chip, which includes an optical lens, an event-driven pixel array and an embedded preprocessing circuit, and the output format is a standard event stream, each event including a timestamp (microsecond level precision), a pixel coordinate (x, y) and a polarity (positive / negative, indicating that the light intensity increases or decreases). The event stream is transmitted in real time to an edge computing unit through a USB 3.0 interface.

[0022] The edge computing unit first preprocesses the original event stream. In order to improve the clarity of subsequent visualization, the system performs a dynamic contrast enhancement operation: the event stream is accumulated into an event frame according to a time window (for example, 50 milliseconds), and the event density in the current window is counted; if the density is lower than a threshold T1, the event point brightness gain coefficient a is adaptively enlarged, where d is the event density, k is the adjustment factor, and ε is a small constant to prevent division by zero, is the basic gain; then the enhanced event frame is mapped to the display brightness range [0, 255], the positive polarity event is represented by a white point, the negative polarity is represented by a black point, and the background is set to dark gray, thereby significantly improving the visual recognition of weak motion areas.

[0023] Subsequently, the enhanced event stream is input to a lightweight spatio-temporal convolutional neural network model. The model takes the sequence of event frames within a sliding time window (e.g., 100 milliseconds) as input and outputs three types of structured action features: the first type is full-body pose keypoints, a total of 17, corresponding to head, shoulder, elbow, wrist, hip, knee, ankle, etc.; the second type is hand keypoints, 21 for each hand, including palm, finger root, finger joint, and fingertip, and the hand palm and five-finger connection structure is generated through a pre-defined topological relationship; the third type is face keypoints, a total of 19, covering both eyes, inner and outer corners, nose tip, nose wings, mouth corners, lower lip, jaw contour, and cheekbones. The model uses knowledge distillation technology to compress the parameter size to 1.8 MB while maintaining accuracy, and can run in real time on edge devices at 30 FPS.

[0024] The system dynamically enables keypoint output of corresponding granularity according to the current service order type. For example, when the order is "bathing service", only the full-body pose and hand keypoints are activated, and the face keypoints are turned off; when the order is "facial massage", the face and hand keypoints are enabled, and the full-body pose is down-sampled. This strategy is controlled by a scene semantic recognition module deployed on the edge, which determines whether the user is in the bathroom, bedroom, etc. by the spatial distribution pattern in the event stream, and makes decisions in combination with service type metadata.

[0025] The visualization rendering module receives the event stream enhancement frame, the optical flow vector field, and the structured action features, and performs multi-layer superposition synthesis. The optical flow is calculated based on the event cumulative gradient method: take the time derivative ∂I / ∂t and spatial gradient ∇I of the event frame I(t), solve the optical flow constraint equation ∇I·v + ∂I / ∂t = 0 to get the velocity vector v; the vector is superimposed on the picture in the form of an arrow, and the length and color coding speed size. The skeleton line is drawn with a fixed width (e.g., 3 pixels), the full-body skeleton is blue, the hand structure is red, and the face contour is green, and all lines do not contain texture filling. The final synthesized picture is encoded as an H.264 video stream for local storage or remote review.

[0026] The behavior semantic analysis module continuously monitors the structured action feature stream. For example, fall detection is achieved by determining whether the trunk inclination rate exceeds a threshold and is accompanied by rapid downward movement of the lower limbs; feeding action verification requires that the right hand hold an object (the hand keypoint distance conforms to the spoon holding pattern), the trajectory moves from the dish area to the mouth area, and stays in the mouth area for ≥1.5 seconds; facial massage compliance is determined by calculating the coverage rate and duration of the hand trajectory within the area enclosed by the face keypoint. The analysis results are generated in JSON format, including timestamp, action_type, duration, body_parts, compliance_status fields, and are uploaded to the management platform after encryption.

[0027] The system supports one-key generation of interpretable backtracking clips. When an abnormal event such as a fall or long-term stillness is triggered, the enhanced event stream, optical flow and skeleton sequence of the previous 5 seconds to the next 5 seconds are automatically intercepted, a time watermarked MP4 video is synthesized, and a structured log summary is attached, forming a complete evidence package for insurance claims or responsibility identification.

[0028] The above functional modules are encapsulated as a standardized SDK, providing C++ and Python API interfaces, and supporting Linux and RTOS operating systems. The SDK has been successfully integrated into the embedded host of a nursing bed, a wearable service recorder (worn on the shoulder of a nursing staff), a bathing robot control unit and an intelligent table lamp internal computing module, verifying its cross-platform deployment capability and engineering practicability.

[0029] Figure 1 The overall architecture of the privacy protection type service behavior monitoring system based on event-driven vision provided by the embodiment of the application is shown. The system adopts a horizontal data flow layout from left to right, clearly showing a complete processing chain from raw perception to upper layer application. The entire architecture is composed of four sequentially connected functional modules, namely an event type vision sensor 101, an edge intelligent processing unit 102, a visual output unit 103 and an SDK encapsulation and application interface layer 104, and the modules are connected by solid lines with arrows, indicating one-way flow of data or control signals.

[0030] The event type vision sensor 101 is located at the leftmost side of the architecture, serving as the perception entrance of the system and being responsible for collecting dynamic light intensity changes in the service scene. The sensor works in a pixel-level asynchronous manner, and only outputs events containing timestamps, pixel coordinates and change polarity when the local light intensity changes by more than a preset threshold. All events form an event stream data in chronological order. This process does not generate, rely on or output any continuous image frame containing the complete scene texture, thereby eliminating the acquisition and leakage of identity information at the transmission source.

[0031] After the event stream data is output from the event type vision sensor 101, it enters the edge intelligent processing unit 102. The unit integrates a lightweight neural network inference engine to perform real-time analysis on the event stream locally on the device, execute dynamic contrast enhancement processing, and extract structured action features related to human behavior therefrom. These features include user face key points, service staff hand key points, human body trunk key points and the simplified skeleton line segments formed by the topological connection relationship thereof, and simultaneously calculate optical flow vectors to represent the motion direction and speed. The edge intelligent processing unit 102 also performs behavior semantic analysis based on the above features, generates a structured service behavior log, and ensures that the original event stream and intermediate feature data do not leave the device.

[0032] The processed enhanced event stream, keypoint coordinates, skeleton line segment data, optical flow vector information, and structured behavior log are sent to the visualization output unit 103. This unit fuses the multi-source data to render a desensitized geometric figure picture, in which the enhanced event stream presents the motion details of the monitored object in the form of discrete points, the skeleton line segment outlines the human structure in high-contrast lines, and the optical flow indicates the motion trend with vector arrows. The entire visualization content does not contain any original image texture, only retaining the interpretable behavior semantic information, supporting both local display and encoding as a video stream for remote access.

[0033] Finally, the data generated by the visualization output unit 103, together with the behavior analysis results, are delivered to the SDK packaging and application interface layer 104. This layer packages the perception, processing, visualization, and analysis functions into a standardized software development kit, providing a unified interface for upper-layer applications to call. Through this SDK, system capabilities can be seamlessly embedded into various terminal devices, including special security cameras, wearable service recorders, service robots, intelligent lamps, nursing beds, or bath assistance devices, thereby enabling flexible deployment and large-scale application in high-privacy sensitive places such as elderly care institutions, medical institutions, home bathrooms, or rehabilitation hospitals.

[0034] Figure 2 The picture example of the event stream, optical flow, and multi-granularity skeleton superimposed visualization in the embodiment of the present application, the whole adopts the fusion rendering mode, the event stream, the optical flow vector field and the multi-granularity human key point skeleton after dynamic contrast enhancement are cooperated in the same picture, to realize the behavior visualization effect of high readability and complete desensitization.

[0035] The event stream image 201 shows the display result after the original event stream is processed by dynamic contrast enhancement. The system renders the accumulated event points in the time window in the form of discrete pixels, positive polarity events (brightness increase) are represented by light color points, negative polarity events (brightness decrease) are represented by dark color points, and through an adaptive gain algorithm, the brightness of event sparse areas is improved, and the dynamic range of dense areas is moderately compressed, thereby avoiding the "white vastness" or "black sheet" phenomenon caused by local overexposure or underexposure in traditional visualization. This processing makes the motion subject, such as the hand of the service personnel, the face or torso of the person being cared for, stand out significantly in the gray-black background, clearly reflecting the details of weak movements.

[0036] The event stream image 202 superimposes the accurate positioning marks of the human body key points on the basis of 201. The system adopts a key point system conforming to the general human body posture estimation standard, and uses 18 trunk and upper limb key points under the basic configuration, which specifically include: nose tip (0), left eye center (1), right eye center (2), left ear (3), right ear (4), upper neck (5), lower neck (6), left shoulder (7), right shoulder (8), left elbow (9), right elbow (10), left wrist (11), right wrist (12), upper back center (13), lower back center (14), left hip (15), and right hip (16). All key points are marked with solid circles of a uniform size, and the colors are distinguished according to the sources (for example, yellow for service personnel and white for the person being cared for), so that the operation subject and the service object can be quickly identified during manual monitoring. It should be emphasized that the key point set is not fixed, but is dynamically enabled according to different granularities in different service scenarios: for example, in the facial massage scenario, only the head and hand key points are enabled; in the bathing or fall detection scenario, the complete upper body key points including the hip are enabled; and in the service involving whole body movement, the leg key points (such as the 25-point model of the full version of OpenPose) can also be extended, but the principle of “minimum necessity” is always followed to avoid collecting irrelevant body part information.

[0037] The event stream image 203 further draws bone segments on the basis of 202, and the connection mode is strictly based on the anatomical topology, but the connection rules can be dynamically adjusted according to the scene strategy. In the default home care mode, the connection rules are as follows: the nose tip (0) is connected to the upper neck (5), and then to the lower neck (6); the lower neck (6) is connected to the left shoulder (7) and the right shoulder (8) respectively; each shoulder is connected to the elbow (9 / 10) and wrist (11 / 12) on the same side in turn; the upper back center (13) and the lower back center (14) are connected vertically to form the main trunk of the spine; and the lower back center (14) is connected to the left hip (15) and the right hip (16) respectively. All the line segments are in high-contrast white with a line width of 2 pixels, and are clear and distinguishable in the enhanced event stream background. However, this connection logic is not unique: in the facial service scenario, the system can automatically suppress the trunk line segments other than shoulder-elbow-wrist, and only keep the hand-head spatial relationship; in the violent behavior early warning mode, the virtual connection line between the hand and the trunk can be strengthened to judge the intrusion distance; and in the fall detection mode, the hip-knee-ankle connection line can be enabled to evaluate the center of gravity deviation. This configurable skeleton topology mechanism enables the same set of underlying key point data to adapt to multiple monitoring needs, improving the flexibility and practicality of the system.

[0038] The event stream image 204 adds facial key positioning points and their connections on the basis of 203. The face defines 19 anatomical key points, including bilateral eyebrow peaks, eyebrow tails, inner / outer corners of the eyes, nasal roots, nasal tips, bilateral nasal alar points, upper philtrum points, upper and lower lip midpoint points, bilateral corner points of the mouth, bilateral lower jaw edge end points, bilateral pre-ear screen points, and forehead center points. These points are presented in light gray small circles without filling and without outline, only retaining coordinate information. The face connection is also not fixed, but is dynamically enabled according to the service type: for example, in facial massage, the system automatically connects the outer corners of the eyes-sunken areas, the nasal alar-cheekbone areas, and the mouth corners-mandibular areas to form preset service partition boundaries; in abnormal contact detection, the spatial proximity relationship between the hand key points and the mouth and eye key points may be highlighted, but no entity connection is drawn to avoid implying contact. All facial elements do not restore the shape of the real five organs and do not constitute an identifiable face image.

[0039] The whole Figure 2 Through the above four-layer progressive visualization (201→202→203→204), the precise expression of hand-face interaction, posture compliance, and motion risk in the service process is realized without any original image texture. The enhanced event stream ensures that the motion subject is prominently visible, the key point positioning provides a spatial coordinate reference, the skeleton line segment reveals the structural relationship, and the optical flow vector (indicated by arrows in the figure, and the black and white line drawing in the patent drawing can be distinguished by solid / dashed lines) supplements the motion direction and intensity information. More importantly, the key point selection and skeleton connection support dynamic configuration according to the scene strategy, which not only meets the regulatory accuracy requirements of different service types, but also strictly follows the principle of privacy minimization. This visualization scheme is designed for remote safety monitors, enabling them to efficiently and accurately complete service behavior interpretation, quality evaluation, and abnormal backtracking while fully protecting user privacy.

[0040] Figure 3 The visualization effect schematic diagram of the embodiment of the present application in the fall detection scene shows the continuous frame sequence of the event stream highlighted motion area, the optical flow vector arrow indicating the motion direction and speed, and the stickman skeleton changing with the posture during the process of the human body from standing, walking to falling. The diagram is constructed based on event-driven vision technology, and its core is to use the asynchronous event stream output by the event camera, combined with optical flow calculation and skeleton modeling, to generate a highly desensitized and behavior semantic-rich visualization picture specially used for remote safety monitors to quickly interpret high-risk events.

[0041] An event camera is a kind of bionic vision sensor, whose working principle is completely different from that of a traditional frame camera. Instead of outputting a complete image at a fixed frame rate, it independently perceives the relative change of local brightness for each pixel: when the light intensity of a certain pixel point changes beyond a preset threshold, an "event" containing a timestamp, a two-dimensional coordinate and a polarity (positive / negative, indicating lightening or darkening) is output. This mechanism enables the event camera to have microsecond-level time resolution, high dynamic range (>120 dB) and extremely low data redundancy, and it can still stably capture dynamic details in environments with sudden light changes, fast motion or weak light, and is particularly suitable for sensing key actions such as slow getting up, sudden imbalance or falling of the elderly in the home environment.

[0042] In the system, the original event stream is first subjected to dynamic contrast enhancement processing, and the events accumulated in a short time window are rendered in the form of discrete points, so that the moving body (such as limbs and trunk) is prominently highlighted in the gray-black background, avoiding information loss caused by overexposure or underexposure in traditional visualization. On this basis, the system further performs event stream-based optical flow estimation. Since the event stream itself is sparse and asynchronous, the invention uses lightweight algorithms such as spatiotemporal gradient method or spiking neural network (SNN) to reconstruct the local motion vector field from the time-space distribution of events. Specifically, the system analyzes the correlation of adjacent events in time and position, deduces the instantaneous velocity vector of each significant motion region, and visualizes it as an arrow with direction and length - the arrow points to the direction of motion, and the length of the arrow is proportional to the speed of motion, and the color or line type can encode the risk level (for example, a red solid line represents high-speed falling, and a blue dashed line represents slow movement).

[0043] In the visualization effect 301, the human body is in a standing and normal walking state. The event stream clearly outlines the dynamic regions of arm swing and leg stride, and the optical flow arrows extend outward from the trunk and limbs, consistent with the gait direction and moderate in length, indicating that the motion is stable and controllable. The stick figure skeleton is composed of 18 key points, connected according to anatomical topology, accurately reflecting the upright posture and coordinated gait.

[0044] In the visualization effect 302, the system switches to a side view (which can be achieved by multi-device deployment or single-device wide-angle coverage), further revealing the change of the center of gravity. At this time, the optical flow arrows not only show the horizontal forward direction, but also present a slight forward tilting trend in the trunk area, with the arrows slightly tilted downward and forward, consistent with the biomechanical characteristics of normal walking. This view is crucial for determining whether to lose balance soon.

[0045] In the visualization effect 303, the human body has fallen down. The event stream bursts densely in the torso falling path and the arm spreading area in a short time, forming a high-density event cluster. The optical flow arrow presents a typical falling mode: the arrow in the torso area sharply points downward and significantly increases in length, indicating a rapid vertical acceleration; the arrow in the arm area extends radially outward and upward, reflecting an instinctive support reaction. At the same time, the matchstick skeleton quickly collapses from the upright configuration to the horizontal or lateral lying posture, the distance between the hip and shoulder key points decreases sharply, and the spine connection line changes from vertical to nearly horizontal. The three signals of event stream + optical flow arrow + skeleton deformation are coordinated, making the falling event have a very high visual recognition degree.

[0046] It should be emphasized that, Figure 3 The optical flow arrow in the above figure is not from the traditional optical flow calculation between video frames, but is a motion vector directly derived from the event stream in real time, which has lower delay, higher time accuracy and stronger anti-fuzzing capability. In the case of a transient event such as falling, a traditional camera may miss the key moment due to insufficient frame rate, while an event camera combined with optical flow visualization can generate clear direction and speed prompts within hundreds of milliseconds of the fall.

[0047] In practical applications, a safety monitor can watch such a sequence of continuous frames without relying on real images, and can efficiently judge whether the current situation is normal walking, unstable standing up, or falling down, only by the arrow direction, skeleton deformation, and event density distribution. Experiments show that after introducing the optical flow arrow, the recognition accuracy of the monitor for the falling event is improved to more than 94%, and the average response time is shortened to less than 6 seconds. This visualization method not only meets the strict restrictions of the Personal Information Protection Law on biometric information, but also provides reliable technical support for active safety intervention in high-privacy-sensitive scenarios.

[0048] Figure 4 The visualization schematic diagram of the fine operation of the hand in the embodiment of the present application is intended to clearly present the topological structure of the palm and the five fingers and the dynamic interaction relationship between the topological structure and the operation object when a service personnel performs high-precision service actions such as facial massage, feeding assistance, or tool holding. The diagram realizes high-fidelity and desensitization expression of the hand posture and operation intention by two-stage progressive rendering, which deeply integrates event-driven perception and geometric modeling, and realizes high-fidelity and desensitization expression of the hand posture and operation intention without relying on image texture.

[0049] The hand visualization effect diagram 401 shows the basic effect of the event type visual sensor output raw event stream after dynamic contrast enhancement. In this picture, all pixel-level light intensity changes are presented in the form of discrete points: positive polarity events (brightness increase) are represented by light gray points, and negative polarity events (brightness decrease) are represented by dark gray points. The system uses an adaptive time window accumulation and local gain adjustment algorithm to significantly improve the visibility of weak motion areas (such as finger tapping, tool sliding), and avoid visual blind spots caused by sparse events. At this time, although there is no structural information in the picture, the dynamic trajectory of the hand outline and tool edge can still be clearly identified, especially in the moment when the fingers contact the face, utensils or nursing equipment, the event density increases significantly, forming a highlighted area, which intuitively reflects the location and timing of the operation.

[0050] The hand visualization effect diagram 402 superimposes the topological connection model of the hand structure on the basis of 401. The model is based on 21 key points: each hand contains 5 fingertip points (from the thumb to the little finger end), 5 distal interphalangeal joints (PIP), 5 proximal interphalangeal joints (MCP), 5 metacarpophalangeal connection points (located at the head of metacarpal bone), and 1 palm center point (defined as the geometric center of intersection of five metacarpal axis). The system draws the skeleton line segment according to the strict anatomical connection rule: each finger is connected in the order of "fingertip, distal joint, proximal joint"; the five proximal joints are connected to the corresponding metacarpophalangeal connection points respectively; all metacarpophalangeal connection points converge to the palm center point, forming a radial palm structure; the thumb is independent due to its anatomical structure, and its proximal joint is directly connected to the palm center offset to the radial side, to accurately reflect the palm function.

[0051] All the lines are bright yellow solid lines with a uniform line width of 2 pixels, which are highly visible in the gray-black enhanced event stream background. When the service personnel hold the massage rod to operate the zygomatic region, the event stream highlights the sliding trajectory of the contact surface between the tool and the skin, and the hand skeleton simultaneously presents the coordinated flexion and extension state of the thumb and the four fingers - for example, if only the index finger and the middle finger are involved in pressing, the key points of the remaining three fingers still exist but the lines are thin or semi-transparent, reflecting the granularity control strategy of "on-demand activation". This fusion visualization of event stream + fine topological skeleton enables the remote safety officer to clearly determine whether the operation is completed by the specified finger, whether there is unauthorized part contact, and whether the tool usage conforms to the standard path.

[0052] Overall, Figure 4 The hand visualization mechanism shown not only meets the verification needs of long-term care insurance for service authenticity, but also provides quantifiable, traceable and auditable technical basis for operation quality evaluation in rehabilitation training, old and disabled assistance and other scenarios, while strictly following the principle of privacy minimization to ensure zero leakage of user identity information.

[0053] In an optional embodiment, a wearable monitoring device named "privacy service recorder" is provided for process recording and compliance verification in mobile service scenarios such as long-term care insurance home service, home beauty, rehabilitation treatment, etc. The overall size of the device is 83mm x 56mm x 26mm, and the weight is not more than 115 grams, which can be worn on the shoulder of the service personnel or integrated on the mobile carrier such as a bath car, a companion robot, etc. to achieve unobtrusive recording during the whole process of the service personnel's work.

[0054] The device contains an event-driven asynchronous visual sensor module, which uses a dynamic visual sensor chip with an optical field angle of 78°, a lens aperture of f / 1.79, and a depth of field range of 43.4 cm to 59 cm, suitable for close-range human-computer interaction scenarios. The sensor works in a pixel-level asynchronous manner, outputting events only when the local light intensity changes exceed a threshold, each event containing a microsecond-level timestamp, pixel coordinates, and polarity information, forming an event stream data. The effective resolution of the module in EVS mode is about 0.5 megapixels, the equivalent event output rate can reach up to 4000 fps, the minimum working illuminance is lower than 1 lux, and the typical power consumption is 80 mW. The sensor does not generate, cache, or output any continuous image frame containing complete scene texture, avoiding the collection of identifiable identity information such as face, body contour, or home environment from the source.

[0055] The device has a built-in AI processing unit using Rockchip RV1106 chip with a computing power of 0.7 TOPS running a lightweight neural network model. The model receives event stream data as input and extracts structured action features in real time. The structured action features are in geometric representation form, including whole body posture key points (17), hand joint points (21 for each hand, including predefined topological connection relationship), and facial dissection key points (19), which form a multi-granularity human skeleton. The system dynamically enables the corresponding granularity according to the current service type: in the bath service, enable the whole body and hand skeleton, and close the facial key points; in the facial massage service, enable the hand and facial skeleton, and downsample the output of the whole body posture.

[0056] In the visualization rendering stage, the system performs multi-channel fusion display, specifically including the following three levels:

[0057] The first layer is event stream visualization. The original event stream is accumulated into an event frame according to a time window (e.g. 50 milliseconds), and then dynamic contrast enhancement processing is performed: calculate the event density d per unit area in the current window; if d is lower than a preset threshold T, apply a gain coefficient to the event point brightness, where For base gain, k is the adjustment factor, and ε is a small constant to prevent division by zero; the enhanced event points are mapped to different gray levels or color phases according to the polarity - positive polarity events are mapped to high-light colors (such as white or light yellow), negative polarity events are mapped to low-light colors (such as dark gray or blue), and the background is set to a neutral dark color, thereby significantly improving the visual recognition of weak motion areas.

[0058] The second layer is the visualization of the optical flow vector field. Based on the time derivative and spatial gradient of the event frame, the system solves the optical flow constraint equation ∇I·v + ∂I / ∂t = 0 to obtain the pixel-level motion velocity vector v. This vector is rendered in the form of an arrow, with the arrow length and color coding the motion speed, and the direction indicating the motion trend, which is used to assist in judging dynamic behaviors such as falls and rapid movements.

[0059] The third layer is the geometric skeleton rendering of structured action features. The system generates a skeleton graph connected by line segments based on the enabled key point set. The line segment color is not fixed, but is dynamically assigned as part of the visualization strategy according to the following rules:

[0060] According to the body part: the whole body trunk uses the first color channel, the hand structure uses the second color channel, and the face contour uses the third color channel;

[0061] According to the service semantics: the skeleton line segments in the compliance operation area use the green color system, and the unauthorized area uses the red color system;

[0062] According to the user privacy level: in high privacy mode, sensitive part skeletons are displayed with low saturation or semi-transparent color;

[0063] The color mapping table can be dynamically updated by a configuration file or a remote strategy, supporting unified visual specifications of the monitoring platform.

[0064] The line segment width is uniformly set to 3 pixels, and the endpoints are processed with rounded corners to improve readability. All lines are without filling and texture, only retaining the geometric topological relationship.

[0065] The above three layers of content are alpha-blended and superimposed in the GPU or dedicated graphics engine, and a single picture is synthesized and output to a local 2.0-inch LCD display or encoded as a video stream. The entire visualization process does not introduce any original image pixels, and all elements are synthetic graphics generated based on event streams and structured features.

[0066] The device records the content after edge encryption and stores it locally. The original event stream and intermediate feature data are retained throughout the device and do not leave the device. After the service is completed, the structured behavior log and enhanced visual video are automatically uploaded to the cloud server through the USB Type-C interface connection to the special data acquisition station. The uploaded content includes time-stamped behavior semantic records (such as "hand in face area continuous operation for 120 seconds, covering the lower jaw and cheek") and corresponding period of time explainable backtracking segments, which are used for insurance verification or dispute tracing.

[0067] The device has a horizontal field of view angle of ≥140°, a built-in 2600mAh lithium battery, supports continuous operation for 12 hours, and has a protection level of IP6X and a working temperature range of -20°C to +55°C. Its core function has been packaged as an SDK and can be transplanted to police law enforcement recorders, property inspection terminals, and other wearable devices.

[0068] In long-term care insurance home service, the device is deployed on the chest of the caregiver, and the operation of bathing, feeding, etc. is recorded throughout the process. Since no image frames are collected, user privacy is protected; and the contrast-enhanced event stream, optical flow, and configurable color skeleton of the visual video enable the reviewer to clearly interpret the service personnel's action trajectory, contact area, and duration, providing objective evidence of service authenticity and supporting automated compliance verification.

[0069] In an optional embodiment, the confidential service recorder is deployed in various high-privacy-sensitive mobile service scenarios, including long-term care insurance home care, home beauty care, and rehabilitation therapy.

[0070] In the long-term care insurance home care scenario, service personnel (such as caregivers or nurses) wear the device to enter the home of the insured and perform services such as bathing, feeding, turning over, and oral cleaning. The device is automatically activated before the service begins and continuously collects changes in ambient light intensity through event-driven sensors. For example, during the bathing process, the user is in the bathroom with part of their body exposed, and traditional cameras are likely to cause resistance due to privacy concerns. However, this device only outputs event streams and does not generate image frames, so users do not feel like they are being "recorded". The AI model extracts hand and torso key points in real time and constructs a multi-granularity skeleton: when the hand is detected to move continuously in the user's back area for more than 3 minutes, the system determines that "back scrubbing is completed"; if the hand does not enter the lower limb area, it is marked as "lower limb cleaning not performed". The visualized picture superimposes enhanced event streams (highlighting water splashes and hand movements), optical flow (indicating scrubbing direction), and hand-torso skeleton (distinguishing operation areas with configurable colors) for post-service review. All raw data are retained locally on the device, and only structured logs (such as "back scrubbing: yes; lower limb scrubbing: no; total duration: 8 minutes and 23 seconds") are encrypted and uploaded for long-term care insurance work order verification, effectively preventing "clock-in and go" fraud.

[0071] In the home beauty care scenario, a private beautician provides facial massage, eyebrow shaping, skin care, etc. Such services rely heavily on facial contact, but users are extremely sensitive to their faces being photographed. The device automatically switches to the "facial service mode" in this scenario: turn off the full-body pose output, and only enable the 19 anatomical key points of the face and the 42 joint nodes of the hands. When the beautician's fingers make a circular motion in the zygomatic to the mandibular region, the system determines whether it covers the preset massage path through the spatial relationship between the hand trajectory and the facial key points. In the visualization picture, the facial key points are presented as discrete points in low-saturation green, and the hand skeleton is displayed as a high-contrast red line, superimposed with a local enhanced event stream (reflecting the micro-light changes caused by finger pressing). If the service duration is insufficient or the path deviates, the system records it as "incomplete standard process". This record can serve as an objective basis for service quality disputes, and also be used for internal training evaluation of beauty institutions.

[0072] In the home rehabilitation therapy scenario, a rehabilitation therapist provides upper limb function training guidance for stroke patients. The device records the therapist's demonstration and the patient's imitation process. The system simultaneously tracks the hand skeletons of both people and evaluates the training effectiveness by comparing the similarity of motion trajectories. For example, when the therapist guides the patient to complete the "grasp-lift-place" action chain, the event stream captures the rapid hand movement, the optical flow indicates the motion direction, and the skeleton line segment restores the finger opening and closing state. If the patient's hand does not reach the target position or the action is interrupted, the system marks "training not up to standard". The visualization picture supports the display of two skeletons on the same screen, with color channels distinguished by identity (such as warm colors for therapists and cold colors for patients), facilitating remote rehabilitation supervisors to review and analyze. Since there is no image collection throughout the process, patients do not need to worry about their home environment or physical condition being leaked, improving service acceptance.

[0073] In all the above scenarios, the device is deployed on the service personnel's moving operation path (such as the shoulder, chest, or tool handle), adapting to complex conditions such as indoor weak light, rapid motion, and occlusion. When entering high-privacy areas such as bathrooms and bedrooms, the system identifies the scene semantics through the spatial distribution characteristics of the event stream and automatically triggers privacy protection strategies: for example, hiding facial key points, reducing skeleton update frequency, or only retaining visualization of the hand operation area. This mechanism meets the requirements of the "Personal Information Protection Law" on "minimum necessary" and "context adaptation".

[0074] After the service is completed, the device connects to a special data acquisition station through the USB Type-C interface, automatically uploads the time-stamped structured behavior log and the enhanced visualization video clips of the corresponding period. These contents constitute a complete and tamper-proof service evidence chain, which can be used for insurance claim auditing, service institution quality control, user complaint handling, or litigation evidence. The device itself does not store any image data other than the original event stream, fundamentally eliminating the risk of privacy leakage.

[0075] In an optional embodiment, a fixed "privacy camera" is provided, which is dedicated to static or semi-static places with high privacy sensitivity and long-term behavior monitoring. The device adopts wall-mounted installation or table-top placement, and typical deployment locations include: 1.5 meters above the ground on the wall of a nursing home room, above the bedside cabinet in a hospital room, on the top of the TV cabinet in a single elderly person's living room, on the side wall of the washbasin in the bathroom, and on the corner bracket of a nursing room in a nursing home. When installed, ensure that the field of view covers the main activity area (such as the bed, the bathing area, and the rehabilitation training mat), while avoiding direct facing of extremely private points such as dressing and toilet, in line with the principle of minimal privacy intervention.

[0076] The privacy camera module is 31 mm x 31 mm x 6.92 mm in size, uses precision packaging, and is connected to a local edge host (such as an embedded NVR or a health care terminal) through a USB2.0 interface, powered by the host and receiving data. The device is built-in Rockchip RV1106 AI chip, with computing power of 0.7 TOPS (INT8), supporting dual-vision fusion architecture: APS mode can output high-definition image frames, which is only used for device debugging or temporary viewing with explicit authorization of the user; in the normal service monitoring state, the system is forced to enable EVS mode, i.e. event-driven asynchronous vision sensing based on events. In this mode, the sensor outputs event stream in pixel-level asynchronous mode, with an effective resolution of about 0.5 million pixels, an equivalent event rate of up to 4000 fps, a minimum working illuminance of less than 1 lux, and a typical power consumption of 80 mW. The key is that the EVS mode does not generate, cache, or output any continuous image frame containing texture details throughout the process, avoiding the collection of identity information from the source.

[0077] The device is mainly applied to three typical service scenarios, solving the problem of full-scene monitoring that cannot be covered by single wearable recorders.

[0078] The first type is multi-person collaborative bathing service. In nursing homes or home environments, two caregivers often need to assist disabled elderly people to complete bed bathing or showering. If only relying on the service recorder worn by a single person, the visual angle is limited, and the operation of the other caregiver or the overall posture change of the user is easily missed, resulting in the authenticity of the service cannot be completely verified. While the fixed camera is deployed obliquely above the bathroom, it can cover the entire bathing area. The system extracts the full-body posture skeleton of the two service personnel and the user in real time through event stream, and tracks the hand key points respectively. When detecting that both hands perform back wiping at the same time, or one person holds the torso and the other cleans the lower limbs, the behavior semantic analysis module determines that it is "compliant collaborative operation"; if only one person moves and the other is still, it is marked as abnormal. The visualized picture superimposes the event stream with dynamic contrast enhancement (highlighting water splashing and limb movement), optical flow vector (indicating wiping direction) and multi-person body skeleton (distinguishing identity by different color channels), for post-checking. Since there is no image frame output, even if the user's body is exposed, there is no risk of privacy leakage.

[0079] The second type is housekeeping cleaning and environmental arrangement service. In the homes of elderly people living alone, housekeepers regularly clean up, arrange the bed, and change the bedding. Service agencies need to verify whether the service is truly performed and whether it covers all areas. The traditional scheme relies on service personnel to take self-shots for clocking in and out, which is easy to fake. The device is fixed on the wall of the living room or bedroom, and captures dynamic features such as fast-moving mops, cloths, and bedding shaking through event stream. The AI model identifies hand trajectories and object interaction relationships to determine whether "ground cleaning is completed" and "mattress is flipped". For example, when high-frequency event streams appear in the ground area and are accompanied by hand-tool skeleton movement, the system records it as "effective cleaning"; if only a short stay is marked as "suspected fake service". The visualized content includes enhanced event stream (reflecting dust disturbance), optical flow (indicating cleaning path), and simplified skeleton (only showing hands and torso), without involving faces or home details, meeting the stringent privacy requirements of home users.

[0080] The third type is fall monitoring and rehabilitation training recording. In hospital wards or rehabilitation centers, patients have a risk of falling when standing for balance or walking for training. The fixed camera is deployed in front of the training area, continuously monitoring the user's posture. When the event stream shows that the torso moves down quickly, the limbs spread out, and the optical flow spreads radially, the system triggers a fall warning and automatically intercepts the enhanced visualization segment from 5 seconds before to 5 seconds after. The segment includes contrast-enhanced event stream (clearly showing the falling trajectory), optical flow (indicating the impact direction), and stickman skeleton (restoring the falling posture), forming an explainable evidence for medical accident analysis or insurance claims. In daily rehabilitation training, the system can also record the number and duration of actions such as "sit-to-stand transition" and "gait cycle" completed by the patient, generating a structured log to assist therapists in assessing progress. There is no image collection throughout the process, and patients do not need to worry about their state being transmitted outside the ward.

[0081] In all the above scenarios, the original event stream and intermediate feature data are stored on the local edge host, without uploading to the device or the cloud. Only the de-identified structured behavior log (such as "Fall event: Yes; Time: 14:23:05; Number of participants: 2") can be accessed by designated roles through an encrypted channel after authorization by the user or guardian. The device supports the data policy of "full encryption recording, local storage, no default upload, and on-demand authorized viewing", which meets the privacy compliance audit requirements of government agencies for health care facilities.

[0082] The privacy camera makes up for the perspective limitations of wearable recorders through fixed deployment, solves practical pain points such as multi-person service, full-scene coverage, and long-time unattended monitoring, and provides a trusted, complete, and traceable service process evidence chain for service agencies, medical insurance departments, and family users without sacrificing privacy.

[0083] In an optional embodiment, a privacy-protected behavior monitoring system integrated into an automatic bathing robot is provided. This system is suitable for large-scale deployment in automated bathing scenarios in future nursing homes, community day care centers for the elderly, and home-based elderly care service stations. As the number of disabled elderly people increases and the nursing workforce continues to be in short supply, more and more institutions are using fixed intelligent bathing cabins or mobile bathing robots to achieve an efficient service mode of "one caregiver supervising multiple devices remotely". This mode significantly reduces labor costs and improves service coverage, but at the same time, it also brings new safety risks.

[0084] When using an automatic bathing robot, users are usually in a state of complete or partial nudity. If an emergency situation such as head submersion in water, sudden illness leading to violent struggle, or unstable sitting posture leading to slipping occurs, and there is a lack of on-site personnel to intervene in time, it is easy to cause serious safety accidents. Traditional video monitoring can record images, but it involves sensitive information such as faces and body outlines, which does not meet privacy protection regulations and is likely to cause user resistance, making it difficult to promote in practice. In addition, there are problems such as water surface reflection, steam atomization, and low illumination in the bathing environment, and traditional cameras often cannot accurately identify key behaviors due to overexposure or blurring.

[0085] To solve the above problems, this embodiment embeds a privacy visual module at a key perspective position inside the bathing robot. The module has a size of 31 mm × 31 mm × 6.92 mm and is connected to the main control unit through a USB 2.0 interface. The main control unit is equipped with a Rockchip RV1106 AI chip with a computing power of 0.7 TOPS (INT8). The system forcibly enables EVS mode during operation, i.e., event-driven asynchronous visual sensing, which does not generate, cache, or output any continuous image frames containing texture or identity information, thereby eliminating the risk of privacy leakage from the source.

[0086] Event cameras have microsecond-level response and high dynamic range characteristics, which can effectively deal with water surface reflection interference. When the light shines on the water surface to form a strong reflection, the traditional camera loses details due to overexposure in the static high-light area, while the event sensor only triggers events at the moment of light intensity change, and the static reflection area does not generate redundant signals, thereby retaining effective events caused by real human motion, ensuring the reliability of perception data.

[0087] The system performs dynamic contrast enhancement processing on the original event stream. Events are accumulated into event frames in 50 millisecond windows, and the brightness gain is adaptively adjusted according to the local event density: in event sparse areas (such as calm water period or weak breathing motion), the gain is increased to make subtle dynamics clearly visible; in event dense areas (such as violent struggle), the dynamic range is moderately compressed to avoid saturation. After this processing, the head micro-motion, arm trajectory and other key motion areas are significantly strengthened, and the overall picture signal-to-noise ratio is greatly improved.

[0088] On this basis, the system runs a lightweight neural network to extract the user's structured action features in real time, including head, torso and upper limb key points, and constructs a simplified human skeleton. These skeleton segments are rendered in a high brightness and high contrast manner (such as white or bright yellow), and are superimposed on the enhanced event stream picture. In particular, for drowning risk, the system accurately locates and highlights the head key points, even in the case of sparse events near the water surface, it can maintain stable tracking through the spatio-temporal context. The facial key points are only used to determine the orientation, and do not restore any texture, only represented by geometric points, to ensure complete desensitization.

[0089] The final generated visualization picture contains three layers of information: the bottom layer is the dynamic contrast enhanced event stream, reflecting the real motion details; the middle layer is the optical flow vector, indicating the water flow disturbance and limb movement direction; the upper layer is the highlighted skeleton segment, expressing the human structure and posture. After the fusion of the three, the remote safety officer can clearly identify the key behavior states such as "whether the head is submerged in the water surface", "whether the arms are violently waving", "whether the torso is tilted", etc., to realize efficient manual review and emergency response.

[0090] The system focuses on monitoring three types of high-risk behaviors: one is the head submerged in the water surface, if the head key point is continuously below the preset water surface threshold for more than 8 seconds and there is no breathing micro-motion event, it is determined as "suspected drowning precursor"; two is the limbs violent struggle, which is characterized by high-frequency, non-periodic rapid swinging of the limbs, the optical flow is in disordered radial, and the hand skeleton shows the action of clenching or scratching; three is the sitting posture instability, which means the torso quickly tilts or falls, accompanied by a sudden increase in event stream density in the lower limb area.

[0091] When the system determines abnormal behavior, a three-level response mechanism is triggered immediately: first, the spraying, drainage and mechanical arm action are suspended to prevent secondary injury; second, a local buzzer is started to issue sound and light warnings, and a structured alarm log is pushed to the remote nursing station, including timestamp, abnormal type, duration and confidence; finally, the enhanced visualization fragments and original event stream data of 5 seconds before and after the abnormality are automatically intercepted and stored in the robot's built-in solid state disk, forming an auditable and traceable evidence package.

[0092] In addition to abnormal monitoring, the system also records compliance indicators for routine service processes. For example, when the mechanical arm performs back washing, the event stream highlights the water impact area, and the system verifies whether the service covers the target area from scapula to lumbar vertebrae (T3-L2) by simulating the spatial relationship between the hand skeleton (representing the end of the mechanical arm) and the user's torso skeleton. If the service duration is insufficient or the path deviates, the log will mark "back washing not up to standard" for service quality evaluation or long-term care insurance work order verification.

[0093] To verify the feasibility of the system under real service logic, a series of experiments were conducted in a controlled laboratory environment. The experiment simulated a single-person bathing room for the elderly, deploying a fixed shower robot with the aforementioned privacy visual module embedded inside. A total of 12 healthy adult volunteers (ages 25-60, half male and half female) were recruited to perform normal bathing, head submersion, and limb struggle actions according to the script, with each type repeated 10 times, forming a total of 360 samples. All experiments used the EVS mode throughout, without outputting image frames.

[0094] The enhanced visualization image is transmitted to the background monitoring terminal through an encrypted local area network, and two rotating security officers watch it simultaneously in separate operating rooms. The security officer's responsibility is only to record the consistency of the system alarm time and the volunteer's action start time, and cannot take screenshots, record videos or export any data. The original event stream and intermediate feature data are retained in the robot's local storage unit throughout, without connecting to the Internet.

[0095] The experimental results show that the average event output rate is 1.2 × 10 4 events / second under normal bathing conditions, and the peak value reaches 8.7 × 10 4 events / second during violent struggle; the average delay of abnormal detection is 320 milliseconds (standard deviation ± 45 ms); the head submersion detection rate is 92.5% (37 / 40) with a false positive rate of 5.0%; the violent struggle detection rate is 96.7% (58 / 60) with a false positive rate of 1.7%; and there is no false trigger in all 120 normal action tests. In 100 randomly selected enhanced visualization videos, 94% were judged by security officers to be "clearly identifiable head position, arm movement direction and torso posture".

[0096] In addition, in the simulation of back washing service, the system has an accuracy of 98.3% in judging the service coverage area, and the average position error is ±2.1 cm (based on the event stream reconstruction coordinate system). The system AI reasoning occupies an average of 38% of the CPU, 180 MB of memory, and the whole machine power consumption is stable at 82 mW ± 3 mW, meeting the long-term continuous operation requirements.

[0097] The entire experimental process is completed in a closed laboratory, and all data processing, storage and monitoring are strictly limited to the local intranet. The raw perception data does not leave the device, and the visual image is only watched by authorized security personnel under controlled conditions, ensuring that the entire process complies with privacy protection and data security specifications.

[0098] The system proposed in this embodiment can provide clear and interpretable behavior visualization pictures for remote security personnel without collecting any images, effectively supporting the safe operation of "few people" bathing service. This scheme solves the practical problem of balancing privacy protection and safety monitoring in the automatic bathing scene, and provides a feasible technical path for the large-scale deployment of intelligent bathing robots in future in nursing homes, community service centers and home environments.

[0099] In an optional embodiment, a privacy-protected behavior observation system for daily life and personal care product research and development testing is provided. The system takes a secret camera as the core device and is deployed in a controlled experimental environment to observe the daily behavior details of volunteers using skin care products, washing and caring products, and female hygiene products without infringing on their privacy. Traditional research and development testing usually relies on high-definition video recording, requires volunteers to sign strict confidentiality agreements, and researchers analyze the indicators such as application area, frequency, method and duration by frame-by-frame review. However, such methods have obvious limitations: on the one hand, high-definition pictures contain sensitive information such as faces, body features and home environment, even if the agreement is signed, there is still a risk of data leakage; on the other hand, volunteers feel psychological pressure due to being "recorded all the time", and their behavior is distorted, affecting the authenticity of the test results.

[0100] To solve the above problems, the embodiment adopts a secret observation device based on event-driven vision. The device is small in size and can be hung on the bathroom wall, above the dressing table or beside the bathroom mirror, or placed on the table stand to ensure coverage of the main operation area. The device is forced to use EVS mode (event-driven vision sensing) during operation, does not output any continuous image frames, and only generates pixel-level asynchronous event streams. All raw data is stored locally and does not leave the device or upload to the cloud, fundamentally avoiding the risk of privacy leakage. At the same time, volunteers feel less "being photographed" and their behavior is more natural, making the data more representative.

[0101] In the skin care observation scene, volunteers sit in front of the dressing table to perform daily skin care procedures such as taking cream, applying essence, and patting lotion. The system captures the interaction dynamics of the hands and face through event streams and extracts 42 hand joints and 19 facial key points to construct a simplified skeleton. After dynamic contrast enhancement, the visualization screen only displays highlighted hand trajectories and discrete facial key points (without texture, skin color, or facial outline), superimposed with local event streams to reflect micro-motions such as pressing and rubbing. Researchers can clearly observe whether the volunteers evenly cover areas such as the forehead, cheeks, and chin; the duration of a single application; the number of repeated applications; and whether there are missed areas (such as the lower jawline or around the eyes). This information, which previously relied on high-definition video and manual annotation, can now be obtained in a completely desensitized picture, significantly improving observation efficiency and ethical compliance.

[0102] In the bathing product usage scenario, volunteers use shampoo, conditioner, and shower gel in a simulated bathroom. A secure camera is installed above the shower area to cover the head to the shoulder area. The system records the hand movement trajectories and dwell time on the scalp, hair, back, and limbs. For example, when the volunteer applies conditioner to the ends of the hair and stays for 2 minutes, the event stream forms a continuous low-frequency signal in the hair end area, and the hand skeleton shows slow combing actions. If it is only quickly brushed, it is determined as "insufficient action". Researchers can determine from the visualization screen whether the shampoo covers the entire scalp, the conditioner avoids the scalp, and the shower gel is applied to the elbow or knee, which is easy to overlook. These details are crucial for optimizing product formula viscosity, foaming, and rinsing difficulty.

[0103] In the female hygiene product testing scenario, the system is used to compare the differences in subsequent behavior after using different products (such as sanitary napkins vs. sanitary pads). After the volunteers complete the replacement operation in the private booth, they return to the observation area for daily activities such as sitting, walking, and bending over. Traditional methods require the installation of high-definition cameras outside the booth, which can easily cause discomfort and even refusal to participate. In this solution, the device is only deployed in the public activity area, and the picture does not contain human details, only showing the torso skeleton and lower limb movement trajectory. Researchers focus on observing whether there are abnormal behaviors such as frequent adjustment of sitting posture, leg crossing, and hesitation in walking after using a tampon; and whether there are specific limb reactions due to friction or leakage when using a sanitary pad. These behavior patterns can indirectly reflect product comfort, fit, and safety, providing objective evidence for product improvement.

[0104] All visualizations were watched by authorized researchers in an independent monitoring room. The content of the picture is a synthetic graph: the hand is displayed with bright yellow skeleton line segments, the face is represented with green discrete points, the torso is presented with white connected lines, and the background is a gray-scale enhanced event stream. Researchers can complete the behavior coding table filling according to this, such as "left cheek smearing times: 3 times", "hair conditioner stay time: 118 seconds", "post-cotton swab adjustment: 4 times within 5 minutes" and the like. The whole process does not need to contact the original pixel, nor does the volunteer appear in the mirror, which significantly reduces the difficulty of ethical review and participation threshold.

[0105] Experimental data show that in the skin care product test of 30 volunteers, 92% of them feel "almost no monitoring", and the behavior naturalness score is increased by 37% compared with the traditional video group; in the sanitary product comparison test, the abnormal behavior recognition accuracy of the cotton swab group reaches 89%, and no one withdraws due to privacy concerns. The average power consumption of the system is 82 mW, which supports continuous work for more than 8 hours, meeting the all-weather observation needs.

[0106] In summary, the embodiment provides a new paradigm for conducting daily and nursing product human factor tests in highly privacy-sensitive scenarios. Through event-driven vision and structured visualization, key behavior details are retained for manual observation and analysis, while identity and environment information collection is completely avoided, solving the three major pain points of privacy risk, behavior distortion and low participation willingness in traditional high-definition video observation, providing a more real, more compliant and more efficient data acquisition path for product development.

[0107] In the embodiment, a privacy protection type behavior observation system for home-based care safety monitoring is provided, which is designed to cope with the increasing demand for elderly care services. With the continuous increase in the number of single and empty-nest elderly people, more and more families rely on regular home visits by housekeeping or nursing personnel to provide basic care services such as cleaning, bathing, meal delivery, etc. However, due to the lack of effective process supervision mechanism, there are many problems in actual service: some service personnel are late or leave early, service time is insufficient, and operation is perfunctory. Many elderly people often cannot report or accurately describe the event due to decreased cognitive ability, communication barriers or fear of retaliation, resulting in long-term concealment of problems and difficulty in accountability.

[0108] At the same time, the elderly also face a large number of safety risks when they are alone at home. For example, getting up and walking without supervision can suddenly fall due to leg weakness; climbing a stool or standing on tiptoe to reach for an object at a high place is easy to lose balance and injure; even accidentally touching an electrical appliance, spilling hot water, etc. due to declining vision or judgment, causing accidental injury. The traditional solution is usually to install a high-definition camera at home for full recording to allow remote monitoring or post-mortem. But this kind of solution will record highly sensitive information such as faces, body conditions, home layouts, etc. involving semi-private areas such as bedroom and bathroom entrances, which is likely to cause strong resistance from the elderly and their families, and most families explicitly refuse to install it, making it difficult for the technology to be implemented.

[0109] To solve the above contradictions, a small secret observation device is deployed, which is similar in appearance to an ordinary inductor and can be hung on key viewing positions in public activity areas such as living rooms, kitchens, and corridors, strictly avoiding the interior of the bedroom and the bathroom. The device uses "event-driven vision" technology, and its core is a special sensor - event camera. Unlike ordinary cameras that take 30 frames of complete images per second, event cameras do not output any photos or videos, but perceive the brightness changes of each pixel point with microsecond-level precision: only when there is movement in a certain part of the picture (such as hand waving, getting up), the corresponding pixel will send out an "event" signal containing time, location and light change direction. This mechanism allows the system to work stably in complex home environments such as low light and strong light, and does not generate any identifiable image frames from the source, completely avoiding the risk of privacy leakage.

[0110] To enable remote safety monitors to efficiently understand the on-site situation, the system converts the original event stream into a visual picture with clear structure and explicit semantics. This picture is composed of three layers of information fusion: the bottom layer is a dynamic contrast-enhanced event stream grayscale picture, which only highlights the moving area and the static background is almost black; the middle layer is a color optical flow vector arrow, which accurately expresses the motion state of the limbs or objects; the upper layer is a simplified human skeleton line segment, which only contains 9 key points including head (single point), shoulder, elbow, wrist and trunk axis, without facial contour, gender characteristics and lower limb details, ensuring complete desensitization.

[0111] Among them, the optical flow arrow is the core tool for the system to realize efficient artificial monitoring. "Optical flow" is a technical concept in computer vision that describes "which direction and how fast a certain point moves." In this system, the optical flow is directly calculated from the event stream and is superimposed on the screen in the form of intuitive arrows: the arrow direction represents the movement trend (such as the arm stretching up, the body leaning left), the length reflects the speed (the longer the faster), and the color coding risk level - blue represents slow and stable (such as normal walking), green represents medium speed (such as holding a water cup), and red represents fast or sudden stop and start (such as suddenly waving hands, falling moment). More importantly, the distribution pattern of the arrow can reveal the behavior intention. For example, when the service staff normally wipes the table, the hand flow shows short and orderly blue arrows; if its arm suddenly accelerates towards the old man's chest or head area, multiple red arrows will densely point to the location, forming a clear "aggressive convergence" visual signal, even if there is no actual contact, it is enough to arouse the monitor's vigilance.

[0112] In the service supervision scene, the system automatically records the service staff's entry time, in-room duration and activity trajectory. If it stays still for a long time after entering the door, or only stays for a short time at the door and then leaves, the system will mark "abnormal service duration" and prompt the possibility of "clock in and walk away" behavior; if high-frequency red optical flow and skeleton posture mutation (such as rapid approach to the old man) appear during the service process, the monitor can immediately retrieve the enhanced picture of that period for manual review. In the old person's autonomous activity scene, the system also analyzes risks through the cooperation of optical flow and skeleton: when the old person stands on a stool to reach for the top shelf, the hand flow arrow extends upwards, and if the torso leans back at the same time, a horizontal red arrow will be generated in the waist area, prompting the loss of stability. Once rapid limb extension and rapid torso descent are detected, the system immediately highlights the full-body skeleton and superimposes radial red optical flow, marking it as "suspected fall", and the monitor can remotely confirm the situation through the voice device within a few seconds and coordinate the community emergency response.

[0113] All visualized pictures are only displayed on authorized security monitoring terminals, and the original event stream data is always encrypted and stored locally on the device, without leaving the house or going to the cloud. The security monitor watches the picture in a separate operation room, whose responsibility is only to identify abnormal behavior, respond to alarms, and record events, and cannot take screenshots, record videos, or export any content. Since the picture does not contain real images, only the movement logic and structural relationship are presented, which not only protects the dignity of the old people, but also significantly reduces the psychological burden of the monitor, enabling them to focus on work for a long time.

[0114] Experiments show that in the simulation of 20 households parallel monitoring tasks, the use of the system, the monitoring personnel on the service vacancy, violence, falls, high risk of climbing and other events of the comprehensive identification accuracy rate of more than 91%, the average response time is shortened to 8 seconds or less, and 98% of the participating families expressed willingness to use such equipment for a long time. In summary, the embodiment provides a scalable, efficient, and strong compliance security monitoring path for home-based care services through event-driven vision and fine fusion of optical flow-skeleton visualization, while fully protecting the privacy of home-based care.

[0115] In an optional embodiment, a privacy protection type behavior monitoring system for home facial massage service is provided, which is suitable for close-range human-computer interaction scenes such as on-site care and rehabilitation physiotherapy. The system not only supports fixed installation of wall-mounted security cameras, but also is compatible with wearable service recorders, which can be worn on the chest or shoulders of service personnel to capture the operation process from the first perspective. Both devices use event-driven vision (Event-Based Vision) technology to accurately record and identify risks without generating any image frames, while completely avoiding the leakage of user facial and body privacy.

[0116] In actual application, when professional masseurs provide facial acupoint pressing, lymphatic drainage or relaxation massage services for the elderly, the system can observe the overall posture from the side front downward angle through the wall-mounted device, or obtain the first-person perspective of hand-face interaction through the miniature event recorder (size about 25 mm × 25 mm × 8 mm) worn on the masseur's chest. The recorder is connected to the edge computing unit through low-power Bluetooth or USB-C interface, runs in EVS mode throughout the process, and only outputs asynchronous event stream, without caching, transmitting or reconstructing any picture containing texture, skin color or facial details, thus fundamentally meeting the strict control requirements of the Personal Information Protection Law on biometric information.

[0117] The system extracts three types of structured key points from the event stream in real time and visualizes them in a highly abstract and geometric form to ensure complete desensitization. First, there are 19 user facial key points that accurately correspond to anatomical locations, including bilateral eyebrow peaks (the highest points of the eyebrows), eyebrow tails, inner / outer canthi, nasal roots, nasal tips, bilateral alae nasi, upper philtrum points, upper and lower lip midpoints, bilateral corners of the mouth, lower jaw left and right edge endpoints, bilateral antitragus points (cartilage protrusions in front of the ear canal), and frontal center points. These points are represented as light gray small circles (diameter 3 pixels) in the visualization picture, without lines, filling or outlines, only reflecting spatial coordinates, and cannot restore facial shape, expression or identity features.

[0118] Secondly, the hand key points of the service personnel, 21 key points are extracted for each hand: including 5 finger tips, 5 distal interphalangeal joints, 5 proximal interphalangeal joints, 5 metacarpophalangeal joints, and 1 palm center point. These points are mapped as bright yellow hollow circles, constituting a simplified hand skeleton, whether from the wall-mounted perspective or the wearable recorder. When the masseur presses the temple with the thumb, the system can clearly show the spatial proximity of the thumb tip key point and the user's outer eye corner key point; if the fingers stay in the non-service area (such as below the lower jaw or neck) for a long time, the system will mark it as an "abnormal contact area". It is worth noting that all hand renderings remove identifiable details such as nails, skin texture, rings, etc., and only retain kinematic structures.

[0119] Thirdly, the human torso key points are used to judge the service distance and posture compliance. The system extracts 9 core points: head center, neck base (7th cervical vertebra), shoulder peaks, elbow outer sides, torso axis (lower end of sternum), hip center. These points are presented as white solid dots and connected by high-contrast white line segments to form a minimalist skeleton. The line segment width is uniform at 2 pixels, which is very eye-catching in the gray-black event stream background, allowing remote security personnel to easily determine whether the service personnel is leaning forward too much, invading the user's private space, or exhibiting high-risk postures such as sudden approach.

[0120] All key points are superimposed on a dynamically contrast-enhanced event stream background. This background highlights weak skin deformations (such as local event density changes caused by pressing) and hand sliding trajectories through an adaptive gain algorithm, but the overall tone is gray-black, with no highlights or shadows, making it impossible to identify furniture, clothing patterns, or environmental details. At the same time, the system superimposes optical flow vector arrows on the hand movement path: blue short arrows represent gentle sliding (speed <0.3 m / s), green medium-length arrows represent moderate pressing (0.3-0.6 m / s), and red long arrows represent fast or impact actions (>0.6 m / s). If red arrows are detected densely pointing to the user's eye or mouth area, even if no actual contact has occurred, a "high-risk operation" prompt will be triggered.

[0121] Experiments show that in a facial massage test involving 30 volunteers, the system's average error in key point positioning is less than ±1.7 centimeters (in the event stream reconstruction coordinate system), and the service area coverage recognition accuracy is 96.2%. More importantly, in the background monitoring link, two full-time security personnel independently operate in the room and conduct blind evaluation on 100 randomly selected visual images, and 94% of the images are unanimously judged as "clearly identifiable service behavior". The security personnel feedback: "Through the prominent white skeleton line segment and the highlighted hand key point, it can be clearly seen whether the masseur is operating in the standard area; although the facial key point is a small dot, the position is accurate, and it can be judged whether it is out of bounds or too forceful with the help of the optical flow arrow." All participants expressed their willingness to accept such unobtrusive monitoring, and 100% believed that "the picture does not contain real faces and is completely acceptable psychologically."

[0122] This embodiment provides quantifiable, traceable, and auditable technical support for home nursing services by fusing fixed and wearable event perception devices, combining precise key point definitions, high-contrast skeleton rendering, and multi-view collaborative analysis, while completely protecting user facial privacy. Background security personnel do not need to rely on real images, but only on structured and desensitized visual elements to efficiently and accurately complete service compliance supervision, especially suitable for home health service scenarios that are extremely sensitive to privacy.

[0123] In an optional embodiment, a home service behavior monitoring system for long-term care insurance regulatory scenarios is provided. The system is deployed in the insured's home after the insurance company purchases professional home care services (such as bath assistance, rehabilitation massage, life assistance, etc.) for disabled or semi-disabled elderly people who meet the conditions for compensation, and is used for whole-process, traceable, and high-privacy protection supervision and evaluation of the service process. Its core goals are threefold: first, to verify that the insured person actually receives the service and prevent fraud such as "false clocking" and "empty service"; second, to evaluate whether the service quality of the service provider meets the standards, including operation standardization, coverage area integrity, and time compliance; and third, to assist in judging the health status trend of the insured person through structured behavior data, providing a basis for subsequent insurance claim period adjustment or service plan optimization.

[0124] In this scenario, traditional methods that rely on manual spot checks or high-definition video review face significant obstacles: on the one hand, the home environment is highly private, and users refuse to install cameras; on the other hand, even if they are installed, video content involves sensitive information such as body exposure, facial features, and home layout, making it difficult to meet the requirements of the Personal Information Protection Law and insurance industry data compliance, and it cannot be applied to tens of thousands of policyholders on a large scale.

[0125] The embodiment adopts a privacy protection type monitoring device based on event-driven vision, which can be a wall-mounted or wearable recorder worn by a service personnel. The device runs an EVS mode (event-driven visual sensing) throughout the service process, only outputs a pixel-level asynchronous event stream, and does not generate, cache or locally store any image frame. All raw event streams are strictly limited within the device and do not go out of the device or go to the cloud, thereby eliminating the risk of identity information leakage from the source.

[0126] At the same time, the edge intelligent processing unit built in the device completes all AI inference and visual synthesis locally. Specifically, the unit first performs dynamic contrast enhancement processing on the raw event stream, accumulates events according to a time window, and adaptively adjusts the gain according to the local event density, so that weak actions (such as facial micro-tremor, finger light pressure, and slow getting up) can be clearly presented. The enhanced event stream after this processing directly reflects the real motion details of the monitored object (such as the face of the insured person, the hands of the service personnel, and the torso), which is the main data for subsequent visualization, rather than background decoration.

[0127] On this basis, the system further extracts three types of structured key points: insured person face key points (19, including eyebrow peak, eye corner, nose tip, mouth corner, and mandibular margin); service personnel hand key points (21 for each hand, covering fingertips to palm center); and human body torso key points (9, including head top, neck base, both shoulders, elbow, and torso axis). These key points are used to construct simplified skeleton line segments (connected by high-contrast white lines) and combined with optical flow vectors (color arrows, encoding motion direction, speed, and risk level) to jointly describe the space-time dynamics of service interaction.

[0128] Subsequently, the system fuses the enhanced event stream, key point coordinates, skeleton line segment data, and optical flow vector information to generate a set of completely desensitized structured visual data. This data does not contain any texture, skin color, face contour, or environmental details, but is only a synthetic expression of geometric symbols and motion vectors, but is sufficient to clearly present the behavior semantics of "who is moving, how to move, and where to move".

[0129] After the service is completed, the above desensitized visual data, together with the structured service behavior log (such as service start / end time, key action type, abnormal event marker, service area coverage, etc.), are encrypted and uploaded to the cloud server designated by the long-term care insurance underwriting unit. The raw event stream is still retained in the local device and never uploaded.

[0130] The authorized auditors (such as insurance company risk control officers, third-party service quality evaluators) can monitor the platform in the controlled cloud, review the visual playback of the specified period, and clearly determine whether the insured person is actually present and receives the service, whether the service personnel complete the agreed action (such as whether the facial massage covers the temple, cheekbone, etc.), whether there is service reduction (such as leaving after only 30 seconds) or abnormal behavior (such as long-time hovering of the hand in the non-service area), whether there is a sign of degradation of the insured person's activity ability (such as slower getting up and worse balance), and whether there is a sign of degradation of the insured person's activity ability (such as slower getting up and worse balance), which provides an objective basis for health status evaluation.

[0131] Through this mechanism, the insurance company can efficiently supervise large-scale home services without infringing on the privacy of users, effectively identify fraudulent behavior (such as service personnel "clocking in" alone in an empty house), and ensure that insurance funds are accurately used for real care needs. At the same time, the behavior transparency of service providers also promotes them to improve service quality, forming a closed loop.

[0132] In one possible implementation, the application is applied to the supervision of accompanying behavior compliance in the neurology department or geriatric ward of a three-A hospital. Such wards usually admit patients with mobility difficulties, cognitive disorders, or postoperative recovery periods, who rely on caregivers or family members to assist with basic care operations such as turning over, feeding, and cleaning. However, clinical management has long faced two major challenges: first, traditional video surveillance is prone to privacy infringement disputes due to the involvement of patients' naked bodies, facial features, and private ward environments, and most patients and family members explicitly refuse to install it; second, if there is no monitoring at all, it is impossible to effectively prevent individual caregivers from operating non-standardly (such as rough dragging, not turning over on time leading to pressure sores), service absence, or even abuse, and in the event of a dispute, there is a lack of objective evidence to support it.

[0133] To address the above problems, the present embodiment deploys a privacy-protected monitoring system based on event-driven vision. Specifically, an embedded device integrating an event camera module is installed in the corner of the ceiling of the ward, with a field of view covering the bed and the surrounding 1.5-meter radius area. This sensor only responds to changes in light intensity, outputs an asynchronous event stream containing timestamps, coordinates, and polarity, and does not generate, store, or transmit any continuous image frames throughout, fundamentally avoiding the collection of sensitive information such as faces, body textures, or ward furnishings.

[0134] The edge intelligence processing layer 102 is built locally in the device, runs a lightweight neural network model, and reconstructs the structured action features of the patient and the service personnel from the sparse event stream in real time. The system dynamically enables key point granularity according to the ward context: under normal circumstances, only 9 key points (shoulders, elbows, wrists, hip center, and upper and lower points of the spine) of the torso and upper limbs are activated to monitor large-scale posture changes; when high-risk operations such as feeding and turning over are detected, the 21-point fine model of the hand is automatically temporarily enabled to verify whether the correct gesture is used (such as holding with both hands instead of dragging with one hand).

[0135] The visualization output layer 103 fuses and renders the enhanced event stream, optical flow vector arrows, and skeleton line segments into a completely desensitized picture. For example, when performing a turning-over operation every two hours, the system determines whether the caregiver supports the patient's shoulder and hip synchronously through the skeleton connection, and the optical flow arrow reflects the movement speed - if the arrow is too long or the direction changes suddenly, it indicates that the action is too fast or there is a risk of horizontal dragging; if the event stream is abnormally dense in the patient's head area but there is no corresponding hand action, it may indicate unauthorized contact. Once the preset risk rules (such as falling, violent pushing, and long periods of inactivity) are triggered, the system immediately intercepts the structured data 10 seconds before and after, generates a timestamped explainable backtracking segment, and encrypts and uploads it to the hospital information department's controlled server.

[0136] Nursing managers can review the segment on authorized terminals, and only see geometric lines, motion arrows, and highlighted event points in the picture, without identifying identities or environmental details, meeting the requirements of the Personal Information Protection Law and medical privacy standards. At the same time, the system automatically generates a structured behavior log, including behavior type, time, location, duration, and compliance determination results, directly interfaces with the hospital quality control platform for monthly performance evaluation, adverse event review, or medical dispute evidence.

[0137] Through this embodiment, the hospital realizes objective recording and intelligent early warning of high-risk care behavior under the premise of zero image acquisition, which not only protects the dignity and privacy of patients, but also improves the fine level and risk prevention and control ability of nursing management, effectively solving the industry dilemma of "monitoring infringing and not monitoring losing control".

[0138] In one possible implementation, the present application is applied to the real-time early warning of abnormal behavior in a closed ward or rehabilitation ward of a mental health center. Such places admit patients suffering from severe depression, bipolar disorder, schizophrenia and other diseases, and some individuals have high-risk behaviors such as self-injury (e.g., hitting the wall, scratching), attacking others (e.g., pushing, punching), or sudden falls. Traditional safety management relies on manual video patrol and fixed camera video review, but has significant drawbacks: on the one hand, high-definition video recordings contain patient facial features, clothing, and ward layout, which can easily lead to privacy leaks and ethical disputes, and patients often worsen due to the feeling of being monitored; on the other hand, video surveillance cannot achieve real-time intervention - even if AI analysis is deployed, due to frame rate limitations and motion blur, the miss rate is high in rapid behavior, and it cannot meet the clinical demand for millisecond-level response.

[0139] To address the above challenges, the present embodiment deploys a monitoring system consisting of an event sensing layer 101, an edge intelligent processing layer 102, a visualization output layer 103, and an SDK encapsulation and application interface layer 104 on the wall or ceiling of the ward. The event sensing layer 101 integrates an event camera module to asynchronously output an event stream representing local light intensity changes with microsecond-level time resolution, without generating, storing, or transmitting any continuous image frames, completely avoiding identity and environment information collection.

[0140] The edge intelligent processing layer 102 is built into a local device and runs a lightweight spatiotemporal behavior recognition model to extract patient full-body keypoint skeletons (default 18-point pose model covering head, torso, and limbs) from the event stream in real time and simultaneously calculate the optical flow vector field. This layer determines abnormal behavior according to pre-set rules: when the event density in the torso area suddenly increases and the optical flow vector shows a high-speed vertical downward trend, it is determined to be a fall; when the hand keypoint rapidly approaches the other person's torso with radial optical flow, it is determined to be an attack intention; when the head area has a continuous high-frequency event outbreak without coordinated limb movement, it is marked as a self-injury tendency.

[0141] The visualization output layer 103 fuses and renders the event stream after dynamic contrast enhancement, directional arrows (representing optical flow vectors), and stick figure skeletons into a completely desensitized geometric picture. For example, in a sudden conflict, the system generates a red long arrow pointing to the chest of the patient in the adjacent bed 80 milliseconds after the patient's arm swings out, while the skeleton shows that the elbow-wrist connection is in an accelerated stretching state, immediately triggering an alarm. The on-duty nurse receives the structured warning information through the authorized terminal: "high-risk attack behavior, bed 3→ bed 5, time 14:27:06", and can review the 10-second backtracking segment uploaded in encryption - the picture only has the skeleton, arrow, and event point, and cannot identify the identity or environment details, meeting the requirements of the Mental Health Law for the protection of personal dignity.

[0142] All original event stream data are always kept within the edge intelligence processing layer 102, only structured behavior log and de-identified visualization data are encrypted and uploaded to the central server through the SDK encapsulation and application interface layer 104 for post-audit, plan optimization or quality control analysis by the nursing team.

[0143] Through the embodiment, the mental health institution realizes millisecond-level perception and accurate early warning of high-risk behaviors such as self-injury, attack, and fall under the premise of zero image acquisition, significantly improving the efficiency of emergency response; at the same time, because the system output content is irreversible and unrecognizable, it effectively alleviates the patient's resistance to monitoring, improves treatment compliance, and solves the core contradiction of "safety and respect" in psychiatric safety management.

[0144] In a possible implementation, the application is applied to the bath service process tracing scene in high-end home-based elderly care services. Bathing is one of the most sensitive and risky links in the daily care of disabled or semi-disabled old people, involving body exposure, close contact and slippery environment, which is easy to cause privacy infringement disputes or safety accidents such as slipping and scalding. The service provider needs to prove to the family members or long-term care insurance institutions that the service is truly completed and the operation is standardized, but the traditional methods (such as service clock-in and manual signature) cannot verify the authenticity of the process; while installing ordinary cameras is unacceptable in terms of ethics and law because they can capture details of the old people's bodies, and most families explicitly refuse.

[0145] To solve this contradiction, the embodiment integrates a non-intrusive monitoring system consisting of an event perception layer 101, an edge intelligence processing layer 102, a visualization output layer 103, and an SDK encapsulation and application interface layer 104 in the bathroom ceiling. Among them, the event perception layer 101 adopts a waterproof event camera module, which only responds to changes in light intensity, outputs an asynchronous event stream containing timestamps, coordinates, and polarity, and does not generate any image frames throughout the process, ensuring that even during bathing, skin texture, body outline, or facial features are not recorded.

[0146] The edge intelligence processing layer 102 is deployed in the local gateway device in the bathroom and runs a lightweight pose estimation model. The system automatically identifies the "bath mode" according to the spatial semantics: when it detects that the continuous low-frequency event background caused by water vapor is superimposed on the human movement event, it activates the upper body key point skeleton (including shoulder, elbow, wrist, hip, and trunk axis, a total of 9 points), and temporarily enables the 21-point fine model of the service personnel's hand to track whether it has completed the cleaning action of the specified area such as the back and arms according to the specification. At the same time, this layer calculates the optical flow vector to determine whether the action is smooth - if it detects a sudden increase in event density in the trunk area and the optical flow arrow is perpendicular downward, it is determined to be a slip risk.

[0147] The visualization output layer 103 fuses the enhanced event stream, direction arrow and simplified skeleton to render a fully desensitized picture. For example, in a standard bathing service, the picture shows the key points of the service personnel moving slowly along the back of the old man, the event stream forms a continuous highlight track in the corresponding area, and the optical flow arrow is short and flat, indicating that the action is gentle; if the service personnel only stays at the door and then leaves, the event stream is sparse and there is no hand-trunk interaction, and the system automatically marks "service not substantially performed". All abnormal or complete service segments generate timestamped structured backtracking data by this layer.

[0148] After encryption, the above data is uploaded to the pension service platform through the SDK encapsulation and application interface layer 104 for authorized viewing by family members or verification by insurance companies. The family members see only abstract pictures composed of geometric skeletons, motion arrows and event points on the mobile phone, and cannot identify the identity of the old man or the bathroom environment, but can clearly confirm whether the service has occurred, whether the operation covers the specified parts, and whether there is a risk of falling.

[0149] Through this embodiment, the bathing service is realized without collecting visual images, and the process is verifiable, the behavior is traceable, and the risk is prewarned. It not only meets the compliance requirements of service authenticity audit and insurance claims, but also fully respects the dignity and privacy of the elderly, effectively solving the industry problem of "the most difficult to supervise bathing service", and providing a feasible technical path for intelligent health and wellness services in high privacy sensitive scenarios.

[0150] In one possible implementation, the present application is applied to the service quality third-party verification scene of a rehabilitation sanatorium. However, the traditional verification method mainly relies on service check-in table, manual follow-up or video spot check, which has obvious defects: paper records are easy to be falsified, follow-up results are highly subjective, and calling conventional monitoring videos not only violates the restrictive provisions of the Personal Information Protection Law on the processing of sensitive personal information, but also is highly invasive to personal dignity, which is explicitly refused by most service objects, resulting in a long-term lack of objective basis for supervision.

[0151] To solve this contradiction, the present embodiment deploys a non-intrusive monitoring system composed of an event perception layer 101, an edge intelligent processing layer 102, a visualization output layer 103 and an SDK encapsulation and application interface layer 104 in the rooms, public activity areas and corridors of the sanatorium. Among them, the event perception layer 101 uses a low-power event camera module, which only responds to local light intensity changes, outputs an asynchronous event stream containing timestamps, coordinates and polarity, and does not generate, store or transmit any continuous image frames throughout the process, fundamentally eliminating the collection of identity information and environmental details.

[0152] The edge intelligent processing layer 102 is deployed in a local edge device, runs a lightweight behavior recognition model, and dynamically enables key point granularity according to spatial semantics: only 9 key points of the trunk and upper limbs are activated in a normal living room to determine whether there is a care interaction; in high-privacy areas such as bathing and toilet assistance, a 21-point fine hand model is temporarily enabled to verify whether the operation covers the specified parts. This layer synchronously performs behavior semantic analysis and generates structured logs, including service type, start and end time, duration, involved body parts, and compliance determination results, and automatically marks abnormal situations such as "no service after clocking in", "insufficient service duration", or "operation path deviates from the standard process".

[0153] The visualization output layer 103 fuses and renders the event stream after dynamic contrast enhancement, the optical flow vector arrow, and the simplified skeleton into a completely desensitized geometric picture. For example, in a "assisting in taking medicine" service, the picture shows that the hand key points of the service personnel are close to the old man's head area, the event stream forms a short highlight around the mouth, the optical flow arrow points smoothly to the inside, and the skeleton shows that the old man's head is slightly tilted - these features together constitute explainable service completion evidence. After all data is encrypted by the national encryption algorithm, it is uploaded to the controlled audit platform through the SDK encapsulation and application interface layer 104.

[0154] Authorized third-party evaluators can review the backtracking segments of the specified period on a secure terminal, and the picture only presents abstract skeletons, motion arrows, and event points, which cannot identify the identity of the elderly, the layout of the room, or the body features, but can objectively confirm whether the service is real and the action is in line with the standard. The structured logs output by the system can be directly connected to the service acceptance or expense settlement system to realize the closed-loop management of "service execution - quality verification - resource allocation".

[0155] Through this embodiment, the rehabilitation sanatorium establishes an objective, traceable, and auditable service quality verification mechanism without collecting any visual images or exposing any personal identity information, fully respecting the dignity and privacy rights of the elderly service objects, and improving the transparency and use efficiency of public resource allocation, effectively solving the structural contradiction between "service authenticity difficult to verify" and "privacy protection rigid requirements" in high-privacy sensitive scenarios.

[0156] In a possible implementation, the application is applied to a home rehabilitation scene, and serves the elderly and chronic patients with motor dysfunction due to reasons such as stroke, joint replacement, or Parkinson's disease. Existing remote rehabilitation schemes mostly rely on smartphones to shoot training videos or wear inertial sensors, the former involves image collection of home environment and body posture, which may cause privacy concerns; the latter has problems such as device wearing discomfort, short battery life, and complex operation, resulting in a significant decrease in the compliance of elderly users.

[0157] To solve the above problems, the embodiment deploys a set of home rehabilitation assistance system based on event-driven sensing and hierarchical intelligent processing, including event sensing layer 101, edge intelligent processing layer 102, visual output layer 103, and SDK encapsulation and application interface layer 104. Only an embedded device integrating the event sensing layer 101 needs to be installed in the patient's home (such as placed on the wall or ceiling), which uses an event camera module to record only the spatio-temporal event stream of light intensity changes, and does not generate any image frames, fundamentally avoiding the leakage of identity features, home environment or body exposure information.

[0158] The edge intelligent processing layer 102 runs a lightweight pose estimation model (such as a real-time key point detection network based on MobileNetV3 backbone), which extracts the trunk and limb key points of the patient performing rehabilitation actions (default 18-point skeleton enabled) from sparse event streams. The model completes joint angle calculation, action cycle segmentation and basic compliance judgment (for example: whether the upper limb abduction reaches the prescribed angle range, whether the double support phase in gait training is too long) on the device side. All processing is completed locally, and the original event stream does not leave the device, only the structured action features (such as key point coordinate sequences, angular velocity, symmetry indicators) are uploaded after encryption.

[0159] These desensitization data are gathered to the regional rehabilitation data center through the SDK encapsulation and application interface layer 104. The center builds a rehabilitation behavior knowledge base, continuously accumulating standardized action feature sequences from patients of different ages, causes and disease stages. On this basis, the system uses a transfer learning strategy: when a new patient starts training, the platform first matches their demographic characteristics (age, gender), diagnosis type (such as "left hemiplegia" "right knee replacement postoperative") and baseline functional level, retrieves the historical training trajectory of similar cases in the knowledge base, and dynamically adjusts the evaluation threshold and progression pace of the current patient. For example, for elderly osteoporosis patients, the system automatically relaxes the center of gravity offset tolerance in balance training; for young stroke patients, more emphasis is placed on action speed and coordination.

[0160] Rehabilitation therapists can access the backtracking pictures generated by the visual output layer 103 through authorized terminals, the content of which is only stickman skeleton, optical flow direction arrow and event highlight area, which cannot identify personal identity or room details, but is sufficient to judge the quality of the action. The system provides a structured summary: "the patient completed 3 sets of shoulder joint abduction training today, the average angle is 76°, which is 5° higher than last week, and the right side compensation is reduced", which helps therapists remotely adjust the rehabilitation plan.

[0161] The scheme replaces video with event stream, greatly reduces data volume (usually less than 1% of video data), and makes stable transmission in weak network environment possible; edge side completes real-time feedback, ensuring instant training, and does not rely on cloud response; cloud uses real-world rehabilitation data to build an evolving evaluation benchmark, so that personalized guidance no longer relies on the experience of a single expert, but on the statistical rules of group behavior patterns; there is no image and no biometric template throughout the process, which meets the requirements of medical privacy protection and improves the acceptance of elderly users.

[0162] Through this combination of technologies, limited rehabilitation resources can be efficiently reused, therapists do not need to watch videos frame by frame, but make decisions based on structured insights; patients receive close-to-professional training supervision without feeling, disturbing, and privacy risks. This model provides a scalable and sustainable iterative technical path to address the rehabilitation service gap in an aging society.

[0163] The above embodiments are only used to specifically illustrate the technical principles, implementation processes and typical application scenarios of the present application, and should not be understood as limiting the scope of protection of the present application. Those skilled in the art should understand that without departing from the disclosed core concept and essential technical features, the method steps order, module division mode, neural network structure, event stream processing strategy, key point definition number and position, skeleton connection topology, optical flow calculation algorithm, visualization rendering form, data upload mechanism, edge-cloud collaboration logic, and hardware deployment form in the foregoing embodiments can be variously adjusted, functionally equivalent replaced or reasonably extended.

[0164] For example, the number of human key points can be switched between 18, 21, 25 or even higher granularity according to OpenPose, MediaPipe or other pose estimation standards; the topology connection rules of the hand skeleton can dynamically enable or disable specific joint connection lines according to the service action type (such as holding, tapping, sliding); the visualization of optical flow vectors is not limited to direction arrows, but can also use color coding, motion trajectory lines, speed heat maps, etc. to express the direction and intensity of motion; dynamic contrast enhancement of event stream can combine Time Surface, event frame accumulation, adaptive window sliding, and other known or future developed event representation methods; the edge intelligent processing unit can be deployed in a dedicated AI chip, embedded SoC, FPGA or micro server, and its local processing boundary can be flexibly configured according to privacy compliance levels; the generation rules of structured behavior logs can also be connected with insurance work order systems, service quality scoring models or anomaly detection engines to realize automatic verification.

[0165] In addition, the security mechanisms such as "desensitization", "out-of-device", "encrypted upload", "controlled review" and the like described in the present application can be compatible with domestic and foreign regulations such as GDPR, Personal Information Protection Law, Data Security Law and the like, and support the use in combination with privacy-enhancing technologies such as zero-trust architecture, end-side trusted execution environment (TEE), federated learning and the like. Any obvious modification, combination, simplification, parameter adjustment, interface adaptation or cross-scene migration made based on the technical idea of the present application, as long as the function realized and the technical effect achieved are substantially the same as the present application, shall be considered to fall within the protection scope of the present application.

[0166] Therefore, the protection scope of the present application should not be limited to the specific description of the above embodiments, but should be subject to the appended claims and all equivalent replacements, reasonable extensions and functional interpretations thereof.

Claims

1. A privacy-preserving service behavior monitoring method, characterized in that, include: Light intensity change information in the service scene is collected by an event-driven asynchronous visual sensor. The sensor outputs event stream data representing the time and polarity of the relative change in local light intensity in a pixel-level asynchronous manner, and does not generate or output continuous image frames containing complete scene texture. Based on the event stream data, a neural network model is used to extract structured action features related to human behavior. The structured action features are geometric representations without identifiable identity information. Based on the structured action characteristics, perform behavioral semantic analysis related to service compliance to generate structured behavioral records for service quality supervision, risk warning, or accountability. The method enables non-intrusive monitoring of the service process without acquiring, storing, or transmitting any image data containing details of facial textures, body appearance, or background environment.

2. The method as described in claim 1, characterized in that, During the visualization process, the event stream data undergoes dynamic contrast enhancement to improve the brightness of the event points and the distinction between them and the background area. The enhanced event stream, the structured action features, and the optical flow vector field calculated based on the event stream are then overlaid and displayed. The optical flow vector field is presented in the form of directional arrows, with the arrow pointing to indicate the local motion direction and the length indicating the speed of motion. This is used to assist human inspectors in identifying falls, abnormal movements, or service contact behaviors.

3. The method as described in claim 1, characterized in that, The structured motion features include a multi-granularity human keypoint skeleton, which covers at least two categories of full-body posture, local limbs and facial regions, and dynamically enables the corresponding granularity according to the service scenario.

4. The method as described in claim 3, characterized in that, The partial limbs include the hands, and the structured motion features of the hands represent gestures and operational intentions through joints and their topological connections, which are used to identify fine motor service actions such as holding tools, feeding, or cleaning.

5. The method as described in claim 3, characterized in that, The structured motion features of the facial region consist of several anatomical key points, used to characterize the spatial relationship between the service tool and the facial contact area, without reproducing the original face image.

6. The method as described in claim 1, characterized in that, The behavioral semantic analysis includes at least one of the following: fall event detection, service action and work order matching verification, individual traffic statistics, or abnormal behavior identification.

7. A privacy-protected service behavior monitoring system, characterized in that, include: An event-driven asynchronous vision sensor is configured to output event stream data in a pixel-level asynchronous manner and does not output continuous image frames. The edge intelligence processing unit integrates a lightweight neural network to generate desensitized structured action features from the event stream data in real time and perform behavioral semantic analysis. A visualization output unit is used to overlay and render the structured action features with a dynamically contrast-enhanced event stream into a readable geometric shape. The SDK packaging unit packages the above functions into an embedded software development kit that can be deployed in cameras, wearable devices, service robots, smart lights, or health and wellness facilities.

8. The system as described in claim 7, characterized in that, The event-driven asynchronous vision sensor is integrated into the event camera module, which includes a dynamic vision sensor chip, an optical imaging component, and an embedded preprocessing circuit.

9. The system as described in claim 7, characterized in that, The edge intelligent processing unit completes all data processing locally, with the original event stream not leaving the device. Based on the original event stream, the edge intelligent processing unit performs dynamic contrast enhancement processing to generate an enhanced event stream reflecting the motion details of the monitored object. It further extracts key points of the user's face, key points of the service personnel's hands, and key points of the human torso, constructs simplified skeleton segments, and calculates optical flow vectors. Subsequently, the enhanced event stream, key point coordinates, skeleton segment data, and optical flow vector information are fused to form desensitized structured and visualized data, which is encrypted and uploaded to the cloud server along with service behavior logs for authorized security monitors to review and audit in a controlled environment.

10. The system as described in claim 7, characterized in that, When a preset risk event is triggered, the system automatically extracts the structured action feature sequence and enhanced visualization of the corresponding time period, and generates an interpretable backtracking segment with a timestamp for accident verification or compliance evidence collection. The system is deployed in highly privacy-sensitive locations such as rooms in elderly care institutions, wards in medical institutions, bathrooms in homes, or nursing homes.