Service process monitoring method and system based on tof radar imaging

CN122513431APending Publication Date: 2026-08-04JIAXING WANFENG INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
JIAXING WANFENG INFORMATION TECH CO LTD
Filing Date
2026-05-13
Publication Date
2026-08-04

AI Technical Summary

Technical Problem

1、传统视频监控依赖可见光成像,可采集人脸、身体、环境等隐私信息,无法在卧室、浴室等私密区域部署,隐私与记录不可兼得;

Benefits of technology

1、全隐私保护:TOF雷达传感器无可见光成像,从物理源头杜绝隐私泄露;

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122513431A_ABST
    Figure CN122513431A_ABST
Patent Text Reader

Abstract

This invention discloses a service process monitoring method and system based on Time-of-Flight (TOF) radar imaging. The system includes a front-end TOF radar monitoring device, a cloud-based AI model layer, a data processing layer, and an application management layer. The method includes steps S1: multi-source data acquisition; step S2: data preprocessing and de-identification; step S3: data transmission; and step S4: cloud-based behavior recognition. The cloud server identifies service behaviors through a model and performs a comprehensive evaluation of service compliance, ultimately generating a service report. This invention provides a fully privacy-preserving service process recording solution; achieves accurate identification of service actions; supports 3D reconstruction of the service site, forming unalterable objective evidence for easy dispute resolution; and integrates multi-mode positioning and one-click assistance, simultaneously achieving service monitoring and personnel safety assurance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of intelligent service process supervision technology, specifically relating to a service process supervision method and system based on TOF radar imaging. Background Technology

[0002] In one-on-one private service scenarios such as in-home housekeeping, home care, institutional elderly care, and long-term care insurance, the authenticity, compliance, and personnel safety supervision of the service process are core pain points in the industry. Current mainstream regulatory methods have significant shortcomings: 1. Traditional video surveillance relies on visible light imaging, which can collect private information such as faces, bodies, and the environment. However, it cannot be deployed in private areas such as bedrooms and bathrooms, and privacy and recording cannot be achieved at the same time. 2. Millimeter-wave radar can only detect the presence and rough movement of people, and cannot distinguish specific service actions such as assisting to turn over, massaging, or feeding, resulting in insufficient accuracy in behavior recognition; 3. Recording audio and taking photos can only record fragmented information and cannot restore the spatial structure of the service site, making it difficult to provide evidence in disputes; 4. Most devices do not integrate multi-mode positioning and one-click emergency assistance, leaving service personnel without safety guarantees in private environments.

[0003] Existing solutions cannot achieve a unified approach to privacy protection, accurate behavior recognition, on-site reconstruction, and personnel safety protection, making it difficult to meet the compliance and regulatory needs of scenarios such as long-term care insurance, domestic services, and elderly care. Summary of the Invention

[0004] The main objective of this invention is to provide a service process monitoring method and system based on TOF radar imaging, offering a fully privacy-preserving service process recording solution that eliminates the collection of sensitive information such as faces, bodies, and environments from the physical source; enabling accurate identification of service actions, automatically determining standard service items such as turning over, massage, feeding, and cleaning; supporting three-dimensional reconstruction of the service site to form tamper-proof objective evidence for easy dispute resolution; and integrating multi-mode positioning and one-click assistance to simultaneously achieve service monitoring and personnel safety assurance.

[0005] To achieve the above objectives, this invention provides a service process monitoring system based on TOF radar imaging, comprising a front-end TOF radar monitoring device, a cloud-based AI model layer, a data processing layer, and an application management layer, wherein: Front-end TOF radar monitoring device: used to collect depth point cloud data, locally de-identify data, compress and upload data, and provide button, voice, remote control interaction and one-click help function; Cloud-based AI model layer: Employing a Transformer architecture that fuses spatiotemporal features, it constructs a sample library of actions for long-term care insurance, domestic services, and institutional elderly care services, and continuously optimizes recognition accuracy through incremental learning and knowledge distillation; Data processing layer: used for deep data parsing, point cloud denoising, human key point extraction, service behavior recognition, service item matching, and compliance scoring; Application Management Layer: Used to generate service compliance reports, service duration statistics, compliance audits, anomaly alerts, location trajectory playback, remote viewing by family members, and display on the monitoring platform.

[0006] As a further preferred embodiment of the above technical solution, the front-end TOF radar monitoring device includes a computing power control module, a sensing module, a human-computer interaction module, a positioning module, a communication transmission module, a power management module, and a local storage module, wherein: The computing power control module adopts an ARM architecture processor and integrates an NPU neural network acceleration unit. The perception module includes a TOF radar sensor and an audio acquisition sensor. The TOF radar sensor is used to emit lasers and receive reflected signals to obtain depth image information of the service scene. The audio acquisition sensor is used to record ambient sound and human voice in real time for voice interaction and environmental sound analysis. The human-computer interaction module includes a color LCD screen, an audio speaker unit, a button unit, an offline voice unit, and a remote control unit, supporting service start / stop, one-click help, and voice control; The positioning module integrates a BeiDou satellite positioning module, a WiFi-assisted positioning module, and a Bluetooth indoor positioning module; The communication transmission module uses 4G as the primary communication method and Bluetooth as the auxiliary communication method, and supports real-time data transmission in environments without WiFi. The local storage module is used for offline data caching and emergency log retention.

[0007] This invention also provides a service process monitoring method based on TOF radar imaging, comprising the following steps: Step S1: Multi-source data acquisition, including sound data acquisition, depth data acquisition, and location data acquisition; Step S2: Data preprocessing and desensitization. Denoising and normalization are performed on the depth data, and the depth data is converted to the color space according to the four-segment linear color mapping algorithm to achieve visualization display without privacy information. Step S3: Data transmission. The depth data, sound data, and location data are compressed using a balanced lossless compression algorithm and uploaded to the cloud server in real time via a wireless network. Step S4: Cloud-based behavior recognition. The cloud server uses the model to recognize service behaviors and completes a comprehensive assessment of service compliance, ultimately generating a service report.

[0008] As a further preferred technical solution to the above technical solution, for step S2, when the service personnel begin their service, the device is positioned facing the service personnel and the service recipient at an appropriate distance. The built-in TOF radar sensor continuously scans the service area, and the collected depth data is converted into color space data using a four-segment linear color mapping algorithm and displayed on the device's configured color LCD screen. The principle of the four-segment linear color mapping algorithm is as follows: The depth values ​​are divided into four segments: 0-L / 4, L / 4-L / 2, L / 2-3L / 4, and 3L / 4-L. These segments are mapped to a gradient color space from red to yellow, yellow to green, green to cyan, and cyan to blue, respectively. While preserving the spatial structure and movement changes, privacy details, including those of faces, bodies, and the environment, are removed from a physical perspective.

[0009] As a further preferred technical solution to the above technical solution, step S4 is specifically implemented as follows: Step S4.1: Model architecture design, adopting a Transformer-based spatiotemporal feature fusion network, and custom design for the service behavior recognition task of TOF deep point cloud data; Step S4.2: Spatial feature extraction module, using the improved PointNet++ structure, performs local feature aggregation on each frame of point cloud; Step S4.3: Temporal feature aggregation module, introduces temporal sequence information into each frame of point cloud data by temporal position encoding, enabling the model to perceive the relative positional relationship of each frame in the sequence and capture the inter-frame dependency relationship through temporal self-attention calculation, thereby realizing continuous action modeling; Step S4.4: Construct a service behavior sample library and perform data augmentation strategies including random rotation, translation, scaling, and Gaussian noise addition; Step S4.5: Feature extraction algorithm, including depth point cloud preprocessing, human pose feature extraction, and interaction relationship feature extraction; Step S4.6: Behavioral temporal analysis algorithm, including action segmentation, temporal action modeling, and action sequence similarity calculation; Step S4.7: Service item matching algorithm, including multi-level matching strategy, feature matching calculation and service item matching probability calculation; Step S4.8: Comprehensive service compliance assessment, including multi-dimensional assessment indicators, comprehensive compliance score, and service report generation.

[0010] As a further preferred technical solution to the above technical solution, step S4.5 is specifically implemented as follows: Step S4.5.1: Depth point cloud preprocessing, noise filtering uses a statistical outlier removal algorithm: like Then remove that point; in: Indicates the point to be judged; express The j-th nearest neighbor; Indicates the number of nearest neighbors; Point The average distance to its k nearest neighbors; Point The mean distance between all pairs of points in the neighborhood; The standard deviation represents the neighborhood distance; This represents the outlier detection threshold coefficient; if the condition is met, the point is marked as an outlier and removed; after traversing all points, filtered clean point cloud data is obtained, thereby removing isolated points caused by sensor noise and environmental interference; Step S4.5.2: Human pose feature extraction. Based on TOF depth data, an improved 3D human pose estimation network is used to detect preset human key points and calculate skeletal vectors. Specifically: Step S4.5.2.1: Extract the key point locations of the human skeleton from the point cloud data, and calculate the skeleton vectors between adjacent joints based on the topological structure of the human skeleton; Skeleton vectors: ; , Let represent the three-dimensional coordinate vectors of the i-th and j-th human body key points, respectively; From key points Pointing to key points The skeletal vector represents the spatial connection between two joints; Step S4.5.2.2: Calculate the bending angle of each joint using vector operations, and the angle between the bone vector and the vertical direction: ; It is a unit vector in the vertical direction in the spatial coordinate system; The magnitude of the skeletal vector; Step S4.5.2.3: Calculate the motion velocity of each joint using temporal difference. Velocity: ; Let be the three-dimensional coordinates of the i-th key point at the current time t; Let be the three-dimensional coordinates of the i-th key point at the previous time t-1; The time interval between two frames; Step S4.5.2.4: Concatenate the skeleton vector, angle features, and velocity features to form a complete pose feature vector. eigenvectors: ; Step S4.5.3: Interaction relationship feature extraction, specifically: Step S4.5.3.1: Calculate the interaction distance between human bodies, and measure the degree of interaction by the Euclidean distance between the two central joints of the human body; , The distance between human bodies during interaction; Indicates the location of the central joint of the human body; To represent the L2 norm; Step S4.5.3.2: Extract the interaction features between the human body and objects and calculate the interaction strength between the human body and objects through a neural network; ; The intensity of interaction between the human body and objects; , To represent the learnable weight matrix and bias vector; Features of the hands; Features of an object; Step S4.5.3.3: Determine the contact state based on the distance between the hand and the object. When the distance... Less than the threshold It is considered that contact has occurred; Step S4.5.3.4: Combine the interaction distance, interaction intensity, and contact state to form an interaction relationship feature vector.

[0011] As a further preferred technical solution to the above technical solution, step S4.6 is specifically implemented as follows: Step S4.6.1: Action segmentation. Temporal action boundary detection uses a dual-stream network. For TOF data, depth frame difference is used instead of optical flow, specifically: Step S4.6.1.1: Calculate the motion difference between adjacent frames. ; Step S4.6.1.2: Perform smoothing filtering on the motion difference sequence to eliminate the influence of noise; Step S4.6.1.3: Detect local minima of motion differences as candidate points for action boundaries; Step S4.6.1.4: Verify the action boundary by combining the rate of change of human posture, and ensure that the segmentation point is located at the moment of action transition; Step S4.6.1.5: Divide the continuous frame sequence into independent action segments based on the verified boundary points; each action segment contains a complete action cycle for subsequent action recognition and matching; Step S4.6.2: Temporal action modeling, each service action is modeled as a Hidden Markov Model and decoded using Viterbi; Step S4.6.3: Calculate action sequence similarity. Dynamic time warping is used to calculate action sequence similarity. Specifically: Step S4.6.3.1: Construct the distance matrix D; Step S4.6.3.2: Construct the cumulative distance matrix C; Step S4.6.3.3: Backtrack from the bottom right corner to the top left corner to find the optimal alignment path; Step S4.6.3.4: Calculate the sum of all local distances on the path to obtain the DTW distance; the smaller the DTW distance, the more similar the two action sequences are.

[0012] As a further preferred technical solution to the above technical solution, step S4.7 is specifically implemented as follows: Step S4.7.1: The multi-level matching strategy adopts a three-level matching strategy, specifically as follows: First layer: coarse-grained category matching; Second layer: Fine-grained item matching; Third level: Motion quality assessment; Step S4.7.2: Feature matching calculation, specifically: Cosine similarity Cosine similarity measures the degree of similarity between two vectors in a direction; the larger the value, the more similar they are. Euclidean distance similarity Euclidean distance similarity uses a Gaussian kernel function to map the Euclidean distance to the [0,1] interval; the smaller the distance, the higher the similarity. Overall matching score The comprehensive matching score, by weighting and fusing multiple similarity indicators, comprehensively evaluates the similarity of action sequences from different perspectives, thereby improving the accuracy and robustness of matching. Step S4.7.3: Calculate the service item matching probability and obtain the judgment result, which includes confirmed service item, pending manual review, and matching failure.

[0013] As a further preferred technical solution to the above technical solution, step S4.8 is specifically implemented as follows: Step S4.8.1: Multi-dimensional evaluation indicators include service duration compliance Completeness of movement and motion quality ; Step S4.8.2: The overall compliance score is ; Indicates the probability of a service item matching; , , , Indicates weight; Step S4.8.3: Generate a service process compliance assessment report, including: Basic information: Service personnel ID, service recipient ID, service time, service location; Service item identification results: identified service item, matching probability, and confidence level; Action sequence recording: timestamps of key action nodes, action type, and duration; Compliance score: Overall compliance score and scores in each dimension; Anomaly labeling: low confidence periods, missing actions, and abnormal behaviors; Location trajectory: Records of location changes during the service process.

[0014] The beneficial effects of this invention are as follows: 1. Full privacy protection: TOF radar sensors do not use visible light for imaging, thus eliminating privacy leaks at the physical source; 2. Precise behavior recognition: It can distinguish specific service actions such as turning over, massaging, and feeding, ensuring that supervision is not merely a formality; 3. Quantifiable services: Automatic verification of service items improves efficiency and credibility; 4. On-site reconstruction is possible: The 3D point cloud preserves the spatial structure, allowing for objective evidence in disputes; 5. Safety and supervision in one: Multi-mode positioning and one-click emergency assistance protect the safety of service personnel; Attached Figure Description

[0015] Figure 1 This is a schematic diagram of the front-end TOF radar monitoring device of the present invention.

[0016] Figure 2 This is a schematic diagram of the regulatory method of the present invention.

[0017] Figure 3 This is a service item matching determination logic diagram of the present invention (showing the complete determination logic from deep data input to service item matching probability output).

[0018] Figure 4 This is a schematic diagram illustrating the conversion of depth data to color space according to the present invention.

[0019] Figure 5 This is a schematic diagram of the detection of 17 key points of the human body according to the present invention. Detailed Implementation

[0020] The following description is intended to disclose the present invention and enable those skilled in the art to implement it. The preferred embodiments described below are merely examples, and other obvious variations will occur to those skilled in the art. The basic principles of the invention defined in the following description can be applied to other embodiments, modifications, improvements, equivalents, and other technical solutions that do not depart from the spirit and scope of the invention.

[0021] In the preferred embodiments of the present invention, those skilled in the art should note that the TOF radar and the like involved in the present invention can be considered as prior art.

[0022] Preferred embodiment.

[0023] like Figure 1 As shown, this invention discloses a service process monitoring system based on TOF radar imaging, including a front-end TOF radar monitoring device, a cloud-based AI model layer, a data processing layer, and an application management layer, wherein: Front-end TOF radar monitoring device: used to collect depth point cloud data, locally de-identify data, compress and upload data, and provide button, voice, remote control interaction and one-click help function; Cloud-based AI model layer: Employing a Transformer architecture that fuses spatiotemporal features, it constructs a sample library of actions for long-term care insurance, domestic services, and institutional elderly care services, and continuously optimizes recognition accuracy through incremental learning and knowledge distillation; Data processing layer: used for deep data parsing, point cloud denoising, human key point extraction, service behavior recognition, service item matching, and compliance scoring; Application Management Layer: Used to generate service compliance reports, service duration statistics, compliance audits, anomaly alerts, location trajectory playback, remote viewing by family members, and display on the monitoring platform.

[0024] Specifically, the front-end TOF radar monitoring device includes a computing power control module, a sensing module, a human-computer interaction module, a positioning module, a communication transmission module, a power management module, and a local storage module, wherein: The computing power control module adopts an ARM architecture processor and integrates an NPU neural network acceleration unit. The perception module includes a TOF radar sensor and an audio acquisition sensor. The TOF radar sensor is used to emit lasers and receive reflected signals to obtain depth image information of the service scene (i.e., three-dimensional point cloud data. The TOF radar sensor does not have visible light RGB imaging function, thus preventing privacy leakage from the physical source). The audio acquisition sensor is used to record ambient sound and human voice in real time for voice interaction and environmental sound analysis. The human-computer interaction module includes a color LCD screen (for real-time display of the current depth image, allowing service personnel to confirm the image perspective and family members to confirm the non-privacy characteristics of the image), an audio speaker unit (the device is equipped with a speaker for necessary voice reminders at the service site), a button unit (the device is equipped with a mechanical button for powering on / off and for physical control during the service process (such as starting or ending service records)), an offline voice unit (the device has a built-in offline AI voice module for voice-controlled actions during the service process, such as starting or ending service), and a remote control unit (supporting service personnel to remotely control necessary functions during the service process, such as starting or ending service, and supporting one-click emergency calls for help in case of emergencies). It supports service start / stop, one-click help requests, and voice control. The positioning module integrates a BeiDou satellite positioning module, a WiFi-assisted positioning module, and a Bluetooth indoor positioning module; The communication transmission module uses 4G as the primary communication method (ensuring that the collected multi-mode data can be transmitted to the cloud server in real time even in the absence of WiFi), and Bluetooth as the auxiliary communication method (it has a built-in Bluetooth master-slave integrated module, which can communicate with Bluetooth beacons to achieve indoor auxiliary positioning, and can also work with mobile phones to achieve local device configuration and other functions), supporting real-time data transmission in the absence of WiFi. The local storage module is used for offline data caching and emergency log retention.

[0025] like Figure 2-5 As shown, the present invention also provides a service process monitoring method based on TOF radar imaging, comprising the following steps: Step S1: Multi-source data acquisition, including sound data acquisition, depth data acquisition, and location data acquisition (sound data acquisition involves collecting ambient sound and human voices at the service site through a built-in microphone; depth data acquisition involves continuously scanning the service area using a TOF radar sensor to collect 3D point cloud data. Each frame of the point cloud contains N points, and each point has 3D coordinates (x, y, z); location data acquisition involves obtaining service location information through satellite positioning, WiFi positioning, and Bluetooth positioning). Step S2: Data preprocessing and desensitization. Denoising and normalization are performed on the depth data (Gaussian filtering). The depth data is then converted to the color space (RGB565) according to the four-segment linear color mapping algorithm to achieve visualization display without privacy information. Specifically, in step S2, when the service personnel begin their service, the (TOF radar imaging-based service process monitoring) device is positioned directly facing the service personnel and the service recipient at an appropriate distance. The built-in TOF radar sensor continuously scans the service area, and the collected depth data is converted into color space data using a four-segment linear color mapping algorithm and displayed on the device's color LCD screen. (This conversion avoids privacy issues while completely recording the service process, eliminating irrelevant human details and retaining only behavioral characteristics such as service actions, spatial position changes, and contact relationships.) The principle of the four-segment linear color mapping algorithm is as follows: The depth values ​​are divided into four segments: 0-L / 4, L / 4-L / 2, L / 2-3L / 4, and 3L / 4-L (variable values ​​are mapped linearly and proportionally, with L being the farthest depth value collected by the device). These segments are mapped to a gradient color space from red to yellow (red(31,0,0)→yellow(31,63,0)), yellow to green (yellow(31,63,0)→green(0,63,0)), green to cyan (green(0,63,0)→cyan(0,63,31)), and cyan to blue (cyan(0,63,31)→blue(0,0,31)). This approach preserves the spatial structure and movement variations while physically eliminating privacy details, including those related to faces, bodies, and the environment.

[0026] Step S2 also preprocesses the location data, including decision-making for selecting one from multiple source location data: Decision logic: Satellite positioning takes precedence over WiFi-assisted positioning (accuracy of 10 meters); If no satellites can be found, try to obtain WiFi-assisted positioning data (accuracy 50 meters). Choose one of the two options to upload latitude and longitude information to the cloud server; If legitimate Bluetooth beacon information (including Bluetooth beacon MAC address, device name, battery information, and signal strength information) can be found, this information will also be uploaded to the cloud server simultaneously. The server uses latitude and longitude information and Bluetooth beacon signal strength information, and compares and matches them with the pre-stored actual location information of Bluetooth beacons to achieve location confirmation with an accuracy of less than 10 meters. Step S3: Data transmission. The depth data, sound data, and location data are compressed using a balanced lossless compression algorithm and uploaded to the cloud server in real time via a wireless network. Step S4: Cloud-based behavior recognition. The cloud server uses the model to recognize service behaviors and completes a comprehensive assessment of service compliance, ultimately generating a service report.

[0027] More specifically, step S4 is implemented as follows: Step S4.1: Model architecture design. A Spatio-Temporal-Transformer-Network (STTN) based on Transformer is adopted. The STTN is customized for the service behavior recognition task of TOF deep point cloud data (the Transformer architecture has powerful long sequence modeling capabilities and is suitable for processing time-series point cloud data; the PointNet++ structure can effectively extract local features of point clouds. The combination of the two can achieve accurate recognition of service behavior). The overall architecture formula is: ; in: : Input temporal point cloud sequence (T frames, N points, 3D coordinates); Spatial feature extraction module; : Time-series feature aggregation module; : Classification layer weight matrix; Step S4.2: Spatial feature extraction module, using the improved PointNet++ structure, performs local feature aggregation on each frame of point cloud; Point cloud spatial encoding formula: ; in: :point The local neighborhood point set; Multilayer perceptron (MLP); Spatial self-attention mechanism: Standard Transformer attention calculation is used: ; in: Q = F_s × W_Q; K = F_s × W_K; V = F_s × W_V; F_s ∈ R^(N × d): spatial features of the point cloud; Step S4.3: Temporal feature aggregation module, introduces temporal sequence information into each frame of point cloud data by temporal position encoding, enabling the model to perceive the relative positional relationship of each frame in the sequence and capture the inter-frame dependency relationship through temporal self-attention calculation, thereby realizing continuous action modeling; The timing position coding method is as follows: set up: pos: The position index of the current frame in the sequence (0 ≤ pos) <T); i: Feature dimension index; d_model: Feature dimension size; Position encoding is defined as: Even-numbered dimensions: ; Odd-dimensional: ; The specific implementation steps of timing position coding are as follows: (1) For the input time-series feature sequence, obtain the position index pos of each time frame; (2) For each dimension i of the feature vector, calculate the sine or cosine code value according to its parity; (3) Even-numbered dimensions (2i) are encoded using a sine function, and odd-numbered dimensions (2i+1) are encoded using a cosine function; (4) Add the location encoding vector to the original feature vector to obtain the enhanced feature with location information; (5) This encoding method enables the model to learn relative position information, that is, PE_(pos+k) can be expressed as a linear function of PE_pos.

[0028] Temporal self-attention computation: 1. Input features: ; 2. Linear mapping: ; 3. Attention Calculation: ; 4. Bullish Attention: ; 5. Residual Connectivity and Normalization: ; 6. Feedforward network: ; Parameter description: , , Trainable weights; Feature dimension; FFN: Two-layer fully connected network and ReLU; in: This represents the feature matrix after spatial feature extraction; MultiHead represents the multi-head self-attention mechanism, used to capture the dependencies between temporal frames; LayerNorm represents the layer normalization operation, used to stabilize the training process.

[0029] The specific steps for calculating temporal self-attention are as follows: (1) Input features The query matrix Q, the key matrix K, and the value matrix V are obtained through three linear transformations, respectively. (2) Calculate attention weights: ,in The dimension of the key vector; (3) Multi-head attention divides the feature space into h heads, and each head independently calculates attention before splicing; (4) Add the attention output to the original input through residual connections; (5) After layer normalization, we obtain ; (6) Then, a nonlinear transformation is performed through a feedforward network FFN (containing two fully connected layers and an activation function); (7) Final output Used for subsequent action recognition tasks.

[0030] Step S4.4: Construction of Service Behavior Sample Database (Taking the long-term care insurance service standard as an example, a hierarchical service behavior sample database is constructed, covering 8 service categories and 20 service items; the service categories include daily living care (assistance with eating, drinking, and excretion), cleaning care (assistance with bathing, hair washing, and body wiping), postural care (assistance with turning over, moving, and sitting up), rehabilitation care (limb massage, joint mobilization, and rehabilitation training), psychological care (companionship and conversation, and psychological counseling), safety care (fall prevention and pressure ulcer prevention), medical care (blood pressure measurement, blood glucose measurement, and medication guidance), and end-of-life care (palliative care and family support). Each service requires the collection of no less than 5,000 labeled samples, including the definition of key action sequences.) Data augmentation strategies including random rotation, translation, scaling, and Gaussian noise addition are also implemented. Data augmentation strategies included: random rotation, translation, scaling, and Gaussian noise addition. ; 1. Random rotation: : Random rotation matrix (around the x / y / z axes); 2. Translation: ; : Random translation vector; 3. Scaling: ; Scaling factor (0.8~1.2); 4. Gaussian noise: ; Parameter description: P represents the original point cloud data; R represents the rotation matrix, used to simulate different viewpoints; t represents the translation vector; s represents the scaling factor. This indicates that the mean is 0 and the variance is 0. Gaussian noise.

[0031] The specific steps for implementing the data augmentation strategy are as follows: (1) Random rotation: During training, a random rotation matrix R is applied to the point cloud data with a rotation angle range of [-15°, 15°] to enhance the robustness of the model to changes in viewpoint; (2) Random translation: Add a random translation vector t to the point cloud coordinates, with the translation range controlled within [-0.1m, 0.1m], to simulate the sensor position deviation; (3) Random scaling: Apply a scaling factor s with a value range of [0.9, 1.1] to enhance the model's adaptability to distance changes; (4) Gaussian noise addition: Add Gaussian noise to the coordinates of each point. ,in =0.01m, simulating sensor measurement noise; (5) During the training process, one or more of the above enhancement methods are randomly selected and applied in combination to enhance the generalization ability of the model.

[0032] Step S4.5: Feature extraction algorithm, including deep point cloud preprocessing, human pose feature extraction, and interaction relationship feature extraction; Step S4.5 is specifically implemented as follows: Step S4.5.1: Depth point cloud preprocessing, noise filtering uses a statistical outlier removal algorithm: like Then remove that point; in: Indicates the point to be judged; express The j-th nearest neighbor; Indicates the number of nearest neighbors; Point The average distance to its k nearest neighbors; Point The mean distance between all pairs of points in the neighborhood; The standard deviation represents the neighborhood distance; This represents the outlier detection threshold coefficient (recommended value is 1.5); if the condition is met, the point is marked as an outlier and removed; after traversing all points, the filtered clean point cloud data is obtained, thereby removing isolated points caused by sensor noise and environmental interference; The specific steps for data removal are as follows: For each point in the point cloud Calculate the set of distances from a given point to its k nearest neighbors; Calculate the mean of this distance set. and standard deviation ; Judgment conditions Is it greater than ; Step S4.5.2: Human pose feature extraction. Based on TOF depth data, an improved 3D human pose estimation network is used to detect 17 human keypoints (keypoint categories include head (top of head, neck), upper limbs (left and right shoulders, left and right elbows, left and right wrists), trunk (center of thoracic spine, center of lumbar spine), and lower limbs (left and right hips, left and right knees, left and right ankles)) and calculate skeletal vectors. Specifically: Step S4.5.2.1: Extract the key point locations of the human skeleton from the point cloud data, and calculate the skeleton vectors between adjacent joints based on the topological structure of the human skeleton; Skeleton vectors: ; , These represent the three-dimensional coordinate vectors of the i-th and j-th key points of the human body (e.g., joints such as shoulder, elbow, wrist, hip, and knee). From key points Pointing to key points The skeletal vector represents the spatial connection (i.e., the skeleton) between two joints. Step S4.5.2.2: Calculate the bending angle of each joint using vector operations, and the angle between the bone vector and the vertical direction: ; It is a unit vector in the vertical direction (Z-axis) of the spatial coordinate system (such as the vertical upward direction); The magnitude (length) of the bone vector; Step S4.5.2.3: Calculate the motion velocity of each joint using temporal difference. Velocity: ; Let be the three-dimensional coordinates of the i-th key point at the current time t; Let be the three-dimensional coordinates of the i-th key point at the previous time t-1; The time interval between two frames; Step S4.5.2.4: Concatenate the skeleton vector, angle features, and velocity features to form a complete pose feature vector. eigenvectors: ; in: The pose feature vector is represented by the skeleton vector, which is calculated by the position difference between adjacent joints, such as upper arm vector = elbow joint position - shoulder joint position; the angle feature is obtained by calculating the joint angle through the vector dot product, such as elbow joint angle = arccos((upper arm vector · forearm vector) / (|upper arm vector||forearm vector|)); the motion velocity feature is calculated by the joint position difference between adjacent frames.

[0033] Step S4.5.3: Interaction relationship feature extraction, specifically: Step S4.5.3.1: Calculate the interaction distance between human bodies, and measure the degree of interaction by the Euclidean distance between the two central joints of the human body; , The distance between human bodies during interaction; Indicates the location of the central joint of the human body; To represent the L2 norm (Euclidean distance); Step S4.5.3.2: Extract human body (hand features) and object interaction features (including position, shape, and motion state) and calculate the interaction intensity between human body and object through a neural network; ; The interaction strength between the human body and the object (value range is [0,1]); , To represent the learnable weight matrix and bias vector; Features of the hands; Features of an object; Step S4.5.3.3: Determine the contact state based on the distance between the hand and the object. When the distance... Less than the threshold It is considered that contact has occurred ( Recommended value: 0.05m. Step S4.5.3.4: Combine the interaction distance, interaction intensity, and contact state to form an interaction relationship feature vector.

[0034] Step S4.6: Behavioral temporal analysis algorithm, including action segmentation, temporal action modeling, and action sequence similarity calculation; Step S4.6 is specifically implemented as follows: Step S4.6.1: Action segmentation. Temporal action boundary detection uses a dual-stream network. For TOF data, depth frame difference is used instead of optical flow, specifically: Step S4.6.1.1: Calculate the motion difference between adjacent frames. ; ; Indicates the motion difference characteristics at time t; A depth image representing time t; Step S4.6.1.2: Perform smoothing filtering on the motion difference sequence to eliminate the influence of noise; Step S4.6.1.3: Detect local minima of motion differences as candidate points for action boundaries; Step S4.6.1.4: Verify the action boundary by combining the rate of change of human posture, and ensure that the segmentation point is located at the moment of action transition; Step S4.6.1.5: Divide the continuous frame sequence into independent action segments based on the verified boundary points. ; This represents the i-th action segment (composed of continuous time-series data from the k-th frame to the m-th frame); each action segment contains a complete action cycle for subsequent action recognition and matching. Step S4.6.2: Temporal action modeling, each service action is modeled as a Hidden Markov Model (HMM) and decoded using Viterbi; Hidden Markov Model A represents the state transition probability matrix; B represents the observation probability distribution; The initial state distribution is represented; taking assisted feeding as an example, the states are defined as follows: S1: preparation stage; S2: feeding stage; S3: delivery stage; S4: feeding stage; S5: retrieval stage; Using Viterbi decoding: ; in: This represents the maximum probability of being in state i at time t; This represents the probability of transitioning from state j to state i; This indicates that it was observed in state i. The probability of; This represents the observational characteristics at time t.

[0035] The specific steps for Viterbi decoding are as follows: (1) Initialization: ,in The initial state probability; (2) Recursion: For t=2 to T, calculate ; (3) Record the predecessor states of the optimal path; (4) Termination: Find the state with the highest probability at the final moment; (5) Backtracking: Based on the recorded predecessor states, obtain the optimal state sequence and complete the action segmentation.

[0036] Step S4.6.3: Calculate action sequence similarity. Dynamic Time Warping (DTW) is used to calculate action sequence similarity. Specifically: Step S4.6.3.1: Construct the distance matrix D; , Indicates local distance; ; Indicates posture characteristics; Indicates interactive features; This represents the weight coefficient of the interaction feature; a recommended value is 0.3.

[0037] Step S4.6.3.2: Construct the cumulative distance matrix C; ; Step S4.6.3.3: Backtrack from the bottom right corner to the top left corner to find the optimal alignment path; Step S4.6.3.4: Calculate the sum of all local distances on the path to obtain the DTW distance; ; The DTW distance represents the dynamic time warping distance between sequences X and Y; the smaller the DTW distance, the more similar the two action sequences are.

[0038] Step S4.7: Service item matching algorithm, including multi-level matching strategy, feature matching calculation, and service item matching probability calculation; Step S4.7 is specifically implemented as follows: Step S4.7.1: The multi-level matching strategy adopts a three-level matching strategy, specifically as follows: First layer: Coarse-grained category matching (identifies 8 major service categories including daily living care, cleaning care, postural care, and rehabilitation care; the calculation formula is as follows). , This represents the probability distribution of an action belonging to each service category, in vector form, with each component corresponding to the confidence level of a category. This is the weight matrix for the category matching layer, used to map global features to the category space; This is the global feature vector of the action segment, obtained by aggregating the spatiotemporal features of the entire action segment; (This is a bias term for the category matching layer, used to adjust the offset of the category probabilities). The second layer: fine-grained service matching (matching 20 standard service items such as assisted turning over, feeding, hair washing, massage, and blood pressure measurement, calculated using the following formula). ; This represents the probability distribution of an action belonging to each fine-grained service item under the k-th service category. Match a weight matrix to the project corresponding to the k-th type of service; It is a temporal feature sequence of action segments, containing dynamic change information of the action; Match bias terms to the items corresponding to the k-th service category; k is the index of the service category identified in the first layer). The third layer: Movement quality assessment (by comparing with standard movement sequences, calculating movement completeness and execution standardization scores, using the following formula). N represents the number of key nodes (or keyframes) used for comparison in the action sequence. This represents the pose features (such as bone vectors and joint angles) of the actual action performed at the i-th key node. (represents the reference pose feature of the standard service action at the i-th critical node). Step S4.7.2: Feature matching calculation, specifically: Cosine similarity Cosine similarity measures the degree of similarity between two vectors in a direction, with a value range of [-1, 1]. A larger value indicates greater similarity. This metric is insensitive to the magnitude of the feature vectors and is suitable for comparing feature similarity at different scales. ; and These represent the two feature vectors to be compared; Euclidean distance similarity Euclidean distance similarity uses a Gaussian kernel function to map Euclidean distance to the [0,1] interval; the smaller the distance, the higher the similarity. This metric reflects the numerical similarity of feature vectors. ; This represents the bandwidth parameter of the Gaussian kernel function; a recommended value is 1.0. Overall matching score The comprehensive matching score, by weighting and fusing multiple similarity indicators, comprehensively evaluates the similarity of action sequences from different perspectives, thereby improving the accuracy and robustness of matching. ; Indicates temporal similarity (calculated via DTW); , , Let each represent a weight coefficient for one of the three similarities, satisfying the following conditions: .

[0039] Step S4.7.3: Calculate the service item matching probability and obtain the judgment result, which includes confirmed service item, pending manual review, and matching failure. The final matching probability is normalized using Softmax. T represents the temperature parameter (preferably 0.1), and M represents the number of candidate service items. Confidence calculation: ; Greater than or equal to 0.9 and If the value is greater than or equal to 0.15, it is considered a confirmed service item; If the value is greater than or equal to 0.6 and less than 0.9, it is determined to be subject to manual review. If the value is less than 0.6, the match is considered to have failed.

[0040] Step S4.8: Comprehensive service compliance assessment, including multi-dimensional assessment indicators, comprehensive compliance score, and service report generation. Step S4.8 is implemented as follows: Step S4.8.1: Multi-dimensional evaluation indicators include service duration compliance (If the actual duration is within the range of [T_min, T_max], then it is 1; otherwise, it is calculated proportionally.) Action completeness (for (Number of standard actions identified in actual service / Total number of all standard actions required to be completed in this service item) and action quality ; N represents the total number of standard actions included in the service item, i.e. ; This represents the individual quality score for the i-th service action (i.e., in the third-layer matching strategy). Step S4.8.2: The overall compliance score is ; Indicates the probability of a service item matching; , , , The values ​​represent the weights, which are preferably 0.4, 0.25, 0.2, and 0.15, respectively.

[0041] Step S4.8.3: Generate a service process compliance assessment report, including: Basic information: Service personnel ID, service recipient ID, service time, service location; Service item identification results: identified service item, matching probability, and confidence level; Action sequence recording: timestamps of key action nodes, action type, and duration; Compliance score: Overall compliance score and scores in each dimension; Anomaly labeling: low confidence periods, missing actions, and abnormal behaviors; Location trajectory: Records of location changes during the service process.

[0042] Based on the preferred embodiments, the present invention also includes Embodiment 1: Long-term care insurance home care services: Application scenario: Caregivers provide home care services to disabled elderly people. Implementation steps: Hardware Deployment: Before starting work, nursing staff should wear the TOF monitoring device described in this invention on their chest or place it in their service tool bag. The device should be fully charged to ensure a 10-hour battery life.

[0043] Service begins: When the caregiver enters the client's home, the device automatically completes indoor positioning calibration via Bluetooth / WiFi, and the TOF radar starts low-power continuous scanning.

[0044] Data transmission and recognition: The device captures depth images in real time of actions such as turning the elderly over, wiping their body, and measuring their blood pressure. After local anonymization, the data is uploaded to the cloud in real time via a 4G network.

[0045] Matching and Reporting: The cloud-based AI model compares the collected motion features with the long-term care insurance standard care item database. If the system recognizes a 95% probability of a turning-over motion, it confirms that the service has been performed.

[0046] Results Generation: After the service is completed, the system automatically generates a "Service Process Compliance Assessment Report", which includes service duration, key action nodes, and location trajectory, for insurance companies to review and settle fees.

[0047] Emergency Response: If a caregiver encounters danger during the service, pressing the one-click help button will immediately trigger the system to contact the police and family members for location information.

[0048] Based on the preferred embodiment, the present invention also includes an embodiment 2: supervision of the domestic service process. Application scenario: Domestic service personnel provide in-home cleaning, cooking, and other services; Implementation steps: Hardware deployment: Domestic service workers wear TOF monitoring devices, which are used to bind their identities through facial recognition (optional function) or scanning of their work badges.

[0049] Service begins: When the service personnel arrive at the customer's home, the device uses GPS / WiFi to locate and confirm the service location, and the "Start Service" button is pressed to start recording.

[0050] Behavior recognition: The device collects deep data during the service process, and the cloud model recognizes service actions such as cleaning, cooking, and tidying up.

[0051] Service Verification: The system automatically matches and identifies service items based on the service order and generates a service completion report. For example, if the order requires cleaning the living room and kitchen, the system will automatically mark the service as completed after recognizing the relevant actions.

[0052] Dispute resolution: If a customer has any objection to the service quality, they can retrieve the anonymized depth image records and reconstruct the service site for verification.

[0053] Based on the preferred embodiment, the present invention also includes an embodiment 3: Supervision of institutional elderly care services. Application scenario: Caregivers in elderly care facilities provide daily care services to the elderly; Implementation steps: Hardware deployment: Caregivers wear TOF monitoring devices, which achieve indoor positioning via Bluetooth beacons and can distinguish different room areas.

[0054] Service process: Caregivers provide care services to multiple elderly people in different rooms, and the equipment automatically records the service time, location, and service content.

[0055] Real-time monitoring: Institutional managers can view the work status and service progress of caregivers in real time through the monitoring platform.

[0056] Quality assessment: The system automatically assesses the quality of nursing services, such as whether the turning service is performed on time and whether the feeding service is standardized.

[0057] Family members can view the service report remotely via a mobile app to understand the nursing care services provided.

[0058] It is worth mentioning that the technical features such as TOF radar involved in this patent application should be regarded as prior art. The specific structure, working principle and possible control methods and spatial arrangement of these technical features can be adopted using conventional choices in the field, and should not be regarded as the inventive point of this patent. This patent will not be further elaborated in detail.

[0059] For those skilled in the art, modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this invention should be included within the protection scope of this invention.

Claims

1. A service process monitoring system based on TOF radar imaging, characterized in that, It includes a front-end TOF radar monitoring device, a cloud-based AI model layer, a data processing layer, and an application management layer, among which: Front-end TOF radar monitoring device: used to collect depth point cloud data, locally de-identify data, compress and upload data, and provide button, voice, remote control interaction and one-click help function; Cloud-based AI model layer: Employing a Transformer architecture that fuses spatiotemporal features, it constructs a sample library of actions for long-term care insurance, domestic services, and institutional elderly care services, and continuously optimizes recognition accuracy through incremental learning and knowledge distillation; Data processing layer: used for deep data parsing, point cloud denoising, human key point extraction, service behavior recognition, service item matching, and compliance scoring; Application Management Layer: Used to generate service compliance reports, service duration statistics, compliance audits, anomaly alerts, location trajectory playback, remote viewing by family members, and display on the monitoring platform.

2. The service process monitoring system based on TOF radar imaging according to claim 1, characterized in that, The front-end TOF radar monitoring device includes a computing power control module, a sensing module, a human-computer interaction module, a positioning module, a communication transmission module, a power management module, and a local storage module, wherein: The computing power control module adopts an ARM architecture processor and integrates an NPU neural network acceleration unit. The perception module includes a TOF radar sensor and an audio acquisition sensor. The TOF radar sensor is used to emit lasers and receive reflected signals to obtain depth image information of the service scene. The audio acquisition sensor is used to record ambient sound and human voice in real time for voice interaction and environmental sound analysis. The human-computer interaction module includes a color LCD screen, an audio speaker unit, a button unit, an offline voice unit, and a remote control unit, supporting service start / stop, one-click help, and voice control; The positioning module integrates a BeiDou satellite positioning module, a WiFi-assisted positioning module, and a Bluetooth indoor positioning module; The communication transmission module uses 4G as the primary communication method and Bluetooth as the auxiliary communication method, and supports real-time data transmission in environments without WiFi. The local storage module is used for offline data caching and emergency log retention.

3. A service process monitoring method based on TOF radar imaging, implemented in the service process monitoring system based on TOF radar imaging as described in any one of claims 1-2, characterized in that, Includes the following steps: Step S1: Multi-source data acquisition, including sound data acquisition, depth data acquisition, and location data acquisition; Step S2: Data preprocessing and desensitization. Denoising and normalization are performed on the depth data, and the depth data is converted to the color space according to the four-segment linear color mapping algorithm to achieve visualization display without privacy information. Step S3: Data transmission. The depth data, sound data, and location data are compressed using a balanced lossless compression algorithm and uploaded to the cloud server in real time via a wireless network. Step S4: Cloud-based behavior recognition. The cloud server uses the model to recognize service behaviors and completes a comprehensive assessment of service compliance, ultimately generating a service report.

4. The service process monitoring method based on TOF radar imaging according to claim 3, characterized in that, For step S2, when the service personnel begin their service, the device is positioned directly facing them and the service recipient at an appropriate distance. The built-in TOF radar sensor continuously scans the service area, and the collected depth data is converted into color space data using a four-segment linear color mapping algorithm and displayed on the device's configured color LCD screen. The principle of the four-segment linear color mapping algorithm is as follows: The depth values ​​are divided into four segments: 0-L / 4, L / 4-L / 2, L / 2-3L / 4, and 3L / 4-L. These segments are mapped to a gradient color space from red to yellow, yellow to green, green to cyan, and cyan to blue, respectively. While preserving the spatial structure and movement changes, privacy details, including those of faces, bodies, and the environment, are removed from a physical perspective.

5. A service process monitoring method based on TOF radar imaging according to claim 4, characterized in that, Step S4 is specifically implemented as follows: Step S4.1: Model architecture design, adopting a Transformer-based spatiotemporal feature fusion network, and custom design for the service behavior recognition task of TOF deep point cloud data; Step S4.2: Spatial feature extraction module, using the improved PointNet++ structure, performs local feature aggregation on each frame of point cloud; Step S4.3: Temporal feature aggregation module, introduces temporal sequence information into each frame of point cloud data by temporal position encoding, enabling the model to perceive the relative positional relationship of each frame in the sequence and capture the inter-frame dependency relationship through temporal self-attention calculation, thereby realizing continuous action modeling; Step S4.4: Construct a service behavior sample library and perform data augmentation strategies including random rotation, translation, scaling, and Gaussian noise addition; Step S4.5: Feature extraction algorithm, including depth point cloud preprocessing, human pose feature extraction, and interaction relationship feature extraction; Step S4.6: Behavioral temporal analysis algorithm, including action segmentation, temporal action modeling, and action sequence similarity calculation; Step S4.7: Service item matching algorithm, including multi-level matching strategy, feature matching calculation and service item matching probability calculation; Step S4.8: Comprehensive service compliance assessment, including multi-dimensional assessment indicators, comprehensive compliance score, and service report generation.

6. A service process monitoring method based on TOF radar imaging according to claim 5, characterized in that, Step S4.5 is specifically implemented as follows: Step S4.5.1: Depth point cloud preprocessing, noise filtering uses a statistical outlier removal algorithm: like Then remove that point; in: Indicates the point to be judged; express The j-th nearest neighbor; Indicates the number of nearest neighbors; Point The average distance to its k nearest neighbors; Point The mean distance between all pairs of points in the neighborhood; The standard deviation represents the neighborhood distance; This represents the outlier detection threshold coefficient; if the condition is met, the point is marked as an outlier and removed; after traversing all points, filtered clean point cloud data is obtained, thereby removing isolated points caused by sensor noise and environmental interference; Step S4.5.2: Human pose feature extraction. Based on TOF depth data, an improved 3D human pose estimation network is used to detect preset human key points and calculate skeletal vectors. Specifically: Step S4.5.2.1: Extract the key point locations of the human skeleton from the point cloud data, and calculate the skeleton vectors between adjacent joints based on the topological structure of the human skeleton; Skeleton vectors: ; , Let represent the three-dimensional coordinate vectors of the i-th and j-th human body key points, respectively; From key points Pointing to key points The skeletal vector represents the spatial connection between two joints; Step S4.5.2.2: Calculate the bending angle of each joint using vector operations, and the angle between the bone vector and the vertical direction: ; It is a unit vector in the vertical direction in the spatial coordinate system; The magnitude of the skeletal vector; Step S4.5.2.3: Calculate the motion velocity of each joint using temporal difference. Velocity: ; Let be the three-dimensional coordinates of the i-th key point at the current time t; Let be the three-dimensional coordinates of the i-th key point at the previous time t-1; The time interval between two frames; Step S4.5.2.4: Concatenate the skeleton vector, angle features, and velocity features to form a complete pose feature vector. eigenvectors: ; Step S4.5.3: Interaction relationship feature extraction, specifically: Step S4.5.3.1: Calculate the interaction distance between human bodies, and measure the degree of interaction by the Euclidean distance between the two central joints of the human body; , The distance between human bodies during interaction; Indicates the location of the central joint of the human body; To represent the L2 norm; Step S4.5.3.2: Extract the interaction features between the human body and objects and calculate the interaction strength between the human body and objects through a neural network; ; The intensity of interaction between the human body and objects; , To represent the learnable weight matrix and bias vector; Features of the hands; Features of an object; Step S4.5.3.3: Determine the contact state based on the distance between the hand and the object. When the distance... Less than the threshold It is considered that contact has occurred; Step S4.5.3.4: Combine the interaction distance, interaction intensity, and contact state to form an interaction relationship feature vector.

7. A service process monitoring method based on TOF radar imaging according to claim 6, characterized in that, Step S4.6 is specifically implemented as follows: Step S4.6.1: Action segmentation. Temporal action boundary detection uses a dual-stream network. For TOF data, depth frame difference is used instead of optical flow, specifically: Step S4.6.1.1: Calculate the motion difference between adjacent frames. ; Step S4.6.1.2: Perform smoothing filtering on the motion difference sequence to eliminate the influence of noise; Step S4.6.1.3: Detect local minima of motion differences as candidate points for action boundaries; Step S4.6.1.4: Verify the action boundary by combining the rate of change of human posture, and ensure that the segmentation point is located at the moment of action transition; Step S4.6.1.5: Divide the continuous frame sequence into independent action segments based on the verified boundary points; each action segment contains a complete action cycle for subsequent action recognition and matching; Step S4.6.2: Temporal action modeling, each service action is modeled as a Hidden Markov Model and decoded using Viterbi; Step S4.6.3: Calculate action sequence similarity. Dynamic time warping is used to calculate action sequence similarity. Specifically: Step S4.6.3.1: Construct the distance matrix D; Step S4.6.3.2: Construct the cumulative distance matrix C; Step S4.6.3.3: Backtrack from the bottom right corner to the top left corner to find the optimal alignment path; Step S4.6.3.4: Calculate the sum of all local distances on the path to obtain the DTW distance; the smaller the DTW distance, the more similar the two action sequences are.

8. A service process monitoring method based on TOF radar imaging according to claim 7, characterized in that, Step S4.7 is specifically implemented as follows: Step S4.7.1: The multi-level matching strategy adopts a three-level matching strategy, specifically as follows: First layer: coarse-grained category matching; Second layer: Fine-grained item matching; Third level: Motion quality assessment; Step S4.7.2: Feature matching calculation, specifically: Cosine similarity Cosine similarity measures the degree of similarity between two vectors in a direction; the larger the value, the more similar they are. Euclidean distance similarity Euclidean distance similarity uses a Gaussian kernel function to map the Euclidean distance to the [0,1] interval; the smaller the distance, the higher the similarity. Overall matching score The comprehensive matching score, by weighting and fusing multiple similarity indicators, comprehensively evaluates the similarity of action sequences from different perspectives, thereby improving the accuracy and robustness of matching. Step S4.7.3: Calculate the service item matching probability and obtain the judgment result, which includes confirmed service item, pending manual review, and matching failure.

9. A service process monitoring method based on TOF radar imaging according to claim 8, characterized in that, Step S4.8 is specifically implemented as follows: Step S4.8.1: Multi-dimensional evaluation indicators include service duration compliance Completeness of movement and motion quality ; Step S4.8.2: The overall compliance score is ; Indicates the probability of a service item matching; , , , Indicates weight; Step S4.8.3: Generate a service process compliance assessment report, including: Basic information: Service personnel ID, service recipient ID, service time, service location; Service item identification results: identified service item, matching probability, and confidence level; Action sequence recording: timestamps of key action nodes, action type, and duration; Compliance score: Overall compliance score and scores in each dimension; Anomaly labeling: low confidence periods, missing actions, and abnormal behaviors; Location trajectory: Records of location changes during the service process.