Methods, devices, and electronic equipment for monitoring the status of users in a room.
By generating a 3D point cloud image sequence using LiDAR and employing a point cloud data analysis model for user status identification, this approach solves the complexity and security issues of existing room user monitoring methods, achieving highly privacy-protected user behavior and status assessment.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHANGHAI HANZHONG TECH CO LTD
- Filing Date
- 2026-03-30
- Publication Date
- 2026-05-26
AI Technical Summary
Existing technologies for monitoring user conditions in rooms are complex and have poor security, and relying on the fusion analysis of data from cameras and multiple sensors poses a risk of privacy leaks.
A three-dimensional point cloud image sequence is generated by LiDAR scanning. The point cloud data analysis model is used to identify user status, including target detection and posture analysis. User behavior and status are judged through non-visual perception and data analysis models.
Without collecting visual images, it achieves highly privacy-preserving user behavior and status judgment, reducing the complexity of the monitoring system and improving security.
Smart Images

Figure CN122085296A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of automation technology, and more specifically, to a method, apparatus, and electronic device for monitoring the status of users in a room. Background Technology
[0002] In the field of indoor automation technology, to achieve the perception of users' status in a room, the system needs to accurately identify users' postures, activity trajectories, and abnormal behaviors. However, related technologies rely on cameras and millimeter-wave radar for room monitoring. But high-definition video data collected by cameras is prone to privacy leaks, and the process of fusing and analyzing data collected by multiple sensors is quite complex. This leads to the technical problems of high complexity and poor security in the methods of monitoring users' status in rooms.
[0003] There is currently no effective solution to the above problems. Summary of the Invention
[0004] This invention provides a method, apparatus, and electronic device for monitoring the status of users in a room, thereby addressing the technical problems of high complexity and poor security in related methods for monitoring the status of users in a room.
[0005] According to one aspect of the present invention, a method for monitoring the status of users in a room is provided, comprising: scanning a preset space in the room using a lidar to obtain a three-dimensional point cloud image sequence corresponding to the preset space; using a point cloud data analysis model to identify the three-dimensional point cloud image sequence to obtain status information of users existing in the preset space; and executing response measures corresponding to the status information of users.
[0006] Furthermore, using a point cloud data analysis model, the three-dimensional point cloud image sequence is identified to obtain the status information of the user existing in the preset space. This includes: inputting the three-dimensional point cloud image sequence into the point cloud data analysis model, performing target detection on the three-dimensional point cloud image sequence through the point cloud data analysis model to determine whether there is a user in the preset space; and responding to the existence of a user in the preset space, performing status detection on the user through the point cloud data analysis model to obtain the user's status information.
[0007] Furthermore, target detection is performed on the three-dimensional point cloud image sequence using a point cloud data analysis model to determine whether a user exists within a preset space. This includes: determining whether a target object exists within the preset space based on a preset target clustering algorithm and the three-dimensional point cloud image sequence; and determining whether the target object is a user based on a preset human target detection algorithm in response to the presence of a target object within the preset space, thereby determining whether a user exists within the preset space.
[0008] Furthermore, based on a preset target clustering algorithm and a 3D point cloud image sequence, it is determined whether a target object exists within a preset space, including: cropping the region of interest from the 3D point cloud image sequence to obtain the target region of interest; and performing target detection on the target region of interest based on the preset target clustering algorithm to determine whether a target object exists within the preset space.
[0009] Furthermore, the user's condition is detected through a point cloud data analysis model to obtain the user's condition information, including: extracting key points of the user's posture from the 3D point cloud image sequence to obtain user posture key points; and determining the user's condition information based on a preset trajectory tracking algorithm and user posture key points.
[0010] Furthermore, based on a preset trajectory tracking algorithm and user posture key points, the user's status information is determined, including: performing multi-frame temporal fusion of user posture key points in a 3D point cloud image sequence to obtain a posture key point sequence; and tracking the user's behavior trajectory based on the preset trajectory tracking algorithm and posture key point sequence to obtain the user's status information.
[0011] Furthermore, the preset space in the room is scanned by a lidar to obtain a sequence of three-dimensional point cloud images corresponding to the preset space, including: scanning the preset space by a lidar to obtain three-dimensional point cloud data corresponding to the preset space; and performing point cloud rendering on the three-dimensional point cloud data to obtain a sequence of three-dimensional point cloud images.
[0012] According to another aspect of the present invention, a monitoring device for the status of users in a room is also provided, comprising: a scanning module for scanning a preset space in the room using a lidar to obtain a three-dimensional point cloud image sequence corresponding to the preset space; an analysis module for identifying the three-dimensional point cloud image sequence using a point cloud data analysis model to obtain status information of users existing in the preset space; and an execution module for executing response measures corresponding to the status information of users.
[0013] According to another aspect of the present invention, an electronic device is also provided, comprising: a memory storing an executable program; and a processor for running the program, wherein the program, when executed, implements the methods of various embodiments of the present application.
[0014] The aforementioned memory can refer to devices inside a computer used to store data and programs, including RAM, hard disks, etc. RAM can be used to temporarily store running programs and data, while hard disks can be used to store programs and data long-term. Memory enables the computer to read and write data and execute programs. The aforementioned processor is responsible for executing instructions in computer programs and performing data processing. It can also be responsible for controlling and executing various operations, including arithmetic operations, logical operations, and data transmission.
[0015] According to another aspect of the embodiments of this application, a computer-readable storage medium is also provided, the computer-readable storage medium including a stored executable program, wherein, when the executable program is running, it controls the device where the computer-readable storage medium is located to perform the methods of various embodiments of this application.
[0016] The aforementioned computer storage media can refer to the media used in computer memory to store certain discontinuous physical quantities. Computer storage media mainly include semiconductors, magnetic cores, magnetic drums, magnetic tapes, laser discs, etc. Computer-readable storage media include stored programs, which can be a set of instructions that a computer can recognize and execute, running on an electronic computer to meet certain information needs.
[0017] According to another aspect of the embodiments of this application, a computer program product is also provided, including a computer program that, when executed by a processor, implements the methods of various embodiments of this application.
[0018] According to another aspect of the embodiments of this application, a computer program product is also provided, including a computer program that, when executed by a processor, implements the methods of various embodiments of this application.
[0019] According to another aspect of the embodiments of this application, a computer program product is also provided, including a non-volatile computer-readable storage medium storing a computer program that, when executed by a processor, implements the methods in various embodiments of this application.
[0020] The aforementioned non-volatile computer-readable storage medium can refer to a medium for storing data. Non-volatile computer-readable storage media can retain data without loss when power is off and can be used to store long-term data, such as operating systems, applications, and user files. Non-volatile storage media can include hard disk drives, solid-state drives, optical disks, and flash memory storage devices, etc.
[0021] According to another aspect of the embodiments of this application, a computer program is also provided, which, when executed by a processor, implements the methods of the various embodiments of this application.
[0022] The aforementioned computer program can refer to a set of instructions used to tell the computer to perform specific tasks or operations. Computer programs can be written by programmers using specific programming languages and can include algorithms, data structures, logic, and control flow. Computer programs can be used for a variety of purposes, including application software, operating systems, etc.
[0023] In this embodiment of the invention, a preset space within a room is scanned using a LiDAR to obtain a sequence of three-dimensional point cloud images corresponding to the preset space. A point cloud data analysis model is then used to identify the three-dimensional point cloud image sequence to obtain the user's status information within the preset space. Response measures corresponding to the user's status information are then executed. This embodiment employs a non-visual perception and data analysis model, rendering the data acquired by the LiDAR into a sequence of three-dimensional point cloud images, providing a relatively intuitive data foundation for subsequent user status analysis. Subsequently, this embodiment uses a point cloud data analysis model to identify the user's status within the three-dimensional point cloud image sequence, further accurately obtaining response measures corresponding to the user's status information. This achieves the goal of judging user behavior and status without image acquisition, thus realizing high privacy protection and eliminating the need for visual images or multimodal fusion. The analysis of the user's status information within the room can be completed solely based on the three-dimensional point cloud image sequence, thereby solving the technical problems of high complexity and poor security in related technologies for monitoring user status within rooms. Attached Figure Description
[0024] The accompanying drawings, which are included to provide a further understanding of the invention and form part of this application, illustrate exemplary embodiments of the invention and, together with their description, serve to explain the invention and do not constitute an undue limitation thereof. In the drawings:
[0025] Figure 1 This is a flowchart of a method for monitoring the status of users in a room according to an embodiment of the present invention;
[0026] Figure 2 This is a schematic diagram of a lidar scanning imaging according to an embodiment of the present invention;
[0027] Figure 3 This is a schematic diagram of a three-dimensional imaging behavior analysis according to an embodiment of the present invention;
[0028] Figure 4 This is a schematic diagram of an optional method for monitoring the status of users in a room according to an embodiment of the present invention;
[0029] Figure 5 This is a schematic diagram of an optional method for monitoring the status of users in a room according to an embodiment of the present invention;
[0030] Figure 6 This is a schematic diagram of a monitoring device for the status of users in a room according to an embodiment of the present invention. Detailed Implementation
[0031] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.
[0032] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0033] According to an embodiment of the present invention, a method for monitoring the status of users in a room is provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.
[0034] Figure 1 This is a flowchart of a method for monitoring the status of users in a room according to an embodiment of the present invention, such as... Figure 1 As shown, the method includes the following steps:
[0035] Step S102: Scan the preset space in the room using a lidar to obtain a sequence of three-dimensional point cloud images corresponding to the preset space.
[0036] The aforementioned lidar is an active three-dimensional sensing sensor that calculates the spatial distance between the target and the lidar by emitting short-pulse lasers and measuring the time difference between the lidar's reflection from the target object and its return.
[0037] The aforementioned preset space refers to a three-dimensional detection area that is predefined and digitally modeled in a closed or public area.
[0038] The aforementioned 3D point cloud image sequence is a dynamic data stream composed of multiple consecutive time frames of 3D point cloud data arranged in chronological order. Each frame represents a 3D snapshot of a preset space at a certain moment and is associated with a timestamp. The aforementioned 3D point cloud data is a collection of a large number of discrete spatial points acquired by the lidar within a single scanning cycle.
[0039] In one optional embodiment, the lidar performs a rotational scan in a fixed installation posture, simultaneously acquiring echo signals from various angles to generate an original set of points containing spatial three-dimensional coordinates and reflection intensity. Subsequently, a monitoring system for the user's condition within the room unifies the discrete point cloud data to a global reference frame through coordinate transformation and inter-frame registration, and performs density equalization and surface interpolation processing to form a sequence of three-dimensional point cloud images with continuous spatial semantics. The aforementioned coordinate transformation refers to the process of converting the representation of a point or group of points in space from one reference coordinate system to another through a mathematical mapping relationship. This mathematical mapping relationship can be rotation, translation, scaling, or a combination thereof. The aforementioned inter-frame registration refers to the process of aligning the point clouds in a unified space by calculating and estimating the relative pose changes between adjacent frames in a continuously acquired sequence of multiple three-dimensional point cloud images through an algorithm. These relative pose changes can be translation and rotation, etc. The aforementioned density equalization refers to a preprocessing process that spatially resamples and adjusts the uneven point distribution in the original three-dimensional point cloud image sequence, making the point density in the three-dimensional point cloud image sequence tend to be uniformly distributed within a preset area. The aforementioned surface interpolation processing refers to a processing method that, if the surface of an object represented by a three-dimensional point cloud image sequence has holes, missing areas, or insufficient sampling, it estimates and supplements the missing point positions based on the geometric distribution information of surrounding points through a mathematical model in order to reconstruct a continuous and complete surface structure.
[0040] In another optional embodiment, the lidar employs a multi-region scanning mode, sequentially and independently scanning multiple sub-regions within a preset space at high frequency. The 3D point cloud image sequences of each sub-region are accumulated synchronously according to timestamps. The monitoring system for user conditions within the room then uses spatial stitching and semantic normalization to fuse the discontinuously acquired local 3D point cloud image sequences into a unified 3D point cloud image sequence, ensuring coverage integrity and data consistency. The aforementioned spatial stitching refers to the process of merging local 3D point cloud image sequences acquired by multiple lidars at different positions or perspectives into a relatively complete, continuous, and distortion-free global point cloud spatial representation through spatial reference frame alignment and overlapping area matching. The aforementioned semantic normalization refers to mapping the geometric features and motion patterns extracted from the 3D point cloud image sequence to a unified semantic expression space, eliminating representational biases caused by sensor differences, environmental disturbances, or individual body shape variations, ensuring that point cloud data from different users and at different times are comparable and consistent at the semantic level.
[0041] For example, a lidar system scans a preset space at a fixed frequency using multiple rotating angles. Each emitted laser pulse returns after encountering an object's surface. A monitoring system for the users' conditions in the room calculates the three-dimensional coordinates of each reflection point based on the echo time and scanning angle. As the scan continues, multiple scan frames are accumulated and fused sequentially over time, forming a continuously changing set of three-dimensional points—a three-dimensional point cloud image sequence. This sequence fully expresses the geometric shape and relative positional relationships of objects within the space.
[0042] This application embodiment uses LiDAR to scan a preset space and generate a three-dimensional point cloud image sequence, which realizes accurate perception of the spatial position and dynamic contour of indoor personnel, and at the same time provides a high-quality, structured, and time-continuous non-identification data foundation for subsequent behavior analysis.
[0043] Step S104: Using a point cloud data analysis model, the three-dimensional point cloud image sequence is identified to obtain the status information of users existing in the preset space.
[0044] The point cloud data analysis model described above is a dedicated computing model built on artificial intelligence algorithms, specifically designed for processing and understanding the spatial structure and dynamic behavior features in 3D point cloud image sequences.
[0045] The aforementioned users refer to living individuals with biological characteristics who appear within a pre-defined space. In a 3D point cloud, users are represented as target objects with dynamic postures, human body outlines, and irregular motion patterns.
[0046] The aforementioned status information refers to non-identified semantic results inferred from a 3D point cloud image sequence by a point cloud data analysis model, which are related to the user's physiological state or behavioral performance within a preset space. This user status information includes, but is not limited to, whether the user exists, and whether the user is standing, sitting, lying down, fallen, stationary, or abnormally clustered.
[0047] In one optional embodiment, the point cloud data analysis model first performs background filtering and foreground clustering on the input 3D point cloud image sequence to separate potential human targets. Subsequently, the point cloud data analysis model extracts the 3D contour features and motion trends of the targets, infers the relative position changes of key human joints through a pre-trained skeletal structure mapping network, and outputs the user's current dynamic behavioral state by combining temporal smoothing and posture consistency constraints. The aforementioned background filtering refers to the process of identifying and removing static or long-term unchanging environmental structure points from the 3D point cloud image sequence acquired by LiDAR, retaining only dynamic or variable region point clouds. The aforementioned foreground clustering refers to the process of dividing the remaining dynamic 3D point cloud data image sequence after background filtering into several independent target clusters based on spatial proximity, density distribution, or geometric morphological features. The aforementioned pre-trained skeletal structure mapping network refers to a deep learning-based neural network model that uses a large amount of labeled 3D human point cloud data for end-to-end learning during the training phase to establish a non-linear mapping relationship between the original 3D point cloud image sequence and the abstract human skeletal structure (such as joint coordinates and limb connectivity). The aforementioned abstract skeletal structure of the human body can be joint coordinates, limb connectivity, etc. The aforementioned temporal smoothing refers to the process of introducing filtering or improvement mechanisms in the temporal dimension of a sequence of three-dimensional pose keypoints in consecutive frames to suppress instantaneous abrupt changes in joint positions caused by sensor noise, LiDAR sampling jitter, or local occlusion, thereby generating a stable, continuous trajectory output that conforms to the laws of physical motion. The aforementioned pose consistency constraint refers to introducing geometric and kinematic constraints based on prior knowledge of human biomechanics during the three-dimensional pose estimation process, forcing the output pose structure to conform to the anatomical structure and motion laws of the real human body, thereby eliminating unreasonable pose predictions that do not conform to physiological principles.
[0048] In another optional embodiment, the point cloud data analysis model inputs a 3D point cloud sequence into a multi-scale convolutional temporal network, automatically learning the morphological evolution of the point cloud in the spatial dimension and its motion pattern in the temporal dimension. Furthermore, through anomaly clustering detection and a state classification head, it identifies the user's state and outputs highly robust semantic status information. The aforementioned multi-scale convolutional temporal network refers to a deep neural network architecture that integrates spatial local perception and multi-temporal-scale dynamic modeling capabilities, used for spatiotemporal feature extraction and behavioral semantic classification of the user's pose keypoint sequence. The aforementioned anomaly clustering detection refers to the process of automatically identifying outlier clusters deviating from normal behavioral patterns in the high-dimensional feature space of the user's pose keypoint sequence using unsupervised or semi-supervised clustering algorithms, thereby discovering potential abnormal states. The aforementioned state classification head refers to a lightweight classifier module connected to the end of the multi-scale convolutional temporal network, used to map temporal behavioral features to a predefined finite state category space, achieving semantic recognition of user behavior.
[0049] This embodiment uses a point cloud data analysis model to intelligently analyze a 3D point cloud image sequence, enabling the identification of user presence and behavioral characteristics without relying on visual images, thus fundamentally protecting privacy. At the same time, it transforms complex spatial dynamics into understandable semantic information, providing a reliable decision-making basis for subsequent automated responses, achieving dual protection of security and privacy.
[0050] Step S106: Execute the response measures corresponding to the user's status information.
[0051] The aforementioned response measures refer to the preset maintenance or security actions automatically triggered by the room user status monitoring system based on the user status information output by the point cloud data analysis model.
[0052] The aforementioned preset maintenance or security actions refer to a series of standardized response measures pre-configured and automatically triggered by the system based on user status information identified by the point cloud data analysis model. These actions aim to achieve contactless and privacy-free intelligent management. Pre-set maintenance or security actions include, but are not limited to, pushing abnormal status alarms to security personnel's terminals, proactively inquiring about user needs through linked voice prompt devices, notifying cleaning or engineering personnel to intervene in specific areas, automatically adjusting environmental equipment, activating the emergency call system, locking restricted areas, or shutting down potentially dangerous equipment.
[0053] In one optional embodiment, after the point cloud data analysis model outputs the user's status information, the monitoring system for the user's status in the room calls the corresponding handling logic according to the preset status and response mapping rules, triggering a multimodal response mechanism including but not limited to notification, linkage control, or log recording.
[0054] This application embodiment triggers response measures based on the identified user status, realizing timely intervention and service linkage in abnormal states, improving security operation and maintenance efficiency, and enhancing the level of intelligent management within the preset area.
[0055] In this embodiment of the invention, a preset space within a room is scanned using a LiDAR to obtain a sequence of three-dimensional point cloud images corresponding to the preset space. A point cloud data analysis model is then used to identify the three-dimensional point cloud image sequence to obtain the user's status information within the preset space. Response measures corresponding to the user's status information are then executed. This embodiment employs a non-visual perception and data analysis model, rendering the data acquired by the LiDAR into a sequence of three-dimensional point cloud images, providing a relatively intuitive data foundation for subsequent user status analysis. Subsequently, this embodiment uses a point cloud data analysis model to identify the user's status within the three-dimensional point cloud image sequence, further accurately obtaining response measures corresponding to the user's status information. This achieves the goal of judging user behavior and status without image acquisition, thus realizing high privacy protection and eliminating the need for visual images or multimodal fusion. The analysis of the user's status information within the room can be completed solely based on the three-dimensional point cloud image sequence, thereby solving the technical problems of high complexity and poor security in related technologies for monitoring user status within rooms.
[0056] Optionally, a point cloud data analysis model is used to identify the three-dimensional point cloud image sequence to obtain the status information of the user existing in the preset space, including: inputting the three-dimensional point cloud image sequence into the point cloud data analysis model, performing target detection on the three-dimensional point cloud image sequence through the point cloud data analysis model to determine whether there is a user in the preset space; and responding to the existence of a user in the preset space, performing status detection on the user through the point cloud data analysis model to obtain the user's status information.
[0057] The aforementioned target detection refers to the process by which a point cloud data analysis model, based on the geometric distribution and spatial density characteristics of three-dimensional point clouds, identifies and locates whether there are non-environmental static objects with human structural features within a preset space.
[0058] In one optional embodiment, the room user status monitoring system inputs a sequence of 3D point cloud images into a point cloud data analysis model. The model first performs multi-scale feature extraction and spatial density analysis to cluster and structurally group discrete points in the input sequence, identifying point clusters that conform to human geometric features. This allows the system to determine whether a target object conforming to human form exists within a preset space. The multi-scale feature extraction refers to the process of simultaneously capturing local fine-grained geometric features and global macroscopic structural features from the 3D point cloud image sequence clusters generated by LiDAR using parallel or cascaded multi-level neural network structures. The spatial density analysis refers to a method of quantitatively calculating the local distribution density of points in the 3D point cloud to identify the spatial aggregation intensity of key human body regions, and thereby inferring the user's posture, movement state, and potential anomalies. These key human body regions can be the torso, head, limbs, etc.
[0059] In response to the detection of a user's presence, the point cloud data analysis model further estimates the user's pose key points from the point cloud sequence. By extracting candidate joint points and constructing a human skeletal topology, combined with temporal smoothing and motion consistency constraints, the model continuously tracks the spatial displacement relationships of various parts of the user's body, thereby inferring the user's pose changes and dynamic behavior patterns. The point cloud data analysis model ultimately outputs a semantic classification result of the user's state, i.e., the user's status information. The aforementioned user pose changes and dynamic behavior patterns can include being stationary, moving, tilting, or abnormally swaying. The aforementioned pose key point estimation refers to the process of automatically inferring the coordinate positions of human joints in three-dimensional space based on a 3D point cloud sequence acquired by LiDAR, thereby constructing the human skeletal structure. The aforementioned candidate joint points refer to the set of candidate positions corresponding to human joints initially selected from the LiDAR 3D point cloud image sequence through geometric feature analysis and local structure reasoning, serving as prior input for subsequent pose key point estimation. The aforementioned human skeleton topology refers to a rigid or quasi-rigid node-edge connection graph model constructed based on prior knowledge of human anatomy in a three-dimensional point cloud image sequence, used to constrain and guide the spatial relationships of key points of posture.
[0060] This application embodiment first confirms the presence of the user through two-stage point cloud analysis, and then identifies the user's state, achieving accurate perception. Without collecting any visual privacy information, it improves the reliability of behavior recognition and the ability to control the false alarm rate, providing a clear and reliable decision-making basis for subsequent responses.
[0061] Optionally, target detection is performed on the three-dimensional point cloud image sequence using a point cloud data analysis model to determine whether a user exists in a preset space. This includes: determining whether a target object exists in the preset space based on a preset target clustering algorithm and the three-dimensional point cloud image sequence; and determining whether the target object is a user based on a preset human target detection algorithm in response to the presence of a target object in the preset space, thereby determining whether a user exists in the preset space.
[0062] The aforementioned pre-defined target clustering algorithm refers to a computational method that automatically divides a set of neighboring points into independent regions based on the spatial density and geometric continuity of a 3D point cloud. This method is used to initially separate static or dynamic objects in the environment, eliminate background noise and fixed facilities, and provide a set of candidate targets for subsequent human recognition.
[0063] The target objects mentioned above refer to a set of point clouds with finite volume and spatial distribution characteristics identified by clustering algorithms.
[0064] The aforementioned preset human target detection algorithm refers to a model that classifies and judges targets based on the inherent morphological features of the human body. It identifies whether an object is a human body by analyzing the contour distribution and dynamic change patterns of point clouds. The inherent morphological features of the human body can include height proportions, limb structure, and movement trajectory.
[0065] In one optional embodiment, the room user status monitoring system performs spatial partitioning processing on a 3D point cloud image sequence based on a preset target clustering algorithm. It then groups discrete points into several 3D point clusters using a density-aware clustering method, eliminating background noise and static object interference, retaining only candidate target objects with dynamic or irregular geometric structures. This clustering process does not rely on prior shape models; it automatically divides target areas based solely on the statistical characteristics of point spacing and local point density, thus filtering potential moving objects in the environment. The density-aware clustering method refers to a local density estimation and adaptive clustering mechanism based on the spatial distribution of 3D point clouds, used to automatically detect abnormal patterns of human gathering behavior in a preset area without relying on visual semantics or identifying individual identities.
[0066] In response to the detection of candidate targets, the room user status monitoring system employs a pre-defined human target detection algorithm to analyze the morphological features of each target. By extracting geometric indicators, a human-specific discriminant function is constructed. Combined with multi-dimensional feature thresholds and topological consistency verification, the system determines whether the target conforms to human structural characteristics. These geometric indicators can include the height distribution of point clusters, base area, symmetry, and limb extension trends. The human-specific discriminant function is a multi-dimensional discriminant model constructed based on the geometric morphological features, motion dynamics, and prior structural constraints of a 3D point cloud image sequence. The multi-dimensional feature thresholds and topological consistency verification constitute a dual verification mechanism used in a room user status monitoring system driven by LiDAR 3D point cloud image sequences to ultimately confirm targets initially identified as human, ensuring they conform to the geometric constraints of human physiological structure and the dynamic rationality of behavioral semantics. This eliminates the possibility of non-human targets being misidentified as human, as well as erroneous behavior recognition caused by noise interference with human posture.
[0067] This application embodiment uses a clustering screening and human feature discrimination hierarchical detection mechanism to gradually filter out non-human interference without relying on visual images, confirming the presence of users, effectively reducing the false detection rate, and improving the robustness and privacy security of the monitoring system for users in the room in complex environments.
[0068] Optionally, based on a preset target clustering algorithm and a 3D point cloud image sequence, determining whether a target object exists within a preset space includes: cropping the region of interest from the 3D point cloud image sequence to obtain the target region of interest; and performing target detection on the target region of interest based on the preset target clustering algorithm to determine whether a target object exists within the preset space.
[0069] The aforementioned region of interest (ROI) clipping refers to extracting a pre-defined activity area from the original 3D point cloud image sequence based on the physical layout and monitoring needs of a pre-defined room scene. Irrelevant point cloud data is then removed to reduce computational redundancy, improve processing efficiency, and focus on the spatial range of user activity. The pre-defined activity area can be a corridor, gym, lobby, etc. Irrelevant areas can refer to walls, ceilings, fixed furniture, etc.
[0070] The aforementioned target region of interest refers to the 3D point cloud subset that is retained after cropping and contains only the potential human activity space. This region has clear spatial boundaries and semantic meaning and serves as the input for subsequent target clustering algorithms.
[0071] In one optional embodiment, the room user status monitoring system performs region of interest cropping on the three-dimensional point cloud image sequence. Based on preset spatial constraints, environmental areas unrelated to the monitoring target are eliminated, and only the dynamic spatial range of user activities is retained, forming a highly focused target region of interest.
[0072] The in-room user status monitoring system uses a pre-defined target clustering algorithm to group point cloud data within the target's region of interest. Through local density analysis and spatial connectivity assessment, it identifies clusters of points exhibiting aggregation, non-planarity, and motion correlation as candidate sets of potential targets. The aforementioned local density analysis and spatial connectivity assessment first calculate the point density distribution of each point in its local neighborhood to identify high-density clustered areas, and then constructs a connectivity graph based on the spatial proximity relationships between points.
[0073] This application's embodiments focus on the effective monitoring range through region-of-interest clipping and extract potential targets using a target clustering algorithm. This reduces redundant computation and environmental interference, improves the response speed and accuracy of target detection, and achieves efficient and stable user presence prediction while protecting privacy.
[0074] Optionally, the user's condition is detected by a point cloud data analysis model to obtain the user's condition information, including: extracting key points of the user's posture in the three-dimensional point cloud image sequence to obtain user posture key points; and determining the user's condition information based on a preset trajectory tracking algorithm and user posture key points.
[0075] The user's posture mentioned above refers to the relative position and angle relationship of various parts of the user's body in three-dimensional space. It is a skeletal structure representation restored by point cloud data, which does not contain identifiable biometric features and only reflects non-privacy movement states such as limb orientation and movement pattern.
[0076] The aforementioned key point extraction refers to automatically identifying and locating geometric feature points such as shoulders, elbows, hips, and knees that are representative of movement from a sequence of three-dimensional point cloud images, based on a prior model of human body structure. The user's limb posture is quantified through a spatial coordinate sequence to achieve anonymized action expression without visual images.
[0077] The aforementioned user posture key points refer to a set of three-dimensional coordinate sequences that change over time after key point extraction. These key points are used to characterize the dynamic features of user actions. Each key point only records spatial location information and is not associated with identity, appearance, or behavioral intent.
[0078] The aforementioned preset trajectory tracking algorithm refers to a temporal analysis method that identifies user behavior states based on the spatiotemporal continuity of user posture key points, through motion pattern modeling and trajectory prediction. It does not rely on external identifiers and determines the state solely based on the evolutionary patterns of 3D point cloud image sequences. These user behavior states can include standing, falling, or remaining stationary.
[0079] In one optional embodiment, the in-room user status monitoring system performs temporal continuous analysis on the user point cloud in the 3D point cloud image sequence. Through skeleton fitting and joint candidate point screening, it extracts the 3D spatial position sequence representing the main motion nodes of the human body, forming a set of key points for user posture. The skeleton fitting refers to the computational process of geometrically aligning and parameterizing the detected human target point cloud data with a preset human biomechanical skeleton model in the 3D point cloud image sequence acquired by LiDAR, in order to reconstruct a virtual 3D skeleton that approximates the real human body structure. The joint candidate point screening refers to a preprocessing mechanism that automatically identifies the candidate point set corresponding to the human joints from the original 3D point cloud image sequence before completing skeleton fitting.
[0080] The in-room user status monitoring system, based on a preset trajectory tracking algorithm, performs spatiotemporal correlation and motion path modeling on a multi-frame continuous sequence of posture key points. It estimates the motion trajectory of each key point through Kalman filtering and, combined with the posture change rate, joint angle evolution trend, and overall center of gravity shift pattern, infers the user's current dynamic state, thereby determining the user's status information. The aforementioned spatiotemporal correlation and motion path modeling refers to a privacy-aware technology that uses a continuous frame sequence of 3D point cloud images from a LiDAR system to perform structured modeling and semantic reasoning of the target human body's motion trajectory through spatiotemporal joint modeling. The aforementioned Kalman filtering refers to an improved dynamic system algorithm based on a state-space model and Bayesian recursive estimation theory. In this invention, it is used to suppress noise, smooth the trajectory, and predict future states from the human joint coordinate sequence extracted from continuous LiDAR frames without visual images, thereby constructing a stable, continuous, and physically consistent motion trajectory.
[0081] This application embodiment achieves continuous and anonymized analysis of user behavior states through user posture key point extraction and preset trajectory tracking algorithms, thereby improving the real-time performance and reliability of intelligent response.
[0082] Optionally, based on a preset trajectory tracking algorithm and user posture key points, the user's status information is determined, including: performing multi-frame temporal fusion of user posture key points in a 3D point cloud image sequence to obtain a posture key point sequence; and performing behavioral trajectory tracking on the user based on the preset trajectory tracking algorithm and posture key point sequence to obtain the user's status information.
[0083] The aforementioned multi-frame temporal fusion refers to spatially aligning and correcting motion consistency of user pose key points in a sequence of consecutive multi-frame 3D point cloud images according to the time sequence, eliminating single-frame noise and jitter, enhancing the continuity and stability of pose information, and providing high-quality temporal input for subsequent behavior analysis.
[0084] The aforementioned posture key point sequence refers to the set of three-dimensional coordinates of user body key points arranged in chronological order, generated by multi-frame temporal fusion, representing the user's motion trajectory and posture evolution process over a period of time, and containing only structured spatial data.
[0085] In one optional embodiment, the in-room user status monitoring system performs multi-frame temporal fusion of user posture key points in a 3D point cloud image sequence, aligns and interpolates the key point positions of consecutive frames through a time sliding window, eliminates single-frame noise and sampling jitter, and constructs a posture key point sequence with continuity and smoothness.
[0086] The in-room user status monitoring system dynamically models the fused posture key point sequence based on a preset trajectory tracking algorithm, analyzes the relative displacement relationship, motion velocity distribution and trajectory curvature change between posture key point sequences, identifies the temporal characteristics of typical behavior patterns, and thus obtains user status information.
[0087] This application embodiment improves the stability and continuity of attitude key points through multi-frame temporal fusion, and combines trajectory tracking algorithm to realize dynamic analysis of user behavior, effectively identify abnormal states, and improve detection accuracy and real-time response.
[0088] Optionally, a preset space within the room is scanned using a lidar to obtain a sequence of three-dimensional point cloud images corresponding to the preset space, including: scanning the preset space using a lidar to obtain three-dimensional point cloud data corresponding to the preset space; and performing point cloud rendering on the three-dimensional point cloud data to obtain a sequence of three-dimensional point cloud images.
[0089] The aforementioned three-dimensional point cloud data is a set of original discrete data consisting of a large number of three-dimensional spatial coordinate points, obtained by lidar through emitting laser pulses and receiving echoes based on the Time of Flight (ToF) principle. Each point contains position and reflection intensity information, truly reflecting the geometric structure and object outline of the measured space.
[0090] In one optional embodiment, the in-room user status monitoring system scans a preset space using lidar, emitting directional laser pulses and receiving reflected signals. Based on the time-of-flight principle, it calculates the spatial coordinates of each sampling point, constructing an original point cloud dataset composed of a large number of discrete 3D points. The in-room user status monitoring system then renders the original 3D point cloud data, transforming the sparse point cloud into a sequence of 3D point cloud images with visual continuity and structural consistency through preprocessing steps such as gridded resampling, normal vector estimation, and density normalization.
[0091] This application embodiment acquires raw 3D point cloud data using LiDAR and performs point cloud rendering, transforming sparse point clouds into a clearly structured visual image sequence. This preserves the complete geometric structure while avoiding the leakage of facial and body features, thus achieving environmental perception without compromising privacy.
[0092] Figure 2 This is a schematic diagram of a lidar scanning imaging according to an embodiment of the present invention, such as... Figure 2 As shown, a laser pulse is emitted and a beam is scanned in a directional manner. After reflecting off the target, the monitoring system for the user's condition inside the room receives the reflected laser signal, calculates the three-dimensional coordinates, and generates three-dimensional point cloud data. The monitoring system then performs post-processing on the raw three-dimensional point cloud data. Finally, based on the processed point cloud data, three-dimensional imaging and perception output are achieved. The monitoring system constructs the spatial morphology and motion characteristics of the target within the room, providing a reliable data foundation for subsequent behavior recognition and status assessment.
[0093] Figure 3 This is a schematic diagram of a three-dimensional imaging behavior analysis according to an embodiment of the present invention, such as... Figure 3 As shown, after the LiDAR 3D point cloud data is input, it first undergoes point cloud data preprocessing. Then, the room user status monitoring system performs target clustering and user target detection based on the preprocessed point cloud data. Next, for the detected users, the system extracts key points of their poses. Then, the system combines multi-frame temporal fusion with trajectory tracking. Finally, the system outputs the corresponding response measures based on the results.
[0094] Figure 4This is a schematic diagram of an optional method for monitoring user conditions in a room according to an embodiment of the present invention, such as... Figure 4 As shown, a preset space within a room is scanned using a LiDAR scanner to obtain a sequence of 3D point cloud images corresponding to that space. This sequence is then input into a point cloud data analysis model, which performs target detection to determine the presence of a user within the preset space. In response to the presence of a user, the point cloud data analysis model detects the user's condition, obtains their status information, and executes corresponding response measures.
[0095] Figure 5 This is a schematic diagram of an optional method for monitoring user conditions in a room according to an embodiment of the present invention, such as... Figure 5 As shown, a preset space within a room is scanned using a LiDAR scanner to obtain a sequence of 3D point cloud images corresponding to that space. Key points of the user's posture are extracted from the 3D point cloud image sequence to obtain user posture key points. Based on a preset trajectory tracking algorithm and the user posture key points, the user's status information is determined, and response measures corresponding to that status information are executed.
[0096] Figure 6 This is a schematic diagram of a room user status monitoring device according to an embodiment of the present invention, such as... Figure 6 As shown, the device includes: a scanning module 602, an analysis module 604, and an execution module 606.
[0097] The scanning module 602 is also used to scan the preset space in the room using a lidar to obtain a three-dimensional point cloud image sequence corresponding to the preset space; the analysis module 604 is used to use a point cloud data analysis model to identify the three-dimensional point cloud image sequence and obtain the status information of the user existing in the preset space; the execution module 606 executes response measures corresponding to the user's status information.
[0098] The analysis module is also used to input the three-dimensional point cloud image sequence into the point cloud data analysis model, and to perform target detection on the three-dimensional point cloud image sequence through the point cloud data analysis model to determine whether there is a user in the preset space; in response to the presence of a user in the preset space, the point cloud data analysis model is used to perform status detection on the user to obtain the user's status information.
[0099] The analysis module is also used to determine whether there is a target object in a preset space based on a preset target clustering algorithm and a 3D point cloud image sequence; in response to the presence of a target object in the preset space, it determines whether the target object is a user based on a preset human target detection algorithm, thereby determining whether there is a user in the preset space.
[0100] The analysis module is also used to crop the region of interest in the 3D point cloud image sequence to obtain the target region of interest; based on the preset target clustering algorithm, it performs target detection on the target region of interest to determine whether there is a target object in the preset space.
[0101] The analysis module is also used to extract key points of the user's posture in the 3D point cloud image sequence to obtain the user posture key points; based on the preset trajectory tracking algorithm and the user posture key points, the user's status information is determined.
[0102] The analysis module is also used to perform multi-frame temporal fusion of user pose key points in the 3D point cloud image sequence to obtain a pose key point sequence; based on the preset trajectory tracking algorithm and the pose key point sequence, the user's behavior trajectory is tracked to obtain the user's status information.
[0103] The device is also used to scan a preset space using a lidar to obtain three-dimensional point cloud data corresponding to the preset space; and to perform point cloud rendering on the three-dimensional point cloud data to obtain a three-dimensional point cloud image sequence.
[0104] According to another aspect of the embodiments of this application, a method for monitoring the status of users in a room is also provided, the method comprising the following steps:
[0105] First, the preset space in the room is scanned using a lidar to obtain a sequence of three-dimensional point cloud images corresponding to the preset space.
[0106] Next, the three-dimensional point cloud image sequence is input into the point cloud data analysis model. The point cloud data analysis model performs target detection on the three-dimensional point cloud image sequence to determine whether there is a user in the preset space. In response to the presence of a user in the preset space, the point cloud data analysis model performs status detection on the user to obtain the user's status information.
[0107] Next, based on a preset target clustering algorithm and a 3D point cloud image sequence, it is determined whether there is a target object in the preset space; in response to the presence of a target object in the preset space, based on a preset human target detection algorithm, it is determined whether the target object is a user, so as to determine whether there is a user in the preset space.
[0108] Next, the region of interest is cropped from the 3D point cloud image sequence to obtain the target region of interest; based on a preset target clustering algorithm, target detection is performed on the target region of interest to determine whether there is a target object in the preset space.
[0109] Finally, implement response measures corresponding to the user's status information.
[0110] According to another aspect of the embodiments of this application, a method for monitoring the status of users in a room is also provided, the method comprising the following steps:
[0111] First, the preset space in the room is scanned using a lidar to obtain a sequence of three-dimensional point cloud images corresponding to the preset space.
[0112] Next, the three-dimensional point cloud image sequence is input into the point cloud data analysis model. The point cloud data analysis model performs target detection on the three-dimensional point cloud image sequence to determine whether there is a user in the preset space. In response to the presence of a user in the preset space, the point cloud data analysis model performs status detection on the user to obtain the user's status information.
[0113] Next, key points of the user's pose are extracted from the 3D point cloud image sequence to obtain the user pose key points.
[0114] Next, multi-frame temporal fusion is performed on the user pose key points in the 3D point cloud image sequence to obtain the pose key point sequence; based on the preset trajectory tracking algorithm and the pose key point sequence, the user's behavior trajectory is tracked to obtain the user's status information.
[0115] Finally, implement response measures corresponding to the user's status information.
[0116] Embodiments of this application also provide an electronic device, including: a memory storing an executable program; and a processor for running the program, wherein the program executes the methods in various embodiments of this application when it runs.
[0117] Embodiments of this application also provide a computer-readable storage medium including a stored executable program, wherein, when the executable program is running, it controls the device where the computer-readable storage medium is located to perform the methods of various embodiments of this application.
[0118] The aforementioned computer storage media can refer to the media used in computer memory to store certain discontinuous physical quantities. Computer storage media mainly include semiconductors, magnetic cores, magnetic drums, magnetic tapes, laser discs, etc. Computer-readable storage media include stored programs, which can be a set of instructions that a computer can recognize and execute, running on an electronic computer to meet certain information needs.
[0119] Embodiments of this application also provide a computer program product, including a computer program that, when executed by a processor, implements the methods of various embodiments of this application.
[0120] The aforementioned computer program products can refer to software programs that have been written, tested, and released, and can run on computers or other devices. Computer program products can include application programs, operating systems, utility software, etc., used to achieve specific functions or solve specific problems.
[0121] Embodiments of this application also provide a computer program product, including a non-volatile computer-readable storage medium for storing a computer program that, when executed by a processor, implements the methods in various embodiments of this application.
[0122] The aforementioned non-volatile computer-readable storage medium can refer to a medium for storing data. Non-volatile computer-readable storage media can retain data without loss when power is off and can be used to store long-term data, such as operating systems, applications, and user files. Non-volatile storage media can include hard disk drives, solid-state drives, optical disks, and flash memory storage devices, etc.
[0123] Embodiments of this application also provide a computer program that, when executed by a processor, implements the methods described in the various embodiments of this application.
[0124] The aforementioned computer program can refer to a set of instructions used to tell the computer to perform specific tasks or operations. Computer programs can be written by programmers using specific programming languages and can include algorithms, data structures, logic, and control flow. Computer programs can be used for a variety of purposes, including application software, operating systems, etc.
[0125] In the above embodiments of the present invention, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0126] The sequence numbers of the above embodiments of the present invention are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.
[0127] In the above embodiments of the present invention, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0128] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. The device embodiments described above are merely illustrative; for example, the division of units can be a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the displayed or discussed mutual couplings, direct couplings, or communication connections may be through some interfaces; indirect couplings or communication connections between units or modules may be electrical or other forms.
[0129] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0130] Furthermore, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0131] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, read-only memory (ROM), random access memory (RAM), portable hard drives, magnetic disks, or optical disks.
[0132] The above are merely preferred embodiments of the present invention. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principle of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.
Claims
1. A method for monitoring the status of users in a room, characterized in that, include: A three-dimensional point cloud image sequence corresponding to the preset space is obtained by scanning the preset space with a lidar. Using a point cloud data analysis model, the three-dimensional point cloud image sequence is identified to obtain the status information of the users existing in the preset space; Execute response measures corresponding to the user's status information.
2. The method for monitoring user status in a room according to claim 1, characterized in that, Using a point cloud data analysis model, the three-dimensional point cloud image sequence is identified to obtain user status information within the preset space, including: The three-dimensional point cloud image sequence is input into the point cloud data analysis model, and the point cloud data analysis model is used to perform target detection on the three-dimensional point cloud image sequence to determine whether there is a user in the preset space. In response to the presence of a user within the preset space, the user's condition is detected using the point cloud data analysis model to obtain the user's condition information.
3. The method for monitoring the status of users in a room according to claim 2, characterized in that, The point cloud data analysis model is used to perform target detection on the 3D point cloud image sequence to determine whether a user exists within the preset space, including: Based on a preset target clustering algorithm and the three-dimensional point cloud image sequence, it is determined whether there is a target object in the preset space; In response to the presence of a target object within the preset space, a preset human target detection algorithm is used to determine whether the target object is a user, thereby determining whether a user exists within the preset space.
4. The method for monitoring the status of users in a room according to claim 3, characterized in that, Based on a preset target clustering algorithm and the 3D point cloud image sequence, determine whether a target object exists within the preset space, including: The region of interest is cropped from the three-dimensional point cloud image sequence to obtain the target region of interest; Based on the preset target clustering algorithm, target detection is performed on the target region of interest to determine whether a target object exists within the preset space.
5. The method for monitoring the status of users in a room according to claim 2, characterized in that, The user's condition is detected using the point cloud data analysis model to obtain the user's condition information, including: Key points are extracted from the user's pose in the three-dimensional point cloud image sequence to obtain user pose key points; Based on a preset trajectory tracking algorithm and the user's posture key points, the user's status information is determined.
6. The method for monitoring user status in a room according to claim 5, characterized in that, Based on a preset trajectory tracking algorithm and the user's posture key points, the user's status information is determined, including: Multi-frame temporal fusion is performed on the user pose key points in the three-dimensional point cloud image sequence to obtain a pose key point sequence. Based on the preset trajectory tracking algorithm and the posture key point sequence, the user's behavior trajectory is tracked to obtain the user's status information.
7. The method for monitoring the status of users in a room according to any one of claims 1 to 6, characterized in that, A preset space within the room is scanned using a lidar scanner to obtain a sequence of three-dimensional point cloud images corresponding to the preset space, including: The preset space is scanned by the lidar to obtain the three-dimensional point cloud data corresponding to the preset space; The three-dimensional point cloud data is rendered to obtain the three-dimensional point cloud image sequence.
8. A device for monitoring the status of users in a room, characterized in that, include: The scanning module is used to scan a preset space in the room using a lidar to obtain a three-dimensional point cloud image sequence corresponding to the preset space. The analysis module is used to identify the three-dimensional point cloud image sequence using a point cloud data analysis model to obtain the status information of the users existing in the preset space. The execution module is used to execute response measures corresponding to the user's status information.
9. An electronic device, characterized in that, include: Memory, which stores executable programs; A processor for running the program, wherein the program, when running, performs the method for monitoring the status of users in a room as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored executable program, wherein, when the executable program is executed, it controls the device on which the storage medium is located to perform the room monitoring method according to any one of claims 1 to 7.