An early warning method and device based on an abstract human body model

By using an early warning method based on an abstract human body model, an abstract human body model is constructed using cameras and distance sensors, mapped to a digital twin space, and monitored and generated to solve the problem of privacy infringement on the elderly, thus achieving efficient monitoring and privacy protection.

CN117334008BActive Publication Date: 2026-07-31XINDA PROPERTY MANAGEMENT (BEIJING) TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
XINDA PROPERTY MANAGEMENT (BEIJING) TECH CO LTD
Filing Date
2023-10-23
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

In existing technologies, webcams installed in the private spaces of the elderly infringe on privacy, reduce user experience, and fail to provide effective monitoring without infringing on privacy.

Method used

An early warning method based on an abstract human body model is used to collect image data by a camera, combine it with a distance sensor to determine depth information, construct an abstract human body model, map it to a digital twin space, monitor motion status, and generate early warnings when anomalies occur.

Benefits of technology

It enables efficient monitoring of elderly people's activities and postures without infringing on their privacy, improving user experience and ensuring safety and health.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117334008B_ABST
    Figure CN117334008B_ABST
Patent Text Reader

Abstract

This application discloses an early warning method and apparatus based on an abstract human body model. The method includes: determining image data collected by a target camera over a monitored area; determining depth information corresponding to the image data using a distance sensor; determining multiple skeletal node information of the target user in the image data based on a preset reference model; determining an abstract human body model corresponding to the target user based on the coordinates of each skeletal node and the reference model; mapping the abstract human body model to a digital twin space corresponding to the monitored area through depth information and a preset position transformation relationship, and determining the motion state of the abstract human body model in the digital twin space; generating early warning information when the motion state meets preset abnormal conditions. This application enables the monitoring of elderly people's activities and postures using an abstract human body model, ensuring both efficiency and accuracy of monitoring while protecting privacy without infringing on the privacy of the elderly, thus improving the user experience.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of monitoring and early warning, and in particular to an early warning method and device based on an abstract human body model. Background Technology

[0002] With societal progress and improved living standards, average life expectancy has significantly increased. However, with age, the probability of sudden illnesses also rises markedly, which is particularly concerning for elderly people living alone. Because they live alone without the daily companionship of relatives or caregivers, the installation of webcams is currently a common way to better monitor their lives.

[0003] However, privacy is very important to everyone, especially for the elderly. As people age, many become more sensitive and anxious, particularly about webcams installed in their private homes, which may make them feel their privacy is being violated and reduce their user experience.

[0004] Application content

[0005] This application provides an early warning method and device based on an abstract human body model to overcome the above-mentioned problems or at least partially solve the problem mentioned above where users feel their privacy is violated due to network cameras installed in private spaces, thereby reducing the user experience.

[0006] Firstly, this application provides an early warning method based on an abstract human body model, including:

[0007] Determine the image data collected by the target camera on the monitored area; and use a distance sensor to determine the depth information corresponding to the image data.

[0008] Based on a pre-defined reference model, multiple skeletal node information of the target user is determined in the image data;

[0009] Based on the information of each skeletal node and the reference model, determine the abstract human body model corresponding to the target user;

[0010] By using depth information and preset position transformation relationships, the abstract human body model is mapped to the digital twin space corresponding to the monitoring area, and the motion state of the abstract human body model in the digital twin space is determined.

[0011] When the motion state meets the preset abnormal conditions, an early warning message is generated.

[0012] Secondly, this application provides an early warning device based on an abstract human body model, comprising:

[0013] The depth information determination module is used to determine the image data collected by the target camera for the monitored area; and uses a distance sensor to determine the depth information corresponding to the image data.

[0014] The skeletal node information determination module is used to determine multiple skeletal node information of the target user in image data based on a preset reference model.

[0015] The abstract human body model determination module is used to determine the abstract human body model corresponding to the target user based on the coordinates of each skeletal node and the reference model.

[0016] The motion state determination module is used to map the abstract human body model to the digital twin space corresponding to the monitoring area through depth information and preset position transformation relationship, and to determine the motion state of the abstract human body model in the digital twin space.

[0017] The early warning information generation module is used to generate early warning information when the motion state meets preset abnormal conditions.

[0018] Thirdly, this application provides a readable medium including executable instructions, which, when executed by a processor of an electronic device, cause the electronic device to perform any of the methods described in the first aspect.

[0019] Fourthly, this application provides an electronic device including a processor and a memory storing execution instructions, wherein when the processor executes the execution instructions stored in the memory, the processor performs the method as described in any of the first aspects.

[0020] The beneficial effects of this application embodiment compared with the prior art are as follows: This application provides an early warning method based on an abstract human body model. It determines the image data collected by the target camera on the monitored area; uses a distance sensor to determine the depth information corresponding to the image data; determines multiple skeletal node information of the target user in the image data based on a preset reference model; determines the abstract human body model corresponding to the target user based on the coordinates of each skeletal node and the reference model; maps the abstract human body model to the digital twin space corresponding to the monitored area through the depth information and a preset position transformation relationship, and determines the motion state of the abstract human body model in the digital twin space; when the motion state meets preset abnormal conditions, an early warning message is generated. This application embodiment realizes the monitoring of the activities and postures of the elderly using an abstract human body model, ensuring the efficiency and accuracy of monitoring while protecting privacy without infringing on the privacy of the elderly, thus improving the user experience. Attached Figure Description

[0021] To more clearly illustrate the embodiments of this application or the existing technical solutions, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0022] Figure 1 A flowchart illustrating an early warning method based on an abstract human body model, provided as an embodiment of this application;

[0023] Figure 2 This is a schematic diagram of the structure of an early warning device based on an abstract human body model provided in an embodiment of this application;

[0024] Figure 3 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0025] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions of this application will be clearly and completely described below in conjunction with specific embodiments and corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0026] Figure 1 This is a flowchart illustrating an early warning method based on an abstract human body model provided in an embodiment of this application. Figure 1 As shown, this early warning method based on an abstract human body model includes:

[0027] S101, determine the image data collected by the target camera for the monitored area; and use the distance sensor to determine the depth information corresponding to the image data;

[0028] S102, based on a preset reference model, determine multiple skeletal node information of the target user in the image data;

[0029] S103, Based on the information of each skeletal node and the reference model, determine the abstract human body model corresponding to the target user;

[0030] S104, through depth information and preset position transformation relationship, maps the abstract human body model to the digital twin space corresponding to the monitoring area, and determines the motion state of the abstract human body model in the digital twin space;

[0031] S105, when the motion state meets the preset abnormal conditions, generate early warning information.

[0032] Specifically, traditional methods of monitoring the elderly via webcams typically involve installing cameras in the elderly person's living area to capture their daily activities. This method provides an effective means of monitoring, allowing family members, caregivers, or monitoring systems to view video streams in real time or subsequently to ensure the safety and well-being of the elderly. However, this method comes with some inconveniences and problems.

[0033] However, older adults may feel that this form of surveillance infringes on their privacy. Cameras capturing their every activity, including private moments, can cause discomfort and resistance. This is especially important for older adults or those with health problems, as they value a degree of freedom and privacy at home.

[0034] For example, an elderly person may occasionally need to use the bathroom, change clothes, or perform personal hygiene activities—all very private activities. If traditional surveillance cameras are installed in the home, the elderly person may feel their privacy is being violated because these cameras could record their activities in the bathroom or while changing clothes, places they wish to keep private.

[0035] Furthermore, the purpose of monitoring the elderly is usually not to keep track of their every move in real time, but primarily to provide early warnings and assistance in case of abnormal activity or emergencies. The key objective of such early warning systems is to ensure the safety and health of the elderly, not to infringe on their privacy or monitor every detail of their lives.

[0036] Therefore, considering the privacy of the elderly and respecting their personal space, a balance between monitoring needs and privacy protection can be achieved by abstracting the users to be monitored into a human body model. This method not only enables the monitoring of abnormal activities and provides early warnings, but also significantly reduces the risk of privacy violations. In this embodiment, an early warning method based on an abstract human body model includes:

[0037] Furthermore, a monitoring area is a specific geographical area or spatial range used to monitor and detect the activities of older adults. This area is typically located within the older adult's living environment, such as their home, room, or specific room area. Monitoring areas clearly define the places that need to be monitored to ensure the safety and well-being of older adults.

[0038] The size and shape of the monitoring area can vary depending on specific monitoring needs and the needs of the elderly person. Typically, this area covers the elderly person's main activity areas, such as the living room, bedroom, and bathroom, as well as potentially hazardous areas such as hallways or staircases. The monitoring area is specifically designed to detect locations where unusual activity may occur, such as falls, prolonged inactivity, or other potential emergencies.

[0039] The selection of cameras should be based on the characteristics of the monitored area, such as its size, shape, and the location to be monitored. Cameras can be standard surveillance cameras or purpose-specific cameras, such as depth cameras to capture depth information. The location and angle of the camera should be chosen to ensure comprehensive coverage of the monitored area and capture the activities of the target user.

[0040] Once the camera is installed and configured, it will begin collecting image data of the monitored area. This data can be captured as a continuous video stream or periodically captured still images, depending on the design and requirements of the monitoring system. The camera records all activity within the monitored area, including that of elderly people, providing crucial data for subsequent analysis.

[0041] Depth information is three-dimensional coordinate data that helps determine the position and distance of objects, rather than just two-dimensional image information. This depth information provides more information about activity in the monitored area, which is helpful for subsequent analysis and the construction of abstract human models.

[0042] Depth information can be determined in various ways, including using dedicated distance sensors or the depth-of-field functionality built into a depth camera. Distance sensors are devices specifically designed to measure the distance from an object to the sensor. They use various techniques, such as structured light, time-of-flight, or binocular vision, to acquire depth information. These sensors provide very accurate depth information, typically expressed in pixels or as actual distance.

[0043] Some cameras have depth-of-field capabilities, allowing them to measure the distance to objects. These cameras typically use dual cameras or other technologies to capture depth information. The built-in depth-of-field functionality of a camera can provide a certain level of depth information, especially useful for special effects when taking photos or videos. However, their accuracy may be lower compared to dedicated distance sensors, especially at long distances or in complex scenes.

[0044] When monitoring elderly individuals, the choice of which method to use to determine depth information typically depends on the specific needs and budget of the monitoring system. Dedicated distance sensors usually provide more accurate depth data and are suitable for applications requiring high precision. However, if the camera has sufficient depth-of-field capabilities and meets the requirements of the monitoring system, it can also be used to reduce cost and complexity.

[0045] The preset reference model is typically a computer-generated two-dimensional model representing the two-dimensional skeletal structure of a standard human body, including key skeletal nodes such as the head, shoulders, elbows, hips, and knees. This two-dimensional model primarily considers the position and relative relationships of the skeletal structures on the image plane.

[0046] Skeletal node information refers to specific key points or locations in the human biological structure, usually represented by numerical coordinates or other methods, used to describe the posture, movement, and structure of the human body. These nodes are reference points distributed throughout various parts of the human body and can be used to analyze, identify, and track the body's movements, position, and state.

[0047] Once the monitoring system collects image data, it uses computer vision and image processing techniques to identify and locate the two-dimensional skeletal node information of the target user. By analyzing feature points in the image, such as the head, arms, and legs, the system can determine their positions and relationships within the image.

[0048] This skeletal node information is crucial for monitoring the activities of the elderly. It provides detailed data that helps the monitoring system understand the target user's posture, activities, and location. The results of this step will be used as input for subsequent steps to build an abstract human body model and detect anomalies, thereby enabling the monitoring and safety of the elderly.

[0049] This skeletal node information is matched against a pre-defined reference model. The reference model is a known planar template that describes a standard human skeletal structure, including the position and connection relationships of each skeletal node. By comparing the extracted skeletal node information of the target user with the reference model, the system can determine how the user's skeletal nodes are associated with corresponding nodes in the reference model, and the relationships between them.

[0050] Once the skeletal node information is successfully matched with the reference model, the system can create an abstract human body model of the target user. This model uses a two-dimensional coordinate mathematical representation or data structure to describe the user's body structure, including information such as the connections between nodes, the position and angle of joints. This abstract model represents the user's body posture and structure on a plane.

[0051] An abstract human body model is a two-dimensional model used to describe the human body's movement states, typically represented using a two-dimensional coordinate system. This model usually includes the positions of various skeletal nodes, such as the head, arms, and legs, which can be represented by x and y coordinates.

[0052] For example, a two-dimensional stick figure model can be used as an abstract human body model. The first step in constructing a two-dimensional stick figure model is to acquire body shape data containing the human body using video or images. The second step is to use pose prediction algorithms to detect the two-dimensional coordinate information of key points of the human body in each image frame. Then, the model framework of the stick figure can be drawn directly using these two-dimensional coordinate information by connecting the corresponding key points with simple lines. This constitutes a two-dimensional abstract stick figure model.

[0053] This model can be used to represent a user's posture and movement, despite limitations in a two-dimensional plane. By continuously updating the model, the system can track the user's dynamic changes to capture their movement, such as walking or raising a hand. This two-dimensional abstract human body model provides a simple and efficient way to understand and monitor a user's actions and postures while protecting user privacy because it does not require directly displaying an image of the user's body.

[0054] The digital twin space corresponding to the monitoring area is a virtual environment that is precisely mapped to the actual monitoring area in order to more accurately simulate and reflect the physical environment and scene of that area. The digital twin space is created by establishing a virtual three-dimensional space that corresponds one-to-one with the actual monitoring area, so as to enable real-time monitoring, analysis, and simulation within it.

[0055] A digital twin space is typically a computer-generated virtual environment whose structure and characteristics mimic the actual monitored area as closely as possible, including rooms, buildings, equipment, and other objects. This space may include various elements such as walls, floors, furniture, and the locations of surveillance cameras and other sensors.

[0056] To map a two-dimensional abstract human body model to a digital twin space, the system needs depth information. Depth information provides data about the object's distance and position in three-dimensional space. This is typically obtained through distance sensors.

[0057] Preset position transformation relationships are a set of rules or equations used to map two-dimensional coordinates to three-dimensional space. These rules can be determined based on the system's camera configuration, scene characteristics, and depth sensor performance. Position transformation relationships help the system convert node coordinates in a two-dimensional abstract human body model into three-dimensional coordinates. After determining the three-dimensional coordinates, the system can apply the preset position transformation relationships to map this three-dimensional model to a twin digital space.

[0058] Once the depth information and positional transformation relationships are determined, the system can map the abstract human body model to a digital twin space. The digital twin space is a virtual environment corresponding to the actual monitoring area. Within this space, the node positions of the abstract model are accurately repositioned to reflect the user's position and movement in the real world. This mapping method allows the system to monitor, analyze, and simulate within this virtual environment, enabling it to more accurately determine the user's motion state.

[0059] Motion state refers to the real-time state of a user's body movements and postures. In monitoring and analysis systems, this typically refers to the user's body position, movement, and activity at a specific point in time. Motion state monitoring usually includes the user's body postures, such as sitting, standing, lying down, bending, and stretching. It can also include the user's motor activities, such as walking, running, raising arms, and swinging. Simultaneously, the user's location is also part of the motion state, including the user's actual coordinates within the monitoring area. The monitoring system can track the user's location to determine if they have deviated from a safe area or are in an abnormal position. Furthermore, time information is also crucial for monitoring motion state because it can be used to detect the duration of the user's activity, for example, whether the user remained stationary or active within a specific time period.

[0060] Since image data can be static images or continuous video data, when using a single static image, the primary object of monitoring is usually the user's current posture or position. By analyzing human features and skeletal nodes in the image, the user's posture and position at that instant can be determined. When the image data is video data, the monitoring system can capture the user's dynamic motion state. In video monitoring, the system can track changes in the user's movements, activities, and postures over time to analyze the user's behavior and position in real time.

[0061] Monitoring systems need to clearly define and set preset abnormal conditions. These conditions can vary depending on the specific application scenario and monitoring objectives. Abnormal conditions typically include situations where the user may be in a dangerous posture, their movements violate certain safety standards, or the user requires medical assistance. These conditions may be a single condition or a combination of multiple factors, such as posture, movement, time, and location.

[0062] For example, abnormal conditions could include triggering an alert when the monitoring system detects a user's fall; this might include specific body postures and movement patterns indicating a potential injury. An alert could also be triggered when an elderly person is almost completely inactive for a specific period to prevent them from needing assistance. Alternatively, an abnormal condition could be triggered when an elderly person leaves their living area to notify caregivers.

[0063] When the system detects that a user's movement meets predefined abnormal conditions, it will generate an alert. This typically includes sending notifications via various communication methods, such as audible alarms, SMS messages, emails, and app notifications. The alert will then be sent to relevant operators, caregivers, family members, or emergency services so they can take appropriate action to handle the situation.

[0064] In addition to generating early warning information, the system can also implement corresponding response measures. These may include guiding users to take safety measures, sending emergency requests, or triggering automated systems to assist users, depending on the goals and functions of the monitoring system.

[0065] As can be seen from the above technical solution, the beneficial effects of this embodiment are: determining the image data collected by the target camera for the monitored area; determining the depth information corresponding to the image data using a distance sensor; determining multiple skeletal node information of the target user in the image data based on a preset reference model; determining the abstract human body model corresponding to the target user based on the coordinates of each skeletal node and the reference model; mapping the abstract human body model to the digital twin space corresponding to the monitored area through the depth information and a preset position transformation relationship, and determining the motion state of the abstract human body model in the digital twin space; generating early warning information when the motion state meets preset abnormal conditions. This embodiment realizes the monitoring of the activities and postures of the elderly using an abstract human body model, ensuring the efficiency and accuracy of monitoring while protecting privacy without infringing on the privacy of the elderly, thus improving the user experience.

[0066] In some embodiments, the method further includes: determining reference nodes based on preset key human body parts; determining first node identifiers corresponding to each reference node; determining logical connection methods between each first node identifier based on human body structural information; and determining a reference model based on the reference nodes, first node identifiers, and logical connection methods.

[0067] Specifically, key parts of the human body usually refer to important components of the human body structure, typically including major body parts and organs. These parts are particularly important in various monitoring and analysis tasks because they contain key information about human posture, movement, and other aspects.

[0068] The selection and definition of key human body parts will depend on the specific monitoring task and application area. These typically include the head, neck, knees, elbows, and hips. These parts are often used as reference nodes to help the system understand the user's posture and movements. They play a crucial role in monitoring and analyzing the user's motion status, health status, and in performing specific tasks.

[0069] A first-node identifier is a unique mark or identifier used to identify a reference node. In a monitoring system, these identifiers are used to distinguish different reference nodes, ensuring that each node can be accurately identified and tracked. Each reference node is assigned a specific first-node identifier, enabling the system to recognize its identity.

[0070] Furthermore, human anatomy information is used to describe the relationships, locations, connections, and range of motion of the human skeleton and joints. Then, based on the actual connection structure of the human skeleton and joints, the logical connections between these first-node identifiers are determined.

[0071] The logical connection method reflects the biological connections between human skeletons and expresses the abstract connection relationships between reference nodes. It is determined based on the actual biological structure of the human body, reflecting the connection order between bones and joint constraints. The connection order of the first node is determined according to the natural growth direction of the human skeleton, such as head connecting to neck, neck connecting to torso, etc.

[0072] Each uniquely identified reference node is connected according to a connection logic to form a complete node structure from the head, neck, torso to limbs, serving as the final human skeletal reference model. This model typically uses a graphical, tree-structured, or other suitable data representation method to describe the reference nodes, the first node identifier, and the logical connections between them. The reference model is the foundation of the monitoring system; it helps the system understand the connections and interactions between different parts of the human body structure, thereby accurately analyzing and tracking the user's movement status.

[0073] In some embodiments, determining multiple skeletal node information of a target user in image data based on a preset reference model includes: using image recognition technology to determine the body data of the target user in the image data; matching the body data with reference nodes to determine a second node identifier corresponding to the skeletal node of the target user; determining the two-bit pixel coordinates of the skeletal node in the current image of the image data as the skeletal node coordinates; and determining the skeletal node information based on the second node identifier and the skeletal node coordinates.

[0074] Specifically, the morphology and structure of the human body are similar across different individuals. By using a reference model, a consistent skeletal structure calibration can be established across different users. This means that the system can more easily identify and track skeletal nodes of different users without requiring independent training or calibration on each user.

[0075] Using a pre-defined reference model can accelerate the extraction of skeletal node information. Compared to analysis from scratch, the reference model provides a known framework that can be quickly matched to image data, enabling real-time monitoring.

[0076] The reference models provide information about the expected locations and connections of skeletal nodes. This helps improve the accuracy of skeletal node information because it is based on knowledge of human anatomy, rather than relying entirely on image recognition technology. This can reduce errors and improve monitoring precision.

[0077] Furthermore, image recognition technology is used to identify and extract the target user's body shape data from the image. This can include information such as height, weight distribution, body proportions, and contours. Image recognition technology can analyze features, shapes, and textures in an image to generate this body shape data. This is fundamental information that helps in understanding the target user's overall body structure.

[0078] Once the body measurements are determined, the next step is to match these measurements against a pre-defined reference model. This matching process helps determine which body features and shapes correspond to skeletal nodes in the reference model. For example, specific height and limb proportions can help determine the locations of skeletal nodes in the torso and limbs.

[0079] Secondary node identifiers are flags or identifiers used in the monitoring system to identify and track user skeletal nodes. These identifiers are determined by matching extracted body shape data with primary node identifiers. Secondary node identifiers are used to associate body features or body parts in the body shape data with skeletal nodes in a reference model. They are based on the body shape data to determine which body features correspond to which nodes.

[0080] By associating the first node identifier with the second node identifier, the monitoring system can accurately identify and track the user's skeletal nodes in real-time monitoring and analysis, enabling precise posture analysis and motion monitoring. This relationship helps the system better understand and interpret the user's body structure and movements.

[0081] Once the second node identifier is determined, the system can map the positions of these skeletal nodes to pixel coordinates in the current image. This is achieved by detecting body features and shapes in the image to determine the positions of the skeletal nodes within the current image. These two-dimensional pixel coordinates represent the specific locations of the nodes in the current image.

[0082] By combining the second node identifier with the pixel coordinates of the skeletal nodes, the system can determine information about each skeletal node. This includes the node's location, relationships, and its specific representation in the image. This node information can serve as the basis for subsequent skeletal model construction and motion analysis.

[0083] In summary, this process combines image recognition technology with a reference model to extract the target user's body shape data and skeletal node information from image data. This provides the data needed by the system to build an abstract human body model and achieve precise monitoring and early warning functions.

[0084] In some embodiments, determining the abstract human body model corresponding to the target user based on the information of each skeletal node and the reference model includes: connecting the coordinates of each node according to a logical linking method to obtain the abstract human body model corresponding to the target user.

[0085] Specifically, abstract human models provide an abstract representation of a user's posture and movements. By understanding the relationships between a user's body structure and skeletal nodes, the system can accurately analyze the user's actions and postures. In monitoring systems, abstract human models can help detect abnormal behavior or dangerous situations. By understanding the user's posture, the system can detect unusual postures or movements, thereby triggering alarms or taking appropriate actions. This is particularly useful in security monitoring and elderly care.

[0086] Logical linking rules are the rules and methods used to connect the coordinates of individual skeletal nodes to each other. These rules are determined based on definitions in the reference model and knowledge of human anatomy. They specify how to connect the coordinates of skeletal nodes to construct an abstract human body model. For example, logical linking rules might specify that the elbow node connects to the wrist node. Logical linking rules guide how to connect the coordinates of individual skeletal nodes to construct an abstract human body model. These logical linking rules are determined based on definitions in the reference model and knowledge of human anatomy.

[0087] Furthermore, since the node coordinates include the second node identifier, and the logical connection method includes the first node identifier, the node coordinates are connected based on the first node identifier, the second node identifier, and the logical connection method to construct an abstract human body model. This model will represent the user's body structure and posture, helping the system to better understand and analyze the user's movements.

[0088] By following a logical linking method, the system connects the coordinates of various skeletal nodes together to form an abstract human body model. This model can be viewed as a simplified, geometric representation of the user's overall posture and body structure. This abstract model is the foundation for the system to analyze and track user movements.

[0089] For example, an abstract human body model can be likened to a stick figure, used to illustrate a simplified representation of human anatomy. A stick figure is a graphic composed of short lines and small dots, used to represent the head, torso, limbs, and joints. This abstract stick figure model does not contain detailed facial features or details, but it is sufficient to represent the basic structure and posture of the human body. The stick figure can represent different actions and postures based on its joint connections, such as raising an arm, running, bending over, etc. This abstract representation is very useful for quickly understanding and visualizing human movements.

[0090] In some embodiments, the method further includes: determining at least three non-collinear reference feature points in the image data; determining the three-dimensional spatial coordinates corresponding to the reference feature points based on depth information; determining the digital feature points corresponding to the reference feature points in the digital twin space; and determining the digital spatial coordinates of the digital feature points in the digital coordinate system corresponding to the digital twin space; calculating the rotation matrix and translation matrix required when each three-dimensional spatial coordinate and the corresponding digital spatial coordinates overlap simultaneously; and determining the position transformation relationship based on the rotation matrix and translation matrix.

[0091] Specifically, the position transformation relationship describes how to map an abstract human body model from one coordinate system to another, ensuring data consistency. This is achieved by establishing a spatial coordinate system to map the abstract human body model from image data into a digital twin space. Static objects can be used as reference feature points, ensuring a consistent coordinate system across different time points and scenarios for reliable data analysis and comparison.

[0092] Reference feature points are specific points or markers that are clearly located in an image or surveillance scene and used as a reference. These feature points are typically used to establish coordinate systems, perform tracking, measure object positions, or perform other related tasks. The selection of reference feature points usually depends on the application requirements and the characteristics of the scene. They should generally be easily identifiable, not easily affected by environmental changes, and able to provide the necessary information for further analysis or measurement.

[0093] Furthermore, at least three reference feature points need to be identified in the surveillance image. These feature points can typically be prominent objects or markers in the image. These points should be selected to be non-collinear to ensure that an accurate coordinate system can be constructed in three-dimensional space.

[0094] Using depth information, the coordinates of these reference feature points in three-dimensional space can be calculated. Depth information provides information about the distance of these points from the observer or sensor, thus constructing a three-dimensional coordinate system.

[0095] In a digital twin space, it is necessary to determine the digital feature points corresponding to these reference feature points. These digital feature points are typically part of the digital twin model, used to represent different parts. The coordinates of these digital feature points in the digital coordinate system within the digital twin space are then determined. This ensures that the coordinate system of the digital feature points is consistent with that of the digital model.

[0096] To map image data to a digital twin space, rotation and translation matrices need to be computed. These matrices describe how coordinates are transformed from one coordinate system to another. The computation of these matrices typically involves linear algebra and geometric transformations.

[0097] Based on the calculated rotation and translation matrices, the positional transformation relationships are determined. These relationships describe how to map the abstract human body model from image data to a digital twin space for further analysis and application.

[0098] In some embodiments, mapping an abstract human body model to a digital twin space corresponding to a monitoring area using depth information and a preset position transformation relationship includes: determining the coordinates of a two-dimensional model corresponding to the abstract human body model; determining the coordinates of a three-dimensional model corresponding to the abstract human body model based on the two-dimensional model coordinates and depth information; and mapping the abstract human body model to the digital twin space based on the three-dimensional model coordinates and the position transformation relationship.

[0099] Specifically, a digital twin provides a virtual environment where broader analysis and applications can be performed. Mapping an abstract human model to a digital twin ensures the model's consistency across different environments, making analysis and experimentation in the digital environment more reliable. Mapping an abstract human model to digital space allows for more accurate pose estimation and motion tracking.

[0100] First, it is necessary to determine the two-dimensional coordinates of the abstract human model in the monitoring image. These are coordinates on the image plane, typically represented using pixel coordinates. These coordinates describe the model's position in the monitoring image. By combining the two-dimensional model coordinates with depth information, the coordinates of the abstract human model in three-dimensional space can be calculated. Depth information provides information about the distance between the model and the observer or sensor, thus mapping the model from a plane to three dimensions.

[0101] Using the previously calculated positional transformation relationships, the abstract human model can be mapped from the 3D coordinates of the monitoring image to the digital twin space. This ensures the consistency of the model across different coordinate systems, enabling further analysis and applications.

[0102] In summary, the purpose of this process is to map the model from two-dimensional image data to three-dimensional coordinates in the digital twin space, enabling further analysis, pose estimation, motion tracking, and other tasks. This step ensures that the data mapping from image data of the real-world scene to the digital twin space is feasible and accurate.

[0103] In some embodiments, where the image data is an image frame in video data, determining the motion state of the abstract human model in the digital twin space includes: monitoring the human posture of the target user in each image frame according to any one of claims 1-6; when the human posture in the image frame satisfies a preset reference posture, determining the motion state of the abstract human model in the digital twin space based on the reference posture; determining human actions based on the human posture in consecutive image frames; and when the human actions satisfy a preset action pattern, determining the motion state of the abstract human model in the digital twin space based on the action pattern.

[0104] Specifically, an image frame is a single image in video data, which is part of a continuous video stream. These image frames are typically captured at fixed time intervals, usually measured in frames per second (e.g., 30 frames per second). Each image frame contains visual information about a moment in time, which can be used to analyze and understand scenes, objects, and actions in the video.

[0105] Video data consists of a series of consecutive image frames that capture the real-time dynamics of a target user. This continuity allows the system to track the user's posture and movements in real time to obtain more comprehensive information. By analyzing the consecutive image frames in the video data, the user's motion trajectory can be generated, including speed, direction, and changes in motion.

[0106] Besides continuous image frames, motion can also be determined based on the target user's posture in a single frame, especially in situations where real-time requirements are not high and static analysis is sufficient. When the monitored scene is relatively static and the user's movements are relatively slow or limited, single-frame image analysis may be adequate. For example, when monitoring elderly people indoors, whose movements may be infrequent, single-frame image analysis can provide useful information.

[0107] Furthermore, the primary task is to detect, track, and identify the target user's human pose in each image frame. This may include detecting key human landmarks, such as the head, hands, shoulders, hips, etc., as well as the location of skeletal nodes.

[0108] In the system, preset reference postures are predefined poses that may represent normal, safe, or desired postures. These postures can be defined according to the specific needs of the application. For example, for monitoring the body posture of the elderly, a reference posture could be the correct posture while sitting in a chair.

[0109] The system compares the detected human pose in an image frame with a preset reference pose. This may involve calculating similarity metrics between poses, such as Euclidean distance or angular differences. If the current pose matches the reference pose sufficiently well, the system considers the current pose appropriate.

[0110] Once the human pose in an image frame meets a preset reference pose, the system can determine the motion state of the abstract human model in the digital twin space based on this information. This means that the system considers the user's motion state to be normal and matching the reference pose.

[0111] Alternatively, the system can continuously analyze image frames from a camera to detect and track the target user's human posture. This may include the detection of key human landmarks, such as hands, head, feet, elbows, and knees, to understand changes in posture. By comparing human postures in consecutive image frames, the system can capture the user's movements and actions. This can be used to detect whether the user is walking, raising their hand, bending, jumping, or performing other actions.

[0112] In the system, preset action patterns are a series of human movements predefined according to specific application needs, which can represent a certain activity or state. The definition of these action patterns can be adjusted according to specific circumstances. For example, a special action pattern can be defined to monitor whether an elderly person has fallen. If the system detects a fall-related action, it can immediately generate an alarm or notify relatives or caregivers.

[0113] The system compares detected human movements from consecutive image frames with preset movement patterns. This may involve calculating the degree of matching between the features, sequence, or posture of the movement. If the detected movement matches the preset movement pattern with a sufficiently high degree of matching, the system assumes that the user is performing a specific action.

[0114] Once the detected human movement matches a preset movement pattern, the system can determine the motion state of the abstract human model in the digital twin space. This means that the system considers the user's movement to match the preset movement pattern.

[0115] All of the above-mentioned optional technical solutions can be combined in any way to form the optional embodiments of this application, and will not be described in detail here.

[0116] The following are embodiments of the apparatus described in this application, which can be used to execute the embodiments of the method described in this application. For details not disclosed in the apparatus embodiments of this application, please refer to the embodiments of the method described in this application.

[0117] Figure 2 This is a schematic diagram of the structure of an early warning device based on an abstract human body model provided in an embodiment of this application. Figure 2 As shown, the early warning device based on an abstract human body model includes:

[0118] The depth information determination module 201 is used to determine the image data collected by the target camera for the monitored area; and to determine the depth information corresponding to the image data using a distance sensor.

[0119] The skeletal node information determination module 202 is used to determine multiple skeletal node information of the target user in image data based on a preset reference model.

[0120] Abstract human body model determination module 203 is used to determine the abstract human body model corresponding to the target user based on the coordinates of each skeletal node and the reference model.

[0121] The motion state determination module 204 is used to map the abstract human body model to the digital twin space corresponding to the monitoring area through depth information and preset position transformation relationship, and to determine the motion state of the abstract human body model in the digital twin space.

[0122] The early warning information generation module 205 is used to generate early warning information when the motion state meets preset abnormal conditions.

[0123] In some embodiments, Figure 2 The skeletal node information determination module 202 uses image recognition technology to determine the body data of the target user in the image data; it matches the body data with reference nodes to determine the second node identifier corresponding to the skeletal node of the target user; it determines the two-pixel coordinates of the skeletal node in the current image of the image data as the skeletal node coordinates; and it determines the skeletal node information based on the second node identifier and the skeletal node coordinates.

[0124] In some embodiments, Figure 2 The abstract human body model determination module 203 connects the coordinates of each node according to the logical linking method to obtain the abstract human body model corresponding to the target user.

[0125] In some embodiments, Figure 2 The motion state determination module 204 determines the coordinates of the two-dimensional model corresponding to the abstract human body model; based on the coordinates of the two-dimensional model and the depth information, it determines the coordinates of the three-dimensional model corresponding to the abstract human body model; based on the coordinates of the three-dimensional model and the position transformation relationship, it maps the abstract human body model to the digital twin space.

[0126] In some embodiments, Figure 2 The motion state determination module 204 monitors the human posture of the target user in each image frame; when the human posture in the image frame meets the preset reference posture, the motion state of the abstract human model in the digital twin space is determined according to the reference posture; the human action is determined according to the human posture in the continuous image frames; when the human action meets the preset action pattern, the motion state of the abstract human model in the digital twin space is determined according to the action pattern.

[0127] Figure 3 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. At the hardware level, the electronic device includes a processor, and optionally also includes an internal bus, a network interface, and a memory. The memory may include RAM, such as high-speed random-access memory (RAM), or non-volatile memory, such as at least one disk storage device. Of course, the electronic device may also include other hardware required for other services.

[0128] The processor, network interface, and memory can be interconnected via an internal bus, which can be an ISA (Industry Standard Architecture) bus, a PCI (Peripheral Component Interconnect) bus, or an EISA (Extended Industry Standard Architecture) bus, etc. This bus can be divided into address bus, data bus, control bus, etc. For ease of representation, Figure 3 The symbol is represented by a single double-headed arrow, but this does not mean that there is only one bus or one type of bus.

[0129] Memory is used to store instructions for execution. Specifically, instructions for execution are computer programs that can be executed. Memory can include main memory and non-volatile memory, and it provides the processor with execution instructions and data.

[0130] In one possible implementation, the processor reads the corresponding execution instructions from non-volatile memory into main memory and then executes them. Alternatively, it may obtain the corresponding execution instructions from other devices to logically form an early warning device based on an abstract human body model. The processor executes the execution instructions stored in the memory to implement the early warning method based on an abstract human body model provided in any embodiment of this application.

[0131] The above is as stated in this application. Figure 2The method for executing an early warning device based on an abstract human body model, as provided in the illustrated embodiment, can be applied to or implemented by a processor. The processor may be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above method can be completed through integrated logic circuits in the processor's hardware or through software instructions. The processor can be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; it can also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this application. The general-purpose processor can be a microprocessor or any conventional processor.

[0132] The steps of the method disclosed in the embodiments of this application can be directly manifested as being executed by a hardware decoding processor, or executed by a combination of hardware and software modules in the decoding processor. The software modules can reside in random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, or other mature storage media in the art. This storage medium is located in memory, and the processor reads information from the memory and, in conjunction with its hardware, completes the steps of the above method.

[0133] This application also proposes a readable medium storing execution instructions. When these instructions are executed by the processor of an electronic device, the electronic device can perform a warning method based on an abstract human body model provided in any embodiment of this application, and specifically execute, for example... Figure 1 The method shown.

[0134] The electronic devices in the foregoing embodiments may be computers.

[0135] Those skilled in the art will understand that the embodiments of this application can be provided as methods or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or a combination of software and hardware.

[0136] The various embodiments in this application are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the device embodiments are basically similar to the method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions of the method embodiments.

[0137] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.

[0138] The above are merely embodiments of this application and are not intended to limit the scope of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of the claims of this application.

Claims

1. An early warning method based on an abstract human body model, characterized in that, include: Determine the image data collected by the target camera in the monitored area; The depth information corresponding to the image data is determined using a distance sensor; Based on a preset reference model, multiple skeletal node information of the target user is determined in the image data; Based on the skeletal node information and the reference model, the abstract human body model corresponding to the target user is determined; Using the depth information and the preset position transformation relationship, the abstract human body model is mapped to the digital twin space corresponding to the monitoring area, and the motion state of the abstract human body model in the digital twin space is determined. When the motion state meets the preset abnormal conditions, an early warning message is generated; Also includes: Identify at least three non-collinear reference feature points in the image data; The three-dimensional spatial coordinates corresponding to the reference feature point are determined based on the depth information. In the digital twin space, determine the digital feature point corresponding to the reference feature point; and determine the digital spatial coordinates of the digital feature point in the digital coordinate system corresponding to the digital twin space. Calculate the rotation and translation matrices required when each of the three-dimensional spatial coordinates and the corresponding digital spatial coordinates overlap simultaneously; The position transformation relationship is determined based on the rotation matrix and the translation matrix; The step of mapping the abstract human body model to the digital twin space corresponding to the monitoring area using the depth information and a preset position conversion relationship includes: Determine the coordinates of the two-dimensional model corresponding to the abstract human body model; Based on the two-dimensional model coordinates and the depth information, determine the three-dimensional model coordinates corresponding to the abstract human body model; Based on the coordinates of the 3D model and the position transformation relationship, the abstract human body model is mapped to the digital twin space.

2. The method according to claim 1, characterized in that, Also includes: Based on the preset key parts of the human body, reference nodes are determined; And determine the first node identifier corresponding to each of the aforementioned reference nodes; The logical connection method between each of the first node identifiers is determined based on human body structure information; The reference model is determined based on the reference node, the first node identifier, and the logical connection method.

3. The method according to claim 2, characterized in that, The determination of multiple skeletal node information of the target user in the image data based on the preset reference model includes: Using image recognition technology, the body size data of the target user is determined from the image data; The body data is matched with the reference node to determine the second node identifier corresponding to the skeletal node of the target user; The two-bit pixel coordinates of the skeletal node in the current image of the image data are determined as the coordinates of the skeletal node. The bone node information is determined based on the second node identifier and the bone node coordinates.

4. The method according to claim 3, characterized in that, The step of determining the abstract human body model corresponding to the target user based on the skeletal node information and the reference model includes: The coordinates of each node are connected according to the logical connection method to obtain the abstract human body model corresponding to the target user.

5. The method according to claim 1, characterized in that, If the image data is an image frame from video data, then determining the motion state of the abstract human model in the digital twin space includes: The method according to any one of claims 1-4 monitors the human posture of the target user in each of the image frames; When the human body pose in the image frame satisfies the preset reference pose, the motion state of the abstract human body model in the digital twin space is determined according to the reference pose. Determine human movement based on the human posture in consecutive image frames; When the human body movement meets the preset movement pattern, the movement state of the abstract human body model in the digital twin space is determined according to the movement pattern.

6. An early warning device based on an abstract human body model, characterized in that, include: The depth information determination module is used to determine the image data collected by the target camera for the monitored area; The depth information corresponding to the image data is determined using a distance sensor; The skeletal node information determination module is used to determine multiple skeletal node information of the target user in the image data based on a preset reference model. The abstract human body model determination module is used to determine the abstract human body model corresponding to the target user based on the skeletal node information and the reference model. The motion state determination module is used to map the abstract human body model to the digital twin space corresponding to the monitoring area through the depth information and the preset position transformation relationship, and to determine the motion state of the abstract human body model in the digital twin space. The early warning information generation module is used to generate early warning information when the motion state meets preset abnormal conditions; Also includes: Identify at least three non-collinear reference feature points in the image data; The three-dimensional spatial coordinates corresponding to the reference feature point are determined based on the depth information. In the digital twin space, determine the digital feature point corresponding to the reference feature point; and determine the digital spatial coordinates of the digital feature point in the digital coordinate system corresponding to the digital twin space. Calculate the rotation and translation matrices required when each of the three-dimensional spatial coordinates and the corresponding digital spatial coordinates overlap simultaneously; The position transformation relationship is determined based on the rotation matrix and the translation matrix; The step of mapping the abstract human body model to the digital twin space corresponding to the monitoring area using the depth information and a preset position conversion relationship includes: Determine the coordinates of the two-dimensional model corresponding to the abstract human body model; Based on the two-dimensional model coordinates and the depth information, determine the three-dimensional model coordinates corresponding to the abstract human body model; Based on the coordinates of the 3D model and the position transformation relationship, the abstract human body model is mapped to the digital twin space.

7. A computer-readable storage medium storing a computer program for executing an early warning method based on an abstract human body model as described in any one of claims 1-5.

8. An electronic device, the electronic device comprising: processor; Memory used to store the processor's executable instructions; The processor is configured to read the executable instructions from the memory and execute the instructions to implement the early warning method based on an abstract human body model as described in any one of claims 1-5.