Sitting posture detection method, electronic equipment and storage medium
By performing target detection processing on sitting posture detection videos to generate human perception data, and conducting cross-frame correlation analysis and temporal feature fusion, the accuracy problem of existing sitting posture detection algorithms is solved, and accurate judgment of poor sitting posture is achieved.
Patent Information
- Application Number
- CN202511078614.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-01
- Publication Date
- 2025-11-18
AI Technical Summary
Existing posture detection algorithms have difficulty guaranteeing accuracy, are prone to false detections or missed detections, and cannot reliably detect the true poor posture. Furthermore, sensor-based methods are costly and inconvenient to install.
By acquiring sitting posture detection videos of the target detection area, target detection processing is performed to generate human perception data. Based on the human perception data, cross-frame correlation analysis is performed to extract facial spatial features and limb posture features frame by frame, and temporal feature fusion analysis is performed to determine whether the sitting posture is poor.
It improves the accuracy of sitting posture detection, reduces false positives and false negatives, ensures accurate detection of specific targets, and avoids detection errors caused by multiple people being present at the same time.
Smart Images

Figure CN120976973A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates to the technical field of visual recognition, and in particular to a sitting posture detection method, an electronic device and a storage medium. BACKGROUND
[0002] With the acceleration of modern life pace and the change of people's work and study mode, long-time sitting has become a common practice. However, long-term bad sitting posture is causing many serious harm to people's health. In the work scene, people often look down at the computer for a long time, causing the neck and shoulder to bear a huge pressure, which is easy to cause cervical spondylosis, periarthritis of shoulder and other diseases; in the learning process, students have bad habits such as writing with tilted head and doing homework with high and low shoulders, which not only affects the normal development of the body, but also distracts attention and reduces learning efficiency; in daily life, the phenomenon of playing mobile phones with low head is also very common, which further aggravates the damage to the cervical spine, eyes and other parts.
[0003] To improve this situation, a sitting posture detection algorithm has emerged, which can provide an alarm mode to help gradually correct bad sitting posture and cultivate good sitting habits. At present, the existing sitting posture detection algorithm is mainly divided into a detection method based on sensor features and a detection method based on visual features. The sensor-based method, such as the sitting posture detection algorithm based on laser radar, the detection algorithm based on TOF features and the detection algorithm based on binocular depth map, although it can detect the sitting posture to some extent, this kind of method often needs additional sensor equipment, which is high in cost and not convenient to install and use. The method based on visual features, such as the detection algorithm based on face frame, the detection algorithm based on head posture and the detection algorithm based on human key points, although it does not need additional complex equipment and only uses image information obtained by the camera for detection, these algorithms generally have defects. They mostly use only partial information of the body or the head to judge the sitting posture, and the accuracy is difficult to guarantee, which is prone to false detection or missed detection problems, and cannot reliably detect the real bad sitting posture, which is difficult to meet people's demand for accurate sitting posture detection. Therefore, it is urgent to develop a more accurate and reliable sitting posture detection method. SUMMARY
[0004] The present disclosure provides a sitting posture detection method, an electronic device and a storage medium. The main purpose is to solve the problem of low recognition accuracy of bad sitting posture.
[0005] According to a first aspect of the present disclosure, a sitting posture detection method is provided, comprising:
[0006] obtaining a sitting posture detection video of a target detection area, performing target detection processing on the sitting posture detection video to generate human perception data of at least one detection object;
[0007] Based on the human perception data, cross-frame correlation analysis is performed on the sitting posture detection video, and the target detection object in the at least one detection object that needs to be detected for sitting posture is tracked and determined;
[0008] The human features of the target detection object are extracted frame by frame in the sitting posture detection video, and the face space features and the limb posture features corresponding to each video frame are obtained;
[0009] The face space features and the limb posture features are subjected to time sequence feature fusion analysis, and it is judged whether the sitting posture of the target detection object is an unhealthy sitting posture.
[0010] Optionally, the sitting posture detection video of the target detection area is obtained, the target detection processing is performed on the sitting posture detection video, and the human perception data of the at least one detection object is generated, including:
[0011] Face recognition detection is performed on each video frame in the sitting posture detection video respectively, the face detection area corresponding to different detection objects in each video frame is determined, and the confidence of each face detection area is generated;
[0012] The face detection area is marked with an identity according to the detection object to which the face detection area belongs, and the identity of each face detection area is generated;
[0013] The human perception data of the at least one detection object is generated for the face detection area in each video frame and the confidence and identity corresponding thereto.
[0014] Optionally, based on the human perception data, cross-frame correlation analysis is performed on the sitting posture detection video, and the target detection object in the at least one detection object that needs to be detected for sitting posture is tracked and determined, including:
[0015] Based on the human perception data, the largest face detection area in the initial frame of the sitting posture detection video is selected, and the largest face detection area is marked as a target face detection area;
[0016] According to the identity of the target face detection area, cross-frame correlation analysis is performed on the target face detection area in the sitting posture detection video, and the number of disappearances of the target face detection area is obtained;
[0017] If the number of disappearances is less than the failure threshold, it is determined that the detection object corresponding to the target face detection area is the target detection object.
[0018] Optionally, the method for detecting sitting posture further comprises:
[0019] If the number of disappearances is greater than or equal to the failure threshold, the cross-frame correlation analysis of the target face detection area is terminated, and the target detection object that needs to be detected for sitting posture is determined again in the subsequent video frames.
[0020] Optionally, the human features of the target detection object are extracted frame by frame in the sitting posture detection video to obtain the facial spatial features and the limb posture features corresponding to each video frame, including:
[0021] According to the target identity of the target detection object, the sitting posture detection video is traversed to determine the target video frame where the target identity is located;
[0022] Based on the first number of facial feature points and the second number of limb feature points in each target video frame, the facial spatial features and the limb posture features corresponding to each video frame are respectively generated.
[0023] Optionally, the facial spatial features and the limb posture features are subjected to time sequence feature fusion analysis to determine whether the sitting posture of the target detection object is an unhealthy sitting posture, including:
[0024] According to the frame sequence of the sitting posture detection video, the facial spatial features and the limb posture features in different video frames are subjected to facial posture detection, limb posture detection and posture fusion detection in sequence;
[0025] The facial posture detection results, the limb posture detection results and the posture fusion detection results of different time sequences are fused to determine whether the sitting posture of the target detection object is an unhealthy sitting posture.
[0026] Optionally, the facial spatial features and the limb posture features in different video frames are subjected to facial posture detection, limb posture detection and cross-modal posture detection in sequence, including:
[0027] At least one target facial feature point and at least one target facial reference line in the facial spatial features are obtained, and at least one target limb feature point and at least one target limb reference line in the limb posture features are obtained;
[0028] According to the spatial topological relationship of the at least one target facial feature point, the facial posture detection of the target detection object is performed to generate a detection abnormality number;
[0029] Based on the at least one target limb reference line, it is determined whether the posture angle of the target detection object exceeds an angle threshold, and according to the spatial topological relationship of the at least one target limb feature point, the limb posture detection of the target detection object is performed to generate a detection abnormality number;
[0030] The spatial relationship between the at least one target facial reference line and the at least one target limb reference line is detected, and whether the target limb feature point is in the facial region is detected to perform posture fusion detection of the target detection object to generate a detection abnormality number.
[0031] Optionally, the facial posture detection results, the limb posture detection results and the posture fusion detection results of different time sequences are fused to determine whether the sitting posture of the target detection object is an unhealthy sitting posture, including:
[0032] The statistical face posture detection result, the limb posture detection result and the posture fusion detection result are used to detect the number of times of abnormality;
[0033] Based on the number of times of abnormality, it is determined whether the sitting posture of the target detection object is an unhealthy sitting posture.
[0034] According to a second aspect of the present disclosure, a device for detecting a sitting posture is provided, comprising:
[0035] A detection unit is configured to acquire a sitting posture detection video of a target detection region, perform target detection processing on the sitting posture detection video, and generate human perception data of at least one detection object;
[0036] A first determination unit is configured to perform cross-frame correlation analysis on the sitting posture detection video based on the human perception data, and track and determine a target detection object that needs to be detected for a sitting posture among the at least one detection object;
[0037] An extraction unit is configured to extract human features of the target detection object frame by frame in the sitting posture detection video, and obtain face spatial features and limb posture features corresponding to each video frame;
[0038] An analysis unit is configured to perform time sequence feature fusion analysis on the face spatial features and the limb posture features, and determine whether the sitting posture of the target detection object is an unhealthy sitting posture.
[0039] Optionally, the detection unit comprises:
[0040] A first detection module is configured to perform face recognition detection on each video frame in the sitting posture detection video respectively, determine face detection regions corresponding to different detection objects in each video frame, and generate a confidence level corresponding to each face detection region;
[0041] A labeling module is configured to label the face detection regions according to the detection objects to which the face detection regions belong, and generate an identity identifier corresponding to each face detection region;
[0042] A first generation module is configured to generate human perception data of at least one detection object for the face detection region in each video frame and the confidence level and the identity identifier corresponding thereto.
[0043] Optionally, the first determination unit comprises:
[0044] A marking module is configured to select a largest face detection region in an initial frame of the sitting posture detection video based on the human perception data, and mark the largest face detection region as a target face detection region;
[0045] The analysis module is configured to perform cross-frame correlation analysis on the target face detection region in the sitting posture detection video according to the identity of the target face detection region, and obtain a number of disappearances of the target face detection region.
[0046] The first determination module is configured to determine that the detection object corresponding to the target face detection region is the target detection object when the number of disappearances is less than the invalidation threshold.
[0047] Optionally, the device for detecting a sitting posture further includes:
[0048] The second determination unit is configured to terminate the cross-frame correlation analysis on the target face detection region when the number of disappearances is greater than or equal to the invalidation threshold, and to determine the target detection object that needs to be detected for a sitting posture in subsequent video frames.
[0049] Optionally, the extraction unit includes:
[0050] The second determination module is configured to traverse the sitting posture detection video according to the target identity of the target detection object, and determine a target video frame in which the target identity is located.
[0051] The second generation module is configured to generate a face space feature and a body posture feature corresponding to each video frame, respectively, based on the first number of face feature points and the second number of body feature points in each target video frame.
[0052] Optionally, the analysis unit includes:
[0053] The second detection module is configured to perform face posture detection, body posture detection, and posture fusion detection on the face space feature and the body posture feature in different video frames in sequence according to the frame sequence of the sitting posture detection video.
[0054] The judgment module is configured to fuse the face posture detection result, the body posture detection result, and the posture fusion detection result at different time sequences, and determine whether the sitting posture of the target detection object is an unhealthy sitting posture.
[0055] Optionally, the second detection module is further configured to:
[0056] At least one target face feature point and at least one target face reference line in the face space feature are obtained, and at least one target body feature point and at least one target body reference line in the body posture feature are obtained.
[0057] The face posture detection is performed on the target detection object according to the spatial topological relationship of the at least one target face feature point to generate a detection abnormality number.
[0058] determine whether the posture angle of the target detection object exceeds an angle threshold based on the at least one target limb reference line, and perform limb posture detection on the target detection object to generate a detection abnormality number according to the spatial topological relationship of the at least one target limb feature point;
[0059] detect the spatial relationship between the at least one target face reference line and the at least one target limb reference line, and detect whether the target limb feature point is in the face region, to perform posture fusion detection on the target detection object to generate a detection abnormality number.
[0060] Optionally, the judging module is further configured to:
[0061] count the abnormality number of the detection abnormality of the face posture detection result, the limb posture detection result and the posture fusion detection result;
[0062] determine whether the sitting posture of the target detection object is an unhealthy sitting posture based on the abnormality number.
[0063] According to a third aspect of the present disclosure, an electronic device is provided, comprising:
[0064] at least one processor; and
[0065] a memory in communication with the at least one processor; wherein
[0066] the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method of the sitting posture detection of the first aspect.
[0067] According to a fourth aspect of the present disclosure, a non-transitory computer readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable the computer to perform the method of the sitting posture detection of the first aspect.
[0068] According to a fifth aspect of the present disclosure, a computer program product is provided, comprising a computer program which, when executed by a processor, implements the method of the sitting posture detection of the first aspect.
[0069] The method, the electronic device and the storage medium for sitting posture detection provided by the present disclosure relate to the technical field of computer vision, and compared with the related art, the embodiment of the present disclosure can generate human perception data by performing target detection processing on a sitting posture detection video; and can accurately track and determine a target detection object that needs to be subjected to sitting posture detection by performing cross-frame correlation analysis based on the human perception data. The target detection object in the target detection region can be perceived, and the target detection object that needs to be subjected to sitting posture detection can be determined, so that the problem of multi-object confusion can be solved, accurate sitting posture detection of a specific target can be ensured, and detection errors caused by the simultaneous presence of multiple people can be avoided. The human features of the target detection object are extracted frame by frame in the sitting posture detection video, and time sequence feature fusion analysis is performed. The extracted facial spatial features and limb posture features can be fused and analyzed, the sitting posture state of the target detection object can be comprehensively reflected, and the accuracy of the judgment of the bad sitting posture can be improved, and the occurrence of false detection and missed detection can be reduced.
[0070] It should be understood that the content described in this part is not intended to identify key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become apparent from the following description. BRIEF DESCRIPTION OF DRAWINGS
[0071] The accompanying drawings are used to better understand the present scheme and do not constitute a limitation on the present disclosure. Among them:
[0072] Figure 1 A flowchart of a method for sitting posture detection provided by an embodiment of the present disclosure;
[0073] Figure 2 A flowchart of another method for sitting posture detection provided by an embodiment of the present disclosure;
[0074] Figure 3 A structural diagram of a device for sitting posture detection provided by an embodiment of the present disclosure;
[0075] Figure 4 A structural diagram of another device for sitting posture detection provided by an embodiment of the present disclosure;
[0076] Figure 5 A schematic block diagram of an example electronic device provided by an embodiment of the present disclosure. DETAILED DESCRIPTION
[0077] Exemplary embodiments of the present disclosure are described below with reference to the accompanying drawings, which include various details of the embodiments of the present disclosure to help understanding, and should be considered as merely exemplary. Therefore, those skilled in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Also, in order to be clear and concise, the description below omits the description of well-known functions and structures.
[0078] The method for detecting a sitting posture, the electronic device, and the storage medium are described below with reference to the accompanying drawings.
[0079] Figure 1 A flowchart of a method for detecting a sitting posture provided by an embodiment of the present disclosure.
[0080] As shown in the method includes the following steps: Figure 1
[0081] In step 101, a sitting posture detection video of a target detection area is acquired, target detection processing is performed on the sitting posture detection video, and human perception data of at least one detection object is generated.
[0082] In an embodiment of the present disclosure, the target detection area refers to a specific spatial range that is pre-set and needs to be subjected to sitting posture detection. The range can be a seat area in a classroom or an office, or a monitorable range covered by a specific device (such as a learning desk and chair or an office chair equipped with a camera). The present disclosure does not limit this. The sitting posture detection video is video data acquired by an image acquisition device (for example, a camera) arranged in the target detection area. The video data contains human sitting posture information in the area. The acquired sitting posture detection video is subjected to target detection processing. The purpose is to identify specific detection objects that need to be subjected to sitting posture detection in the sitting posture detection video and determine the position and category and the like. The detection object is a target individual that needs to be subjected to sitting posture detection, which can be one or more personnel in the target detection area.
[0083] After the target detection processing, human perception data is generated. The human perception data is a digital description of the human state of the detection object, which contains various information related to the human body, such as the position information of the human body in the video picture, the contour information of the human body, the preliminary posture feature information, and the like. These data provide a basis for further analyzing the human sitting posture. The specific form of the human perception data is not limited in the embodiment of the present disclosure. For example, it can be face recognition related data or posture recognition.
[0084] The acquisition of the sitting posture detection video and the target detection processing to generate the human perception data lay a data foundation for subsequent accurate analysis and judgment of the sitting posture of the detection object. It can quickly and comprehensively identify the human target (detection object) in the target detection area and convert it into a data form that can be processed and analyzed by a computer, so that the subsequent monitoring and judgment of the sitting posture have feasibility and accuracy, and the overall performance and reliability of the sitting posture detection system are improved.
[0085] Step 102, based on the human perception data, performing cross-frame correlation analysis on the sitting posture detection video to track and determine the target detection object in the at least one detection object that needs to be subjected to sitting posture detection.
[0086] In an embodiment of the present disclosure, since the video is composed of a sequence of continuous frames, the state of the same detection object may change in different frames. The core purpose of cross-frame correlation analysis is to establish the correspondence between detection objects in different frames through a specific algorithm, so as to track the dynamic changes of the detection objects on the video time axis. In the sitting posture detection scenario, the analysis compares the human perception data in each frame, and determines whether the human bodies in different frames belong to the same detection object according to factors such as the similarity of human features and the continuity of motion trajectories.
[0087] Based on the results of cross-frame correlation analysis, the target detection object is selected from the at least one detection object. The target detection object refers to a specific human target selected by cross-frame correlation analysis in the sitting posture detection video, which needs to perform sitting posture state judgment. The target detection object is a specific detection object that really needs detailed sitting posture analysis in the current sitting posture detection task. This selection process may be performed according to pre-set rules or task requirements, for example, in some scenarios, the detection object in a specific area or meeting specific conditions (such as first entering the monitoring range) is preferentially selected as the target detection object.
[0088] Cross-frame correlation analysis ensures accurate identification and tracking of each detection object in the dynamic changes of the video, avoiding misjudgment or omission due to the movement of the detection object or the difference between video frames. By accurately determining the target detection object, the subsequent sitting posture detection can focus on a specific individual, improving the detection specificity and accuracy. It can effectively deal with the situation of multiple detection objects existing at the same time in complex scenes, optimize the process of sitting posture detection, and improve the reliability and practicality of the whole system.
[0089] Step 103, extracting the human features of the target detection object frame by frame in the sitting posture detection video to obtain the facial spatial features and limb posture features corresponding to each video frame.
[0090] In embodiments of the present disclosure, each frame of video is sequentially subjected to image processing in the time order of the video. Since a video is composed of a series of continuous static images (i.e., frames), frame-by-frame processing can comprehensively and meticulously analyze the state changes of the target detection object throughout the entire video process. The “human body features” are a collection of various information describing the body state of the target detection object, mainly including “facial spatial features” and “limb posture features” in this step. The “facial spatial features” are a collection of features reflecting the position, shape, angle, and other information of the face of the target detection object in three-dimensional space. These features can reflect the spatial relationship of the face relative to other parts of the body or a reference coordinate system, such as the pitch angle and yaw angle of the face, which are of great significance for analyzing the head posture and sitting posture. The “limb posture features” are features used to describe the position, stretching degree, joint angle, and other information of each limb of the target detection object. These features can intuitively show the actions and postures of the limbs, such as the placement position of the arms and the bending degree of the legs, which are the key basis for judging whether the sitting posture is correct.
[0091] In actual operation, specific feature extraction algorithms are used to analyze each frame of image in the sitting posture detection video. These algorithms are based on image processing, pattern recognition, and other related technologies, and can accurately separate the target detection object from the complex image background and extract the corresponding facial spatial features and limb posture features.
[0092] Step 104, time sequence feature fusion analysis is performed on the facial spatial features and limb posture features to determine whether the sitting posture of the target detection object is an improper sitting posture.
[0093] In embodiments of the present disclosure, the facial spatial features and limb posture features obtained in the previous step are analyzed in depth to determine whether the sitting posture of the target detection object is an improper sitting posture. Since the facial spatial features and limb posture features are obtained in different video frames, they have a time sequence. Time sequence feature fusion analysis is to use specific algorithms to integrate and comprehensively analyze these time-varying feature data. It not only considers the individual information of the facial spatial features and limb posture features in each frame, but also focuses on the change trend and mutual correlation of these features in the time dimension, organically combines the feature information at different times, and forms a more comprehensive and representative feature set.
[0094] Based on the results of the time sequence feature fusion analysis, it is determined whether the sitting posture of the target detection object is an improper sitting posture. In the determination, a series of rules or standards are used, which are developed based on the research and analysis of normal and improper sitting postures. The fused features are compared with these rules and standards, and then a conclusion is drawn as to whether the current sitting posture of the target detection object is an improper sitting posture.
[0095] Through the time sequence feature fusion analysis, the information of the facial space features and the body posture features in the time dimension is fully utilized, the limitation of judging only according to a single moment or a single type of feature is avoided, and the accuracy and reliability of the judgment are greatly improved. Accurate judgment of the bad sitting posture can timely find the bad sitting posture behavior of the target detection object.
[0096] The method for detecting a sitting posture provided by the present disclosure can accurately track and determine the target detection object that needs to be detected for a sitting posture, can perceive the target detection object in the target detection area, and determine the target detection object that needs to be detected for a sitting posture, can solve the problem of multi-object confusion, ensure accurate detection of the sitting posture of a specific target, and avoid detection errors caused by the simultaneous presence of multiple people. The human body features of the target detection object are extracted frame by frame in the sitting posture detection video, and time sequence feature fusion analysis is performed. The extracted facial space features and body posture features can be fused and analyzed, the sitting posture state of the target detection object can be comprehensively reflected, the accuracy of the bad sitting posture judgment can be improved, and the occurrence of false positives and false negatives can be reduced.
[0097] In order to clearly illustrate the embodiments of the present disclosure, the present embodiment provides a flowchart of another method for detecting a sitting posture.
[0098] As shown in Figure 2 the method comprises the following steps:
[0099] Step 201, performing face recognition detection on each video frame in the sitting posture detection video respectively, determining the face detection area corresponding to different detection objects in each video frame, and generating the confidence of each face detection area.
[0100] Specifically in step 201, the face recognition detection in the sitting posture detection video is performed, the purpose is to accurately determine the face detection area corresponding to different detection objects in each video frame, and generate the confidence data for each area.
[0101] Each video frame in the sitting posture detection video is processed by using a target detection algorithm such as yolov5. When processing each video frame, the algorithm analyzes the pixel information in the image, extracts possible face features through a multi-layer convolutional neural network, and determines the position and range of the face in the image according to these features, so as to obtain a “face detection area”. Each “face detection area” corresponds to a detection object in the video, i.e. the region where the face of a different person may exist.
[0102] At the same time of determining the face detection region, the system will also generate a "confidence" for each region. The confidence is a quantitative value that reflects the reliability of the detected face region as a real face. This value is calculated based on multiple factors such as the matching degree of face features, image quality, and the confidence model of the detection algorithm itself. For example, when the face features detected by the algorithm are highly matched with the standard face features in the training data, and the image is clear and unobstructed, the generated confidence will be higher; on the contrary, if the image is blurred and part of the face is obstructed, the confidence will be reduced.
[0103] Step 202, according to the detection object to which the face detection region belongs, identity labeling is performed on the face detection region, and the identity identifier corresponding to each face detection region is generated. For each face detection region in each video frame and its corresponding confidence and identity identifier, human perception data of at least one detection object is generated.
[0104] Specifically in step 202, the face detection region obtained in step 201 is further processed to generate human perception data for subsequent analysis, laying a more solid foundation for accurate sitting posture detection.
[0105] It is achieved by using a tracking algorithm such as deepSort algorithm. In complex scenarios, multiple detection objects may appear in the video at the same time, and their positions and postures are constantly changing. The deepSort algorithm uses the motion information and appearance features of the target to associate and match the face detection regions in different video frames with the corresponding detection objects through multi-target tracking technology. In this process, the algorithm assigns a unique "identity identifier" to each detection object, which is like everyone's "digital identity", enabling the system to accurately identify and distinguish different detection objects in consecutive video frames, avoiding confusion caused by the movement or obstruction of personnel.
[0106] Human perception data is a comprehensive data set that integrates multiple information. The face detection region determines the position and range of the detection object's face in the video frame, providing a basis for subsequent accurate analysis of facial features; the confidence reflects the reliability of the face detection region, which helps to filter out effective detection results in subsequent processing and improves the accuracy of data processing; and the identity identifier binds the detection information in different video frames with a specific detection object, ensuring the coherence and traceability of the data. The organic integration of these information generates human perception data that can more comprehensively and accurately describe the state of the detection object in the video, not only containing the position information of the face, but also covering the identity recognition information of the detection object and the reliability information of the detection result.
[0107] In step 203, based on the human perception data, the largest face detection region in the initial frame of the sitting posture detection video is selected, and the largest face detection region is marked as the target face detection region.
[0108] Specifically, in step 203, a specific face detection region is selected from the initial frame of the sitting posture detection video as the target for subsequent monitoring.
[0109] The initial frame refers to the first frame of the video, which serves as the starting point for the entire video analysis and carries important initial information. In this frame, there may be multiple detection objects, resulting in multiple face detection regions of different sizes. According to the coordinate information of the face detection region in the image, the number of pixels occupied by each region is calculated to determine its area size. Then, the areas of all regions are sorted to find the largest one. In actual application scenarios, such as in a classroom or office sitting posture monitoring scenario, the face of the detection object closer to the camera usually displays a larger area in the video frame. Selecting the largest face detection region as the target can prioritize those objects that may be more conducive to accurate detection and analysis, as larger face detection regions contain more facial details, which is beneficial for subsequent accurate facial feature extraction.
[0110] When performing time series supervision, the largest target box (target face detection region) ID number (identity) is initially locked for related analysis. In other embodiments of the present disclosure, the target face detection region can be further filtered based on factors such as face clarity and posture stability. For example, when there are multiple large face detection regions with similar areas, the region with high clarity and relatively stable posture is preferred as the target to improve the accuracy and stability of subsequent analysis.
[0111] In step 204, based on the identity of the target face detection region, cross-frame correlation analysis of the target face detection region in the sitting posture detection video is performed to obtain the number of disappearances of the target face detection region.
[0112] Specifically, in step 204, using deepSort and other tracking algorithms, all frames containing the target face detection region in the sitting posture detection video are processed. These algorithms will establish the association between frames based on the position, size, appearance features, and other information of the target face detection region in different frames, combined with its identity. Specifically, in each frame, the algorithm searches for a region matching the identity of the target face detection region, and through calculating position changes, feature similarities, and other indicators, it determines whether the target in the current frame and the target in the previous frame are the same object.
[0113] In the process of cross-frame association analysis, the system monitors the state of the target face detection region in real time. When the algorithm fails to find a region matching the identity of the target face detection region in a frame, it is determined that the target face detection region "disappears" in this frame, and the "disappearance times" are accumulated. The "disappearance times" is an important statistical data, which reflects the frequency of the interruption of the target face detection region in the video stream. In actual application scenarios, the disappearance of the target face detection region may be caused by the temporary departure of the detection object from the monitoring range, being blocked by other objects, etc.
[0114] In some embodiments of the present disclosure, a variety of feature information can be further combined to improve the accuracy of cross-frame association analysis. For example, in addition to using the position, size and appearance features of the face, dynamic features such as the motion trajectory and speed of the head can also be introduced. At the same time, by setting a reasonable threshold to determine whether the target has truly disappeared, it is possible to avoid misjudging the disappearance times due to temporary obstruction or detection errors. For example, when the target face detection region does not appear in several consecutive frames, but reappears within a short time, the statistics of the disappearance times can be adjusted according to the specific circumstances to ensure the accuracy of the data.
[0115] In step 205, the disappearance times are compared with the failure threshold to determine whether the detection object corresponding to the target face detection region is the target detection object.
[0116] If the disappearance times are less than the failure threshold, it is determined that the detection object corresponding to the target face detection region is the target detection object. If the disappearance times are greater than or equal to the failure threshold, the cross-frame association analysis of the target face detection region is terminated, and the target detection object that needs to be detected for the sitting posture is determined again in the subsequent video frames.
[0117] In step 205, in the process of video frame-by-frame analysis, the disappearance times refer to the cumulative number of frames in which the target face detection region is not successfully identified, reflecting the stability of the detection object state. The failure threshold is a key parameter set based on a large amount of experimental data, actual scene requirements, and detection accuracy and stability, and is used to measure the severity of the disappearance of the target face detection region and to determine whether the detection object is a valid target.
[0118] By comparing the disappearance times with the failure threshold, if the disappearance times are less than the failure threshold, it indicates that the disappearance of the target face detection region is within an acceptable range, and the detection object is likely to be within the monitoring range. The system determines that it is the target detection object and continues the analysis related to the sitting posture detection.
[0119] If the disappearance times are greater than or equal to the failure threshold, it indicates that the target face detection region disappears seriously, and the detection object may have left the monitoring range or the system is difficult to track. At this time, the system terminates the cross-frame association analysis of the current target face detection region, no longer takes it as the main monitoring object, stops the subsequent related analysis and calculation, and restarts the process of determining the target detection object in the subsequent video frames. When the target detection object is determined again, the method used in steps 203 to 204 is adopted.
[0120] In step 206, according to the target identity of the target detection object, the sitting posture detection video is traversed to determine the target video frame where the target identity is located.
[0121] Specifically, in step 206, during the traversal process, the system reads the information carried by each frame of image, especially the part about the identity of the detection object. Based on the target identity determined in advance, the system searches and compares in each frame, and judges whether the current frame contains the target detection object by searching the identity information. When the system successfully matches the information consistent with the target identity in a certain frame, it is determined that the frame is a "target video frame".
[0122] In step 207, based on the first number of facial feature points and the second number of limb feature points in each target video frame, the corresponding facial space features and limb posture features of each video frame are generated respectively.
[0123] Specifically, in step 207, the first number of facial feature points and the second number of limb feature points are the key data for describing the face and limb state of the target detection object. 68 facial feature points are obtained, which are used as the first number of facial feature points. They are accurately distributed in various key positions of the face, such as eyes, nose, mouth, cheeks, etc., and can comprehensively reflect the shape, position and posture information of the face. For the limb feature points, 17 key points of the human body are obtained, which are used as the second number of limb feature points, covering the main joints and skeletal endpoints of the human body, such as shoulders, elbows, wrists, hips, knees, ankles, etc. Through these points, the stretching, bending and relative position relationship of the limbs can be accurately described.
[0124] For facial space features, according to the three-dimensional coordinate information of facial feature points, the distances, angles and other geometric relationships between different parts of the face are calculated to determine the position, orientation and distortion degree of the face in three-dimensional space. For example, the distance between the two eyes, the angle of the line connecting the tip of the nose and the chin, and other information are calculated to comprehensively describe the facial space features. For the limb posture features, the relative positions and joint angles between the limb feature points are used to construct the posture model of the limbs. For example, the angle between the shoulder and the elbow, the relative position of the wrist to the body axis, etc. are calculated to determine the posture state of the limbs.
[0125] In step 208, the frame sequence of the sitting posture detection video is detected, and the face spatial features and the body posture features in different video frames are sequentially detected for face posture detection, body posture detection, and posture fusion detection.
[0126] As a more specific implementation of the embodiments of the present disclosure, when sequentially detecting the face spatial features and the body posture features in different video frames for face posture detection, body posture detection, and cross-modal posture detection, the following methods can be used, but are not limited to: obtaining at least one target face feature point and at least one target face reference line in the face spatial features, and obtaining at least one target body feature point and at least one target body reference line in the body posture features; according to the spatial topological relationship of the at least one target face feature point, performing face posture detection on the target detection object to generate a detection abnormality number; based on the at least one target body reference line, determining whether the posture angle of the target detection object exceeds an angle threshold, and according to the spatial topological relationship of the at least one target body feature point, performing body posture detection on the target detection object to generate a detection abnormality number; detecting the spatial relationship between the at least one target face reference line and the at least one target body reference line, and detecting whether the target body feature point is in the face region, to perform posture fusion detection on the target detection object to generate a detection abnormality number.
[0127] Specifically in step 208, the posture of the target detection object in different video frames is comprehensively and meticulously detected and analyzed, and the multi-dimensional detection method is used to accurately determine whether the sitting posture is abnormal.
[0128] In face posture detection, first, “at least one target face feature point and at least one target face reference line in the face spatial features” are obtained. Some of these face feature points (such as key points of eyes, nose tip, and mouth corner) can be target face feature points. The target face reference line can be set according to the structure characteristics of the face, such as the line connecting the two eyes, the line connecting the nose tip and the chin, etc. These points and lines can accurately reflect the key parts and basic outlines of the face.
[0129] The spatial relationships such as the relative positions, distances, and angles between these target face feature points are analyzed to determine whether the face posture is abnormal. For example, if the distance and angle between the two eyes are within a certain range under normal circumstances, when it is detected that these parameters exceed the preset range in a certain frame, it is determined that there is one face posture detection abnormality, and the detection abnormality number is increased accordingly. This process is based on the anatomical principles of the face and a large number of normal sample data to establish a normal spatial topological relationship model, which is used as a basis for judgment.
[0130] For limb posture detection, at least one target limb feature point and at least one target limb reference line in the limb posture feature need to be obtained. For example, the shoulder, elbow, wrist, hip, knee, ankle and other joints in the 17 human body limb feature points can be used as the target limb feature point, and the target limb reference line can be a straight line connecting these joints, such as the line connecting the shoulder peak and the elbow tip, the line connecting the hip center and the knee center, etc., which is used to describe the basic architecture of the limb. By calculating the angle between the target limb reference line and the reference coordinate system (such as the horizontal or vertical direction), it is determined whether the posture angle exceeds the preset angle threshold. At the same time, the spatial topological relationship between the target limb feature points is analyzed, such as the relative position change of the elbow and the wrist, the distance relationship between the leg joints, etc. If these angles or spatial relationships are abnormal, it is also determined as one limb posture detection abnormality, and the number of detection abnormalities is accumulated.
[0131] In the posture fusion detection stage, the overall posture information of the face and the limb is comprehensively considered. For example, under normal sitting posture, the face and the limb should maintain a certain coordination relationship, and the relative position and angle between the target face reference line and the target limb reference line also have a certain rule. When it is detected that these spatial relationships deviate obviously, it is considered as one abnormality. In addition, if the target limb feature point (such as the wrist key point) is found in the face area, this also does not conform to the characteristics of the normal sitting posture, and is also recorded as one detection abnormality, so as to comprehensively evaluate the rationality of the sitting posture.
[0132] In some embodiments of the present disclosure, a deep learning model can be used for posture detection. The deep learning model can automatically learn more complex and accurate face and limb posture features and their relationships, and has higher accuracy and robustness compared with traditional rule-based detection methods. At the same time, multi-sensor data fusion technology can be introduced to combine the image information obtained by the camera with the posture information obtained by other sensors (such as acceleration sensors, gyroscopes, etc.), to more comprehensively detect the posture from multiple dimensions, and further improve the accuracy and reliability of the detection.
[0133] In step 209, the face posture detection results, the limb posture detection results and the posture fusion detection results of different time sequences are fused to determine whether the sitting posture of the target detection object is an unhealthy sitting posture.
[0134] As a more specific implementation manner of the embodiments of the present disclosure, when the face posture detection results, the limb posture detection results and the posture fusion detection results of different time sequences are fused to determine whether the sitting posture of the target detection object is an unhealthy sitting posture, the following manner can be used but is not limited thereto: the number of abnormalities of the face posture detection results, the limb posture detection results and the posture fusion detection results is counted; and based on the number of abnormalities, it is determined whether the sitting posture of the target detection object is an unhealthy sitting posture.
[0135] Specifically, in step 209, the results of the previous multiple detection links are comprehensively analyzed to accurately determine whether the sitting posture of the target detection object belongs to an unhealthy sitting posture.
[0136] In step 208, the face spatial features and the limb posture features in different video frames have been subjected to face posture detection, limb posture detection, and posture fusion detection, and the corresponding detection abnormal times have been generated. These detection results reflect the posture abnormality of the target detection object at different times and in different dimensions, but the individual detection results cannot comprehensively and accurately determine the overall sitting posture state, and therefore need to be analyzed.
[0137] The "fusion of face posture detection results, limb posture detection results, and posture fusion detection results at different time sequences" refers to the detection at different time points of the sitting posture detection video. Since the video continuously records the posture changes of the target detection object, the detection results at each time point contain the posture information at that time. By integrating the detection results at different time points, the posture change trend of the target detection object within a period of time can be obtained. Specifically, the detection abnormal times generated by the face posture detection, the limb posture detection, and the posture fusion detection in each video frame are summarized in chronological order.
[0138] The "statistical abnormal times of the face posture detection results, the limb posture detection results, and the posture fusion detection results for detection abnormalities" refers to the statistics of the summarized detection abnormal times. For example, the face posture detection appears 5 times, the limb posture detection appears 3 times, and the posture fusion detection appears 2 times. The system will record these abnormal times respectively and further calculate the total abnormal times. This statistical process is the basis for subsequent judgment, which quantifies the degree of posture abnormality of the target detection object. When the total abnormal times obtained by the statistics exceed the set threshold, the system determines that the sitting posture of the target detection object is unhealthy; if the total abnormal times do not reach the threshold, it is considered that the current sitting posture is within the normal range.
[0139] In some other embodiments of the present disclosure, a machine learning algorithm can be used to more accurately model the relationship between the abnormal times and the sitting posture state. For example, a logistic regression model or a decision tree model can be used to train a large number of sample data with known sitting posture states (good or unhealthy) to learn the complex relationship between the abnormal times and the sitting posture state, so as to more accurately determine the sitting posture. In addition, other related information such as the duration of the target detection object maintaining a certain posture can be combined to further improve the judgment basis.
[0140] It should be noted that the embodiments of the present disclosure can include a plurality of steps, in order to facilitate description, these steps are numbered, but these numbers are not a limitation on the execution time slot, execution order between steps; these steps can be implemented in any order, the embodiments of the present disclosure do not limit this.
[0141] Corresponding to the above-mentioned method of sitting posture detection, the present disclosure also proposes a device for detecting sitting posture. Since the device embodiments of the present disclosure correspond to the above-mentioned method embodiments, for the details not disclosed in the device embodiments, please refer to the above-mentioned method embodiments, which will not be described in detail in the present disclosure.
[0142] Figure 3 The structural schematic diagram of a device for detecting sitting posture provided by the embodiments of the present disclosure is shown in Figure 3 , which includes:
[0143] The detection unit 31 is configured to obtain a sitting posture detection video of a target detection area, perform target detection processing on the sitting posture detection video, and generate human perception data of at least one detection object;
[0144] The first determination unit 32 is configured to perform cross-frame correlation analysis on the sitting posture detection video based on the human perception data, and track and determine a target detection object that needs to be detected for sitting posture among the at least one detection object;
[0145] The extraction unit 33 is configured to extract human features of the target detection object frame by frame in the sitting posture detection video, and obtain facial spatial features and limb posture features corresponding to each video frame;
[0146] The analysis unit 34 is configured to perform time sequence feature fusion analysis on the facial spatial features and the limb posture features, and judge whether the sitting posture of the target detection object is an unhealthy sitting posture.
[0147] Further, in a possible implementation manner of the present embodiment, as shown in Figure 4 , the detection unit 31 includes:
[0148] The first detection module 311 is configured to perform face recognition detection on each video frame in the sitting posture detection video respectively, determine a face detection area corresponding to different detection objects in each video frame, and generate a confidence degree corresponding to each face detection area;
[0149] The labeling module 312 is configured to label the identity of the face detection area according to the detection object to which the face detection area belongs, and generate an identity identifier corresponding to each face detection area;
[0150] The first generation module 313 is configured to generate human perception data of at least one detection object for the face detection area in each video frame and the confidence degree and the identity identifier corresponding thereto.
[0151] Further, in a possible implementation manner of the embodiment, as shown in Figure 4 the first determining unit 32 comprises:
[0152] The marking module 321 is configured to select a maximum face detection region in an initial frame of the sitting posture detection video based on the human perception data, and mark the maximum face detection region as a target face detection region.
[0153] The analysis module 322 is configured to perform cross-frame correlation analysis on the target face detection region in the sitting posture detection video according to an identity of the target face detection region, and obtain a disappearance number of the target face detection region.
[0154] The first determining module 323 is configured to determine that a detection object corresponding to the target face detection region is a target detection object when the disappearance number is less than a failure threshold.
[0155] Further, in a possible implementation manner of the embodiment, as shown in Figure 4 the device for sitting posture detection further comprises:
[0156] The second determining unit 35 is configured to terminate the cross-frame correlation analysis on the target face detection region when the disappearance number is greater than or equal to the failure threshold, and determine a target detection object that needs to be subjected to sitting posture detection in a subsequent video frame.
[0157] Further, in a possible implementation manner of the embodiment, as shown in Figure 4 the extraction unit 33 comprises:
[0158] The second determining module 331 is configured to traverse the sitting posture detection video according to a target identity of the target detection object, and determine a target video frame in which the target identity is located.
[0159] The second generation module 332 is configured to generate a face space feature and a body posture feature corresponding to each video frame based on the first number of face feature points and the second number of body feature points in each target video frame.
[0160] Further, in a possible implementation manner of the embodiment, as shown in Figure 4 the analysis unit 34 comprises:
[0161] The second detection module 341 is configured to perform face posture detection, body posture detection and posture fusion detection on the face space feature and the body posture feature in different video frames in sequence according to a frame sequence of the sitting posture detection video.
[0162] The judgment module 342 is configured to fuse the face posture detection result, the body posture detection result and the posture fusion detection result at different time sequences, and judge whether a sitting posture of the target detection object is an unhealthy sitting posture.
[0163] Further, in a possible implementation of the embodiment, the second detection module 341 is further configured to:
[0164] obtain at least one target facial feature point and at least one target facial reference line in the facial spatial feature, and obtain at least one target limb feature point and at least one target limb reference line in the limb posture feature;
[0165] perform facial posture detection on the target detection object according to the spatial topological relationship of the at least one target facial feature point to generate a detection abnormality number;
[0166] determine whether the posture angle of the target detection object exceeds an angle threshold based on the at least one target limb reference line, and perform limb posture detection on the target detection object according to the spatial topological relationship of the at least one target limb feature point to generate a detection abnormality number;
[0167] detect the spatial relationship between the at least one target facial reference line and the at least one target limb reference line, and detect whether the target limb feature point is in the facial region, to perform posture fusion detection on the target detection object to generate a detection abnormality number.
[0168] Further, in a possible implementation of the embodiment, the judging module 342 is further configured to:
[0169] count the abnormality number of the detection abnormality in the facial posture detection result, the limb posture detection result, and the posture fusion detection result;
[0170] determine whether the sitting posture of the target detection object is an unhealthy sitting posture based on the abnormality number.
[0171] It should be noted that the foregoing explanation and description of the method embodiment also apply to the device of the present embodiment, and the principle is the same, which will not be limited in the present embodiment.
[0172] According to the embodiments of the present disclosure, the present disclosure further provides an electronic device, a readable storage medium, and a computer program product.
[0173] Figure 5A schematic block diagram of an example electronic device 400 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0174] like Figure 5 As shown, the electronic device 400 includes a computing unit 401, which can perform various appropriate actions and processes based on a computer program stored in ROM (Read-Only Memory) 402 or a computer program loaded from storage unit 408 into RAM (Random Access Memory) 403. The RAM 403 can also store various programs and data required for the operation of the electronic device 400. The computing unit 401, ROM 402, and RAM 403 are interconnected via a bus 404. An I / O (Input / Output) interface 405 is also connected to the bus 404.
[0175] Multiple components in electronic device 400 are connected to I / O interface 405, including: input unit 406, such as keyboard, mouse, etc.; output unit 407, such as various types of displays, speakers, etc.; storage unit 408, such as disk, optical disk, etc.; and communication unit 409, such as network card, modem, wireless transceiver, etc. Communication unit 409 allows electronic device 400 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0176] The computing unit 401 can be various general and / or special purpose processing components with processing and computing capabilities. Some examples of the computing unit 401 include, but are not limited to, a CPU (Central Processing Unit), a GPU (Graphic Processing Units), various special-purpose AI (Artificial Intelligence) computing chips, various computing units running machine learning model algorithms, a DSP (Digital Signal Processor), and any appropriate processor, controller, microcontroller, etc. The computing unit 401 performs various methods and processes described above, such as the method of sitting posture detection. For example, in some embodiments, the method of sitting posture detection can be implemented as a computer software program, which is tangibly embodied in a machine-readable medium, such as the storage unit 408. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 400 via the ROM 402 and / or the communication unit 409. When the computer program is loaded onto the RAM 403 and executed by the computing unit 401, one or more steps of the methods described above can be performed. Alternatively, in other embodiments, the computing unit 401 can be configured to perform the aforementioned method of sitting posture detection by any other appropriate means, such as by means of firmware.
[0177] Various implementations of the systems and techniques described above herein can be realized in digital electronic circuitry, integrated circuitry, a Field Programmable Gate Array (FPGA), an Application-Specific Integrated Circuit (ASIC), an Application Specific Standard Product (ASSP), a System on Chip (SOC), a Complex Programmable Logic Device (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include implementation in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be special or general purpose, coupled to receive data and instructions from, and to transmit data and instructions to, a storage system, at least one input device, and at least one output device.
[0178] Program code for carrying out methods of the present disclosure can be written in any combination of one or more programming languages. The program code can be provided to a processor or controller of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the program code, when executed by the processor or controller, produces the functions / operations specified in the flowcharts and / or the block diagrams. The program code can be executed entirely on a machine, partially on a machine, partially on a machine as a stand-alone software package, partially on a machine and partially on a remote machine or entirely on a remote machine or server.
[0179] In the context of the present disclosure, a machine-readable medium can be a tangible medium that contains or stores a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include but is not limited to an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium would include one or more lines of electrical connections, portable computer disks, hard disk drives, RAM, ROM, EPROM (Electrically Programmable Read-Only-Memory), or flash memory, fiber optics, CD-ROM (Compact Disc Read-Only Memory), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0180] To provide for interaction with a user, the systems and techniques described here can be implemented on a computer having a display device (e.g., a CRT (Cathode Ray Tube) or LCD (Liquid Crystal Display) monitor) for displaying information to the user and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form, including acoustic, speech, or tactile input.
[0181] The systems and techniques described herein can be implemented in a computing system that includes a back end component (e.g., as a data server), or that includes a middleware component (e.g., an application server), or that includes a front end component (e.g., a user computer having a graphical user interface or a Web browser through which a user can interact with an implementation of the systems and techniques described herein), or any combination of such back end, middleware, or front end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a LAN (Local Area Network), a WAN (Wide Area Network), the Internet, and a blockchain network.
[0182] The computer system can include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The relationship of client and server is generally established using computer programs running on the respective computers and having a client-server relationship to each other. A server can be a cloud server, also known as a cloud computing server or cloud host, which is a host product in the cloud computing service system, to solve the defects of large management difficulty and weak business scalability in traditional physical host and VPS (Virtual Private Server, or VPS for short) services. The server can also be a server of a distributed system, or a server combined with a blockchain.
[0183] It should be noted that artificial intelligence is a discipline that studies enabling computers to simulate some thinking processes and intelligent behaviors of humans (such as learning, reasoning, thinking, planning, etc.), and has both hardware and software technologies. Artificial intelligence hardware technologies generally include technologies such as sensors, special artificial intelligence chips, cloud computing, distributed storage, big data processing, etc.; artificial intelligence software technologies mainly include computer vision technology, speech recognition technology, natural language processing technology, and machine learning / deep learning, big data processing technology, knowledge graph technology, etc.
[0184] The first, second, and various other numerical designations involved in the present disclosure are only used for differentiation for convenience of description, and do not limit the scope of the embodiments of the present disclosure, nor represent a sequence.
[0185] At least one of the present disclosure can also be described as one or more, multiple can be two, three, four or more, the present disclosure does not make restrictions. In the embodiments of the present disclosure, for a technical feature, the technical features in the technical feature are distinguished by "first", "second", "third", "A", "B", "C" and "D" and the like. The technical features described by "first", "second", "third", "A", "B", "C" and "D" have no order or size order.
[0186] It should be understood that the steps shown above can be reordered, added or deleted. For example, the steps described in the present disclosure can be executed in parallel, sequentially or in different order, as long as the desired results of the technical solutions disclosed in the present disclosure can be achieved, which is not limited herein.
[0187] The above specific embodiments do not constitute a limitation on the protection scope of the present disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modifications, equivalent replacements and improvements made within the spirit and principles of the present disclosure shall be included in the protection scope of the present disclosure.
Claims
1. A method for detecting sitting posture, characterized in that, include: Acquire a sitting posture detection video of the target detection area, perform target detection processing on the sitting posture detection video, and generate human perception data of at least one detection object; Based on the human perception data, cross-frame correlation analysis is performed on the sitting posture detection video to track and determine the target detection object that needs to be detected in the at least one detection object; In the sitting posture detection video, the human body features of the target detection object are extracted frame by frame to obtain the facial spatial features and limb posture features corresponding to each video frame; The facial spatial features and the limb posture features are subjected to temporal feature fusion analysis to determine whether the sitting posture of the target detection object is poor sitting posture.
2. The method for detecting sitting posture according to claim 1, characterized in that, The process of acquiring a seated posture detection video of the target detection area, performing target detection processing on the seated posture detection video, and generating human perception data of at least one detected object includes: Face recognition detection is performed on each video frame in the sitting posture detection video to determine the face detection region corresponding to different detection objects in each video frame, and the confidence level corresponding to each face detection region is generated. The face detection regions are labeled with identities according to the detection objects to which they belong, and an identity identifier is generated for each face detection region. For each video frame, generate human perception data for at least one detected object based on the face detection region and its corresponding credibility and identity identifier.
3. The method for detecting sitting posture according to claim 1, characterized in that, The step of performing cross-frame correlation analysis on the sitting posture detection video based on the human perception data, and tracking and determining the target detection object that needs to be detected among the at least one detection object, includes: Based on the human perception data, the largest face detection region in the initial frame of the sitting posture detection video is selected, and the largest face detection region is marked as the target face detection region. Based on the identity identifier of the target face detection region, cross-frame correlation analysis is performed on the target face detection region in the sitting posture detection video to obtain the number of times the target face detection region disappears; If the number of disappearances is less than the failure threshold, then the detection object corresponding to the target face detection region is determined to be the target detection object.
4. The method for detecting sitting posture according to claim 3, characterized in that, The method for detecting sitting posture also includes: If the number of disappearances is greater than or equal to the failure threshold, the cross-frame correlation analysis of the target face detection region is terminated, and the target detection object that needs to be detected for sitting posture is re-determined in subsequent video frames.
5. The method for detecting sitting posture according to claim 1, characterized in that, The step of extracting the human body features of the target detection object frame by frame from the seated posture detection video to obtain the facial spatial features and limb posture features corresponding to each video frame includes: Based on the target identity identifier of the target detection object, the sitting posture detection video is traversed to determine the target video frame where the target identity identifier is located; Based on a first number of facial feature points and a second number of limb feature points in each target video frame, facial spatial features and limb posture features corresponding to each video frame are generated respectively.
6. The method for detecting sitting posture according to claim 1, characterized in that, The step of performing temporal feature fusion analysis on the facial spatial features and the limb posture features to determine whether the sitting posture of the target detection object is poor sitting posture includes: According to the frame sequence of the sitting posture detection video, facial posture detection, limb posture detection and posture fusion detection are performed sequentially on the facial spatial features and limb posture features in different video frames. By integrating facial posture detection results, limb posture detection results, and posture fusion detection results from different time sequences, it is determined whether the sitting posture of the target detection object is an improper sitting posture.
7. The method for detecting sitting posture according to claim 6, characterized in that, The sequential processing of facial spatial features and limb pose features in different video frames includes facial pose detection, limb pose detection, and cross-modal pose detection, comprising: Obtain at least one target facial feature point and at least one target facial baseline from the facial spatial features, and obtain at least one target limb feature point and at least one target limb baseline from the limb posture features; Based on the spatial topological relationship of the at least one target facial feature point, perform facial pose detection on the target detection object to generate the number of anomaly detections; Based on the at least one target limb baseline, determine whether the pose angle of the target detection object exceeds the angle threshold, and perform limb pose detection on the target detection object to generate the number of detection anomalies according to the spatial topological relationship of at least one target limb feature point; The spatial relationship between the at least one target facial baseline and the at least one target limb baseline is detected, and whether the target limb feature points are located in the facial region is detected, so as to perform pose fusion detection on the target detection object to generate the number of detection anomalies.
8. The method for detecting sitting posture according to claim 7, characterized in that, The method of fusing facial pose detection results, limb pose detection results, and pose fusion detection results from different time sequences to determine whether the sitting posture of the target detection object is poor posture includes: The number of abnormalities detected is calculated based on the facial pose detection results, the limb pose detection results, and the pose fusion detection results. Based on the number of anomalies, it is determined whether the sitting posture of the target detection object is poor sitting posture.
9. An electronic device, characterized in that, include: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, which, when executed by the at least one processor, enables the at least one processor to perform the sitting posture detection method according to any one of claims 1-8.
10. A non-transitory computer-readable storage medium storing computer instructions, characterized in that, The computer instructions are used to cause the computer to perform the sitting posture detection method according to any one of claims 1-8.