Sitting Posture Recognition Method, Device and Storage Medium
By acquiring and processing multi-frame character images, building a feature matrix and inputting a seating posture recognition model, the problem of sitting posture recognition in the prior art relying on wearable devices is solved, and efficient sitting posture recognition without equipment is achieved.
Patent Information
- Application Number
- CN202210474105.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-29
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2042-04-29
AI Technical Summary
In the prior art, sitting posture recognition requires a wearable device, which is costly and has poor user experience, and can only limit the user's range of activities.
By obtaining multi-frame character images, extracting key point information, building feature matrix, and inputting a pre-trained sitting posture recognition model to identify the character's sitting posture without relying on wearable devices.
It realizes effective recognition of sitting posture without wearing devices, reduces recognition costs, and expands user activity range.
Smart Images

Figure CN114821652B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of image recognition, and in particular, to a sitting posture recognition method, apparatus, and storage medium. Background Art
[0002] In a home learning scenario or a face-to-face course scenario, accurate recognition of a user's sitting posture helps analyze the student's learning concentration, and further helps improve the content presentation method of a learning machine or the knowledge teaching method of a teacher, making it more efficient and user-friendly.
[0003] In related technologies, the detection of sitting postures needs to be based on wearable devices. However, the method based on wearable detection requires the user to wear corresponding devices or install corresponding sensors, which is costly and has a poor user experience, and can only limit the user's activity range to places where these devices are installed. Summary of the Invention
[0004] To solve the problems existing in related technologies, the present disclosure provides a sitting posture recognition method, apparatus, and storage medium.
[0005] To achieve the above object, the first aspect of the present disclosure provides a sitting posture recognition method, the method including:
[0006] Obtain multiple frames of human images;
[0007] Extract features from each frame of the human image to obtain key point information;
[0008] Based on the key point information, determine a feature matrix;
[0009] Input the feature matrix into a pre-trained sitting posture recognition model to obtain the probabilities of various sitting postures corresponding to the human image.
[0010] Optionally, the determining the feature matrix based on the key point information includes:
[0011] Filter the key point information to obtain first key point information with redundant key points removed;
[0012] Based on the first key point information, determine the feature matrix.
[0013] Optionally, the determining the feature matrix based on the key point information includes:
[0014] Perform position encoding on the key point information to obtain second key point information including position encoding information;
[0015] Based on the second key point information, determine the feature matrix.
[0016] Optionally, determining the feature matrix based on the key point information includes:
[0017] Normalize the key point information to obtain the normalized third key point information;
[0018] Determine the feature matrix based on the third key point information.
[0019] Optionally, determining the feature matrix based on the key point information includes:
[0020] Perform feature stitching on the key point information of each frame of the person image to determine the feature matrix.
[0021] Optionally, the sitting posture recognition model includes an attention module and a sitting posture classification module. Inputting the feature matrix into the pre-trained sitting posture recognition model to obtain the probabilities of each sitting posture corresponding to the person image includes:
[0022] Input the feature matrix into the attention module to determine the attention scores of each sub-feature in the feature matrix;
[0023] Process the feature matrix based on the attention scores to obtain the first feature matrix;
[0024] Input the first feature matrix into the sitting posture classification module to obtain the probabilities of each sitting posture corresponding to the person image.
[0025] Optionally, the training of the sitting posture recognition model includes:
[0026] Obtain sample person images, where the sample person images include sitting posture type annotation information, and each sample person image includes multiple frames of images;
[0027] Extract features from each frame of the sample person image to obtain sample key point information;
[0028] Determine a sample feature matrix based on the sample key point information;
[0029] Input the sample feature matrix into the initial sitting posture recognition model to determine the sample probabilities of each sitting posture corresponding to the sample person image;
[0030] Adjust the parameters of the initial sitting posture recognition model according to the sample probabilities and the sitting posture type annotation information.
[0031] Optionally, the initial sitting posture recognition model includes an attention module and a sitting posture classification module. Inputting the sample feature matrix into the initial sitting posture recognition model to obtain the sample probabilities of each sitting posture corresponding to the sample person image includes:
[0032] Input the sample feature matrix into the attention module to obtain the sample attention scores of each sub - feature in the sample feature matrix;
[0033] Process the sample feature matrix based on the sample attention scores to obtain a first sample feature matrix;
[0034] Input the first sample feature matrix into the sitting - posture classification module to obtain the sample probabilities of each sitting - posture corresponding to the sample person image;
[0035] Adjusting the parameters of the initial sitting - posture recognition model according to the sample probabilities and the sitting - posture type annotation information includes:
[0036] Adjust the parameters of the attention module and the sitting - posture classification module according to the sample probabilities and the sitting - posture type annotation information.
[0037] A second aspect of the present disclosure provides a sitting - posture recognition device, the device includes:
[0038] An acquisition device, configured to acquire multiple frames of person images;
[0039] An extraction module, configured to extract features from each frame of the person image to obtain key - point information;
[0040] A determination module, configured to determine a feature matrix based on the key - point information;
[0041] A recognition module, configured to input the feature matrix into a pre - trained sitting - posture recognition model to obtain the probabilities of each sitting - posture corresponding to the person image.
[0042] A third aspect of the present disclosure provides a non - transitory computer - readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the steps of the method described in any item of the first aspect of the present disclosure are implemented.
[0043] Through the above - mentioned technical solution, by acquiring multiple frames of person images, extracting key - point information in the person images, and constructing a feature matrix corresponding to the person image, the sitting - posture recognition model can recognize the sitting - posture of the person in the person image based on the feature matrix, without relying on wearable devices, which can effectively achieve sitting - posture recognition and reduce the cost of sitting - posture recognition.
[0044] Other features and advantages of the present disclosure will be described in detail in the subsequent specific implementation section. Description of the Drawings
[0045] The accompanying drawings are used to provide a further understanding of the present disclosure and constitute a part of the specification. Together with the following detailed description, they are used to explain the present disclosure, but do not limit the present disclosure. In the accompanying drawings:
[0046] Figure 1 is a flowchart of a sitting posture recognition method shown according to an exemplary embodiment;
[0047] Figure 2 is a flowchart of a training method of a sitting posture recognition model shown according to an exemplary embodiment;
[0048] Figure 3 is another flowchart of a sitting posture recognition method shown according to an exemplary embodiment;
[0049] Figure 4 is a block diagram of a sitting posture recognition device shown according to an exemplary embodiment;
[0050] Figure 5 is a block diagram of an electronic device shown according to an exemplary embodiment;
[0051] Figure 6 is a block diagram of another electronic device shown according to an exemplary embodiment. Detailed Description of the Invention
[0052] The following provides a detailed description of the specific embodiments of the present disclosure with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are only used to illustrate and explain the present disclosure, and do not limit the present disclosure.
[0053] Figure 1 is a flowchart of a sitting posture recognition method shown according to an exemplary embodiment. This method can be applied to electronic devices with information processing capabilities such as mobile phones, personal computers, servers, etc. The present disclosure does not make specific limitations in this regard. For example, Figure 1 as shown, the method includes:
[0054] S101. Obtain multiple frames of human images.
[0055] Among them, the number of frames of the human image can be pre-calibrated. For example, the human image can be composed of 6 consecutive frames or 12 frames of images collected by a camera. The time interval between two consecutive frames of images can be 100 milliseconds to 500 milliseconds. The specific number of frames and time interval can be calibrated according to the specific situation of the sitting posture to be recognized. For example, in the case of recognizing students' rough play actions, the number of frames and time interval can be made smaller, such as 3 frames and 100 milliseconds respectively, so as to capture useful information and exclude useless information.
[0056] S102. Extract features from each frame of the human image to obtain key point information.
[0057] Among them, the key point information includes the position information of different parts of the human body. For example, the position information of the left eye, the position information of the nose, and so on.
[0058] In a possible implementation manner, the key point information can be obtained by OpenPose. Or, in some other possible implementation manners, if higher requirements are imposed on the accuracy of the key point information, a top-down method can be adopted, that is, first detect the area where the human body is located, and then perform key point regression on the human body in this area to obtain the key point information.
[0059] S103. Determine a feature matrix based on the key point information.
[0060] Among them, the feature matrix can be obtained after performing operations such as normalizing the key point information, position encoding, and feature fusion. The specific operation execution method will be described in detail in the optional implementation manners below and will not be elaborated here.
[0061] S104. Input the feature matrix into a pre-trained sitting posture recognition model to obtain the probabilities of various sitting postures corresponding to the human image.
[0062] Among them, in specific implementation, considering the health of students and teaching, the sitting postures can include categories such as hunchback, tilted head, chin in hand, lying on the desk, and concentration. For example, the probabilities of various sitting postures corresponding to the human image can be expressed as {hunchback: 0.8; tilted head: 0.1; chin in hand: 0.1; lying on the desk: 0; concentration: 0}. The present disclosure does not make specific limitations on the specific sitting posture types.
[0063] Further, in the case where the probability that the sitting posture of the person in the human image is in a bad sitting posture is greater than a preset threshold, the sound generating unit can be controlled to send a prompt sound effect to remind the user to adjust the sitting posture.
[0064] In the embodiments of the present disclosure, by acquiring multiple frames of human images, extracting the key point information in the human images, and constructing a feature matrix corresponding to the human images, the sitting posture recognition model can recognize the sitting posture of the person in the human image based on the feature matrix without relying on wearable devices, effectively realizing sitting posture recognition and reducing the cost of sitting posture recognition.
[0065] In some optional embodiments, the determining a feature matrix based on the key point information includes:
[0066] Screen the key point information to obtain first key point information with redundant key points removed;
[0067] Determine the feature matrix based on the first key point information.
[0068] Those skilled in the art should be aware that conventional detectors for extracting key points will output a relatively large number of human key point information. However, some key point information, such as the key points of the lower body and the key points of the left and right ears, may be less helpful for sitting posture recognition. Therefore, through the above solution, the key point information corresponding to redundant key points can be eliminated. In a possible implementation manner, only 9 key point information may be retained, that is, the key point information corresponding to the left eye, right eye, nose, left shoulder, right shoulder, left elbow, right elbow, left wrist, and right wrist.
[0069] By adopting this solution, by screening the key point information and eliminating the key point information corresponding to redundant key points that are less helpful for sitting posture recognition, the dimension and data volume of the feature matrix can be greatly reduced, and the processing speed of the sitting posture recognition model can be effectively improved.
[0070] In still other alternative embodiments, the determining the feature matrix based on the key point information includes:
[0071] Perform position encoding on the key point information to obtain second key point information including position encoding information; determine the feature matrix based on the second key point information.
[0072] Exemplarily, if only the above 9 key point information is retained and position encoding is performed in the order of the left eye, right eye, nose, left shoulder, right shoulder, left elbow, right elbow, left wrist, and right wrist, the position encoding information corresponding to the key point information of the nose can be expressed as [0, 0, 1, 0, 0, 0, 0, 0, 0], and the position encoding information corresponding to the left shoulder can be expressed as [0, 0, 0, 1, 0, 0, 0, 0, 0].
[0073] Another exemplarily, encoded in the same order, the position encoding information corresponding to the key point information of the left eye can be expressed as [0, 0, 0, 1], and the position encoding information corresponding to the key point information of the right wrist can be expressed as [1, 0, 0, 1], that is, the value corresponding to the four-bit binary number corresponding to the position encoding is the corresponding sorting. Alternatively, other position encoding methods can also be used to perform position encoding on the key point information, and the present disclosure does not limit the specific position encoding method.
[0074] Those skilled in the art should understand that when performing sitting posture recognition based on key point information, if the input key point information is disordered, it is difficult for the sitting posture recognition model to learn regular features.
[0075] With the above solution, by adding position encoding information to the key point information of each key point, the sitting posture recognition model can sort the positions according to the position encoding information, so as to be able to pay attention to the correlation between key points, which is more conducive to the learning of the sitting posture recognition model, and enables the sitting posture recognition model to more accurately recognize the sitting posture.
[0076] In some further optional embodiments, determining the feature matrix based on the key point information includes:
[0077] Normalize the key point information to obtain the normalized third key point information; determine the feature matrix based on the third key point information.
[0078] It can be understood that the range of the x coordinate of the key point in the key point information is [0, W], and the range of the y coordinate is [0, H], where W and H are the width and height of the image. However, the values of W and H are both relatively large, and the data volume gap between data is large. During the training of the sitting posture recognition model, the situation of gradient explosion may occur due to the chain reaction. Therefore, the key point information can be normalized so that the value range in the key point information is within [0, 1].
[0079] Specifically, normalization can be performed through the following formula: where pts represents the set of all key point information, such as the 9 key point information corresponding to the above 9 key points, and pt can be any key point information in the set of this key point information, x pt can represent the actual x coordinate of any key point in the set of this key point information, x pt* can represent the normalized x coordinate of any key point in the set of this key point information, y pt can represent the actual y coordinate of any key point in the set of this key point information, y pt* can represent the normalized y coordinate of any key point in the set of this key point information.
[0080] With the above solution, before constructing the feature matrix, normalizing the key point information reduces the data quantity set gap in each key point information, can effectively avoid the situation of gradient explosion during the training of the sitting posture recognition model, and improves the sitting posture discrimination accuracy of the sitting posture recognition model.
[0081] In some further optional embodiments, determining the feature matrix based on the key point information includes:
[0082] Perform a feature splicing operation on the key point information of each frame of the person image to determine the feature matrix.
[0083] Among them, the key point information corresponding to each key point of multiple frames of images can be feature - spliced to obtain a unique feature including all the key point information contained in the multiple frames of images.
[0084] Exemplarily, before splicing, the above - mentioned redundant key point elimination operation and position encoding operation can be performed on the key point information corresponding to each frame of the image.
[0085] Specifically, when the key point information includes the x - coordinate, y - coordinate, and confidence level, after performing the redundant key point elimination operation, a feature matrix with a feature dimension of [N, 3] can be obtained, where 3 corresponds to the dimensions of the x - coordinate, y - coordinate, and confidence level in the key point information, and N represents the number of key points, for example, it can be 9. When performing position encoding in the form of [0, 0, 1, 0, 0, 0, 0, 0, 0], a feature matrix with a feature dimension of [N, 3 + N] can be obtained. After further performing the above - mentioned feature splicing, a feature matrix with a dimension of [T, N, 3 + N] can be obtained.
[0086] By adopting the above - mentioned scheme, the features corresponding to multiple frames of images can be fused into a feature matrix, which includes all the key point information of the multiple frames of images. Inputting this feature matrix into the sitting - posture classification model can obtain the sitting - posture detection result corresponding to the multiple frames of images.
[0087] In some possible implementation manners, the sitting - posture recognition model includes an attention module and a sitting - posture classification module. The obtaining the probabilities of each sitting - posture corresponding to the person image by inputting the feature matrix into the pre - trained sitting - posture recognition model includes:
[0088] Inputting the feature matrix into the attention module to determine the attention scores of each sub - feature in the feature matrix; processing the feature matrix based on the attention scores to obtain a first feature matrix; inputting the first feature matrix into the sitting - posture classification module to obtain the probabilities of each sitting - posture corresponding to the person image.
[0089] It can be understood that the feature matrix can be obtained after operations such as the above - mentioned redundant key point elimination, normalization, position encoding, and feature splicing.
[0090] Specifically, the attention module includes a Global Average Pooling (GAP) layer and Fully connected layers (FC). Based on the [T, N, N + 3]-dimensional feature matrix obtained after the above processing, after passing through the global average pooling layer, a feature of dimension [1, 1N] can be obtained, and then the attention score can be obtained after full connection. Multiplying the attention score by the feature matrix input to the sitting posture recognition model can obtain the final feature.
[0091] Adopting the above solution, by designing the attention module, processing the feature matrix, assigning different weights to the features of multiple frames of images, and finally inputting the obtained features into the classification branch, the sitting postures of students can be recognized more accurately.
[0092] Figure 2 It is a flowchart of a training method of a sitting posture recognition model shown according to an exemplary embodiment, as Figure 2 shown, the method includes the steps:
[0093] S201. Obtain sample human images.
[0094] The sample human images include sitting posture type annotation information, and each sample human image includes multiple frames of images.
[0095] Among them, there can be multiple sample human images, and each sample human image can correspond to a sitting posture type annotation information. The sitting posture type annotation information corresponding to a sample human image can be, for example, any one of hunchback, head tilt, chin on hand, lying on the table, and concentration.
[0096] S202. Extract features from each frame of the sample human image to obtain sample key point information.
[0097] S203. Determine a sample feature matrix based on the sample key point information.
[0098] Among them, the sample feature matrix can be obtained after the redundant key point removal, normalization, position encoding, and feature splicing operations described in the above sitting posture recognition method.
[0099] It can be understood that training by performing position encoding on the key points of each frame in a certain order can enable the model to learn the relative spatial information of other positions based on the starting position, and thus the trained sitting posture recognition model can perform sitting posture recognition more accurately.
[0100] S204. Input the sample feature matrix into the initial sitting posture recognition model to determine the sample probabilities of the sitting postures corresponding to the sample human images.
[0101] S205. Adjust the parameters of the initial sitting posture recognition model according to the sample probability and the sitting posture type annotation information.
[0102] Among them, steps S201 to S205 can be executed multiple times until the number of executions meets a preset threshold, or stop executing when the parameters of the initial sitting posture recognition model converge. Then, a trained sitting posture recognition model is obtained.
[0103] Adopting the above solution, it is possible to train a sitting posture recognition model based on a sample person image including sitting posture type annotation information to adjust the parameters in the sitting posture recognition model, so that the trained sitting posture recognition model can accurately recognize the sitting posture types corresponding to multiple frames of images.
[0104] In some optional embodiments, the initial sitting posture recognition model includes an attention module and a sitting posture classification module. The obtaining of the sample probabilities of each sitting posture corresponding to the sample person image by inputting the sample feature matrix into the initial sitting posture recognition model includes:
[0105] Input the sample feature matrix into the attention module to obtain the sample attention scores of each sub-feature in the sample feature matrix;
[0106] Based on the sample attention scores, process the sample feature matrix to obtain a first sample feature matrix; input the first sample feature matrix into the sitting posture classification module to obtain the sample probabilities of each sitting posture corresponding to the sample person image;
[0107] The adjusting of the parameters of the initial sitting posture recognition model according to the sample probability and the sitting posture type annotation information includes:
[0108] Adjust the parameters of the attention module and the sitting posture classification module according to the sample probability and the sitting posture type annotation information.
[0109] Among them, the attention module includes a global average pooling layer and a fully connected layer. If the sample feature matrix is a feature matrix of [T, N, N + 3] dimensions, after passing through the global average pooling layer, a feature of [1, 1N] dimensions can be obtained, and after passing through the fully connected layer, the attention scores can be obtained. Multiply the attention scores by the feature matrix input into the sitting posture recognition model to obtain the final feature.
[0110] Adopting the above solution, since the above processes are all differentiable, the parameters of the attention module and the sitting posture classification module can be adjusted by means of backpropagation, so that the attention scores output by the attention module are more accurate, can more effectively exclude the interference of irrelevant frames, and make the recognition more accurate.
[0111] Based on the same inventive concept, in order to enable those skilled in the art to better understand the technical solutions provided by the present disclosure, the present disclosure provides a flowchart of a sitting posture recognition method shown in accordance with an exemplary embodiment as follows Figure 3 shown, the method includes: Figure 3 as
[0112] S301. Collect continuous T frames of human images.
[0113] S302. Extract features from each frame of human image to obtain key point information.
[0114] The key point information of each key point may include the x coordinate, y coordinate, and confidence level of the key point.
[0115] S303. Screen the key point information to obtain first key point information with redundant key points removed.
[0116] Exemplarily, if the key point information can include the x coordinate, y coordinate, and confidence level of the key point, and there are N key points after removing redundant key points, the dimension of the first key point information can be [N, 3].
[0117] S304. Perform position encoding on the first key point information to obtain second key point information.
[0118] In the case of performing position encoding in the form of [0, 0, 1, 0, 0, 0, 0, 0, 0], the second key point information can be [N, 3 + N]. Or, in the case of performing position encoding in the form of [0, 0, 0, 1], the dimension of the second key point information can be [N, 3 + 4].
[0119] S305. Normalize the second key point information to obtain third key point information.
[0120] S306. Perform a feature splicing operation on the third key point information of each frame of human image to obtain a feature matrix.
[0121] In the case where the dimension of the second key point is [N, 3 + N], a feature matrix with the dimension of [T, N, 3 + N] can be obtained.
[0122] S307. Input the feature matrix into the attention module to obtain attention scores.
[0123] S308. Obtain a target feature matrix according to the attention scores.
[0124] S309. Input the target feature matrix into the sitting posture classification module to obtain the probabilities of each sitting posture corresponding to N frames of human images.
[0125] It is understandable that the execution order of the above steps is only exemplary. For example, the normalization of the key point information in step S305 can be performed before removing redundant key points or before position encoding.
[0126] Based on Figure 3 the flowchart of a sitting posture recognition method shown, the training method of the sitting posture recognition model is similar to the above sitting posture recognition method, except that the person image also has pre-annotated sitting posture type annotation information, and steps for loss calculation and backpropagation of the loss are added, which will not be elaborated here.
[0127] Figure 4 is a schematic diagram of a sitting posture recognition device 40 shown according to an exemplary embodiment. As Figure 4 shown, the device 40 includes:
[0128] An acquisition module 41, configured to acquire multiple frames of person images;
[0129] An extraction module 42, configured to perform feature extraction on each frame of the person image to obtain key point information;
[0130] A determination module 43, configured to determine a feature matrix based on the key point information;
[0131] An identification module 44, configured to input the feature matrix into a pre-trained sitting posture recognition model to obtain the probabilities of various sitting postures corresponding to the person image.
[0132] Optionally, the determination module is specifically configured to:
[0133] Filter the key point information to obtain first key point information with redundant key points removed;
[0134] Determine the feature matrix based on the first key point information.
[0135] Optionally, the determination module 43 is further configured to:
[0136] Perform position encoding on the key point information to obtain second key point information including position encoding information;
[0137] Determine the feature matrix based on the second key point information.
[0138] Optionally, the determination module 43 is further configured to:
[0139] Normalize the key point information to obtain third key point information after normalization;
[0140] Determine the feature matrix based on the third key point information.
[0141] Optionally, the determining module 43 is further configured to:
[0142] Perform a feature splicing operation on the key point information of each frame of the human image to determine the feature matrix.
[0143] Optionally, the sitting posture recognition model includes an attention module and a sitting posture classification module, and the recognition module is further configured to:
[0144] Input the feature matrix into the attention module to determine the attention scores of each sub-feature in the feature matrix;
[0145] Process the feature matrix based on the attention scores to obtain a first feature matrix;
[0146] Input the first feature matrix into the sitting posture classification module to obtain the probabilities of each sitting posture corresponding to the human image.
[0147] Optionally, the device 40 is further configured to:
[0148] Obtain a sample human image, where the sample human image includes sitting posture type annotation information, and each sample human image includes multiple frames of images;
[0149] Extract features from each frame of the sample human image to obtain sample key point information;
[0150] Determine a sample feature matrix based on the sample key point information;
[0151] Input the sample feature matrix into the initial sitting posture recognition model to determine the sample probabilities of each sitting posture corresponding to the sample human image;
[0152] Adjust the parameters of the initial sitting posture recognition model according to the sample probabilities and the sitting posture type annotation information.
[0153] Optionally, the device 40 is further configured to:
[0154] Input the sample feature matrix into the attention module to obtain the sample attention scores of each sub-feature in the sample feature matrix;
[0155] Process the sample feature matrix based on the sample attention scores to obtain a first sample feature matrix;
[0156] Input the first sample feature matrix into the sitting posture classification module to obtain the sample probabilities of each sitting posture corresponding to the sample human image;
[0157] Adjusting the parameters of the initial sitting posture recognition model according to the sample probability and the sitting posture type annotation information includes:
[0158] Adjusting the parameters of the attention module and the sitting posture classification module according to the sample probability and the sitting posture type annotation information.
[0159] Regarding the device in the above embodiments, the specific manners in which each module performs operations have been described in detail in the embodiments related to the method, and will not be elaborated herein.
[0160] Figure 5 is a block diagram of an electronic device 500 shown according to an exemplary embodiment. As Figure 5 shown, the electronic device 500 may include: a processor 501, a memory 502. The electronic device 500 may further include one or more of a multimedia component 503, an input / output (I / O) interface 504, and a communication component 505.
[0161] Among them, the processor 501 is used to control the overall operation of the electronic device 500 to complete all or part of the steps in the above-mentioned sitting posture recognition method. The memory 502 is used to store various types of data to support the operation of the electronic device 500. These data may include, for example, instructions for any application or method operating on the electronic device 500, as well as application-related data, such as a person image, a sample person image, and so on. The memory 502 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random Access Memory (SRAM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Erasable Programmable Read-Only Memory (EPROM), Programmable Read-Only Memory (PROM), Read-Only Memory (ROM), magnetic memory, flash memory, a magnetic disk, or an optical disc. The multimedia component 503 may include a screen and an audio component. Among them, the screen may be a touch screen, and the audio component is used to output and / or input an audio signal. For example, the audio component may include a microphone for receiving an external audio signal. The received audio signal may be further stored in the memory 502 or sent through the communication component 505. The audio component further includes at least one speaker for outputting an audio signal. The I / O interface 504 provides an interface between the processor 501 and other interface modules. The above-mentioned other interface modules may be a keyboard, a mouse, buttons, etc. These buttons may be virtual buttons or physical buttons. The communication component 505 is used for wired or wireless communication between the electronic device 500 and other devices. Wireless communication, such as Wi-Fi, Bluetooth, Near Field Communication (NFC), 2G, 3G, 4G, NB-IOT, eMTC, or other 5G, etc., or a combination of one or more of them, is not limited herein. Therefore, the corresponding communication component 505 may include: a Wi-Fi module, a Bluetooth module, an NFC module, and so on.
[0162] In an exemplary embodiment, the electronic device 500 can be implemented by one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components for performing the above-mentioned sitting posture recognition method.
[0163] In another exemplary embodiment, a computer-readable storage medium including program instructions is also provided. When the program instructions are executed by a processor, the steps of the above-mentioned sitting posture recognition method are implemented. For example, the computer-readable storage medium can be the above-mentioned memory 502 including program instructions, and the above program instructions can be executed by the processor 501 of the electronic device 500 to complete the above-mentioned sitting posture recognition method.
[0164] Figure 6 FIG. is a block diagram of an electronic device 600 shown according to an exemplary embodiment. For example, the electronic device 600 can be provided as a server. Referring to Figure 6 , the electronic device 600 includes a processor 622, the number of which can be one or more, and a memory 632 for storing computer programs executable by the processor 622. The computer programs stored in the memory 632 can include one or more modules each corresponding to a set of instructions. In addition, the processor 622 can be configured to execute the computer program to perform the above-mentioned sitting posture recognition method.
[0165] In addition, the electronic device 600 can further include a power supply component 626 and a communication component 650. The power supply component 626 can be configured to perform power management of the electronic device 600, and the communication component 650 can be configured to implement communication of the electronic device 600, for example, wired or wireless communication. In addition, the electronic device 600 can further include an input / output (I / O) interface 658. The electronic device 600 can operate based on an operating system stored in the memory 632, such as Windows Server TM , Mac OSX TM , Unix TM , Linux TM and so on.
[0166] In another exemplary embodiment, a computer-readable storage medium including program instructions is also provided. When the program instructions are executed by a processor, the steps of the above-mentioned sitting posture recognition method are implemented. For example, the non-transitory computer-readable storage medium may be the above-mentioned memory 632 including program instructions, and the above-mentioned program instructions may be executed by the processor 622 of the electronic device 600 to complete the above-mentioned sitting posture recognition method.
[0167] In another exemplary embodiment, a computer program product is also provided. The computer program product includes a computer program that can be executed by a programmable device, and the computer program has a code portion for executing the above-mentioned sitting posture recognition method when executed by the programmable device.
[0168] The preferred embodiments of the present disclosure have been described in detail above in conjunction with the accompanying drawings. However, the present disclosure is not limited to the specific details in the above embodiments. Within the scope of the technical concept of the present disclosure, various simple modifications can be made to the technical solutions of the present disclosure, and these simple modifications all fall within the protection scope of the present disclosure.
[0169] In addition, it should be noted that, in the above-mentioned specific embodiments, the various specific technical features described can be combined in any suitable manner without contradiction. To avoid unnecessary repetition, the present disclosure does not separately describe various possible combination methods.
[0170] Furthermore, any combination can be made between various different embodiments of the present disclosure as long as it does not violate the idea of the present disclosure, and it should also be regarded as the content disclosed by the present disclosure.
Claims
1. A sitting posture recognition method, characterized in that, The method includes: Obtaining multiple frames of human images; Performing feature extraction on each frame of the multiple frames of human images to obtain key point information; Determining a feature matrix based on the key point information; The determining the feature matrix based on the key point information includes: For each frame of human image, performing position encoding on the N key point information included in the human body in a preset order of human body parts to obtain position encoding information corresponding to each key point information; For each frame of human image, splicing the abscissa, ordinate, confidence level, and the position encoding information corresponding to each key point information to obtain a feature matrix corresponding to each frame of human image; Fusing the feature matrices corresponding to the multiple frames of human images to obtain a fused feature matrix; Inputting the fused feature matrix into a pre-trained sitting posture recognition model to obtain the probabilities of each sitting posture corresponding to the human image.
2. The method according to claim 1, characterized in that, The determining the feature matrix based on the key point information includes: Filtering the key point information to obtain first key point information with redundant key points removed; Determining the feature matrix based on the first key point information.
3. The method according to claim 1, characterized in that, Based on the key point information, determining the feature matrix includes: Normalizing the key point information to obtain third key point information after normalization; Determining the feature matrix based on the third key point information.
4. The method according to claim 1, characterized in that, The determining the feature matrix based on the key point information includes: Performing a feature splicing operation on the key point information of each frame of the multiple frames of human images to determine the fused feature matrix.
5. The method according to claim 1, characterized in that, The sitting posture recognition model includes an attention module and a sitting posture classification module. The inputting the fused feature matrix into the pre-trained sitting posture recognition model to obtain the probabilities of each sitting posture corresponding to the human image includes: Inputting the fused feature matrix into the attention module to determine the attention scores of each sub-feature in the fused feature matrix; Processing the fused feature matrix based on the attention scores to obtain a first feature matrix; Inputting the first feature matrix into the sitting posture classification module to obtain the probabilities of each sitting posture corresponding to the human image.
6. The method according to claim 1, characterized in that, The training of the sitting posture recognition model includes: Obtaining sample human images, where the sample human images include sitting posture type annotation information, and each sample human image includes multiple frames of images; Performing feature extraction on each frame of the sample human images to obtain sample key point information; Determining a sample feature matrix based on the sample key point information; Inputting the sample feature matrix into an initial sitting posture recognition model to determine the sample probabilities of each sitting posture corresponding to the sample human image; Adjusting the parameters of the initial sitting posture recognition model according to the sample probabilities and the sitting posture type annotation information.
7. The method according to claim 6, characterized in that, The initial sitting posture recognition model includes an attention module and a sitting posture classification module. The inputting the sample feature matrix into the initial sitting posture recognition model to obtain the sample probabilities of each sitting posture corresponding to the sample human image includes: Input the sample feature matrix into the attention module to obtain the sample attention scores of each sub-feature in the sample feature matrix; Based on the sample attention scores, process the sample feature matrix to obtain a first sample feature matrix; Input the first sample feature matrix into the sitting posture classification module to obtain the sample probabilities of each sitting posture corresponding to the sample person image; The adjustment of the parameters of the initial sitting posture recognition model according to the sample probabilities and the sitting posture type annotation information includes: Adjust the parameters of the attention module and the sitting posture classification module according to the sample probabilities and the sitting posture type annotation information.
8. A sitting posture recognition device, characterized in that The device includes: An acquisition device for acquiring multiple frames of person images; An extraction module for extracting feature points from each frame of the multiple frames of person images to obtain key point information; A determination module for determining a feature matrix based on the key point information; The determination module is further configured to: For each frame of person image, perform position encoding on the N key point information included in the human body in a preset order of human body parts to obtain position encoding information corresponding to each key point information; for each frame of person image, splice the abscissa, ordinate, confidence level, and the position encoding information corresponding to each key point information to obtain the feature matrix corresponding to each frame of person image; Fuse the feature matrices corresponding to the multiple frames of person images to obtain a fused feature matrix; An identification module for inputting the fused feature matrix into a pre-trained sitting posture recognition model to obtain the probabilities of each sitting posture corresponding to the person image.
9. A non - transitory computer - readable storage medium, on which a computer program is stored, characterized in that When the program is executed by a processor, it implements the steps of the method according to any one of claims 1-7.
Citation Information
Patent Citations
Target posture recognition method and device, and camera
CN111104816A
Human body sitting posture recognition method, device, equipment and storage medium
CN111178280A