Three-dimensional human body joint point positioning method for monocular color video

Through the three-dimensional human joint positioning method of monocular color video, the human body skeleton diagram is extracted and the motion characteristics is fusion, which solves the problem of low behavioral misclassification and recognition accuracy of human posture estimation in the prior art, achieving higher recognition accuracy and lower computing complexity.

CN120183002AActive Publication Date: 2025-06-20TAISHAN SPORTS IND GRP CO LTD +4
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510660043.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-22
Publication Date
2025-06-20
Estimated Expiration
2045-05-22

AI Technical Summary

Technical Problem

Existing neural network models have problems with low behavioral misclassification and recognition accuracy in human posture estimation, especially when the arms are blocked, it is difficult to accurately identify movements.

Method used

The three-dimensional human joint positioning method of monocular color video is adopted. By obtaining monocular color video images for preprocessing, human targets are extracted and binary processing is performed. The diamond-shaped four-neighborhood template is used for corrosion and expansion operations to segment the human body area, determine the coordinate information of the shoulder and hip joints, extract the human body skeleton diagram and input it into the motion recognition model, and generate high-level motion characteristics through the fusion of static and dynamic features.

Benefits of technology

It improves the recognition accuracy of human posture estimation, especially when the arm is blocked, which reduces the number of parameters and calculation complexity of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120183002A_ABST
    Figure CN120183002A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of computer vision, in particular to a three-dimensional human body joint point positioning method for a monocular color video, and the method comprises the steps: obtaining a monocular color video image, and carrying out the preprocessing of the obtained monocular color video image; extracting a human body target monocular color video image from the preprocessed monocular color video image, and performing binarization processing on the human body target monocular color video image to obtain a binarized depth human body target monocular color video image; inputting the obtained binarized depth human body target monocular color video image into a skeleton feature enhanced image convolutional network to extract human body skeleton data, and obtaining static features and time dynamic features based on the extracted human body skeleton data; the static features and the time dynamic features are fused, the high-level motion features are extracted based on the fusion result, the behavior recognition result is generated, and the recognition accuracy of the behavior recognition result is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision, and particularly to a method and storage medium for three-dimensional human joint point positioning in a monocular color video. Background Art

[0002] All along, human pose estimation has been an important research direction in the field of computer vision and is also the basis for many other research fields. For example, behavior analysis, gait recognition, person tracking, etc. all require further analysis and judgment by combining the coordinate positions of precise joint points.

[0003] Currently, in the existing neural network models, there are often some misclassifications in the action behavior classification results, resulting in the correct classification being misclassified into other behavior categories, and the recognition accuracy is too low. For example, for behaviors such as right hand wave (Wav), right hand knock on the door (Kno), and right hand reach to grab an object (RCa), the key features of these behaviors are concentrated in the arm, and the arm is easily blocked by the body, resulting in the inability to distinguish the motion actions well. Summary of the Invention

[0004] In order to solve the technical problems existing in the above background art, the present invention provides a method for three-dimensional human joint point positioning in a monocular color video, which helps to extract the human skeleton diagram when the arm is blocked during the movement, thereby improving the recognition accuracy.

[0005] To achieve the above technical solution, in the first aspect, the present invention provides a method for three-dimensional human joint point positioning in a monocular color video, including: Step 1: Obtain a monocular color video image and preprocess the obtained monocular color video image; Step 2: Extract a human target monocular color video image from the preprocessed monocular color video image, and perform binarization processing on the human target monocular color video image to obtain a binarized depth human target monocular color video image; Step 3: Extract a human skeleton diagram including joint points from the binarized depth human target monocular color video image; Step 4: Input the obtained human skeleton diagram into a motion recognition model to determine the motion recognition result; The said Step 3 includes: Using a diamond-shaped four-neighborhood template, perform the same number of erosion and dilation operations on the binarized depth human target monocular color video image to achieve human background segmentation, thereby obtaining the human region range; Statistically analyze the width of the obtained human region range, and scan the determined human region range to determine the human shoulder joint coordinate information and hip key coordinate information; Based on the coordinate information of the hip joint, perform skeleton extraction on the binary depth human target monocular color video image to extract the overall human skeleton diagram; Based on the coordinate information of the two shoulder joints, extract the skeleton of the arm part; Combine the obtained arm skeleton diagram and the overall human skeleton diagram to obtain the human skeleton diagram; Based on the determined coordinate information of the human shoulder joint and hip key coordinates and the human skeleton diagram, determine the coordinate information of other human joints; Step four includes: extracting human skeleton data from the binary depth human target monocular color video image through a human pose estimation algorithm, and obtaining static features and temporal dynamic features based on the extracted human skeleton data; Fuse the static features and temporal dynamic features, and extract high-level motion features based on the fusion result; Generate a behavior recognition result based on the high-level motion features.

[0006] Furthermore, step two includes: Perform filtering processing on the obtained monocular color video image; Perform image enhancement processing on the filtered monocular color video image.

[0007] Furthermore, step two includes: Define a rectangular frame containing the human target in the monocular color video image after image enhancement processing; Use a Gaussian mixture model for background modeling within the defined rectangular frame to obtain the human target monocular color video image; Perform binary processing on the segmented human target monocular color video image to obtain the binary depth human target monocular color video image.

[0008] Furthermore, the performing filtering processing on the obtained monocular color video image includes: A: Construct a filtering window Sxy with a size of S; B: Statistically calculate the gray values of all pixels within the filtering window, and determine the gray median value, the gray maximum value, and the gray minimum value of all pixels; C: Calculate the difference A1 between the gray median value and the gray maximum value and the difference A2 between the gray median value and the gray minimum value; D: If the difference A1 is not greater than zero and / or the difference A2 is not less than zero, and when S is less than or equal to Smax, then construct a filtering window with a size of S + 2 and repeat step B; E: If the difference A1 is greater than zero and the difference A2 is less than zero, then perform the following steps: (1)Determine the pixel gray value at the position (x, y) within the filtering window Sxy with a window size of S; (2)Calculate the difference B1 between the pixel gray value at the position (x, y) and the maximum gray value, and the difference B2 between the pixel gray value at the position (x, y) and the minimum gray value; If the difference B1 is not greater than zero and / or the difference B2 is not less than zero, then the median pixel gray value is used as the pixel gray value at the position (x, y).

[0009] Furthermore, the image enhancement processing of the filtered monocular color video image includes: Segment the filtered monocular color video image into several sub-blocks; According to the preset contrast threshold of the histogram, crop each sub-block, and redistribute the cropped part evenly across the entire gray level range.

[0010] In a second aspect, the present invention provides a computer-readable storage medium. The computer-readable storage medium includes a stored program, wherein when the program runs, it controls the device where the computer-readable storage medium is located to execute the above-mentioned three-dimensional human joint point positioning method for a monocular color video.

[0011] The beneficial effects of the present invention are as follows: (1)By using a diamond-shaped four-neighborhood template, the present invention performs the same number of erosion and dilation operations on the binary depth human target monocular color video image to obtain the human body area range; and scans the determined human body area range to determine the human shoulder joint coordinate information and hip key coordinate information; extracts the overall human skeleton diagram and the skeleton of the arm part based on the coordinate information of the hip joint and the coordinate information of the two shoulder joints, and then combines them to obtain the human skeleton diagram, which helps to avoid the inability to recognize arm movements due to human occlusion, and performs joint point positioning on the human skeleton diagram to obtain the human joint point skeleton diagram, providing reliable data for the recognition of the motion state and helping to improve the recognition accuracy.

[0012] (2)By obtaining static features (joint positions, bone vectors, bone cosine angles) and temporal dynamic features (joint velocities, bone velocities, skeleton diagram velocities), then using a two-stream branch input structure to perform feature fusion in a spatio-temporal graph convolutional block, and then integrating the fused features into a three-layer network for training to extract high-level motion features, and generating behavior recognition results based on the high-level motion features, not only reduces the number of model parameters and computational complexity, but also further helps to improve the recognition accuracy of the behavior recognition results. Description of the Drawings

[0013] The accompanying drawings forming a part of this invention are used to provide a further understanding of the invention. The schematic embodiments and descriptions thereof of the invention are used to explain the invention and shall not unduly limit the invention.

[0014] Figure 1 It is a flowchart of a method for three-dimensional human joint point positioning of a monocular color video of the present invention; Figure 2 It is a flowchart of human body background segmentation of the present invention. Among them, (a) is a binary depth human target monocular color video image; (b) is an image after erosion of the binary depth human target monocular color video image; (c) is an image after dilation of the binary depth human target monocular color video image; Figure 3 It is a human body skeleton diagram of the present invention; Figure 4 It is a flowchart of arm skeleton extraction of the present invention. Among them, (a) is a preliminary arm silhouette of the present invention; (b) is an arm silhouette of the present invention; (c) is an arm skeleton diagram of the present invention; Figure 5 It is a combined skeleton diagram of the body and arm skeletons of the present invention; Figure 6 It is a flowchart of human body skeleton positioning of the present invention. Among them, (a) is an effect diagram of the spatial coordinate information of the human hand and elbow; (b) is an effect diagram of the spatial coordinate information of the human body skeleton; (c) is a human body skeleton diagram of the joint points; Figure 7 It is a schematic structural diagram of a motion recognition model of the present invention; Figure 8 It is a schematic structural diagram of the dual-stream branch input of the motion recognition model of the present invention; Figure 9 It is a comparison diagram of the training loss value change curves of the motion recognition model of the present invention and the existing model with the change of Epoch; Figure 10 It is a comparison diagram of the test accuracy rate change curves of the motion recognition model of the present invention and the existing model with the change of Epoch. Detailed implementation manners

[0015] The present invention will be further described below in conjunction with the accompanying drawings and embodiments.

[0016] It should be noted that the following detailed descriptions are all illustrative and are intended to provide a further description of the present invention. Unless otherwise specified, each technical and scientific term used in this embodiment has the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs.

[0017] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present invention. As used herein, unless the context clearly indicates otherwise, the singular forms are also intended to include the plural forms. In addition, it should be understood that when the terms "comprising" and / or "including" are used in this specification, they specify the presence of features, steps, operations, devices, components, and / or combinations thereof.

[0018] In the present invention, terms such as "upper", "lower", "left", "right", "front", "rear", "vertical", "horizontal", "side", "bottom", etc. indicate the orientation or positional relationship based on the orientation or positional relationship shown in the drawings. They are only relational terms determined for the convenience of describing the structural relationship of each component or element of the present invention and do not specifically refer to any component or element in the present invention. It should not be construed as a limitation to the present invention.

[0019] In the present invention, terms such as "fixed connection", "connected", "connected to" should be understood in a broad sense, which may mean a fixed connection, an integral connection or a detachable connection; it may be directly connected or indirectly connected through an intermediate medium. For those related scientific research or technical personnel in the field, the specific meanings of the above terms in the present invention can be determined according to specific circumstances and should not be construed as a limitation to the present invention.

[0020] Example 1: As Figure 1 shown, this embodiment provides a method for three-dimensional human joint point positioning of monocular color video, including the following steps: S1: Obtain a monocular color video image.

[0021] S2: Preprocess the obtained monocular color video image.

[0022] Specifically, it includes the following steps: S2-1: Perform filtering processing on the obtained monocular color video image.

[0023] Specifically, it includes the following steps: A: Construct a filtering window Sxy with a size of S; B: Statistically calculate the gray values of all pixels within the filtering window and determine the gray median value, the gray maximum value, and the gray minimum value of all pixels; C: Calculate the difference A1 between the gray median value and the gray maximum value and the difference A2 between the gray median value and the gray minimum value; the calculation formula is as follows: A1 = Zmed - Zmin; A2 = Zmed - Zmax; D: If the difference A1 is not greater than zero and / or the difference A2 is not less than zero, and when S is less than or equal to Smax, then construct a filtering window with a size of S + 2, and repeat step B; E: If the difference A1 is greater than zero and the difference A2 is less than zero, then perform the following steps: (1) Determine the pixel gray value at the position (x, y) within the filtering window Sxy with a window size of S; (2) Calculate the difference B1 between the pixel gray value at the position (x, y) and the maximum gray value, and the difference B2 between the pixel gray value at the position (x, y) and the minimum gray value; The specific calculation formula is as follows: B1 = Zxy - Zmin; B2 = Zxy - Zmax; If the difference B1 is not greater than zero and / or B2 is not less than zero, then the median pixel gray value is used as the pixel gray value at the position (x, y).

[0024] Through the adaptive median filtering process, the size of the filtering window can be adaptively changed according to the noise density. A large window is used in the area where the noise appears frequently to improve the image purity, and a small window is used in the position where the noise appears less frequently, which helps to retain the details of the monocular color video image.

[0025] S2 - 2: Perform image enhancement processing on the filtered monocular color video image.

[0026] Specifically, it includes the following steps: A: Divide the filtered monocular color video image into several sub - blocks.

[0027] B: Perform histogram equalization processing on each sub - block separately.

[0028] Specifically, it includes the following steps: (1) Preset the contrast threshold of the histogram; (2) According to the preset contrast threshold of the histogram, crop each sub - block, and redistribute the cropped part evenly across the entire gray level range.

[0029] For example, assume that the contrast threshold (cropping value) is T. Sum the parts of each sub - block with gray levels greater than T, denoted as Sum. Then distribute Sum evenly across the entire gray level range, so that the overall filtered monocular color video image rises by a height L = Sum / N, where N represents the total number of pixels in the monocular color video image. Perform the following processing on the histogram with H = T - L as the boundary: 1) If the histogram amplitude is greater than T, then directly set its value to T; 2) If the histogram amplitude is between H and T, fill it up to T. 3) If the histogram amplitude is less than H, directly fill L pixels.

[0030] S3: Extract the monocular color video image of the human target from the preprocessed monocular color video image, and perform binarization processing on the monocular color video image of the human target.

[0031] Specifically, it includes the following steps: S3-1: Customize a rectangular box containing the human target.

[0032] S3-2: Use the Gaussian Mixture Model (GMM) to perform background modeling within the customized rectangular box to obtain the segmentation result of the background and foreground, that is, obtain the monocular color video image of the human target.

[0033] Specifically, in the monocular color video image, the pixels outside the rectangular box are defaulted to be the background; each pixel in the monocular color video image is considered to be connected to the surrounding pixels through an edge, and each edge has a probability of belonging to the foreground or the background according to the color similarity with the surrounding pixels. Finally, if the edge between two pixels belongs to different classifications (foreground and background), the edge between them is cut off, that is, the monocular color video image is segmented.

[0034] S3-3: Perform binarization processing on the segmented monocular color video image of the human target to obtain the binarized depth monocular color video image of the human target.

[0035] S4: Extract the human skeleton diagram including joint points from the binarized depth monocular color video image of the human target.

[0036] Extract human skeleton data from the binarized depth monocular color video image of the human target through the human pose estimation algorithm; among them, the human skeleton data includes the spatial coordinate information of each joint point of the human body in different frames, specifically including the following steps: A-1: Use the four-neighborhood template of a rhombus to perform the same number of erosion and dilation operations on the binarized depth monocular color video image of the human target to achieve human background segmentation, so as to obtain the range of the human body area (as Figure 2 shown, where Figure 2 in (a) represents the binarized depth monocular color video image of the human target, Figure 2 in (b) represents the image of the binarized depth monocular color video image of the human target after erosion, Figure 2 in (c) represents the image of the binarized depth monocular color video image of the human target after dilation).

[0037] Among them, through the erosion operation, thinner legs, heads, and unoccluded hand silhouettes will be eroded away, while the body part, due to its larger area, will retain the central part. After erosion, a dilation operation of the same duration is performed to restore the original body part, and the parts outside the other body part regions are completely eliminated, thus effectively helping to remove the influence of internal noise and boundary non-smooth regions. The extracted skeleton contains fewer spurious skeleton branches.

[0038] It should be noted that the erosion operation and dilation operation are implemented as follows: Q1: Suppose X and Y are two sets in the two-dimensional Euclidean space Ώ, and x, y, and z are points in the Euclidean space. When taking y as the origin, the set of elements in X after transformation is defined by the following formula: .

[0039] Q2: The dilation operation is performed by the following formula: ; Among them, U represents the result of X dilated by Y. Dilation is the union formed by translating the figure X by each y value. If Y is symmetric about the origin, it means moving Y at the boundary of X and incorporating the covered area into U. The intuitive effect is that X becomes one circle larger, so it is called the dilation operation.

[0040] Q3: The erosion operation is performed by the following formula: ; Among them, Z represents the result of X eroded by Y. The operation effect of erosion is that the union of the sets obtained by translating the elements in Z by the template Y forms A. If the template Y is symmetric about the origin, then the final effect is to remove some elements on the periphery of A to form Z. The intuitive effect is that X becomes one circle smaller.

[0041] A-2: Statistically analyze the width of the obtained human body region range and scan the determined human body region range to determine the human shoulder joint coordinate information and hip key coordinate information.

[0042] Specifically, it includes the following steps: (1) Statistically analyze the width of the obtained human body region range to obtain the average width value W. Take the average width value W as the width of the human shoulder joint and, based on the proportional relationship between pixels and the actual image, obtain the pixel value W + n (for example, W + 6).

[0043] Among them, when statistically analyzing the width of the human body region range, significant minimum and maximum values are automatically removed.

[0044] (2) Scan the human body area from top to bottom. When the width of the human body area in several consecutive rows is greater than the average width value W, determine the longitudinal coordinate position of the shoulder joint, that is, determine the height position of the shoulder joint.

[0045] (3) According to the sharp corner at the highest point, determine the horizontal coordinate position of the midpoint of the line connecting the two shoulders, that is, determine the horizontal position of the shoulder joint.

[0046] (4) Based on the determined longitudinal coordinate position and central horizontal coordinate position of the shoulder joint, determine the coordinate information of the two shoulder joints of the human body area.

[0047] (5) In a similar way, scan the human body area from top to bottom. When it is detected that the width of the human body area in several consecutive rows is close to the average width value W, determine the longitudinal coordinate position of the hip joint, that is, determine the height position of the hip joint. The boundary of the human body area at the height position is the coordinate information of the hip joint.

[0048] A-3: Based on the coordinate information of the hip joint, perform skeleton extraction on the binary depth human target monocular color video image to extract the overall human skeleton diagram ( Figure 3 as shown).

[0049] A-4: Based on the coordinate information of the two shoulder joints, extract the skeleton of the arm part.

[0050] As Figure 4 shown, it specifically includes the following steps: (1) Statistically analyze the depth values within the human body area range (that is, Figure 2 the human body range shown in (c) in it), and remove the depth values that differ from the average value by more than 3. That is, remove the influence of the arm in front. Then, take the average of the remaining depth values and record this average depth value as AVE_body.

[0051] (2) Set the pixels within the human body area range whose depth values differ from AVE_body by less than or equal to 4 to zero, and keep other values to obtain a preliminary arm silhouette (such as Figure 4 (a) in it).

[0052] (3) Perform several erosion operations on the arm silhouette to remove small interference areas and obtain the arm silhouette (such as Figure 4 (b) in it).

[0053] (4) Perform skeleton extraction on the obtained arm silhouette to obtain the arm skeleton diagram (such as Figure 4 (c) in it).

[0054] A-5: Combine the obtained arm skeleton diagram and the overall human skeleton diagram to obtain the human skeleton diagram (such as Figure 5 ).

[0055] A-6: Based on the determined coordinate information of the human shoulder joints, hip joints, and the human skeleton diagram, determine the coordinate information of other human joints.

[0056] Among them, other joints include: head, neck, hip, two feet, two legs, elbows, hands, and knees.

[0057] Specifically, it includes the following steps: (1) Based on the determined coordinate information of the two shoulder joints, hip joint, and the human skeleton model, the coordinate information of the head, neck, hip, two feet, and two legs can be directly determined.

[0058] Specifically, since the neck is the midpoint of the line connecting the two shoulder joints, the head is the endpoint of the skeleton above the neck, the hip is the bifurcation point where the two legs and the body part are connected; and the two feet are the two endpoints at the bottom of the skeleton; and since the position of the hip joint is the width of the body, with the hip joint information known, the hip is located on the center line of the hip joint.

[0059] (2) Based on the determined coordinate information of the two shoulder joints and the hip joint, use the method of the largest triangle to determine the coordinate information of the elbows, hands, and knees.

[0060] Specifically, it includes the following steps: K1: Start searching from the shoulder joint, find a point closer to the shoulder joint as a point on the upper arm, and the farther point as the hand.

[0061] K2: Calculate the coordinates of all points on the skeleton line corresponding to the actual space, and find a point on the skeleton line such that the area of the triangle determined by this point, the hand, and the shoulder in the actual three-dimensional space is the largest, and use this as the turning point on the skeleton line of the hand, that is, the coordinate information of the elbow joint point. If the areas are all very small, then the arm is basically in a straight state, and take the midpoint of the skeleton line as the elbow joint coordinate information (as shown in (a) of Figure 6 ).

[0062] K3: Use the above steps to determine the coordinate information of the knee joint point of the leg.

[0063] (3) Based on the determined coordinate information of the head, neck, hip, two feet, two legs, elbows, hands, and knees, obtain the human skeleton space coordinate information (as shown in (b) of Figure 6 ).

[0064] (4) Connect the obtained human skeleton space coordinate information with straight lines to form a human skeleton diagram including joint points (as shown in (c) of Figure 6 ).

[0065] Step S4 enables the correct segmentation of the human body region in a complex background environment, helps eliminate the influence of internal noise and boundary uneven areas in the human body region, and at the same time avoids the phenomenon of incorrect joint human skeletons due to body occlusion.

[0066] S5: Input the obtained human skeleton graph including joint points into the motion recognition model to determine the motion recognition result.

[0067] Among them, as Figure 7 shown, the motion recognition model includes: four layers of networks, and the four layers of networks are stacked; the four layers of networks are the first layer network, the second layer network, the third layer network, and the fourth layer network respectively.

[0068] The first layer network is a spatio-temporal graph convolution module for extracting static features and temporal dynamic features of skeleton behavior data from the human skeleton graph. Its basic network structure mainly includes a spatial graph convolution block and a temporal convolution block, which extract static features and dynamic characteristics of skeleton behavior respectively. Among them, the dynamic features include: the velocity feature of joint points, the velocity feature of bone edges, and the body velocity feature of the entire human skeleton.

[0069] The second layer network, the third layer network, and the fourth layer network are global adaptive and local enhancement graph convolution network modules (GL-GCN).

[0070] And the extracted static features and dynamic features output by the first layer network are fused before being input into the second layer network, and the fused result is then sent into the second layer network, the third layer network, and the fourth layer network to extract high-level motion features.

[0071] Specifically, the following steps: S5-1: Extract static features and dynamic features of skeleton behavior data from the human skeleton graph.

[0072] Specifically, it includes the following steps: E1: Extracting static features of skeleton behavior data from the human skeleton graph specifically includes the following steps: (1) Assume that the spatial coordinate information of joint points is represented as S: ; where C represents the number of channels; T represents the number of frames of the human skeleton sequence, and V represents the number of joint points in the human skeleton graph.

[0073] (2) Normalize the joint point coordinate information through a normalization module, which specifically includes the following steps: Determine that P is the average Euclidean distance from one joint point to another joint point in the human skeleton from the nth frame to the mth frame. The formula is as follows: .

[0074] (3) Set of joint point spatial coordinate information after normalization: .

[0075] E2: Extracting the dynamic features of the skeleton behavior data from the human skeleton diagram specifically includes the following steps: (1) Calculate the velocity features of the joint points and the velocity features of the bones by taking the difference between an adjacent frame and a frame separated by one frame.

[0076] Among them, the velocity changes of the joints and bones can only show the local motion trend of the skeleton behavior.

[0077] (2) Represent the global skeleton diagram velocity by the difference between the bone vectors formed by the shoulder center joint point and the hip center joint point in the previous frame and the next frame of the skeleton diagram.

[0078] S5-2: Initially fuse the dynamic features and static features through a dual-stream branch structure, and sequentially send the fusion result into the second layer network, the third layer network, and the fourth layer network.

[0079] Among them, the dual-stream branch input structure is as Figure 8 shown.

[0080] Among them, the number of input and output channels and the stride of the first layer network are 6, 16, and 1 respectively.

[0081] In the present invention, a dual-stream branch input structure is adopted for early feature-level fusion, greatly reducing the calculation of the number of parameters.

[0082] Global average pooling module: Used to unify the motion features output by the fourth layer network.

[0083] The fully connected layer and the softmax function layer constitute a prediction network, which is used to perform motion recognition and prediction based on the unified motion features after global average pooling, so as to output the motion and prediction results.

[0084] In the present invention, static features (joint position, bone vector, bone cosine angle) and time dynamic features (joint velocity, bone velocity, skeleton diagram velocity) are obtained, then a dual-stream branch input structure is adopted for feature fusion in the spatio-temporal graph convolution block, and then the fusion features are incorporated into a three-layer network for training, reducing the number of parameters and computational complexity of the model, and improving the recognition accuracy of the behavior recognition result.

[0085] Experimental results and analysis: The model stacks a total of four layers of networks. The first layer is a spatio-temporal graph convolution block, and the remaining 3 layers are all global adaptive and local enhanced graph convolution blocks (GL-GCN). The detailed settings of the input and output channels of the network layer and the stride are asFigure 5 As shown. The hardware platform for all experiments has a CPU of AMD R5-3500X with 6 cores and a working frequency of 3.59 GHz; 16 GB of RAM, a graphics card of Nvidia GeForce RTX 2060 with a video memory of 6 GB; model training and testing are both carried out in the PyTorch 1.5.1 deep learning framework under the Windows 10 system environment.

[0086] Among them, the correct recognition rate of this experiment is compared with that of the Lie group method of traditional manual feature extraction, the multi-view depth motion map (STACOG) method in the prior art, and the GL-GCN network model composed of three-layer global adaptive and local enhanced graph convolutional blocks (GL-GCN) for motion recognition. The comparison results are shown in Table 1.

[0087] Table 1 Comparison of the recognition accuracy of the motion recognition model and the existing methods

[0088] Generally speaking, the present invention performs feature-level fusion of rich skeleton behavior features in a shallow network through a two-stream branch input structure, and its recognition accuracy has been improved to a certain extent. From the above table, the recognition accuracy of the GL-GCN network model composed of three-layer global adaptive and local enhanced graph convolutional blocks (GL-GCN) is the closest to that of the motion recognition model of this embodiment. Taking these two models as examples, their training loss values and test accuracies are compared. Specifically Figure 9 The curve of the training loss value changing with Epoch as shown and Figure 10 The curve of the test accuracy changing with Epoch as shown.

[0089] From Figure 9 As shown, for the loss value change curves of the motion recognition model of this embodiment and the GL-GCN network model in the same period, when the motion recognition model of this embodiment is training, the loss value changes more rapidly, and its loss value drops below 1.0 at the 13th Epoch, while the GL-GCN network model drops to 0.9015 until the 24th Epoch. This shows that the two-stream branch input structure with early feature fusion of skeleton high-order information has a better fitting effect on skeleton data.

[0090] From Figure 10 As shown, the motion recognition model of this embodiment has achieved a recognition accuracy of 94.7% at the 45th Epoch, and then achieved the best recognition accuracy of 96.69% at the 85th Epoch. Compared with the recognition accuracy of the GL-GCN network model that fluctuates greatly in the early stage, the motion recognition model of this embodiment is more stable during training.

[0091] In this specification, for the same or similar parts among various embodiments, reference can be made to each other. In particular, for the terminal embodiments, since they are basically similar to the method embodiments, the description is relatively simple. For the relevant parts, reference can be made to the descriptions in the method embodiments.

[0092] In several embodiments provided by the present invention, it should be understood that the disclosed systems and methods can be implemented in other ways. For example, the system embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of the systems or units can be in electrical, mechanical or other forms.

[0093] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0094] In addition, it should be noted that the flowcharts in the accompanying drawings show the methods of the embodiments of the present disclosure. In the corresponding descriptions in the flowcharts or block diagrams in the accompanying drawings, the operations or steps corresponding to different blocks may also occur in a sequence different from that disclosed in the description. Sometimes there is no specific sequence between different operations or steps. For example, two consecutive operations or steps can actually be executed substantially in parallel, and sometimes they can also be executed in the reverse order, which can depend on the functions involved. Each block in the block diagram and / or flowchart, as well as the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system that performs the specified functions or actions, or can be implemented by a combination of dedicated hardware and computer instructions.

[0095] The above are only the preferred embodiments of the present invention and are not used to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A method for three-dimensional human joint point localization in monocular color video, characterized in that, Including: Step 1: Obtain a monocular color video image and preprocess the obtained monocular color video image; Step 2: Extract a monocular color video image of a human target from the preprocessed monocular color video image, and perform binarization processing on the monocular color video image of the human target to obtain a binarized depth monocular color video image of the human target; Step 3: Extract a human skeleton diagram including joint points from the binarized depth monocular color video image of the human target; Step 4: Input the obtained human skeleton diagram into a motion recognition model to determine the motion recognition result; The said Step 3 includes: Using a four-neighborhood template of a rhombus, perform the same number of erosion and dilation operations on the binarized depth monocular color video image of the human target to achieve human background segmentation, so as to obtain the range of the human body area; Statistically analyze the width of the obtained range of the human body area, and scan the determined range of the human body area to determine the coordinate information of the human shoulder joint and the coordinate information of the hip key; Based on the coordinate information of the hip joint, perform skeleton extraction on the binarized depth monocular color video image of the human target to extract the overall human skeleton diagram; Based on the coordinate information of the two shoulder joints, extract the skeleton of the arm part; Combine the obtained arm skeleton diagram and the overall human skeleton diagram to obtain the human skeleton diagram; Based on the determined coordinate information of the human shoulder joint and the hip key and the human skeleton diagram, determine the coordinate information of other joints of the human body; The said Step 4 includes: Extract human skeleton data from the binarized depth monocular color video image of the human target through a human pose estimation algorithm, and obtain static features and time dynamic features based on the extracted human skeleton data; Fuse the static features and the time dynamic features, and extract high-level motion features based on the fusion result; Generate a behavior recognition result based on the high-level motion features.

2. The method for three-dimensional human joint point localization in monocular color video according to claim 1, characterized in that, The said Step 2 includes: Perform filtering processing on the obtained monocular color video image; Perform image enhancement processing on the monocular color video image after the filtering processing.

3. The method for three-dimensional human joint point localization in monocular color video according to claim 1, characterized in that, The said Step 2 includes: Customize a rectangular frame containing a human target in the monocular color video image after the image enhancement processing; Use a Gaussian mixture model to perform background modeling within the customized rectangular frame to obtain a monocular color video image of the human target; Perform binarization processing on the segmented monocular color video image of the human target to obtain a binarized depth monocular color video image of the human target.

4. The method for three-dimensional human joint point localization in monocular color video according to claim 2, characterized in that, The said performing filtering processing on the obtained monocular color video image includes: A: Construct a filtering window Sxy with a size of S; B: Statistically analyze the gray values of all pixels within the filtering window, and determine the median gray value, the maximum gray value, and the minimum gray value of all pixels; C: Calculate the difference A1 between the median gray value and the maximum gray value and the difference A2 between the median gray value and the minimum gray value; D: If the difference A1 is not greater than zero and / or the difference A2 is not less than zero, then when S is less than or equal to Smax, construct a filtering window with a size of S plus 2, and repeat Step B; E: If the difference A1 is greater than zero and the difference A2 is less than zero, then perform the following steps: (1) Determine the pixel gray value at the position (x, y) within the filtering window Sxy with a window size of S; (2) Calculate the difference B1 between the pixel gray value at the position (x, y) and the maximum gray value, and the difference B2 between the pixel gray value at the position (x, y) and the minimum gray value; If the difference B1 is not greater than zero and / or the difference B2 is not less than zero, then the median pixel gray value is used as the pixel gray value at the position (x, y).

5. The method for three-dimensional human joint point localization in monocular color video according to claim 2, characterized in that, The image enhancement processing of the filtered monocular color video image includes: Segment the filtered monocular color video image into several sub-blocks; According to the preset contrast threshold of the histogram, crop each sub-block, and redistribute the cropped part evenly across the entire gray level interval.

6. A computer-readable storage medium, characterized in that, A computer-readable storage medium includes a stored program, wherein, when the program runs, it controls the device where the computer-readable storage medium is located to execute the three-dimensional human joint point positioning method of the monocular color video according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Automatic labeling method for human joint based on monocular video

    CN102609683A

  • Method for recovering real-time three-dimensional body posture based on multimodal fusion

    CN102800126A

  • Human body behavior recognition algorithm based on multi-stream graph convolution residual network

    CN115346264A

  • Action recognition method based on skeleton and image data fusion

    CN115841697A

  • Method and system for tracking human body skeleton point in two-dimensional video stream

    WO2017084204A1