Continuous action positioning method and system and man-hour normalization analysis method
By performing object detection and node information processing on the analytical video, the three-dimensional pose information of the target object is obtained, which solves the problem of insufficient accuracy of action category recognition in traditional working hours normative analysis, and achieves efficient and accurate continuous action positioning.
Patent Information
- Application Number
- CN202510100508.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-22
- Publication Date
- 2025-05-13
AI Technical Summary
In the traditional standardized working hours analysis method, the accuracy of action category identification and positioning is insufficient, the analysis efficiency is low, and there are problems such as human error and high monitoring costs.
By performing object detection on the video to be analyzed, the target detection results are obtained, and inputting them into the preset human detection model and hand detection model, obtaining human joint node information and hand joint node information, combining these information to obtain the three-dimensional position information of the target object, and then performing continuous action positioning.
It improves the accuracy of continuous action positioning, is efficient and has low cost, and can effectively identify and correct behaviors that deviate from standard operations, ensuring efficient and compliant operations.
Smart Images

Figure CN119991807A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular to a continuous action positioning method, system and working time standardization analysis method. Background Art
[0002] As global competition becomes increasingly fierce, the competitiveness of enterprises increasingly depends on efficient production operations and accurate cost control. In order to achieve this goal, accurate work time standardization analysis has become a core element to improve overall efficiency. Work time standardization analysis refers to the analysis process of whether workers perform their work in accordance with predetermined standards, processes, and operating requirements during the production or service process. It is not only about monitoring the working time, but more importantly, it is about identifying and correcting behaviors that deviate from standard operations, thereby ensuring efficient and compliant operations.
[0003] Traditional work time standardization analysis usually relies on manual recording and manual comparison. However, this method has many limitations, such as insufficient accuracy of action category recognition (action classification) and positioning, low analysis efficiency, human errors, and high monitoring costs. Summary of the invention
[0004] The present application provides a continuous motion positioning method, system and work time standardization analysis method to solve the problems of insufficient accuracy of motion category recognition (motion classification) and positioning, as well as low efficiency of continuous motion positioning in related technologies.
[0005] The present application provides a continuous action positioning method, the method comprising: performing target detection on a video to be analyzed to obtain a target detection result, the target detection result comprising a target positioning frame where a target object in the video to be analyzed is located;
[0006] Inputting the target detection results into a preset human body detection model and a hand detection model respectively, obtaining the human body joint point information of the target object output by the human body detection model, and the hand joint point information of the target object output by the hand detection model;
[0007] Based on the human body joint point information and the hand joint point information, obtaining three-dimensional position and posture information of the joint points of the target object;
[0008] Based on the three-dimensional posture information, continuous action positioning is performed.
[0009] In one embodiment of the present application, target detection is performed on the video to be analyzed to obtain a target detection result, including:
[0010] Performing target detection on any image frame in the video to be analyzed to obtain a plurality of first positioning frames, where the first positioning frame refers to a frame containing a detection object, and the detection object refers to a detected object or human body;
[0011] Determine a first positioning frame located in a preset region of interest in the image frame as a second positioning frame, and determine the second positioning frame with the highest confidence as a target positioning frame of the current image frame;
[0012] The target positioning frame of the current image frame is determined as a reference positioning frame, and the positioning frame located in the region of interest in the next image frame is determined as a third positioning frame;
[0013] The target positioning frame in the next image frame is determined by acquiring the similarity between the third positioning frame and the reference positioning frame.
[0014] In one embodiment of the present application, obtaining the similarity between the third positioning frame and the reference positioning frame includes:
[0015] Obtaining a target intersection-and-union ratio, where the target intersection-and-union ratio refers to an intersection-and-union ratio between the third positioning frame and the reference positioning frame; determining the complement of the target intersection-and-union ratio as a frame overlap loss;
[0016] Acquire the Euclidean distance between the center point of the third positioning frame and the center point of the reference positioning frame; determine the square value of the Euclidean distance as the first intermediate value;
[0017] Obtaining a diagonal length of a target area, wherein the target area refers to a minimum closed area including the third positioning frame and the reference positioning frame; determining a square value of the diagonal length as a second intermediate value;
[0018] determining a ratio between the first intermediate value and the second intermediate value as a box distance loss;
[0019] Determine a difference in aspect ratio between the third positioning frame and the reference positioning frame as a first size loss;
[0020] Determine a product of the first size loss and a preset trade-off parameter as a second size loss;
[0021] A sum of the frame overlap loss, the frame distance loss, and the second size loss is determined as a target loss, and a complement of the target loss is determined as the similarity.
[0022] In one embodiment of the present application, based on the human body joint point information and the hand joint point information, obtaining the three-dimensional position information of the joint points of the target object includes:
[0023] Performing shape transformation on the hand joint point information to align the dimensions of the hand joint point information with the dimensions of the human body joint point information;
[0024] Performing a matrix transformation on the dimensionally aligned hand joint point information so that the hand joint point information and the human body joint point information are spatially aligned, wherein the matrix transformation includes a rotation transformation and a translation transformation;
[0025] Determine the hand joint point information after dimension alignment and spatial alignment as pre-processed hand information;
[0026] Splicing the pre-processed hand information with human body joint point information to obtain initial posture information;
[0027] The three-dimensional posture information is obtained by performing outlier processing and filtering on the initial posture information.
[0028] In one embodiment of the present application, outlier processing is performed on the initial posture information, including:
[0029] If the speed of the joint point in the initial posture information exceeds a preset speed threshold, the corresponding joint point is determined as an abnormal joint point;
[0030] Determine the position information of the abnormal joint point as the first position information, and determine the position information of the joint point adjacent to the abnormal joint point as the second position information;
[0031] Performing linear interpolation according to the first position information and the second position information to obtain position information of an interpolated joint point between the abnormal joint point and its adjacent joint point;
[0032] The abnormal joint points are replaced by using the interpolated joint points to complete the abnormal value processing.
[0033] In one embodiment of the present application, continuous action positioning is performed based on the three-dimensional posture information, including:
[0034] Based on the three-dimensional position information and the video to be analyzed, tracking the movement trajectory of the joint points of the target object to obtain the three-dimensional movement trajectory of the joint points of the target object;
[0035] Projecting the three-dimensional movement trajectory onto multiple two-dimensional planes respectively to obtain two-dimensional trajectory information of the joint point on different two-dimensional planes;
[0036] Based on the two-dimensional trajectory information of the joint points on different two-dimensional planes, continuous action positioning is performed to obtain the action position, action category, and confidence of the action position of the target object. The action position refers to the position of the bounding box where the detected action is located, and the action time is obtained based on the action position.
[0037] In one embodiment of the present application, based on the two-dimensional trajectory information of the joint points on different two-dimensional planes, continuous action positioning is performed to obtain the action position, action category, and confidence level of the action position of the target object, including:
[0038] splicing the plurality of two-dimensional trajectory information of the joint points to obtain information to be input;
[0039] Input the information to be input into a positioning sub-model in a preset continuous action positioning model, perform feature extraction and continuous action positioning, and obtain an action positioning result of the current joint point output by the positioning sub-model, wherein the action positioning result includes: action position, action category, and confidence of the action position;
[0040] Based on the action category in the action positioning result, a plurality of joint points associated with the current action category are obtained; and the plurality of joint points associated with the current action category are determined as associated joint points;
[0041] The action localization results of the multiple associated joint points are input into the adapter in the continuous action localization model for joint prediction to obtain a final localization result output by the adapter, wherein the final localization result includes the action position, action category, and confidence level of the action position of the target object that is finally determined.
[0042] In one embodiment of the present application, the training step of the continuous action positioning model includes:
[0043] Acquire a training set, wherein the training set includes a plurality of training samples, and the training samples include: information samples to be input, and corresponding real positioning results;
[0044] Inputting the information sample to be input into the continuous action positioning model, performing continuous action prediction, and obtaining a prediction result, wherein the prediction result includes: action position prediction information of the target object, action category prediction information, and a confidence prediction value of the action position, wherein the action position prediction information includes position information of a prediction bounding box where the action is located;
[0045] Obtaining a distance intersection-over-union loss between the predicted bounding box and the corresponding real bounding box in the real positioning result; obtaining a classification loss based on the gap between the action category prediction information and the real action category in the real positioning result; obtaining a confidence loss based on the gap between the confidence prediction value and the real confidence in the real positioning result;
[0046] Determine a final loss according to the distance intersection loss, the classification loss, and the confidence loss;
[0047] Based on the final loss, the continuous action localization model is trained.
[0048] The present application also provides a continuous motion positioning system, comprising:
[0049] A target detection module, used to perform target detection on the video to be analyzed and obtain a target detection result, wherein the target detection result includes a target positioning frame where the target object in the video to be analyzed is located;
[0050] A joint point detection module, used to input the target detection result into a preset human body detection model and a hand detection model respectively, to obtain the human body joint point information of the target object output by the human body detection model, and the hand joint point information of the target object output by the hand detection model;
[0051] A three-dimensional posture information acquisition module, used to obtain the three-dimensional posture information of the joint points of the target object based on the human body joint point information and the hand joint point information;
[0052] The continuous action positioning module is used to perform continuous action positioning based on the three-dimensional posture information.
[0053] This application also provides a method for analyzing the standardization of working hours, including:
[0054] Based on the positioning result of the video to be analyzed, a work time standardization analysis is performed, and the positioning result includes: the action position, action category, and confidence of the action position of the target object in the video to be analyzed, and the positioning result is obtained using any of the continuous action positioning methods described above.
[0055] Beneficial effects of the embodiments of the present application: The continuous action positioning method, system and work-hour normative analysis method provided by the embodiments of the present application, the continuous action positioning method obtains the target detection result by performing target detection on the video to be analyzed, and the target detection result includes the target positioning frame where the target object in the video to be analyzed is located; the target detection result is respectively input into the preset human body detection model and the hand detection model to obtain the human body joint point information of the target object output by the human body detection model, and the hand joint point information of the target object output by the hand detection model; based on the human body joint point information and the hand joint point information, the three-dimensional position and posture information of the joint point of the target object is obtained; based on the three-dimensional position and posture information, continuous action positioning is performed. The continuous action positioning method can help improve the accuracy of continuous action positioning, with high efficiency and low cost. BRIEF DESCRIPTION OF THE DRAWINGS
[0056] Figure 1 A schematic diagram of a flow chart of a continuous motion positioning method provided in an embodiment of the present application;
[0057] Figure 2A schematic diagram of the effects of outlier processing and Kalman filtering processing in the continuous action positioning method provided in one embodiment of the present application;
[0058] Figure 3 A schematic diagram of the trajectory projection of the ankle joint point on the xy two-dimensional plane in the continuous motion positioning method provided in one embodiment of the present application;
[0059] Figure 4 A schematic diagram of the structure of a continuous action positioning model in a continuous action positioning method provided in an embodiment of the present application;
[0060] Figure 5 An exemplary schematic diagram of a predicted bounding box and a real bounding box in a continuous action positioning method provided in an embodiment of the present application;
[0061] Figure 6 A schematic diagram of the structure of a continuous motion positioning system provided in one embodiment of the present application;
[0062] Figure 7 A schematic diagram of the structure of an electronic device provided in one embodiment of the present application. DETAILED DESCRIPTION
[0063] The following describes the embodiments of the present application through specific examples, and those skilled in the art can easily understand other advantages and effects of the present application from the contents disclosed in this specification. The present application can also be implemented or applied through other different specific embodiments, and the details in this specification can also be modified or changed in various ways based on different viewpoints and applications without departing from the spirit of the present application. It should be noted that the following embodiments and features in the embodiments can be combined with each other without conflict.
[0064] It should be noted that the illustrations provided in the following embodiments are only schematic illustrations of the basic concept of the present application, and thus the drawings only show components related to the present application rather than being drawn according to the number, shape and size of components in actual implementation. In actual implementation, the type, quantity and proportion of each component may be changed at will, and the component layout may also be more complicated.
[0065] In the following description, a large number of details are discussed to provide a more thorough explanation of the embodiments of the present application. However, it is obvious to those skilled in the art that the embodiments of the present application can be implemented without these specific details. In other embodiments, well-known structures and devices are shown in the form of block diagrams rather than in detail to avoid making the embodiments of the present application difficult to understand.
[0066] In order to facilitate understanding of the continuous motion positioning method, system and working time standardization analysis method provided by this application, some technical principles involved in this application are explained below.
[0067] Object detection technology: Object detection is a key task in the field of computer vision, which aims to accurately detect and locate all target objects in an image or video. Unlike traditional image classification tasks, object detection requires not only determining whether there is an object in the image, but also accurately determining the bounding box of the object location and classifying the object to determine which predefined category the object belongs to.
[0068] Target localization technology: Target localization is the use of algorithms to find the only target to be searched among all objects detected by the target detection algorithm, so it is a sub-problem of image retrieval.
[0069] Target positioning can be applied in security, personal positioning, and in shopping malls, it can cooperate with recommendation systems to build personalized recommendation services. Target positioning in this application refers to locating the target object we need to analyze based on all the detected objects detected by the target. The target object can be a worker or a robot, etc.
[0070] Motion capture technology: Motion capture is a technology that records the movements of humans or other organisms in the real world and converts them into digital form (posture). Motion capture can generally be divided into three aspects: human limb motion capture, facial expression capture, and hand motion capture. This technology is widely used in many fields such as movies, animation, video games, sports analysis, and virtual reality. Through motion capture, very realistic digital character movements can be produced, visual effects can be improved, and the efficiency of animation production can be increased. In this application, the main information of human joints (posture information of human joints) and hand joints (posture information of hand joints) are captured.
[0071] Motion trajectory generation technology: Motion trajectory generation technology uses various motion capture to record the position and posture data of an object during movement.
[0072] Continuous action localization technology: The purpose of continuous action localization is not only to identify the actions or events occurring in the video, but also to accurately determine the spatiotemporal location of these actions. In other words, continuous action localization involves identifying the type of each action, finding the action time (the start and end time of the action) in the video frame sequence, and the specific location of the action in each image frame.
[0073] Combine the following Figures 1 to 7 , the continuous action positioning method, system and working time normative analysis method provided in this application are explained.
[0074] See also Figure 1 , Figure 1A flow chart of a continuous motion positioning method provided in an embodiment of the present application is shown as follows: Figure 1 As shown, the method includes:
[0075] S110: Performing target detection on the video to be analyzed to obtain a target detection result, wherein the target detection result includes a target positioning frame where a target object in the video to be analyzed is located.
[0076] In some examples of this embodiment, the target object may be a person or a robot, etc. The target positioning frame where the target object is located refers to a frame containing the target object, and the target object refers to an object that needs to be positioned for continuous actions, and the frame may be a rectangular frame, etc.
[0077] S120: Input the target detection result into a preset human body detection model and a hand detection model respectively to obtain the human body joint point information of the target object output by the human body detection model and the hand joint point information of the target object output by the hand detection model.
[0078] In some examples of this embodiment, the human body detection model may be a SMPLer-X body model (an extension of the SMPL (Skinned Multi-Person Linear Model) human body model), etc., and the hand detection model may be a Willow hand model (a deep learning model specifically used for hand modeling), etc. The human body joint point information refers to the position information of the joint points of the human body of the target object, and the hand joint point information refers to the position information of the joint points of the hand of the target object. It can be understood that the position information includes: position information, velocity, acceleration and other information of the joint points.
[0079] S130: Based on the human body joint point information and the hand joint point information, obtain three-dimensional position and posture information of the joint points of the target object.
[0080] It should be noted that by integrating or splicing the human joint point information and the hand joint point information, the three-dimensional pose information of the joint point of the target object can be obtained. The three-dimensional pose information obtained in the above manner better integrates the human joint point information and the hand joint point information, and has high accuracy.
[0081] S140: Perform continuous action positioning based on the three-dimensional posture information.
[0082] It should be noted that by performing continuous motion positioning based on the above three-dimensional posture information, it can help obtain a positioning result with higher accuracy and improve the accuracy of continuous motion positioning. Moreover, compared with the method of manual motion recognition and classification, the above method can effectively improve the efficiency of continuous motion positioning, with lower cost and stronger feasibility.
[0083] In some embodiments, performing target detection on the video to be analyzed to obtain a target detection result includes:
[0084] 1. Perform target detection on any image frame in the video to be analyzed to obtain multiple first positioning frames, where the first positioning frame refers to a frame containing a detection object, and the detection object refers to a detected object or human body.
[0085] It should be noted that the image frame refers to any frame of the video to be analyzed. In target detection, the detected target may be an object in the image frame or a human body.
[0086] Second, a first positioning frame located in a preset region of interest (ROI) in the image frame is determined as a second positioning frame, and the second positioning frame with the highest confidence is determined as a target positioning frame of the current image frame.
[0087] It should be noted that by presetting the region of interest, it is easy to filter out the second positioning frame. The region of interest refers to any area in the image frame, which can be set and adjusted according to actual needs. For example: suppose the image frame is an image of a production workshop, there is an aisle in the middle of the image, and both sides of the aisle are operating tables of production equipment, then the region of interest can be the area corresponding to the aisle in the image frame, etc., so as to facilitate the positioning and monitoring of the actions of employees in the aisle.
[0088] It should be mentioned that if the currently identified image frame is the first frame of the video to be analyzed, the second positioning frame with the highest confidence level can be determined as the target positioning frame of the current image frame.
[0089] 3. Determine the target positioning frame of the current image frame as the reference positioning frame, and determine the positioning frame located in the region of interest in the next image frame as the third positioning frame.
[0090] Fourth, the target positioning frame in the next image frame is determined by obtaining the similarity between the third positioning frame and the reference positioning frame.
[0091] It should be noted that by executing the above steps, the target object in the video to be analyzed can be accurately and continuously located. It is understandable that by looping through the above steps three and four, the target object in each image frame in the video to be analyzed can be accurately located.
[0092] In some embodiments, obtaining the similarity between the third positioning frame and the reference positioning frame includes:
[0093] 1. Obtain a target intersection-and-union ratio, where the target intersection-and-union ratio refers to an intersection-and-union ratio between the third positioning frame and the reference positioning frame; and determine the complement of the target intersection-and-union ratio as a frame overlap loss.
[0094] It can be understood that the intersection-to-union ratio between the third positioning frame and the reference positioning frame refers to the ratio of the intersection area of the third positioning frame and the reference positioning frame to their union area.
[0095] The larger the target intersection-union ratio is, the more similar the third positioning frame is to the reference positioning frame, and the smaller the frame overlap loss is.
[0096] 2. Obtaining the Euclidean distance between the center point of the third positioning frame and the center point of the reference positioning frame; and determining the square value of the Euclidean distance as the first intermediate value.
[0097] 3. Obtaining the diagonal length of the target area, where the target area refers to the minimum closed area including the third positioning frame and the reference positioning frame; and determining the square value of the diagonal length as the second intermediate value.
[0098] Fourth, a ratio between the first intermediate value and the second intermediate value is determined as a box distance loss.
[0099] 5. Determine the difference in aspect ratio between the third positioning frame and the reference positioning frame as a first size loss.
[0100] 6. Determine the product of the first size loss and a preset trade-off parameter as the second size loss.
[0101] 7. Determine the sum of the frame overlap loss, the frame distance loss, and the second size loss as the target loss, and determine the complement of the target loss as the similarity.
[0102] It should be noted that by determining the similarity in the above manner, the accuracy is relatively high.
[0103] The mathematical expression of the target loss is:
[0104]
[0105] Among them, TKid represents the target loss, IOU represents the target intersection-over-union ratio, ρ(b, b gt ) indicates the center point of the third positioning frame b and the reference positioning frame b gt , c represents the diagonal length of the minimum closed area containing the third positioning frame and the reference positioning frame, a represents a preset trade-off parameter, and by adjusting the size of a, the importance of the aspect ratio difference in the total loss (target loss) can be achieved, v represents the aspect ratio difference between the third positioning frame and the reference positioning frame, which is used to measure the consistency of the aspect ratio of the third positioning frame and the reference positioning frame, and "·" represents the dot product.
[0106] Additionally, in some embodiments, the mathematical expression of v is:
[0107]
[0108] Among them, π represents pi, arctan represents the inverse tangent function, and w gt 、h gt It represents the width and height of the reference positioning frame, w and h represent the width and height of the third positioning frame.
[0109] It is understandable that common mathematical model-based target positioning methods can usually only locate characters in videos with a small number of characters and less interaction between characters. If there are more interactions between characters, the character ID (unique identifier of the target character) may jump, and the target positioning may fail. The continuous action positioning method in the above embodiment can better avoid the above problems and effectively improve the positioning accuracy of the target object or target character by adopting technical means such as setting an area of interest, determining a reference positioning frame, and determining the similarity between the reference positioning frame and the third positioning frame.
[0110] In addition, compared with REID (Re-Identification, pedestrian re-identification technology) positioning technology, the continuous motion positioning method in the above embodiment has a higher accuracy in continuous positioning of the target object. It can be understood that REID positioning technology relies on the surface features of the target object, such as clothing style, color, etc. However, in an industrial environment, workers usually wear the same clothes. Therefore, this method is prone to positioning errors, etc. The continuous motion positioning method in the above embodiment can better overcome this problem, avoid target positioning errors, and achieve accurate positioning of the target object in the video.
[0111] In some embodiments, obtaining three-dimensional position information of the joints of the target object based on the human body joint information and the hand joint information includes:
[0112] 1. Reshape the hand joint point information to align the dimensions of the hand joint point information with the dimensions of the human body joint point information.
[0113] In some examples of this embodiment, the hand joint point information is equivalent to an array containing multiple elements. By using the reshape method (a method of the NumPy (an open source library for programming languages) array object), the elements in the array are rearranged, and the dimension alignment of the hand joint point information and the human body joint point information can be better completed. For example: in reshape (-1, 3, 3), the meanings of the three parameters in the brackets are:
[0114] "-1": This is a special parameter that tells NumPy to automatically calculate the size of this dimension so that the reshaped array keeps the total number of elements in the original array unchanged. In the context of multidimensional arrays, "-1" tells NumPy to infer the size of the dimension at this position so that the new array has the same number of elements as the original array.
[0115] The first "3": The second parameter is used to specify the size of the second dimension of the new array. Here it is set to 3, which means that the number of rows in each two-dimensional sub-array ("matrix") in the new array is 3.
[0116] The second "3": The third parameter is used to specify the size of the third dimension of the new array. The "3" here means that the number of columns of each two-dimensional sub-array in the new array is also 3.
[0117] 2. Performing a matrix transformation on the dimensionally aligned hand joint point information so that the hand joint point information and the human body joint point information are spatially aligned, wherein the matrix transformation includes a rotation transformation and a translation transformation.
[0118] It should be noted that the hand joint point information is equivalent to the model composed of the hand joint points (hereinafter referred to as the hand joint point model), and the human body joint point information is equivalent to the model composed of the human body joint points (hereinafter referred to as the human body joint point model).
[0119] Maintaining alignment in space includes coordinate system alignment and joint alignment, among which, coordinate system alignment refers to performing rotation transformation, translation transformation, etc. on the hand joint information so that the hand joint model and the human body joint model are aligned in the same coordinate system. It can be understood that the hand joint model and the human body joint model may each correspond to different coordinate systems. For example, the human body joint model may use a coordinate system with the center of the human body as the origin, while the hand joint model may be a local coordinate system relative to the hand. Therefore, by performing a translation transformation on the hand joint information, the origin of the hand joint model can be moved to the origin position of the human body joint model to achieve origin alignment. Joint alignment refers to aligning the hand joints with the corresponding parts of the human body and the corresponding joints in three-dimensional space, so that the hand joints are correctly integrated with the human body joint model visually or physically.
[0120] In some examples of this embodiment, the Rodriguez formula may be solved to obtain a corresponding rotation vector, thereby performing a rotation transformation, wherein the mathematical expression of the Rodriguez formula is:
[0121] L=I+sin(θ)K+(1-cos(θ))K 2
[0122] Among them, L represents the hand joint point information after rotation transformation, I represents the unit matrix, K represents the skew-symmetric matrix corresponding to the rotation vector, θ represents the rotation angle, sin represents the sine function, and cos represents the cosine function.
[0123] 3. Determine the hand joint point information after dimension alignment and spatial alignment as the pre-processed hand information.
[0124] Fourth, the pre-processed hand information is spliced with the human body joint point information to obtain initial posture information.
[0125] 5. The three-dimensional posture information is obtained by performing outlier processing and filtering on the initial posture information.
[0126] It should be noted that, through the above steps, three-dimensional pose information with high accuracy can be obtained.
[0127] It should be mentioned that before performing the above-mentioned shape transformation and matrix transformation, the human joint point information output by the human body detection model and the hand joint point information output by the hand detection model can also be sliced and indexed to facilitate subsequent shape transformation and matrix transformation, etc.
[0128] It can be understood that the continuous action positioning method in the above embodiment can achieve the alignment and unification of human joint information and hand joint information by slicing and indexing the human joint information output by the human body detection model and the hand joint information output by the hand detection model, shape transformation, and matrix transformation. In addition, by fusing the coarse-grained human joint information and the fine-grained aligned hand joint information, detailed three-dimensional posture information can be obtained.
[0129] In some embodiments, performing outlier processing on the initial pose information includes:
[0130] 1. If the speed of a joint point in the initial posture information exceeds a preset speed threshold, the corresponding joint point is determined as an abnormal joint point.
[0131] In some examples of this embodiment, the speed threshold can be set or adjusted according to actual needs. Through the above method, it is easy to determine the abnormal joint point.
[0132] 2. Determine the position information of the abnormal joint point as the first position information, and determine the position information of the joint point adjacent to the abnormal joint point as the second position information.
[0133] 3. Perform linear interpolation based on the first position information and the second position information to obtain position information of an interpolated joint point between the abnormal joint point and its adjacent joint point.
[0134] In some examples of this embodiment, the linear interpolation formula is:
[0135] P=P1+(P2-P1)·t
[0136] Among them, P represents the interpolated joint point, P1 represents the abnormal joint point, P2 represents the adjacent joint point of the abnormal joint point, and t represents the interval between the abnormal joint point and the adjacent joint point.
[0137] In addition, assuming that the position information of P1 is (x1, y1, z1) and the position information of P2 is (x2, y2, z2), then the above linear interpolation can be written as:
[0138] P=(x1,y1,z1)+[(x2-x1), (y2-y1), (z2-z1)]·t
[0139] Fourth, using the interpolation joint points, the abnormal joint points are replaced to complete the abnormal value processing.
[0140] It can be understood that the abnormal joint point is deleted, and the difference joint point is determined as a new joint point, so as to correct the abnormal joint point. Repeating the above steps can correct all abnormal joint points in the initial posture information with high accuracy and strong rationality.
[0141] In some embodiments, the filtering process may be a Kalman filtering process. The process of the Kalman filtering process includes:
[0142] First, assume that the Kalman filter system in this embodiment is represented by the following two equations:
[0143] State transition equation: x k =F k x k-1 +B k u k +w k
[0144] Observation equation: z k =H k x k +v k
[0145] Among them, x k represents the state vector at time k, which includes the position and posture information of the joint points, etc. k represents the state transfer matrix, x k-1 represents the state vector at time k-1, B k is the input control matrix, u k represents the control vector, w krepresents the process, which can usually be assumed to be Gaussian distributed. k represents the observation vector at time k, H k represents the observation matrix, v k represents the observation noise, which can also be generally assumed to be Gaussian distributed.
[0146] Next, perform the prediction step:
[0147] 1. Prediction state estimation:
[0148] in, represents the predicted state estimate at time k, represents the predicted state estimate at time k-1.
[0149] 2. Forecast Error Covariance:
[0150] Among them, P k|k-1 represents the prediction error covariance at time k, which is used to represent the uncertainty of the prediction state estimate, P k-1|k-1 The forecast error covariance at time k-1, Indicates F k The transpose of Q k represents the covariance of the process noise.
[0151] Then, perform the update steps:
[0152] 1. Calculate Kalman gain:
[0153] Among them, K k represents the Kalman gain, Indicates H k The transpose of R k represents the covariance of the observation noise.
[0154] 2. Update the state estimate using the observed values:
[0155] in, represents the updated state estimate.
[0156] 3. Update error covariance: P k|k =(IK k H k ) k|k-1
[0157] Among them, P k|k represents the updated error covariance.
[0158] It should be noted that by executing the above steps, the filtering optimization of the initial posture information after outlier processing can be better achieved to improve the accuracy.
[0159] Figure 2 This is a schematic diagram of the effect of outlier processing and Kalman filtering processing in the continuous action positioning method provided in an embodiment of the present application. Please refer to Figure 2 , Figure 2 The optimization effect of a joint point on an arm is demonstrated in an exemplary manner. The x, y, and z in the legend represent the initial pose information of the joint point without outlier processing and Kalman filtering ( Figure 2 Only the three-dimensional coordinate information in the pose information is exemplarily shown), zx, zy, zz represent the initial pose information of the joint point after outlier processing, and kx, ky, kz represent the initial pose information of the three joint points after Kalman filtering. Figure 2 The horizontal axis represents time, and the vertical axis represents the coordinate value. Figure 2 It can be seen that the posture information after outlier processing and Kalman filtering has been well optimized.
[0160] It should be mentioned that the prediction results of most motion capture models are not 100% reliable. There are usually some outliers in their prediction results, or the difference between the output values before and after the prediction results is large, causing jitters, etc. The continuous motion positioning method in the above embodiment can effectively reduce the outliers and jitters by performing the above-mentioned outlier processing and Kalman filter processing operations.
[0161] In some embodiments, continuous action positioning is performed based on the three-dimensional posture information, including:
[0162] 1. Based on the three-dimensional posture information and the video to be analyzed, the movement trajectory of the joint points of the target object is tracked to obtain the three-dimensional movement trajectory of the joint points of the target object.
[0163] In some examples of this embodiment, a full pixel point tracking model (co-tracker) can be used to track the movement trajectory of the joint points of the target object in the video to be analyzed, so as to obtain the complete three-dimensional movement trajectory r(t)=(x(t), y(t), z(t)) of the joint points of the target object, where t represents time. x(t), y(t), z(t) represent the position function of the curve in three-dimensional space.
[0164] 2. Projecting the three-dimensional movement trajectory onto multiple two-dimensional planes respectively to obtain two-dimensional trajectory information of the joint points on different two-dimensional planes.
[0165] In some examples of the present embodiment, the three-dimensional movement trajectory can be projected onto the xy plane (horizontal plane or transverse plane), the yz plane (longitudinal plane or vertical plane), and the xz plane (lateral plane). On the xy plane, the z coordinate is ignored. Therefore, the projection on the xy plane only includes the x(t) and y(t) coordinates. On the yz plane, the x coordinate is ignored. Therefore, the projection on the yz plane only includes the y(t) and z(t) coordinates. On the xz plane, the y coordinate is ignored. Therefore, the projection on the xz plane only includes the x(t) and z(t) coordinates.
[0166] The parametric equation for the projection on the xy plane is:
[0167] p xy (t)=(x(t),y(t))
[0168] The parametric equation for the projection on the yz plane is:
[0169] p yz (t)=(y(t),z(t))
[0170] The parametric equation for the projection onto the xz plane is:
[0171] p xz (t)=(x(t),z(t))
[0172] Among them, p xy (t) represents the parametric equation of the projection on the xy plane, p yz (t) represents the parametric equation of the projection on the yz plane, p xz (t) represents the parametric equation of the projection onto the xz plane.
[0173] Figure 3 This is a schematic diagram of the trajectory projection of the ankle joint point on the xy two-dimensional plane in the continuous action positioning method provided in an embodiment of the present application. Please refer to Figure 3 , Figure 3 The Coordinate in the figure indicates the coordinate value, and Time indicates the time. The legend "X-Coordinate" indicates the horizontal coordinate. Figure 3 The trajectory projection of the ankle joint point on the xy plane can be clearly seen.
[0174] 3. Based on the two-dimensional trajectory information of the joint points on different two-dimensional planes, continuous action positioning is performed to obtain the action position, action category, and confidence of the action position of the target object. The action position refers to the position of the bounding box where the detected action is located, and the action time is obtained based on the action position.
[0175] It should be noted that, by performing continuous motion positioning based on the two-dimensional trajectory information of the joint points on different two-dimensional planes, it can help to improve the accuracy of continuous motion positioning and help to reduce the difficulty of positioning.
[0176] It should be mentioned that the two-dimensional trajectory information of the above-mentioned multiple joints on different two-dimensional planes can form a data set, which can be used for subsequent continuous action positioning. It is understandable that the common displacement trajectory analysis method usually adopts a key point detection model to generate a two-dimensional displacement trajectory. However, this method lacks depth dimension information and cannot well reflect the motion state of the target object. The continuous action positioning method in the above embodiment, by obtaining the three-dimensional posture information based on the human body joint point information and the hand joint point information, and performing two-dimensional projection on the three-dimensional posture information, can not only take into account the depth dimension information of the target object's motion trajectory, but also reduce the processing difficulty and reduce resource consumption. It is understandable that compared with continuous action positioning directly based on three-dimensional posture information, the two-dimensional projection method can effectively reduce the difficulty of continuous action positioning. In addition, compared with the 3D (three-dimensional) convolution spatiotemporal positioning method, the continuous action positioning method in the above embodiment can greatly reduce resource consumption and improve positioning accuracy. In addition, the 3D convolutional spatiotemporal positioning method usually lacks action timing information, while the continuous action positioning method in the above embodiment can better overcome this problem, that is, positioning can obtain information such as action time and action category.
[0177] In some embodiments, based on the two-dimensional trajectory information of the joint points on different two-dimensional planes, continuous action positioning is performed to obtain the action position, action category, and confidence level of the action position of the target object, including:
[0178] 1. Splicing the two-dimensional trajectory information of the joint points to obtain the information to be input.
[0179] 2. Input the information to be input into the positioning sub-model in the preset continuous action positioning model, perform feature extraction and continuous action positioning, and obtain the action positioning result of the current joint point output by the positioning sub-model, and the action positioning result includes: action position, action category, and confidence of the action position.
[0180] 3. Based on the action category in the action positioning result, a plurality of joint points associated with the current action category are obtained; and the plurality of joint points associated with the current action category are determined as associated joint points.
[0181] 4. Input the action localization results of the multiple associated joint points into the adapter in the continuous action localization model for joint prediction to obtain the final localization result output by the adapter, wherein the final localization result includes the action position, action category, and confidence level of the action position of the target object finally determined.
[0182] It should be mentioned that by adopting the above-mentioned positioning (prediction) method, accurate positioning of the action of the target object can be achieved.
[0183] Figure 4 For a structural diagram of a continuous action positioning model in a continuous action positioning method provided in an embodiment of the present application, please refer to Figure 4 , the continuous action positioning model includes a positioning sub-model and an adapter. The input of the positioning sub-model is the above-mentioned information to be input, that is, the information obtained by concatenating multiple two-dimensional trajectory information of the joint point (such as the trajectory information on the three two-dimensional planes of the xy plane, the yz plane, and the xz plane). The positioning sub-model includes: an input layer, a backbone network (Backbone), a neck network (Neck), and a prediction network (prediction). The backbone network includes multiple downsampling layers to extract features of the input information (downsampling layer by layer). The neck network includes an upsampling layer and a fusion layer, and the fusion layer is used to output three sizes of output heads, respectively (19*19, 38*38, 76*76), where a large feature map detects small changes in action, and a small feature map detects large changes in action. The prediction network is used to perform continuous action positioning based on the output heads of three sizes to obtain the action positioning result of the current joint point, and the action positioning result includes: action position, action category, and confidence of the action position.
[0184] For a certain action, the joint motion of multiple joints is usually involved. Therefore, when the action positioning result of the current joint is obtained, the action positioning results of multiple joints associated with the current action category are input into the adapter for joint prediction to obtain the final positioning result. In some embodiments, the adapter includes a fully connected layer and a Sfotmax (output function) layer.
[0185] It should be mentioned that the above fusion layer can fuse the images output by the lower different sampling layers with the images output by the upsampling layer.
[0186] In some embodiments, the training step of the continuous action localization model includes:
[0187] 1. Obtain a training set, wherein the training set includes multiple training samples, and the training samples include: information samples to be input and corresponding real positioning results.
[0188] 2. Input the information sample to be input into the continuous action positioning model, perform continuous action prediction, and obtain a prediction result, wherein the prediction result includes: action position prediction information of the target object, action category prediction information, and confidence prediction value of the action position, and the action position prediction information includes the position information of the prediction boundary box where the action is located.
[0189] 3. Obtain the distance intersection-over-union loss between the predicted bounding box and the corresponding real bounding box in the real positioning result; obtain the classification loss based on the gap between the action category prediction information and the real action category in the real positioning result; obtain the confidence loss based on the gap between the confidence prediction value and the real confidence in the real positioning result.
[0190] Figure 5 For an exemplary schematic diagram of a predicted bounding box and a real bounding box in a continuous action positioning method provided in an embodiment of the present application, please refer to Figure 5 , Figure 5 The orange box in the middle represents the true bounding box, and the green box represents the predicted bounding box. The blue box represents the minimum enclosing box (minimum enclosed area) that contains the true bounding box and the predicted bounding box, and the black line represents the Euclidean distance between the center point of the true bounding box and the center point of the predicted bounding box. The red dotted line represents the diagonal line of the minimum enclosing box of the true bounding box and the predicted bounding box.
[0191] In some examples of this embodiment, the mathematical expression of the distance intersection-over-union ratio loss is:
[0192] L DIOY =1-L IOU +ρ DIOU
[0193]
[0194] Among them, L DIOU represents the distance intersection loss, L IOU represents the intersection-over-union error, ρ DIOU Denotes the distance error, B∩B gt Represents the predicted bounding box B and the true bounding box B gt The area intersection of gt Represents the predicted bounding box B and the true bounding box B gt The area union of ρ(B,B gt ) represents the Euclidean distance, and c represents the diagonal length of the minimum enclosing box.
[0195] In some examples of this embodiment, a preset first binary cross entropy loss function may be used to obtain the classification loss. The mathematical expression of the first binary cross entropy loss function is:
[0196]
[0197] Among them, L class represents the classification loss, N represents the number of categories, x i represents action category prediction information, e represents a natural constant, y i represents the probability corresponding to the current category, represents the real action category, and log represents the logarithmic function.
[0198] In some examples of this embodiment, a preset second binary cross entropy loss function may be used to obtain the confidence loss, and the mathematical expression of the second binary cross entropy loss function is:
[0199]
[0200] Among them, L conf represents the confidence loss, N represents the number of images, M represents the number of bounding boxes predicted in a certain image, and C i represents the confidence prediction value, Indicates the true confidence. The default indicator.
[0201] 4. Determine the final loss based on the distance intersection loss, the classification loss, and the confidence loss.
[0202] 5. Based on the final loss, training the continuous action localization model.
[0203] In summary, the continuous motion positioning method in the above embodiment has the following advantages: 1. It can realize accurate positioning of the target object in the video; 2. By fusing the coarse-grained human joint point information and the fine-grained aligned hand joint point information, detailed three-dimensional posture information can be obtained; 3. It better solves the abnormal value and output jitter problems of the motion capture model output (three-dimensional posture information); 4. By performing two-dimensional projection on the three-dimensional posture information, a data set containing two-dimensional trajectory information of the relevant nodes on different two-dimensional planes can be obtained, which is convenient for model training; 5. By designing the above-mentioned continuous motion positioning model, the continuous motion positioning of the target object is better achieved.
[0204] The continuous motion positioning system provided by the present application is described below. The continuous motion positioning system described below and the continuous motion positioning method described above can be referenced to each other.
[0205] Please refer to Figure 6 , the continuous motion positioning system provided in this embodiment includes:
[0206] The target detection module 610 is used to perform target detection on the video to be analyzed and obtain a target detection result, wherein the target detection result includes a target positioning frame where the target object in the video to be analyzed is located;
[0207] The joint point detection module 620 is used to input the target detection result into a preset human body detection model and a hand detection model respectively, to obtain the human body joint point information of the target object output by the human body detection model, and the hand joint point information of the target object output by the hand detection model;
[0208] A three-dimensional pose information acquisition module 630, configured to obtain three-dimensional pose information of the joint points of the target object based on the human body joint point information and the hand joint point information;
[0209] The continuous motion positioning module 640 is used to perform continuous motion positioning based on the three-dimensional posture information. The continuous motion positioning system in this embodiment can help improve the accuracy of continuous motion positioning, with high efficiency, low cost and strong feasibility.
[0210] It should be noted that the continuous action positioning method and the continuous action positioning system provided in the above embodiment belong to the same concept, and the specific manner in which each module performs the operation has been described in detail in the method embodiment, which will not be repeated here. In practical applications, the continuous action positioning system provided in the above embodiment can allocate the above functions to different functional modules as needed, that is, divide the internal structure of the system into different functional modules to complete all or part of the functions described above, and this is not limited here.
[0211] This embodiment also provides a method for analyzing the standardization of working hours, including:
[0212] Based on the positioning result of the video to be analyzed, a work time standardization analysis is performed, and the positioning result includes: the action position, action category, and confidence of the action position of the target object in the video to be analyzed, and the positioning result is obtained by using any of the continuous action positioning methods described above. It should be noted that by using any of the continuous action positioning methods described above to perform continuous action positioning, it is possible to facilitate subsequent work time standardization analysis, effectively improve the accuracy of the work time standardization analysis, and greatly improve the analysis efficiency of the work time standardization analysis.
[0213] In some embodiments, an electronic device is also provided, which may be a server, and its internal structure is shown in FIG. Figure 7As shown. The electronic device includes a processor, a memory, a network interface and a database connected via a system bus. The processor of the electronic device is used to provide computing and control capabilities. The memory of the electronic device includes a non-volatile and / or volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the electronic device is used to communicate with an external client via a network connection. When the computer program is executed by the processor, the functions or steps on the server side of the above method are implemented.
[0214] In some embodiments, an electronic device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the following steps are implemented when the processor executes the computer program: performing target detection on a video to be analyzed to obtain a target detection result, wherein the target detection result includes a target positioning frame where a target object in the video to be analyzed is located; inputting the target detection result into a preset human body detection model and a hand detection model respectively to obtain human body joint point information of the target object output by the human body detection model, and hand joint point information of the target object output by the hand detection model; obtaining three-dimensional posture information of the joint points of the target object based on the human body joint point information and the hand joint point information; and performing continuous motion positioning based on the three-dimensional posture information.
[0215] In some embodiments, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented: target detection is performed on the video to be analyzed to obtain a target detection result, the target detection result includes a target positioning frame where the target object in the video to be analyzed is located; the target detection result is input into a preset human body detection model and a hand detection model respectively to obtain human body joint point information of the target object output by the human body detection model, and hand joint point information of the target object output by the hand detection model; based on the human body joint point information and the hand joint point information, three-dimensional posture information of the joint points of the target object is obtained; based on the three-dimensional posture information, continuous action positioning is performed.
[0216] It should be noted that the above functions or steps that can be implemented by the computer-readable storage medium or electronic device can refer to the relevant descriptions on the server side and the client side in the aforementioned method embodiment. To avoid repetition, they will not be described one by one here.
[0217] The flow chart and block diagram in the accompanying drawings illustrate the possible implementation architecture, function and operation of the method and computer program product according to various embodiments of the present disclosure. In this regard, each square box in the flow chart or block diagram can represent a module, a program segment, or a part of a code, and the module, program segment, or a part of a code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some implementations as replacements, the functions marked in the square box can also occur in a sequence different from that marked in the accompanying drawings. For example, two square boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each square box in the block diagram and / or flow chart, and the combination of the square boxes in the block diagram and / or flow chart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0218] The above embodiments are merely illustrative of the principles and effects of the present application and are not intended to limit the present application. Anyone familiar with the technology may modify or change the above embodiments without violating the spirit and scope of the present application. Therefore, all equivalent modifications or changes made by a person of ordinary skill in the art without departing from the spirit and technical ideas disclosed in the present application shall still be covered by the claims of the present application.
Claims
1. A continuous action positioning method, characterized in that: include: Performing target detection on the video to be analyzed to obtain a target detection result, wherein the target detection result includes a target positioning frame where a target object in the video to be analyzed is located; Inputting the target detection results into a preset human body detection model and a hand detection model respectively, obtaining the human body joint point information of the target object output by the human body detection model, and the hand joint point information of the target object output by the hand detection model; Based on the human body joint point information and the hand joint point information, obtaining three-dimensional position and posture information of the joint points of the target object; Based on the three-dimensional posture information, continuous action positioning is performed.
2. The continuous action positioning method according to claim 1, characterized in that: Perform target detection on the video to be analyzed and obtain target detection results, including: Performing target detection on any image frame in the video to be analyzed to obtain a plurality of first positioning frames, where the first positioning frame refers to a frame containing a detection object, and the detection object refers to a detected object or human body; Determine a first positioning frame located in a preset region of interest in the image frame as a second positioning frame, and determine the second positioning frame with the highest confidence as a target positioning frame of the current image frame; The target positioning frame of the current image frame is determined as a reference positioning frame, and the positioning frame located in the region of interest in the next image frame is determined as a third positioning frame; The target positioning frame in the next image frame is determined by acquiring the similarity between the third positioning frame and the reference positioning frame.
3. The continuous action positioning method according to claim 2, characterized in that: Acquiring the similarity between the third positioning frame and the reference positioning frame includes: Obtaining a target intersection-and-union ratio, where the target intersection-and-union ratio refers to an intersection-and-union ratio between the third positioning frame and the reference positioning frame; determining the complement of the target intersection-and-union ratio as a frame overlap loss; Acquire the Euclidean distance between the center point of the third positioning frame and the center point of the reference positioning frame; determine the square value of the Euclidean distance as the first intermediate value; Obtaining a diagonal length of a target area, wherein the target area refers to a minimum closed area including the third positioning frame and the reference positioning frame; determining a square value of the diagonal length as a second intermediate value; determining a ratio between the first intermediate value and the second intermediate value as a box distance loss; Determine a difference in aspect ratio between the third positioning frame and the reference positioning frame as a first size loss; Determine a product of the first size loss and a preset trade-off parameter as a second size loss; A sum of the frame overlap loss, the frame distance loss, and the second size loss is determined as a target loss, and a complement of the target loss is determined as the similarity.
4. The continuous action positioning method according to claim 1, characterized in that: Based on the human body joint point information and the hand joint point information, obtaining the three-dimensional position information of the joint points of the target object includes: Performing shape transformation on the hand joint point information to align the dimensions of the hand joint point information with the dimensions of the human body joint point information; Performing a matrix transformation on the dimensionally aligned hand joint point information so that the hand joint point information and the human body joint point information are spatially aligned, wherein the matrix transformation includes a rotation transformation and a translation transformation; Determine the hand joint point information after dimension alignment and spatial alignment as pre-processed hand information; Splicing the pre-processed hand information with human body joint point information to obtain initial posture information; The three-dimensional posture information is obtained by performing outlier processing and filtering on the initial posture information.
5. The continuous action positioning method according to claim 4, characterized in that: Performing outlier processing on the initial posture information includes: If the speed of the joint point in the initial posture information exceeds a preset speed threshold, the corresponding joint point is determined as an abnormal joint point; Determine the position information of the abnormal joint point as the first position information, and determine the position information of the joint point adjacent to the abnormal joint point as the second position information; Performing linear interpolation according to the first position information and the second position information to obtain position information of an interpolated joint point between the abnormal joint point and its adjacent joint point; The abnormal joint points are replaced by using the interpolated joint points to complete the abnormal value processing.
6. The continuous action positioning method according to claim 1, characterized in that: Based on the three-dimensional posture information, continuous action positioning is performed, including: Based on the three-dimensional position information and the video to be analyzed, tracking the movement trajectory of the joint points of the target object to obtain the three-dimensional movement trajectory of the joint points of the target object; Projecting the three-dimensional movement trajectory onto multiple two-dimensional planes respectively to obtain two-dimensional trajectory information of the joint point on different two-dimensional planes; Based on the two-dimensional trajectory information of the joint points on different two-dimensional planes, continuous action positioning is performed to obtain the action position, action category, and confidence of the action position of the target object. The action position refers to the position of the bounding box where the detected action is located, and the action time is obtained based on the action position.
7. The continuous action positioning method according to claim 6, characterized in that: Based on the two-dimensional trajectory information of the joint points on different two-dimensional planes, continuous action positioning is performed to obtain the action position, action category, and confidence level of the action position of the target object, including: splicing the plurality of two-dimensional trajectory information of the joint points to obtain information to be input; Input the information to be input into a positioning sub-model in a preset continuous action positioning model, perform feature extraction and continuous action positioning, and obtain an action positioning result of the current joint point output by the positioning sub-model, wherein the action positioning result includes: action position, action category, and confidence of the action position; Based on the action category in the action positioning result, a plurality of joint points associated with the current action category are obtained; and the plurality of joint points associated with the current action category are determined as associated joint points; The action localization results of the multiple associated joint points are input into the adapter in the continuous action localization model for joint prediction to obtain a final localization result output by the adapter, wherein the final localization result includes the action position, action category, and confidence level of the action position of the target object that is finally determined.
8. The continuous action positioning method according to claim 7, characterized in that: The training steps of the continuous action positioning model include: Acquire a training set, wherein the training set includes a plurality of training samples, and the training samples include: information samples to be input, and corresponding real positioning results; Inputting the information sample to be input into the continuous action positioning model, performing continuous action prediction, and obtaining a prediction result, wherein the prediction result includes: action position prediction information of the target object, action category prediction information, and a confidence prediction value of the action position, wherein the action position prediction information includes position information of a prediction bounding box where the action is located; Obtaining a distance intersection-over-union loss between the predicted bounding box and the corresponding real bounding box in the real positioning result; obtaining a classification loss based on the gap between the action category prediction information and the real action category in the real positioning result; obtaining a confidence loss based on the gap between the confidence prediction value and the real confidence in the real positioning result; Determine a final loss according to the distance intersection loss, the classification loss, and the confidence loss; Based on the final loss, the continuous action localization model is trained.
9. A continuous motion positioning system, characterized in that: include: A target detection module, used to perform target detection on the video to be analyzed and obtain a target detection result, wherein the target detection result includes a target positioning frame where the target object in the video to be analyzed is located; A joint point detection module, used to input the target detection result into a preset human body detection model and a hand detection model respectively, to obtain the human body joint point information of the target object output by the human body detection model, and the hand joint point information of the target object output by the hand detection model; A three-dimensional posture information acquisition module, used to obtain the three-dimensional posture information of the joint points of the target object based on the human body joint point information and the hand joint point information; The continuous action positioning module is used to perform continuous action positioning based on the three-dimensional posture information.
10. A method for analyzing the standardization of working hours, characterized in that: include: Based on the positioning result of the video to be analyzed, a work time standardization analysis is performed, and the positioning result includes: the action position, action category, and confidence of the action position of the target object in the video to be analyzed, and the positioning result is obtained using the continuous action positioning method as described in any one of claims 1 to 8.