Human motion action detection method and device, electronic equipment and storage medium
By utilizing the binocular triangulation principle and a multi-agent model, two ordinary cameras are used to detect ball manipulation actions, solving the problem that existing technologies cannot detect ball manipulation actions. This enables the generation of low-cost and efficient training suggestion reports, thereby improving the training effect of ball sports.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BEIJING YUANSHENG BALL SCIENCE TECHNOLOGY CO LTD
- Filing Date
- 2026-01-22
- Publication Date
- 2026-05-01
AI Technical Summary
Existing marker-based motion capture technology based on infrared cameras cannot effectively detect athletes' ball-handling movements and has poor applicability.
Using the principle of binocular triangulation, two ordinary cameras are used to simultaneously capture dual-view videos of the target trainee's actions in manipulating a target ball, obtaining the three-dimensional spatial coordinates of the human body's preset key posture points and the center of the ball, and combining this with a multi-agent model to output a training suggestion report.
This system enables motion detection in ball sports scenarios using two ordinary cameras, reducing hardware costs, simplifying the detection process, improving the efficiency and quality of training suggestion reports, and enhancing the training effect of trainees in ball sports.
Smart Images

Figure CN121963315A_ABST
Abstract
Description
Human motion detection methods, devices, electronic equipment and storage media Technical Field
[0001] This application relates to the field of computer technology, and in particular to a method, device, electronic device, and storage medium for detecting human motion. Background Technology
[0002] With the rapid development of artificial intelligence, computer vision and big data analysis technologies, human motion detection and automated analysis systems have shown broad application prospects in sports training, rehabilitation assessment and human-computer interaction.
[0003] Currently, the detection of athletes' movements mainly relies on marker-based motion capture technology based on infrared cameras, such as the Vicon motion capture system. This system constructs the athlete's movement trajectory in three-dimensional space by attaching reflective markers to key joints of the athlete's body and using multiple high-precision infrared cameras to collect data synchronously.
[0004] However, existing detection systems only detect human movements and postures, and cannot detect athletes when they are performing ball-handling actions, making them less applicable. Summary of the Invention
[0005] The purpose of this application is to address the shortcomings of the prior art by providing a method, device, electronic device, and storage medium for detecting human movement, which can reduce detection costs, improve the efficiency of generating training suggestion reports, ensure the quality of training suggestion reports, and help improve the training effect of trainees in ball sports.
[0006] To achieve the above objectives, the technical solution adopted in this application is as follows: Firstly, the present invention provides a method for detecting human motion, the method comprising: simultaneously capturing dual-view videos of a target student performing target ball manipulation actions using two cameras; based on the principle of binocular triangulation, obtaining three-dimensional spatial coordinates of a preset key human posture point and the center of the ball at at least some frame timestamps according to the dual-view video; determining the three-dimensional spatial coordinates of the preset key human posture point and the center of the ball at each frame timestamp according to the three-dimensional spatial coordinates of the preset key human posture point and the center of the ball at at least some frame timestamps; calculating target motion detection indicators corresponding to the target student performing target ball manipulation actions according to the three-dimensional spatial coordinates of the preset key human posture point and the center of the ball at each frame timestamp; and outputting a training suggestion report for the target student performing target ball manipulation actions through a multi-agent model based on the target motion detection indicators, wherein the multi-agent model includes a preset retrieval knowledge base, the preset retrieval knowledge base including: standard action descriptions corresponding to various ball manipulation actions.
[0007] In an optional implementation, determining the three-dimensional spatial coordinates of the sphere's center at each frame timestamp based on the three-dimensional spatial coordinates of the sphere's center at at least some frame timestamps includes: obtaining the three-dimensional spatial coordinates of a preset foot key posture point at at least some frame timestamps based on the three-dimensional spatial coordinates of the preset key posture points of the human body at at least some frame timestamps; determining the spatial relative position of the sphere's center and the preset foot key posture point at at least some frame timestamps based on the three-dimensional spatial coordinates of the sphere's center and the preset foot key posture point at at least some frame timestamps, the spatial relative position of the sphere's center and the preset foot key posture point being used to indicate whether the sphere and the foot are in contact; and determining the three-dimensional spatial coordinates of the sphere's center at each frame timestamp based on the spatial relative position of the sphere's center and the preset foot key posture point at at least some frame timestamps, according to the three-dimensional spatial coordinates of the sphere's center at at least some frame timestamps.
[0008] In an optional implementation, determining the three-dimensional spatial coordinates of the sphere's center at each frame timestamp based on the spatial relative position of the sphere's center and the preset key foot posture point at at least some frame timestamps includes: if the spatial relative position of the sphere's center and the preset key foot posture point at the first target frame timestamp indicates that the sphere and the foot are not in contact, then, based on the three-dimensional spatial coordinates of the sphere's center at at least some of the first frame timestamps, a preset smoothing filtering algorithm is used to determine the three-dimensional spatial coordinates of the sphere's center at at least some of the first frame timestamps within the first frame time range.
[0009] In an optional implementation, if the spatial relative position of the center of the sphere and the preset foot key posture point at the first target frame timestamp indicates that the sphere and the foot are in contact, the target motion detection index includes: the relative motion velocity of the center of the sphere and the preset foot key posture point at the first target frame timestamp; the step of calculating the target motion detection index corresponding to the target student's target ball manipulation action based on the three-dimensional spatial coordinates of the preset human body key posture point and the center of the sphere at each frame timestamp includes: based on the three-dimensional spatial coordinates of the center of the sphere and the preset foot key posture point at the first target frame timestamp, obtaining the first motion velocity and the second motion velocity of the center of the sphere and the preset foot key posture point at the first target frame timestamp based on a preset time sliding window; and calculating the relative motion velocity of the center of the sphere and the preset foot key posture point at the first target frame timestamp based on the first motion velocity and the second motion velocity of the center of the sphere and the preset foot key posture point at the first target frame timestamp.
[0010] In an optional implementation, the step of calculating the target motion detection index corresponding to the target student's target ball manipulation action based on the three-dimensional spatial coordinates of the human body's preset key posture points and the ball's center at each frame timestamp includes: determining the target motion detection index type corresponding to the target motion identifier in a preset data index requirement library based on the target motion identifier corresponding to the target ball manipulation action, wherein the preset data index requirement library includes: motion detection index types corresponding to multiple motion identifiers; and calculating the target motion detection index corresponding to the target motion identifier based on the target motion detection index type corresponding to the target motion identifier, according to the three-dimensional spatial coordinates of the human body's preset key posture points and the ball's center at each frame timestamp.
[0011] In an optional implementation, after determining the three-dimensional spatial coordinates of the human body's preset key posture points and the sphere's center at each frame timestamp based on the three-dimensional spatial coordinates corresponding to at least some frame timestamps, the method further includes: rendering a three-dimensional simulation animation of the target student performing target ball manipulation actions based on the three-dimensional spatial coordinates of the human body's preset key posture points and the sphere's center at each frame timestamp. If the spatial relative position of the sphere's center and the preset foot key posture points indicates that the sphere and foot are in contact at the first target frame timestamp, then the sphere is displayed separately in the first target frame timestamp of the three-dimensional simulation animation.
[0012] In an optional implementation, the step of obtaining the three-dimensional spatial coordinates of the human body's preset key pose points and the center of the sphere at at least some frame timestamps based on the binocular triangulation principle includes: obtaining the two-dimensional pixel coordinates of the human body's standard key pose points and the center of the sphere at at least some frame timestamps based on the binocular triangulation principle; simultaneously capturing multiple chessboard images of a preset chessboard grid under various poses using two cameras; obtaining the intrinsic parameter matrix, distortion coefficient, and relative extrinsic parameters between the two cameras based on the chessboard images of the preset chessboard grid under each pose; determining the three-dimensional spatial coordinates of the human body's standard key pose points and the center of the sphere at at least some frame timestamps based on the two-dimensional pixel coordinates of the human body's standard key pose points and the center of the sphere at at least some frame timestamps based on the three-dimensional spatial coordinates of the human body's standard key pose points at at least some frame timestamps; and determining the three-dimensional spatial coordinates of the human body's preset key pose points at at least some frame timestamps based on the three-dimensional spatial coordinates of the human body's standard key pose points at at least some frame timestamps, wherein the number of the human body's preset key pose points is greater than the number of human body preset key pose points.
[0013] Secondly, the present invention provides a human motion detection device, comprising: a shooting module for simultaneously capturing dual-view video of a target student performing target ball manipulation actions using two cameras; an acquisition module for acquiring, based on the principle of binocular triangulation, the three-dimensional spatial coordinates of a preset key human posture point and the center of the ball at at least some frame timestamps from the dual-view video; a determination module for determining, based on the three-dimensional spatial coordinates of the preset key human posture point and the center of the ball at each frame timestamp; a calculation module for calculating, based on the three-dimensional spatial coordinates of the preset key human posture point and the center of the ball at each frame timestamp, a target motion detection index corresponding to the target student performing the target ball manipulation actions; and an output module for outputting a training suggestion report for the target student performing the target ball manipulation actions through a multi-agent model based on the target motion detection index, wherein the multi-agent model includes a preset retrieval knowledge base, the preset retrieval knowledge base including standard action descriptions corresponding to various ball manipulation actions.
[0014] Thirdly, the present invention provides an electronic device, comprising: a processor, a storage medium, and a bus, wherein the storage medium stores machine-readable instructions executable by the processor, and when the electronic device is running, the processor communicates with the storage medium via the bus, and the processor executes the machine-readable instructions to perform the steps of the human motion detection method as described in any of the foregoing embodiments.
[0015] Fourthly, the present invention provides a computer-readable storage medium storing a computer program, wherein the computer program, when executed by a processor, performs the steps of the human motion detection method as described in any of the foregoing embodiments.
[0016] The beneficial effects of this application are as follows: The human motion detection method, device, electronic device, and storage medium provided in the embodiments of this application simultaneously capture dual-view videos of a target student performing target ball manipulation actions using two cameras; based on the principle of binocular triangulation, the three-dimensional spatial coordinates of the human body's preset key posture points and the ball's center at at least some frame timestamps are obtained from the dual-view video; based on the three-dimensional spatial coordinates of the human body's preset key posture points and the ball's center at at least some frame timestamps, the three-dimensional spatial coordinates of the human body's preset key posture points and the ball's center at each frame timestamp are determined; based on the three-dimensional spatial coordinates of the human body's preset key posture points and the ball's center at each frame timestamp, the target motion detection index corresponding to the target student performing target ball manipulation actions is calculated; based on the target... The motion detection metrics output training suggestion reports for target learners' ball control actions through a multi-agent model. The multi-agent model includes a preset retrieval knowledge base, which contains standard action descriptions corresponding to various ball control actions. Using the embodiments of this application, compared to the existing Vicon motion capture system, it achieves the detection of target learners' ball control actions in ball sports scenarios using two ordinary cameras. The hardware setup cost is lower, and there is no need to attach reflective markers to the athletes during the detection process. The detection method is simple, and the training suggestion report can be automatically generated. Compared to existing technologies, this improves the efficiency of training suggestion report generation and ensures the quality of the training suggestion reports, which is beneficial to improving the training effect of learners in ball sports. Attached Figure Description
[0017] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of this application and should not be regarded as a limitation of the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.
[0018] Figure 1 is a flowchart illustrating a human motion detection method according to an embodiment of this application; Figure 2 is a flowchart illustrating another human motion detection method according to an embodiment of this application; Figure 3 is a flowchart illustrating another human motion detection method according to an embodiment of this application; Figure 4 is a flowchart illustrating another human motion detection method according to an embodiment of this application; Figure 5 is a flowchart illustrating another human motion detection method according to an embodiment of this application; Figure 6 is a flowchart illustrating another human motion detection method according to an embodiment of this application; Figure 7 is a functional module diagram illustrating a human motion detection device according to an embodiment of this application; Figure 8 is a structural diagram illustrating an electronic device according to an embodiment of this application. Detailed Implementation
[0019] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. The components of the embodiments of this application described and shown in the accompanying drawings can generally be arranged and designed in various different configurations.
[0020] Therefore, the following detailed description of the embodiments of this application provided in the accompanying drawings is not intended to limit the scope of the claimed application, but merely to illustrate selected embodiments of the application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without inventive effort are within the scope of protection of this application.
[0021] It should be noted that similar labels and letters in the following figures indicate similar items. Therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures.
[0022] In related technologies, the detection of athletes' movements mainly relies on marker-based motion capture technology based on infrared cameras, such as the Vicon motion capture system. This system uses reflective markers attached to key joints of the athlete's body and multiple high-precision infrared cameras to simultaneously collect data to construct the athlete's movement trajectory in three-dimensional space. However, existing detection systems only detect human body postures and cannot detect when athletes are performing ball control movements, resulting in poor applicability.
[0023] In view of this, the embodiments of this application provide a method for detecting human motion movements, which can reduce detection costs, improve the efficiency of generating training suggestion reports, ensure the quality of training suggestion reports, and help improve the training effect of trainees in ball sports.
[0024] Figure 1 is a flowchart illustrating a human motion detection method provided in an embodiment of this application. The execution subject of this method can be an electronic device such as a computer, server, or processor. Optionally, this method can be applied to training scenarios for ball sports (such as football, basketball, etc.). As shown in Figure 1, the method includes: S101, simultaneously capturing dual-view video of the target trainee performing target ball control actions using two cameras.
[0025] In order to better understand this application, we will take the football scenario as an example for explanation. Accordingly, the target ball control actions may include: passing with the inside of the foot, passing with the inside of the foot, passing with the front of the foot, passing with the heel, pushing with the inside of the foot, shooting with the front of the foot, heading the ball, protecting the ball, juggling the ball, etc., without limitation.
[0026] Optionally, during filming, two cameras (a first camera and a second camera) can be symmetrically positioned on either side of the center line of the target detection scene. The trainee performs target ball manipulation actions within this scene, and the two cameras can simultaneously capture dual-view videos of the trainee's actions. These dual-view videos can include a first-view video captured by the first camera and a second-view video captured by the second camera. The first and second cameras can be ordinary cameras.
[0027] S102. Based on the principle of binocular triangulation, obtain the three-dimensional spatial coordinates of the human body's preset key pose points and the center of the sphere at at least some frame timestamps from the dual-view video.
[0028] Optionally, the preset key posture points of the human body may include multiple key posture points such as the head, hands, left knee, right knee, left ankle, right ankle, left elbow, and right elbow, which are not limited here.
[0029] The core of binocular triangulation is determining the three-dimensional coordinates of a point by the intersection of rays from the optical centers of two cameras in space. Optionally, based on the binocular triangulation principle, dual-view video can be analyzed to obtain the three-dimensional spatial coordinates of preset key pose points of the human body at at least some frame timestamps, and the three-dimensional spatial coordinates of the center of a sphere at at least some frame timestamps.
[0030] It should be noted that, considering the influence of lighting in the shooting scene and the potential dynamic occlusion between the human body and the ball when the target student performs the target ball manipulation action, it may not be possible to obtain the three-dimensional spatial coordinates of the human body's preset key posture points and / or the three-dimensional spatial coordinates of the ball's center for a certain frame timestamp.
[0031] S103. Based on the three-dimensional spatial coordinates of the human body's preset key pose points and the sphere's center at at least some frame timestamps, determine the three-dimensional spatial coordinates of the human body's preset key pose points and the sphere's center at each frame timestamp.
[0032] In some embodiments, considering that based on the principle of binocular triangulation, it may be impossible to obtain the three-dimensional spatial coordinates corresponding to the preset key pose points of the human body at a certain frame timestamp, and / or the three-dimensional spatial coordinates corresponding to the center of the sphere, the three-dimensional spatial coordinates of the preset key pose points of the human body at each frame timestamp can be obtained by using a preset human pose correction algorithm based on the three-dimensional spatial coordinates corresponding to the preset key pose points of the human body at at least some frame timestamps; and the three-dimensional spatial coordinates of the center of the sphere at each frame timestamp can be obtained by using a preset sphere correction algorithm based on the three-dimensional spatial coordinates corresponding to the center of the sphere at at least some frame timestamps.
[0033] Optionally, the preset human pose correction algorithm can be based on the temporal modeling model of an artificial neural network (ANN). For example, it can be implemented using a hybrid architecture that includes LSTM layers and a Transformer encoder-decoder structure, which is not limited here. The preset sphere correction algorithm can include at least one of the following: interpolation algorithm, filtering algorithm, fitting algorithm, etc., which is not limited here.
[0034] S104. Based on the preset key posture points of the human body and the three-dimensional spatial coordinates of the center of the ball at each frame timestamp, calculate the target motion detection index corresponding to the target student's target ball manipulation action.
[0035] Optionally, each ball control action can correspond to preset motion detection indicators. For example, preset motion detection indicators for a certain ball control action may include: left lunge depth, right lunge depth, pelvic descent amplitude, ball contact frequency, peak joint angles (such as knee angle and hip angle), range of joint angle changes, and average trajectory deviation, etc., which are not limited here and may vary depending on the ball control action. Of course, it should be noted that this application does not limit the specific content of the preset motion detection indicators.
[0036] Based on the above explanation, after obtaining the three-dimensional spatial coordinates of the human body's preset key posture points and the ball's center at each frame timestamp, the target motion detection index corresponding to the target student's target ball manipulation action can be calculated. It can be understood that the quantitative detection of the target student's target ball manipulation action can be achieved through this target motion detection index.
[0037] For example, when performing calculations, a three-dimensional trajectory file of the human body and the ball's key points can be generated based on the three-dimensional spatial coordinates of the human body's preset key posture points and the ball's center at each frame's timestamp (e.g., in .trc format). Based on the target action detection index type corresponding to the target action identifier, a preset calculation model is used to analyze the three-dimensional trajectory file of the human body and the ball's key points, thereby calculating the target motion detection index corresponding to the target student's target ball manipulation action.
[0038] Optionally, the preset calculation model can be built based on a large language model and pre-injected with a preset data indicator requirement library built by a team of ball sports training experts. It has the function of generating corresponding target motion detection indicators based on the analysis of the three-dimensional trajectory files of the human body and the ball's key points.
[0039] Optionally, the preset data indicator requirement library may include: action detection indicator types corresponding to various action identifiers, and indicator calculation algorithms corresponding to each action identifier.
[0040] S105. Based on the target motion detection index, output a training suggestion report for the target student to perform target ball control actions through a multi-agent model.
[0041] The multi-agent model includes a pre-set retrieval knowledge base, which contains standard action descriptions corresponding to various ball control actions.
[0042] Optionally, the standard action description corresponding to each ball control action may include at least one of the following: a description of the joint movements involved in the standard action corresponding to the ball control action, a description of the interaction between body parts and the ball (such as touching the ball with the inside or outside of the foot, the ratio of left and right feet), a description of the temporal stages of the action, etc., which are not limited here.
[0043] In some implementations, the multi-agent model can be built on a large language model and pre-injected with a preset retrieval knowledge base built by a team of ball sports training experts. It has the function of inferring and generating corresponding training suggestion reports based on the target motion detection indicators corresponding to the target student's target ball sports manipulation actions.
[0044] Optionally, the retrieval knowledge base may include: standard action descriptions for various ball control actions, interpretation information of motion detection indicators for various ball control actions, and improvement suggestions for various ball control actions under different motion detection indicators.
[0045] Based on the above description, after obtaining the target motion detection indicators corresponding to the target student's target ball manipulation actions, the target standard action description corresponding to the target ball manipulation actions can be retrieved from a preset knowledge base. The target motion detection indicators can then be analyzed using a multi-intelligent model to infer and output a training suggestion report for the target student's target ball manipulation actions. In some implementations, the generated training suggestion report can be in a preset format.
[0046] Optionally, the training suggestion report may include at least one of the following: a detailed interpretation of the meaning of each target motion detection indicator, suggestions for improving movements for each target motion detection indicator, a list of movement training, and statistical charts for each target motion detection indicator. These are not limited here and may vary depending on the actual application scenario. For example, it may also include an introduction to the indicator, a summary of suggestions, and a list of training suggestions.
[0047] In summary, this application provides a method for detecting human motion, which includes: simultaneously capturing dual-view video of a target student performing target ball manipulation actions using two cameras; based on the principle of binocular triangulation, obtaining the three-dimensional spatial coordinates of preset key human posture points and the center of the ball at at least some frame timestamps from the dual-view video; determining the three-dimensional spatial coordinates of the preset key human posture points and the center of the ball at each frame timestamp based on the three-dimensional spatial coordinates of the preset key human posture points and the center of the ball at at least some frame timestamps; calculating the target motion detection index corresponding to the target student performing target ball manipulation actions based on the three-dimensional spatial coordinates of the preset key human posture points and the center of the ball at each frame timestamp; and based on the target motion detection index, ... This application outputs training suggestion reports for target learners' ball control actions through a multi-agent model. The multi-agent model includes a preset retrieval knowledge base, which contains standard action descriptions corresponding to various ball control actions. Compared to the existing Vicon motion capture system, this application enables the detection of target learners' ball control actions in ball sports scenarios using two ordinary cameras. The hardware setup cost is low, and there is no need to attach reflective markers to athletes during the detection process. The detection method is simple, and the training suggestion report can be automatically generated. Compared to existing technologies, this application improves the efficiency of training suggestion report generation and ensures the quality of training suggestion reports, which is beneficial to improving the training effect of learners in ball sports.
[0048] Figure 2 is a flowchart illustrating another human motion detection method provided in an embodiment of this application. In an optional implementation, as shown in Figure 2, determining the three-dimensional spatial coordinates of the sphere's center at each frame timestamp based on the three-dimensional spatial coordinates of the sphere's center at at least some frame timestamps includes: S201, obtaining the three-dimensional spatial coordinates of a preset foot key posture point at at least some frame timestamps based on the three-dimensional spatial coordinates of the preset key posture points of the human body at at least some frame timestamps.
[0049] Among them, the preset key posture points of the human body may include: preset key posture points of the feet. Optionally, the preset key posture points of the feet may include: ankle joint, heel, midpoint of arch of foot, toes, etc., which are not limited here.
[0050] In some implementations, the three-dimensional spatial coordinates of the preset foot key posture points can be filtered out based on the three-dimensional spatial coordinates of the preset key posture points of the human body at at least some frame timestamps.
[0051] S202. Determine the spatial relative positions of the sphere's center and the preset foot key attitude points at at least some frame timestamps based on the three-dimensional spatial coordinates corresponding to the sphere's center and the preset foot key attitude points at at least some frame timestamps.
[0052] The spatial relative position of the sphere's center and the preset key foot posture point is used to indicate whether the sphere and the foot are in contact.
[0053] Optionally, taking the center of a sphere as an example, the three-dimensional spatial coordinates of the center of the sphere at a certain frame timestamp may include: the coordinate positions of the center of the sphere on the X-axis, Y-axis, and Z-axis in a preset world coordinate system. This application does not limit the setting method of the preset world coordinate system; it can be flexibly set according to the actual application scenario.
[0054] For a given frame timestamp, the spatial relative position of the sphere's center and the preset foot key pose points at that frame timestamp can be calculated based on the 3D spatial coordinates of the sphere's center and the preset foot key pose points at that frame timestamp.
[0055] It is understandable that if the spatial relative position of the center of the sphere and the preset key pose point of the foot in a certain frame timestamp is less than the preset distance threshold, then the sphere and the foot can be considered to be in contact; otherwise, the sphere and the foot can be considered not to be in contact.
[0056] S203. Based on the spatial relative position of the sphere's center and the preset key foot posture points at at least some frame timestamps, determine the three-dimensional spatial coordinates of the sphere's center at each frame timestamp according to the three-dimensional spatial coordinates of the sphere's center at at least some frame timestamps.
[0057] Specifically, based on the spatial relative position of the sphere's center and the preset key foot posture points at a certain frame timestamp, it can be determined whether the sphere and the foot are in contact. Based on the contact status between the sphere and the foot, different algorithms can be used to determine the three-dimensional spatial coordinates of the sphere's center at each frame timestamp, according to the three-dimensional spatial coordinates of the sphere's center at at least some frame timestamps.
[0058] By applying the embodiments of this application, when the three-dimensional spatial coordinates corresponding to a certain frame timestamp in a dual-view video are missing, the three-dimensional spatial coordinates of the sphere's center in each frame timestamp can be calculated based on the three-dimensional spatial coordinates corresponding to the sphere's center in at least some frame timestamps, thus avoiding the problem of missing data in subsequent calculations.
[0059] In an optional implementation, the above-mentioned determination of the three-dimensional spatial coordinates of the sphere's center at each frame timestamp based on the spatial relative position of the sphere's center and the preset key foot posture points at at least some frame timestamps includes: if the spatial relative position of the sphere's center and the preset key foot posture points at the first target frame timestamp indicates that the sphere and the foot are not in contact, then based on the three-dimensional spatial coordinates of the sphere's center at at least some of the first frame timestamps, a preset smoothing filtering algorithm is used to determine the three-dimensional spatial coordinates of the sphere's center at at least some of the first frame timestamps.
[0060] The first frame time range includes the timestamps of the N frames before and after the timestamp of the first target frame. The value of N can be 2, 3, etc., which is not limited here. At least some of the frame timestamps in the first frame time range can be any frame timestamps in the first frame time range, or they can be all the frame timestamps in the first frame time range, which is not limited here.
[0061] Optionally, if the spatial relative position of the sphere's center and the preset key foot posture point in the first target frame timestamp indicates that the sphere and foot are not in contact, then the sphere is mainly affected by gravity and friction, and its motion can be approximated as a smooth quadratic curve. Optionally, at this time, based on the three-dimensional spatial coordinates of the sphere's center corresponding to at least some of the first frame timestamps, a preset smoothing filtering algorithm (e.g., zero-phase smoothing filtering) and a local quadratic fitting algorithm can be used to determine the three-dimensional spatial coordinates of the sphere's center corresponding to at least some of the frame timestamps within the first frame time range.
[0062] In some implementations, if the sphere and foot come into contact at a certain frame timestamp, it is understood that the sphere's velocity and direction may change abruptly due to the impact force from the foot. Therefore, the three-dimensional spatial coordinates of the sphere's center corresponding to at least some of the frame timestamps in the second frame time range can be calculated using linear interpolation or based on a Long Short-Term Memory (LSTM) network, based on the three-dimensional spatial coordinates of the sphere's center at at least some of the second target frame timestamps.
[0063] Optionally, the second frame time range includes: the timestamps of the M frames before and after the second target frame timestamp, where M can be 2, 3, etc. At least some of the frame timestamps in the second frame time range can be any frame timestamps in the second frame time range, or they can be all frame timestamps in the second frame time range, which is not limited here.
[0064] By applying the embodiments of this application, it is possible to flexibly calculate the three-dimensional spatial coordinates of the ball's center at at least some frame timestamps based on whether the ball and the foot are in contact. This not only preserves the impact characteristics at the moment of contact but also smooths the trajectory during the free movement phase. It effectively avoids the energy attenuation and trajectory lag caused by traditional single filtering methods, significantly improving the physical consistency and dynamic accuracy of the three-dimensional reconstruction results. This provides a good data foundation for subsequent technical action analysis, especially the human-ball combination ability and the human-ball relationship.
[0065] In an optional implementation, if the spatial relative position of the center of the sphere and the preset key foot posture point at the first target frame timestamp indicates that the sphere and the foot are in contact, the target motion detection index includes: the relative motion speed of the center of the sphere and the preset key foot posture point at the first target frame timestamp.
[0066] Of course, if the spatial relative position of the center of the ball and the preset key foot posture point in the first target frame timestamp indicates that the ball and the foot are in contact, then the target motion detection index also includes other index parameters, such as the target student's left lunge depth, right lunge depth, pelvic descent amplitude, ball contact frequency, etc., which are not limited here.
[0067] Figure 3 is a flowchart illustrating another human motion detection method provided in an embodiment of this application. In an optional implementation, as shown in Figure 3, the above-mentioned calculation of the target motion detection index corresponding to the target student's target ball manipulation action based on the three-dimensional spatial coordinates of the preset key posture points of the human body and the center of the ball at each frame timestamp includes: S301, based on the three-dimensional spatial coordinates of the center of the ball and the preset key posture points of the feet at the first target frame timestamp, obtaining the first motion velocity and the second motion velocity of the center of the ball and the preset key posture points of the feet at the first target frame timestamp based on a preset time sliding window.
[0068] Optionally, the number of frames corresponding to the preset time sliding window can be 3 frames, 5 frames, etc., and is not limited here.
[0069] In some implementations, if the spatial relative position of the center of the sphere and the preset key foot posture point at the first target frame timestamp indicates that the sphere and the foot are in contact, then based on the preset time sliding window, the first motion velocity of the center of the sphere at the first target frame timestamp and the second motion velocity of the preset key foot posture point at the first target frame timestamp can be calculated respectively according to the three-dimensional spatial coordinates corresponding to the center of the sphere and the preset key foot posture point at the first target frame timestamp.
[0070] In this example, the number of frames corresponding to the preset time sliding window is 5. Optionally, for the first target frame timestamp t3, the start frame timestamp can be determined as t1 and the end frame timestamp as t5 within the preset time sliding window. Based on the three-dimensional spatial coordinates of the sphere's center at the start frame timestamp t1 and the three-dimensional spatial coordinates of the sphere's center at the end frame timestamp t5, the spatial displacement of the sphere's center corresponding to the preset time sliding window can be calculated. Based on the spatial displacement of the sphere's center corresponding to the preset time sliding window and the total time corresponding to the preset time sliding window, the first velocity of the sphere's center at the first target frame timestamp t3 can be calculated.
[0071] Of course, it should be noted that the number of frames corresponding to the preset time sliding window is not limited to the example above; secondly, the second motion velocity of the preset foot key posture point at the first target frame timestamp t3 can be calculated by referring to the calculation process of the first motion velocity of the sphere center at the first target frame timestamp t3, which will not be repeated here.
[0072] S302. Calculate the relative motion velocity between the center of the sphere and the preset key foot posture point at the first target frame timestamp based on the first and second motion velocities of the sphere's center and the preset key foot posture point.
[0073] Based on the above explanation, after obtaining the first motion velocity of the sphere's center at the first target frame time and the second motion velocity of the preset foot key posture point at the first target frame time, the relative motion velocity of the sphere's center and the preset foot key posture point at the first target frame time can be calculated based on the velocity difference between the first motion velocity and the second motion velocity.
[0074] By applying this application, the spatial relative position of the sphere's center and the preset key foot posture points at the first target frame timestamp indicates when the sphere and foot are in contact. Furthermore, the relative motion velocity of the sphere's center and the preset key foot posture points at the first target frame timestamp can be calculated, which can refine the target motion detection indicators and provide basic data support for the generation of subsequent training suggestion reports.
[0075] Figure 4 is a schematic diagram of a three-dimensional simulation effect provided by an embodiment of this application. In an optional embodiment, after determining the three-dimensional spatial coordinates of the human body's preset key posture points and the ball's center at each frame timestamp based on the three-dimensional spatial coordinates corresponding to at least some frame timestamps, the method further includes: rendering a three-dimensional simulation animation of the target student performing target ball manipulation actions based on the three-dimensional spatial coordinates of the human body's preset key posture points and the ball's center at each frame timestamp. If the spatial relative position of the ball's center and the preset foot key posture points indicates that the ball and the foot are in contact at the first target frame timestamp, then the ball is displayed separately in the first target frame timestamp of the three-dimensional simulation animation.
[0076] In some implementations, this application also supports rendering and generating three-dimensional simulation animations of target trainees performing target ball control actions.
[0077] Among them, after obtaining the three-dimensional spatial coordinates of the human body's preset key posture points and the three-dimensional spatial coordinates of the sphere's center under each frame's timestamp, it is also possible to render and generate a three-dimensional simulation animation of the target student performing target ball manipulation actions.
[0078] For example, as shown in Figure 4, the right side shows a first-view video captured by a first camera when the target student performs target ball manipulation actions within the target detection scene, and the left side shows a rendered 3D simulation animation. This 3D simulation animation may include: a virtual human body composed of multiple preset human body key posture points P1, and a virtual ball P2. Of course, it should be noted that, depending on the actual application scenario, this 3D simulation animation may also include the movement trajectory of multiple preset human body key posture points under at least some frame timestamps within the third frame time range. Among them, the multiple preset human body key posture points may include preset foot key posture points (e.g., the midpoint of the arch of the foot). The third frame time range may be L frames before the current frame timestamp, and the value of L may be 3, 5, etc., which is not limited here.
[0079] Referring to Figure 4, the movement trajectory corresponding to a preset key foot posture point (e.g., the midpoint of the left foot arch) with an L value of 3 is used as an example. This movement trajectory can include multiple trajectory points, such as S1 and S2. It is understood that, based on this movement trajectory, the target learner's step changes can be intuitively understood, which is beneficial for professional analysts to provide movement guidance by replaying the 3D simulation animation.
[0080] Optionally, during rendering, if the spatial relative position of the sphere's center and the preset key foot posture point at the first target frame timestamp indicates that the sphere and foot are in contact, then the sphere is displayed differently in the 3D simulation animation corresponding to the first target frame timestamp, so that the student can intuitively know the time of the ball's contact through the 3D simulation animation. For example, the sphere can be highlighted, or the sphere can be displayed with a preset color; there are no limitations here.
[0081] By applying the embodiments of this application, a three-dimensional simulation animation of the target trainee performing target ball manipulation actions is displayed, so that subsequent action guidance can be provided by replaying the three-dimensional simulation animation, thereby improving the training effect.
[0082] Figure 5 is a flowchart illustrating another human motion detection method provided in an embodiment of this application. In an optional implementation, as shown in Figure 5, the above-mentioned method, based on the principle of binocular triangulation, obtains the three-dimensional spatial coordinates of the human body's preset key pose points and the center of the sphere at at least some frame timestamps from dual-view video, including: S501, obtaining the two-dimensional pixel coordinates of the human body's standard key pose points and the center of the sphere at at least some frame timestamps from dual-view video.
[0083] By analyzing each video frame in the dual-view video, the two-dimensional pixel coordinates corresponding to the standard key pose points of the human body and the two-dimensional pixel coordinates corresponding to the center of the sphere can be obtained at least some of the frame timestamps.
[0084] Optionally, for the detection of standard human pose points in dual-view videos, a High-Resolution Network (HRNet) model can be used. The HRNet model can employ multi-resolution parallel convolutional feature extraction and cross-scale feature fusion mechanisms, progressively optimizing the prediction of standard human pose points through multi-stage residual modules. In some implementations, the output layer of the HRNet model can use a heatmap to represent the probability distribution of each key point, and finally, post-processing is used to obtain the two-dimensional pixel coordinates of the standard human pose points at at least some frame timestamps.
[0085] The HRNet model can be pre-trained and fine-tuned on public datasets (such as COCO Keypoints, MPII Human PoseDataset, etc.), which enables it to maintain high-resolution representation throughout the feature extraction process, taking into account both local details and global structural information. This significantly improves the accuracy and robustness of keypoint detection and ensures the algorithm's generalization ability in different scenarios and poses.
[0086] In some implementations, the detection of the center of a sphere in each frame of a dual-view video can be achieved using a YOLO model. This YOLO model can be trained based on a sphere sample set, which may include multiple sphere sample images. Each sphere sample image may be labeled with a sphere region and the corresponding sphere category label (e.g., soccer ball, basketball, etc.).
[0087] S502: Simultaneously captures multiple chessboard images corresponding to preset chessboard patterns in various poses using two cameras.
[0088] S503. Based on the chessboard images corresponding to each pose of the preset chessboard, obtain the intrinsic parameter matrix, distortion coefficient, and relative extrinsic parameters between the two cameras for each camera.
[0089] Optionally, during the actual shooting, two cameras can be used to simultaneously capture multiple chessboard images corresponding to a preset chessboard at different rotation angles. For each chessboard image, a corner detection algorithm (such as Harris corner) can be used to automatically find the pixel coordinates of the chessboard corners in each chessboard image. Using the world coordinates of these corners (known because the size of the chessboard is known) and pixel coordinates, the homography matrix H is solved using the linear least squares method (or SVD). Based on the multiple homography matrices corresponding to the multiple chessboard images, the initial values of the intrinsic parameters and distortion coefficients of each camera are linearly solved. Finally, the intrinsic parameter matrix and distortion coefficients of each camera are adjusted through nonlinear optimization (minimizing reprojection error).
[0090] In some implementations, after obtaining the intrinsic parameter matrices and distortion coefficients of each camera, a preset checkerboard pattern is simultaneously captured by the first and second cameras. The relative extrinsic parameters (rotation matrix R and translation vector t) between the two cameras are obtained by calculating the relative pose. For example, as shown in Figure 4, the pattern of the preset checkerboard pattern can be as indicated by Q in Figure 4.
[0091] S504. Based on the intrinsic parameter matrix and distortion coefficient of each camera and the relative extrinsic parameters between the two cameras, determine the three-dimensional spatial coordinates of the human standard key pose points and the sphere center at at least some frame timestamps according to the two-dimensional pixel coordinates of the human standard key pose points and the sphere center at at least some frame timestamps.
[0092] The following explanation uses the determination of the three-dimensional spatial coordinates of the human body's standard key pose points at a specific frame timestamp as an example. For each human body's standard key pose point at each frame timestamp, after obtaining the intrinsic parameter matrix, distortion coefficients, and relative extrinsic parameters between the two cameras, the intrinsic parameter matrix and distortion coefficients of the first camera can be used to perform distortion correction on the two-dimensional pixel coordinates of the detected human body's standard key pose points in the first camera image, resulting in corrected first image coordinates. Similarly, the intrinsic parameter matrix and distortion coefficients of the second camera can be used to perform distortion correction on the two-dimensional pixel coordinates of the detected human body's standard key pose points in the second camera image, resulting in corrected second image coordinates. Based on the relative extrinsic parameters between the two cameras, a binocular projection matrix group is constructed. Using the first image coordinates and the second image coordinates as corresponding point pairs, triangulation is employed, and the three-dimensional spatial coordinates of the first human body's standard key pose point at that frame timestamp are determined according to the binocular projection matrix group.
[0093] It should be noted that the process of determining the three-dimensional spatial coordinates of the center of the sphere at at least some frame timestamps can be found in the process of determining the three-dimensional spatial coordinates of the human body standard key pose points at a certain frame timestamp, and will not be repeated here.
[0094] By applying the embodiments of this application, and combining the intrinsic parameter matrix, distortion coefficient, and relative extrinsic parameters between cameras, distortion correction and 3D reconstruction are performed on the human body standard key pose points and the center of the sphere detected from multiple perspectives. This enables high-precision and robust 3D spatial positioning of human body key points and the center of the sphere without the need for external markers.
[0095] S505. Determine the three-dimensional spatial coordinates of the preset key pose points of the human body at at least some frame timestamps based on the three-dimensional spatial coordinates of the human body standard key pose points at at least some frame timestamps.
[0096] The number of preset key pose points for the human body is greater than the number of preset key pose points for the human body. Standard key pose points for the human body can be preset standard key pose points, and the number of standard key pose points for the human body can be less than the number of preset key pose points for the human body. For example, the number of standard key pose points for the human body can be 22, and the number of preset key pose points for the human body can be 40. There is no limitation here.
[0097] Optionally, based on the three-dimensional spatial coordinates of the human body standard key pose points at at least some frame timestamps, a preset temporal neural network model can be used to enhance the sparse human body standard key pose points to obtain dense human body preset key pose points. This achieves high-precision mapping from sparse to dense poses, significantly improving the spatial coherence, movement naturalness, and physical rationality of human body three-dimensional poses, and providing more stable and reliable basic data for subsequent kinematic analysis, action recognition, and performance evaluation.
[0098] The preset temporal neural network module can be built based on LSTM and Transformer, and trained using a preset training sample set. By learning the dynamic dependencies between human joints, this preset temporal neural network module can achieve a high-precision mapping from sparse standard human key pose points to dense preset human key pose points. The preset training sample set can include multiple training samples, which can include sample images of the person in various poses.
[0099] Figure 6 is a flowchart illustrating another human motion detection method provided in an embodiment of this application. In an optional implementation, as shown in Figure 6, the above-mentioned calculation of the target motion detection index corresponding to the target student's target ball manipulation action based on the three-dimensional spatial coordinates of the human body's preset key posture points and the ball's center at each frame timestamp includes: S601, determining the target motion detection index type corresponding to the target motion identifier in a preset data index requirement library based on the target motion identifier corresponding to the target ball manipulation action.
[0100] The preset data indicator requirement library includes: action detection indicator types corresponding to various action identifiers. Each action identifier may include one or more action detection indicator types, which are not limited here.
[0101] Optionally, based on the target action identifier corresponding to the target ball manipulation action, the target action detection indicator type corresponding to the target action identifier can be filtered and searched in the preset data indicator requirement library.
[0102] S602. Based on the target action detection index type corresponding to the target action identifier, calculate the target action detection index corresponding to the target student's target ball manipulation action according to the three-dimensional spatial coordinates of the human body's preset key posture points and the ball's center at each frame timestamp.
[0103] Based on the target action detection index type corresponding to the target action identifier, the target action detection index corresponding to each target action detection index type can be calculated according to the three-dimensional spatial coordinates of the human body's preset key posture points and the ball's center at each frame timestamp.
[0104] It should be noted that this application does not limit the specific representation and storage format of the target action detection index corresponding to the target ball manipulation action. Optionally, the specific representation format of the target action detection index can be numerical, chart, image, etc., and the storage format of the target action detection index can be JSON, XML, etc., which is not limited here.
[0105] For example, in a certain scenario, the target action detection index corresponding to the target ball manipulation action may include: a time-series curve diagram of the joint angle, where the horizontal axis of the time-series curve diagram of the joint angle can be time and the vertical axis can be the joint angle.
[0106] In summary, the human motion detection method provided in this application does not require an expensive infrared camera motion capture system, nor does it require dedicated sports training experts to manually analyze data, resulting in low detection costs. Furthermore, it does not require attaching reflective markers to athletes during the detection process, making it easy to deploy. In addition, through a multi-agent model, it can automatically output training suggestion reports for target trainees to perform target ball control actions, which not only improves the efficiency of training suggestion report generation but also ensures the quality of the training suggestion reports.
[0107] Figure 7 is a functional module diagram of a human motion detection device provided in an embodiment of this application. The basic principle and technical effects of the device are the same as those of the corresponding method embodiment described above. For the sake of brevity, parts not mentioned in this embodiment can be referred to the corresponding content in the method embodiment. As shown in Figure 7, the human motion detection device 100 includes: a shooting module 110, used to simultaneously capture dual-view video of the target student performing target ball manipulation actions using two cameras; an acquisition module 120, used to acquire, based on the principle of binocular triangulation, the three-dimensional spatial coordinates of the human body's preset key posture points and the ball's center at at least some frame timestamps from the dual-view video; a determination module 130, used to determine, based on the three-dimensional spatial coordinates of the human body's preset key posture points and the ball's center at at least some frame timestamps from the three-dimensional spatial coordinates of the human body's preset key posture points and the ball's center at each frame timestamp; a calculation module 140, used to calculate the target motion detection index corresponding to the target student performing the target ball manipulation actions based on the three-dimensional spatial coordinates of the human body's preset key posture points and the ball's center at each frame timestamp; and an output module 150, used to output a training suggestion report for the target student performing target ball manipulation actions based on the target motion detection index through a multi-agent model, wherein the multi-agent model includes a preset retrieval knowledge base, and the preset retrieval knowledge base includes standard action descriptions corresponding to various ball manipulation actions.
[0108] In an optional implementation, the determining module 130 is specifically configured to: obtain the three-dimensional spatial coordinates of a preset foot key posture point corresponding to at least a portion of the frame timestamps based on the three-dimensional spatial coordinates of the preset human body key posture points corresponding to at least a portion of the frame timestamps; determine the spatial relative position of the sphere center and the preset foot key posture point at at least a portion of the frame timestamps based on the three-dimensional spatial coordinates of the sphere center and the preset foot key posture point, wherein the spatial relative position of the sphere center and the preset foot key posture point is used to indicate whether the sphere and the foot are in contact; and determine the three-dimensional spatial coordinates of the sphere center at each frame timestamp based on the spatial relative position of the sphere center and the preset foot key posture point at at least a portion of the frame timestamps, according to the three-dimensional spatial coordinates of the sphere center at at least a portion of the frame timestamps.
[0109] In an optional implementation, the determining module 130 is specifically used to determine the three-dimensional spatial coordinates of the sphere's center corresponding to at least a portion of the frame timestamps within the first frame time range, based on the three-dimensional spatial coordinates of the sphere's center corresponding to at least a portion of the first frame timestamps, using a preset smoothing filtering algorithm, if the spatial relative position of the sphere's center and the preset foot key posture point indicates that the sphere and foot are not in contact in the first target frame timestamp.
[0110] In an optional implementation, if the spatial relative position of the center of the sphere and the preset foot key posture point at the first target frame timestamp indicates that the sphere and the foot are in contact, the target motion detection index includes: the relative motion velocity of the center of the sphere and the preset foot key posture point at the first target frame timestamp; the calculation module 140 is specifically used to obtain the first motion velocity and the second motion velocity of the center of the sphere and the preset foot key posture point at the first target frame timestamp based on the three-dimensional spatial coordinates corresponding to the center of the sphere and the preset foot key posture point at the first target frame timestamp, respectively, based on a preset time sliding window; and to calculate the relative motion velocity of the center of the sphere and the preset foot key posture point at the first target frame timestamp based on the first motion velocity and the second motion velocity of the center of the sphere and the preset foot key posture point at the first target frame timestamp.
[0111] In an optional implementation, the determining module 130 is further configured to render and generate a three-dimensional simulation animation of the target student performing target ball manipulation actions based on the three-dimensional spatial coordinates of the preset key posture points of the human body and the center of the ball at each frame timestamp. If the spatial relative position of the center of the ball and the preset key posture points of the feet indicates that the ball and the feet are in contact at the first target frame timestamp, then the ball is displayed in a differentiated manner in the first target frame timestamp of the three-dimensional simulation animation.
[0112] In an optional implementation, the acquisition module 120 is specifically configured to: acquire, based on the dual-view video, two-dimensional pixel coordinates corresponding to the standard key pose points of the human body and the center of the sphere at at least some frame timestamps; simultaneously capture multiple chessboard images corresponding to a preset chessboard pattern under various poses using two cameras; acquire the intrinsic parameter matrix, distortion coefficient, and relative extrinsic parameters between the two cameras based on the chessboard images corresponding to the preset chessboard pattern under each pose; determine the three-dimensional spatial coordinates corresponding to the standard key pose points of the human body and the center of the sphere at at least some frame timestamps based on the two-dimensional pixel coordinates corresponding to the standard key pose points of the human body and the center of the sphere at at least some frame timestamps; and determine the three-dimensional spatial coordinates corresponding to the preset key pose points of the human body at at least some frame timestamps based on the three-dimensional spatial coordinates corresponding to the standard key pose points of the human body at at least some frame timestamps, wherein the number of preset key pose points of the human body is greater than the number of preset key pose points of the human body.
[0113] In an optional implementation, the calculation module 140 is specifically used to determine the target action detection index type corresponding to the target action identifier in a preset data index requirement library based on the target action identifier corresponding to the target ball manipulation action. The preset data index requirement library includes: action detection index types corresponding to various action identifiers. Based on the target action detection index type corresponding to the target action identifier, and according to the three-dimensional spatial coordinates of the preset key posture points of the human body and the center of the ball at each frame timestamp, the target action detection index corresponding to the target student performing the target ball manipulation action is calculated.
[0114] The above-described device is used to execute the method provided in the foregoing embodiments, and its implementation principle and technical effect are similar, so they will not be described again here.
[0115] These modules can be one or more integrated circuits configured to implement the above methods, such as one or more Application Specific Integrated Circuits (ASICs), one or more microprocessors, or one or more Field Programmable Gate Arrays (FPGAs). Alternatively, when a module is implemented using processing element scheduler code, the processing element can be a general-purpose processor, such as a Central Processing Unit (CPU) or other processor capable of calling program code. Furthermore, these modules can be integrated together as a system-on-a-chip (SOC).
[0116] Figure 8 is a schematic diagram of an electronic device structure provided in an embodiment of this application. This electronic device can be integrated into the aforementioned human motion detection device. As shown in Figure 8, the electronic device may include: a processor 210, a storage medium 220, and a bus 230. The storage medium 220 stores machine-readable instructions executable by the processor 210. When the electronic device is running, the processor 210 communicates with the storage medium 220 via the bus 230, and the processor 210 executes the machine-readable instructions to perform the steps of the above-described method embodiment. The specific implementation and technical effects are similar and will not be described in detail here.
[0117] Optionally, this application also provides a storage medium storing a computer program, which, when run by a processor, executes the steps of the above-described method embodiments. The specific implementation and technical effects are similar and will not be repeated here.
[0118] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.
[0119] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0120] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or in a combination of hardware and software functional units.
[0121] The integrated units implemented as software functional units described above can be stored in a computer-readable storage medium. These software functional units, stored in a storage medium, include several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) or processor to execute some steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0122] It should be noted that, in this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element. The above descriptions are merely preferred embodiments of this application and are not intended to limit the application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of protection of this application. It should be noted that similar reference numerals and letters in the following figures indicate similar items. Therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures. The above description is merely a preferred embodiment of this application and is not intended to limit this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.
Claims
1. A method for detecting human movement, characterized in that, The method includes: simultaneously capturing dual-view videos of a target student performing target ball manipulation actions using two cameras; based on the principle of binocular triangulation, obtaining the three-dimensional spatial coordinates of preset key human posture points and the center of the ball at at least some frame timestamps from the dual-view videos; determining the three-dimensional spatial coordinates of the preset human posture points and the center of the ball at each frame timestamp based on the three-dimensional spatial coordinates of the preset human posture points and the center of the ball at at least some frame timestamps; calculating the target motion detection index corresponding to the target student performing the target ball manipulation actions based on the three-dimensional spatial coordinates of the preset human posture points and the center of the ball at each frame timestamp; and outputting a training suggestion report for the target student performing the target ball manipulation actions through a multi-agent model based on the target motion detection index. The multi-agent model includes a preset retrieval knowledge base, which includes standard action descriptions corresponding to various ball manipulation actions.
2. The method according to claim 1, characterized in that, The step of determining the three-dimensional spatial coordinates of the sphere's center at each frame timestamp based on the three-dimensional spatial coordinates of the sphere's center at at least some frame timestamps includes: obtaining the three-dimensional spatial coordinates of a preset foot key posture point at at least some frame timestamps based on the three-dimensional spatial coordinates of the preset human body key posture points at at least some frame timestamps; determining the spatial relative position of the sphere's center and the preset foot key posture point at at least some frame timestamps based on the three-dimensional spatial coordinates of the sphere's center and the preset foot key posture point at at least some frame timestamps, the spatial relative position of the sphere's center and the preset foot key posture point being used to indicate whether the sphere and the foot are in contact; and determining the three-dimensional spatial coordinates of the sphere's center at each frame timestamp based on the spatial relative position of the sphere's center and the preset foot key posture point at at least some frame timestamps.
3. The method according to claim 2, characterized in that, The step of determining the three-dimensional spatial coordinates of the sphere's center at each frame timestamp based on the spatial relative position of the sphere's center and the preset key foot posture point at at least a portion of the frame timestamps includes: if the spatial relative position of the sphere's center and the preset key foot posture point at the first target frame timestamp indicates that the sphere and the foot are not in contact, then based on the three-dimensional spatial coordinates of the sphere's center at at least a portion of the first frame timestamps, a preset smoothing filtering algorithm is used to determine the three-dimensional spatial coordinates of the sphere's center at at least a portion of the first frame timestamps.
4. The method according to claim 3, characterized in that, If the spatial relative position of the ball's center and the preset foot key posture point at the first target frame timestamp indicates that the ball and the foot are in contact, the target motion detection index includes: the relative motion velocity of the ball's center and the preset foot key posture point at the first target frame timestamp; the step of calculating the target motion detection index corresponding to the target student's target ball manipulation action based on the three-dimensional spatial coordinates of the preset human body key posture point and the ball's center at each frame timestamp includes: based on the three-dimensional spatial coordinates of the ball's center and the preset foot key posture point at the first target frame timestamp, obtaining the first motion velocity and the second motion velocity of the ball's center and the preset foot key posture point at the first target frame timestamp based on a preset time sliding window; and calculating the relative motion velocity of the ball's center and the preset foot key posture point at the first target frame timestamp based on the first motion velocity and the second motion velocity of the ball's center and the preset foot key posture point at the first target frame timestamp.
5. The method according to claim 1, characterized in that, The step of calculating the target motion detection index corresponding to the target student's target ball manipulation action based on the three-dimensional spatial coordinates of the human body's preset key posture points and the ball's center at each frame timestamp includes: determining the target motion detection index type corresponding to the target motion identifier in a preset data index requirement library based on the target motion identifier corresponding to the target ball manipulation action, wherein the preset data index requirement library includes: motion detection index types corresponding to multiple motion identifiers; and calculating the target motion detection index corresponding to the target motion identifier based on the target motion detection index type corresponding to the target motion identifier, according to the three-dimensional spatial coordinates of the human body's preset key posture points and the ball's center at each frame timestamp.
6. The method according to claim 2, characterized in that, After determining the three-dimensional spatial coordinates of the human body's preset key posture points and the ball's center at each frame timestamp based on the three-dimensional spatial coordinates corresponding to at least some frame timestamps, the method further includes: rendering a three-dimensional simulation animation of the target student performing target ball manipulation actions based on the three-dimensional spatial coordinates of the human body's preset key posture points and the ball's center at each frame timestamp. If the spatial relative position of the ball's center and the preset foot key posture points indicates that the ball and the foot are in contact at the first target frame timestamp, then the ball is displayed separately in the first target frame timestamp of the three-dimensional simulation animation.
7. The method according to any one of claims 1-6, characterized in that, The method based on binocular triangulation, which obtains the three-dimensional spatial coordinates of the human body's preset key pose points and the center of the sphere at at least some frame timestamps from the dual-view video, includes: obtaining the two-dimensional pixel coordinates of the human body's standard key pose points and the center of the sphere at at least some frame timestamps from the dual-view video; simultaneously capturing multiple chessboard images of a preset chessboard grid under various poses using two cameras; obtaining the intrinsic parameter matrix, distortion coefficient, and relative extrinsic parameters between the two cameras from each chessboard image corresponding to the preset chessboard grid under each pose; determining the three-dimensional spatial coordinates of the human body's standard key pose points and the center of the sphere at at least some frame timestamps based on the two-dimensional pixel coordinates of the human body's standard key pose points and the center of the sphere at at least some frame timestamps based on the three-dimensional spatial coordinates of the human body's standard key pose points at at least some frame timestamps; and determining the three-dimensional spatial coordinates of the human body's preset key pose points at at least some frame timestamps based on the three-dimensional spatial coordinates of the human body's standard key pose points at at least some frame timestamps, wherein the number of the human body's preset key pose points is greater than the number of human body preset key pose points.
8. A human motion detection device, characterized in that, The human motion detection device includes: a shooting module for simultaneously capturing dual-view video of a target student performing target ball manipulation actions using two cameras; an acquisition module for acquiring, based on the principle of binocular triangulation, the three-dimensional spatial coordinates of a preset key human posture point and the center of the ball at at least some frame timestamps from the dual-view video; a determination module for determining, based on the three-dimensional spatial coordinates of the preset key human posture point and the center of the ball at each frame timestamp; a calculation module for calculating, based on the three-dimensional spatial coordinates of the preset key human posture point and the center of the ball at each frame timestamp, the target motion detection index corresponding to the target student performing the target ball manipulation actions; and an output module for outputting a training suggestion report for the target student performing target ball manipulation actions based on the target motion detection index through a multi-agent model, wherein the multi-agent model includes a preset retrieval knowledge base, the preset retrieval knowledge base including standard action descriptions corresponding to various ball manipulation actions.
9. An electronic device, characterized in that, include: The device includes a processor, a storage medium, and a bus. The storage medium stores machine-readable instructions executable by the processor. When the electronic device is running, the processor communicates with the storage medium via the bus, and the processor executes the machine-readable instructions to perform the steps of the human motion detection method as described in any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, which, when executed by a processor, performs the steps of the human motion detection method as described in any one of claims 1-7.
Citation Information
Patent Citations
Fitting processing method of three-dimensional trajectory data and optical motion capture method
CN110753930A
Ball target detection and positioning calculation method based on binocular vision
CN114511633A
Method and apparatus for detecting posture of player in net-separated ball game
CN118967793A
Method for analyzing ball game motion, electronic device, and storage medium
US20250391038A1