Dynamic target-oriented gravity center track prediction method and system and construction method thereof
By combining an optical axis orthogonal camera and a machine learning model, the accuracy problem of human center of gravity trajectory detection in dynamic scenes is solved, and high-precision center of gravity trajectory prediction is achieved, which is suitable for applications such as rehabilitation medicine and health assessment.
Patent Information
- Application Number
- CN202510571794.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-06
- Publication Date
- 2025-09-05
AI Technical Summary
Existing technologies make it difficult to accurately detect the trajectory of the human body's center of gravity in dynamic scenes. Traditional static balance instruments have safety risks and high costs, while machine vision-based methods have low detection accuracy.
At least two cameras with orthogonal optical axes are used for frame alignment and spatial alignment. Combined with a machine learning model, the center of gravity position is predicted by extracting motion vectors of different active parts of the human body, a center of gravity trajectory prediction system is constructed, and the model accuracy is improved through hyperparameter optimization and denoising algorithms.
It achieves high-precision center of gravity trajectory prediction in dynamic environments, reduces equipment costs, and improves detection accuracy. It is suitable for scenarios such as rehabilitation medicine and health assessment.
Smart Images

Figure CN120598993A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field related to machine vision, and more specifically, relates to a method, system and construction method of center of gravity trajectory prediction for dynamic targets. Background Art
[0002] Accurate acquisition of the center of gravity trajectory is of great value in evaluating balance ability. Traditional methods mainly rely on balancing instruments and machine vision technology.
[0003] Static balance instruments calculate the coordinates of a person's center of gravity at rest using force-sensitive sensors distributed at the four corners. They are widely used in static balance assessments, but they are suitable for static scenes and cannot capture the trajectory of the center of gravity during dynamic movements such as walking and turning. While dynamic balance instruments can simulate dynamic environments, they face high safety risks and expensive equipment, making them difficult to popularize. Non-contact center of gravity trajectory acquisition methods based on machine vision use image processing technology to extract human motion characteristics. While they are applicable to dynamic scenes, their detection accuracy is low.
[0004] Therefore, it is urgent to propose a solution that can accurately detect the center of gravity trajectory of dynamic targets. Summary of the Invention
[0005] In response to the above defects or improvement needs of the prior art, the present invention provides a center of gravity trajectory prediction method, system and construction method for dynamic targets, the purpose of which is to accurately detect the center of gravity trajectory of dynamic targets.
[0006] To achieve the above objectives, according to a first aspect of the present invention, a method for predicting center of gravity trajectory of a dynamic target is provided, which comprises:
[0007] S1. performing frame alignment processing on image frames of the same target captured by at least two cameras, where the optical axes of the two cameras are orthogonal;
[0008] S2. spatially align the frame-aligned image frames to obtain the spatial coordinates of the centers of different moving parts of the target at the time of frame alignment, and obtain the motion vectors of the centers of the corresponding moving parts at different times based on the spatial coordinates of the centers of the moving parts at different times;
[0009] S3. Taking all motion vectors at the same moment as input, the machine learning model is used to predict the center of gravity position of the target at the corresponding moment to obtain the center of gravity trajectory of the target.
[0010] Optionally, in S1, performing frame alignment processing includes:
[0011] One of the cameras is the reference camera A, and the other camera is the camera to be aligned B. The timestamp t of the i-th frame image of camera A is obtained. A,i , find the image frame timestamp of camera B and tA,i The closest timestamp t B,j , and time difference |t A,i -t B,j |Compare with the preset value μ:
[0012] If |t A,i -t B,j |≤μ, then the timestamp t A,i Image frame with timestamp t B,j The image frames are aligned frames at the same moment;
[0013] Otherwise, get the camera B at timestamp t A,i The most adjacent timestamps t on both sides B,j and t B,q , calculate the time proportional coefficient The virtual frame is constructed by the scaling factor α as the time stamp t A,i The image frame is aligned with the alignment frame, and the pixels of the virtual frame are F B,virtual =(1-α)·F B,j +α·F B,q ;
[0014] Wherein, i is the frame index of camera A, i=1, 2, ..., M, M is the number of image frames of camera A, j and q are the frame indexes of camera B.
[0015] Optionally, in S2, performing spatial alignment includes:
[0016] Extract the image coordinates of the center of the same active part in different image frames;
[0017] Taking the spatial coordinates as the quantity to be determined, a set of equations is constructed based on the conversion relationship between image coordinates and world coordinates;
[0018] The image coordinates of the center of the same active part in different image frames are substituted into the equation group and solved to obtain the spatial coordinates of the center of the active part.
[0019] Optionally, the different active parts include the head, neck, shoulders, chest, abdomen, hips, thighs, knees and ankles.
[0020] According to a second aspect of the present invention, a center of gravity trajectory prediction system for a dynamic target is provided, comprising:
[0021] a frame alignment unit, configured to perform frame alignment processing on image frames of the same target captured by at least two cameras, wherein the optical axes of the two cameras are orthogonal;
[0022] A spatial alignment unit is used to perform spatial alignment on the frame-aligned image frames, obtain the spatial coordinates of the centers of different moving parts of the target at the frame alignment time, and obtain the motion vectors of the centers of the corresponding moving parts at different times based on the spatial coordinates of the centers of the moving parts at different times;
[0023] The prediction unit is used to use all motion vectors at the same moment as input, use the machine learning model to predict the center of gravity position of the target at the corresponding moment, and obtain the center of gravity trajectory of the target.
[0024] According to a third aspect of the present invention, a method for constructing a center of gravity trajectory prediction system for a dynamic target is provided. The method for constructing the center of gravity trajectory prediction system according to the second aspect includes:
[0025] Acquire image frames of different targets, where each target image frame is acquired by at least two cameras, and the optical axes of the two cameras are orthogonal;
[0026] determining a frame alignment unit, and performing frame alignment processing on image frames of the same target through the frame alignment unit;
[0027] Determine a spatial alignment unit, perform image alignment on frame-aligned image frames of the same target using the spatial alignment unit, obtain spatial coordinates of centers of different moving parts of the corresponding target at the time of frame alignment, and obtain motion vectors of the centers of different moving parts of each target at different times based on the spatial coordinates of the centers of the moving parts at different times;
[0028] All motion vectors of the same target at the same time are taken as a data point, and multiple data points form a training data set. The machine learning model is trained with the data points in the training data set as input and the actual center of gravity position of the corresponding target at the corresponding time as the label to determine the prediction unit.
[0029] Optionally, after training the machine learning model, hyperparameters of the machine learning model are further constructed;
[0030] The process of building hyperparameters for a machine learning model involves:
[0031] Obtaining an original data set, the original data set comprising a plurality of data points consisting of motion vectors, inputting the data points in the original data set into a trained machine learning model to obtain a predicted point data set, wherein the predicted points in the predicted point data set are predicted center of gravity positions generated by the machine learning model based on the input data points;
[0032] Calculate the mean square error (MSE) between the predicted point dataset and the expected point dataset, where the expected point in the expected point dataset is the actual center of gravity position;
[0033] Calculate the squared deviation between each predicted point and its corresponding expected point, and select data points from the original data set whose squared deviation of the corresponding predicted point exceeds U times the mean square error (MSE) as noise points, where U>1;
[0034] An optimal denoising ratio is determined by a progressive adjustment strategy within a preset ratio range, and noise points with the highest squared deviation value are removed from the original data set according to the optimal denoising ratio to obtain a purified data set after denoising;
[0035] Based on the purified data set, an optimization method combining hyperparameter grid search and cross-validation is used to automatically search and determine the optimal hyperparameter combination of the machine learning model.
[0036] Optionally, the training of the machine learning model to determine the prediction unit includes: training a plurality of different machine learning models respectively, and then selecting the machine learning model with the best performance as the learning model in the prediction unit;
[0037] The process of selecting a machine learning model includes:
[0038] Calculate the root mean square error RMSE and mean absolute error MAE of the k-th machine learning model and use them as the first confidence score and the second confidence score K is the number of machine learning models;
[0039] Respectively and Input the prediction task fuzzy evaluation function and get the Calculated fuzzy score and based on Calculated fuzzy score The prediction task fuzzy evaluation function includes the function function and function multiplication;
[0040] Fuzzy fraction and fuzzy scores Sum up and get the comprehensive score FS of the kth machine learning model k ;
[0041] The machine learning model with the lowest comprehensive score is selected as the machine learning model in the prediction unit.
[0042] According to a fourth aspect of the present invention, there is provided an electronic device comprising a memory and a processor, wherein the memory stores a computer program and the processor implements the steps of any one of the above methods when executing the computer program.
[0043] According to a fifth aspect of the present invention, there is provided a computer-readable storage medium having a computer program stored thereon, wherein the computer program implements the steps of any of the above methods when executed by a processor.
[0044] In general, the above technical solutions conceived by the present invention have the following beneficial effects compared with the prior art:
[0045] 1. The present invention uses at least two cameras with orthogonal optical axes to capture image frames. Through frame alignment and spatial alignment, the spatial coordinates of each target part can be accurately obtained. Furthermore, the present invention divides the target into multiple distinct moving parts, uses the motion vectors of the centers of each moving part as input parameters, and utilizes a machine learning model to learn and predict the target's center of gravity position, achieving relatively accurate center of gravity prediction. Overall, the present invention combines frame alignment, spatial alignment, and the extraction of motion vectors of different moving parts as model input, ultimately enabling the precise detection of the center of gravity trajectory of a dynamic target.
[0046] 2. Furthermore, a specific frame alignment scheme is provided. When the actual image does not meet the alignment conditions, a virtual frame is constructed by calculating the proportional coefficient and using the proportional coefficient to fuse the two frames of image. Frame alignment is achieved with the virtual frame, thereby improving the matching degree between the aligned image frames.
[0047] 3. Furthermore, when building the center of gravity trajectory prediction system, after the model training is completed, a denoising algorithm is introduced to construct the hyperparameters of the machine learning model, which can further improve the prediction accuracy of the model.
[0048] 4. Furthermore, when determining the prediction unit, a variety of different machine learning models are trained separately, and then a fuzzy evaluation mechanism is introduced to select the machine learning model with the best performance as the learning model in the prediction unit. The use of the fuzzy evaluation mechanism can more accurately evaluate the model, thereby accurately selecting the model with the best performance, thereby further improving the prediction accuracy of the system. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] Figure 1 is a flowchart of the steps of a method for predicting center of gravity trajectory in one embodiment of the present invention;
[0050] Figure 2 is a schematic diagram of capturing an image by a camera in one embodiment of the present invention;
[0051] Figure 3 is a schematic structural diagram of a center of gravity trajectory prediction system in one embodiment of the present invention;
[0052] Figure 4is a flowchart of the steps of a method for constructing a center of gravity trajectory prediction system in one embodiment of the present invention;
[0053] Figure 5 It is a schematic diagram of capturing image frames by a camera and obtaining the actual center of gravity position by a balancing instrument in one embodiment of the present invention. DETAILED DESCRIPTION
[0054] In order to make the objectives, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely for the purpose of explaining the present invention and are not intended to limit the present invention. In addition, the technical features involved in the various embodiments of the present invention described below may be combined with each other as long as they do not conflict with each other.
[0055] Example 1
[0056] The present invention provides a method for predicting center of gravity trajectory based on machine vision, such as Figure 1 FIG2 is a flowchart of a method for predicting center of gravity trajectory according to an embodiment of the present invention, and the steps are described in detail below.
[0057] S1. Perform frame alignment processing on image frames of the same target captured by at least two cameras, where the optical axes of the two cameras are orthogonal.
[0058] like Figure 2 Figure 1 shows a schematic diagram of image acquisition using cameras in one embodiment of the present invention. The optical axes of two cameras are orthogonal, and the human body is used as the target. The two cameras are used to capture images of the human body for subsequent center of gravity trajectory prediction. Specifically, one camera captures the front view, and the other captures the side view.
[0059] In the present invention, the optical axes of the cameras are made orthogonal, so that image data of the target object can be obtained from two perpendicular directions, avoiding the occlusion problem; and when the optical axes of the cameras are orthogonal, the conversion between coordinate systems is more direct and simple.
[0060] During the acquisition period, test instructions can be sent to the subject. Taking medical health monitoring as an example, multiple test scenarios can be set up. Doctors can select appropriate actions based on the patient's physical condition and provide guidance on completion, such as standing on one leg with hands on hips and eyes open, or standing on two legs with arms outstretched and eyes closed. The subject completes the balance test according to the instructions in the test area, and the camera records the entire process, generating a series of image frames during the test. It can be understood that the test area is the field of view that can be captured by all cameras.
[0061] More specifically, the camera and client can communicate, for example, via Bluetooth. The doctor can perform operations on the client, such as searching by patient ID and selecting an action item. Then, the doctor can guide the patient into the test area and initiate the test. During this process, the client generates a temporary data file and receives image data captured by the camera. After the test, the temporary file is automatically stored in a preset path, with the file name uniformly managed using the "patient ID - test ID" format. This image data can then be used for subsequent center of gravity trajectory prediction.
[0062] After obtaining a series of image frames from multiple cameras, frame alignment processing is required to achieve frame synchronization between different cameras and build a unified timing benchmark.
[0063] In one embodiment, frame alignment is performed based on a timestamp matching strategy. The specific process is as follows:
[0064] One of the cameras is the reference camera A, and the other camera is the camera to be aligned B. The timestamp t of the i-th frame image of camera A is obtained. A,i , find the image frame timestamp of camera B and t A,i The closest timestamp t B,j , and time difference |t A,i -t B,j |Compare with the preset value μ:
[0065] If |t A,i -t B,j |≤μ, then the timestamp t A,i Image frame with timestamp t B,j The image frames are aligned frames at the same moment;
[0066] Otherwise, get the camera B at timestamp t A,i The most adjacent timestamps t on both sides B,j and t B,q , calculate the time scale factor α, and construct a virtual frame with the scale factor α as the time stamp t A,i The image frame is aligned with the alignment frame, and the pixels of the virtual frame are F B,virtual ;
[0067] Among them, the calculation formula of the proportional coefficient α is:
[0068]
[0069] Pixel F of the virtual frame B,virtual The calculation formula is:
[0070] F B,virtual =(1-α)·F B,j +α·F B,q (2);
[0071] Wherein, i is the frame index of camera A, i=1, 2, ..., M, M is the number of image frames of camera A, j and q are the frame indexes of camera B.
[0072] Furthermore, the preset value μ is less than or equal to half a frame period. For example, when the frame rate is 30 Hz, μ≤16 ms.
[0073] Through the above method, we can obtain data pairs with consistent timing.
[0074] In the above embodiment, when the actual image does not meet the alignment condition, the scale coefficient is calculated and used to fuse the two frames of image to construct a virtual frame, and the virtual frame is used to achieve frame alignment. In this way, the matching degree between the aligned image frames can be improved.
[0075] S2. Perform spatial alignment on the frame-aligned image frames to obtain the spatial coordinates of the centers of different active parts of the target at the frame alignment time, and obtain the motion vectors of the centers of the corresponding active parts at different times based on the spatial coordinates of the centers of the active parts at different times.
[0076] Taking the human body as an example, it can be divided into nine active parts based on anatomical characteristics, including head, neck, shoulders, chest, abdomen, hips, thighs, knees, and ankles. It can also be divided more finely.
[0077] Since image frames at different perspectives are collected by different cameras at the same time, multi-perspective images are obtained. The multi-perspective images can be spatially aligned to obtain the three-dimensional spatial coordinates of the center of each active part at the moment of frame alignment.
[0078] In one embodiment, spatial alignment is performed on the frame-aligned image frames of camera A and camera B to obtain the spatial coordinates of the centers of different moving parts of the target at the time of frame alignment, including:
[0079] Extracting the image coordinates of the center of the same active part in different image frames; in one embodiment, after obtaining the image frames, conventional image prediction methods can be used to predict each active part and determine the image coordinates of its center. For example, background modeling and edge detection algorithms can be used to extract the contour features of the target. Then, independent masks are constructed for each active part, and the image coordinates of the center of each region are extracted through maximum connected component analysis and bounding box calculation.
[0080] Taking the spatial coordinates as the quantity to be determined, a set of equations is constructed based on the conversion relationship between image coordinates and world coordinates;
[0081] The image coordinates of the center of the same active part in different image frames are substituted into the equation group and solved to obtain the spatial coordinates of the center of the active part.
[0082] Furthermore, the image coordinates can be normalized and then substituted into the equations for solution, thus eliminating the size differences between individuals.
[0083] For example, suppose p A =(u A ,v A ) and p B =(u B ,v B ) are the centers of the same active part in view I A and I B The coordinates in the image need to be converted to a point P = (X, Y, Z) in three-dimensional space. Normalized coordinate transformation can be performed first. The coordinates of the image point obtained after the transformation are as follows:
[0084]
[0085] The normalized image point coordinates and the three-dimensional point P = (X, Y, Z) in the global coordinate system satisfy the relationship described by equation (4). This relationship can be further transformed into a linear equation system, namely equation (5), and the least squares method is used to solve the equation system, and the result is shown in equation (6):
[0086]
[0087] Among them, K A and K B are the internal parameters of camera A and camera B respectively, R A 、T A is the external parameter of camera A, R B 、T B is the external parameter of camera B.
[0088] After obtaining the spatial coordinates of the center of each active part through the above method, the motion vector of the center of each active part can be obtained. The motion vector of a certain active part at time t is (Δx t ,Δy t ,Δz t ), Δx t ,Δy t ,Δz t They are the corresponding x, y, and z-axis displacements of the spatial coordinates of the active part at time t compared to the spatial coordinates at time t-1.
[0089] S3. Taking all motion vectors at the same moment as input, the machine learning model is used to predict the center of gravity position of the target at the corresponding moment to obtain the center of gravity trajectory of the target.
[0090] Specifically, the machine learning model can adopt a conventional regression model. It can be understood that the machine learning model is a trained model, and its construction method can refer to the introduction of Example 3.
[0091] Specifically, the motion vectors from time t1 to time tm are sequentially input into the machine learning model to determine the center of gravity positions at those times, and subsequently the center of gravity trajectory from those times. The resulting center of gravity trajectory exhibits high stability and spatial accuracy, with a prediction error of less than 1 cm, meeting the practical needs of high-precision human posture and balance assessment.
[0092] After obtaining the target's center of gravity trajectory, the balance ability can be evaluated based on the center of gravity trajectory. Table 1 below shows various indicators for balance ability evaluation based on the center of gravity trajectory.
[0093] Table 1 Six balance ability evaluation indicators and calculation formulas
[0094]
[0095] Among them, Count refers to the counting function, Distinct is a function used to remove duplicate elements in a point set, and Ceil represents a function that rounds the value upward. Assume that the predicted center of gravity trajectory dataset is Where N is the total number of center of gravity points, and f is the default sampling frequency of the balance instrument, which is 10Hz.
[0096] Ultimately, the analysis results can be visualized graphically, generating a well-structured and clearly defined balance assessment report. This report allows doctors to quickly understand a patient's center of gravity control ability, providing a scientific basis for developing personalized rehabilitation or training plans. This method not only significantly improves analysis efficiency but also reduces subjective judgment errors, demonstrating excellent clinical practicality.
[0097] Therefore, this method can achieve high-precision prediction of the human body's center of gravity trajectory in a dynamic environment, and output quantitative indicators of balance ability based on the prediction results. It is suitable for scenarios such as rehabilitation medicine, health assessment of the elderly, and intelligent motion assistance. It has significant advantages such as low equipment cost, simple system deployment, and high algorithm reasoning efficiency. It has broad application value and promotion prospects.
[0098] Example 2
[0099] The present invention also provides a center of gravity trajectory prediction system for dynamic targets, such as Figure 3 FIG2 is a schematic structural diagram of a center-of-gravity trajectory prediction system in an embodiment of the present invention, which includes a frame alignment unit, a space alignment unit, and a prediction unit.
[0100] The frame alignment unit is used to perform frame alignment processing on image frames of the same target captured by at least two cameras, where the optical axes of the two cameras are orthogonal;
[0101] The spatial alignment unit is used to perform spatial alignment on the frame-aligned image frames, obtain the spatial coordinates of the centers of different moving parts of the target at the frame alignment time, and obtain the motion vectors of the centers of the corresponding moving parts at different times based on the spatial coordinates of the centers of the moving parts at different times;
[0102] The prediction unit is used to take all motion vectors at each same moment as input, use the input machine learning model to predict the center of gravity position of the target at the corresponding moment, and obtain the center of gravity trajectory of the target.
[0103] Specifically, the frame alignment unit is used to implement the process of step S1 in Example 1, the spatial alignment unit is used to implement the process of step S2 in Example 1, and the prediction unit is used to implement the process of step S2 in Example 3. The specific details can be referred to the introduction of Example 1 and will not be repeated here.
[0104] Example 3
[0105] The present invention also mentions a method for constructing a center of gravity trajectory prediction system for dynamic targets, which is used to construct the center of gravity trajectory prediction system in Example 2, such as Figure 4 The figure shows a flowchart of the steps of a method for constructing a center of gravity trajectory prediction system in one embodiment of the present invention. The construction method includes:
[0106] S01. Acquire image frames of different targets, where each target's image frame is captured by at least two cameras, and the optical axes of the two cameras are orthogonal;
[0107] S02, determining a frame alignment unit, and performing frame alignment processing on image frames of the same target through the frame alignment unit;
[0108] S03. Determine a spatial alignment unit, and use the spatial alignment unit to perform image alignment on the frame-aligned image frames of the same target, obtain the spatial coordinates of the centers of different moving parts of the corresponding target at the frame alignment time, and obtain the motion vectors of the centers of different moving parts of each target at different times based on the spatial coordinates of the centers of the moving parts at different times;
[0109] S04. Take all motion vectors of the same target at the same time as a data point, and multiple data points form a training data set. Use the data points in the training data set as input and the actual center of gravity position of the corresponding target at the corresponding time as a label to train the machine learning model to determine the prediction unit.
[0110] Specifically, the actual center of gravity position can be directly obtained using a balance instrument, such as Figure 5Figure 1 shows a schematic diagram of an embodiment of the present invention, in which a camera captures image frames and a gyroscope is used to determine the actual center of gravity position. It should be noted that after training, the gyroscope is no longer required. Instead, the camera captures images synchronously, extracts motion vector data, and uses these motion vectors as input to a machine learning model, ultimately predicting the center of gravity trajectory.
[0111] In one embodiment, after the machine learning model is trained, hyperparameters of the machine learning model are further constructed.
[0112] The process of building hyperparameters for a machine learning model is as follows.
[0113] S041. Obtain an original data set, where the original data set contains multiple data points consisting of motion vectors. Input the data points in the original data set into a trained machine learning model to obtain a predicted point data set. The predicted points in the predicted point data set are predicted center of gravity positions generated by the machine learning model based on the input data points.
[0114] S042. Calculate the mean square error (MSE) between the predicted point dataset and the expected point dataset, where the expected point in the expected point dataset is the actual center of gravity position.
[0115] Specifically, the center of gravity position is the two-dimensional coordinate on the projection surface, and the prediction point dataset is recorded as The expected point dataset is denoted as N is the number of elements in the dataset, x n 、y n are the positions of the nth predicted points, x′ n , y′ n is the position of the nth expected point. The calculation formula of mean square error MSE is:
[0116]
[0117] S043. Calculate the squared deviation between each predicted point and its corresponding expected point, and select from the original data set data points whose squared deviations from their corresponding predicted points exceed U times the mean square error (MSE) as noise points, where U>1.
[0118] Specifically, the squared deviation DS of the i-th prediction point i The calculation formula is:
[0119] DS i =(x i -x′ i ) 2 +(y i -y′ i ) 2 (8).
[0120] Specifically, U can be 3.
[0121] S044. Determine the optimal denoising ratio using a progressive adjustment strategy within a preset ratio range, and remove the noise point with the highest square deviation value from the original data set according to the optimal denoising ratio to obtain a purified data set after denoising.
[0122] Specifically, the preset ratio range can be 0% to 2%, and the step size used in the progressive adjustment strategy can be 0.4%. The denoising ratio is adjusted in 0.4% steps within the range of 0% to 2%, and the optimal denoising ratio is determined through experiments. After the optimal ratio is determined, the corresponding number of noise points is removed according to the ratio.
[0123] S045. Based on the purified dataset, an optimization method combining hyperparameter grid search and cross-validation is used to automatically search and determine the optimal hyperparameter combination of the machine learning model.
[0124] Specifically, an optimization method combining hyperparameter grid search and 5-fold cross-validation is applied to automatically search and determine the optimal hyperparameter combination of the machine learning model, thereby constructing the final prediction model.
[0125] In one embodiment, multiple different machine learning models may be trained separately, and then the machine learning model with the best performance may be selected as the learning model in the prediction unit.
[0126] The process of selecting a machine learning model is as follows.
[0127] The first step is to calculate the root mean square error RMSE and mean absolute error MAE of the k-th machine learning model and use them as the first confidence score and the second confidence score K is the number of machine learning models.
[0128] Among them, the calculation formulas for the root mean square error RMSE and the mean absolute error MAE are:
[0129]
[0130] RMSE reflects the degree of deviation between the model prediction value and the true value by calculating the root mean square of the prediction error, and MAE gives the same weight to all prediction errors and calculates their average absolute value.
[0131] The second step is to and Input the prediction task fuzzy evaluation function and get the Calculated fuzzy score and based on Calculated fuzzy score The fuzzy evaluation function of the prediction task includes the function f1(p k ), function f2(p k ) and function f3(p k ) multiplied.
[0132] Among them, the function f1(p k ), function f2(p k ) and function f3(p k ) is:
[0133]
[0134] Specifically, the function f1(p k ) is an increasing function, which increases with the increase of confidence score. This is a reward function. The closer the confidence score is to 1, the more rewards there are. Function f2(p k ) is a decreasing function that captures deviations from 1, and the function f3(p k ) is a S-shaped function used to normalize the overall function. f1(p k )*f2(p k )*f3(p k ) is a decreasing function that decreases as the confidence score increases.
[0135] The above three nonlinear functions are the exponential decay function, the tanh function, and the sigmoid function. These three nonlinear functions have different concavities and will produce complementary results, thereby improving the accuracy of model evaluation.
[0136] Step 3: Fuzzy Score and fuzzy scores Sum up and get the comprehensive score FS of the kth machine learning model k .
[0137] Step 4: Select the machine learning model with the lowest comprehensive score as the machine learning model in the prediction unit.
[0138] In a specific embodiment, the multiple different machine learning models include LightGBM (Light Gradient Boosting Machine), XGBoost (eXtreme Gradient Boosting), Random Forest, and Gradient Boosting Decision Trees (GBDT).
[0139] Example 4
[0140] The present invention also relates to an electronic device, comprising a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps of the above method when executing the computer program.
[0141] The electronic device may be a computing device such as a desktop computer, a notebook, a PDA, or a cloud server. The processor may be a central processing unit (CPU), or other general-purpose processors, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The memory may be used to store computer programs and / or modules, and the processor may perform various functions of the electronic device by running or executing the computer programs and / or modules stored in the memory, and calling the data stored in the memory.
[0142] Example 5
[0143] The present invention also relates to a computer-readable storage medium having a computer program stored thereon, which implements the steps of the above method when the computer program is executed by a processor.
[0144] Specifically, the memory may include a high-speed random access memory, and may also include a non-volatile memory, such as a hard disk, a memory, a plug-in hard disk, a smart memory card (Smart Media Card, SMC), a secure digital (Secure Digital, SD) card, a flash card (Flash Card), at least one disk storage device, a flash memory device, or other volatile solid-state storage device.
[0145] The technical features of the above embodiments can be combined in any manner. To simplify the description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification. It should be noted that the phrases "in one embodiment", "for example", "and another example", etc. of the present invention are intended to illustrate the present invention and are not intended to limit the present invention.
[0146] The above embodiments merely illustrate several implementations of the present invention, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the patent application. It should be noted that a person skilled in the art would be able to make numerous variations and improvements without departing from the spirit of the present invention, and all such variations and improvements fall within the scope of protection of the present invention.
Claims
1. A method for predicting center of gravity trajectory of a dynamic target, characterized in that: include: S1. performing frame alignment processing on image frames of the same target captured by at least two cameras, where the optical axes of the two cameras are orthogonal; S2. spatially align the frame-aligned image frames to obtain the spatial coordinates of the centers of different moving parts of the target at the time of frame alignment, and obtain the motion vectors of the centers of the corresponding moving parts at different times based on the spatial coordinates of the centers of the moving parts at different times; S3. Taking all motion vectors at the same moment as input, the machine learning model is used to predict the center of gravity position of the target at the corresponding moment to obtain the center of gravity trajectory of the target.
2. The method for predicting the center of gravity trajectory according to claim 1, wherein: In S1, the frame alignment process includes: One of the cameras is the reference camera A, and the other camera is the camera to be aligned B. The timestamp t of the i-th frame image of camera A is obtained. A,i , find the image frame timestamp of camera B and t A,i The closest timestamp t B,j , and time difference |t A,i -t B,j |Compare with the preset value μ: If |t A,i -t B,j |≤μ, then the timestamp t A,i Image frame with timestamp t B,j The image frames are aligned frames at the same moment; Otherwise, get the camera B at timestamp t A,i The most adjacent timestamps t on both sides B,j and t B,q , calculate the time proportional coefficient The virtual frame is constructed by the scaling factor α as the time stamp t A,i The image frame is aligned with the alignment frame, and the pixels of the virtual frame are F B,virtual =(1-α)·F B,j +α·F B,q ; Wherein, i is the frame index of camera A, i=1, 2, ..., M, M is the number of image frames of camera A, j and q are the frame indexes of camera B.
3. The method for predicting center of gravity trajectory according to claim 1, wherein: In S2, the performing of spatial alignment includes: Extract the image coordinates of the center of the same active part in different image frames; Taking the spatial coordinates as the quantity to be determined, a set of equations is constructed based on the conversion relationship between image coordinates and world coordinates; The image coordinates of the center of the same active part in different image frames are substituted into the equation group and solved to obtain the spatial coordinates of the center of the active part.
4. The method for predicting center of gravity trajectory according to claim 1, wherein: The target is a human body, and the different active parts include the head, neck, shoulders, chest, abdomen, hips, thighs, knees and ankles.
5. A center of gravity trajectory prediction system for dynamic targets, characterized in that: include: a frame alignment unit, configured to perform frame alignment processing on image frames of the same target captured by at least two cameras, wherein the optical axes of the two cameras are orthogonal; A spatial alignment unit is used to perform spatial alignment on the frame-aligned image frames, obtain the spatial coordinates of the centers of different moving parts of the target at the frame alignment time, and obtain the motion vectors of the centers of the corresponding moving parts at different times based on the spatial coordinates of the centers of the moving parts at different times; The prediction unit is used to use all motion vectors at the same moment as input, use the machine learning model to predict the center of gravity position of the target at the corresponding moment, and obtain the center of gravity trajectory of the target.
6. A method for constructing a center of gravity trajectory prediction system for dynamic targets, characterized in that: Constructing the center of gravity trajectory prediction system as claimed in claim 5, the construction method includes: Acquire image frames of different targets, where each target image frame is acquired by at least two cameras, and the optical axes of the two cameras are orthogonal; determining a frame alignment unit, and performing frame alignment processing on image frames of the same target through the frame alignment unit; Determine a spatial alignment unit, perform image alignment on frame-aligned image frames of the same target using the spatial alignment unit, obtain spatial coordinates of centers of different moving parts of the corresponding target at the time of frame alignment, and obtain motion vectors of the centers of different moving parts of each target at different times based on the spatial coordinates of the centers of the moving parts at different times; All motion vectors of the same target at the same time are taken as a data point, and multiple data points form a training data set. The machine learning model is trained with the data points in the training data set as input and the actual center of gravity position of the corresponding target at the corresponding time as the label to determine the prediction unit.
7. The construction method according to claim 6, wherein: After training the machine learning model, the hyperparameters of the machine learning model are further constructed; The process of building hyperparameters for a machine learning model involves: Obtaining an original data set, the original data set comprising a plurality of data points consisting of motion vectors, inputting the data points in the original data set into a trained machine learning model to obtain a predicted point data set, wherein the predicted points in the predicted point data set are predicted center of gravity positions generated by the machine learning model based on the input data points; Calculate the mean square error (MSE) between the predicted point dataset and the expected point dataset, where the expected point in the expected point dataset is the actual center of gravity position; Calculate the squared deviation between each predicted point and its corresponding expected point, and select data points from the original data set whose squared deviation of the corresponding predicted point exceeds U times the mean square error (MSE) as noise points, where U>1; An optimal denoising ratio is determined by a progressive adjustment strategy within a preset ratio range, and noise points with the highest squared deviation value are removed from the original data set according to the optimal denoising ratio to obtain a purified data set after denoising; Based on the purified data set, an optimization method combining hyperparameter grid search and cross-validation is used to automatically search and determine the optimal hyperparameter combination of the machine learning model.
8. The construction method according to claim 6, wherein: The training of the machine learning model to determine the prediction unit includes: training a plurality of different machine learning models separately, and then selecting the machine learning model with the best performance as the learning model in the prediction unit; The process of selecting a machine learning model includes: Calculate the root mean square error RMSE and mean absolute error MAE of the k-th machine learning model and use them as the first confidence score and the second confidence score K is the number of machine learning models; Respectively and Input the prediction task fuzzy evaluation function and get the Calculated fuzzy score and based on Calculated fuzzy score The prediction task fuzzy evaluation function includes the function function and function multiplication; Fuzzy fraction and fuzzy scores Sum up and get the comprehensive score FS of the kth machine learning model k ; The machine learning model with the lowest comprehensive score is selected as the machine learning model in the prediction unit.
9. An electronic device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 or any one of claims 6 to 8 are implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 or any one of claims 6 to 8 are implemented.