Multi-source feature collaborative personalized model generation method and device, equipment and medium
By generating personalized models within action cameras and utilizing incremental learning algorithms and multi-source feature collaboration technology, the systematic bias problem of action cameras under the heterogeneity of individual physiological characteristics and movement habits is solved, achieving high-precision, real-time response personalized adaptation and calibration.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHENRE TECHNOLOGY (SHANGHAI) CO LTD
- Filing Date
- 2026-04-15
- Publication Date
- 2026-07-21
AI Technical Summary
Existing general AI models for action cameras exhibit significant systematic biases due to heterogeneity in individual physiological characteristics and movement habits, as well as non-ideal shooting angles. They lack dynamic calibration and spatiotemporal alignment strategies for multi-source data, resulting in insufficient model accuracy and adaptability.
Personalized models are generated within the action camera using incremental learning algorithms. Parameter compensation is performed using data from the camera terminal and external sensors to achieve multi-source feature collaboration. This includes spatiotemporal alignment of motion quantization indicators of the attitude estimation model with external sensor data and calculation of deviation vectors, generating a personalized model adapted to the target object.
It improves the model's inference accuracy, robustness, and environmental adaptability, ensures real-time response of calibration and inference, reduces communication latency, and enhances the accuracy and stability of personalized adaptation.
Smart Images

Figure CN122435501A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of AI camera model personalized calibration technology, and relates to a method, device, equipment and medium for generating personalized models through multi-source feature collaboration. Background Technology
[0002] Current mainstream solutions for action cameras rely on general artificial intelligence models for pose recognition and power consumption prediction, but significant systematic biases are prevalent in real-world applications. This stems from the heterogeneity of individual physiological characteristics (such as basal metabolic rate) and the specificity of movement habits, coupled with the loss of depth information due to non-ideal shooting angles, severely limiting model accuracy. Furthermore, existing calibration techniques are largely limited to static factory parameters, lacking a dynamic model evolution mechanism based on long-term user data, and failing to construct effective multi-source data spatiotemporal alignment and compensation strategies, making personalized adaptation difficult to achieve.
[0003] Therefore, how to generate a target-specific model within an action camera to output high-fidelity physiological data under the constraints of personalized adaptation and dynamic calibration has become a key technical problem that urgently needs to be solved. Summary of the Invention
[0004] This application provides a method, apparatus, device, and medium for generating personalized models based on multi-source feature collaboration, which is used to generate a target-oriented model within a motion camera to output high-fidelity physiological data under the constraints of personalized adaptation and dynamic calibration.
[0005] In a first aspect, this application provides a method for generating a personalized model through multi-source feature collaboration. The method includes: acquiring a motion video stream of a target object based on a camera terminal; outputting a motion quantization index corresponding to the motion video stream based on a pose estimation model; acquiring external sensor data synchronized with the motion video stream; determining a prediction deviation vector between the motion quantization index and the external sensor data in response to a preset calibration trigger condition; and performing parameter compensation on the pose estimation model using an incremental learning algorithm based on the prediction deviation vector to generate a personalized model adapted to the target object.
[0006] In this application, an incremental learning algorithm is used to continuously absorb bias data for online self-correction. As the usage period increases, the model adaptively approaches the user's movement habits and physiological characteristics, breaking through the limitations of traditional static models. Relying on edge-based lightweight incremental learning, local parameter compensation is achieved, avoiding cloud data uploads. This significantly reduces communication latency while ensuring data privacy, ensuring real-time response for calibration and inference. With the help of the parameter compensation mechanism, the high-frequency and high-precision characteristics of external sensors are transferred to the visual model, making up for the shortcomings of single visual perception, achieving deep fusion of multi-source data, and significantly improving the model's inference accuracy, robustness, and environmental adaptability.
[0007] In one implementation of the first aspect, the step of outputting the motion quantization index corresponding to the motion video stream based on the attitude estimation model includes: extracting the spatiotemporal trajectory of key points in the motion video stream based on the attitude estimation model; performing data processing operations on the spatiotemporal trajectory of the key points; and outputting the motion quantization index.
[0008] In one implementation of the first aspect, determining the prediction deviation vector between the motion quantization index and the external sensor data in response to a preset calibration trigger condition includes: performing a time-domain alignment operation on the motion quantization index and the external sensor data to obtain a first motion quantization index vector corresponding to the motion quantization index and a first external data vector corresponding to the external sensor data; performing a spatial alignment operation on the first motion quantization index vector and the first external data vector to obtain a second motion quantization index vector corresponding to the first motion quantization index and a second external data vector corresponding to the first external data; and determining the prediction deviation vector based on the second motion quantization index vector and the second external data vector.
[0009] In one implementation of the first aspect, the step of using an incremental learning algorithm to compensate the parameters of the attitude estimation model based on the predicted deviation vector to generate a personalized model adapted to the target object includes: extracting multiple first deviation vectors containing a preset action cycle from the predicted deviation vector; performing statistical smoothing on the multiple first deviation vectors to obtain multiple first smoothed vectors; using the multiple first smoothed vectors as residual supervision signals, performing batch-dimensional parameter compensation on the attitude estimation model via the incremental learning algorithm to generate a personalized model.
[0010] In one implementation of the first aspect, the preset calibration trigger condition includes: detecting an explicit synchronization command generated by the camera terminal's interactive interface.
[0011] In one implementation of the first aspect, the preset calibration triggering condition further includes: the prediction deviation vector between the motion quantization index and the external sensor data exceeds a preset threshold for N consecutive frames.
[0012] In one implementation of the first aspect, the parameters include: a spatial scale scaling factor, a dynamic conversion coefficient, and a viewpoint weight bias, wherein the parameters are nonlinearly coupled.
[0013] Secondly, this application provides a personalized model generation device based on multi-source feature collaboration, characterized in that the device comprises: a first data acquisition module for acquiring motion video streams of a target object based on a camera terminal; a motion quantization index output module for outputting motion quantization indexes corresponding to the motion video stream based on a posture estimation model; a second data acquisition module for acquiring external sensor data synchronized with the motion video stream; a deviation vector determination module for determining a prediction deviation vector between the motion quantization index and the external sensor data in response to a preset calibration trigger condition; and a personalized model generation module for performing parameter compensation on the posture estimation model using an incremental learning algorithm based on the prediction deviation vector to generate a personalized model adapted to the target object.
[0014] Thirdly, embodiments of this application provide a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the personalized model generation method for multi-source feature collaboration as described in any of the first aspects of this application.
[0015] Fourthly, embodiments of this application provide an electronic device, the electronic device comprising: a memory storing a computer program; and a processor communicatively connected to the memory, which, when the computer program is invoked, executes the personalized model generation method for multi-source feature collaboration as described in the first aspect of this application.
[0016] As described above, the method, apparatus, device, and medium for generating personalized models based on multi-source feature collaboration described in this application have the following beneficial effects:
[0017] 1) This application utilizes an incremental learning algorithm, enabling the model to continuously absorb bias data for online self-correction. As the usage period increases, the model adaptively approaches the user's movement habits and physiological characteristics, overcoming the limitations of traditional static models. It relies on lightweight incremental learning on the device side to achieve localized parameter compensation, avoiding cloud data uploads. This significantly reduces communication latency while ensuring data privacy and guarantees real-time response for calibration and inference. Through the parameter compensation mechanism, the high-frequency and high-precision characteristics of external sensors are transferred to the visual model, compensating for the shortcomings of single visual perception, achieving deep fusion of multi-source data, and significantly improving the model's inference accuracy, robustness, and environmental adaptability.
[0018] 2) In this embodiment of the application, by transforming the abstract "key point spatiotemporal trajectory" into calculable and comparable "motion quantification indicators" (such as joint angle, movement speed, and range of motion), a standardized data foundation is provided for subsequent prediction deviation vector calculation with external sensor data. After being transformed into motion quantification indicators such as joint angle and relative speed, the data no longer depends on the absolute pixel position. Regardless of the distance between the user and the camera or how the camera angle is finely adjusted, the range of knee joint angle change in the standard "squatting action" is consistent. This characteristic significantly improves the generalization ability of the posture estimation model in different scenarios and reduces non-systematic errors caused by environmental changes.
[0019] 3) By transforming “incomparable” heterogeneous data into “comparable” homogeneous data through spatiotemporal alignment operations, it minimizes the systematic interference in multi-source data fusion, making the generated prediction bias vector an accurate, reliable, and quantifiable optimization target, laying a solid foundation for the subsequent construction of a high-precision personalized attitude estimation model.
[0020] 4) In this embodiment, by extracting multiple first deviation vectors of a preset action cycle and performing statistical smoothing, the prediction deviation of a single action often includes random errors from visual detection. By eliminating random noise and outliers, the resulting first smoothed vector reflects the user's systematic deviation trend under that action rather than random errors, making model calibration more stable and avoiding oscillations in model parameters caused by single data anomalies. Using the first smoothed vector as a residual supervision signal and performing parameter compensation in a batch dimension, compared to online learning for each original data point, batch learning utilizes aggregated gradient information. The smoothed residual signal has a clearer direction, the loss function surface is smoother, and the model can find the optimal solution faster. The batch dimension update strategy avoids overfitting the model to the latest data points, retaining the model's general characteristics while correcting deviations, ensuring the generalization ability of the pose estimation model when handling other actions. Attached Figure Description
[0021] Figure 1A The diagram shows an application scenario corresponding to the personalized model generation method for multi-source feature collaboration provided in the embodiments of this application.
[0022] Figure 1B The flowchart shown is a process for generating a personalized model using multi-source feature collaboration, as provided in an embodiment of this application.
[0023] Figure 2 The flowchart shown is a motion quantization index corresponding to the motion video stream output by the attitude estimation model provided in the embodiments of this application.
[0024] Figure 3The flowchart shown is a process for determining the prediction deviation vector provided in an embodiment of this application.
[0025] Figure 4 The flowchart shown is a process for generating a personalized model adapted to the target object, as provided in an embodiment of this application.
[0026] Figure 5 This diagram illustrates the evolutionary state of the personalized model provided in this embodiment after parameter compensation.
[0027] Figure 6 The image shown is a comparison chart of the progress jumps before and after calibration, provided in an embodiment of this application.
[0028] Figure 7 This is a personalized model generation device for multi-source feature collaboration provided in an embodiment of this application.
[0029] Figure 8 The diagram shown is a structural diagram of an electronic device provided in an embodiment of this application.
[0030] Component designation explanation
[0031] S11~S15 step 74 Deviation Vector Determination Module S21~S22 step 75 Personalized model generation module S31~S33 step 80 electronic devices S41~S43 step 81 processor 70 Personalized model generation device based on multi-source feature collaboration 82 Non-volatile storage media 71 First data acquisition module 83 System bus 72 Quantitative Indicator Output Module 84 Internal memory 73 Second data acquisition module 85 Network interface Detailed Implementation
[0032] The following specific examples illustrate the implementation of this application. Those skilled in the art can easily understand other advantages and effects of this application from the content disclosed in this specification. This application can also be implemented or applied through other different specific embodiments, and various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of this application. It should be noted that, unless otherwise specified, the following embodiments and features in the embodiments can be combined with each other.
[0033] It should be noted that the illustrations provided in the following embodiments are only schematic representations of the basic concept of this application. Therefore, the drawings only show the components related to this application and are not drawn according to the actual number, shape and size of the components in the actual implementation. In the actual implementation, the form, quantity and proportion of each component can be arbitrarily changed, and the layout of the components may also be more complex.
[0034] like Figure 1AAs shown in the diagram, this embodiment provides an application scenario illustration of a method, apparatus, device, and medium for generating personalized models based on multi-source feature collaboration. Specifically, it includes a camera terminal and an external sensor, wherein the camera terminal and the external sensor are communicatively connected. The camera terminal is an AI camera, and the external sensor is a wearable device. The external sensor is used to collect external sensor data of the target object in real time and send the collected external sensor data to the camera terminal. The camera terminal determines the prediction deviation vector between the motion quantization index output by its own attitude estimation model and the external sensor data through preset calibration trigger conditions, and uses an incremental learning algorithm to compensate the parameters of the attitude estimation model based on the prediction deviation vector, generating a personalized model adapted to the target object.
[0035] The technical solutions in the embodiments of this application will be described in detail below with reference to the accompanying drawings.
[0036] like Figure 1B As shown in the flowchart, this application provides a method for generating personalized models through multi-source feature collaboration. Figure 1B As shown, the personalized model generation method for multi-source feature collaboration provided in this application includes the following steps S11 to S15.
[0037] S11, based on the camera terminal, acquires motion video streams of the target object.
[0038] For example, the camera terminal can be an AI camera.
[0039] For example, sports video streams include playing badminton, playing basketball, running, doing yoga, etc.
[0040] It should be noted that the motion video streams listed in the above examples are only used as illustrative examples. In actual applications, other suitable motion video streams can be selected based on specific application requirements, and this application does not impose any restrictions on this.
[0041] S12, based on the attitude estimation model, output the motion quantization index corresponding to the motion video stream.
[0042] Among them, the quantitative indicators of exercise are a comprehensive set of data, which not only include internal load parameters that characterize the body's physiological load and energy metabolism, such as individual basal metabolic rate, heart rate, and maximum oxygen uptake, but also kinematic parameters that describe the biomechanical characteristics of movement, such as joint angle and angular velocity, as well as performance indicators that reflect competitive performance and external load, such as speed, power, and training volume.
[0043] S13, acquire external sensor data synchronized with the motion video stream.
[0044] For example, external sensor data synchronized with the motion video stream can be acquired based on external sensors.
[0045] It should be noted that external sensors include, but are not limited to, smartwatches, heart rate belts, body-worn inertial sensors, or smart rackets.
[0046] S14, in response to a preset calibration trigger condition, determine the prediction deviation vector between the motion quantification index and the external sensor data.
[0047] In some embodiments, the preset calibration triggering condition includes: detecting an explicit synchronization command generated by the camera terminal's interactive interface.
[0048] For example, the user can actively initiate a synchronization operation signal through the screen of the camera device, which serves as the trigger for time alignment or state consistency between the camera terminal and the external sensor data in the external sensor.
[0049] In some embodiments, the preset calibration triggering condition further includes: the prediction deviation vector between the motion quantization index and the external sensor data exceeds a preset threshold for N consecutive frames.
[0050] For example, in the preset calibration trigger condition, if the prediction deviation vector between the motion quantization index and the external sensor data exceeds a preset threshold for N consecutive frames, the system will no longer rely on the user to manually click to synchronize (such as the aforementioned "explicit synchronization command"), but will automatically sense the error and correct itself.
[0051] S15, Based on the predicted deviation vector, the attitude estimation model is compensated for parameters using an incremental learning algorithm to generate a personalized model adapted to the target object.
[0052] This application effectively addresses the technical bottlenecks of existing general attitude estimation models, such as significant systematic errors, fixed calibration parameters that are difficult to dynamically evolve, and the lack of multi-source data alignment compensation mechanisms. Through incremental learning algorithms, the model continuously absorbs deviation data for online self-correction, adaptively approximating the user's movement habits and physiological characteristics as the usage period increases, overcoming the limitations of traditional static models. It achieves localized parameter compensation through lightweight incremental learning on the edge, avoiding cloud data uploads, significantly reducing communication latency while ensuring data privacy, and ensuring real-time response for calibration and inference. Furthermore, by leveraging the parameter compensation mechanism, the high-frequency and high-precision characteristics of external sensors are transferred to the visual model, compensating for the shortcomings of single-source visual perception, achieving deep fusion of multi-source data, and significantly improving the model's inference accuracy, robustness, and environmental adaptability.
[0053] like Figure 2As shown in the figure, this application embodiment provides a flowchart of outputting motion quantization indicators corresponding to a motion video stream based on a pose estimation model, as follows: Figure 2 As shown in the embodiments of this application, the method for outputting motion quantization indicators corresponding to motion video streams based on attitude estimation models includes the following steps S21 to S22.
[0054] S21, extract the spatiotemporal trajectories of key points in the motion video stream based on the attitude estimation model.
[0055] For example, extracting the spatiotemporal trajectory of key points in a motion video stream based on a pose estimation model is a transformation process from unstructured image data to structured time series data. The specific steps are as follows:
[0056] I. Visual Feature Extraction and Key Point Localization
[0057] 1) The motion video stream is fed frame by frame into the backbone network of the pose estimation model. The backbone network extracts multi-scale feature maps of the images in the motion video stream through convolutional layers, captures low-level visual features such as edges and textures, and then abstracts high-level semantic features.
[0058] 2) For a predefined skeletal topology (e.g., 17 key points: nose, shoulder, elbow, wrist, knee, ankle, etc.), the backbone network generates a probability heatmap for each key point. The peak position of the pixel value in the probability heatmap is the highest probability position of the key point in the image coordinate system.
[0059] 3) The peak values of the heatmap are analyzed by the maximum suppression algorithm and converted into pixel coordinates (x, y). At the same time, the confidence score of each key point is output, and noise points with confidence scores below the threshold (such as key points that are occluded or outside the frame) are removed to ensure the reliability of the key points.
[0060] II. Temporal Relationships and Identity Maintenance
[0061] 1) Multi-target tracking: If there are multiple moving objects in the motion video stream, the system uses target detection boxes or key point feature vectors, combined with the Hungarian algorithm or Kalman filter, to assign a unique identity ID to the skeleton of each detected moving object, ensuring that the skeleton of the moving object with ID=1 always represents the same moving object in consecutive frames, thus preventing trajectory confusion.
[0062] 2) Based on human kinematic constraints, discrete key points are connected according to preset skeletal connections (e.g., "shoulder-elbow-wrist") to construct a single-frame human skeleton structure, forming data in the spatial dimension.
[0063] III. Spatiotemporal Trajectory Construction
[0064] 1) Stack the skeleton data of key points from consecutive T frames along the time axis to generate a three-dimensional tensor matrix. Where T represents the time dimension (number of frames / time steps), K represents the number of keypoints, and R represents the feature dimension.
[0065] 2) For any key point k (e.g., "right wrist"), its coordinate sequence on the time axis t. This constitutes the spatiotemporal trajectory of the key point, and this trajectory fully records the spatial displacement path of the body part.
[0066] S22, perform data processing operations on the spatiotemporal trajectory of the key points and output motion quantification indicators.
[0067] For example, data processing operations on the spatiotemporal trajectories of the key points include noise reduction filtering, geometric modeling, etc.
[0068] For example, in the process of data processing of key point spatiotemporal trajectories, the denoising and filtering steps include: using the polynomial least squares method to fit the data within the sliding window, which can effectively filter out high-frequency noise while better preserving the peak features of the trajectory (such as the turning point of a waving gesture), avoiding signal distortion. A smooth and continuous key point position sequence P(t) can be output, providing a stable foundation for subsequent calculus operations.
[0069] For example, in the process of data processing operations on the spatiotemporal trajectories of key points, the geometric modeling steps include:
[0070] 1) Define bone vectors based on the topological structure of the human skeleton. For example, the upper arm vector. Forearm vector .
[0071] 2) Joint angle calculation: The joint angle is obtained by calculating the angle between two adjacent bone vectors. The corresponding formula is: For example, calculating the elbow flexion angle. When the arm is fully extended, the angle is close to 180 degrees; when fully bent, the angle is less than 90 degrees.
[0072] 3) Trunk tilt and posture angle: Calculate the angle between the spinal vector and the direction of gravity to determine whether the body is leaning forward, bending to the side, or in balance.
[0073] 4) Output motion quantification indicators: joint angle sequence (e.g., elbow angle, knee angle), body orientation angle, relative length ratio of bones, etc.
[0074] In this embodiment, by transforming the abstract "key point spatiotemporal trajectory" into calculable and comparable "motion quantification indicators" (such as joint angles, movement speed, and range of motion), a standardized data foundation is provided for subsequent prediction deviation vector calculation with external sensor data. After being transformed into motion quantification indicators such as joint angles and relative speeds, the data no longer depends on absolute pixel positions. Regardless of the distance between the user and the camera or how the camera angle is finely adjusted, the range of knee joint angle changes in a standard "squat" is consistent. This characteristic significantly improves the generalization ability of the posture estimation model in different scenarios and reduces non-systematic errors caused by environmental changes.
[0075] like Figure 3 As shown in the figure, this application provides a flowchart for determining the prediction deviation vector, as follows: Figure 3 As shown, the method for determining the prediction deviation vector provided in this application includes the following steps S31 to S33.
[0076] S31, perform time-domain alignment operation on the motion quantization index and the external sensor data to obtain the first motion quantization index vector corresponding to the motion quantization index and the first external data vector corresponding to the external sensor data, respectively.
[0077] Specifically, the video frame rate of the motion video stream (e.g., 30fps) and the sampling rate of the external sensor data (e.g., 200Hz) are different, requiring temporal alignment of the motion quantization metrics with the external sensor data.
[0078] For example, the steps for time-domain alignment of motion quantization metrics with external sensor data are as follows:
[0079] 1) Constructing the interpolation function
[0080] Let the sequence corresponding to the quantitative index of exercise be . The corresponding timestamp is: Since human movement is continuous, cubic spline interpolation or piecewise linear interpolation algorithms can be used. Spline interpolation can ensure the smoothness of the curve and better conforms to the physical laws of human joint movement.
[0081] 2) Perform interpolation operations
[0082] Specifically, using high-frequency timestamps from external sensors Using a given set of data points as a baseline, the interpolation function is substituted to calculate the estimated values of the motion quantization index sequence at these high-frequency moments. For example, 30 discrete data points can be expanded into 200 continuous data points, filling the gaps between video frames.
[0083] 3) Vector Construction and Output
[0084] Specifically, the interpolated motion quantization index (e.g., the interpolated joint angle sequence) is constructed as the first motion quantization index vector. Sensor data (e.g., joint angles measured by an IMU) corresponding to the physical meaning of visual indicators are selected to construct a first external data vector of the same dimension. This outputs a vector with two identical dimensions and strictly aligned timestamps.
[0085] S32, perform spatial alignment operation on the first motion quantization index vector and the first external data vector to obtain the second motion quantization index vector corresponding to the first motion quantization index and the second external data vector corresponding to the first external data, respectively.
[0086] Specifically, because the camera terminal and the external sensor are located in different physical spaces and have different measurement systems, even after time alignment, their data still suffer from differences in coordinate system definitions, different reference bases, and inconsistent dimensions. The core purpose of spatial alignment is to eliminate these heterogeneities, enabling accurate comparisons between the two within the same mathematical space.
[0087] For example, the steps for spatial alignment of the first motion quantization index vector and the first external data vector are as follows:
[0088] I. Unified Coordinate System Definition 1) Axial Mapping and Transformation
[0089] Based on the preset external sensor wearing position and the camera terminal's vision, a rotation matrix is established. . The first external data
[0090] Multiplying a vector by a rotation matrix maps its coordinate axes to the coordinate system of the attitude estimation model within the camera terminal. For example, if the attitude estimation model defines "upward as the positive Y-axis" while the external sensor defines "upward as the positive Z-axis," then a matrix transformation is needed to unify the axes.
[0091] 2) Projection Transformation
[0092] Specifically, if the pose estimation model only outputs 2D keypoints, while external sensors provide 3D pose data, dimensionality reduction projection is required. Using the camera intrinsic parameter matrix, the sensor's 3D spatial coordinates are projected onto the image plane to generate 2D coordinates consistent with the visual data perspective, ensuring that the two match in spatial dimensions.
[0093] II. Alignment of Semantic Definitions and Measurement Benchmarks
[0094] Even if the coordinate axes are aligned, the attitude estimation model within the camera terminal and the external sensor may differ in their definitions of the same "quantitative index." Therefore, it is necessary to align the semantic definitions and measurement benchmarks of the attitude estimation model and the external sensor. The specific steps are as follows:
[0095] 1) Zero-point calibration
[0096] The calibration window is selected for the period when the user maintains a static standard posture (such as T-Pose or natural standing). The mean of the visual data stream and the mean of the external sensor data are calculated at this time. The difference between the mean of the visual data stream and the mean of the external sensor data is obtained as the zero-point offset. The first external data vector is subtracted from the offset to align its reference with the first motion quantization index vector.
[0097] 2) Dimensional unification
[0098] Specifically, if speed comparison is required, the visual "pixels / frame" needs to be converted to "meters / second" using camera calibration parameters.
[0099] Alternatively, both can be unified and normalized to the dimensionless interval of [0, 1] or [−1, 1] to eliminate calculation errors caused by differences in physical units.
[0100] III. Spatial Transformation Matrix Generation and Vector Output
[0101] After the coordinate system transformation and datum calibration described above, the system generates a spatial transformation matrix. Used to finally generate aligned
[0102] The vector.
[0103] 1) Generate the second external data vector
[0104] A linear (or nonlinear) transformation is performed on the first external data vector using a transformation matrix. The formula is expressed as:
[0105]
[0106] in, Represents the first external data vector. Represents the second external data vector. This indicates the zero-point offset.
[0107] 2) Generate the second motion quantization index vector
[0108] Specifically, the first motion quantification index vector undergoes necessary unit unification or normalization to maintain its physical meaning, but its numerical scale is adapted to the second external data vector.
[0109] S33, determine the prediction deviation vector based on the second motion quantization index vector and the second external data vector.
[0110] For example, the second motion quantization index vector and the second external data vector are subtracted to obtain the prediction deviation vector.
[0111] In this embodiment, the spatiotemporal alignment operation transforms "incomparable" heterogeneous data into "comparable" homogeneous data. It minimizes the systematic interference in multi-source data fusion, making the generated prediction bias vector an accurate, reliable, and quantifiable optimization target, thus laying a solid foundation for the subsequent construction of a high-precision personalized attitude estimation model.
[0112] like Figure 4 As shown in the figure, this application embodiment provides a flowchart for generating a personalized model adapted to the target object, as follows: Figure 4 As shown, the method for generating a personalized model adapted to a target object provided in this application embodiment includes the following steps S41 to S43.
[0113] S41, extract multiple first deviation vectors containing preset action cycles from the predicted deviation vector.
[0114] For example, a preset action cycle represents the complete technical cycle of a movement action.
[0115] S42, statistical smoothing is performed on the multiple first deviation vectors respectively to obtain multiple first smoothed vectors.
[0116] Specifically, since frame-by-frame compensation introduces sensor noise that causes image jitter, this application collects the first deviation vector of N frames (e.g., 150-300 frames, about one complete technical action cycle) through a sliding window, performs statistical smoothing, and then updates the parameters of the attitude estimation model all at once to ensure the consistency and stability of the inference output.
[0117] For example, statistical smoothing is applied to the first deviation vector to obtain the expression for the first smoothed vector:
[0118]
[0119] in, Indicates the regulating factor. This represents a penalty term based on the user's static characteristics. This represents the second motion quantification index vector. Let represent the second external data vector, where the first deviation vector is obtained by taking the difference between the second motion quantization index vector and the second external data vector.
[0120] S43, using multiple first smooth vectors as residual supervision signals, perform batch dimension parameter compensation on the pose estimation model via the incremental learning algorithm to generate a personalized model.
[0121] For example, the expression corresponding to the parameter compensation is:
[0122]
[0123] in, This represents the parameters after compensation. This represents the parameters before compensation. Indicates the learning rate. This indicates the preset action cycle. Indicates parameters The corresponding first smooth vector.
[0124] In some embodiments, the parameters include: spatial scale scaling factor, dynamic conversion coefficient and viewpoint weight bias, wherein the parameters are nonlinearly coupled.
[0125] For example, when physiological data from external sensors indicates a rapidly increasing heart rate, while the motion quantification indicators in the motion video stream show a decrease in the amplitude of movement, the system will prioritize increasing the "fatigue feature weight" rather than simply increasing the spatial scaling factor. This is a joint compensation based on a multi-dimensional feature matrix.
[0126] Among them, the spatial scale scaling factor is used to correct the estimation error of limb length by the visual model; the dynamic conversion coefficient is used to correct the nonlinear weight when mapping bone angular velocity to energy consumption (for users with different physical fitness levels); and the viewpoint weight bias is used to weight and compensate for the confidence of hidden key points for non-ideal shooting angles (such as non-45-degree side shots).
[0127] In this embodiment, by extracting multiple first deviation vectors from a preset action cycle and performing statistical smoothing, the prediction deviation of a single action often includes random errors from visual detection. By eliminating random noise and outliers, the resulting first smoothed vector reflects the user's systematic deviation trend under that action rather than random errors, making model calibration more stable and avoiding oscillations in model parameters caused by single data anomalies. Using the first smoothed vector as a residual supervision signal and performing parameter compensation in a batch dimension, compared to online learning for each original data point, batch learning utilizes aggregated gradient information. The smoothed residual signal has a clearer direction, the loss function surface is smoother, and the model can find the optimal solution faster. The batch dimension update strategy avoids overfitting the model to the latest data points, preserving the model's general characteristics while correcting deviations, ensuring the generalization ability of the pose estimation model when handling other actions.
[0128] Please see Figure 5 , Figure 5 This diagram illustrates the evolutionary state of the personalized model provided in this embodiment after parameter compensation. Figure 5 It can be seen that, Figure 5 The left side shows the software interface containing the personalized model. Figure 5 The right side displays the software interface screenshots before and after calibration, as well as a graph showing the number of calibrations and the model evolution value. Figure 5 It can be seen that with each calibration, the model evolves a little further until the model shows 100% evolution.
[0129] Please see Figure 6 , Figure 6 The image shown is a comparison chart of the progress jumps before and after calibration, provided in an embodiment of this application. Figure 6 It can be seen that,
[0130] In response to the preset calibration trigger condition, as the number of calibrations increases, the external sensor data synchronized with the motion video stream, i.e., the synchronized physiological data, is transmitted to the posture estimation model for parameter calibration to generate a personalized model adapted to the target object. The motion quantification indicators (including motion power consumption) predicted by the posture estimation model gradually converge to the physiological data corresponding to the external sensor data until significant convergence is achieved, thus optimizing the personalized model of the target object.
[0131] The scope of protection for the multi-source feature collaborative personalized model generation method described in this application is not limited to the execution order of the steps listed in this embodiment. Any solution implemented by adding, subtracting, or replacing steps in the prior art based on the principles of this application is included within the scope of protection of this application.
[0132] This application also provides a personalized model generation device for multi-source feature collaboration. The personalized model generation device for multi-source feature collaboration can implement the personalized model generation method for multi-source feature collaboration described in this application. However, the implementation device for the personalized model generation method for multi-source feature collaboration described in this application includes, but is not limited to, the structure of the personalized model generation device for multi-source feature collaboration listed in this embodiment. All structural modifications and substitutions of the prior art made based on the principles of this application are included within the protection scope of this application.
[0133] like Figure 7 As shown, in one embodiment, the multi-source feature collaborative personalized model generation device 70 of this application includes a first data acquisition module 71, a motion quantification index output module 72, a second data acquisition module 73, a deviation vector determination module 74, and a personalized model generation module 75.
[0134] The first data acquisition module 71 is used to acquire motion video streams of the target object based on a camera terminal. The motion quantization index output module 72 is used to output motion quantization indices corresponding to the motion video stream based on a posture estimation model. The second data acquisition module 73 is used to acquire external sensor data synchronized with the motion video stream. The deviation vector determination module 74 is used to determine the predicted deviation vector between the motion quantization indices and the external sensor data in response to a preset calibration trigger condition. The personalized model generation module 75 is used to perform parameter compensation on the posture estimation model using an incremental learning algorithm based on the predicted deviation vector, generating a personalized model adapted to the target object.
[0135] The structure and principle of the first data acquisition module 71, the motion quantification index output module 72, the second data acquisition module 73, the deviation vector determination module 74, and the personalized model generation module 75 correspond one-to-one with the steps in the above-mentioned personalized model generation method of multi-source feature collaboration, so they will not be described in detail here.
[0136] In the embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, or methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative. For instance, the division of modules / units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple modules or units may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection of apparatuses or modules or units may be electrical, mechanical, or other forms.
[0137] The modules / units described as separate components may or may not be physically separate. The components shown as modules / units may or may not be physical modules; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules / units can be selected to achieve the objectives of the embodiments of this application, depending on actual needs. For example, the functional modules / units in the various embodiments of this application may be integrated into one processing module, or each module / unit may exist physically separately, or two or more modules / units may be integrated into one module / unit.
[0138] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0139] This application also provides an electronic device. Figure 8 The diagram shown is a structural schematic of an electronic device 80 in one embodiment of this application. The multi-source feature collaborative personalized model generation method provided in this embodiment can be applied to... Figure 8 The electronic devices shown are 80, but not limited to these. For example... Figure 8 As shown, the electronic device 80 includes a processor 81, a memory, a system bus 83, and a network interface 85. The memory may include a non-volatile storage medium 82 and internal memory 84.
[0140] The non-volatile storage medium 82 can store an operating system and a computer program. The computer program includes program instructions that, when executed, cause the processor to perform any of the multi-source feature collaborative personalized model generation methods provided in the embodiments of this application.
[0141] The processor provides computing and control capabilities, supporting the operation of the entire computer device.
[0142] The internal memory 84 provides an environment for the execution of a computer program in a non-volatile storage medium. When the computer program is executed by the processor, it enables the processor to execute any of the multi-source feature collaborative personalized model generation methods provided in the embodiments of this application.
[0143] This network interface 85 is used for network communication, such as sending assigned tasks. Those skilled in the art will understand that... Figure 8 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0144] It should be understood that processor 81 can be a Central Processing Unit (CPU), but it can also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among these, the general-purpose processor can be a microprocessor or any conventional processor.
[0145] The electronic device 80 in this application embodiment may include terminal devices such as tablet computers, laptop computers, mobile phones, supercomputers, and smart wearable devices. It can also be applied to databases, servers, and service response systems based on terminal artificial intelligence. This application embodiment does not impose any restrictions on the specific type of electronic device.
[0146] For example, electronic devices can be stations (STAION, ST) in WLANs, cellular phones, cordless phones, Session Initiation Protocol (SIP) phones, Wireless Local Loop (WLL) stations, handheld devices with wireless communication capabilities, computing devices or other processing devices connected to a wireless modem, computers, laptops, handheld communication devices, handheld computing devices, and / or other devices for communicating over wireless systems, as well as next-generation communication systems, such as mobile terminals in 5G networks, mobile terminals in future evolved Public Land Mobile Networks (PLMNs), or mobile terminals in future evolved Non-terrestrial Networks (NTNs).
[0147] This application also provides a computer-readable storage medium. Those skilled in the art will understand that all or part of the steps in the methods of the above embodiments can be implemented by a program instructing a processor. The program can be stored in a computer-readable storage medium, which is a non-transitory medium, such as random access memory, read-only memory, flash memory, hard disk, solid-state drive, magnetic tape, floppy disk, optical disk, and any combination thereof. The storage medium can be any available medium accessible to a computer or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., digital video disc (DVD)), or a semiconductor medium (e.g., solid-state disk (SSD)).
[0148] This application embodiment may also provide a computer program product comprising one or more computer instructions. When the computer instructions are loaded and executed on a computing device, all or part of the processes or functions described in this application embodiment are generated. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions may be transmitted from one website, computer, or data center to another via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means.
[0149] When the computer program product is executed by a computer, the computer performs the method described in the foregoing method embodiments. The computer program product can be a software installation package; when the foregoing method is required, the computer program product can be downloaded and executed on the computer.
[0150] The descriptions of the processes or structures corresponding to the above figures each have their own emphasis. For parts of a process or structure that are not described in detail, please refer to the relevant descriptions of other processes or structures.
[0151] The above embodiments are merely illustrative of the principles and effects of this application and are not intended to limit this application. Any person skilled in the art can modify or alter the above embodiments without departing from the spirit and scope of this application. Therefore, all equivalent modifications or alterations made by those skilled in the art without departing from the spirit and technical concept disclosed in this application should still be covered by the claims of this application.
Claims
1. A method for generating personalized models through multi-source feature collaboration, characterized in that, The method includes: Based on the acquisition of motion video streams of the target object using a camera terminal; The motion quantization index corresponding to the motion video stream is output based on the attitude estimation model. Acquire external sensor data synchronized with the motion video stream; In response to a preset calibration trigger condition, a prediction deviation vector between the motion quantification index and the external sensor data is determined; Based on the predicted deviation vector, the attitude estimation model is compensated for parameters using an incremental learning algorithm to generate a personalized model adapted to the target object.
2. The method according to claim 1, characterized in that, The motion quantization metrics corresponding to the motion video stream output by the pose estimation model include: Based on the attitude estimation model, the spatiotemporal trajectories of key points in the motion video stream are extracted; Data processing operations are performed on the spatiotemporal trajectories of the key points to output motion quantification indicators.
3. The method according to claim 1, characterized in that, The step of determining the prediction deviation vector between the motion quantization index and the external sensor data in response to a preset calibration trigger condition includes: Perform a time-domain alignment operation on the motion quantization index and the external sensor data to obtain the first motion quantization index vector corresponding to the motion quantization index and the first external data vector corresponding to the external sensor data, respectively. Perform a spatial alignment operation on the first motion quantization index vector and the first external data vector to obtain the second motion quantization index vector corresponding to the first motion quantization index and the second external data vector corresponding to the first external data, respectively. The prediction bias vector is determined based on the second motion quantization index vector and the second external data vector.
4. The method according to claim 1, characterized in that, The step of using an incremental learning algorithm to compensate the parameters of the attitude estimation model based on the predicted deviation vector to generate a personalized model adapted to the target object includes: Extract multiple first deviation vectors containing a preset action cycle from the predicted deviation vector; Statistical smoothing is performed on the multiple first deviation vectors respectively to obtain multiple first smoothed vectors; Using multiple first smooth vectors as residual supervision signals, the pose estimation model is subjected to batch-dimensional parameter compensation via the incremental learning algorithm to generate a personalized model.
5. The method according to claim 1, characterized in that, The preset calibration triggering conditions include: detecting an explicit synchronization command generated by the camera terminal's interactive interface.
6. The method according to claim 1, characterized in that, The preset calibration triggering condition also includes: the prediction deviation vector between the motion quantization index and the external sensor data exceeds a preset threshold for N consecutive frames.
7. The method according to claim 1, characterized in that, The parameters include: spatial scale scaling factor, dynamic conversion coefficient and viewpoint weight bias, wherein the parameters are nonlinearly coupled.
8. A personalized model generation device based on multi-source feature collaboration, characterized in that, The device includes: The first data acquisition module is used to acquire motion video streams of the target object based on the camera terminal; The motion quantization index output module is used to output the motion quantization index corresponding to the motion video stream based on the pose estimation model. The second data acquisition module is used to acquire external sensor data synchronized with the motion video stream; The deviation vector determination module is used to determine the predicted deviation vector between the motion quantization index and the external sensor data in response to a preset calibration trigger condition. The personalized model generation module is used to perform parameter compensation on the attitude estimation model based on the prediction deviation vector using an incremental learning algorithm, and generate a personalized model adapted to the target object.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method of any one of claims 1 to 7.
10. An electronic device, characterized in that, The electronic device includes: A memory that stores a computer program; The processor, which is communicatively connected to the memory, executes the method of any one of claims 1 to 7 when the computer program is invoked.