Multi-source data fused user motion portrait construction method and system

By using multi-source data fusion and graph neural network technology, a motion semantic graph is constructed and a dynamic individual state vector is generated, which solves the problem of data loss and jumps in user motion profiles in complex scenarios, and realizes high-precision personalized modeling and expression of user ability status.

CN120974425AInactive Publication Date: 2025-11-18BEIJING TIMES DIGITAL TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202511128550.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-13
Publication Date
2025-11-18
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing user motion profiling technologies struggle to dynamically adapt to changes in user status due to training, fatigue, and health conditions in complex motion scenarios, resulting in data gaps and jumps, failing to accurately capture key motion behaviors, and lacking personalized modeling capabilities.

Method used

By using a multi-source data fusion method, multimodal motion data is collected and time-series aligned to construct a motion semantic graph. A graph neural network is then used for semantic-level completion reasoning, and a motion profile is generated by combining dynamic individual state vectors, thereby achieving accurate identification and completion of action segments.

Benefits of technology

It improves the robustness and personalized modeling capabilities of motion profiles, accurately captures user capability status in complex motion scenarios, enhances data continuity and the completeness of behavioral expression, and has good computational scalability and personalized expression accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120974425A_ABST
    Figure CN120974425A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-source data fused user motion portrait construction method and system, and relates to the technical field of user motion portrait construction, and the method comprises the following steps: collecting multi-modal motion data of a user in a complex motion scene, and carrying out the time sequence alignment; segmenting the multi-modal motion data into structured motion segments, and constructing a motion semantic graph containing motion nodes and motion transfer edges; on the basis of the motion semantic graph, semantic-level completion reasoning is carried out on the action segments with jumping and missing through a graph neural network; a dynamic individual state vector is constructed for a target user, feature fusion is carried out on the repaired structured action fragment sequence and the dynamic individual state vector, and a motion portrait reflecting the current ability state of the user is generated; according to the method, the problem of lack of dynamic adaptability of user capability state portrait construction caused by data asynchronization and action fragment missing and jumping in a complex motion scene is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of user motion profile construction technology, and more specifically, to a method and system for constructing user motion profiles by fusing multi-source data. Background Technology

[0002] The technology for building user activity profiles is a comprehensive approach that integrates multi-source data fusion, behavioral modeling, and personalized analysis. This technology is widely used in smart wearable devices, sports and health management platforms, fitness applications, and smart fitness equipment. Its aim is to create high-precision models of users' activity behaviors, abilities, health conditions, and exercise preferences, thereby supporting functions such as personalized recommendations, sports and health assessments, and risk warnings.

[0003] Existing user motion profiling technologies mainly rely on rule-based or static feature extraction-based modeling methods, which have significant limitations when dealing with complex motion scenarios. On the one hand, for explosive, high-intensity, or non-linearly changing motion forms, such as sprinting, jumping, or high-frequency action switching, sensor data often exhibits jumps, saturation, or missing information. Existing interpolation, filtering, and traditional time-series prediction methods struggle to achieve semantic-level data restoration, resulting in the inability to accurately capture and reconstruct key motion behaviors.

[0004] On the other hand, most current profiling systems use fixed structure models or group average parameters, lacking the ability to model the evolution of individual user status over time. They cannot dynamically adapt to fluctuations in user capabilities caused by training, fatigue, and changes in health status, thus causing the profile to become fixed in the long term and making it difficult to meet core application needs such as personalized training recommendations and capability trend prediction.

[0005] To address the above problems, this invention proposes a solution. Summary of the Invention

[0006] To overcome the aforementioned deficiencies of the prior art, embodiments of the present invention provide a method and system for constructing user motion profiles through multi-source data fusion, which solves the problem of lack of dynamic adaptability in the construction of user capability status profiles due to data asynchrony, missing action segments, and jumps in complex motion scenarios.

[0007] To achieve the above objectives, the present invention provides the following technical solution: Firstly, this application provides a method for constructing a user motion profile by fusing multi-source data. The method includes: collecting multimodal motion data of users in complex motion scenarios and performing temporal alignment; segmenting the multimodal motion data into structured action segments and constructing a motion semantic graph containing action nodes and action transition edges; based on the motion semantic graph, performing semantic-level completion reasoning on action segments with jumps and missing data using a graph neural network; constructing a dynamic individual state vector for the target user; and fusing the repaired structured action segment sequence with the dynamic individual state vector to generate a motion profile reflecting the user's current ability state.

[0008] In one embodiment, multimodal motion data of users in complex motion scenarios is collected and time-series aligned. Specifically, the following steps are taken: raw motion data is collected and multiple modal time series sets are formed, and the standard deviation of data segments in each modal time series set is calculated; if the standard deviation exceeds a preset threshold, it is determined that the modality has undergone significant changes within the corresponding time window; the center moment of time segments in which at least two modalities have undergone significant changes within the same time window is defined as a co-occurrence event, resulting in a timestamp set; the timestamp difference sequence of corresponding co-occurrence events between each non-reference modality and the reference modality in the set is obtained, and a mapping relationship between the time difference and time is established using linear fitting to form a time correction rule; based on the time correction rule, the timestamps of all data points in the non-reference modality are mapped to the reference modality time coordinate system; a unified time axis is divided by a preset fixed time interval, and the data points of each modality on the unified time axis are completed using linear interpolation to generate time-consistent multimodal motion data.

[0009] In one embodiment, the multimodal motion data is segmented into structured action segments, specifically: a semi-hidden Markov model is constructed using the multimodal motion data, and Viterbi path reasoning is performed to extract candidate action boundary points; based on the candidate action boundary points, the multimodal motion data is divided into multiple non-overlapping initial action segments; the initial action segments are input into a preset self-supervised feature learning network to extract low-dimensional embedding vectors; all embedding vectors are obtained, and the cosine similarity between all action segments is calculated to obtain their feature similarity, and a feature similarity matrix is ​​constructed; hierarchical clustering is performed on the matrix to obtain multiple feature aggregation groups; based on the change position of the clustering labels of adjacent action segments in the aggregation group, boundary points are identified; and continuous action segments between boundary points are combined into structured action segments.

[0010] In one embodiment, a motion semantic graph containing action nodes and action transition edges is constructed, specifically by: extracting low-dimensional embedding vectors corresponding to each structured action segment and defining graph nodes to form a set of action nodes; establishing directed edges between adjacent segments according to the temporal order of the structured action segments; establishing auxiliary edges between segments with semantic feature similarity exceeding a preset similarity threshold according to the semantic feature similarity between action segments to form a complete edge set; assigning weights to each edge in the edge set, and combining the action nodes and the weighted edges between them to construct a weighted directed graph; and embedding paths with a frequency higher than a preset frequency threshold in the sequence of action nodes as subgraph structures into the weighted directed graph to obtain a global motion semantic graph.

[0011] In one embodiment, based on the motion semantic graph, semantic-level completion reasoning is performed on action segments with jumps and missing data using a graph neural network. Specifically, the motion semantic graph is transformed into an adjacency matrix; multi-dimensional features are extracted from each node in the motion semantic graph, including but not limited to segment duration, modal amplitude mean, temporal position similarity to neighboring nodes, and data integrity; the multi-dimensional features are combined according to the node sequence to form a node input feature matrix; residual analysis is performed based on the node multi-dimensional features and the statistical characteristics of neighboring nodes to identify jump anomalous nodes and missing anomalous nodes; anomalous nodes are merged into an anomalous target node set; the node input feature matrix and adjacency matrix are input into the graph neural network, and a completion vector is generated using a neighbor feature aggregation mechanism with edge weights; the action segment content is reconstructed based on the completion vector.

[0012] In one embodiment, the action segment content is reconstructed based on the completion vector, specifically by: mapping the completion vector to segment attributes of the structured action segment, including the segment duration range, modal amplitude range, and start and end time index range; extracting the missing segment corresponding to the abnormal target node from the multimodal data based on the time index range; using the segment attributes as target priors to generate context-consistent fitting data from the current multimodal data to construct the reconstructed action segment; embedding the reconstructed action segment into the missing segment and smoothing the transition boundary between the reconstructed segment and adjacent segments to obtain the repaired structured action segment sequence.

[0013] In one embodiment, residual analysis is performed based on the multidimensional features of a node and the statistical characteristics of its neighboring nodes to identify abrupt change nodes and missing nodes. Specifically, multidimensional features are extracted from the direct neighboring nodes of each node; for the target node, its multidimensional features are compared with the multidimensional features of its direct neighboring nodes to calculate the residual index between the node and the neighborhood features; the residual index is compared with the corresponding preset residual index threshold, and abrupt change nodes and missing nodes are marked according to the comparison results.

[0014] In one embodiment, a dynamic individual state vector is constructed for the target user. Specifically, this involves: collecting multimodal motion-related data of the target user and preprocessing the multimodal motion-related data; weighting and summing the preprocessed multimodal motion-related data to obtain a unified fusion feature vector for each modality; and inputting the unified fusion feature vector into a state encoding network constructed based on a meta-learning dynamic adaptive training mechanism to generate an individual state vector.

[0015] In one embodiment, the repaired structured action segment sequence is fused with the dynamic individual state vector to generate a motion profile reflecting the user's current ability state, specifically: A temporal interpolation algorithm is used to temporally extend the individual state vector so that its time step count matches the length of the action segment sequence. A timestamp synchronization mechanism is then used to align the extended individual state vector with the action segment sequence time-by-time, generating a joint representation tensor. A motion profiling modeling network is constructed, using the joint representation tensor as input. This network includes an input adaptation layer, a temporal feature extraction module, an attention feature selection module, a global feature fusion module, and a profiling output layer. The input adaptation layer receives the joint representation tensor and performs dimensionality standardization, regularization, and positional encoding. The temporal feature extraction module includes multiple one-dimensional convolutional layers responsible for extracting local temporal change features in the action sequence; the attention feature selection module is based on a multi-head attention mechanism, which weights and aggregates local temporal change features related to the user's ability state in the action feature sequence according to the individual state vector; the global feature fusion module includes a gated linear unit structure to selectively activate the weighted local temporal change features and form a unified multi-dimensional ability representation vector; the profile output layer includes two fully connected neural networks, which generate a motion profile vector representing the user's current ability state by linearly mapping the multi-dimensional ability representation vector.

[0016] Secondly, this application provides a method and system for constructing user motion profiles through multi-source data fusion, the system comprising: The data acquisition and alignment module is used to collect multimodal motion data of users in complex motion scenarios and perform temporal alignment. The semantic graph construction module is used to segment multimodal motion data into structured action segments and construct a motion semantic graph containing action nodes and action transition edges. The completion module is used to perform semantic-level completion reasoning on action segments with jumps and missing parts based on the motion semantic graph and through a graph neural network. The motion profile generation module is used to construct a dynamic individual state vector for the target user, and to perform feature fusion between the repaired structured action segment sequence and the dynamic individual state vector to generate a motion profile that reflects the user's current ability state.

[0017] As can be seen from the above technical solutions, the embodiments of this application have the following advantages: 1. By constructing a complete multimodal motion data processing and semantic modeling workflow, a systematic expression from temporal alignment and action structure partitioning to semantic graph construction is achieved, possessing the following significant advantages: First, through a time correction mechanism based on co-occurring events, the time synchronization problem caused by sampling frequency differences and clock drift in complex motion scenarios of multi-source sensors is effectively solved, ensuring the consistency and accuracy of data fusion; Second, by introducing a semi-hidden Markov model combined with Gaussian emission and Poisson persistent modeling, high-quality boundary recognition and temporal structure characterization of motion segments are achieved; Third, by integrating self-supervised feature learning and semantic similarity measurement, action segments are mapped to a unified embedding space, enabling action semantics to have low-dimensional expressive capabilities; Finally, by constructing a weighted directed graph containing temporal transitions and semantic auxiliary edges, the logical structured expression of individual behaviors and the extraction of high-frequency behavior patterns are achieved, possessing good computational scalability and strong adaptability to multimodal behavior recognition tasks, significantly improving the robustness, expressiveness, and downstream modeling efficiency of the overall system.

[0018] 2. A multimodal motion semantic modeling method combining graph neural networks and meta-learning achieves accurate identification and semantic-level completion of abnormal action segments, and further integrates individual states to construct highly personalized motion profiles. This method has the following significant advantages: First, it utilizes graph neural networks to perform context-aware completion of transitions and missing nodes in the motion semantic graph, effectively solving data breaks and noise problems caused by sensor anomalies and drastic motion changes, significantly improving data continuity and the completeness of behavioral expression. Second, it introduces a multidimensional residual analysis mechanism, combining indicators such as duration, modal amplitude, similarity, and data completeness to achieve highly sensitive identification of structurally abnormal segments, enhancing the robustness of graph structures. Third, it constructs a dynamic individual state encoding network based on a meta-learning mechanism, possessing the ability to quickly adapt to different user motion characteristics, improving the generalization of state modeling and the accuracy of personalized expression. Fourth, through temporal interpolation alignment and multi-module deep network fusion mechanisms, it constructs a joint representation tensor that coordinates temporal and semantic aspects, effectively capturing local dynamic and static capability trends, achieving capability state modeling with strong semantic interpretability, high timeliness, and sensitivity to individual differences. The overall solution possesses strong robustness, high completion accuracy, and good scalability, providing solid technical support for multimodal motion behavior understanding, personalized intervention decision-making, and intelligent motion assessment. Attached Figure Description

[0019] Figure 1This is a schematic diagram of the user motion profile construction method based on multi-source data fusion provided in the embodiments of this application.

[0020] Figure 2 This is a schematic diagram of the user motion profile construction system based on multi-source data fusion provided in this application embodiment. Detailed Implementation

[0021] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present invention.

[0022] Reference Figure 1 As shown in the diagram, the user motion profile construction method based on multi-source data fusion provided by this invention includes the following steps: S1 collects multimodal motion data of users in complex motion scenarios and performs time-series alignment.

[0023] In this embodiment, the acquisition phase aims to obtain multimodal motion data with temporal sequence, behavioral relevance, and spatial continuity through wearable devices and sensors, providing foundational data support for subsequent motion semantic modeling and individual ability state recognition. The multimodal motion data describes the spatiotemporal evolution characteristics of the user's continuous movements. This multimodal motion data includes acceleration data, gyroscope data, heart rate data, and GPS trajectory data. In S1, multimodal motion data of users in complex motion scenarios is collected and temporally aligned, specifically as follows: Raw motion data is collected from multiple sensors, including accelerometers, gyroscopes, GPS modules, and heart rate monitors. The sampling timestamps and corresponding sensor values ​​of each sensor are recorded, and multiple modal time series sets are formed with each sensor as a mode. Using a sliding time window of preset length, calculate the standard deviation of data segments in each modal time series set; If the standard deviation exceeds a preset threshold, it is determined that the mode has changed significantly within the corresponding time window; Obtain time segments in which at least two modalities change significantly within the same time window, define the center moment of the time segment as a co-occurrence event, and obtain the set of timestamps of the co-occurrence events corresponding to each modality; Obtain the sequence of timestamp differences between the co-occurrence events of each non-reference mode and the reference mode in the set of timestamps of co-occurrence events, and establish the mapping relationship between the time difference and time by linear fitting to form a time correction rule. The fitting uses the least squares method to fit the trend of the difference changing with time. Based on the time correction rule, the timestamps of all data points in the non-reference mode are mapped to the reference mode time coordinate system; A unified time axis is divided by a preset fixed time interval, and a linear interpolation method is used to complete the data points of each mode on the unified time axis, generating multimodal motion data with consistent time.

[0024] In this context, the reference mode refers to the mode selected as the time base in a multimodal time series. It is typically the sensor mode with the most accurate timestamps, the most stable clock, or the most reliable data. The data time axis of this mode is regarded as the standard time base, and the data time of other modes needs to be mapped onto this base.

[0025] Non-reference modes are sensor modes other than the reference mode. Since the timestamps collected by each sensor may have time drift, asynchrony, or errors, the timestamps of these non-reference modes need to be mapped to the time coordinate system of the reference mode through the constructed time correction rules, so as to achieve time alignment of multimodal data.

[0026] Time correction rules are a set of functions or mapping rules that describe how non-reference modal data is mapped to the reference modal time base based on co-occurrence event timestamps. Essentially, they fit the systematic or dynamic offset between the non-reference modal time system and the reference modal time system, ensuring the uniformity and alignment of multimodal data in the time dimension. A co-occurrence event refers to a situation where, within the same time window, the changes of at least two or more modal sensors exceed their respective thresholds; the center moment of this time window is then taken as the time point of the co-occurrence event.

[0027] It should be noted that time-series alignment of multimodal sensor data aims to address the time asynchrony issues caused by factors such as sampling frequency, start-up time, or clock drift among different modalities. In complex motion scenarios, the responses of different sensors to the same physical event exhibit time discrepancies; failure to correct these discrepancies can affect the accuracy of data fusion and subsequent analysis. By detecting significant changes through a sliding window and identifying co-occurring events as alignment anchors, and employing a least-squares method to construct time correction rules for time mapping, accurate alignment of non-reference modalities to the reference modal time axis can be achieved. Finally, through unified time axis interpolation, all modalities have corresponding observations at the same time point, thereby improving the temporal consistency and robustness of multimodal fusion, behavior recognition, and model construction.

[0028] S2 divides the multimodal motion data into structured action segments and constructs a motion semantic graph containing action nodes and action transition edges.

[0029] In this embodiment, in S2, the multimodal motion data is segmented into structured action segments, specifically as follows: A semi-hidden Markov model is constructed using multimodal motion data, wherein: K-means was used to perform initial clustering of the multimodal motion data to obtain the number of hidden states; It should be noted that the K-means clustering algorithm is used. The number of clusters K during clustering is determined using the elbow method, that is, within a range of candidate K values ​​(e.g., from 5 to 20), the sum of squared clustering errors corresponding to each K is calculated, and the K value corresponding to the point where the error decreasing trend slows down is selected as the number of cluster centers, and the number of hidden states is determined accordingly.

[0030] The multimodal motion data at each moment is used as an observation sequence; For each hidden state, a Gaussian distribution model is used to represent the feature emission probability; It should be noted that using a Gaussian distribution model to represent the feature emission probability for each hidden state can be understood as follows: after clustering, each cluster result is considered as an initial hidden state. The mean vector and covariance matrix of all feature vectors belonging to that state are calculated to construct the emission probability model for each state. This emission model is expressed in Gaussian form, that is, the shape and center position of the distribution of a state in the feature space, describing the probability distribution characteristics of the observations in that state.

[0031] A duration parameter is introduced for each state, and the duration of state stay is represented by fitting a Poisson distribution; It should be noted that using a Poisson distribution to represent state dwell time can be understood as follows: for each initial hidden state, statistically analyze the distribution of the duration of continuous maintenance of that state in the training samples, recording the number of frames maintained (i.e., state duration). Frequency statistics are then performed on these duration data, and a discrete distribution model is fitted. The Poisson distribution is used to model the state dwell time, where the parameters are obtained through maximum likelihood estimation, i.e., selecting distribution parameter values ​​that maximize the probability of occurrence of the observed duration data.

[0032] A semi-hidden Markov model is constructed by combining the aforementioned hidden states, Gaussian emission model, and Poisson duration distribution.

[0033] Perform Viterbi path reasoning on the model to extract candidate action boundary points; Based on candidate action boundary points, the multimodal motion data is divided into multiple non-overlapping initial action segments. Each action segment contains a complete motion action unit and maintains the original temporal order. The initial action segments are input into a pre-defined self-supervised feature learning network, and the semantic feature representation of each action segment in multimodal mode is extracted to obtain a low-dimensional embedding vector of uniform dimension, which is used to measure the semantic similarity between different action segments. It should be noted that inputting the initial action fragments into a pre-defined self-supervised feature learning network to extract the semantic feature representations of each action fragment in multimodal modes can be understood as feeding the initial action fragments as input into a feature extraction network built on a self-supervised learning mechanism. This network, through the design of a multimodal fusion structure and unsupervised proxy tasks, automatically extracts low-dimensional embedding vectors for each action fragment. These vectors semantically represent the modal commonalities and temporal structural features of the actions, possessing the ability to perform semantic comparisons and similarity discrimination between different fragments.

[0034] Obtain all embedded vectors, calculate the cosine similarity between all action segments to obtain their feature similarity, and construct a feature similarity matrix; Hierarchical clustering of the matrix yields multiple feature aggregation groups; Based on the location of cluster label changes of adjacent action segments in the aggregated group, potential action boundary points are identified. Combine consecutive action segments between boundary points into structured action segments.

[0035] It should be noted that performing Viterbi path reasoning on the model to extract candidate action boundary points can be understood as using the Viterbi decoding algorithm to perform optimal path reasoning on the entire observation data, searching for the state sequence that maximizes the probability of the observed sequence occurring. This state sequence represents the most likely state of the system at each time point. By comparing the changes in state numbers at adjacent time points, the time points where state transitions occur are marked; these transition points are the candidate action boundaries. The original data is divided into several temporally continuous initial action segments using these boundary points, with each segment corresponding to a possible action unit.

[0036] In S2, a motion semantic graph containing action nodes and action transition edges is constructed, specifically as follows: Extract the low-dimensional embedding vector corresponding to each structured action segment, and define each structured action segment as a graph node to form a set of action nodes. The action nodes use their corresponding embedding vectors as node features. Based on the temporal sequence of structured action segments, directed edges are established between adjacent segments to represent the transition relationships between action segments; Based on the semantic feature similarity between action segments, auxiliary edges are established between segments whose semantic feature similarity exceeds a preset similarity threshold to form a complete edge set; A weight is assigned to each edge in the edge set, and the weight is calculated based on the transition frequency or semantic similarity between segments; Combine the action nodes and the weighted edges between them to construct a weighted directed graph; In the sequence of action nodes, continuous transition paths are extracted by sliding. Paths with a frequency higher than a preset frequency threshold are statistically analyzed and embedded as subgraph structures into the weighted directed graph to obtain a global motion semantic graph.

[0037] By mapping structured action segments to nodes in a graph and constructing transition edges and auxiliary edges based on temporal order and semantic similarity, a weighted directed graph structure with realistic action logic and semantic connections is formed. Node embedding vectors serve as features, enabling compressed representation of action semantics; transition edges capture the temporal dynamics of actions, while auxiliary edges enhance semantic connectivity across time segments; high-frequency path subgraph embeddings extracted based on sliding windows strengthen the structural representation of typical behavioral patterns. Further graph optimization and structural standardization processes give the graph structure good computability and scalability, making it suitable for various downstream graph learning tasks. The overall scheme integrates the advantages of temporal modeling, semantic representation, and structural abstraction, possessing significant advantages such as clear structure, semantic discriminability, and computational friendliness, providing a highly robust and expressive unified representation framework for multimodal behavior modeling.

[0038] S3. Based on the motion semantic graph, semantic-level completion reasoning is performed on action segments with jumps and missing parts using a graph neural network.

[0039] In this embodiment, in S3, based on the motion semantic graph, semantic-level completion reasoning is performed on action segments with jumps and missing parts using a graph neural network, specifically as follows: The motion semantic graph structure is optimized and transformed into an adjacency matrix. The optimization process includes edge trimming, node merging, and label annotation. Multidimensional features are extracted from each node in the motion semantic graph. These multidimensional features include, but are not limited to, segment duration, mean modal amplitude, temporal location and similarity to adjacent nodes, and data integrity. The segment duration refers to the length of time the current structured action segment occupies in the time series, and is obtained by the difference between the end timestamp and the start timestamp of the action segment. It is used to identify transition nodes; for example, an abnormally short duration may indicate a mis-segmented segment or a noisy segment.

[0040] The modal amplitude mean represents the average activity intensity of the current segment across various modalities (such as acceleration, gyroscope, heart rate, EMG, etc.). For each modal data corresponding to the current structured motion segment, the signal amplitude of each mode is extracted at all time points within that segment. Then, within the corresponding time range of the segment, all amplitude values ​​are averaged to obtain the average intensity value for each mode. This is used to determine whether there is modal failure, equipment loosening, or abnormally high-intensity motion segments. Sudden changes in the modal mean may also indicate data errors or abrupt motion transitions.

[0041] Temporal location and neighboring node similarity measures the similarity between a node and its neighboring segments in the feature space, reflecting its continuity with the context. Euclidean distance or cosine similarity is used to calculate the similarity between preceding and following nodes. If the similarity between the current node and its preceding and following segments is significantly lower than the average, it may be a "jump point" or data seam, and should be marked as a potential anomaly.

[0042] Data integrity refers to the proportion of missing raw sensor data in the current segment. Within the time period corresponding to the current structured motion segment, the acquired data for each modality is checked, and the number of time points with invalid data (e.g., missing values, zero values, or data marked as no response) is counted. This number is then compared to the total number of time points in the segment to obtain the data missing proportion over a period of time. A higher missing proportion indicates poorer raw data integrity and lower reliability for that segment. This is used to identify whether a structurally missing motion segment is caused by signal obstruction, sensor failure, etc. Multidimensional features are combined according to node sequences to form a node input feature matrix; Residual analysis is performed based on the multidimensional features of nodes and the statistical characteristics of neighboring nodes to identify abrupt change nodes and missing nodes. Abnormal nodes are merged into a set of abnormal target nodes to be completed, while maintaining their position in the graph and their relationship with their context neighbors. The node input feature matrix and adjacency matrix are input into the graph neural network. By using the neighbor feature aggregation mechanism with edge weights, the contextual semantic modeling of the abnormal target node set is realized, and the completion vector in the semantic completion process of each abnormal node is generated. This vector is used to recover or predict the feature attributes that the node should have, so as to realize the semantic-level completion and repair of abnormal action segments.

[0043] Reconstruct the content of the action segment based on the completion vector.

[0044] It's important to note that the core mechanism of graph neural networks is that a node's representation vector can obtain information from the representations of its neighboring nodes (i.e., neighbor aggregation features). The edge weighting refers to the fact that different neighboring nodes have different influences on the current node. During aggregation, the features of neighbors are weighted and averaged according to the edge weights in the adjacency matrix. This results in more accurate aggregated information that reflects the true contextual dependency structure. The completion vector generated in the semantic completion process for each anomalous node can be understood as follows: after the aforementioned aggregation and modeling process, the graph neural network outputs a completion vector for each target anomalous node. Essentially, this process uses the graph neural network to perform "context-aware" semantic modeling of anomalous segments using information from other normal segments, thereby generating a set of vectors representing the features that the anomalous segment should have. Based on this, segment-level completion and repair are performed to ensure the continuity and semantic integrity of the user's motion profile.

[0045] Furthermore, the action segment content is reconstructed based on the completed vectors, specifically as follows: The completion vector is mapped to the fragment attributes of the structured action fragment, including the fragment duration range, modal amplitude range, and start and end time index range; Based on the time index range, the missing segments corresponding to the abnormal target nodes are extracted from the original multimodal data; Using segment attributes as target priors, we generate fitting data consistent with the context from the current multimodal data, and construct reconstructed action segments with reasonable continuity in time sequence, modal amplitude and motion trend. The reconstructed motion segments are embedded into the missing segments, and the transition boundaries between the reconstructed segments and adjacent segments are smoothed to ensure the continuity of motion content and temporal consistency, resulting in a repaired structured motion segment sequence.

[0046] It should be noted that the segment duration range refers to the time interval occupied by the structured action segment on the time axis. This range is used to define the temporal boundaries of the segment generated during the completion process. The modal amplitude range refers to the range of activity intensity that each modality (such as acceleration, gyroscope, heart rate, etc.) signal should be in during the completed segment, usually expressed as the upper and lower limits of the mean ± standard deviation. The start and end time index range refers to the start and end frame numbers in the original sensor data frames. It is used to locate the position of the missing segment in the original data stream and "embed" the reconstructed segment at that position to achieve temporal reconstruction.

[0047] Furthermore, residual analysis is performed based on the multidimensional features of nodes and the statistical characteristics of neighboring nodes to identify nodes with abrupt changes and nodes with missing nodes. Specifically: Extract multidimensional features from the direct neighbor nodes (the previous and next nodes) of each node; For a target node, its multidimensional features are compared with the multidimensional features of its direct neighbor nodes to calculate the residual index between the node and the neighborhood features. The residual index includes fragment duration residual, modal amplitude mean residual, temporal location similarity residual with neighboring nodes residual, and data integrity residual. Calculating the residual index between the node and the neighborhood features means calculating the difference between the two. The residual indicators are compared with the corresponding preset residual indicator thresholds. If the residual of segment duration, the residual of modal amplitude mean, and the residual of time position and similarity with adjacent nodes exceed the preset jump residual indicator threshold, they are marked as jump abnormal nodes. If the data integrity residual exceeds the preset missing residual indicator threshold, it is marked as a missing abnormal node.

[0048] S4. Construct a dynamic individual state vector for the target user, and fuse the repaired structured action segment sequence with the dynamic individual state vector to generate a motion profile reflecting the user's current ability state.

[0049] In this embodiment, in S4, a dynamic individual state vector is constructed for the target user, specifically as follows: Collect multimodal motion-related data of the target user, including historical motion performance, physiological monitoring indicators, behavioral preference information and environmental context parameters; The multimodal motion-related data is preprocessed, including time alignment, missing value imputation, and normalization. The preprocessed multimodal motion-related data are weighted and summed to obtain a unified fusion feature vector for each modality. The unified fusion feature vector is input into a state encoding network constructed based on a meta-learning dynamic adaptive training mechanism to generate individual state vectors. The state encoding network adopts a dual-branch structure, one branch is a global shared layer for extracting general motion features, and the other branch is a personalized adaptive layer for conditional modeling based on user labels. The two are trained collaboratively under the meta-learning strategy.

[0050] It should be noted that, to enhance the individual adaptability and generalization ability of this fused feature in individual modeling, it is further input into a state encoding network. This state encoding network is constructed based on a meta-learning mechanism, specifically employing the MAML strategy for parameter initialization and rapid adaptive updates: First, based on a historical set of individual tasks, the encoding network is trained to quickly converge to the parameter space representing the optimal state of the individual when receiving different fused feature inputs. Second, for the target user, the fused feature vector generated at the current moment is used to form a support set and a query set. Personalized parameters are obtained through a fine-tuning process, outputting an embedding vector that dynamically reflects the user's ability state. This embedding vector semantically inherits the global representation capability of the fused feature and enhances its rapid adaptability to individual differences through the meta-learning mechanism. This embedding vector serves as a key input for subsequent motion profiling modeling and personalized intervention strategy generation, exhibiting good timeliness, discriminability, and model transferability.

[0051] Furthermore, the repaired structured action segment sequence is fused with the dynamic individual state vector to generate a motion profile reflecting the user's current ability state, specifically: A time interpolation algorithm is used to extend the individual state vector in time so that its time step count is consistent with the length of the action segment sequence. Using a timestamp synchronization mechanism, the expanded individual state vector is aligned with the action segment sequence time by time to generate a joint representation tensor. The joint representation tensor co-encodes user state and action performance in both temporal and semantic dimensions. A motion profile modeling network is constructed, and a joint representation tensor is used as input. The motion profile modeling network includes an input adaptation layer, a temporal feature extraction module, an attention feature selection module, a global feature fusion module, and a profile output layer. The input adaptation layer is used to receive the joint representation tensor and perform dimension normalization, regularization and position encoding processing; The dimensionality standardization unit adopts a batch normalization method to achieve uniform feature scale; The regularized normalized unit uses the L2 regularization method to limit the parameter range and prevent overfitting; The position coding unit generates a temporal position code based on a sine-cosine function. After coding, it is added to the joint representation tensor to retain the temporal position information. The temporal feature extraction module includes multiple one-dimensional convolutional layers, each with a kernel size of 3, a stride of 1, and a total of 3 layers. It uses the ReLU activation function and is responsible for extracting local temporal change features in the action sequence. The attention feature selection module is based on a multi-head attention mechanism, which maps dynamic individual state vectors to queries (Q); maps local temporal change features to keys (K) and values ​​(V); and uses multi-head attention calculation with 8 heads to calculate a weighted score matrix to achieve weighted aggregation of local temporal change features related to the user's ability state in the action sequence. The global feature fusion module includes: The gated activation function uses a gated linear unit (GLU) structure to selectively activate the weighted local temporal variation features; Simultaneously, the static capability trend characteristics in the individual state vector are integrated to form a unified multidimensional capability representation vector; The image output layer consists of two fully connected neural networks. The first layer has 256 nodes, the second layer has 128 nodes, and the activation function is ReLU. Finally, the multidimensional ability representation vector is linearly mapped to generate a motion image vector representing the user's current ability state.

[0052] It should be noted that by introducing temporal interpolation and temporal alignment mechanisms, a precise fusion of dynamic individual state vectors and structured action segment sequences is achieved, constructing a joint representation tensor of temporal and semantic co-coding. Then, a multi-module deep modeling network is used to achieve multi-dimensional joint modeling of local action dynamics and individual ability states, ultimately generating a highly individualized, timely, and semantically interpretable motion profile vector that accurately reflects the user's current ability state, providing high-precision input for downstream tasks such as sports rehabilitation, personalized training, or state assessment. Local temporal change features refer to the subtle changes in action features over time within a short time window in a structured action segment sequence, mainly reflecting micro-level temporal dynamic information such as action continuity, stability, rhythm, and acceleration changes. The multi-dimensional ability representation vector refers to a high-dimensional, structured, and integrated representation that comprehensively reflects the user's motor ability state within the current time window, including multiple semantic dimensions such as dynamic feature performance (extracted from the action sequence) and static ability trends (extracted from individual state embedding).

[0053] Reference Figure 2 As shown in the diagram, the user motion profile construction system based on multi-source data fusion provided by this invention includes a data acquisition and alignment module, a semantic graph construction module, a completion module, and a motion profile generation module. These modules are interconnected. The data acquisition and alignment module is used to collect multimodal motion data of users in complex motion scenarios and perform temporal alignment. The semantic graph construction module is used to segment multimodal motion data into structured action segments and construct a motion semantic graph containing action nodes and action transition edges. The completion module is used to perform semantic-level completion reasoning on action segments with jumps and missing parts based on the motion semantic graph and through a graph neural network. The motion profile generation module is used to construct a dynamic individual state vector for the target user, and to perform feature fusion between the repaired structured action segment sequence and the dynamic individual state vector to generate a motion profile that reflects the user's current ability state.

[0054] The above formulas are all dimensionless calculations. The formulas are derived from software simulations based on a large amount of collected data to obtain the most recent real-world results. The preset parameters in the formulas are set by those skilled in the art according to the actual situation.

[0055] The above embodiments can be implemented, in whole or in part, by software, hardware, firmware, or any other combination thereof. When implemented using software, the above embodiments can be implemented, in whole or in part, in the form of a computer program product.

[0056] Those skilled in the art will recognize that the modules and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0057] In addition, the functional modules in the various embodiments of this application can be integrated into one processing module, or each module can exist physically separately, or two or more modules can be integrated into one module.

[0058] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

[0059] In conclusion, the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A method for constructing user motion profiles through multi-source data fusion, characterized in that, Includes the following steps: Collect multimodal motion data of users in complex motion scenarios and perform time-series alignment; Multimodal motion data is segmented into structured action segments, and a motion semantic graph containing action nodes and action transition edges is constructed. Based on the motion semantic graph, semantic-level completion reasoning is performed on action segments with jumps and missing parts using a graph neural network. A dynamic individual state vector is constructed for the target user. The repaired structured action segment sequence is then fused with the dynamic individual state vector to generate a motion profile that reflects the user's current ability state.

2. The user motion profile construction method based on multi-source data fusion according to claim 1, characterized in that, The process of collecting and temporally aligning multimodal motion data from users in complex motion scenarios specifically involves: Raw motion data is collected and multiple modal time series are generated, and the standard deviation is calculated. If the standard deviation exceeds a preset threshold, it is determined that the mode has changed significantly within the corresponding time window; The center moment of a time segment in which at least two modalities change significantly within the same time window is defined as a co-occurrence event, and a set of timestamps is obtained. Obtain the timestamp difference sequence of the corresponding co-occurrence events between each non-reference mode and the reference mode in the set, and use linear fitting to establish the mapping relationship between the time difference and time, forming a time correction rule; Based on the time correction rule, the timestamps of all data points in the non-reference mode are mapped to the reference mode time coordinate system; A unified time axis is divided by a preset fixed time interval, and the data points of each mode on the unified time axis are completed by linear interpolation to generate multimodal motion data with consistent time.

3. The user motion profile construction method based on multi-source data fusion according to claim 2, characterized in that, The process of segmenting multimodal motion data into structured action segments specifically involves: A semi-hidden Markov model is constructed using multimodal motion data, and Viterbi path reasoning is performed to extract candidate action boundary points. Based on candidate action boundary points, the multimodal motion data is divided into multiple non-overlapping initial action segments; The initial action fragment is input into a pre-defined self-supervised feature learning network to extract a low-dimensional embedding vector; Obtain all embedded vectors and calculate the cosine similarity between all action segments to obtain their feature similarity. And construct a feature similarity matrix; Hierarchical clustering of the matrix yields multiple feature aggregation groups; Boundary points are identified based on the changes in cluster labels of adjacent action segments in the aggregated group; Combine consecutive action segments between boundary points into structured action segments.

4. The user motion profile construction method based on multi-source data fusion according to claim 3, characterized in that, The construction of the motion semantic graph, which includes action nodes and action transition edges, is specifically as follows: Extract the low-dimensional embedding vector corresponding to each structured action segment, and define graph nodes to form a set of action nodes. The action nodes use their corresponding embedding vectors as node features. Based on the temporal order of structured action segments, directed edges are established between adjacent segments; Based on the semantic feature similarity between action segments, auxiliary edges are established between segments that exceed a preset similarity threshold to form a complete edge set; Assign a weight to each edge in the edge set, and combine the action nodes and the weighted edges between them to construct a weighted directed graph; In the sequence of action nodes, paths that appear more frequently than a preset frequency threshold are embedded as subgraph structures into the weighted directed graph to obtain a global motion semantic graph.

5. The user motion profile construction method based on multi-source data fusion according to claim 4, characterized in that, The step of performing semantic-level completion reasoning on action segments with jumps and missing parts based on the motion semantic graph using a graph neural network is as follows: Transform the motion semantic graph into an adjacency matrix; Multidimensional features are extracted from each node in the motion semantic graph. These multidimensional features include, but are not limited to, segment duration, mean modal amplitude, temporal location and similarity to adjacent nodes, and data integrity. Multidimensional features are combined according to node sequences to form a node input feature matrix; Residual analysis is performed based on the multidimensional features of nodes and the statistical characteristics of neighboring nodes to identify abrupt change anomalous nodes and missing anomalous nodes, and anomalous nodes are merged into a set of anomalous target nodes. The node input feature matrix and adjacency matrix are input into the graph neural network. The neighbor feature aggregation mechanism with edge weights is used to generate a completion vector, and the action segment content is reconstructed based on the completion vector.

6. The user motion profile construction method based on multi-source data fusion according to claim 5, characterized in that, The process of reconstructing the action segment content based on the completion vector specifically involves: The completion vector is mapped to the fragment attributes of the structured action fragment, including the fragment duration range, modal amplitude range, and start and end time index range; Based on the time index range, the missing segments corresponding to the abnormal target nodes are extracted from the multimodal data; Using fragment attributes as target priors, we generate context-consistent fitted data from the current multimodal data to construct reconstructed action fragments; The reconstructed action fragments are embedded into the missing fragments, and the transition boundaries between the reconstructed fragments and adjacent fragments are smoothed to obtain the repaired structured action fragment sequence.

7. The user motion profile construction method based on multi-source data fusion according to claim 6, characterized in that, The residual analysis based on the multidimensional features of nodes and the statistical characteristics of neighboring nodes is used to identify abrupt change anomalous nodes and missing anomalous nodes, specifically as follows: Extract multidimensional features from the direct neighbor nodes of each node; For the target node, its multidimensional features are compared with the multidimensional features of its direct neighbor nodes, and the residual index between the node and the neighborhood features is calculated. The residual indexes are compared with the corresponding preset residual index thresholds, and the jump abnormal nodes and missing abnormal nodes are marked according to the comparison results.

8. The user motion profile construction method based on multi-source data fusion according to claim 7, characterized in that, The construction of dynamic individual state vectors for target users specifically involves: Collect multimodal motion-related data of the target user and preprocess the multimodal motion-related data; The preprocessed multimodal motion-related data are weighted and summed to obtain a unified fusion feature vector for each modality. The unified fused feature vector is input into a state encoding network constructed based on a meta-learning dynamic adaptive training mechanism to generate individual state vectors.

9. The method for constructing a user motion profile by fusing multi-source data according to claim 1, characterized in that, The step of fusing the repaired structured action segment sequence with the dynamic individual state vector to generate a motion profile reflecting the user's current ability state specifically involves: A time interpolation algorithm is used to extend the individual state vector in time so that its time step count is consistent with the length of the action segment sequence. By using a timestamp synchronization mechanism, the expanded individual state vector is aligned with the action segment sequence time by time to generate a joint representation tensor; A motion profile modeling network is constructed, and a joint representation tensor is used as input. The motion profile modeling network includes an input adaptation layer, a temporal feature extraction module, an attention feature selection module, a global feature fusion module, and a profile output layer. The input adaptation layer is used to receive the joint representation tensor and perform dimension normalization, regularization and position encoding processing; The temporal feature extraction module includes multiple one-dimensional convolutional layers, which are responsible for extracting local temporal change features in the action sequence. The attention feature selection module is based on a multi-head attention mechanism, which performs weighted aggregation of local temporal change features related to the user's ability state in the action feature sequence according to the individual state vector; The global feature fusion module includes selective activation of weighted local temporal variation features through a gated linear unit structure, forming a unified multidimensional capability representation vector; The image output layer includes two fully connected neural networks, which generate motion image vectors representing the user's current ability status by linearly mapping multidimensional ability representation vectors.

10. A system for constructing a user motion profile using the multi-source data fusion method as described in any one of claims 1-9, characterized in that, It includes a data acquisition and alignment module, a semantic graph construction module, a completion module, and a motion profile generation module, and these modules are interconnected. The data acquisition and alignment module is used to collect multimodal motion data of users in complex motion scenarios and perform temporal alignment. The semantic graph construction module is used to segment multimodal motion data into structured action segments and construct a motion semantic graph containing action nodes and action transition edges. The completion module is used to perform semantic-level completion reasoning on action segments with jumps and missing parts based on the motion semantic graph and through a graph neural network. The motion profile generation module is used to construct a dynamic individual state vector for the target user, and to perform feature fusion between the repaired structured action segment sequence and the dynamic individual state vector to generate a motion profile that reflects the user's current ability state.

Citation Information

Cited By

  • Multi-modal vehicle-mounted AIGC generation method and device based on end-cloud hierarchical collaboration, medium and product

    CN121902074A