Intelligent identification method and system for target behavior data facing simulation training task
By employing a cloud-edge-device collaborative architecture and multimodal data fusion technology, the problems of subjectivity in training evaluation and lag in real-time feedback were solved, enabling multi-dimensional, real-time feedback and efficient evaluation of training behavior, thereby improving training effectiveness.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-01-12
- Publication Date
- 2026-04-10
AI Technical Summary
In competitive sports, emergency rescue, medical rehabilitation, and high-end skills training, current technologies rely on subjective observation for training assessment, lack objective quantitative indicators, and cannot achieve multi-dimensional data fusion and real-time feedback, resulting in inaccurate assessment results and low training efficiency.
Adopting a cloud-edge-device collaborative architecture, visual, auditory, and physiological data are collected synchronously through a heterogeneous sensor network to construct a multimodal heterogeneous graph. Feature extraction and fusion are performed using a modal-specific encoder and a multi-task learning network to achieve a three-dimensional global insight into action, acoustics, and physiology. Edge computing is used to reduce latency and provide real-time feedback on training results.
It enables multi-dimensional and correlated evaluation of training behavior, improves the objectivity and efficiency of evaluation, and can correct errors in a timely manner during training, thereby improving training effectiveness and the accuracy of intelligent guidance.
Smart Images

Figure CN121502236B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of behavior data identification of simulation models, in particular to a target behavior data intelligent identification method and system for simulation training tasks. BACKGROUND
[0002] In the fields of competitive sports, emergency rescue, medical rehabilitation, and high-end skill training (such as precise instrument operation), the scientific, precise, and intelligent training process is the core key to improving the effectiveness of individuals and teams, reducing training risks, and accelerating talent training. However, for a long time, the evaluation and guidance of these complex training processes have been heavily dependent on the subjective observation and immediate experience of coaches, instructors, or experts. This traditional "human eye observation + experience analysis" mode has exposed many fundamental limitations that are difficult to overcome under the demand of modern and high-standard training:
[0003] (1) Subjectivity and non-quantitative bottleneck
[0004] The evaluation results are deeply influenced by the personal experience, attention state, perspective limitations, and even subjective preferences of the evaluators. The lack of unified and objective quantitative indicators leads to different evaluations of the same training action by different evaluators, making it difficult to form a standardized and reproducible evaluation system. The measurement of training effectiveness is limited to vague levels such as "good / medium / poor", which cannot provide data support for the fine-tuning of training programs.
[0005] (2) Single information perception dimension, lack of global insight
[0006] Human observers usually focus more on the visual modality information that is easiest to capture, such as the form and speed of actions. However, training, especially high-load or high-precision training, is a complex system involving physiology, psychology, and environmental interaction. The traditional approach severely neglects two other crucial information dimensions:
[0007] 1) Auditory dimension
[0008] The breathing rhythm, intensity, and stability of the trainee are direct indicators of their physical consumption and physiological load; the clarity of the order and response time reflect their cognitive load and stress state; the sound of instrument operation (such as mechanical clunking) is related to the standardization and proficiency of operation.
[0009] 2) Physiological dimension
[0010] Heart rate and its variability (HRV) is a gold standard to measure autonomic nervous system activity, mental stress and recovery from fatigue; Electromyography (EMG) can accurately reflect the activation timing, intensity and fatigue state of specific muscle groups; Electroencephalogram (EEG) can see the concentration and cognitive resource allocation of the trainee. These internal states are the fundamental driving force of external action performance, but their changes cannot be obtained by naked eye observation.
[0011] (Three) real-time feedback lag, miss the correction golden window
[0012] Human trainers are difficult to track multiple trainees at the same time in dynamic and fast training scenarios, and capture all subtle errors in real time. Most problems need to be pointed out through video playback after training, and this lagging feedback makes the trainee unable to establish correct muscle memory and conditioned reflex at the first time of action, and the training efficiency is greatly reduced.
[0013] (Four) serious data island phenomenon, lack of deep correlation mining
[0014] Even if some digital devices (such as high-speed cameras, heart rate monitoring bracelets) are introduced, the multi-source data (video, audio, physiological signals) generated are often in a state of mutual isolation. The existing technology lacks effective technical means to perform spatio-temporal alignment and deep fusion analysis on these heterogeneous data, so it cannot reveal the deep causal relationship between "external performance" and "internal state". For example, it cannot answer "why is there always a deformation of action in the later stage of training?", is it because a specific muscle group is tired (EMG signal drops), the heart and lung function reaches the limit (rapid breathing, heart rate increases dramatically), or attention is dispersed, without such correlation analysis, the optimization of training program has no way to start. SUMMARY
[0015] Therefore, it is necessary to provide a target behavior data intelligent identification method and system for simulation training tasks, which can improve the objectivity and efficiency of multi-modal training data evaluation.
[0016] A target behavior data intelligent identification method for simulation training tasks is applied to a cloud-edge-end collaborative architecture, which includes an end-side sensing layer, an edge computing layer, a cloud-end intelligent center and an application interaction layer, and the method comprises:
[0017] The heterogeneous sensor network in the end-side sensing layer synchronously collects multi-modal behavior data under simulation training tasks, and the multi-modal behavior data includes visual data, auditory data and physiological data.
[0018] The cloud intelligent center is caused to add millisecond-level timestamps to multi-modal behavior data, and after adaptive preprocessing is respectively performed on each modal data, feature vectors corresponding to visual data, auditory data and physiological data are extracted.
[0019] A multi-modal heterogeneous graph including visual nodes, auditory nodes and physiological nodes is constructed, time sequence edges, cross-modal synchronization edges and semantic association edges are defined to represent relationships between nodes, and node features are preliminarily encoded by a modal-specific encoder.
[0020] A time sequence attention layer is applied to a feature sequence of visual space-time features, auditory feature vectors and physiological state feature vectors, the weighted multi-modal feature sequence is spliced and input into a cross-modal attention module, and a fused multi-modal joint feature is obtained.
[0021] The multi-modal joint feature is input into a multi-task learning network, parallel task heads sharing bottom layer features are used to cooperatively complete an identification and evaluation task of target behavior data.
[0022] An intelligent identification device for target behavior data for simulation training tasks, the device comprises:
[0023] A data acquisition module is configured to construct a heterogeneous sensor network on an end-side sensing layer to synchronously collect multi-modal behavior data under simulation training tasks, and the multi-modal behavior data includes visual data, auditory data and physiological data.
[0024] A feature vector extraction module is configured to cause the cloud intelligent center to add millisecond-level timestamps to multi-modal behavior data, and after adaptive preprocessing is respectively performed on each modal data, feature vectors corresponding to visual data, auditory data and physiological data are extracted.
[0025] An encoding module is configured to construct a multi-modal heterogeneous graph including visual nodes, auditory nodes and physiological nodes, define time sequence edges, cross-modal synchronization edges and semantic association edges to represent relationships between nodes, and preliminarily encode features of each node by a modal-specific encoder.
[0026] A fusion feature module is configured to apply a time sequence attention layer to a feature sequence of visual space-time features, auditory feature vectors and physiological state feature vectors, splice the weighted multi-modal feature sequence and input it into a cross-modal attention module, and obtain a fused multi-modal joint feature.
[0027] An identification and evaluation module is configured to input the multi-modal joint feature into a multi-task learning network, use parallel task heads sharing bottom layer features to cooperatively complete an identification and evaluation task of target behavior data.
[0028] The aforementioned intelligent recognition method and system for target behavior data in simulation training tasks overcomes the limitations of single-dimensional data by synchronously acquiring multimodal data (visual (motion trajectory), auditory (command / equipment interaction sound), and physiological (heart rate / electromyography) data through an edge-side heterogeneous sensor network. Simultaneously, it constructs a multimodal heterogeneous graph containing visual, auditory, and physiological nodes, defining temporal edges, cross-modal synchronization edges, and semantic association edges to represent node relationships. Furthermore, it uses a dedicated modal encoder to initially encode the features of each node, transforming discrete multimodal data into structured information with associated logic. This achieves a three-dimensional global insight into training behavior ("action-acoustic-physiological"), avoiding the biased evaluation caused by single data. Secondly, it leverages a cloud-edge-device collaborative architecture to decompose the computation process: the edge computing layer performs preliminary processing of multimodal data before uploading it to the cloud, significantly reducing data transmission volume and latency; the cloud uses millisecond-level timestamps to achieve efficient synchronization of multimodal data, coupled with a rapid processing chain for subsequent feature extraction and fusion, allowing recognition and evaluation results to be pushed promptly through the application interaction layer, ensuring the feedback rhythm aligns with the training rhythm and accurately covers the golden correction window. Then, spatiotemporal alignment of multimodal data is achieved using millisecond-level timestamps. Modal barriers are broken down through multimodal heterogeneous graphs. Temporal attention highlights key time steps for each modality, and cross-modal attention dynamically weights and fuses features. Finally, relying on the underlying feature sharing mechanism of the multi-task learning network, the coupling relationships between different modalities (such as the correlation between movement standardization and electromyographic intensity) are explored, realizing the synergistic release of data value. Finally, efforts are made synergistically from three aspects: perception dimension, feedback timeliness, and data correlation. This not only improves the objectivity of the assessment with multi-dimensional and correlated data support but also enhances the efficiency of the assessment through cloud-edge-device collaboration, effectively improving the accuracy of intelligent guidance in simulation training. Attached Figure Description
[0029] Figure 1 This is a flowchart illustrating a method for intelligent recognition of target behavior data for simulation training tasks in one embodiment.
[0030] Figure 2 This is a flowchart illustrating the intelligent recognition steps for target behavior data in a simulation training task, as shown in one embodiment.
[0031] Figure 3 This is a structural block diagram of a target behavior data intelligent recognition device for simulation training tasks in one embodiment. Detailed Implementation
[0032] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.
[0033] In one embodiment, such as Figure 1As shown, a target behavior data intelligent identification method for simulation training tasks is provided, which is applied to a cloud-edge-end collaborative architecture, and the cloud-edge-end collaborative architecture includes an end-side sensing layer, an edge computing layer, a cloud-end intelligent center and an application interaction layer, and includes the following steps:
[0034] Step 102, constructing a heterogeneous sensor network in the end-side sensing layer to synchronously collect multi-modal behavior data under the simulation training task, and the multi-modal behavior data includes visual data, auditory data and physiological data.
[0035] Specifically, the synchronous collection includes but is not limited to:
[0036] ① Visual perception unit: composed of multiple high-definition, wide-dynamic-range cameras deployed at different angles in the training site, preferably equipped with infrared modules to cope with low-light environments, for capturing the global motion trajectory and local joint posture of the trainee.
[0037] ② Auditory perception unit: composed of a high-sensitivity directional microphone array, for collecting the trainee's password, instruction response, breathing rhythm, instrument interaction sound and environmental sound.
[0038] ③ Physiological perception unit: composed of lightweight, low-invasive wearable devices such as heart rate bands, surface electromyography sensors, inertial measurement units, etc., for real-time monitoring of physiological and biochemical signals such as heart rate and heart rate variability, muscle activation sequence and intensity, limb acceleration and angular velocity, etc.
[0039] Further, synchronization and preprocessing: all sensors are synchronized through a unified hardware synchronization signal or software synchronization protocol to ensure that the data has a time stamp accurate to milliseconds. Preprocessing includes:
[0040] ① Visual data: lens distortion correction, image enhancement, background subtraction and human skeleton key point extraction.
[0041] ② Auditory data: noise reduction, pre-emphasis, framing and windowing to prepare for subsequent feature extraction.
[0042] ③ Physiological data: band-pass filtering (such as removing power frequency interference and baseline drift), signal smoothing and normalization.
[0043] Step 104, adding a millisecond-level timestamp to the multi-modal behavior data by the cloud-end intelligent center, and after performing adaptive preprocessing on each modal data, extracting the feature vectors corresponding to the visual data, auditory data and physiological data.
[0044] Specifically, visual spatio-temporal feature extraction: the extracted human skeleton key point sequence or the cropped video segment is input into a time series modeling network. Preferably, a spatio-temporal graph convolution network or a Transformer encoder is used to model the spatial topological relationship between the key points and the evolution rule of the action in the time dimension, and a high-dimensional feature vector capable of representing the action mode is output.
[0045] Further, auditory event and state feature extraction: the preprocessed audio data is input into a convolutional neural network on the one hand to extract deep features of its mel spectrum graph for identifying specific sound events (such as ball hitting sound); on the other hand, acoustic features such as mel frequency cepstral coefficient, spectral roll-off point are extracted and input into a bidirectional long short-term memory network to model the continuous change state of breathing sound and password sound, and a joint feature vector representing the sound content and state is output.
[0046] Further, physiological state and mode feature extraction: for the preprocessed physiological signal, not only the conventional time domain features (such as mean, variance), frequency domain features (such as power spectral density) are extracted, but also the nonlinear dynamic features (such as sample entropy, Lyapunov exponent) are extracted, which are extremely sensitive to the changes of states such as fatigue and stress. Finally, a comprehensive physiological state feature vector is formed by fusion.
[0047] Step 106, constructing a multi-modal heterogeneous graph containing visual nodes, auditory nodes, and physiological nodes, defining time series edges, cross-modal synchronous edges, and semantic association edges to represent the relationship between nodes, and preliminarily encoding the features of each node by modal-specific encoder.
[0048] Specifically, a heterogeneous graph G=(V,E) containing three types of nodes is constructed: visual nodes (Vvis): the human skeleton key point sequence is divided into K key frames, and the key point features of each key frame are taken as visual nodes; auditory nodes (Vaud): the audio stream is cut into time windows, and the acoustic features are extracted as auditory nodes; physiological nodes (Vphy): the physiological signals (heart rate, electromyogram, etc.) are cut into time-synchronous slices, and the physiological features are extracted as physiological nodes.
[0049] Further, the feature of each node is preliminarily encoded by a modal-specific encoder:
[0050] ;
[0051] wherein, is the initial encoding feature of the node , is the modal-specific encoder function, is the original feature of the node , a set of allowed learning parameters for a modality-specific encoder, a modality identifier, a visual modality, an auditory modality, a physiological modality.
[0052] Further, three types of edge relationships are defined: a temporal edge: a connection between nodes in the same modality at adjacent time steps; a cross-modality synchronous edge: a connection between nodes in different modalities within the same time window; and a semantic association edge: a cross-time, cross-modality semantic association dynamically discovered based on an attention mechanism.
[0053] At step 108, a temporal attention layer is applied to the feature sequence of the visual spatiotemporal features, the auditory feature vector, and the physiological state feature vector, respectively, and the weighted multi-modal feature sequence is concatenated and input into a cross-modality attention module to obtain a fused multi-modal joint feature.
[0054] Specifically, intra-modality temporal attention: first, a temporal attention layer is applied to the feature sequence of each modality. This layer can automatically learn and highlight the key time steps in each modality sequence that are most relevant to the current recognition task. For example, in the "shooting" action, the moment of pulling the trigger is the most critical for the visual and electromyographic signals.
[0055] Further, a cross-modality attention mechanism is designed to calculate the information flow weight between modalities using a gated attention mechanism:
[0056] ;
[0057] wherein, is the information transmission weight from node to node , is a sigmoid function, is concatenation, is a weight matrix of the attention gate, is the encoded feature vector of node , is the encoded feature vector of node , is a bias term of the attention gate.
[0058] Further, multi-granularity attention pooling is designed with a three-level attention mechanism: node-level attention: aggregating relevant information within the same modality; modality-level attention: calculating the overall importance of different modalities; and time-level attention: capturing cross-modality interactions at key time points.
[0059] Further, the multi-modal graph neural network propagation mechanism adopts a hierarchical message passing mechanism, each layer containing two stages, wherein the first stage is intra-modality feature aggregation:
[0060] ;
[0061] wherein, is the intermediate feature of the node in the network of the th layer, the intermediate feature of the node after intra-modal feature aggregation, is the intra-modal feature aggregation function, is the encoded feature vector of the node in the network of the th layer, the intra-modal neighbor nodes of the node , is the set of neighbor nodes of the node in the same modality;
[0062] The second stage is cross-modal feature aggregation:
[0063] ;
[0064] wherein, is the intermediate feature of the node in the network of the th layer, the intermediate feature of the node after cross-modal feature aggregation, is the information transmission weight from the node to the node , is the intermediate feature of the cross-modal node in the network of the th layer after intra-modal aggregation, is the set of cross-modal neighbor nodes of the node .
[0065] The intermediate features output according to the two stages are used to update the node features, to obtain the target feature vector:
[0066] ;
[0067] wherein, is the updated target feature vector of the node in the network of the th layer, is the feature update function, is the target feature vector of the node in the network of the th layerThe original feature vector without aggregation. Further, cross-modal adaptive attention: Then, the time-attention-weighted feature sequences of each modality are concatenated and input into a cross-modal attention module. This module dynamically calculates the contribution weights (α, β, γ) of visual, auditory, and physiological features to the final decision at a specific time based on the cross-attention mechanism of the Transformer. The mathematical expression can be simplified as:
[0068] ;
[0069] wherein the weight coefficients α, β, and γ dynamically change over time t and are automatically generated by the network according to the current context. This enables the system to flexibly cope with complex scenarios, such as automatically increasing the weights of the auditory and physiological modalities when the vision is blocked.
[0070] Further, dynamically adjust the neighborhood range of the node based on the attention weight: For edges with high attention weights, expand the neighborhood sampling range; for edges with low attention weights, reduce the neighborhood range to reduce computational complexity.
[0071] Step 110: Input the multi-modal joint features into a multi-task learning network, and cooperatively complete the identification and evaluation tasks of the target behavior data through parallel task heads sharing the underlying features.
[0072] Specifically, the unified and robust multi-modal joint features F_fused are input into a multi-task learning network to enable the network to share the underlying features, but parallel multiple specific task heads in the output layer to obtain the target decision.
[0073] Further, four joint training objectives are designed: main task loss: cross-entropy loss for training the state recognition task; correlation prediction loss: contrastive loss for predicting the correlation between modalities; reconstruction loss: self-supervised loss for reconstructing the features of each modality from the multi-modal fusion features; consistency loss: regularization loss to ensure consistency of cross-modal prediction results. The multi-task learning network completes the knowledge transfer of technical score, fatigue state warning, and action specification recognition of the target behavior data by sharing the underlying features.
[0074] The total loss function is:
[0075] ;
[0076] wherein, is the total loss value of the multi-task learning network, which is used to guide the optimization of model parameters and balance the training objectives of each task, , , , are the weight coefficients of each loss term, which are used to adjust the contribution proportion of different losses to the total loss, is the main task loss (cross-entropy loss), corresponding to the prediction error of the training state recognition task (such as action standard recognition, fatigue warning), is the correlation prediction loss (contrast loss), used to optimize the prediction accuracy of the inter-modal correlation, is the reconstruction loss (self-supervised loss), which measures the reconstruction error of the multi-modal joint feature on each original modal feature, and enhances the feature representation ability, is the consistency loss (regularization loss), which ensures the consistency of the cross-modal prediction results and avoids inter-modal decision conflicts.
[0077] Further, the collaborative recognition and evaluation are realized: the architecture of multi-task learning not only improves the computing efficiency, but more importantly, through sharing the representation, the knowledge between different tasks can be mutually promoted, and the generalization ability of the model is improved.
[0078] Further, feedback generation: according to the target decision output by the model, the system does not give the same feedback, but generates personalized guidance strategies combined with the historical data archives of the trainee. For example, for beginners, the feedback may focus on the essentials of basic movements; for experts, the feedback may focus on efficiency improvement and detail optimization.
[0079] Further, feedback presentation: real-time / near real-time feedback through various human-computer interaction interfaces:
[0080] ① Real-time channel: through the trainee's bone conduction earphone or AR glasses, concise correction instructions are provided in the form of voice or text superposition.
[0081] ② Review channel: after a single training, a visual analysis report of multi-dimensional data linkage is automatically generated, which is displayed on a tablet computer or a large screen, and the action picture, physiological data curve and sound spectrum of the abnormal moment are played back synchronously, which intuitively reveals the root cause of the problem.
[0082] In the above method of intelligently identifying target behavior data for simulation training tasks, multi-modal data of vision (motion trajectory), hearing (password / instrument interaction sound), and physiology (heart rate / electromyogram) are synchronously collected through an end-side heterogeneous sensor network, breaking through the limitations of a single dimension. A multi-modal heterogeneous graph containing visual, auditory, and physiological nodes is constructed, defining time sequence edges, cross-modal synchronization edges, and semantic association edges to represent node relationships. Each node feature is preliminarily encoded by a modal-specific encoder, converting discrete multi-modal data into structured information with associated logic, achieving a three-dimensional global insight of "action-acoustic-physiology" of training behavior, and avoiding the one-sidedness of evaluation caused by single data. Second, the computing process is split by a cloud-edge-end collaborative architecture: after preliminary processing of multi-modal data by the edge computing layer, the data is uploaded to the cloud, significantly reducing data transmission volume and delay. The cloud efficiently synchronizes multi-modal data with millisecond-level timestamps, and cooperates with the subsequent fast processing link of feature extraction and fusion, so that the recognition evaluation results can be timely pushed through the application interaction layer, making the feedback rhythm match the training rhythm and accurately cover the correction golden window. Then, the multi-modal data is temporally and spatially aligned with millisecond-level timestamps. The multi-modal heterogeneous graph breaks down the modal barriers, and then highlights the key time steps of each modality through temporal attention, dynamically weights and fuses features through cross-modal attention, and finally relies on the underlying feature sharing mechanism of the multi-task learning network to excavate the coupling relationship of different modal data (such as the correlation between action specification and electromyogram intensity), realizing the collaborative release of data value. Finally, from the aspects of perception dimension, feedback timeliness, and data association, the system synergistically works, not only improving the evaluation objectivity with multi-dimensional and associated data, but also enhancing the evaluation efficiency through cloud-edge-end collaboration, effectively improving the intelligent guidance accuracy of simulation training.
[0083] In one embodiment, the heterogeneous sensor network includes a visual perception unit, an auditory perception unit, and a physiological perception unit. The visual perception unit is composed of multi-angle high-definition wide-dynamic-range cameras, which are used to capture the motion trajectory and joint posture of the trainee. The auditory perception unit is composed of a high-sensitivity directional microphone array, which is used to collect passwords, instrument interaction sounds, and breathing rhythms. The physiological perception unit is composed of lightweight wearable devices, which are used to monitor heart rate variability, muscle activation intensity, and limb inertia parameters.
[0084] In one embodiment, a temporal attention layer is applied to the feature sequence of visual spatio-temporal features, auditory feature vectors, and physiological state feature vectors, highlighting the key time steps related to target behavior recognition in each modality. Then, the weighted feature sequences of each modality are concatenated and input into a cross-modal attention module. The information flow weight is calculated through inter-modal attention gating. A multi-granularity attention pooling is adopted at the node level, modality level, and time level. The hierarchical message passing mechanism and adaptive neighborhood sampling strategy of the multi-modal graph neural network are combined to dynamically learn the weight coefficients of each modality over time. The multi-modal joint features are obtained by fusing the weight coefficients:
[0085] ;
[0086] wherein, is a multi-modal joint feature, 、 、 are respectively the modal weight coefficients corresponding to the visual spatio-temporal feature, the auditory feature vector and the physiological state feature vector which dynamically change over time, is a visual spatio-temporal feature, is an auditory feature vector, is a physiological state feature vector.
[0087] In one of the embodiments, a multi-modal heterogeneous graph is constructed by using a graph neural network; the nodes of the multi-modal heterogeneous graph include: visual nodes, auditory nodes and physiological nodes; the visual nodes are composed of joint node features of key frames divided by human body skeleton key point sequences; the auditory nodes are composed of acoustic features extracted by audio streams after time window division; the physiological nodes are composed of physiological features extracted by physiological signals after time synchronization slicing; in the edge relationship of the multi-modal heterogeneous graph, the time sequence edge is the connection between adjacent time step nodes in the same modal, the cross-modal synchronization edge is the node connection of different modes in the same time window, and the semantic association edge is the cross-time and cross-modal semantic association dynamically discovered based on the attention mechanism.
[0088] In one of the embodiments, the information flow weight is calculated by inter-modal attention gate:
[0089] ;
[0090] wherein, is the information transmission weight from node to node , is a sigmoid function, is splicing, is a weight matrix of the attention gate, is the encoded feature vector of node , is the encoded feature vector of node , is a bias term of the attention gate. The multi-granularity attention pooling is designed by aggregating related information in the same modal by node-level attention, calculating the overall importance of different modes by modal-level attention and capturing cross-modal interaction at key time points by time-level attention. The hierarchical transmission mechanism of the multi-modal graph neural network designed according to the multi-granularity attention pooling includes two stages, wherein the first stage is intra-modal feature aggregation:
[0091] ;
[0092] wherein, is the intermediate feature of the node in the network of the th layer, the intermediate feature of the node after intra-modal feature aggregation, is the intra-modal feature aggregation function, is the encoded feature vector of the node in the network of the th layer, the homomodal neighbor node of the node , is the neighbor node set of the node in the same modality. The second stage is cross-modal feature aggregation:
[0093] ;
[0094] wherein, is the intermediate feature of the node in the network of the th layer, the intermediate feature of the node after cross-modal feature aggregation, is the information transmission weight from the node to the node , is the intermediate feature of the cross-modal node in the network of the th layer after intra-modal aggregation, is the cross-modal neighbor node set of the node . The intermediate features output by the two stages are used to update the node features to obtain the target feature vector:
[0095] ;
[0096] wherein, is the updated target feature vector of the node in the network of the th layer, is the feature update function, is the original feature vector of the node in the network of the th layer without aggregation. The neighbor sampling range is dynamically adjusted according to the attention weight, i.e., the neighbor sampling range is expanded for edges with high attention weight and is reduced for edges with low attention weight. The multi-modal joint feature is obtained by fusing the target feature vector and the dynamically learned modal weight coefficient changing over time.
[0097] In one of the embodiments, the node features are preliminarily encoded by a modality-specific encoder:
[0098] ;
[0099] wherein, is an initial encoding feature of a node , is a modality-specific encoder function, is an original feature of a node , is a set of allowed learning parameters of a modality-specific encoder, is a modality identifier, is a visual modality, is an auditory modality, is a physiological modality.
[0100] In one embodiment, the multi-modal graph neural network is combined with a cross-modal attention mechanism, and a cross-modal correlation matrix that allows interpretation is obtained through correlation matrix learning:
[0101] ;
[0102] wherein, is a correlation strength between modalities, Q is a feature of a query modality, K is a feature of a key modality, and d is a feature dimension. The cloud intelligent center adds a millisecond-level timestamp to the multi-modal behavior data, performs adaptive preprocessing on each modality data respectively, inputs the preprocessed visual data into a spatio-temporal graph convolution network or a Transformer encoder, extracts visual spatio-temporal features representing action patterns. The preprocessed auditory data is input into a convolutional neural network to extract Mel spectrum graph deep features, and a bidirectional long short-term memory network is used to model the continuous change state of acoustic features, and the auditory feature vector is fused. The preprocessed physiological data is extracted to obtain time domain features, frequency domain features, sample entropy, and Lyapunov exponent nonlinear dynamic features, and the physiological state feature vector is fused.
[0103] In one embodiment, the target decision generated by the target behavior data recognition and evaluation task is combined with the historical data archive of the trainee to generate a personalized guidance strategy. According to the personalized guidance strategy, the trainee's basic action essentials are prompted through the feedback of the motion sensing device, and the efficiency optimization details are guided according to the skilled trainee, so that the terminal device presents a multi-dimensional data linkage offline visual review report. The offline visual review report includes: action pictures, physiological data curves, and synchronous playback content of sound spectrum at abnormal moments.
[0104] In one embodiment, as shown in Figure 2 , a target behavior data intelligent recognition step for simulation training tasks is provided, and the specific steps are as follows:
[0105] S1: Multimodal data synchronous acquisition and preprocessing. Synchronize the acquisition of multimodal raw data of the trainee during the training process through a heterogeneous sensor network, and the multimodal raw data at least includes visual data, auditory data and physiological data, and timestamp alignment and preprocessing are performed on the raw data;
[0106] S2: Fine-grained multimodal feature extraction. Fine-grained time sequence features are extracted from each modality data after preprocessing, to obtain visual feature vectors, auditory feature vectors and physiological feature vectors;
[0107] S3: Hierarchical attention mechanism. A fusion network based on hierarchical attention mechanism is adopted to dynamically weight and fuse the visual feature vectors, auditory feature vectors and physiological feature vectors, to generate a unified multimodal joint feature vector; the hierarchical attention mechanism includes an intra-modal time sequence attention layer and a cross-modal adaptive attention layer;
[0108] S4: Multitask learning based collaboration. The multimodal joint feature vector is input into a multitask learning classifier, and multiple recognition and evaluation results are output in parallel, including at least action skill recognition and scoring, training phase recognition, physiological and psychological state evaluation, and abnormality and risk warning;
[0109] S5: Personalized adaptive feedback and decision support. Based on the recognition and evaluation results, combined with the historical data archives of the trainee, personalized guidance strategies are generated, and real-time or offline feedback and presentation are realized through human-computer interaction interface.
[0110] (1) The heterogeneous sensor network comprises:
[0111] 1. Visual perception unit composed of multiple high-definition cameras, used for capturing the global motion and local joint posture of the trainee;
[0112] 2. Auditory perception unit composed of directional microphone array, used for collecting the trainee's password, breathing sound and instrument interaction sound;
[0113] 3. Physiological perception unit composed of wearable devices, used for monitoring heart rate variability, muscle group electrical signals and inertial data.
[0114] (2) A fusion network based on hierarchical attention mechanism is adopted to dynamically weight and fuse the visual feature vectors, auditory feature vectors and physiological feature vectors, to generate a unified multimodal joint feature vector, and the specific steps are as follows:
[0115] 1. The visual spatio-temporal feature extraction is to model the sequence or video segment based on human body key points by using spatio-temporal graph convolution network or time sequence Transformer model;
[0116] 2. The auditory feature extraction is the recognition of sound events combined with a convolutional neural network and the modeling of sound states with a bidirectional long short-term memory network;
[0117] 3. The physiological feature extraction is the integration of its time domain, frequency domain and nonlinear dynamics features to form a physiological state feature vector.
[0118] (3) The hierarchical attention mechanism is specifically:
[0119] 1. The intra-modal temporal attention layer is used to process visual, auditory and physiological feature sequences respectively, to learn and highlight the key time steps in each modality that are most relevant to the current recognition task;
[0120] 2. The cross-modal adaptive attention layer receives the temporal attention weighted features of each modality and dynamically calculates the contribution weight of each modality feature to the final decision at each time. The weight coefficient changes over time and is used to generate the multi-modal joint feature vector.
[0121] (4) The cross-modal adaptive attention layer is realized based on the cross-attention mechanism of Transformer.
[0122] (5) Based on the recognition and evaluation results, combined with the historical data archives of the trainees, personalized guidance strategies are generated, and real-time or offline feedback and presentation are provided through human-computer interaction interfaces. The specific steps are:
[0123] 1. Real-time feedback is provided through the augmented reality glasses or bone conduction earphones of the trainee in the form of visual superposition or voice to provide correction instructions;
[0124] 2. Offline feedback is to generate a multi-dimensional data linked visual analysis report to synchronize and play back the action footage, physiological data curve and sound spectrum at the abnormal moment.
[0125] (6) The multi-modal data fusion intelligent recognition system for training scenarios adopts a cloud-edge-end collaborative architecture, including:
[0126] 1. The end-side sensing layer is composed of a heterogeneous sensor network and is used to collect multi-modal raw data;
[0127] 2. The edge computing layer is deployed on the training site side and is used to run preliminary processing tasks with extremely high delay requirements, including human skeleton key point extraction and simple anomaly real-time detection;
[0128] 3. The cloud intelligent center is the core processing unit of the system, including:
[0129] 1) Data management and synchronization engine for receiving, storing and aligning multi-modal data streams;
[0130] 2) Multi-modal feature extraction and fusion server, for performing the above steps ② and ③;
[0131] 3) Multi-task recognition and evaluation engine, for performing the above step ④;
[0132] 4) User portrait and model updating module, for storing the historical data of the trainee and optimizing the recognition model;
[0133] 4. Application interaction layer, including the coach command large screen, the trainee AR / VR device and the mobile terminal, for presenting the feedback and report of the above step ⑤.
[0134] (7) The edge computing layer is equipped with a graphics processing unit for accelerating the calculation process of the human body skeleton key point extraction, and uploading the processed feature data to the cloud intelligent center.
[0135] (8) The user portrait and model updating module can use new training data to perform incremental learning on the multi-task learning classifier and the fusion network of the hierarchical attention mechanism, to realize continuous optimization and personalized adaptation of the model.
[0136] It should be understood that, although Figures 1-2 the steps in the flowchart are shown in sequence according to the arrows, these steps are not necessarily executed in sequence according to the arrows. Unless otherwise specified herein, the execution of these steps is not strictly limited in sequence, and these steps can be executed in other orders. Moreover, Figures 1-2 at least part of the steps in may include multiple sub-steps or multiple stages, which are not necessarily executed at the same time, but can be executed at different times, and the execution order of these sub-steps or stages is not necessarily sequential, but can be alternately executed with other steps or sub-steps or stages of other steps.
[0137] Figure 3 In one embodiment, as shown in
[0138] The data acquisition module 302 is configured to construct a heterogeneous sensor network on the terminal side to synchronously collect multi-modal behavior data under the simulation training task, and the multi-modal behavior data includes visual data, auditory data and physiological data.
[0139] The feature vector extraction module 304 is configured to add millisecond-level timestamps to the multi-modal behavior data by the cloud intelligent center, and extract feature vectors corresponding to visual data, auditory data and physiological data after performing adaptive preprocessing on the respective modal data.
[0140] The encoding module 306 is configured to construct a multi-modal heterogeneous graph including visual nodes, auditory nodes and physiological nodes, define time sequence edges, cross-modal synchronization edges and semantic association edges to represent the relationships between the nodes, and perform preliminary encoding on the features of the nodes by using a modal-specific encoder.
[0141] The fusion feature module 308 is configured to apply a time sequence attention layer to the feature sequences of the visual spatio-temporal features, the auditory feature vectors and the physiological state feature vectors respectively, splice the weighted multi-modal feature sequences to input a cross-modal attention module, and obtain fused multi-modal joint features.
[0142] The recognition and evaluation module 310 is configured to input the multi-modal joint features into a multi-task learning network, and cooperatively complete the recognition and evaluation task of the target behavior data by using parallel task heads sharing underlying features.
[0143] The specific limitations of the target behavior data intelligent recognition device for the simulation training task can refer to the limitations of the target behavior data intelligent recognition method for the simulation training task described above, and will not be repeated here. Each module in the target behavior data intelligent recognition device for the simulation training task described above can be realized by software, hardware and combinations thereof in whole or in part. The above modules can be embedded in or independent of the processor in the computer device in hardware form, or can be stored in the memory in the computer device in software form, so as to be called and executed by the processor to perform the operations corresponding to each module.
[0144] Those skilled in the art can understand that, Figure 3 The structure shown in the figure is only a block diagram of part of the structure related to the scheme of the present application, and does not constitute a limitation on the computer device to which the scheme of the present application is applied. The specific computer device can include more or fewer components than those shown in the figure, or combine certain components, or have a different component arrangement.
[0145] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer readable storage medium, and when the computer program is executed, the processes of the above-mentioned embodiment methods can be included. Any reference to memory, storage, database or other medium used in the embodiments provided by the present application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. As an illustration but not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink), DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0146] Any combination of the technical features of the above embodiments can be made, and in order to make the description simple, all possible combinations of the technical features in the above embodiments are not described, however, as long as the combination of the technical features does not exist, it should be considered as the scope of the present application.
[0147] The above embodiments only express several embodiments of the present application, and the description is more specific and detailed, but it should not be understood as a limitation on the scope of the application. It should be pointed out that for ordinary skilled in the art, without departing from the concept of the present application, a number of modifications and improvements can be made, which are within the scope of the present application. Therefore, the protection scope of the present application should be subject to the appended claims.
Claims
1. A method for intelligent identification of target behavior data for simulation training tasks, characterized in that, The method is applied to a cloud edge end collaborative architecture, the cloud edge end collaborative architecture includes an end side sensing layer, an edge computing layer, a cloud end intelligent center and an application interaction layer, and the method includes: The end side sensing layer constructs a heterogeneous sensor network to synchronously collect multi-modal behavior data under a simulation training task, and the multi-modal behavior data is uploaded to the cloud end intelligent center through the computing layer; the multi-modal behavior data includes visual data, auditory data and physiological data; The cloud end intelligent center adds a millisecond level time stamp to the multi-modal behavior data, extracts feature vectors corresponding to the visual data, the auditory data and the physiological data after performing adaptive preprocessing on each modal data respectively, and inputs the feature vectors into a multi-task learning network; A multi-modal heterogeneous graph including visual nodes, auditory nodes and physiological nodes is constructed, time sequence edges, cross-modal synchronous edges and semantic association edges are defined to represent the relationship between nodes, and each node feature is preliminarily encoded by a modal dedicated encoder; A time sequence attention layer is applied to a feature sequence of visual space-time features, auditory feature vectors and physiological state feature vectors, the weighted multi-modal feature sequence is spliced and input into a cross-modal attention module to obtain fused multi-modal joint features; The multi-modal joint features are input into a multi-task learning network, and a parallel task head sharing bottom layer features is used to cooperatively complete a target behavior data recognition and evaluation task.
2. The method of claim 1, wherein, The heterogeneous sensor network includes a visual perception unit, an auditory perception unit and a physiological perception unit; The visual perception unit is composed of multi-angle high-definition wide-dynamic-range cameras, and is used to capture the motion trajectory and joint posture of a trainee; The auditory perception unit is composed of a high-sensitivity directional microphone array, and is used to collect passwords, instrument interaction sounds and breathing rhythms; The physiological perception unit is composed of lightweight wearable devices, and is used to monitor heart rate variability, muscle activation intensity and limb inertia parameters.
3. The method of claim 1, wherein, A time sequence attention layer is applied to a feature sequence of visual space-time features, auditory feature vectors and physiological state feature vectors, the weighted multi-modal feature sequence is spliced and input into a cross-modal attention module to obtain fused multi-modal joint features, including: A time sequence attention layer is applied to a feature sequence of visual space-time features, auditory feature vectors and physiological state feature vectors, the weighted multi-modal feature sequence is spliced and input into a cross-modal attention module to obtain fused multi-modal joint features, including: wherein, is a multi-modal joint feature, , , are modal weight coefficients corresponding to the visual spatio-temporal feature, the auditory feature vector and the physiological state feature vector respectively, is a visual spatio-temporal feature, is an auditory feature vector, is a physiological state feature vector.
4. The method of claim 3, wherein, A multi-modal heterogeneous graph including visual nodes, auditory nodes and physiological nodes is constructed, time sequence edges, cross-modal synchronous edges and semantic association edges are defined to represent the relationship between nodes, and each node feature is preliminarily encoded by a modal dedicated encoder, including: A graph neural network is used to construct a multi-modal heterogeneous graph; The nodes of the multi-modal heterogeneous graph include visual nodes, auditory nodes and physiological nodes; The visual node is composed of key frame joint features divided by human body skeleton key point sequence; The auditory node is composed of acoustic features extracted from audio stream according to time window division; The physiological node is composed of physiological features extracted from physiological signal according to time synchronization slicing; In the edge relationship of the multi-modal heterogeneous graph, the time sequence edge is the connection between the nodes in the same modal at adjacent time steps, the cross-modal synchronization edge is the connection between the nodes in different modal at the same time window, and the semantic association edge is the cross-time and cross-modal semantic association dynamically discovered based on the attention mechanism.
5. The method of claim 3, wherein, The information flow weight is calculated through inter-modal attention gate control, multi-granularity attention pooling is designed at node level, modal level and time level, the hierarchical message passing mechanism and adaptive neighborhood sampling strategy of the multi-modal graph neural network are combined, the modal weight coefficient changing over time is dynamically learned, and the multi-modal joint feature is obtained by fusing the weight coefficient, including: The information flow weight is calculated through inter-modal attention gate control: wherein, is the information passing weight from node to node , is a sigmoid function, is a concatenation, is a weight matrix of the attention gate, is the encoded feature vector of node , is the encoded feature vector of node , is a bias term of the attention gate; Multi-granularity attention pooling is designed by aggregating related information in the same modal through node-level attention, calculating the overall importance of different modal through modal-level attention, and capturing cross-modal interaction at key time points through time-level attention; The hierarchical transmission mechanism of the multi-modal graph neural network designed according to the multi-granularity attention pooling includes two stages, wherein the first stage is intra-modal feature aggregation: wherein, is the layer network, the node intermediate feature aggregated from intra-modal features, is an intra-modal feature aggregation function, is the layer network, the node encoded feature vector of the node is the node same-modal neighbor node of the node is a set of neighbor nodes of the node in the same modality; The second stage is cross-modal feature aggregation: wherein, is the first layer network, intermediate feature of the node after cross-modal feature aggregation, is a cross-modal feature aggregation function, is the information passing weight from the node to the node is the first layer network, intermediate feature of the cross-modal node after intra-modal aggregation, is the set of cross-modal neighbor nodes of the node The node features are updated according to the intermediate features output by the two stages to obtain a target feature vector: wherein, is the first layer network, the node updated target feature vector, is the feature update function, is the first layer network, the node raw feature vector without aggregation; The neighborhood sampling range of the node is dynamically adjusted according to the attention weight, the neighborhood sampling range is expanded for the edge with high attention weight, and the neighborhood sampling range is reduced for the edge with low attention weight; The multi-modal joint feature is obtained by fusing the target feature vector and the modal weight coefficient changing over time.
6. The method of claim 5, wherein, Each node feature is preliminarily encoded by a modal-specific encoder, including: Each node feature is preliminarily encoded by a modal-specific encoder: wherein, is an initial encoding feature of the node , is a modality-specific encoder function, is an original feature of the node , is a set of allowed learning parameters of the modality-specific encoder, is a modality identifier, is a visual modality, is an auditory modality, is a physiological modality.
7. The method of claim 6, wherein, So that the cloud intelligent center adds millisecond-level timestamps to the multi-modal behavior data, performs adaptive preprocessing on each modal data respectively, extracts feature vectors corresponding to the visual data, the auditory data and the physiological data, including: In the combination process of the multi-modal graph neural network and the cross-modal attention mechanism, the cross-modal association matrix allowed to be explained is learned through the association matrix: wherein, is the correlation strength between modalities, Q is the feature of the query modality, K is the feature of the key modality, and d is the feature dimension; So that the cloud intelligent center adds millisecond-level timestamps to the multi-modal behavior data, performs adaptive preprocessing on each modal data respectively, inputs the preprocessed visual data into a spatio-temporal graph convolution network or a Transformer encoder, and extracts visual spatio-temporal features representing action patterns; The preprocessed auditory data is respectively input into a convolutional neural network to extract mel spectrum graph deep features, and input into a bidirectional long short-term memory network to model the continuous change state of acoustic features, and the auditory feature vector is obtained by fusion; And the preprocessed physiological data is extracted to obtain time domain features, frequency domain features, sample entropy and Lyapunov exponent nonlinear dynamic features, and the physiological state feature vector is obtained by fusion.
8. A target behavior data intelligent identification device for simulation training tasks, characterized in that, The device comprises: The data acquisition module is configured to synchronously collect multi-modal behavior data under a simulation training task by constructing a heterogeneous sensor network on an end-side sensing layer, wherein the multi-modal behavior data comprises visual data, auditory data and physiological data; The feature vector extraction module is configured to add a millisecond-level timestamp to the multi-modal behavior data of the cloud intelligent center, extract feature vectors corresponding to the visual data, the auditory data and the physiological data after performing adaptive preprocessing on each modality data, respectively. The encoding module is configured to construct a multi-modal heterogeneous graph comprising a visual node, an auditory node and a physiological node, define a time sequence edge, a cross-modal synchronization edge and a semantic association edge to represent the relationship between the nodes, and perform preliminary encoding on the feature of each node by using a modality-specific encoder. The fusion feature module is configured to apply a time sequence attention layer to the feature sequence of the visual spatio-temporal feature, the auditory feature vector and the physiological state feature vector, respectively, splice the weighted multi-modal feature sequence to input a cross-modal attention module, and obtain a fused multi-modal joint feature. The recognition and evaluation module is configured to input the multi-modal joint feature into a multi-task learning network, complete a recognition and evaluation task of target behavior data by using a parallel task head sharing a bottom layer feature.
Citation Information
Patent Citations
Behavior identification method and system based on buried leakage cable and multi-mode sensing
CN118747321A
User behavior prediction system and method based on multi-modal data fusion
CN120832498A