Spaceflight exercise equipment human-computer interaction optimization method and system based on multi-modal sensing

By employing recursive optimization algorithms and cross-modal semantic association graph technology, the problems of temporal alignment and intent recognition of multi-source sensor data in aerospace training equipment were solved, enabling personalized intelligent human-computer interaction for astronauts and improving training effectiveness and safety.

CN121578893BActive Publication Date: 2026-04-21HUNAN VOCATIONAL INST OF TECH +1
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
HUNAN VOCATIONAL INST OF TECH
Filing Date
2026-01-26
Publication Date
2026-04-21

AI Technical Summary

Technical Problem

The existing human-computer interaction system of aerospace training equipment cannot effectively handle the temporal alignment and semantic association of multi-source heterogeneous sensor data, lacks the ability to recognize forward-looking intentions, and cannot make dynamic adjustments according to individual differences and real-time physiological states of astronauts, thus affecting training effectiveness and safety.

Method used

By using a recursive optimization algorithm to deeply analyze multi-source heterogeneous sensor data, a time baseline axis is constructed, multimodal feature sequences are generated, and deep feature enhancement and dynamic evolution are performed within the cross-modal semantic association graph to generate intent representation vectors. Combined with fatigue accumulation assessment, a forward-looking interactive intent prediction is generated to optimize the response of the interactive system.

Benefits of technology

It improves the accuracy and stability of multimodal data fusion, enhances the precision of interactive intent recognition, transforms the system response from passive to proactive, significantly improves the naturalness and intelligence of human-computer interaction, and adapts to individual differences and state changes among astronauts.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121578893B_ABST
    Figure CN121578893B_ABST
Patent Text Reader

Abstract

This invention provides a method and system for optimizing human-computer interaction in aerospace training equipment based on multimodal sensing, belonging to the field of human-computer interaction technology. The method includes acquiring and parsing multi-source heterogeneous sensor data to construct a time-series correlated multimodal feature sequence; adaptively constructing a cross-modal semantic association graph to generate an intent representation vector; parsing intent state features and combining them with fatigue assessment to generate a forward-looking interaction intent prediction; and generating control commands based on the prediction and feeding back to optimize the association graph. This invention improves the accuracy of human-computer interaction and enhances the real-time performance and personalized adaptability of equipment response.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to human-computer interaction technology, and more particularly to a method and system for optimizing human-computer interaction in aerospace training equipment based on multimodal sensing. Background Technology

[0002] During long-duration space missions, astronauts need specialized training equipment to combat the physiological deterioration effects of microgravity and maintain muscle strength, bone density, and cardiopulmonary function. Space training equipment has become a crucial facility on space stations and spacecraft, and its human-computer interaction performance directly impacts the training effectiveness and experience of astronauts. With the increasing duration of space exploration missions, the demands for the intelligence and personalization of space training equipment are constantly rising, making human-computer interaction optimization an important research direction in the field of aerospace medical engineering.

[0003] Traditional human-computer interaction systems for aerospace training equipment primarily rely on single or limited sensor data sources, such as force sensors and displacement sensors, for simple monitoring and feedback. These systems typically employ preset standard parameters and fixed interaction modes, failing to dynamically adjust based on individual astronaut differences and real-time physiological states. With the development of multimodal sensing technology, comprehensive analysis combining biomechanical parameters, physiological indicators, and operational behaviors has become possible, providing a new technological path for intelligent human-computer interaction in aerospace training equipment.

[0004] Existing technologies struggle to effectively handle the temporal alignment and semantic association issues of multi-source heterogeneous sensor data. Physiological signals, motion parameters, and environmental data generated during space training have different sampling frequencies and data structures, lacking a unified time reference and feature extraction method, thus hindering a comprehensive understanding of astronauts' training status.

[0005] Current interactive systems lack the ability to proactively recognize intentions and primarily rely on passive responses to astronauts' explicit actions. They cannot predict potential changes in astronauts' needs or fatigue transitions, thus missing optimal interaction opportunities and adjustment windows, affecting training effectiveness and safety. Summary of the Invention

[0006] This invention provides a method and system for optimizing human-computer interaction in aerospace training equipment based on multimodal sensing, which can solve the problems in the prior art.

[0007] A first aspect of this invention provides a method for optimizing human-computer interaction in aerospace training equipment based on multimodal sensing, comprising:

[0008] Acquire multi-source heterogeneous sensor data when astronauts use exercise equipment, deeply analyze the multi-source heterogeneous sensor data through a recursive optimization algorithm, dynamically construct a time reference axis based on physiological event nodes, calculate the high-dimensional projection relationship of sensor signals with different sampling frequencies, and generate a time-series correlated multimodal feature sequence.

[0009] Based on the multimodal feature sequence, an adaptive cross-modal semantic association graph is constructed. By performing deep feature enhancement and dynamic evolution within the cross-modal semantic association graph, an intent representation vector with spatiotemporal features is generated.

[0010] The intent representation vector is parsed to obtain intent state features. The intent state features are decomposed using the cross-modal semantic association graph and combined with real-time fatigue accumulation evaluation results to generate a forward-looking interactive intent prediction.

[0011] Based on the predicted forward interaction intent, precise control commands are intelligently generated and actual response data of execution feedback is obtained. By quantifying the multidimensional consistency deviation between the predicted forward interaction intent and the actual response data, the edge weight distribution and state transition probability value of the cross-modal semantic association graph are continuously iteratively optimized.

[0012] The multi-source heterogeneous sensor data is deeply analyzed using a recursive optimization algorithm. A time reference axis is dynamically constructed based on physiological event nodes. The high-dimensional projection relationship of sensor signals with different sampling frequencies is calculated, and a time-series correlated multimodal feature sequence is generated, including:

[0013] By using a recursive optimization algorithm to deeply analyze multi-source heterogeneous sensor data, physiological event feature points in each modal sensor signal are identified, and a time reference axis is dynamically constructed based on the distribution pattern of the physiological event feature points in the time series.

[0014] Analyze the timing relationship between the sensor signals with different sampling frequencies and the time reference axis in the multi-source heterogeneous sensor data, calculate the time offset of each sensor signal relative to the time reference axis, and establish a time mapping function from each sensor signal to the time reference axis based on the time offset;

[0015] The time mapping function is used to calculate the high-dimensional projection relationship of sensor signals with different sampling frequencies under a unified time scale. The high-dimensional projection relationship is applied to the temporal position alignment of each modal sensor signal to optimize the spatiotemporal consistency of each modal sensor signal, and finally generate a multimodal feature sequence that maintains cross-modal causal correlation.

[0016] By using a recursive optimization algorithm to deeply analyze multi-source heterogeneous sensor data, physiological event feature points in each modal sensor signal are identified. Based on the distribution pattern of these physiological event feature points over time, a time reference axis is dynamically constructed, including:

[0017] The recursive optimization algorithm is used to deeply analyze multi-source heterogeneous sensor data. In each recursive iteration, the sensor signals of each modality are decomposed into multi-scale time series and local time series feature vectors at each scale are extracted. Based on the local time series feature vectors, the temporal gradient change rate and frequency domain energy distribution are calculated. The physiological event feature points in each modality sensor signal are identified by the joint discrimination criterion of the temporal gradient change rate and the frequency domain energy distribution.

[0018] Based on the occurrence times of the physiological event feature points in the original time series, sort them to generate a time series set of event feature points for each modality. Calculate the time interval between adjacent physiological event feature points in the time series set of event feature points. Obtain a time interval distribution histogram through statistical analysis of the time intervals. Determine the periodic rhythm parameters of each modality's physiological events based on the peak position of the time interval distribution histogram.

[0019] A cross-modal physiological event correlation matrix is ​​constructed. The element values ​​of the correlation matrix are filled based on the temporal correspondence of physiological event feature points between different modalities. The correlation matrix is ​​used to identify physiological event feature point pairs with causal correlation. The average occurrence time of the physiological event feature point pairs is determined as the cross-modal synchronization reference point.

[0020] A time reference axis is dynamically constructed by combining the distribution pattern of physiological event feature points with the cross-modal synchronization reference points.

[0021] Based on the multimodal feature sequence, a cross-modal semantic association graph is adaptively constructed. Through deep feature enhancement and dynamic evolution within the cross-modal semantic association graph, an intent representation vector with spatiotemporal features is generated, including:

[0022] Based on the adaptive construction of a cross-modal semantic association graph using multimodal feature sequences, the feature vectors at each time step in the multimodal feature sequences are used as nodes in the cross-modal semantic association graph. The physiological-motor causal strength between different modal feature vectors is calculated. Based on the magnitude of the physiological-motor causal strength, it is determined whether to establish a connection edge between nodes and the weight coefficient of the connection edge, thus forming a graph structure containing cross-modal dependencies.

[0023] Deep feature enhancement is performed within the cross-modal semantic association graph. The feature vectors of each node are aggregated in the neighborhood by graph convolution operation. The feature vectors of each node are weighted and fused with the feature vectors of their neighboring nodes according to the weight coefficients of the connecting edges to obtain an enhanced feature vector that fuses neighborhood information.

[0024] The enhanced feature vector is dynamically evolved, and the change path of each node feature vector in multiple graph convolution operations is recorded. The temporal pattern of feature evolution is extracted based on the change path, and the temporal pattern is encoded as a time dimension feature.

[0025] The enhanced feature vector is concatenated and fused with the temporal dimension feature, and the concatenated feature is mapped to a unified semantic space through feature projection transformation to generate an intent representation vector with spatiotemporal features.

[0026] Based on the adaptive construction of a cross-modal semantic association graph using multimodal feature sequences, the feature vectors at each time step in the multimodal feature sequences are used as nodes in the cross-modal semantic association graph. The physiological-motor causal strength between different modal feature vectors is calculated, including:

[0027] Based on the adaptive construction of a cross-modal semantic association graph using multimodal feature sequences, feature vectors of different modalities are extracted from the multimodal feature sequences. The feature vectors of each modality are arranged in chronological order, and the feature vectors at each time point are used as nodes of the cross-modal semantic association graph. A modality type label and a timestamp label are attached to each node to form a node set.

[0028] Based on the modality type label and timestamp label in the node set, all node pairs are traversed and their corresponding node feature vectors are extracted. The semantic similarity between node pairs is calculated using the node feature vectors. The correlation between physiological attributes and motion attributes is analyzed by combining the modality attribute information recorded in each node label.

[0029] Based on the timestamp labels, the temporal sequence relationship is determined, and the physiological signal feature vectors and motion signal feature vectors with temporal correlation are analyzed. By performing correlation analysis on the temporal change trends of the physiological signal feature vectors and the motion signal feature vectors, the influence of physiological signal changes on the subsequent evolution of motion signals is quantified, and the physiological-motor causal strength is obtained.

[0030] The intent representation vector is parsed to obtain intent state features. These features are then decomposed using the cross-modal semantic association graph. Finally, a forward-looking interactive intent prediction is generated by combining the real-time fatigue accumulation assessment results.

[0031] The intent representation vector is input into a feature parser and subjected to multi-layer nonlinear transformation to extract higher-order semantic information from the intent representation vector, thereby obtaining the intent state feature.

[0032] The intention state features are decomposed using a cross-modal semantic association graph. Based on the modality type label of each node in the cross-modal semantic association graph, the feature components corresponding to different modalities in the intention state features are identified. The causal transmission path between each feature component is traced along the connection edge in the cross-modal semantic association graph. Based on the causal transmission path, the intention state features are decomposed into physiological modality contribution components and motor modality contribution components.

[0033] Fatigue-sensitive features are extracted from the physiological modality contribution component and the motion modality contribution component respectively. The fatigue-sensitive features are correlated and mapped with the real-time fatigue accumulation assessment results. The fatigue impact correction is performed on the physiological modality contribution component and the motion modality contribution component in combination with the fatigue accumulation assessment results.

[0034] Based on the physiological modal contribution components and motion modal contribution components after fatigue correction, an intention evolution trajectory is constructed. By analyzing the changing trend of the intention evolution trajectory, the direction of future intention state changes is inferred, and a forward-looking interactive intention prediction is generated.

[0035] Based on the predicted forward interaction intent, precise control commands are intelligently generated and actual response data of the execution feedback is obtained. By quantifying the multidimensional consistency deviation between the predicted forward interaction intent and the actual response data, the edge weight distribution and state transition probability values ​​of the cross-modal semantic association graph are continuously iteratively optimized, including:

[0036] Based on the forward-looking interactive intent prediction, the control command sequence of the training equipment is generated, the control command sequence is executed, and the physiological data and motion state data of the astronauts are collected in real time. The physiological data and motion state data are used as the actual response data.

[0037] The temporal consistency, state matching degree, and physiological feedback deviation between the predicted forward interaction intent and the actual response data are calculated to obtain a multidimensional consistency deviation value. Based on the multidimensional consistency deviation value, the optimization direction of the cross-modal semantic association graph is determined, the edge weight distribution between nodes in the cross-modal semantic association graph is dynamically adjusted, and the state transition probability value is updated to improve the prediction accuracy of the system.

[0038] A second aspect of the present invention provides a human-computer interaction optimization system for aerospace training equipment based on multimodal sensing, comprising:

[0039] The analysis unit is used to acquire multi-source heterogeneous sensor data when astronauts use training equipment. It deeply analyzes the multi-source heterogeneous sensor data through a recursive optimization algorithm, dynamically constructs a time reference axis based on physiological event nodes, calculates the high-dimensional projection relationship of sensor signals with different sampling frequencies, and generates a time-series associated multimodal feature sequence.

[0040] The generation unit is used to adaptively construct a cross-modal semantic association graph based on the multimodal feature sequence, and generate an intent representation vector with spatiotemporal features by performing deep feature enhancement and dynamic evolution within the cross-modal semantic association graph.

[0041] The decomposition unit is used to parse the intent representation vector to obtain intent state features, decompose the intent state features using the cross-modal semantic association graph, and generate a forward-looking interactive intent prediction by combining the real-time fatigue accumulation evaluation results.

[0042] The iterative unit is used to intelligently generate precise control commands based on the predicted forward interaction intent and obtain the actual response data of the execution feedback. By quantifying the multidimensional consistency deviation between the predicted forward interaction intent and the actual response data, it continuously iteratively optimizes the edge weight distribution and state transition probability value of the cross-modal semantic association graph.

[0043] A third aspect of the present invention provides an electronic device, comprising:

[0044] processor;

[0045] Memory used to store processor-executable instructions;

[0046] The processor is configured to invoke instructions stored in the memory to execute the aforementioned method.

[0047] A fourth aspect of the present invention provides a computer-readable storage medium having stored thereon computer program instructions that, when executed by a processor, implement the aforementioned method.

[0048] The beneficial effects of this application are as follows:

[0049] By using a recursive optimization algorithm to deeply analyze multi-source heterogeneous sensor data and dynamically constructing a time reference axis based on physiological event nodes, the problem of timing alignment between sensor signals with different sampling frequencies is solved, effectively improving the accuracy and stability of multimodal data fusion.

[0050] The adaptively constructed cross-modal semantic association graph can effectively capture the complex relationships between different perceptual modalities. Through deep feature enhancement and dynamic evolution mechanisms, the intent representation vector has richer spatiotemporal feature information, which improves the accuracy of interactive intent recognition.

[0051] Based on the results of intention state feature decomposition and fatigue accumulation assessment, the forward-looking interactive intention prediction can detect the potential changes in astronauts' needs in advance, enabling the system response to shift from passive to proactive, and significantly improving the naturalness and intelligence of human-computer interaction.

[0052] By quantifying the multidimensional consistency deviation between the predicted forward-looking interaction intent and the actual response data, continuous iterative optimization of the cross-modal semantic association graph is achieved, enabling the system to have self-learning capabilities, adapt to individual differences and state changes of astronauts, and continuously optimize the interactive experience. Attached Figure Description

[0053] Figure 1 This is a flowchart illustrating the human-computer interaction optimization method for aerospace training equipment based on multimodal sensing, according to an embodiment of the present invention.

[0054] Figure 2 This is a flowchart illustrating the generation of cross-modal intent representations based on physiological-motor causal intensity in an embodiment of the present invention. Detailed Implementation

[0055] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0056] The technical solution of the present invention will be described in detail below with reference to specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments.

[0057] Figure 1 This is a flowchart illustrating the human-computer interaction optimization method for aerospace training equipment based on multimodal sensing, as described in an embodiment of the present invention. Figure 1 As shown, the method includes:

[0058] Acquire multi-source heterogeneous sensor data when astronauts use exercise equipment, deeply analyze the multi-source heterogeneous sensor data through a recursive optimization algorithm, dynamically construct a time reference axis based on physiological event nodes, calculate the high-dimensional projection relationship of sensor signals with different sampling frequencies, and generate a time-series correlated multimodal feature sequence.

[0059] Based on the multimodal feature sequence, an adaptive cross-modal semantic association graph is constructed. By performing deep feature enhancement and dynamic evolution within the cross-modal semantic association graph, an intent representation vector with spatiotemporal features is generated.

[0060] The intent representation vector is parsed to obtain intent state features. The intent state features are decomposed using the cross-modal semantic association graph and combined with real-time fatigue accumulation evaluation results to generate a forward-looking interactive intent prediction.

[0061] Based on the predicted forward interaction intent, precise control commands are intelligently generated and actual response data of execution feedback is obtained. By quantifying the multidimensional consistency deviation between the predicted forward interaction intent and the actual response data, the edge weight distribution and state transition probability value of the cross-modal semantic association graph are continuously iteratively optimized.

[0062] In one optional implementation, the multi-source heterogeneous sensor data is deeply analyzed using a recursive optimization algorithm. A time reference axis is dynamically constructed based on physiological event nodes. The high-dimensional projection relationship of sensor signals with different sampling frequencies is calculated to generate a time-series correlated multimodal feature sequence, including:

[0063] By using a recursive optimization algorithm to deeply analyze multi-source heterogeneous sensor data, physiological event feature points in each modal sensor signal are identified, and a time reference axis is dynamically constructed based on the distribution pattern of the physiological event feature points in the time series.

[0064] Analyze the timing relationship between the sensor signals with different sampling frequencies and the time reference axis in the multi-source heterogeneous sensor data, calculate the time offset of each sensor signal relative to the time reference axis, and establish a time mapping function from each sensor signal to the time reference axis based on the time offset.

[0065] The time mapping function is used to calculate the high-dimensional projection relationship of sensor signals with different sampling frequencies under a unified time scale. The high-dimensional projection relationship is applied to the temporal position alignment of each modal sensor signal to optimize the spatiotemporal consistency of each modal sensor signal, and finally generate a multimodal feature sequence that maintains cross-modal causal correlation.

[0066] The recursive optimization algorithm for deep analysis of multi-source heterogeneous sensor data employs an adaptive sliding window mechanism, setting the window length to 512 sampling points and the overlap ratio to 50%. Deep analysis is achieved through convergence judgment based on iterative error. The algorithm maintains a circular buffer of size 1024 to store multimodal sensor data, including ECG signals, pulse wave signals, respiration signals, and blood oxygen saturation signals. The buffer uses a dual-pointer structure, with read and write pointers tracking data consumption and production positions, respectively. Data compression is triggered when the buffer utilization exceeds 75%. During the recursive optimization process, the maximum number of iterations is set to 20, and the convergence threshold is 0.001. Iteration terminates when the mean square error change of three consecutive iterations is less than this threshold.

[0067] Physiological event feature point identification is achieved through multi-scale wavelet decomposition, with a decomposition level of 6 layers and the Daubechies4 wavelet basis function. R-wave detection in ECG signals employs a differential thresholding method, setting the threshold to 1.2 times the mean signal amplitude and a minimum interval of 0.6 seconds. Pulse wave signal peak detection uses first-derivative zero-crossing point judgment combined with second-derivative negative value confirmation, with peak interval constraints set between 0.4 and 1.5 seconds. Respiratory signal cycle detection uses an envelope extraction method, with a low-pass filter cutoff frequency of 2Hz, and peak-valley alternation verification to ensure detection accuracy. Oxygen saturation signal anomaly detection sets a change rate threshold of 5% per second; data points exceeding this threshold are marked as event feature points.

[0068] The time baseline is dynamically constructed based on the timestamp sequence of detected physiological event feature points, and a weighted average method is used to determine the baseline time nodes. Weight calculation considers signal quality indicators: ECG signal weight is set to 0.4, pulse wave signal weight to 0.3, respiration signal weight to 0.2, and blood oxygen signal weight to 0.1. The interval between baseline time nodes is controlled within the range of 0.5 seconds to 2 seconds, and missing time nodes are supplemented by linear interpolation. The time baseline is updated every 5 seconds, and a moving average window length of 30 seconds is used to ensure the smoothness and stability of the baseline.

[0069] The processing of sensor signal sampling frequency differences is achieved through a frequency detection module, supporting automatic identification of sampling frequencies within the range of 100Hz to 1000Hz. Frequency detection employs autocorrelation function peak analysis with a calculation window length of 2048 sampling points and a detection accuracy of 0.1Hz. When a sampling frequency change exceeding 1% is detected, a recalibration process is triggered. The time-series relationship analysis between signals of different sampling frequencies and the time reference axis uses a cross-correlation function, with a calculation delay window range of -100 milliseconds to +100 milliseconds and a step size of 1 millisecond.

[0070] The time offset is calculated using a phase difference estimation method, extracting the instantaneous phase of the signal through Hilbert transform, with a phase difference accuracy requirement of 0.01 radians. The offset calculation results undergo median filtering with a filtering window length of 7 sampling points to eliminate sudden offset estimation errors. The time offset is stored using a circular queue structure with a queue length of 100 elements, supporting historical offset queries and trend analysis.

[0071] The time mapping function is established using cubic spline interpolation with a node interval of 0.1 seconds and natural boundary constraints. The mapping function parameters include linear, quadratic, and cubic coefficients, all requiring double-precision floating-point numbers. The mapping function is updated incrementally, recalculating interpolation parameters only for time intervals where changes occur, reducing computational overhead. The effectiveness of the mapping function is validated through a reverse mapping error assessment, with an error threshold set at 2 milliseconds.

[0072] The high-dimensional projection relationship calculation employs principal component analysis (PCA) for dimensionality reduction, retaining principal components with a cumulative variance contribution rate of up to 95%. The projection matrix dimension is dynamically adjusted based on the number of input signal channels, supporting 2 to 16 channel sensor signal processing. Projection calculations are implemented using singular value decomposition (SVD), and numerical stability is guaranteed through condition number checks, with a condition number threshold set to 1000. Projection relationships are stored in a sparse matrix format, achieving a compression ratio typically exceeding 60%.

[0073] The projection relationship under a unified time scale is implemented through matrix multiplication, and the calculation precision uses single-precision floating-point numbers to meet real-time requirements. Data validity checks are implemented during the projection process, including numerical range verification and outlier detection. Outlier replacement uses linear interpolation of nearest valid values. The projection result cache adopts an LRU strategy, with a cache capacity of 512 time segments, each containing projection data for 100 sampling points.

[0074] Temporal alignment is achieved using a dynamic time warping algorithm, with the search window width set to 10% of the sampling frequency and the step size constrained by Sakoe-Chiba banding. During alignment, a cumulative distance matrix is ​​calculated, with its size dynamically allocated based on the signal length, and memory usage limited to 25% of available memory. Alignment path backtracking employs a greedy strategy, prioritizing the path node with the shortest distance.

[0075] Spatiotemporal consistency optimization is achieved through iterative least squares, with initial values ​​based on linear alignment results. Convergence is determined when the rate of change of the objective function is less than 0.1%. A regularization parameter of 0.01 is set during optimization to prevent overfitting. The consistency evaluation uses the Pearson correlation coefficient with a threshold of 0.85; signal segments below this threshold are marked as requiring realignment.

[0076] Multimodal feature sequence generation employs a feature-level fusion method, with the weights of each modality feature dynamically adjusted based on signal quality. The feature sequence sampling rate is uniformly set to 250Hz, with a temporal resolution of 4 milliseconds. The feature sequence length is set to a fixed window size of 10 seconds, with an overlap of 50%. Cross-modal causal relationships are maintained through a causal graph construction, where nodes represent feature variables, edges represent causal relationship strength, and a strength threshold is set to 0.3.

[0077] In one optional implementation, a recursive optimization algorithm is used to deeply analyze multi-source heterogeneous sensor data, identify physiological event feature points in each modal sensor signal, and dynamically construct a time reference axis based on the distribution pattern of the physiological event feature points in the time series, including:

[0078] The recursive optimization algorithm is used to deeply analyze multi-source heterogeneous sensor data. In each recursive iteration, the sensor signals of each modality are decomposed into multi-scale time series and local time series feature vectors at each scale are extracted. Based on the local time series feature vectors, the temporal gradient change rate and frequency domain energy distribution are calculated. The physiological event feature points in each modality sensor signal are identified by the joint discrimination criterion of the temporal gradient change rate and the frequency domain energy distribution.

[0079] Based on the occurrence times of the physiological event feature points in the original time series, sort them to generate a time series set of event feature points for each modality. Calculate the time interval between adjacent physiological event feature points in the time series set of event feature points. Obtain a time interval distribution histogram through statistical analysis of the time intervals. Determine the periodic rhythm parameters of each modality's physiological events based on the peak position of the time interval distribution histogram.

[0080] A cross-modal physiological event correlation matrix is ​​constructed. The element values ​​of the correlation matrix are filled based on the temporal correspondence of physiological event feature points between different modalities. The correlation matrix is ​​used to identify physiological event feature point pairs with causal correlation. The average occurrence time of the physiological event feature point pairs is determined as the cross-modal synchronization reference point.

[0081] A time reference axis is dynamically constructed by combining the distribution pattern of physiological event feature points with the cross-modal synchronization reference points.

[0082] A recursive optimization algorithm is employed for deep analysis of multi-source heterogeneous sensor data. In each recursive iteration, multi-scale temporal decomposition is performed on each modal sensor signal. Specifically, for input physiological signals such as electrocardiograms, photoplethysmography (PPG), and blood pressure waveforms, discrete wavelet transform is used to decompose the signals into sub-band signals of different frequency bands. The Daubechies wavelet basis function is selected, and the decomposition level is set to five levels to obtain low-frequency approximate components and high-frequency detail components. Local temporal feature vectors are extracted for each component, including statistical features such as signal amplitude, first derivative, second derivative, local kurtosis, and local skewness.

[0083] Based on the extracted local temporal feature vectors, the temporal gradient rate of change is calculated. This rate of change reflects the speed of signal change near a specific time point. The calculation method involves normalizing the absolute value of the signal's first derivative within a sliding window. The window size is set according to the characteristics of different physiological signals; for example, it is set to 250 milliseconds for electrocardiogram (ECG) signals. Simultaneously, the frequency domain energy distribution is calculated using short-time Fourier transform. The analysis window length is set to 512 points with an overlap rate of 50%, and the Hamming window function is used to reduce spectral leakage.

[0084] Physiological event feature points are identified using a joint criterion of temporal gradient rate of change and frequency domain energy distribution. The joint criterion is defined as follows: when the temporal gradient rate of change exceeds an adaptive threshold (signal mean plus twice the standard deviation) and the frequency domain energy shows a significant peak within a target frequency band (e.g., the 5-15Hz band corresponding to the ECG R wave), that time point is marked as a candidate feature point. The true physiological event feature points are then selected by verifying the morphological characteristics of the candidate feature points (e.g., the sharpness and symmetry of the ECG R wave).

[0085] After identifying the feature points of each modality of physiological signal, these feature points are sorted based on their occurrence times in the original time series to generate a time series set of event feature points for each modality. For example, for electrocardiogram (ECG) signals, a time series set containing all R-wave peak times is generated; for photoplethysmography (PPG) pulse wave signals, a time series set containing all pulse wave peak times is generated. The time interval between adjacent physiological event feature points in each time series is calculated to obtain an interval sequence. Statistical analysis is performed on these interval sequences to generate a histogram of time interval distribution, with the bar width set to 50 milliseconds and the statistical interval to be 0-2000 milliseconds. Based on the peak positions in the histogram, the periodic rhythm parameters of each modality of physiological event are determined. For example, if the peak of the ECG RR interval histogram appears at 800 milliseconds, the heartbeat cycle is approximately 800 milliseconds, and the heart rate is approximately 75 beats per minute.

[0086] A cross-modal physiological event correlation matrix is ​​constructed, with dimensions N×M, where N and M represent the number of feature points of the two different modalities, respectively. The elements of the correlation matrix are filled based on the temporal correspondence of the physiological event feature points. Specifically, if the time difference between the i-th feature point of the first modality and the j-th feature point of the second modality is less than a preset threshold (e.g., 200 milliseconds), then the element (i, j) in the correlation matrix is ​​assigned a value of 1; otherwise, it is assigned a value of 0.

[0087] The padded correlation matrix is ​​used to identify causally related pairs of physiological event features. The correlation matrix is ​​then normalized in rows and columns, and a spectral clustering algorithm is applied to group strongly correlated features into the same category. Within each category, the average time delay of feature points from different modalities is calculated. Pairs of feature points that meet physiological rationality (e.g., the peak pulse wave should lag behind the ECG R wave) and have stable delays are identified as causally related pairs. The average occurrence time of these pairs is then determined as the cross-modal synchronization reference point.

[0088] By combining the distribution patterns of physiological event feature points with cross-modal synchronization reference points, a coarse-grained time axis is dynamically constructed, establishing the cross-modal synchronization reference points as anchors and setting the intervals to the most important physiological cycles (such as the heartbeat cycle). Then, physiological event feature points from each modality are inserted to form a fine-grained time axis. Nonlinear correction is applied to the time axis to compensate for periodic fluctuations caused by physiological variability. The final generated time axis retains the original time series information and establishes precise time correspondences between different modal signals, providing a unified time reference standard for subsequent multimodal data fusion and analysis.

[0089] In application examples, this method is used to process synchronously acquired electrocardiogram (ECG), pulse wave, and respiratory waveform signals. By identifying the QRS complex and T wave terminations in the ECG, the systolic peak and diastolic trough in the pulse wave, and the inspiratory peak and expiratory trough in the respiratory waveform, a comprehensive time reference axis incorporating cardiac electrical activity, peripheral vascular response, and respiratory activity is constructed. Experimental results show that the constructed time reference axis can accurately reflect the temporal relationships between different physiological systems, particularly capturing the coupling pattern between heart rate variability and the respiratory cycle, providing a more comprehensive physiological function assessment basis for clinical diagnosis.

[0090] In one optional implementation, a cross-modal semantic association graph is adaptively constructed based on the multimodal feature sequence. Deep feature enhancement and dynamic evolution are performed within the cross-modal semantic association graph to generate an intent representation vector with spatiotemporal features, including:

[0091] Based on the adaptive construction of a cross-modal semantic association graph using multimodal feature sequences, the feature vectors at each time step in the multimodal feature sequences are used as nodes in the cross-modal semantic association graph. The physiological-motor causal strength between different modal feature vectors is calculated. Based on the magnitude of the physiological-motor causal strength, it is determined whether to establish a connection edge between nodes and the weight coefficient of the connection edge, thus forming a graph structure containing cross-modal dependencies.

[0092] Deep feature enhancement is performed within the cross-modal semantic association graph. The feature vectors of each node are aggregated in the neighborhood by graph convolution operation. The feature vectors of each node are weighted and fused with the feature vectors of their neighboring nodes according to the weight coefficients of the connecting edges to obtain an enhanced feature vector that fuses neighborhood information.

[0093] The enhanced feature vector is dynamically evolved, and the change path of each node feature vector in multiple graph convolution operations is recorded. The temporal pattern of feature evolution is extracted based on the change path, and the temporal pattern is encoded as a time dimension feature.

[0094] The enhanced feature vector is concatenated and fused with the temporal dimension feature, and the concatenated feature is mapped to a unified semantic space through feature projection transformation to generate an intent representation vector with spatiotemporal features.

[0095] like Figure 2 As shown, the method includes:

[0096] The cross-modal semantic association graph construction process employs a dynamic graph data structure. The graph node storage structure includes four core fields: a feature vector array, a timestamp, a modality type identifier, and a node status. Feature vectors are stored as floating-point one-dimensional arrays with a length of 256 dimensions, supporting a configurable range from 16 bits to 1024 dimensions. Timestamps are recorded using 64-bit long integers, achieving microsecond-level precision. Modality type identifiers use 8-bit unsigned integer encoding, supporting the differentiation of 256 different modality types. The node status field is defined using an enumeration type, containing three status values: active, dormant, and pending deletion.

[0097] The multimodal feature sequence input interface is designed for batch processing, with a maximum of 512 feature vectors per batch and a batch interval of 10 milliseconds. Feature vectors at each time step in the feature sequence are used to create graph nodes in chronological order, with a time step of 2 milliseconds corresponding to a sampling frequency of 500Hz. Duplicate detection is performed during node creation, determining node uniqueness based on a combined hash value of the timestamp and modality identifier. Graph node storage uses a hash table structure with an initial capacity of 4096 nodes and a load factor threshold of 0.75. When the threshold is exceeded, the table automatically expands to twice its original capacity.

[0098] The physiological-motor causal strength was calculated using Granger causality analysis, with a calculation window of 80 consecutive time steps, a window sliding step size of 40 time steps, and an overlap rate of 50%. The causal strength calculation was based on a linear regression model with an 8th order, optimized using the Bayesian information criterion. The causal relationships between different modal eigenvectors were determined using conditional independence tests, with a significance level set at 0.05, and p-values ​​calculated to six decimal places.

[0099] The numerical calculation of causal strength employs the F-statistic method, where the F-value is calculated by the ratio of the sum of squared residuals of the constrained and unconstrained models. Causal strength normalization uses a sigmoid function to map the original F-value to a range of 0 to 1, with mapping parameters set to a slope of 0.1 and a center point of 5.0. Causal strength calculation supports parallel processing using a multi-threaded pool model, with the number of threads set to 1.5 times the number of CPU cores, and a maximum of 100 node pairs processed per thread.

[0100] The connection edge establishment rules are based on a causal strength threshold judgment mechanism, with a primary threshold set to 0.25 and a secondary threshold set to 0.15. Node pairs with causal strength exceeding the primary threshold are established as strong connections, with the weight coefficient of the strong connection edge directly using the normalized causal strength value. Node pairs with causal strength between the secondary and primary thresholds are established as weak connections, with the weight coefficient of the weak connection edge being the normalized causal strength value multiplied by a decay factor of 0.6. Connection edges are stored in a sparse matrix format, and matrix compression uses a compressed sparse row format, achieving a memory usage optimization rate of over 85%.

[0101] The graph structure maintenance employs an incremental update strategy, performing a full update every 200 time steps. During incremental updates, the connection edges of newly added nodes are determined by calculating the causal strength with historical nodes, keeping the computational complexity at the O(n) level. Edge weights are dynamically adjusted based on a moving average method, with a sliding window length set to 10 update cycles. The weight update formula is the current weight multiplied by 0.7 plus the newly calculated weight multiplied by 0.3.

[0102] Graph convolution is implemented using spectral domain convolution. The Laplacian matrix is ​​calculated based on the graph's degree matrix and adjacency matrix. The diagonal elements of the degree matrix represent the node degrees, and the elements of the adjacency matrix represent the edge weights. The normalized Laplacian matrix is ​​calculated by multiplying the degree matrix by its negative 1 / 2 power by the adjacency matrix. Spectral decomposition uses the Lanczos algorithm for eigenvalue decomposition, with the number of eigenvalues ​​set to 10% of the total number of graph nodes, and the eigenvalue calculation accuracy maintained at the 1e-8 level.

[0103] The neighborhood aggregation process employs a weighted summation method, where the aggregation weights are derived from the product of the weight coefficients of the connecting edges and the spatial distance decay factor. The spatial distance decay factor is calculated based on the Euclidean distance of the nodes in the feature space, and the decay function adopts an exponential decay form with a decay parameter set to 0.1. Neighborhood aggregation of each node's feature vectors uses a message-passing mechanism. The message aggregation buffer size is set to the node degree multiplied by the feature dimension, and a nearest neighbor truncation strategy is used when the buffer overflows.

[0104] The weighted fusion computation employs an attention weight adjustment mechanism, where attention weights are calculated through the dot product attention of the query vector, key vector, and value vector. The query vector is the feature vector of the current node, while the key and value vectors are the feature vectors of neighboring nodes. Attention calculation uses scaled dot product attention, with the scaling factor being the reciprocal of the square root of the feature dimension. The multi-head attention mechanism uses eight attention heads, each with a feature dimension one-eighth of the original feature dimension.

[0105] The enhanced feature vector generation employs a residual connection structure with a residual connection weight set to 0.8 to ensure the preservation of original feature information. The feature enhancement process includes two graph convolutional layers with the GELU activation function and a dropout ratio of 0.1. Batch normalization is applied to the output of each convolutional layer, with normalization parameters including a scaling factor and an offset. The initial value of the scaling factor is 1.0, and the initial value of the offset is 0.0.

[0106] The dynamic evolution record employs a feature change trajectory tracking method, maintaining a feature history queue of length 20 for each node. The feature change path record includes three dimensions: feature vector difference, change rate, and change direction. The feature vector difference is calculated using the L2 norm, the change rate is obtained by dividing the difference by the time interval, and the change direction is quantified using cosine similarity. The change path storage uses a circular buffer structure; when the buffer is full, a first-in, first-out (FIFO) strategy is used to overwrite historical records.

[0107] The graph convolution operation iterates for 5 iterations, each generating a new feature vector state. Convergence is determined by the average L2 norm of the feature vector changes between adjacent iterations, with a convergence threshold of 0.001. If convergence fails, the process is forced to end at the 5th iteration to prevent infinite loops. The iteration process is asynchronous and parallel, and node updates use a red-black coloring scheme to avoid data races, with red and black nodes updating alternately.

[0108] Temporal pattern extraction is achieved using a sequence encoder. The encoder structure is a bidirectional long short-term memory network with 128 hidden layers and 3 layers. After bidirectional processing, the output dimension is 256. The input to the sequence encoder is a sequence of feature change paths with a fixed length of 20 time steps. The encoder is trained using a teacher-mandated strategy, and the training data consists of historical change path samples with a minimum of 10,000 samples.

[0109] The Long Short-Term Memory (LSTM) configuration includes parameter settings for the forget gate, input gate, and output gate. The forget gate bias is set to 1.0 to promote long-term memory retention. The input and output gate biases are set to 0.0, and the weight matrix uses an orthogonal initialization method. Cell state updates use the tanh activation function, and the gating mechanism uses the sigmoid activation function. The gradient clipping threshold is set to 1.0 to prevent gradient explosion.

[0110] The temporal dimension feature encoding employs a combination of positional encoding and learned encoding. Positional encoding is generated based on sine and cosine functions, with a 64-dimensional encoding dimension and frequency parameters distributed logarithmically from 1 to 1000. Learned encoding is implemented using a trainable embedding layer with a 64-dimensional embedding dimension and an embedding table size of 1000. The final temporal dimension feature is 128-dimensional, obtained through the concatenation of positional encoding and learned encoding.

[0111] Feature concatenation and fusion are implemented using tensor join operations. The concatenation dimension is the last dimension of the feature vector, the augmented feature vector has a dimension of 256, the temporal feature vector has a dimension of 128, and the concatenated feature vector has a dimension of 384. Dimension alignment is checked before the concatenation operation to ensure consistency between batch and sequence dimensions. After concatenation, layer normalization is applied, with normalization parameters including learnable scaling and offset parameters.

[0112] The feature projection transformation is implemented using a multilayer perceptron architecture, consisting of three fully connected layers. The first layer has an input dimension of 384 and an output dimension of 512, using ReLU activation. The second layer has an input dimension of 512 and an output dimension of 256, also using ReLU activation. The third layer has an input dimension of 256 and an output dimension of 128, with no activation function. Batch normalization and dropout regularization are added between each layer, with dropout ratios set to 0.1, 0.2, and 0.1, respectively.

[0113] The unified semantic space mapping is achieved through feature normalization and projection matrix transformation. Feature normalization employs the z-score normalization method, with the mean and standard deviation calculated statistically based on the training data. The projection matrix is ​​generated using principal component analysis, retaining 95% of the feature variance. The semantic space is set to 128 dimensions, supporting similarity calculations using cosine similarity and Euclidean distance.

[0114] The intent representation vector generation process includes a final L2 norm normalization step to ensure the vector magnitude is 1.0. The normalization process employs a numerically stable implementation, adding a small constant of 1e-12 to the denominator to prevent division by zero errors. The generated intent representation vectors support batch output, with the output format being a floating-point two-dimensional array. The first dimension represents the batch size, and the second dimension represents the feature size.

[0115] In one optional implementation, a cross-modal semantic association graph is adaptively constructed based on the multimodal feature sequence. The feature vectors at each time step in the multimodal feature sequence are used as nodes in the cross-modal semantic association graph. The calculation of the physiological-motor causal strength between different modal feature vectors includes:

[0116] Based on the adaptive construction of a cross-modal semantic association graph using multimodal feature sequences, feature vectors of different modalities are extracted from the multimodal feature sequences. The feature vectors of each modality are arranged in chronological order, and the feature vectors at each time point are used as nodes of the cross-modal semantic association graph. A modality type label and a timestamp label are attached to each node to form a node set.

[0117] Based on the modality type label and timestamp label in the node set, all node pairs are traversed and their corresponding node feature vectors are extracted. The semantic similarity between node pairs is calculated using the node feature vectors. The correlation between physiological attributes and motion attributes is analyzed by combining the modality attribute information recorded in each node label.

[0118] Based on the timestamp labels, the temporal sequence relationship is determined, and the physiological signal feature vectors and motion signal feature vectors with temporal correlation are analyzed. By performing correlation analysis on the temporal change trends of the physiological signal feature vectors and the motion signal feature vectors, the influence of physiological signal changes on the subsequent evolution of motion signals is quantified, and the physiological-motor causal strength is obtained.

[0119] Multimodal feature sequences are extracted from raw multimodal data. Taking a fitness training scenario as an example, the system can simultaneously collect users' physiological signals (such as heart rate, electromyography, and blood oxygen saturation) and motion signals (such as acceleration, angular velocity, and joint angles). For physiological signals, wavelet transform is used to extract frequency domain features, and for motion signals, a temporal convolutional network is used to extract motion features, forming a multimodal feature sequence.

[0120] The feature vectors of each modality are arranged in time series order. For example, the heart rate feature sampled once per second and the acceleration feature sampled ten times per second are sorted according to their timestamps. To maintain the time alignment of data from different modalities, interpolation is performed on data with inconsistent sampling frequencies to synchronize the features of each modality in the time dimension.

[0121] The feature vectors at each time step are used as nodes in a cross-modal semantic association graph. For example, at time t, the heart rate feature vector forms the node v_heart_t, and the acceleration feature vector forms the node v_acc_t. Simultaneously, a modality type label and a timestamp label are attached to each node. The modality type label identifies whether the node represents a physiological signal or a motion signal, and the timestamp label records the acquisition time corresponding to the node. In this way, a set of nodes containing all time steps and all modal features is formed.

[0122] Based on the formed set of nodes, all node pairs are traversed, and the semantic similarity between nodes is calculated. Specifically, for node i and node j, their corresponding feature vectors vi and vj are extracted, and the semantic correlation between the two feature vectors is calculated using cosine similarity.

[0123] Cosine similarity = dot product of two vectors / (product of the magnitudes of two vectors);

[0124] By combining the modality type labels of nodes, the association types of node pairs can be further distinguished: physiological-physiological association, motor-motor association, and physiological-motor association. For physiological-motor associations, the focus is on analyzing the mutual influence between physiological and motor attributes. For example, by analyzing the trend of similarity changes between heart rate and acceleration features, the impact of heart rate changes on athletic performance can be inferred.

[0125] Based on timestamp labels, node pairs with temporal correlation are further identified. For the physiological signal feature vector v_physio_t and motion signal feature vector v_motion_(t+Δt) with time sequence t and t+Δt, their temporal change trends are analyzed. Specifically, the sliding window method is used to extract feature change patterns within a continuous time window, such as the rising / falling trend of physiological signals and the strengthening / weakening pattern of motion signals.

[0126] To quantify the impact of physiological signal changes on the subsequent evolution of motion signals, a time-series cross-information analysis method is employed. Taking the heart rate feature sequence HR and acceleration feature sequence ACC as examples, the mutual information value MI(HR_t, ACC_(t+τ)) under different time delays is calculated, where τ represents the time delay. A larger mutual information value indicates a stronger predictive ability of heart rate changes on future acceleration changes, i.e., a more significant causal relationship. By comparing the mutual information values ​​under different time delays, the optimal delay time τ_opt can be determined, which represents the time delay at which the physiological signal influences the motion signal. Finally, based on the mutual information value under the optimal delay time, the physiological-motor causal strength is defined.

[0127] In practical applications, the constructed cross-modal semantic association graphs visualize the correlation strength and causal relationships between signals of different modalities, helping to understand the mechanisms by which physiological states affect athletic performance. For example, in long-distance running training, by analyzing the causal relationship between heart rate changes and cadence adjustments, early signs of athlete fatigue can be identified, allowing for timely adjustments to training intensity and preventing sports injuries caused by overtraining.

[0128] For different users, there are individual differences in their physiological-motor causal relationship. By adaptively adjusting the causal strength calculation parameters, such as the time window size and similarity threshold, the cross-modal semantic association graph can adapt to the characteristics of different users and provide personalized health monitoring and exercise guidance.

[0129] Based on the adaptive construction of cross-modal semantic association graphs using multimodal feature sequences, the semantic similarity between feature vectors of different modalities is calculated by using the feature vectors at each time step in the multimodal feature sequence as graph nodes, and their temporal change trends are analyzed. Finally, the physiological-motor causal strength is quantified, providing an effective method for the joint analysis of multimodal physiological and motion signals.

[0130] In one optional implementation, the intent representation vector is parsed to obtain intent state features, the intent state features are decomposed using the cross-modal semantic association graph, and a forward-looking interaction intent prediction is generated by combining real-time fatigue accumulation evaluation results, including:

[0131] The intent representation vector is input into a feature parser and subjected to multi-layer nonlinear transformation to extract higher-order semantic information from the intent representation vector, thereby obtaining the intent state feature.

[0132] The intention state features are decomposed using a cross-modal semantic association graph. Based on the modality type label of each node in the cross-modal semantic association graph, the feature components corresponding to different modalities in the intention state features are identified. The causal transmission path between each feature component is traced along the connection edge in the cross-modal semantic association graph. Based on the causal transmission path, the intention state features are decomposed into physiological modality contribution components and motor modality contribution components.

[0133] Fatigue-sensitive features are extracted from the physiological modality contribution component and the motion modality contribution component respectively. The fatigue-sensitive features are correlated and mapped with the real-time fatigue accumulation assessment results. The fatigue impact correction is performed on the physiological modality contribution component and the motion modality contribution component in combination with the fatigue accumulation assessment results.

[0134] Based on the physiological modal contribution components and motion modal contribution components after fatigue correction, an intention evolution trajectory is constructed. By analyzing the changing trend of the intention evolution trajectory, the direction of future intention state changes is inferred, and a forward-looking interactive intention prediction is generated.

[0135] The intent representation vector is input into a feature parser for processing. The feature parser employs a multi-layer neural network structure to achieve non-linear transformation. This neural network consists of four fully connected layers, each using a different number of neurons (512, 256, 128, and 64, respectively) to transform the feature dimensions. A ReLU activation function is added after each fully connected layer to enhance the network's non-linear expressive power, while a residual connection is added between the second and third layers to prevent the gradient vanishing problem. In this way, the feature parser can extract high-order semantic information from the intent representation vector, generating a compact feature representation containing the driver's current intent state, i.e., intent state features.

[0136] Intent state features are decomposed using a pre-constructed cross-modal semantic association graph, which consists of multiple nodes and connecting edges. Each node is labeled with a specific modality type, such as physiological modality nodes and motor modality nodes. Physiological modality nodes include physiological signals such as eye movement features, heart rate changes, and skin conductance responses, while motor modality nodes include body movement features such as steering wheel operation, pedal control, and head rotation. By reading the modality type label of each node, the feature components corresponding to different modalities in the intent state features can be identified.

[0137] In the cross-modal semantic association graph, connecting edges represent causal relationships between different feature components. A graph traversal algorithm is used to trace the causal transmission paths between feature components along the connecting edges. Specifically, the implementation is as follows: First, the intention state features are mapped to each node in the graph; then, a bidirectional depth-first search is performed, starting from the intention feature node and tracing the transmission paths related to physiological and motor modalities along the directed edges respectively; finally, the contributions of each node are aggregated according to the path weights, decomposing the intention state features into physiological modal contribution components and motor modal contribution components. These two components reflect the degree of influence of the driver's physiological state and behavioral patterns on intention formation, respectively.

[0138] Fatigue-sensitive features were extracted from the physiological and motor modal contribution components obtained from the decomposition. An attention mechanism was employed for feature extraction, generating a fatigue sensitivity matrix by learning the weights of elements highly correlated with fatigue state in the feature vector. For the physiological modality, features such as eye closure frequency and blink duration were emphasized; for the motor modality, indicators such as steering wheel fine-tuning frequency and reaction delay were emphasized. The extracted fatigue-sensitive features were correlated and mapped with real-time cumulative fatigue assessment results, and an adaptive weight adjustment mechanism was used to calculate the fatigue correction coefficient.

[0139] Fatigue accumulation assessment results are typically provided in real time by a fatigue monitoring system, including fatigue level (mild, moderate, severe) and duration. Based on this information, fatigue impact corrections are applied to the physiological modal contribution components and the motor modal contribution components. The correction process employs a nonlinear mapping function; as fatigue intensifies, the correction coefficient for physiological modal characteristics increases (reflecting the enhanced influence of physiological state on intent), while the correction coefficient for motor modal characteristics decreases (reflecting a decline in operational precision). The corrected contribution components more accurately reflect the driver's true intent characteristics under current fatigue conditions.

[0140] The intention evolution trajectory is constructed based on the physiological modal contribution components and motion modal contribution components after fatigue correction. Specifically, the corrected contribution components from the most recent N time points (e.g., the past 10 seconds) are arranged chronologically to form a temporal feature sequence. A recurrent neural network (e.g., LSTM or GRU) is used to process this temporal sequence to capture the dynamic changes in the intention state. By analyzing the hidden state output of the RNN, the direction vector and velocity factor of the intention evolution are extracted to predict the changes in the intention state at the next T time points (e.g., the next 3 seconds).

[0141] In analyzing the trajectory of intent evolution, the focus is on three key characteristics: curvature, velocity, and acceleration. Curvature reflects the drastic change in intent, velocity indicates the speed of intent shift, and acceleration refers to the strengthening or weakening of the trend of change in the trajectory. By comprehensively considering these factors, it is possible to predict the driver's actions in the near future, such as lane-changing intentions, deceleration intentions, or attention shifts. This proactive interactive intent prediction can identify the driver's behavioral intentions in advance, providing decision support for vehicle intelligent assistance systems.

[0142] In practical applications, this method can dynamically adjust the time window and sensitivity of intent prediction according to different driving scenarios (such as highway driving, urban roads, and traffic congestion), improving the accuracy and practicality of prediction through an adaptive mechanism. When high driver fatigue is detected, the system will generate early warning information and activate corresponding driver assistance functions to ensure driving safety.

[0143] In one optional implementation, based on the predicted forward interaction intent, precise control commands are intelligently generated and actual response data of the execution feedback is obtained. The edge weight distribution and state transition probability values ​​of the cross-modal semantic association graph are continuously iteratively optimized by quantifying the multidimensional consistency deviation between the predicted forward interaction intent and the actual response data, including:

[0144] Based on the forward-looking interactive intent prediction, the control command sequence of the training equipment is generated, the control command sequence is executed, and the physiological data and motion state data of the astronauts are collected in real time. The physiological data and motion state data are used as the actual response data.

[0145] The temporal consistency, state matching degree, and physiological feedback deviation between the predicted forward interaction intent and the actual response data are calculated to obtain a multidimensional consistency deviation value. Based on the multidimensional consistency deviation value, the optimization direction of the cross-modal semantic association graph is determined, the edge weight distribution between nodes in the cross-modal semantic association graph is dynamically adjusted, and the state transition probability value is updated to improve the prediction accuracy of the system.

[0146] Based on the predictive interaction intent, the system generates control command sequences for training equipment. In practical applications, a trained deep neural network model can be used to predict the astronaut's interaction intent. This model analyzes multimodal input information such as the astronaut's gaze movement trajectory, pre-action hand gestures, and voice commands to comprehensively determine the astronaut's next operational intent. For example, when the system detects that the astronaut's gaze is focused on a specific area on the control panel and a pre-action hand gesture occurs, it predicts that the astronaut needs to adjust the relevant equipment parameters in that area. Based on this prediction, the system automatically generates a series of control command sequences, such as {adjusting the power of device A by 10%, activating the backup system of device B, switching the display screen to mode C}, etc.

[0147] The system executes the control command sequence and collects the astronaut's physiological and motion data in real time. The control command sequence is sent to the training system execution layer through the command interface of the training equipment, simultaneously activating the multi-channel physiological data acquisition subsystem and the motion capture system. Physiological data acquisition includes, but is not limited to, indicators such as heart rate variability, respiratory rate, skin conductance response, and pupillary dilation, with a sampling frequency of no less than one hundred times per second to ensure data temporal accuracy. Motion data is recorded using high-precision motion capture equipment, capturing information such as the astronaut's postural changes, joint angles, movement speed, and acceleration before and after command execution, forming a multi-dimensional time-series data stream. This real-time collected physiological and motion data serves as actual response data for subsequent comparative analysis with predicted results.

[0148] The temporal consistency between predicted interactive intent and actual response data is calculated. Temporal consistency primarily evaluates the degree of alignment between predicted and actual actions over time. Specifically, a dynamic time warping algorithm is constructed to align the two time series, and the warped Euclidean distance is calculated as the temporal consistency metric. The temporal consistency calculation considers the order of action execution, duration, and relative temporal relationships. The deviation between predicted and actual actions on the time axis is quantified to obtain the temporal consistency score (ST).

[0149] Simultaneously, state matching degree is calculated to evaluate the spatial similarity between the predicted state and the actual execution state. State matching degree is quantified by vectorizing the predicted action sequence and the actual action sequence in the feature space, and then calculating the cosine similarity between the two sets of feature vectors. For complex action sequences, key pose points are first extracted to construct a pose skeleton model, and then the structural similarity between the skeleton models is calculated. By weighted fusion of matching degree indices at different levels, the comprehensive state matching degree score SM is finally obtained.

[0150] The calculation of physiological feedback bias primarily focuses on the difference between changes in astronauts' physiological indicators before and after operations and their expected responses. A personalized baseline model is established based on historical physiological data of the astronauts. The currently collected physiological indicators are compared with the predicted values ​​of this baseline model, and the standardized deviation value is calculated. Particular attention is paid to abnormal changes in physiological indicators under mission stress, such as sudden changes in heart rate and abnormal skin conductance, which indicate a mismatch between prediction and actual intention. A multi-indicator fusion algorithm is used to comprehensively obtain the physiological feedback bias score (SP).

[0151] The multidimensional consistency deviation value was obtained by combining the results of the three dimensions—temporal consistency score (ST), state matching score (SM), and physiological feedback deviation score (SP)—using a weighted summation method to integrate them into a single comprehensive deviation value (D).

[0152] D = α×ST + β×SM + γ×SP;

[0153] Wherein, α, β, and γ are the weight coefficients of each dimension, and α+β+γ=1. The weight coefficients can be dynamically adjusted according to different task scenarios and needs. For example, in high-precision operation tasks, the weight of state matching degree can be increased, and in high-pressure environments, the weight of physiological feedback deviation can be increased.

[0154] The optimization direction of the cross-modal semantic association graph is determined based on the multidimensional consistency deviation value. The cross-modal semantic association graph is a directed weighted graph structure where nodes represent semantic units from different modalities, and edges represent semantic relationships between modalities. When the multidimensional consistency deviation value exceeds a preset threshold, an adaptive optimization mechanism for the graph structure is triggered. The optimization direction is determined based on the distribution of specific deviation values ​​across dimensions. For example, if temporal consistency is low, the weights of temporal association edges in the graph are adjusted; if state matching is low, the representation of spatial semantic nodes is optimized.

[0155] The edge weight distribution between nodes in the cross-modal semantic association graph is dynamically adjusted. Gradient descent is used to iteratively update the edge weights, with the update formula as follows:

[0156] ;

[0157] Where w represents the edge weight and η is the learning rate. This represents the gradient of the multidimensional consistency deviation value with respect to the weights. To improve optimization efficiency, an adaptive learning rate strategy can be adopted, dynamically adjusting the learning rate based on the importance of different edges and historical updates. Simultaneously, sparse constraints are introduced to encourage the model to retain the most significant cross-modal correlations, thereby improving the model's interpretability and generalization ability.

[0158] The state transition probabilities are updated to improve the system's prediction accuracy. The state transition probability matrix P describes the probability distribution of the system transitioning from the current state to the next state. Based on multidimensional consistency deviation feedback, the state transition probabilities are updated using Bayesian methods.

[0159] P(s_j|s_i)(t+1) = P(s_j|s_i)(t) + λ × (Actual(s_j|s_i) - P(s_j|s_i)(t));

[0160] Where λ is the update step size, and Actual(s_j|s_i) is the frequency of actual observed transitions from state s_i to state s_j. By continuously accumulating actual state transition data, the system continuously adjusts the state transition probability values, making the prediction model more consistent with the actual operating habits and intention expression patterns of astronauts.

[0161] Through the above methods, the cross-modal semantic association graph continuously optimizes itself, and the prediction accuracy is continuously improved, thereby providing astronauts with more accurate understanding of interaction intentions and response support.

[0162] A second aspect of the present invention provides a human-computer interaction optimization system for aerospace training equipment based on multimodal sensing, comprising:

[0163] The analysis unit is used to acquire multi-source heterogeneous sensor data when astronauts use training equipment. It deeply analyzes the multi-source heterogeneous sensor data through a recursive optimization algorithm, dynamically constructs a time reference axis based on physiological event nodes, calculates the high-dimensional projection relationship of sensor signals with different sampling frequencies, and generates a time-series associated multimodal feature sequence.

[0164] The generation unit is used to adaptively construct a cross-modal semantic association graph based on the multimodal feature sequence, and generate an intent representation vector with spatiotemporal features by performing deep feature enhancement and dynamic evolution within the cross-modal semantic association graph.

[0165] The decomposition unit is used to parse the intent representation vector to obtain intent state features, decompose the intent state features using the cross-modal semantic association graph, and generate a forward-looking interactive intent prediction by combining the real-time fatigue accumulation evaluation results.

[0166] The iterative unit is used to intelligently generate precise control commands based on the predicted forward interaction intent and obtain the actual response data of the execution feedback. By quantifying the multidimensional consistency deviation between the predicted forward interaction intent and the actual response data, iteratively optimizes the edge weight distribution and state transition probability value of the cross-modal semantic association graph.

[0167] A third aspect of the present invention provides an electronic device, comprising:

[0168] processor;

[0169] Memory used to store processor-executable instructions;

[0170] The processor is configured to invoke instructions stored in the memory to execute the aforementioned method.

[0171] A fourth aspect of the present invention provides a computer-readable storage medium having stored thereon computer program instructions that, when executed by a processor, implement the aforementioned method.

[0172] This invention can be a method, apparatus, system, and / or computer program product. The computer program product may include a computer-readable storage medium having computer-readable program instructions loaded thereon for performing various aspects of the invention.

[0173] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for optimizing human-computer interaction in aerospace training equipment based on multimodal sensing, characterized in that, include: Acquire multi-source heterogeneous sensor data when astronauts use exercise equipment, deeply analyze the multi-source heterogeneous sensor data through a recursive optimization algorithm, dynamically construct a time reference axis based on physiological event nodes, calculate the high-dimensional projection relationship of sensor signals with different sampling frequencies, and generate a time-series correlated multimodal feature sequence. Based on the multimodal feature sequence, an adaptive cross-modal semantic association graph is constructed. By performing deep feature enhancement and dynamic evolution within the cross-modal semantic association graph, an intent representation vector with spatiotemporal features is generated. The intent representation vector is parsed to obtain intent state features. These features are then decomposed using the cross-modal semantic association graph. Finally, a forward-looking interaction intent prediction is generated by combining the real-time fatigue accumulation evaluation results. This includes: The intent representation vector is input into a feature parser and subjected to multi-layer nonlinear transformation to extract higher-order semantic information from the intent representation vector, thereby obtaining the intent state feature. The intention state features are decomposed using a cross-modal semantic association graph. Based on the modality type label of each node in the cross-modal semantic association graph, the feature components corresponding to different modalities in the intention state features are identified. The causal transmission path between each feature component is traced along the connection edge in the cross-modal semantic association graph. Based on the causal transmission path, the intention state features are decomposed into physiological modality contribution components and motor modality contribution components. Fatigue-sensitive features are extracted from the physiological modality contribution component and the motion modality contribution component respectively. The fatigue-sensitive features are correlated and mapped with the real-time fatigue accumulation assessment results. The fatigue impact correction is performed on the physiological modality contribution component and the motion modality contribution component in combination with the fatigue accumulation assessment results. Based on the physiological modal contribution components and motion modal contribution components after fatigue correction, an intention evolution trajectory is constructed. By analyzing the changing trend of the intention evolution trajectory, the direction of future intention state changes is inferred, and a forward-looking interactive intention prediction is generated. Based on the predicted forward interaction intent, precise control commands are intelligently generated and actual response data of execution feedback is obtained. By quantifying the multidimensional consistency deviation between the predicted forward interaction intent and the actual response data, the edge weight distribution and state transition probability value of the cross-modal semantic association graph are continuously iteratively optimized.

2. The method according to claim 1, characterized in that, The multi-source heterogeneous sensor data is deeply analyzed using a recursive optimization algorithm. A time reference axis is dynamically constructed based on physiological event nodes. The high-dimensional projection relationship of sensor signals with different sampling frequencies is calculated, and a time-series correlated multimodal feature sequence is generated, including: By using a recursive optimization algorithm to deeply analyze multi-source heterogeneous sensor data, physiological event feature points in each modal sensor signal are identified, and a time reference axis is dynamically constructed based on the distribution pattern of the physiological event feature points in the time series. Analyze the timing relationship between the sensor signals with different sampling frequencies and the time reference axis in the multi-source heterogeneous sensor data, calculate the time offset of each sensor signal relative to the time reference axis, and establish a time mapping function from each sensor signal to the time reference axis based on the time offset. The time mapping function is used to calculate the high-dimensional projection relationship of sensor signals with different sampling frequencies under a unified time scale. The high-dimensional projection relationship is applied to the temporal position alignment of each modal sensor signal to optimize the spatiotemporal consistency of each modal sensor signal, and finally generate a multimodal feature sequence that maintains cross-modal causal correlation.

3. The method according to claim 2, characterized in that, By using a recursive optimization algorithm to deeply analyze multi-source heterogeneous sensor data, physiological event feature points in each modal sensor signal are identified. Based on the distribution pattern of these physiological event feature points over time, a time reference axis is dynamically constructed, including: The recursive optimization algorithm is used to deeply analyze multi-source heterogeneous sensor data. In each recursive iteration, the sensor signals of each modality are decomposed into multi-scale time series and local time series feature vectors at each scale are extracted. Based on the local time series feature vectors, the temporal gradient change rate and frequency domain energy distribution are calculated. The physiological event feature points in each modality sensor signal are identified by the joint discrimination criterion of the temporal gradient change rate and the frequency domain energy distribution. Based on the occurrence times of the physiological event feature points in the original time series, sort them to generate a time series set of event feature points for each modality. Calculate the time interval between adjacent physiological event feature points in the time series set of event feature points. Obtain a time interval distribution histogram through statistical analysis of the time intervals. Determine the periodic rhythm parameters of each modality's physiological events based on the peak position of the time interval distribution histogram. A cross-modal physiological event correlation matrix is ​​constructed. The element values ​​of the correlation matrix are filled based on the temporal correspondence of physiological event feature points between different modalities. The correlation matrix is ​​used to identify physiological event feature point pairs with causal correlation. The average occurrence time of the physiological event feature point pairs is determined as the cross-modal synchronization reference point. A time reference axis is dynamically constructed by combining the distribution pattern of physiological event feature points with the cross-modal synchronization reference points.

4. The method according to claim 1, characterized in that, Based on the multimodal feature sequence, a cross-modal semantic association graph is adaptively constructed. Through deep feature enhancement and dynamic evolution within the cross-modal semantic association graph, an intent representation vector with spatiotemporal features is generated, including: Based on the adaptive construction of a cross-modal semantic association graph using multimodal feature sequences, the feature vectors at each time step in the multimodal feature sequences are used as nodes in the cross-modal semantic association graph. The physiological-motor causal strength between different modal feature vectors is calculated. Based on the magnitude of the physiological-motor causal strength, it is determined whether to establish a connection edge between nodes and the weight coefficient of the connection edge, thus forming a graph structure containing cross-modal dependencies. Deep feature enhancement is performed within the cross-modal semantic association graph. The feature vectors of each node are aggregated in the neighborhood by graph convolution operation. The feature vectors of each node are weighted and fused with the feature vectors of their neighboring nodes according to the weight coefficients of the connecting edges to obtain an enhanced feature vector that fuses neighborhood information. The enhanced feature vector is dynamically evolved, and the change path of each node feature vector in multiple graph convolution operations is recorded. The temporal pattern of feature evolution is extracted based on the change path, and the temporal pattern is encoded as a time dimension feature. The enhanced feature vector is concatenated and fused with the temporal dimension feature, and the concatenated feature is mapped to a unified semantic space through feature projection transformation to generate an intent representation vector with spatiotemporal features.

5. The method according to claim 4, characterized in that, Based on the adaptive construction of a cross-modal semantic association graph using multimodal feature sequences, the feature vectors at each time step in the multimodal feature sequences are used as nodes in the cross-modal semantic association graph. The physiological-motor causal strength between different modal feature vectors is calculated, including: Based on the adaptive construction of a cross-modal semantic association graph using multimodal feature sequences, feature vectors of different modalities are extracted from the multimodal feature sequences. The feature vectors of each modality are arranged in chronological order, and the feature vectors at each time point are used as nodes of the cross-modal semantic association graph. A modality type label and a timestamp label are attached to each node to form a node set. Based on the modality type label and timestamp label in the node set, all node pairs are traversed and their corresponding node feature vectors are extracted. The semantic similarity between node pairs is calculated using the node feature vectors. The correlation between physiological attributes and motion attributes is analyzed by combining the modality attribute information recorded in each node label. Based on the timestamp labels, the temporal sequence relationship is determined, and the physiological signal feature vectors and motion signal feature vectors with temporal correlation are analyzed. By performing correlation analysis on the temporal change trends of the physiological signal feature vectors and the motion signal feature vectors, the influence of physiological signal changes on the subsequent evolution of motion signals is quantified, and the physiological-motor causal strength is obtained.

6. The method according to claim 1, characterized in that, Based on the predicted forward interaction intent, precise control commands are intelligently generated and actual response data of the execution feedback is obtained. By quantifying the multidimensional consistency deviation between the predicted forward interaction intent and the actual response data, the edge weight distribution and state transition probability values ​​of the cross-modal semantic association graph are continuously iteratively optimized, including: Based on the forward-looking interactive intent prediction, the control command sequence of the training equipment is generated, the control command sequence is executed, and the physiological data and motion state data of the astronauts are collected in real time. The physiological data and motion state data are used as the actual response data. The temporal consistency, state matching degree, and physiological feedback deviation between the predicted forward interaction intent and the actual response data are calculated to obtain a multidimensional consistency deviation value. Based on the multidimensional consistency deviation value, the optimization direction of the cross-modal semantic association graph is determined, the edge weight distribution between nodes in the cross-modal semantic association graph is dynamically adjusted, and the state transition probability value is updated to improve the prediction accuracy of the system.

7. A human-computer interaction optimization system for aerospace training equipment based on multimodal sensing, used to implement the method of any one of claims 1-6, characterized in that, include: The analysis unit is used to acquire multi-source heterogeneous sensor data when astronauts use training equipment. It deeply analyzes the multi-source heterogeneous sensor data through a recursive optimization algorithm, dynamically constructs a time reference axis based on physiological event nodes, calculates the high-dimensional projection relationship of sensor signals with different sampling frequencies, and generates a time-series associated multimodal feature sequence. The generation unit is used to adaptively construct a cross-modal semantic association graph based on the multimodal feature sequence, and generate an intent representation vector with spatiotemporal features by performing deep feature enhancement and dynamic evolution within the cross-modal semantic association graph. The decomposition unit is used to parse the intent representation vector to obtain intent state features, decompose the intent state features using the cross-modal semantic association graph, and generate a forward-looking interactive intent prediction by combining the real-time fatigue accumulation evaluation results. The iterative unit is used to intelligently generate precise control commands based on the predicted forward interaction intent and obtain the actual response data of the execution feedback. By quantifying the multidimensional consistency deviation between the predicted forward interaction intent and the actual response data, iteratively optimizes the edge weight distribution and state transition probability value of the cross-modal semantic association graph.

8. An electronic device, characterized in that, include: processor; Memory used to store processor-executable instructions; The processor is configured to invoke instructions stored in the memory to execute the method according to any one of claims 1 to 6.

9. A computer-readable storage medium having computer program instructions stored thereon, characterized in that, When the computer program instructions are executed by the processor, they implement the method described in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Vehicle service processing method and device, computer equipment and readable storage medium

    CN121217789A

  • Intelligent AI semantic annotation method and system based on multi-modal analysis

    CN121388601A